← Founder Notes
Archive ·

The inference server that everyone runs just made model restarts nearly free. vllm v0.30.0, out…

01:47 ISTby Yethikrishna R

the inference server that everyone runs just made model restarts nearly free. vllm v0.30.0, out september 22, keeps quantized weights resident in gpu memory via a per-gpu daemon, so a restarting engine maps them over cuda ipc instead of reloading from disk, and ships hybrid-attention paths for kimi k3, deepseek-v4.1-flash and qwen3.8-flash-next. serving infra now treats weights like a cache instead of a cold start.

Share

Embed this note

<iframe src="https://founder.myndlabs.tech/notes/embed/the-inference-server-that-everyone-runs-just-made-DdpJh2ZiPpL" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="The inference server that everyone runs just made model restarts nearly free. vllm v0.30.0, out…"></iframe>

Original

More notes