← Founder Notes
Archive

The fastest ai coding speedup this month came from the serving stack, not a new model.…

Yethikrishna ROriginal on Threads

the fastest ai coding speedup this month came from the serving stack, not a new model. deepseek-v4.1-flash gained 1.9x at low concurrency and 5.3x under load within three weeks, after vllm rewrote attention replay with cuda graphs.

the same weights just got five times cheaper to run.

Context

AlphaSignal (October 7, 2026) reports that vLLM made DeepSeek-V4.1-Flash 1.9x faster at low concurrency and 5.3x higher throughput on AgentX. It says sliding window attention bounded replay reruns only the last 128 tokens, and that CUDA graphs over trimmed layers 21 to 39 cut prefill compute 30 to 40% and dropped time to first token nearly 70% at 100K throughput.

It also lists integrated kernels (MegaAttention with NVFP4 KV, Mega-mHC, Mega-Gate and DeepSelect) and says GSM8K and GPQA show no meaningful accuracy regression from bounded replay, with the feature enabled by default. The vLLM recipe page for the model is the project's own deployment guide.

How it compares

The 1.9x and 5.3x figures match the report. The 5.3x is a throughput figure on one benchmark, AgentX, and the 1.9x is at low concurrency, so the note's wording is close to the report's.

The note says 'the same weights just got five times cheaper to run'. Cost was not stated in the page read. Throughput of 5.3x is not the same as cost, which also depends on hardware and utilization. That step is derived and unsupported, not refuted.

'Within three weeks' and the 'serving stack, not a new model' claim were not tied to dates in the page read beyond the October 7 report. That the speedup beat any model gain this month is the author's opinion.

Related work

Watch next

  • Open the vLLM pull requests behind the change for the benchmark setup. Look for an independent throughput test.

Sources

  1. vLLM speeds up DeepSeek V4.1-Flash serving (AlphaSignal, October 7, 2026)alphasignal.ai

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 9 October 2026 at 00:17 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/the-fastest-ai-coding-speedup-this-month-came-DePnMloDfR7" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="The fastest ai coding speedup this month came from the serving stack, not a new model.…"></iframe>

More notes