The fastest ai coding speedup this month came from the serving stack, not a new model.…
the fastest ai coding speedup this month came from the serving stack, not a new model. deepseek-v4.1-flash gained 1.9x at low concurrency and 5.3x under load within three weeks, after vllm rewrote attention replay with cuda graphs.
the same weights just got five times cheaper to run.
Context
AlphaSignal (October 7, 2026) reports that vLLM made DeepSeek-V4.1-Flash 1.9x faster at low concurrency and 5.3x higher throughput on AgentX. It says sliding window attention bounded replay reruns only the last 128 tokens, and that CUDA graphs over trimmed layers 21 to 39 cut prefill compute 30 to 40% and dropped time to first token nearly 70% at 100K throughput.
It also lists integrated kernels (MegaAttention with NVFP4 KV, Mega-mHC, Mega-Gate and DeepSelect) and says GSM8K and GPQA show no meaningful accuracy regression from bounded replay, with the feature enabled by default. The vLLM recipe page for the model is the project's own deployment guide.
The 1.9x and 5.3x figures match the report. The 5.3x is a throughput figure on one benchmark, AgentX, and the 1.9x is at low concurrency, so the note's wording is close to the report's.
The note says 'the same weights just got five times cheaper to run'. Cost was not stated in the page read. Throughput of 5.3x is not the same as cost, which also depends on hardware and utilization. That step is derived and unsupported, not refuted.
'Within three weeks' and the 'serving stack, not a new model' claim were not tied to dates in the page read beyond the October 7 report. That the speedup beat any model gain this month is the author's opinion.
Related work
- vLLM speeds up DeepSeek V4.1-Flash serving (AlphaSignal, October 7, 2026) ↗Source for the speedup figures, the bounded replay and the CUDA graph changes.
- deepseek-ai/DeepSeek-V4.1-Flash (vLLM Recipes) ↗The vLLM project's own recipe page for this model; surfaced in search and not read in full here.
Watch next
- Open the vLLM pull requests behind the change for the benchmark setup. Look for an independent throughput test.
Sources
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 9 October 2026 at 00:17 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →