← Founder Notes
Archive

The flash model just 5x'd agentic throughput. deepseek-v4.1-flash on vllm, out oct 7, uses decoder…

Yethikrishna ROriginal on Threads

the flash model just 5x'd agentic throughput. deepseek-v4.1-flash on vllm, out oct 7, uses decoder replay with cuda graphs to cut prefill compute by 30 to 40 percent.

the small model now outruns the queue.

Context

Verified event date: the vLLM Blog post 'DeepSeek-V4.1-Flash on vLLM: 5x Agentic Throughput Since Day 0' is dated Oct 7, 2026 (Inferact and the vLLM Team). Its summary says that in the three weeks after the model's release, Inferact and the vLLM team worked on agentic throughput.

A vLLM pull request, 'Compact large decoder prefills with a dependency-preserving halo', says DeepSeek V4.1 Flash runs the full prompt through decoder layers 21 to 39 even though those layers reuse the encoder's global KV and have only a 128-token causal window. The paper on the model (arXiv 2609.19969) describes a 552B-parameter multimodal mixture-of-experts with 16B active parameters and up to one million tokens of context.

How it compares

The Oct 7 date and the 5x agentic throughput match the vLLM blog title. The 30 to 40 percent prefill compute cut, the decoder replay and the CUDA graph detail were not seen in the sources read, so unsupported here, not refuted. The sources read do describe a decoder prefill compaction change in a pull request, which is the nearest match. 'The small model now outruns the queue' is the author's opinion, and a 552B backbone is not small by parameter count, though 16B parameters are active per token.

Related work

Watch next

  • Find the 30 to 40 percent prefill figure and the decoder replay detail in the vLLM blog.

Sources

  1. vLLM Blog: DeepSeek-V4.1-Flash on vLLM, 5x agentic throughput since day 0vllm-project.github.io
  2. vLLM pull request 57281: compact large decoder prefillsgithub.com
  3. arXiv: DeepSeek-V4.1-Flash, pushing the limits of KV cache compressionarxiv.org

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 11 October 2026 at 15:55 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/the-flash-model-just-5x-d-agentic-throughput-DeWcK8Gja86" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="The flash model just 5x'd agentic throughput. deepseek-v4.1-flash on vllm, out oct 7, uses decoder…"></iframe>

More notes