← Founder Notes
Archive

The inference engine just jumped 7.8 times on new silicon. vllm's vera rubin support, detailed oct…

Yethikrishna ROriginal on Threads

the inference engine just jumped 7.8 times on new silicon. vllm's vera rubin support, detailed oct 9, claims 7.8x throughput over the gb200 nvl72 rack with 2.4x bandwidth.

the software now compounds the hardware jump.

Context

Verified date: the vLLM blog post is dated Oct 9, 2026 and titled vLLM Support for NVIDIA Vera Rubin NVL72: 7.8x Throughput over GB200 NVL72. It says vLLM now supports Vera Rubin NVL72, written with Inferact, NVIDIA and Red Hat.

FourWeekMBA reports that on SemiAnalysis AgentX vLLM measured MiniMax M3 at up to 7.84x the throughput per GPU of GB200 NVL72 at matched interactivity, and 5.18x under a 150 tokens-per-second constraint. Magica adds a separate MLPerf Inference v6.1 preview with up to 3.7x the throughput of GB300 on Qwen3-VL.

How it compares

The Oct 9 date and the 7.8x figure over GB200 NVL72 match the vLLM post. The sources scope the 7.84x to MiniMax M3 on one benchmark, per GPU at matched interactivity, with 5.18x under a 150 tokens-per-second limit, so it is not a general speedup. The 2.4x bandwidth claim was not seen in the sources read, so unsupported here, not refuted. 'The software now compounds the hardware jump' is the author's opinion.

Related work

Watch next

  • Read the vLLM post for where the bandwidth figure comes from, if it is there.

Sources

  1. vLLM blog: Support for NVIDIA Vera Rubin NVL72, 7.8x throughput over GB200 NVL72vllm.ai
  2. FourWeekMBA: vLLM on Vera Rubin posts 7.8x GB200 throughput per chipfourweekmba.com
  3. Magica: vLLM reports up to 7.84x GB200 throughputmagica.com

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 11 October 2026 at 19:18 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/the-inference-engine-just-jumped-7-8-times-DeWzYBejM7v" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="The inference engine just jumped 7.8 times on new silicon. vllm's vera rubin support, detailed oct…"></iframe>

More notes