The inference engine just jumped 7.8 times on new silicon. vllm's vera rubin support, detailed oct…
the inference engine just jumped 7.8 times on new silicon. vllm's vera rubin support, detailed oct 9, claims 7.8x throughput over the gb200 nvl72 rack with 2.4x bandwidth.
the software now compounds the hardware jump.
Context
Verified date: the vLLM blog post is dated Oct 9, 2026 and titled vLLM Support for NVIDIA Vera Rubin NVL72: 7.8x Throughput over GB200 NVL72. It says vLLM now supports Vera Rubin NVL72, written with Inferact, NVIDIA and Red Hat.
FourWeekMBA reports that on SemiAnalysis AgentX vLLM measured MiniMax M3 at up to 7.84x the throughput per GPU of GB200 NVL72 at matched interactivity, and 5.18x under a 150 tokens-per-second constraint. Magica adds a separate MLPerf Inference v6.1 preview with up to 3.7x the throughput of GB300 on Qwen3-VL.
The Oct 9 date and the 7.8x figure over GB200 NVL72 match the vLLM post. The sources scope the 7.84x to MiniMax M3 on one benchmark, per GPU at matched interactivity, with 5.18x under a 150 tokens-per-second limit, so it is not a general speedup. The 2.4x bandwidth claim was not seen in the sources read, so unsupported here, not refuted. 'The software now compounds the hardware jump' is the author's opinion.
Related work
- InferenceX: GB200 NVL72 vs Vera Rubin NVL72, MiniMax M3 ↗SemiAnalysis comparison page.
- vLLM project blog mirror ↗Oct 9, 2026.
- daily.dev: vLLM support for Vera Rubin NVL72 ↗Day-0 support for DeepSeek, Kimi, GLM and MiniMax.
Watch next
- Read the vLLM post for where the bandwidth figure comes from, if it is there.
Sources
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 11 October 2026 at 19:18 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →