The benchmark just crowned grok. grok 4.5 hit 94.9 percent on gpqa diamond zero-shot on oct 10,…
the benchmark just crowned grok. grok 4.5 hit 94.9 percent on gpqa diamond zero-shot on oct 10, ahead of gemini 3 pro at 93.4 and gpt-6.1 sol at 92.9.
the frontier order shifted without an announcement.
Context
Verified sources: the GPQA Diamond leaderboards read were updated or dated Oct 9 and 10, 2026 and show different numbers from the post. Kaggle's few-shot leaderboard (updated Oct 9, 2026) lists Gemini 3 Pro Preview at 93.4 percent, and BenchLeader (as of Oct 10) has GPT-6 Astra at 95.8 percent, Claude Sonnet 5.5 at 95.6 percent and GPT-6.1 Sol at 95.4 percent.
Sophon (Jul 8, 2026) lists Grok 4.5's best GPQA Diamond score at 93.1 percent.
The Gemini 3 Pro 93.4 figure matches Kaggle's few-shot board. No source read shows Grok 4.5 at 94.9 percent zero-shot on Oct 10, and the 92.9 for GPT-6.1 Sol was not seen; leaderboards read give GPT-6.1 Sol 95.1 to 95.4 percent depending on the board. Different boards use different settings, so the ranking is not settled by one number. The Grok 94.9 claim is not seen in the sources read, so unsupported here, not refuted.
'the frontier order shifted without an announcement' is the author's opinion.
Related work
- Kaggle: GPQA Diamond few-shot leaderboard ↗Updated Oct 9, 2026.
- BenchLeader: GPQA Diamond leaderboard ↗As of Oct 10, 2026.
- Sophon: Grok 4.5 by xAI ↗Jul 8, 2026.
Watch next
- Look for xAI's own Grok 4.5 GPQA figure and its test settings.
Sources
- Kaggle, GPQA Diamond few-shotkaggle.com
- BenchLeader, GPQA Diamondbenchleader.com
- Artificial Analysis, GPQA Diamondartificialanalysis.ai
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 11 October 2026 at 10:17 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →