The best vulnerability finder in the world right now is an open weight model. mistral's ml4 scored…
the best vulnerability finder in the world right now is an open weight model. mistral's ml4 scored 82% on the cyber index test that makes a model reproduce a real bug and then patch it, the highest of any model, beating closed labs on their home turf.
that inverts the whole safety argument for keeping weights secret.
Context
Mistral's announcement launches a public preview of Mistral Large 4, called ML4, and says the weights will be released by the end of the month. On the Artificial Analysis Cyber Index it says ML4 ranks among the top five models globally and leads open-weight models developed outside China. On one test in that index, which asks a model to reproduce a real vulnerability in open-source software and then patch it, it reports 82% and says that is 'the highest of any model'.
Mistral says several leading closed models, including Claude Opus 5.5 and GPT-6 Astra, score near zero on that test because they refuse the task. It also says it is red-teaming ML4 with cybersecurity leaders, vetted partners and state authorities who get 'reduced moderation and expanded cyber capabilities'.
Threat Frontier (October 6) checked Artificial Analysis's leaderboard and found Mistral Large 4 Preview at 81.7% on the CyberGym-E2E-AA test, first of 18 models, with a composite Cyber Index of 49.5, fifth of 18. It notes that safety blocks are scored as zero, and that ML4 scores 16.0% on the DeepsecBench-AA discovery test against 26.9% for Grok 4.7.
The note says the best vulnerability finder in the world is an open weight model. The 82% figure is one sub-test of the index, the reproduce-and-patch task, and the composite ranking from the same data puts ML4 fifth of 18. On the separate test of finding vulnerabilities it scores lower than the leader. 'Best finder' is the note's wording and is not what the sources show, so it stays unsupported, not refuted.
'Open weight' is forward-looking: the preview runs on Mistral's API and the weights were promised by the end of the month. The sources read do not say they had been published by the time of the note.
'Beating closed labs on their home turf' depends partly on refusals. Threat Frontier reports that two of the closed models were blocked on nearly all tasks of this test and scored zero, and among models that did not refuse the lead is small (81.7% against 78.6% for MiMo-V2.6-Pro and 77.9% for GPT-6 Luna). The Cyber Index figures come from Artificial Analysis, but the claims about leading open models and refusal rates are Mistral's own.
'That inverts the whole safety argument for keeping weights secret' is the author's opinion. Threat Frontier makes the opposite point that a self-hosted model has no provider-side policy once the weights are out, so it applies to attackers as well as defenders.
Related work
- Introducing Mistral Large 4 (Mistral AI) ↗Primary source for the preview, the 82% claim and the partner tier.
- Mistral Large 4 preview ships a reduced-moderation cyber tier, and Artificial Analysis lists its 82% as the top score (Threat Frontier) ↗Source for the Artificial Analysis leaderboard figures and the caveats.
- Mistral Large 4 (Mistral Docs) ↗Model documentation page for ML4.
Watch next
- Check Artificial Analysis's own leaderboard directly once the weights ship. Compare ML4's discovery score with the reproduce-and-patch score.
Sources
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 8 October 2026 at 22:04 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →