The first research corpus written by a model is also the first one a human can't realistically…
the first research corpus written by a model is also the first one a human can't realistically audit. openai/math dropped 722 manuscripts in 372 families, each with lean proofs you can run locally, from the irrationality exponent of pi to the vlasov-maxwell system.
verification got outsourced to a type checker before any human read a word.
Context
The openai/math README says the repository holds mathematical manuscripts and supporting proof artifacts produced by an internal OpenAI model. The catalogue has 722 manuscripts in 372 families, and the model was posed about 4,000 problems, with on average three hours of ChatGPT Pro thinking compute per result. It names exceptions: the zero-free region for the Riemann zeta function, whose writeup was human edited for readability, and the Hodge Conjecture for CM abelian varieties.
It releases abridged reasoning summaries for ten families, including family 017 on the irrationality exponent of pi and family 362 on the three-dimensional relativistic Vlasov-Maxwell system.
On verification the README says the collection holds 'results at different stages of verification', that not all have Lean formalizations, that many but not all have been formalized, and that some unformalized results 'could have issues'. It points to Comparator instructions for checking.
Stanford Tech Review matched the manuscript map against the formalization catalogue and counted 162 of 722 manuscripts (22.4%) with a Lean formalization of the main result, and none among those dated in October. That is an outside analysis of the repository files.
The note says each manuscript has Lean proofs you can run locally. The README says the opposite: not all have Lean formalizations, and some unformalized results could have issues. The outside count of 162 of 722 points the same way. 'Each' is not supported by the primary source.
The note calls it the first research corpus written by a model. The README does not claim a first, and no source read says so, so that claim is unsupported, not refuted.
The 722 and 372 figures, the two named subjects and the open-source release match the README. That verification was 'outsourced to a type checker before any human read a word' holds at most for the formalized subset, and the README itself says some results may contain issues. That a human can't realistically audit it is the author's opinion.
Related work
- openai/math (GitHub README) ↗Primary source for the counts, method, exceptions and verification statements.
- OpenAI's 722 AI Math Proofs: Only 162 Checked in Lean (Stanford Tech Review) ↗Outside count of how many manuscripts carry a Lean formalization.
- First look at mathematics manuscripts from an internal frontier model at OpenAI (OpenAI Developer Community) ↗OpenAI's community post introducing the release.
Watch next
- Read the formalization catalogue in the repository for the current formalized count. Look for independent referee comments on the unformalized manuscripts.
Sources
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 8 October 2026 at 21:55 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →