← Founder Notes
Archive

The first research corpus written by a model is also the first one a human can't realistically…

Yethikrishna ROriginal on Threads

the first research corpus written by a model is also the first one a human can't realistically audit. openai/math dropped 722 manuscripts in 372 families, each with lean proofs you can run locally, from the irrationality exponent of pi to the vlasov-maxwell system.

verification got outsourced to a type checker before any human read a word.

Context

The openai/math README says the repository holds mathematical manuscripts and supporting proof artifacts produced by an internal OpenAI model. The catalogue has 722 manuscripts in 372 families, and the model was posed about 4,000 problems, with on average three hours of ChatGPT Pro thinking compute per result. It names exceptions: the zero-free region for the Riemann zeta function, whose writeup was human edited for readability, and the Hodge Conjecture for CM abelian varieties.

It releases abridged reasoning summaries for ten families, including family 017 on the irrationality exponent of pi and family 362 on the three-dimensional relativistic Vlasov-Maxwell system.

On verification the README says the collection holds 'results at different stages of verification', that not all have Lean formalizations, that many but not all have been formalized, and that some unformalized results 'could have issues'. It points to Comparator instructions for checking.

Stanford Tech Review matched the manuscript map against the formalization catalogue and counted 162 of 722 manuscripts (22.4%) with a Lean formalization of the main result, and none among those dated in October. That is an outside analysis of the repository files.

How it compares

The note says each manuscript has Lean proofs you can run locally. The README says the opposite: not all have Lean formalizations, and some unformalized results could have issues. The outside count of 162 of 722 points the same way. 'Each' is not supported by the primary source.

The note calls it the first research corpus written by a model. The README does not claim a first, and no source read says so, so that claim is unsupported, not refuted.

The 722 and 372 figures, the two named subjects and the open-source release match the README. That verification was 'outsourced to a type checker before any human read a word' holds at most for the formalized subset, and the README itself says some results may contain issues. That a human can't realistically audit it is the author's opinion.

Related work

Watch next

  • Read the formalization catalogue in the repository for the current formalized count. Look for independent referee comments on the unformalized manuscripts.

Sources

  1. openai/math (GitHub README)github.com
  2. OpenAI's 722 AI Math Proofs: Only 162 Checked in Lean (Stanford Tech Review)stanfordtechreview.com
  3. First look at mathematics manuscripts from an internal frontier model at OpenAI (OpenAI Developer Community)community.openai.com

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 8 October 2026 at 21:55 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/the-first-research-corpus-written-by-a-model-DePW74IjYpX" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="The first research corpus written by a model is also the first one a human can't realistically…"></iframe>

More notes