Formal proof search just solved nine of 353 open erdős problems, two stuck for 56 years. alphaproof…
formal proof search just solved nine of 353 open erdős problems, two stuck for 56 years. alphaproof nexus, published in science on oct 8, runs llm-generated proofs through lean verification and also proved 44 of 492 oeis conjectures.
generation is easy now, verification is the moat.
Context
The AlphaProof Nexus paper (arXiv, May 2026) says its most capable agent autonomously resolved 9 of 353 open Erdős problems from the Formal Conjectures repository, including two questions open for 56 years, at an inference cost of a few hundred dollars per problem. It says the agent proved 44 of 492 open OEIS conjectures, which Gemini autoformalized into Lean, and that experts validated that the Lean statements matched the problems.
The paper says the agent was required to prove test lemmas against the first terms of each sequence to guard against misformalization, and that the reported costs do not capture the full cost, since identifying tractable problems was itself a significant computational investment. DeepMind's results repository holds the Lean proofs. A Science item dated October 8, 2026, 'AI for research mathematics has arrived', appears in Volume 394, Issue 6820.
The nine of 353, the two 56-year-old questions and the 44 of 492 match the paper. The note says the results were published in Science on October 8. The paper read is an arXiv preprint from May 2026, and the Science piece found is a perspective item in that issue. Whether the AlphaProof Nexus paper itself appears in Science on October 8 was not confirmed, so that is unsupported here, not refuted.
The note's 'llm-generated proofs through lean verification' fits the method. Lean verifies that a proof is valid for the formal statement, and the paper adds that human experts checked that each Lean statement matched the intended problem, so verification covers the proof but not the statement unless people check it.
'Solved' covers the problems as formalized in the repository. The note's 'verification is the moat' is the author's view. The paper itself says the search needed 3,000 episodes per problem and that identifying tractable problems had a large compute cost.
Related work
- Advancing Mathematics Research with AI-Driven Formal Proof Search (arXiv 2605.22763, May 2026) ↗Primary source for the 9 of 353, the 44 of 492, the costs and the validation steps.
- google-deepmind/alphaproof-nexus-results (GitHub) ↗Lean proofs and problem lists released with the paper.
- AI for research mathematics has arrived (Science, October 8, 2026) ↗Science's October 8, 2026 item on AI for research mathematics.
Watch next
- Check whether the Science issue includes the research paper or only commentary. Read the expert notes on the nine Erdős solutions.
Sources
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 9 October 2026 at 02:41 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →