
OpenAI just dropped 722 mathematical proofs from a model nobody can use — and the math establishment is not happy about it.
On October 6, 2026 (UTC), OpenAI published a GitHub repository containing 722 manuscripts organized into 372 result families, all produced by an unreleased internal frontier model that the company confirms is significantly more capable than GPT-6 Astra. The collection spans number theory, algebraic geometry, mathematical physics, combinatorics, and analysis — including results on the Riemann zeta function, the Hodge Conjecture, and the irrationality exponent of π. It is the largest single release of AI-generated mathematics in history, and it has reignited a fierce debate about what happens when a proprietary model starts solving problems that human mathematicians have spent careers on.
What was released
The repository, hosted at openai/math under an Apache-2.0 license, contains PDFs, LaTeX source files, Lean formal proof artifacts, and 10 abridged reasoning summaries. According to the README, the model was posed approximately 4,000 problems during the evaluation, and the average result consumed the equivalent of roughly three hours of ChatGPT Pro thinking time.
Not every result is fully verified. Many manuscripts have accompanying formalizations in Lean — a proof assistant that checks arguments mechanically — but some do not. OpenAI acknowledges that unformalized results "could have issues" and says it will fix errors while adding more formalizations over time.
| Result area | What was claimed | Formalized? |
|---|---|---|
| Riemann zeta zero-free region | Re(s) > 11/12 (human-edited for readability) | Partial |
| Hodge Conjecture (CM abelian varieties) | Full proof | Yes |
| Irrationality exponent of π | New bound | Yes |
| Mahler conjectures (symmetric & general) | Resolved | Yes |
| Mézard-Parisi formula (diluted spin glasses) | Proven | Yes |
| Quantum Heisenberg ferromagnet | Spontaneous magnetization | Yes |
| Free group factors | Isomorphism result | Yes |
| 3D relativistic Vlasov-Maxwell | Well-posedness / behavior | Yes |
Sources: OpenAI openai/math repository README, Unite.AI, The Verge
Two results received special handling. The zero-free region for the Riemann zeta function at Re(s) > 11/12 was human-edited for readability — a notable concession given that the previous best bound was Re(s) > 1, established over a century ago. The Hodge Conjecture proof for CM abelian varieties addresses a special case of one of the seven Clay Millennium Prize problems, the same class as the Navier-Stokes result OpenAI announced on September 8, 2026 (UTC).
The advisory group that said "stop"
OpenAI did not release these results unilaterally. It consulted the Advisory Group on Mathematics and Artificial Intelligence (AGMAI) at the Institute for Advanced Study in Princeton, which published its responsible-release recommendations on September 29, 2026 (UTC) after surveying more than 600 mathematicians.
The group's verdict was blunt: "We do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models."
AGMAI's recommendations go further. It calls for results to be deposited in scholarly repositories that no AI lab controls, with persistent citable identifiers. It urges labs to refrain from treating result releases as marketing vehicles. For each result, it demands public disclosure of the model name, prompts used, summarized chain of thought, time taken, and estimated compute cost. It also calls for significant funding — controlled by nonprofit institutions, not the labs — to help the mathematical community understand and build on AI-generated output.
OpenAI says it is "continuing to explore other community-hosted alternatives" that meet the committee's guidelines, and that it will fund workshops, conferences, and special programs around AI-generated mathematics.
Why it matters
This release is a watershed moment, and not because of any single proof. The scale is what changes the calculus: 722 manuscripts from 4,000 attempted problems, at roughly three hours of compute each. If those numbers hold up under independent verification, we are looking at a system that can produce publishable-level mathematics at a rate no human mathematician — or even no department — can match.
The compute framing is revealing. Three hours of ChatGPT Pro equivalent per result is not a supercomputer-scale experiment. It suggests the model's mathematical reasoning capability is general and relatively cheap to invoke, rather than dependent on massive bespoke search. That has implications far beyond mathematics: if the same underlying capability transfers to scientific discovery, engineering design, or formal verification, the productivity multiplier could be transformative.
The tension with AGMAI is the real story. The mathematical community is not opposed to AI assistance — many researchers already use Lean and machine-learning tools. What it objects to is a closed system producing results that nobody outside OpenAI can reproduce, inspect, or build on, packaged and released on the lab's own GitHub page under its own brand. The advisory group's demand for "stop testing advanced mathematical problems on proprietary models" is essentially a demand that frontier labs not turn pure mathematics into a competitive benchmark for models the public cannot access.
There is a legitimate counterargument. OpenAI points out that it is releasing the proofs, the code, the formalizations, and the reasoning summaries — more transparency than most corporate research labs provide. The Apache-2.0 license means anyone can use, modify, and redistribute the work. And the Navier-Stokes result from September has already stimulated independent verification efforts. A world in which a private company solves Millennium Prize problems and gives away the proofs is, in some sense, better than one in which it keeps them secret.
Critical lens
The verification gap is the elephant in the room. Lean formalization is the gold standard, but not all 722 manuscripts have it. The unformalized ones are, at this stage, claims — impressive claims, but claims nonetheless. Mathematics has a long history of announced proofs that collapsed under scrutiny, and AI-generated proofs introduce new failure modes: subtle logical gaps that look plausible to both humans and weaker models, or results that depend on unstated assumptions buried in the model's training data.
The "unreleased internal frontier model" framing also raises questions. OpenAI says this model is more capable than GPT-6 Astra, which itself was classified at the Critical threshold for cybersecurity. If a model this powerful is being used to solve Millennium Prize problems, what else is it being used for? The company says it is "working to responsibly release the model," but has given no timeline. The mathematical community is being asked to evaluate outputs from a system it cannot inspect.
The AGMAI survey itself deserves scrutiny. Six hundred respondents is a meaningful sample, but the survey question was framed around a specific scenario — "OpenAI announced the existence of many results without giving details" — which may have primed respondents toward skepticism. The "clear plurality" supporting the recommendations is not the same as a consensus, and there are prominent mathematicians who have welcomed AI assistance.
What to watch
- Independent verification: Which of the 722 results survive third-party Lean formalization? Track the rate at which the community confirms or rejects unformalized manuscripts.
- The Riemann zero-free region: The Re(s) > 11/12 claim is the most concrete, testable result. If verified, it represents the first major improvement on the classical Re(s) > 1 bound in over a century and would be a genuine breakthrough.
- Model release timeline: When — and in what form — does OpenAI release the internal frontier model? AGMAI has demanded broad, equitable access.
- AGMAI compliance: Does OpenAI move the repository to a community-controlled platform? Does it disclose prompts, chain-of-thought, and compute costs per result as recommended?
- Other labs: Will Anthropic, Google DeepMind, or xAI follow with their own math result releases? The risk of an AI-mathematics arms race is real.
- Hodge and Navier-Stokes: Both are Millennium Prize problems. The Clay Institute has not yet indicated how it will treat AI-generated solutions, and OpenAI has said it does not intend to claim the prize money. The institutional response will set precedent.
OpenAI's 722-manuscript drop is not the end of the story — it is the opening salvo in a negotiation between the most capable AI system ever built and a mathematical community that suddenly finds itself competing with something it cannot see. The proofs may or may not all hold up. What is already clear is that the economics of mathematical discovery have changed, and the question of who gets to set the rules for that change is now wide open.
No comments yet