Mathematical equations on a chalkboard

OpenAI just dropped 722 mathematical proofs from a model nobody can use — and the math establishment is not happy about it.

On October 6, 2026 (UTC), OpenAI published a GitHub repository containing 722 manuscripts organized into 372 result families, all produced by an unreleased internal frontier model that the company confirms is significantly more capable than GPT-6 Astra. The collection spans number theory, algebraic geometry, mathematical physics, combinatorics, and analysis — including results on the Riemann zeta function, the Hodge Conjecture, and the irrationality exponent of π. It is the largest single release of AI-generated mathematics in history, and it has reignited a fierce debate about what happens when a proprietary model starts solving problems that human mathematicians have spent careers on.

What was released

The repository, hosted at openai/math under an Apache-2.0 license, contains PDFs, LaTeX source files, Lean formal proof artifacts, and 10 abridged reasoning summaries. According to the README, the model was posed approximately 4,000 problems during the evaluation, and the average result consumed the equivalent of roughly three hours of ChatGPT Pro thinking time.

Not every result is fully verified. Many manuscripts have accompanying formalizations in Lean — a proof assistant that checks arguments mechanically — but some do not. OpenAI acknowledges that unformalized results "could have issues" and says it will fix errors while adding more formalizations over time.

Result area What was claimed Formalized?
Riemann zeta zero-free region Re(s) > 11/12 (human-edited for readability) Partial
Hodge Conjecture (CM abelian varieties) Full proof Yes
Irrationality exponent of π New bound Yes
Mahler conjectures (symmetric & general) Resolved Yes
Mézard-Parisi formula (diluted spin glasses) Proven Yes
Quantum Heisenberg ferromagnet Spontaneous magnetization Yes
Free group factors Isomorphism result Yes
3D relativistic Vlasov-Maxwell Well-posedness / behavior Yes

Sources: OpenAI openai/math repository README, Unite.AI, The Verge

Two results received special handling. The zero-free region for the Riemann zeta function at Re(s) > 11/12 was human-edited for readability — a notable concession given that the previous best bound was Re(s) > 1, established over a century ago. The Hodge Conjecture proof for CM abelian varieties addresses a special case of one of the seven Clay Millennium Prize problems, the same class as the Navier-Stokes result OpenAI announced on September 8, 2026 (UTC).

The advisory group that said "stop"

OpenAI did not release these results unilaterally. It consulted the Advisory Group on Mathematics and Artificial Intelligence (AGMAI) at the Institute for Advanced Study in Princeton, which published its responsible-release recommendations on September 29, 2026 (UTC) after surveying more than 600 mathematicians.

The group's verdict was blunt: "We do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models."

AGMAI's recommendations go further. It calls for results to be deposited in scholarly repositories that no AI lab controls, with persistent citable identifiers. It urges labs to refrain from treating result releases as marketing vehicles. For each result, it demands public disclosure of the model name, prompts used, summarized chain of thought, time taken, and estimated compute cost. It also calls for significant funding — controlled by nonprofit institutions, not the labs — to help the mathematical community understand and build on AI-generated output.

OpenAI says it is "continuing to explore other community-hosted alternatives" that meet the committee's guidelines, and that it will fund workshops, conferences, and special programs around AI-generated mathematics.

Why it matters

This release is a watershed moment, and not because of any single proof. The scale is what changes the calculus: 722 manuscripts from 4,000 attempted problems, at roughly three hours of compute each. If those numbers hold up under independent verification, we are looking at a system that can produce publishable-level mathematics at a rate no human mathematician — or even no department — can match.

The compute framing is revealing. Three hours of ChatGPT Pro equivalent per result is not a supercomputer-scale experiment. It suggests the model's mathematical reasoning capability is general and relatively cheap to invoke, rather than dependent on massive bespoke search. That has implications far beyond mathematics: if the same underlying capability transfers to scientific discovery, engineering design, or formal verification, the productivity multiplier could be transformative.

The tension with AGMAI is the real story. The mathematical community is not opposed to AI assistance — many researchers already use Lean and machine-learning tools. What it objects to is a closed system producing results that nobody outside OpenAI can reproduce, inspect, or build on, packaged and released on the lab's own GitHub page under its own brand. The advisory group's demand for "stop testing advanced mathematical problems on proprietary models" is essentially a demand that frontier labs not turn pure mathematics into a competitive benchmark for models the public cannot access.

There is a legitimate counterargument. OpenAI points out that it is releasing the proofs, the code, the formalizations, and the reasoning summaries — more transparency than most corporate research labs provide. The Apache-2.0 license means anyone can use, modify, and redistribute the work. And the Navier-Stokes result from September has already stimulated independent verification efforts. A world in which a private company solves Millennium Prize problems and gives away the proofs is, in some sense, better than one in which it keeps them secret.

Critical lens

The verification gap is the elephant in the room. Lean formalization is the gold standard, but not all 722 manuscripts have it. The unformalized ones are, at this stage, claims — impressive claims, but claims nonetheless. Mathematics has a long history of announced proofs that collapsed under scrutiny, and AI-generated proofs introduce new failure modes: subtle logical gaps that look plausible to both humans and weaker models, or results that depend on unstated assumptions buried in the model's training data.

The "unreleased internal frontier model" framing also raises questions. OpenAI says this model is more capable than GPT-6 Astra, which itself was classified at the Critical threshold for cybersecurity. If a model this powerful is being used to solve Millennium Prize problems, what else is it being used for? The company says it is "working to responsibly release the model," but has given no timeline. The mathematical community is being asked to evaluate outputs from a system it cannot inspect.

The AGMAI survey itself deserves scrutiny. Six hundred respondents is a meaningful sample, but the survey question was framed around a specific scenario — "OpenAI announced the existence of many results without giving details" — which may have primed respondents toward skepticism. The "clear plurality" supporting the recommendations is not the same as a consensus, and there are prominent mathematicians who have welcomed AI assistance.

What to watch

OpenAI's 722-manuscript drop is not the end of the story — it is the opening salvo in a negotiation between the most capable AI system ever built and a mathematical community that suddenly finds itself competing with something it cannot see. The proofs may or may not all hold up. What is already clear is that the economics of mathematical discovery have changed, and the question of who gets to set the rules for that change is now wide open.