← All posts

Artificial intelligence

OpenAI Has Published 722 Mathematical Manuscripts. The Body It Consulted Asked It to Stop a Week Earlier

OpenAI released 722 manuscripts in 372 result families from an unnamed internal model on October 6, seven days after the advisory group it convened asked frontier labs to stop testing advanced mathematics on proprietary models. The release meets some of the group's recommendations, defers others, and keeps the repository, the model and the problem selection in the lab.

MAI
The official social-share image OpenAI published with its post "Sharing AI progress in mathematics," carrying the post's title artwork.

OpenAI published 722 mathematical manuscripts on October 6, organised into 372 result families and produced by an internal model it has not named. They sit in a GitHub repository the company owns. Seven days earlier, the advisory group OpenAI itself convened had published its recommendations for releases exactly like this one, and opened them by asking the labs to stop.

At present, some frontier AI labs are testing advanced mathematical problems on proprietary models that remain inaccessible to the broader scientific community. Our recommendations are formulated with this practical context in mind. However, ideally, they would not do so. We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models.

That is the frame. OpenAI did not ignore the recommendations — it met several of them, some in detail. It declined the first one by shipping.

What is in the release

ItemFigure, as stated by OpenAI
Manuscripts722
Result families372
Problems posed to the model during evaluationapproximately 4,000
Average compute per resultthree hours of ChatGPT Pro thinking
Lean formalisations"Many, but not all, of the manuscripts"
Reasoning summaries published10 families, abridged
Model"an unreleased internal OpenAI model"

The repository's own caveat is worth quoting, because it is the honest part: "Some of the unformalized results could have issues. We will endeavor to fix any such issues quickly."

Scored against the checklist

The advisory group's September 29 document, informed by more than six hundred survey responses, is specific enough to mark.

What the group asked forWhat the release does
Deposit results in scholarly repositories "not controlled by any AI lab"Published in github.com/openai/math; OpenAI says it is "exploring community-hosted repositories"
Disclose the model name, the prompts, the time taken and the estimated compute costCompute disclosed as a Pro-hours average; model unnamed; prompts not published, ten abridged reasoning summaries instead
Formalise proofs as far as possibleMany formalised in Lean, not all
Rewrite each proof in the conventions of a mathematical paper, citing related literatureDeferred — OpenAI commits to improving citations and exposition in future releases
Publish how many comparable problems the models failed, and how problems were chosenFailure base rate effectively disclosed; selection criteria not
Fund community understanding without controlling itOpenAI says it will fund workshops, conferences and special programs

One met, three partial, two not — and the overarching ask declined by the act of publishing. For a company whose previous disclosure on this subject was a sentence containing a number, that is a real movement. It is also not what the group asked for, and the gap is in the same place every time: the lab keeps control of the artefact, the model and the selection.

The denominator is the new thing

The most useful number in the release is the one that looks like an admission. Roughly 4,000 problems went in; 372 families of sufficiently significant results came out. No lab has previously published its own base rate on open research problems, and the base rate is what converts a headline count into a capability estimate. A model that resolves 372 families from 4,000 attempts is a very different instrument from one that resolves 372 from 400.

It also reframes the compute figure. Three hours of Pro-grade thinking per result, across an evaluation of that size, puts the whole campaign in the region of ten thousand-plus hours of frontier inference — cheap against a mathematician's career, expensive against a grant. The economics of attacking open problems by brute search are now visible, and they are not absurd.

What is claimed, precisely

Precision matters here more than enthusiasm, because the headline numbers invite a misreading.

The two results produced outside the standard procedure are a zero-free region for the Riemann zeta function and a proof of the Hodge Conjecture for CM abelian varieties. Neither is a Millennium Prize problem solved. A zero-free region in the half-plane Re(s) > 11/12 would be a substantial advance in analytic number theory; the Riemann Hypothesis asks for every non-trivial zero on the line Re(s) = 1/2, and a region is not a line. The Hodge Conjecture for CM abelian varieties is a special case of the conjecture, not the conjecture. OpenAI also notes that the write-up of the 11/12 result was human-edited for readability — the one place in the catalogue where a person is acknowledged in the text.

The named reasoning summaries give a fair sense of the rest: the irrationality exponent of π, the symmetric and general Mahler conjectures, Kaplansky's direct-finiteness conjecture in characteristic two, the Mézard–Parisi formula for diluted spin glasses, the three-dimensional relativistic Vlasov–Maxwell system. Real problems, deep in their subfields, mostly not famous outside them. Scientific American's Joseph Howlett reported that nearly all of the results came from a single prompt handed to a single agent, with some requiring multiple attempts.

The read

The verification burden has been transferred, and that was always the point of contention rather than the arithmetic. Lean covers some of this corpus; the rest is 722 PDFs awaiting referees who have other jobs. OpenAI has promised money for workshops to help the field catch up, which is a reasonable offer and also an acknowledgement that the field cannot.

The question the advisory group was formed to answer in September — whether the results would arrive as a number and a date, or as something a mathematician could use — now has a partial answer. They arrived with proofs, a denominator and a caveat. They also arrived in a repository with the lab's name on it, from a model with no name at all, one week after nine mathematicians asked for neither.

Sources: OpenAI: Sharing AI progress in mathematics · openai/math on GitHub · Advisory Group on Mathematics and AI: General recommendations, September 29 · Advisory Group on Mathematics and Artificial Intelligence · Scientific American: OpenAI unleashes hundreds more math results upon a field already in shock · Interesting Engineering: OpenAI's largest math release tackles 4,000 problems with Lean proofs

Keep reading