On 6 October 2026 OpenAI published 722 mathematical manuscripts in a single GitHub repository. The company says they were produced by an unreleased internal model and grouped into 372 "families" of related papers. A family can contain a main result, companion arguments, consequences or alternative proofs.
The papers address open research problems, meaning questions that professional mathematicians had not yet answered. OpenAI says it began testing its models on such problems after they saturated its existing mathematics tests, in other words, after they started scoring close to the maximum.
According to the repository, the model was given roughly 4,000 problems over the course of the evaluation. The average result used about three hours of ChatGPT Pro "thinking" compute, the extended reasoning time OpenAI sells to its most expensive subscribers. OpenAI has not published how many of those 4,000 attempts failed.
What the release contains
Each manuscript comes as a PDF with its source files. For many of them, OpenAI also supplies a formalisation in Lean, a programming language in which a proof can be written so precisely that a computer can check every step.
That matters because traditional proofs are checked by human referees, a slow process. A Lean proof removes the need to trust the author on the logic. It does not, however, guarantee that the formal statement says the same thing as the written paper. An independent index of the collection counts 162 papers for which OpenAI claims a Lean formalisation of the main result, and notes that it reports this claim without checking the code itself.
OpenAI's own repository is candid on this point. It says the collection "includes results at different stages of verification" and that "some of the unformalized results could have issues."
The company says it consulted the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study on how to release the work. An OpenAI spokesperson told Retraction Watch that the group recommended releasing results without waiting for full formalisation, and that roughly 50 percent of the results were released unconfirmed.
The first corrections
Less than a day later, on 7 October, OpenAI withdrew three manuscripts. A sign error invalidated an argument in one paper and the construction used by two papers that depended on it. The company also said it had revised 14 other manuscripts with proof repairs, corrected statements and clearer assumptions. The repository's contents page now lists 719 manuscripts.
Alex Townsend, an associate professor of mathematics at Cornell University, told Retraction Watch he expects more errors to surface. "Given the skepticism of AI in the mathematics community, I think that OpenAI should have announced the manuscripts that were lean verified first," he said.
Andrew Sutherland, a senior research scientist in mathematics at MIT, called the quick withdrawals "the responsible thing to do". He added: "I think the withdrawals and corrections will be viewed positively by the mathematics community, but it will take a lot more than that to earn back the trust they have lost."
Part of that lost trust dates from September, when OpenAI announced that an AI system had solved the Navier-Stokes problem, a famous open question about the equations that describe flowing liquids and gases. A group of leading researchers said that announcement left no time for a proper write-up or for crediting earlier work by others. Retraction Watch reports that more than 8,000 researchers have endorsed those concerns.

Mathematicians react
The maths blog Proofs and Prompts collected more than 100 responses from researchers, summarised by The Decoder. Several were impressed by the content. Terence Tao wrote that the proofs appeared to "introduce clever new ideas that will be fruitful once digested", while saying he was "deeply frustrated" that there was nobody to discuss the work, present it or teach it.
Fields Medallist Peter Scholze urged patience: "We should remember that we are all in this together; mathematics is a marathon, not a sprint; and the goal is and always will be the human understanding of mathematics, which will invariably take time."
Others described the effect on their careers. A PhD student learned that the central problem of his dissertation had been claimed the same morning. Henry Wilton criticised OpenAI for withholding author names and failure rates. The Association for Human Mathematics went further: "Releasing over 700 files at once is not a demonstration of scholarship, but a demonstration of power."
What happens next
OpenAI says it will keep adding Lean formalisations, correct errors promptly and fund workshops and conferences to help the community understand major AI-produced results. It also says it is working to release the model that produced them.
In our view, the release shows where the bottleneck in AI mathematics has moved. Producing candidate proofs is now fast and cheap for a frontier lab. Checking them, understanding them and connecting them to existing work still falls on a limited number of human experts, most of whom did not ask for this workload.
Two steps would help. Publishing the full record of the roughly 4,000 attempts, including failures, would let outsiders judge how selective the catalogue is. Separating fully formalised results from unconfirmed ones, as Townsend suggests, would tell readers which claims they can rely on today. Our conclusion is that a proof becomes part of mathematics only when people can read it, test it and build on it, and for most of these 722 manuscripts that work has only just begun.




