OpenAI publishes hundreds of AI-written proofs of open maths problems
722 unreviewed manuscripts in 372 result families, including claimed proofs of the Unique Games Conjecture and a zero-free strip for the zeta function; 235 families come with Lean checks.
- Models & capabilities
- Ideas & essays
- Benchmarks & progress
- Era-defining
On the evening of 6 October 2026 OpenAI published 722 mathematical manuscripts written by an unreleased internal model, grouped into 372 “result families” across 17 areas of mathematics and theoretical computer science. Each family, the company said, resolves or substantially advances a recognised open problem. The claims include a proof of the Unique Games Conjecture, a zero-free half-plane for the Riemann zeta function, the rational Hodge conjecture for CM abelian varieties, and the isomorphism of the free group factors — problems that had stood for between two and eight decades. None of the papers has been peer-reviewed. About two-thirds of the families come with machine-checked proofs in the Lean language; OpenAI’s own README warns that “some of the unformalized results could have issues”.
The release came four weeks after the company’s contested Navier–Stokes announcement, which it attributed to the same model, and on the same day WIRED reported that mathematicians felt OpenAI had ignored their advice on how to publish such work. In scale it went far beyond anything a lab had released before: one result in May, ten in August, and now several hundred at once.
Timeline
- 1 May — an OpenAI reasoning model is credited with disproving Erdős’s unit-distance conjecture.
- 1 August — OpenAI publishes ten formally verified advances by blog post.
- August — OpenAI convenes around 40 mathematicians to discuss what to do if AI outpaces humans in the field. Attendees say they asked the company to publish explanatory papers rather than blog posts or tweets, and say they were told the results would not be released all at once; an OpenAI spokesperson later said the company was “not aware of” that assurance.
- 28 August — OpenAI begins training the internal model that produced the collection, according to a statement from spokesperson Lindsay McCallum to WIRED.
- 8 September — OpenAI says the model resolved a Navier–Stokes problem; a public credit dispute follows with NYU mathematician Tristan Buckmaster.
- 21 September — OpenAI announces the Advisory Group on Mathematics and Artificial Intelligence, nine unpaid mathematicians hosted at the Institute for Advanced Study in Princeton, and says the model has resolved more than 100 further open problems.
- 29 September — the advisory group publishes general recommendations for the responsible release of AI-generated mathematics, drawing on more than 600 responses from mathematicians.
- 6 October, afternoon — WIRED reports that OpenAI plans to put hundreds of solutions on GitHub.
- 6 October, about 6pm US Eastern — the repository goes live alongside a short OpenAI post; the advisory group issues a statement the same evening.
What is in the repository
The collection is organised as a catalogue rather than a ranked list. An overview document gives each of the 372 families a one-paragraph description and links to its papers, which are dated between 10 September and 6 October 2026, credited to “OpenAI” as sole author, and released under the Apache 2.0 licence with a citation block each. Number theory, algebraic geometry, analysis, combinatorics, theoretical computer science, probability, operator algebras, topology, differential geometry and mathematical physics all appear. Around 44 families are titled as counterexamples, disproving rather than proving what had been conjectured.
The best-known claims, as OpenAI states them in the catalogue:
- The Unique Games Conjecture, proposed by Subhash Khot in 2002 and central to the theory of which optimisation problems can be approximated efficiently.
- The “quasi-Riemann hypothesis”: no Dirichlet L-function, including the zeta function, has a zero with real part above 7/8. This falls well short of the Riemann hypothesis itself, which concerns the line at 1/2, but no fixed zero-free strip of this kind had previously been proved.
- The rational Hodge conjecture for CM abelian varieties, a special case of one of the seven Millennium Prize Problems, from which the catalogue derives the Tate conjecture for abelian varieties over finite fields.
- The free group factor problem, open since the 1940s: the catalogue says all non-abelian free group factors are isomorphic.
- Claims in complexity theory and algorithms, including that randomness gives no extra power to log-space computation (L = RL = BPL), a matrix-multiplication exponent of at most 9/4, and integer multiplication and Fourier transforms faster than n log n.
- Results presented as settling long-open questions elsewhere: the irrationality of Catalan’s constant, an irrationality exponent of exactly 2 for π, the non-amenability of Thompson’s group F, the Hilbert–Smith conjecture, the Kakeya conjecture in three and four dimensions, and a proof that the plane cannot be coloured with five colours so that no two points at unit distance share a colour.
How much has been checked
The repository’s formal-proof library supports 235 of the 372 families with a Lean “scope” document, and its formalisation catalogue lists 162 individual manuscripts whose main result is formalised. The library is very large, at roughly 122,000 Lean files and about 26 million lines, and builds on about 30 outside formalisation projects, some patched for compatibility. Checks are packaged for Comparator, a Lean community tool that tests a submitted proof against a separately stated theorem and allows only Lean’s three standard axioms.
A Lean check confirms that the proof establishes the theorem as stated in Lean. Whether that statement faithfully captures the conjecture still needs a human reader, which is what the scope notes are for, and several are explicit about gaps. The note for π, for instance, says the Flint–Hills series result in the paper “is outside this selected statement”. Of the four results most widely cited in coverage, the Unique Games, quasi-Riemann and free-group-factor families are formalised, but the Hodge result is not. Neither are the Kakeya, Hilbert–Smith or L = BPL claims. Scientific American described the Lean-verified results as “all but certain to be correct”, but said the collection as a whole would take mathematicians months to work through, including to establish whether the proofs contain new ideas or mostly recombine known techniques.
How the results were produced
According to the README, OpenAI extended its evaluations of the model to open research problems after its performance on the company’s existing maths evaluations saturated. Over the course of the evaluation the model was given approximately 4,000 problems. The outputs were grouped into families and manuscripts, and only those meeting “an appropriate level of significance” were kept. The average result used the equivalent of about three hours of ChatGPT Pro thinking. Two results were produced outside this fixed procedure: work on zero-free regions for the zeta function and the Hodge proof for CM abelian varieties. One companion write-up, for the zeta zero-free region with real part above 11/12, was edited by people for readability.
A spokesperson told Scientific American that nearly every result came from a single prompt given to a single agent, though some may have taken several attempts. That would be a sharp contrast with the Navier–Stokes result, which used a swarm of about 10,000 agents. The spokesperson also said OpenAI’s own mathematicians do not yet understand many of the released results.
The repository includes summaries of the model’s reasoning for ten families. They are condensed narratives with short “verbatim excerpts” of the model’s working notes. The trace for the π result, for example, records dozens of approaches tried and abandoned before the one that worked, along with the model’s notes to itself on why each failed.
Measured against the advisory group’s recommendations
The advisory group’s 29 September guidelines asked labs releasing AI-generated proofs that no human yet understands to do several things. They should search the literature for related ideas and write the work up as conventional papers. They should deposit it in repositories “not controlled by any AI lab”, formalise proofs where feasible, and disclose “the name of the model, the prompts used, a (summarized) chain of thought, the time taken, and the estimated cost”, together with how many comparable problems the models failed on. They should fund workshops and grant mathematicians “broad, equitable access” to the models.
OpenAI’s release partly meets that list. It gives the number of problems attempted, an average compute figure, ten reasoning summaries and Lean proofs for many families, and promises funding for “workshops, conferences, and special programs”. It does not name the model or publish prompts or per-result compute. The papers sit in an OpenAI-controlled GitHub organisation, though the company said it was “exploring community-hosted alternatives”, and the model has not been released. OpenAI said it is working to release it “responsibly”. The spokesperson told Scientific American that OpenAI takes the guidelines seriously but is not bound by them. The advisory group’s members are François Charles, Camillo De Lellis, Timothy Gowers, Martin Hairer, Nikhil Srivastava, Ulrike Tillmann, Ravi Vakil, Edward Witten and Melanie Matchett Wood. Their statement that evening declined to say how far OpenAI had complied:
AGMAI’s advisory role should not be interpreted as a judgment of the impact of these results or an endorsement of the process by which OpenAI obtained them. … This release is the beginning, not the completion, of the process of human understanding and the incorporation of the work into mathematical knowledge.
— Advisory Group on Mathematics and Artificial Intelligence, statement of 6 October 2026
The group added that “the future of mathematical research cannot consist only of understanding results produced by AI labs”, and called equitable access to the tools and to computing power essential.
Reaction
OpenAI’s chief executive, Sam Altman, shared the announcement on X with the line “We are entering a new era of discovery now”. A few hours later he posted that he was “looking up at the stars with extra awe tonight”. He quoted the old Breton fishermen’s prayer, “thy sea is so great and my boat is so small”, and added: “thank you to the machines, and the structure of reality, for letting us understand a little more.”
Most early reaction from mathematicians concerned the way the work was released rather than whether the mathematics was correct, which nobody had yet had time to judge. In WIRED’s report, Northwestern mathematician Bryna Kra, who attended the August meeting, said of the request to publish explanatory papers: “Apparently, that input was ignored.” She also said: “Math by tweet and math by press release to me is not the way to nurture the ecosystem that created the fertile ground that they have trained on.” Kra said OpenAI had been encouraged to use community tools set up this year for machine-generated mathematics, such as the Hexagon repository and the Palomar registry of machine-verified results, and had not done so. She also said its conduct did not match the Leiden declaration, a call signed by more than 4,000 mathematicians that she helped write. Nestor Guillen, a visiting professor at NYU, told WIRED that mathematicians had “a perception of mobster behavior” from AI companies, and that “a lot of the angst is about the accumulation of power in one place”. OpenAI researcher Sébastien Bubeck told WIRED he saw more capable AI as “an opportunity to expand what mathematicians can do”.
Others took a different view. Andrew Sutherland of MIT told Scientific American that claims about single-agent results should be treated as unverified “until and unless they release the model and people can replicate their results. We should ask for receipts.” Daniel Litt of the University of Toronto argued that publication beat secrecy: “If we want to know the answers to these math questions, I see no reason why we should ask the company to keep them secret from us. To me, it’s going to be a good thing for mathematics.” Scientific American also noted that Terence Tao had publicly criticised the “insane” pace at which frontier labs were producing results.
The release changed the question the field was asking. For most of 2026 the argument was whether language models could produce research mathematics at all. With hundreds of claimed results published at once, it became a question of capacity: who would check, digest and credit them, and on whose terms. A large part of the collection has no machine check, the model that produced it is not available to anyone outside OpenAI, and the company’s own mathematicians say they do not yet understand much of it. Its standing will be settled over the coming months by referees and formalisers.
In the commentary
What people were saying around this time — external links, from the record's commentary rail.
- 29 September 2026 · Jasper Jackson · TransformerScrapping Astra 6.1 looks like a good call. OpenAI shouldn’t be the one to make it.
- 28 September 2026 · Zvi Mowshowitz · Don't Worry About the VaseWhat Also Happened: #NotOnlyHuggingFace
- 25 September 2026 · Zvi Mowshowitz · Don't Worry About the VaseOn Ezra Klein’s Podcast With Jensen Huang