← All stories
📚

OpenAI's large batch of AI-generated mathematical results

OpenAI released a substantial catalog of AI-produced mathematics: 722 manuscripts organized into 372 “result families,” generated by an unreleased internal model. According to OpenAI, roughly 4,000 problems were posed to the system and the company published the subset of outputs it judged significant. The collection spans number theory, algebraic geometry, analysis and theoretical computer science. Some outputs were accompanied by Lean 4 formal proofs—proof assistant files that mechanically check logical steps—but reporting differs on how many: an earlier account cited 162 Lean-checked manuscripts while a later tally puts the number at about 42% of the set (roughly 300). OpenAI itself warned that not all manuscripts are formalized and that “some of the unformalized results could have issues.”

A central controversy is transparency and verifiability. OpenAI published the results on GitHub but did not release the underlying model or the prompts used to generate them; the repository had Issues turned off and had not accepted pull requests, limiting direct community feedback. The company did provide average compute figures, ten abridged reasoning summaries, and said it consulted an Advisory Group on Mathematics and Artificial Intelligence and guidelines from the Institute for Advanced Study. Prominent mathematicians including Andrew Sutherland and others have called for the model, prompts and additional “receipts” to allow independent replication, while some researchers welcomed the material as a major stimulus for research.

Contextual details matter: OpenAI previously reported work on the Navier–Stokes question involving an effort that reportedly used about 10,000 coordinating agents for 88 hours, and the present release saw at least three manuscripts retracted shortly after publication. The immediate task for the mathematical community is triage: prioritize the roughly 300 results with Lean formalizations for human attention, watch for further retractions, pursue independent verification of results tied to high-profile questions such as Riemann-related claims, and see whether OpenAI will eventually release the model, prompts and fuller methodological disclosures that the advisory group recommended.