When proofs become abundant…OpenAI's 722 maths manuscripts, the argument over what counts as progress, and what this might mean beyond mathematics
These maths results seem pretty spectacular even if not all will be proved out. This event raises questions about the human endeavour in maths which I'm pretty unsure about (Terry Tao is worth reading here). Maths specifically has a language, Lean, which can help verify results, so it doesn't necessarily read across to all other human endeavours. Still, it seems to me to be a clear indicator... worth pondering.
This is what happened, what's important, the critiques, and what it might mean. (LLM assisted)
First: what actually happened?
This week OpenAI released an extraordinary quantity of mathematics in one go: 722 manuscripts, grouped into 372 families of related results, spanning number theory, geometry, theoretical computer science, analysis, probability, mathematical physics and much else. The work was produced mainly by an unreleased internal frontier model.
OpenAI says it posed roughly 4,000 research problems to the model. The retained results used, on average, compute equivalent to about three hours of ChatGPT Pro thinking each. There are important exceptions: the quasi-Riemann work and the CM Hodge result did not follow exactly the same procedure. OpenAI has also released ten abridged reasoning summaries and machine-checkable Lean formalizations for many of the papers.
The first thing to resist is the headline 'AI solved 722 open problems'. The 722 are manuscripts, not 722 independent discoveries. A family can contain a main theorem, a companion paper, consequences or a different proof. And the release deliberately contains work at different stages of verification. OpenAI itself says some unformalized results could have issues.
Even after applying that discount, the list is startling. Several of the claimed results concern questions that would normally deserve their own major announcement. Here they arrive as entries in a catalogue.
I am not going to pretend I can referee any of the deep proofs. What one can do already is separate three questions that are too easily collapsed into one. Is the stated result mathematically important? Is there a machine-checked formal proof of the central claim? And has the human mathematical community actually understood, stress-tested and accepted it? Those are different thresholds.
The findings that look most important
My provisional ranking is below. 'Importance' is the significance the result would have if it survives in essentially its advertised form. The last column describes the public checking I can see at the time of writing. Lean is powerful evidence, but it is not the same thing as expert community acceptance. For the unformalized blockbusters I would keep a large discount until specialists have gone through the papers.
Here’s a list:
Rank / Finding What it says / why it matters
Importance / Evidence today
1
Quasi-Riemann: zero-free half-plane Re(s) > 7/8
For the zeta function and all Dirichlet L-functions, no zeros in a fixed strip to the right of 7/8. This is qualitatively stronger than the classical zero-free regions that narrow with height.
5/5
Lean formalisation; zeta theorem independently replayed through Lean and a second Rust kernel. Human expert digestion still early.
2
Hilbert's tenth problem over Q
Claims there is no algorithm that can decide whether an arbitrary integer-coefficient polynomial has a rational solution. A famous open problem in logic and number theory.
5/5
Preprint claim; no central Lean formalisation located in the release catalogue. Needs serious expert scrutiny.
3
Rational Hodge conjecture for CM abelian varieties
Claims the Hodge conjecture in every codimension for all complex CM abelian varieties, with major Tate-conjecture consequences. A very substantial piece of a Millennium-problem landscape.
5/5
Preprint claim; produced by an exceptional workflow. No central Lean proof located in the catalogue.
4
Unique Games Conjecture
Would settle a central problem in theoretical computer science and pin down optimal hardness thresholds for many approximation problems, including Max-Cut.
5/5
Lean formalisation of the central gap reduction and several consequences. Independent expert review is only beginning.
5
Birch-Swinnerton-Dyer in low Selmer corank / density-one twists
Claims the full BSD leading-term formula in coranks 0 and 1 and, combined with another result, full BSD for a density-one set of quadratic twists of every elliptic curve over Q.
5/5
Headline manuscripts; I would treat this as unconfirmed pending specialist review.
6
Mahler conjectures
Sharp lower bounds for the volume product of a convex body and its polar, with equality cases. One of the central long-standing problems of convex geometry.
4.5/5
Lean formalisation covers the symmetric and general geometric inequalities and equality cases.
7
Free group factor isomorphism problem
Claims all interpolated free group factors with parameters greater than 1, including infinity, are isomorphic. A major operator-algebra problem.
4.5/5
Lean formalisation of the main isomorphism statement.
8
Matrix multiplication exponent <= 9/4
Pushes the asymptotic exponent for complex matrix multiplication to 2.25, a conspicuous improvement in one of algorithms' most famous quantitative races.
4.5/5
Lean formalisation of the 9/4 bound and related rectangular bounds.
9
Kadison's similarity conjecture
Settles a problem dating to the 1950s about bounded homomorphisms of C*-algebras and similarity to *-homomorphisms.
4.5/5
Lean formalisation of the central theorem and supporting uniform estimates.
10
Global solutions for 3-D relativistic Vlasov-Maxwell
Claims global existence and uniqueness for the one-species system for broad smooth data, without smallness or symmetry assumptions. A major nonlinear PDE / mathematical-physics result.
4.5/5
Lean formalisation of the global classical solution statement.
11
Banach's simple Lebesgue-spectrum problem
Constructs a smooth volume-preserving diffeomorphism of the three-torus with simple Lebesgue spectrum on the full mean-zero L2 space.
4/5
Lean formalisation of the central construction statement.
12
The irrationality exponent of pi is exactly 2
Shows pi cannot be approximated unusually well by rational numbers in the asymptotic sense measured by irrationality exponent. A long-standing Diophantine approximation problem.
4/5
Lean formalisation of the exact exponent-2 statement.
13
Catalan's constant is irrational
Settles the irrationality of the classical constant G = 1 - 1/9 + 1/25 - ... , a question going back to the nineteenth century.
4/5
Lean formalisation of the headline irrationality theorem.
The Riemann result is a useful case study
The Riemann hypothesis itself remains unsolved. It predicts that every non-trivial zero of the Riemann zeta function lies on the line Re(s) = 1/2. OpenAI's theorem puts a zero-free boundary at 7/8. That is much weaker than 1/2, and it would be silly to say the Riemann hypothesis is now a quarter solved. Mathematics does not work like a progress bar.
The reason mathematicians care is more structural. For more than a century, unconditional zero-free regions could be pushed slightly left of 1, but the safe strip becomes thinner as one moves higher. A fixed line strictly below 1 had remained out of reach. The OpenAI formalisation says 7/8 works uniformly for every Dirichlet L-function, as well as for zeta itself. It also gives a uniform logarithmic exclusion region for Landau-Siegel zeros.
This is one of the strongest parts of the release because the formal theorem matches the interesting headline closely. There is already an independent technical re-run: the zeta proof was compiled through Lean's kernel and then through an independently written Rust kernel called nanoda. Both accepted it. The proof closure runs to almost half a million lines of OpenAI modules.
There is a caveat worth keeping in view. That external check was carried out by Claude Code on one person's machine. It verified the formal artefact; it did not amount to a roomful of analytic number theorists reading the paper, finding the core idea and saying, yes, this changes the field. We have unusually strong evidence that a precise theorem has been formally derived. We are still at the start of understanding what happened mathematically.
Why Lean matters so much here
Maths has an advantage that most intellectual work does not have. A proof can, in principle, be translated into a formal language where every logical step is checked by a small trusted kernel. Lean is one of the leading systems for doing this.
That changes the economics of AI-generated work. A language model can hallucinate. A second language model can fail to notice the hallucination. A Lean kernel has a much narrower job: does this exact formal statement follow from these exact definitions and accepted axioms? If it does, the machine accepts the proof. If it does not, compilation fails.
This gives maths a plausible architecture for abundance: a high-volume generator on one side and a rigorous filter on the other. Terry Tao has used the image of a firehose that needs a filter. Mathematics is unusually well placed to build one.
Still, Lean does not settle everything people care about. It checks the statement you formalised. A famous conjecture can be subtly weakened in translation. Hypotheses can be stronger than advertised. The formal theorem may cover only one part of a natural-language paper. Formal correctness says nothing about novelty, importance, proper attribution, or whether a human can see the idea.
So I would put the verification ladder roughly like this: a natural-language AI manuscript is a candidate result; a matching Lean proof is much stronger; independent reproduction is stronger again; expert mathematical digestion remains another step. The top OpenAI results currently sit at different places on that ladder.
Why many mathematicians are unhappy
The interesting reaction has not been a simple defence of human jobs. A lot of the strongest critics accept that the capability is real. Their question is what happens to mathematics when proofs become much easier to produce than to understand.
1. A verification and attention debt
Seven hundred papers can be generated much faster than seven hundred papers can be responsibly checked. That creates an asymmetry. OpenAI gets the headline; the mathematical community inherits months or years of unpaid validation work.
The Advisory Group on Mathematics and Artificial Intelligence (AGMAI), an independent group hosted at the Institute for Advanced Study, described the release as an important event and then immediately added that publication was only the beginning. The work still has to be assessed, understood and incorporated into mathematical knowledge. That process cannot be automated away merely by making the PDF public.
There is also a practical curation problem. Which ten of the 372 families deserve a specialist's week? Which are incremental? Which accidentally rediscover something? Which contain a subtle mismatch between the paper and the formal statement? An abundance of candidate knowledge creates a new shortage of expert attention.
2. Tao: generation, verification and digestion
Terry Tao's framework is the most useful one I have found. He separates mathematical progress into generation, verification and digestion. Generation finds a proof. Verification establishes that it is correct. Digestion turns the proof into something the community understands: the key lemma, the reusable technique, the clean example, the connection to another field, eventually the version a graduate student can learn.
Historically these stages moved at roughly human speed. If it took six months to crack a problem, spending another month writing it properly was a manageable extra cost. AI can compress the first stage from months to hours while exposition may still take weeks. The ratio changes dramatically. Tao calls the resulting problem 'proof indigestion'. Each individual researcher can become more productive while the average quality of the total literature deteriorates, because raw output grows faster than careful explanation.
This is a subtle point. A correct theorem is still a correct theorem. Yet mathematics is more than a database of true statements. The proof has value because it gives people new mental equipment. If nobody understands the equipment, some of the value stays locked inside the artefact.
3. Famous problems were landmarks, training grounds and incentives
In September, a declaration initially signed by 25 Fields Medallists — including Tao, Peter Scholze, James Maynard, Martin Hairer, Maryna Viazovska, Manjul Bhargava and Pierre Deligne — argued that famous open problems have functioned as 'landmarks and lighthouses'. They give communities something hard to push against. The failed attempts generate techniques, conjectures, collaborations and trained mathematicians.
A benchmark culture can Goodhart that system. If AI companies optimise directly for the count and prestige of famous problems solved, the old proxy starts to break. A deep theorem used to be reasonably good evidence that deep thinking had happened around it. That relationship may no longer hold in the same way.
Tao pushes the thought further. Good open problems are, in one sense, a scarce educational resource. Once a machine publishes the solution, a young mathematician can still study the problem, but the untouched landscape has changed. Knowing that a long proof exists, and roughly where it goes, affects how you explore. Industrially sweeping through a curated list of open problems may therefore consume something that the field used to use for training and discovery.
4. Failure itself contains information
The Unique Games story makes this concrete. Dor Minzer and his students Yumou Fei and Shuo Wang had spent years on a closely related problem. When rumours spread that OpenAI had solved Unique Games, they rushed a 95-page manuscript online so that their contribution would not be swallowed by the coming announcement.
Minzer's description of their human process is telling. Several approaches failed. Those failures taught them which structures mattered and eventually helped them assemble a proof. His worry is that a system which jumps straight to the successful path removes much of that learning loop. A human who sees only the final route may know where the summit is without learning the mountain.
That does not make the AI proof worse as a proof. It may make the research process less fertile for the humans around it.
5. Credit, priority and provenance become messy
Mathematics has always built aggressively on earlier mathematics. The difference now is speed and opacity. A model has absorbed an enormous literature and can combine ideas across it. It may also finish a research programme that humans have pushed close to completion. Working out where the genuinely new idea enters can be hard even for specialists.
The worry is not merely philosophical. Careers depend on priority. In this release Daniel Litt noticed that one advertised result was a special case of a conjecture of his and followed from stronger work already under way by one of his students. The result can still be useful; it shows why '372 discoveries' is too clean a description.
The earlier Navier-Stokes episode also left some mathematicians angry about how AI labs handled credit and unpublished human work. Whatever one thinks of the disputed details, the episode increased the demand for careful prior-art searches, prompt provenance and slower claims about novelty.
6. Proprietary models create a two-tier science
AGMAI's sharpest recommendation came before this release: frontier labs should stop testing advanced mathematical problems on proprietary internal models that the mathematical community cannot access. Their worry is straightforward. The lab can explore thousands of questions with an instrument more capable than anything available in a university department. The people who formulated the questions then watch the answers arrive from outside their field's normal institutions.
OpenAI says it is working towards a responsible release of the model. That matters. A world where every serious mathematician has access to this kind of collaborator feels very different from a world where a few companies privately harvest the frontier and publish the outputs.
Access is also about compute. Even if the model becomes public, the ability to run thousands of long research attempts may be concentrated in rich institutions. Mathematics has historically been unusually cheap at the point of research: paper, pencil, library, blackboard. AI could make the frontier materially more capital-intensive.
7. Publication norms have not caught up
The ordinary norms of mathematics assume that authors understand their papers, can defend them in seminars, know the nearby literature and take responsibility for errors. An AI lab can now publish a proof that nobody at the lab fully understands in that traditional sense.
AGMAI therefore asked labs to do more than upload files: search for prior art, improve exposition, reveal the model, prompts and compute, document failed attempts, formalise where possible, use independent scholarly repositories, and fund the human work of understanding what has been released. It also asked that mathematical results not be treated primarily as model marketing.
OpenAI has moved part way. It released the papers, revision histories, substantial Lean material, ten reasoning summaries, aggregate compute information and a promise to fund workshops and programmes around understanding the results. The model remains closed, the full attempt history is absent, and GitHub is controlled by OpenAI rather than an independent mathematical repository. The current release feels more transparent than a press release and still some distance from a mature scientific publication system.
8. What happens to the next generation?
This may be the most important long-run concern. A mathematics PhD is simultaneously research and training. You learn judgement by being stuck, trying bad ideas, reading around them, proving small lemmas, explaining your work and discovering that your beautiful shortcut is wrong.
If a model can answer every well-specified technical question faster than the student, the temptation to skip those stages will be enormous. Some of that is healthy cognitive offloading. Nobody wants students to spend a career doing arithmetic that a calculator handles perfectly. The harder question is where the productive struggle ends and deskilling begins.
There is also an incentive problem. Why spend ten years becoming a world expert in a narrow area if a proprietary model may solve the flagship problem next Tuesday? Quanta's account of the Unique Games race captures the psychological effect rather well. Humans sleep, eat, teach and have moods. A trillion-dollar company does not have those constraints.
The positive case is also strong
There is a risk of romanticising scarcity. Maths was slow because it was hard. Some of that slowness was productive; some was simply slowness.
If an open question has resisted us for 80 years and a machine supplies a correct proof tomorrow, humanity knows something tomorrow that it did not know today. Cheap discovery does not make the discovery fake.
Tim Gowers, another Fields Medallist and a member of AGMAI, deliberately did not sign the September declaration. He agrees that mathematics faces a crisis and worries about credit, institutions and training. He is less persuaded that understanding must always be the primary goal with problem-solving as merely a proxy. Mathematics has long had different cultures: theory builders, problem solvers and people somewhere in between.
Gowers also makes a simple quantitative point. Suppose AI produces 1,000 important results and humans only manage to digest 200 of them. If the old world would have produced and digested 100, we are still ahead. There will be more undigested mathematics and more digested mathematics. That could be a pretty good bargain.
There is another upside: machines can search across fields whose literatures have become too large for any human. A result buried in ergodic theory may be exactly the missing tool in number theory. A model with broad recall and endless patience can test combinations no individual would have time to try.
Jeremy Avigad's optimistic version is that humans move up a level. Let the machine fill in more of the technical terrain. We ask harder questions, build better concepts, find simpler proofs, decide what is beautiful and important, and explore spaces that become visible only after today's frontier has moved.
Formalisation could also enable a genuinely new collaborative scale. Large proofs can be split into modules. Contributions can come from humans, AI systems, students and specialists who do not know one another personally, with formal checking supplying a common trust layer. Tao has called this direction 'big mathematics'. It looks less like a solitary genius at a blackboard and more like a large engineering or software project — with mathematicians still setting architecture and meaning.
And there is an appealing possibility that status itself shifts. If raw proof generation becomes cheap, perhaps more credit moves to the person who asks the best question, extracts the reusable idea, connects two fields, writes the explanation everyone remembers, or builds the formal library that lets everybody else work faster. Mathematics could end up valuing taste more explicitly.
What might this mean beyond maths?
This is where I would be careful and also where I think the release matters most.
Mathematics is unusually favourable territory for AI. The literature is formal and cumulative. Many answers are crisp. Most importantly, proof assistants provide a route to mechanical verification. There is no Lean kernel for 'this company has a durable moat', 'this patient will benefit from this treatment', 'this policy will increase social welfare' or 'this play is good'.
So maths probably moves earlier than many other areas. It may also give a cleaner signal of capability because we can separate fluent nonsense from verified output more effectively than elsewhere.
One comforting story about AI has been that routine work will automate while highly creative, abstract work requiring years of specialist training will remain human territory. Research mathematics is a fairly brutal test of that story. These systems are now producing candidate original research across unrelated fields, including work on problems that generations of exceptionally able humans knew about and could not settle.
If even a modest fraction of the strongest results survive, creativity and abstraction no longer look like reliable technological moats around human work.
The more interesting economic point may be that the bottleneck moves. For most of modern research, producing the new result was scarce. Peer review, publication and interpretation were arranged around that scarcity. If candidate discoveries become abundant, other things rise in value: verification, attention, experimental confirmation, explanation, taste, choice of problem, and trust.
That pattern should travel beyond maths even when formal proof does not. Software has tests and formal verification. Chips can be simulated. Chemistry and biology have assays, though often expensive ones. AI research systems will move fastest in fields where proposed answers can be rejected cheaply and reliably.
Fields with weak verification may actually face a nastier version of the problem. A firehose of bad mathematical proofs can eventually meet Lean. A firehose of plausible investment theses, policy arguments or psychological explanations can sound excellent and remain wrong for years.
That makes the maths story both reassuring and unsettling. Reassuring because we can see a route to rigorous checking. Unsettling because the underlying generator appears able to operate at a level of abstraction we were recently inclined to treat as distinctly human.
What happens to mathematicians?
Chess did not end after Deep Blue. People still run despite cars. Photography did not eliminate painting. These analogies are useful and incomplete. Professional chess is different from chess in 1996, and few people are paid to perform arithmetic by hand.
People will continue doing mathematics because understanding something beautiful is enjoyable. I am quite confident about that. The profession of research mathematician is a different question.
Universities and governments fund research partly because researchers produce new knowledge. If frontier models become cheaper and materially better at theorem production, funders will eventually notice. The strongest case for maintaining human mathematics may become less about beating the machines at their own benchmark and more about preserving a population capable of understanding, directing and challenging them.
That requires a pipeline. You cannot preserve expert judgement by keeping a handful of senior people around to read AI output. School pupils, undergraduates, PhDs and postdocs have to practise. They need hard problems and room to fail. We may end up having to protect some forms of inefficient human learning precisely because the efficient machine alternative is so good.
This sounds odd only because we are used to the process of producing mathematics and the process of producing mathematicians being the same activity. AI may separate them.
Where I might land, for now
A day after the release, certainty would be silly. Some manuscripts will probably fail. Some will turn out to be known in stronger form. Some formal statements will prove narrower than casual headlines suggest. Several of the most spectacular claims still lack the kind of verification that would make me comfortable calling them established mathematics.
The reverse error is also possible: concentrating so much on the process controversy that we miss the capability signal. The quasi-Riemann theorem has a matching Lean formalisation and has already survived a second-kernel replay. Unique Games, Mahler, the free-group-factor result, Kadison, Vlasov-Maxwell, matrix multiplication and several other striking claims have formal artefacts tied to their central statements. If even five or ten major results ultimately stand cleanly, this is an extraordinary research output.
And the number I keep returning to is not 722. It is roughly 4,000 questions in; hundreds of significant candidate result families out; average retained-result inference measured in hours. That is a different research production function.
Tao may be right that mathematics needs to defend the slow work of digestion, failure and human formation. Gowers may be right that abundance is still a good bargain and, in any case, hard to stop once these models diffuse. I suspect both views contain part of the answer.
The sensible response is probably to change what we reward. If proofs become abundant, reward understanding. If candidate discoveries become cheap, reward judgement. If machines can search the known landscape faster than us, reward the people who make better maps and ask better questions. Make the labs pay some of the cost of verification and exposition. Give researchers access to the tools.
I would also resist the urge to make this a morality play about humans versus AI. The more interesting question is what human intellectual life looks like when one of its hardest outputs becomes abundant.
Maths lets us see that question early because it has a language for checking the answer.
Worth pondering.
Sources and further reading
OpenAI mathematics repository — 722 manuscripts / 372 result families
'A Severe Misalignment of AI in Mathematics' — Fields Medallists' declaration
Terence Tao, living summary of his views on AI and mathematics
Tim Gowers, 'Why I didn't sign the Fields medallists' letter'
Quanta, 'As AI Closed In on Unique Games Proof, Researchers Raced to Beat the Machines'
WIRED, 'OpenAI Is Pissing Off a Bunch of Mathematicians — Again'
Note: this article reflects the public evidence available at the time of writing. The mathematical status of individual claims may change quickly as specialists review, reproduce, correct or reject them. The article was LLM assisted.
