TL;DR
OpenAI published 722 new math papers in one release, produced by an AI model the public can't use yet.
A year ago, AI math "breakthroughs" were mostly hype. Now they come with machine-checked proofs, and one of the claims is on a $1 million Millennium Prize problem.
The math itself won't affect your life. The speed matters, and so does what this kind of reasoning will do next.
How worried?: 🟡 3/5 — Meaningful
TOP STORY

The AI that wrote 722 math papers
On Tuesday, OpenAI released something no lab has released before: hundreds of new mathematical results at once. The collection contains 722 manuscripts grouped into 372 "families," which bundle a main result with related arguments, consequences, or alternate proofs. (Source: github)
To see why this matters, look at how fast things changed.
In October 2025, AI math claims were embarrassing. OpenAI said GPT-5 had solved 10 Erdős problems, and the claim was debunked within hours. The model hadn't solved them. It had dug up obscure papers where humans already had. (Source: 36kr)
Then came 2026. In May, an OpenAI model found a counterexample to a conjecture Paul Erdős proposed in 1946. On August 1, OpenAI published solutions to ten problems that had been open since at least 2016, generated by an internal version of its Astra model for roughly $2,000 in compute. Each proof came with a formal certificate written in Lean, a programming language that lets a computer check every step. (Source: mathscholarforklog)
In September the claims got much bigger. OpenAI said an internal system had resolved the Navier–Stokes problem, one of the seven Millennium Prize Problems, by showing that the equations for fluid flow can "blow up" in finite time. The effort used a group of around 10,000 AI agents working together, and they reached a solution about 88 hours after launch. OpenAI also said the model involved was significantly more capable than GPT-6 Astra. (Source: On the Navier–Stokes Millennium Prize Problem +2)
This week's 722 papers fit that pattern: more output, more ambition, and less time between releases. OpenAI says it gave the model about 4,000 problems, and each result used about three hours of ChatGPT Pro-level thinking on average. The list reaches into famous territory. It includes work on a zero-free region for the Riemann zeta function and a proof of the Hodge Conjecture for a special class of shapes called CM abelian varieties. That second one is not the full Hodge Conjecture, which is still a Millennium Prize problem, but it's a piece of it.
Not all of the 722 have been checked. OpenAI itself says the collection mixes results at different stages of verification, not every paper has a Lean proof, and some of the unformalized ones may contain errors.
Why mathematicians are angry
The pushback has been loud. A declaration signed by Fields Medalists, posted by Terence Tao, warns that rushed AI results can weaken verification, credit, and the way mathematical ideas get passed from person to person. More than 1,000 mathematicians signed an open letter criticizing AI's growing role in research, and some called OpenAI's approach "research misconduct." (Source: letsdatasciencecryptobriefing)
There's also a credit fight. NYU mathematician Tristan Buckmaster accused OpenAI of not properly accounting for whether his own Navier–Stokes work, which he did using OpenAI's Codex tool, fed into the company's result. OpenAI says its investigation found his prompts could not have influenced the system, including through training. (Source: northeasttimesopenai)
What else?
The backlash has had additional effects. OpenAI pulled its sponsorship of a Caltech "Mathathon" after mathematicians published an open letter against it. It also set up an independent advisory group at the Institute for Advanced Study to help review and communicate its results. For this release, OpenAI says it followed that group's advice, plans to fund workshops on understanding AI-produced results, and is working to responsibly release the model behind them. (Source: OpenAI Withdraws From Caltech Mathathon Sponsorship +2)
OUR TAKE
Math is actually the safest place to watch this happen, because a proof can be checked by a computer. If AI is going to reason in new ways anywhere, we want it to be somewhere errors can be caught.
The concern is what this signals. A system that can produce original arguments at this scale can also be pointed at chemistry, biology, engineering, and code. According to The Information, OpenAI engineers see dominating math as a key path toward AI that improves itself. That's the part an AI-risk reader should notice. (Source: Gizmodo)
This isn't a 4 or 5 yet because nothing here harms anyone directly, the results are being published openly, and OpenAI has slowed down a little in response to criticism. But the pace is fast, and the most capable systems are internal ones the public can't inspect.
How worried?: 🟡 3/5 — Meaningful
A 3? It is just math, why should we be worried??? Check out our ‘What If?’ section below…
WHAT TO KNOW
Here are the key takeaways:
The labs are ahead of what you can use. The models behind these results aren't public. The ChatGPT you use today is not the frontier.
It's getting cheap. Ten decade-old problems cost about $2,000 in compute. When discovery is that cheap, the bottleneck becomes humans checking the work.
"Checked by a computer" isn't the same as "understood." A Lean proof shows each step is valid. It doesn't explain the idea to anyone, and mathematicians worry about a field full of results nobody really understands.
It won't stay in math. The same reasoning ability applies to drug discovery, materials, chip design, and code. Expect claims like these in other fields soon.
Treat headlines with care. "Solved" often means "claimed, verification pending." The Navier–Stokes review isn't done yet.
WHAT IF….?
What if a future proof breaks the math that protects your bank account?
Most online security, including banking, messaging, and passwords, depends on certain math problems being too hard to solve in any reasonable time. Nobody has proven they're hard. We just haven't found a shortcut.
Now picture a model like this pointed at those problems. OpenAI's August batch already included a result in lattice cryptography, which is the family of math behind the "quantum-safe" encryption governments are moving to now. That result didn't break anything. But if a future model found a real shortcut, the lab that found it would hold a key to much of the internet, at least until everyone else updated their systems.
To be clear, experts consider this unlikely anytime soon. AI is already finding real security holes, but it's finding them in software code, not in the math itself. But it shows why "AI does math" isn't only an academic story, and why it matters who finds these results first and what they do with them.
OTHER NEWS
WHAT DO YOU THINK?
If an AI proves something no human fully understands, should we trust it?
Hit reply: are we too worried, or not worried enough? We read every response.
