Can AI Do Original Mathematics? The Honest Answer
Table of Contents

Can AI Do Original Mathematics? The Honest Answer

Can AI do original mathematics? After September 2026 the answer has three layers: what was claimed, what a machine checked, and what nobody has confirmed yet.

The question of AI original mathematics got a concrete test case on September 8, 2026, when OpenAI posted a claimed resolution of the Navier–Stokes problem, one of the seven Millennium Prize Problems. The company’s own write-up says the work “resolves the Navier–Stokes Millennium Prize problem by establishing statement ‘C’ (and also ‘D’),” and that it included a Lean formalization. OpenAI also said it does “not intend to claim the Millennium Prize for this result.”

As of October 2026, Wikipedia’s entry on the Millennium Prize Problems states that the solution “awaits scientific validation.” That sentence is the most important one in this article.

Key Takeaways

  • The claim is specific and narrow. Statements (C) and (D) are the breakdown statements: that there exists an initial condition and a smooth external force for which no smooth solution exists. That is a different claim from “smooth solutions always exist.”
  • The proof was formalized in Lean, a proof assistant, and OpenAI reported an additional 17 hours of verification via GPT-6 Astra. Machine-checked is a real and strong form of correctness.
  • Machine-checked is not community-accepted. Clay Mathematics Institute rules require peer-reviewed publication and the Institute “only considers proposed solutions to Millennium Prize problems after at least two years since publication.”
  • Scale: OpenAI described an internal model more capable than GPT-6 Astra, coordinating roughly 10,000 concurrent agents, 2.7 million messages and about 130 billion output tokens, finishing in around 88 hours.
  • There are priority disputes. Tristan Buckmaster of NYU alleged OpenAI used non-public results of his; OpenAI initially said it “cannot rule out that de-identified data” from his product usage helped improve its models, and later said an investigation found his prompts “could not have influenced the system.” An allegation and a company response are not a finding either way.

What a Machine Actually Does When It “Does Mathematics”

The mechanism first, because every sensible opinion about this depends on it.

A proof assistant is software that checks logic. The Lean community describes a proof assistant as “a piece of software that provides a language for defining objects, specifying properties of these objects, and proving that these specifications hold,” and says the system “checks that these proofs are correct down to their logical foundation.” Lean was developed principally by Leonardo de Moura, and mathlib is “a community-driven effort to build a unified library of mathematics formalized in the Lean proof assistant.”

That checker is the ground truth. It does not guess, it does not have a style preference, and it cannot be persuaded. Either the proof term type-checks or it does not.

Now add a model. The search for a proof is enormous: at every step there are many plausible next moves, and almost all of them lead nowhere. A language model trained on mathematics can rank candidate next steps. Google DeepMind’s AlphaProof paired “a pre-trained language model with the AlphaZero reinforcement learning algorithm” and generated proofs in Lean. The model proposes; the checker disposes.

So “doing mathematics” here means three concrete things: translating a problem into formal language, searching a vast space of proof steps with learned guidance, and having every candidate checked mechanically. There is no step in that pipeline where the system understands fluid dynamics the way a person does. There is also no step where it can get away with a mistake, which is the part people underestimate.

The useful analogy, now that the mechanism is in place: a model plus a proof checker is a climber plus a rope that physically cannot hold a bad anchor. The climber chooses the route. The rope decides whether the route holds.

Can AI Do Original Mathematics? Three Senses of the Word

“Original” does a lot of work in this question, and separating its meanings dissolves most of the argument.

Sense one: a proof nobody had written. Here the answer is yes, and it was yes before September 2026. On July 25, 2024, Google DeepMind reported that AlphaProof and AlphaGeometry 2 solved four of six problems at the International Mathematical Olympiad, scoring 28 of 42 points, one short of the 29-point gold threshold. The solutions were graded by Prof Sir Timothy Gowers and Dr Joseph Myers under official IMO rules. AlphaProof solved the hardest problem, which only five human contestants solved that year. One solution came within minutes; others took up to three days.

Sense two: a new definition, framework or question. Here the answer is no, or at least not demonstrated. Olympiad problems and Millennium statements are posed by humans. Nobody has shown a system deciding that a different question was the interesting one, inventing the vocabulary to ask it, and persuading a field to adopt that vocabulary. That is most of what mathematicians mean by originality, and it remains untouched.

Sense three: a result the community accepts as settling a known open problem. Here the answer is not yet, and the gap is procedural rather than philosophical. Clay’s process requires peer-reviewed publication plus a two-year waiting period before the Institute will even consider a submission. Martin Bridson, president of the Clay Mathematics Institute, emphasised peer-review requirements, and CMI’s public response was that it “shares in the excitement of the global mathematical community as we contemplate the announcement.” Excitement is not endorsement, and the Institute chose its words carefully.

A parent can hold all three at once without contradiction. Something real happened. Something else did not happen. And the thing claimed is in a queue.

The Narrow Thing That Was Claimed, Stated Precisely

This detail is worth getting right because almost no coverage states it.

The Clay problem, in the official formulation by Charles L. Fefferman, comes in four statements. (A) says smooth solutions always exist in Euclidean space with no external force. (B) is the same on the torus. (C) says there is at least one initial condition with a smooth external force that produces no smooth solution in Euclidean space. (D) is the torus version of (C). Wikipedia’s article on Navier–Stokes existence and smoothness notes that (A) and (C) are not logical negations of each other, because the external force requirements differ.

OpenAI’s claim is (C) and (D). In plain language: not “fluid equations always behave,” but “here is a case where they break.” A solution can develop a singularity, where velocity becomes unbounded in finite time while kinetic energy stays bounded. The Wikipedia description compares it to “a spinning top” growing increasingly thin.

Clay’s rules award the prize for solving any one of the four statements, so establishing (C) would resolve the problem. But the headline “AI solves fluid equations” gets the direction backwards, and a kid who learns the actual claim learns something more interesting than the headline.

How to Teach Your Kid About AI and Original Mathematics

Ages 5–8: Finding versus checking

Materials: twelve coins, buttons or dried beans.

Ask your child to arrange all twelve into a perfect rectangle. They will find 3×4. Ask for another. They find 2×6, then 1×12. Ask whether there are any more, and whether they are sure.

Then split the two jobs out loud. “Finding the rectangle is one job. Making sure there are no others is a different job.” Have them check by trying 5 (leftover), 7 (leftover), 8 (leftover). The checking is slow and boring and completely certain.

Name it: “The computer is very good at the boring certain part. The interesting part is deciding which question to ask.” A six-year-old will remember this, and it is the same distinction the whole adult debate turns on.

Ages 9–12: The conjecture notebook

Materials: a notebook, a pencil.

Give your child this pattern and nothing else: 1, then 1+3, then 1+3+5, then 1+3+5+7. Have them compute each total and look for a pattern. They will usually spot squares: 1, 4, 9, 16.

Now three columns in the notebook. Column one, the guess. Column two, every case they tested. Column three, why it works. The third column is the hard one, and the honest answer for a ten-year-old might be a picture: a square growing by an L-shaped border each time, and each border has an odd number of squares in it.

The lesson: testing twenty cases is evidence. Drawing the growing square is a proof. Both are useful and they are not the same thing. Keep the notebook going for a month with a new pattern each week.

Ages 13+: Break a proof, then check one

Two parts, and the first is more fun.

Write out a short algebraic “proof” that contains a division by zero, the classic one that concludes 1 = 2. Hand it over without comment and ask them to find the exact line where it fails. Most teenagers find it in ten minutes and never forget that a persuasive argument can be wrong at exactly one step.

Then, if they are interested, try a proof assistant. Lean has a free web interface and the community site explains the ideas. Proving something trivial, like that addition is commutative for natural numbers, takes an evening and changes how a teenager thinks about the word “proof.” They will discover the thing that matters here: the computer refuses to accept hand-waving, and finding that annoying is the correct reaction.

The question to ask: “If a computer checked the proof and said it was right, what could still be wrong?”

What Has Been Demonstrated, and What Has Not

ClaimStatusEvidence or gap
A machine can produce a proof no human wroteDemonstratedIMO 2024: 28/42, four of six problems, graded by Gowers and Myers
A machine proof can be mechanically verifiedDemonstratedLean formalization; OpenAI reported 17 further hours of verification
A machine can work at enormous scaleDemonstrated~10,000 concurrent agents, 2.7M messages, ~88 hours
A machine resolved a Millennium Prize ProblemClaimed, not confirmedAwaits scientific validation as of October 2026
A machine invented a new mathematical frameworkNot demonstratedProblems were human-posed in every published case
The result will win the prizeNot applicableOpenAI said it does not intend to claim it; Clay requires publication plus two years

Six rows, three different answers. That is the shape of an honest summary, and it is more useful to a curious thirteen-year-old than either triumph or dismissal.

What to Do at Home

Separate “a machine checked it” from “mathematicians agree”

These are genuinely different standards and both are legitimate. Formal verification rules out logical error in the proof as written. Community acceptance also checks whether the statement proved is the statement that matters, whether the formalisation faithfully captures the problem, and whether the definitions were set up fairly. A teenager who can say that sentence understands more than most commentary.

Treat the priority dispute as a lesson in evidence, not gossip

Buckmaster’s allegation and OpenAI’s shifting statements are a clean example of how claims about provenance get resolved: slowly, with internal investigations whose methods are not public. Walk through it with an older kid as a reasoning exercise. What would count as proof here? Usually logs, which only one party holds. That asymmetry is worth noticing and it comes up everywhere.

Use the two-year clock as a patience lesson

Clay’s requirement of peer-reviewed publication plus at least two years before consideration is an institution deliberately slowing itself down. Explaining why, that extraordinary claims have historically needed time to survive scrutiny, is a more durable lesson about science than any single result.

Keep doing arithmetic and proof anyway

The temptation after a headline like this is to conclude that learning mathematics is pointless. Nothing in the September 2026 result supports that. The systems involved required human-posed problems, a human-built formal library, and human judgement about which claim mattered. Stanford’s AI Index notes models that excel at Olympiad problems still struggle with complex reasoning benchmarks like PlanBench, which is a precise way of saying capability is uneven.

What not to do

Do not tell your kid this is hype. It is not. A formalised proof of a long-open statement, if it holds, is a serious mathematical event, and dismissing it will make you unpersuasive when the verification news arrives. Hold the uncertainty honestly instead. “Something big may have happened and we will know in a year or two” is both accurate and more interesting than a verdict.

What to Watch For Over the Next 3 Months

  • Week 4: Look for a preprint on arXiv and a public Lean file. A claim with a downloadable formalisation that anybody can re-verify is in a different category from a blog post. Check whether independent mathematicians have reported running the check themselves.
  • Month 2 red flags: Named specialists in fluid dynamics raising specific objections to the formal statement rather than to the proof steps. That would be the serious kind of problem, because it would mean the formalisation did not capture the Clay statement. Also watch for silence, which in mathematics usually means people are reading.
  • Month 3 self-check: Ask your teenager to explain the difference between statement (A) and statement (C) of the Navier–Stokes problem. If they can say that one claims fluids always behave and the other claims they sometimes break, they understand the result better than most news coverage did.

Frequently Asked Questions

Did an AI solve the Navier–Stokes problem?

OpenAI claimed on September 8, 2026 to have established statements (C) and (D), with a Lean formalization. As of October 2026 that solution awaits scientific validation, and the Clay Mathematics Institute requires peer-reviewed publication plus a two-year waiting period before it will consider a submission. “Claimed and formally checked, not yet community-verified” is the accurate description.

If a computer verified the proof, why is anyone still unsure?

Because verification answers one question: does this argument follow from these axioms. It does not answer whether the formal statement faithfully encodes the Clay problem, whether the definitions were chosen fairly, or whether the result means what the summary says. Those judgements are human and they take time.

Why does OpenAI not want the prize?

The company stated it does “not intend to claim the Millennium Prize for this result.” The reasons are not fully explained in the announcement. Note that claiming it would require a peer-reviewed publication and a two-year wait under Clay’s rules regardless of who or what produced the proof.

Can a machine come up with a new mathematical idea, not just a proof?

No published example demonstrates that. In every case so far, including the IMO results and the Navier–Stokes claim, humans posed the problem and humans built the formal libraries. Producing a proof and choosing what is worth proving are different capabilities.

Should my kid still learn to write proofs?

Yes, and the reason is sharper than before. The valuable human skills in this new arrangement are deciding what to prove, checking that a formal statement matches the informal question, and spotting where an argument assumes something it should not. All three are trained by writing proofs by hand.

What is Lean, in one paragraph?

Lean is a proof assistant: software that provides a language for defining mathematical objects, stating properties and proving them, and that checks proofs “down to their logical foundation.” Its mathematics library, mathlib, is built collaboratively. If a proof type-checks in Lean, it contains no logical gap, which is a narrow but very strong guarantee.


About the author

Ricky Flores is the founder of HiWave Makers and an electrical engineer with 15+ years of experience building consumer technology at Apple, Samsung, and Texas Instruments. He writes about how kids learn to build, think, and create in a tech-saturated world. Read more at hiwavemakers.com.


Sources

  1. OpenAI. (2026, September 8). “Navier–Stokes solution.” https://openai.com/index/navier-stokes-solution/
  2. Wikipedia contributors. (2026). “Millennium Prize Problems.” Wikipedia. https://en.wikipedia.org/wiki/Millennium_Prize_Problems
  3. Clay Mathematics Institute. “Navier–Stokes Equation” and the official problem description by Charles L. Fefferman. https://www.claymath.org/millennium/navier-stokes-equation/
  4. Wikipedia contributors. “Navier–Stokes existence and smoothness.” Wikipedia. https://en.wikipedia.org/wiki/Navier%E2%80%93Stokes_existence_and_smoothness
  5. Google DeepMind. (2024, July 25). “AI achieves silver-medal standard solving International Mathematical Olympiad problems.” https://deepmind.google/discover/blog/ai-solves-imo-problems-at-silver-medal-level/
  6. Lean Community. “Lean and mathlib.” https://leanprover-community.github.io/
  7. Stanford Institute for Human-Centered AI. (2025). AI Index Report 2025, Chapter 2: Technical Performance. https://hai.stanford.edu/ai-index/2025-ai-index-report
  8. Wikipedia contributors. (2026). “2026 in science.” Wikipedia. https://en.wikipedia.org/wiki/2026_in_science

Related reading on HiWave Makers: the Navier–Stokes claim explained for kids, how mathematicians check a 125-page AI proof, and proof assistants your kid can actually try.

Ricky Flores
Written by Ricky Flores

Founder of HiWave Makers and electrical engineer with 15+ years working on projects with Apple, Samsung, Texas Instruments, and other Fortune 500 companies. He writes about how kids learn to build, think, and create in a tech-driven world.