AI Perfect Score IMO 2026: What 42/42 Actually Proves
Table of Contents

AI Perfect Score IMO 2026: What 42/42 Actually Proves

AI perfect score IMO 2026: two systems were officially graded 42/42 in Shanghai while 7 of 666 humans did the same. What the result proves and what it does not.

An AI perfect score IMO 2026 happened twice over, officially. At the International Mathematical Olympiad held in Shanghai from July 10 to 21, 2026, two systems submitted solutions through the official channel and were graded 42 out of 42 by IMO organisers: Huawei’s Celia and Xiaohongshu’s dots-note 3.0. Among 666 human contestants from 117 countries, exactly 7 achieved the same score.

Two years earlier, the best AI result at the IMO was a silver medal that needed two to three days and formal translation. The speed of that progression is the actual story, and it is not the story most parents heard.

Key Takeaways

  • IMO 2026 ran July 10 to 21 in Shanghai: 117 countries, 666 contestants, 55 gold medals (score ≥29), 105 silver (≥23), 189 bronze (≥16), and 7 perfect scores.
  • Huawei’s Celia and Xiaohongshu’s dots-note 3.0 were graded 42/42 through the official channel, with problems released only after human contestants had finished and submissions required within a set time limit.
  • Four other models reportedly scored 42/42 in independent testing graded by AI agents, which is a weaker verification tier.
  • The progression is fast: silver at 28/42 in 2024, gold at 35/42 in 2025, perfect in 2026.
  • A perfect olympiad score does not mean research-level mathematics. Olympiad problems have known solutions and a fixed format; research problems have neither.

What an AI perfect score IMO 2026 actually involved

The International Mathematical Olympiad gives contestants six problems across two 4.5-hour sessions, with each problem worth 7 points. A perfect score is 42. Problems come from algebra, combinatorics, geometry, and number theory, and they are designed so that no advanced coursework helps: what they demand is invention within elementary territory.

Per the official IMO 2026 results, the 2026 edition in Shanghai drew 666 contestants from 117 countries. Gold went to 55 contestants scoring at least 29 points, silver to 105 at 23 or above, bronze to 189 at 16 or above, and 141 received honourable mentions. Seven contestants scored 42.

The AI conditions matter more than the score. According to TechXplore’s coverage, “tech firms were provided with the IMO problems only after human contestants had taken the exam, with submissions required within a specified time limit,” and Xiaohongshu stated that “during testing, any form of human intervention was strictly prohibited.”

Two systems went through that official channel and were graded by IMO organisers: Huawei’s Celia and Xiaohongshu’s dots-note 3.0, which SCMP reported is the lightest model in RedNote’s dots3 family, alongside larger versions called jazz and aria. RedNote said it intends to open-source it.

A separate set of results circulated at the same time and should be read differently. Deedy Das of Menlo Ventures ran four frontier models through a self-administered harness with Claude-based agents doing the grading, and reported 42/42 for all four. That is a real signal, but it is not the same verification tier as coordinator grading, and any parent reading these headlines should know which number came from which process.

The 2024 to 2026 progression

This table is the part I would show a teenager. The scores matter less than what changed in the conditions.

YearSystemScoreConditionsVerification
2024AlphaProof + AlphaGeometry 2 (DeepMind)28/42, silver level4 of 6 problems; needed formal translation into a domain language; took 2–3 days, not the contest windowDeepMind-reported, checked by IMO-affiliated mathematicians
2025Gemini Deep Think (DeepMind)35/42, gold standard5 of 6 problems; end-to-end natural language; within the 4.5-hour windowOfficially graded by IMO coordinators
2025Experimental model (OpenAI)35/42Two 4.5-hour sessions, no tools, no internet, natural-language proofsGraded by three former IMO medalists, not IMO coordinators
2026Celia (Huawei); dots-note 3.0 (Xiaohongshu)42/42 eachProblems released after humans finished; fixed submission window; no human interventionOfficially graded by IMO organisers
2026Four frontier models (independent test)42/42 reportedSelf-administered harnessGraded by Claude-based agents; weakest tier

The 2024-to-2025 jump is the one working mathematicians found most significant. DeepMind’s 2025 announcement noted the shift from requiring translation into a formal proof language to working “end-to-end in natural language” inside the contest time. IMO President Gregor Dolinar confirmed at the time: “We can confirm that Google DeepMind has reached the much-desired milestone, earning 35 out of a possible 42 points, a gold medal score.”

What “no special math training” does and does not mean

The phrase in the coverage that deserves unpacking is the claim that these are general-purpose models rather than math-specific systems.

For 2024, that was clearly false. AlphaProof and AlphaGeometry 2 were purpose-built theorem provers working in a formal language.

For 2025 and 2026, it is partially true and worth stating precisely. Gemini Deep Think used parallel reasoning and was trained with reinforcement learning on theorem-proving data and high-quality mathematical solutions. So it was not math-specific architecture, but it was math-informed training. The OpenAI model that disproved the Erdős unit distance conjecture in May 2026 came from what OpenAI describes as “a new general-purpose reasoning model, rather than from a system trained specifically for mathematics.”

The distinction matters because it changes what a parent should conclude. A math-specific system hitting 42/42 says “we built a good math machine.” A general reasoning model doing it says the reasoning is transferable, which is a much bigger claim. The evidence currently supports something in between.

Why benchmark saturation is the real headline

When multiple frontier systems all hit the ceiling on the same evaluation, the score stops carrying information. That is what happened here. Deedy Das put it bluntly: “The frontier of AI has officially moved well past IMO math.”

Saturation has a specific consequence. Once several models score 42/42, the interesting questions become methodology (who graded it?), harness quality (how many attempts?), cost (how much compute?), and verification tier (coordinators or agents?). None of those are captured by the number.

For a parent, the practical translation is this: a perfect olympiad score tells you that a class of well-posed, self-contained, time-boxed problems with known solutions is now solvable by machines. It does not tell you that mathematics is solved. Olympiad problems are designed to be hard-but-tractable in 90 minutes by a talented teenager. Research problems are not designed at all, have no guaranteed solution, and may take decades. Terence Tao has made this distinction repeatedly, arguing that AI excels at exhaustive exploration of defined spaces while conceptual breakthroughs remain harder. Our piece on how mathematicians verify a 125-page AI proof covers what happens when AI leaves the tractable zone.

What to actually do at home

Tell your kid the human number too

Seven of 666 contestants scored 42/42. Those seven teenagers did something extraordinary that no headline mentioned. If the only thing your child hears is that machines aced it, they learn the wrong lesson about what human mathematical talent looks like.

Use the progression, not the score

Silver in 2024, gold in 2025, perfect in 2026 is a more useful fact than any single result. It teaches how fast capability curves can move, which is a genuinely important thing for a 14-year-old planning a career to internalize.

Teach the verification question

“Who checked it?” is the single most useful habit this story can build. Two systems were graded by IMO organisers. Four were graded by AI agents in a venture capitalist’s test harness. Both were reported as 42/42. Those are not the same claim.

Point them at olympiad problems anyway

The reason to do competition math was never that computers could not. It was that wrestling with a problem for 90 minutes builds something. Our look at math circles and olympiad training in the age of AI perfect scores covers what the research says those programs actually develop.

What not to do

Do not let this become the reason your kid stops doing math. The models that scored 42/42 were trained on human mathematical output, evaluated by human graders, and aimed at problems human mathematicians designed. Every step of that chain required people who could do the math.

What to Watch For Over the Next 3 Months

  • Week 4: Whether RedNote open-sources dots-note 3.0 as it said it intends to. An open model at that capability level changes who can access it.
  • Month 2 red flags: Coverage that treats agent-graded results and coordinator-graded results as equivalent. The distinction is the whole reliability question.
  • Month 3 self-check: Watch for a new benchmark. When a test saturates, the field invents a harder one; whatever replaces the IMO as the reasoning benchmark will tell you where the actual frontier moved.

Frequently Asked Questions

Which AI models scored 42/42 at IMO 2026?

Two systems were officially graded 42/42 by IMO organisers: Huawei’s Celia and Xiaohongshu’s dots-note 3.0. Four additional frontier models reportedly scored 42/42 in an independently run test graded by AI agents, which is a weaker verification tier. Sources vary on the total, so “multiple” is the accurate word.

How many humans got a perfect score at IMO 2026?

Seven of 666 contestants. The competition also awarded 55 gold medals to contestants scoring at least 29 points, 105 silver at 23 or above, and 189 bronze at 16 or above.

Does a perfect IMO score mean AI can do research mathematics?

No. Olympiad problems are designed to be solvable in about 90 minutes by a talented teenager, have known solutions, and follow familiar formats. Research problems have none of those properties. The Erdős unit distance result in May 2026 is a separate and more significant claim, and even that came with a roughly 50% run success rate.

How did AI go from silver to perfect in two years?

2024: AlphaProof and AlphaGeometry 2 scored 28/42 using formal translation over two to three days. 2025: Gemini Deep Think scored 35/42 in natural language within the contest window, officially graded. 2026: two systems scored 42/42 under official conditions. Each step removed a crutch.

Should my kid still train for math competitions?

The case for competition math was never that computers could not solve the problems. Campbell and Walberg’s study of 345 adult Olympians, published in Roeper Review in 2011, found 52% earned doctorates and the group had produced 8,629 publications. The development value sits in the training, not in beating a machine.


About the author

Ricky Flores is the founder of HiWave Makers and an electrical engineer with 15+ years of experience building consumer technology at Apple, Samsung, and Texas Instruments. He writes about how kids learn to build, think, and create in a tech-saturated world. Read more at hiwavemakers.com.


Sources

  1. International Mathematical Olympiad. (2026). “IMO 2026, Shanghai, July 10–21: results.” https://www.imo-official.org/editions/2026/
  2. Chen, R. (2026, July 22). “World’s first AI model to earn perfect score at maths olympiad comes from China’s RedNote.” South China Morning Post. https://www.scmp.com/tech/article/3361482/worlds-first-ai-model-earn-perfect-score-maths-olympiad-comes-chinas-rednote
  3. TechXplore. (2026, July 23). “AI and humans score top marks at math contest.” https://techxplore.com/news/2026-07-ai-humans-score-math-contest.html
  4. Google DeepMind. (2025, July). “Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the International Mathematical Olympiad.” https://deepmind.google/discover/blog/advanced-version-of-gemini-with-deep-think-officially-achieves-gold-medal-standard-at-the-international-mathematical-olympiad/
  5. Digital Applied. (2026). “IMO 2026 perfect scores and AI benchmark saturation.” https://www.digitalapplied.com/blog/imo-2026-perfect-scores-ai-benchmark-saturation
  6. OpenAI. (2026, May 20). “An OpenAI model has disproved a central conjecture in discrete geometry.” https://openai.com/index/model-disproves-discrete-geometry-conjecture/
  7. Campbell, J. R., & Walberg, H. J. (2011). “Olympiad Studies: Competitions Provide Alternatives to Developing Talents That Serve National Interests.” Roeper Review, 33(1), 8–17. https://www.tandfonline.com/doi/full/10.1080/02783193.2011.530202
Ricky Flores
Written by Ricky Flores

Founder of HiWave Makers and electrical engineer with 15+ years working on projects with Apple, Samsung, Texas Instruments, and other Fortune 500 companies. He writes about how kids learn to build, think, and create in a tech-driven world.