Verify the Math: A Weekly Ritual for Checking AI Answers
Table of Contents

Verify the Math: A Weekly Ritual for Checking AI Answers

How to verify AI math answers with kids: the 50% run-success lesson from an OpenAI proof, five check methods ranked, and a 20-minute weekly family ritual.

In May 2026, an OpenAI model disproved a conjecture Paul Erdős posed in 1946. Real result, checked by named mathematicians. Then OpenAI researcher Sébastien Bubeck told Science News something that should be printed on every classroom wall: the team “ran their prompt on the Erdős conjecture through the same model multiple times, and it produced the correct solution in 50 percent of those trials.” Fifty percent. That’s the frontier. So when your kid asks whether they need to verify AI math answers, the answer is yes, and here’s a twenty-minute weekly ritual that makes it automatic.

Key Takeaways

  • OpenAI’s model disproved the Erdős unit-distance conjecture in May 2026, but Bubeck confirmed the correct solution appeared in “50 percent of those trials” when the prompt was rerun.
  • The earlier claim that GPT-5 had “solved 10 previously unsolved Erdős problems” was retracted after mathematician Thomas Bloom called it “a dramatic misrepresentation”; the solutions already existed in the literature.
  • Bloom noted the model’s proof “happened to be relatively easy for a human expert to verify” in this case, and asked the real question: “It could be right. It could be nonsense. Who’s going to be able to check this?”
  • Verification methods are not equal. Plugging the answer back in beats re-asking the AI, which is the method most kids default to.
  • Roediger and Karpicke (2006) found testing beats restudying for retention, and students feel more confident after restudying even though it works worse. Verification doubles as retrieval practice.

The 50% number, and what it actually means

On May 20, 2026, OpenAI announced that one of its models had disproved a conjecture about unit distances that Erdős posed in 1946. Per OpenAI’s framing, “for nearly 80 years, mathematicians believed the best possible solutions looked roughly like square grids,” and the model “disproved that belief, discovering an entirely new family of constructions that performs better.” The company obtained supporting statements from mathematicians including Noga Alon, Melanie Wood, and Thomas Bloom, who maintains the Erdős Problems website.

That part is genuinely impressive, and the proof ran to well over a hundred pages.

But Science News reported on June 8, 2026 that Bubeck said the team reran the same prompt through the same model multiple times and “it produced the correct solution in 50 percent of those trials.” Same prompt. Same model. Coin flip.

There’s context that makes this even more instructive. In October 2025, an OpenAI executive claimed “GPT-5 found solutions to 10 (!) previously unsolved Erdős problems and made progress on 11 others.” That was wrong. The “solutions” already existed in published literature. Bloom called it “a dramatic misrepresentation” and the post was deleted.

And Bloom’s assessment of the real result contains the warning. He noted the proof “happened to be relatively easy for a human expert to verify” in this particular case, then asked the question that generalizes to your kid’s homework: “It could be right. It could be nonsense. Who’s going to be able to check this?”

If a 50% run-success rate is acceptable at the research frontier because experts verify the output, then verification isn’t an optional add-on to AI use. It’s the part that makes the use legitimate.

Which check methods actually work

Not all verification is verification. Kids default to the weakest method, which is asking the AI whether it’s sure.

MethodHow it worksReliabilityTimeBest for
Substitute backPlug the answer into the original equation or conditionHigh. Either it satisfies the condition or it doesn’t30 secondsAlgebra, equations, word problems with a checkable condition
Estimate first, compareGuess the order of magnitude before looking, then compareHigh for catching gross errors20 secondsAnything with units; catches decimal-place disasters
Work backwardStart from the answer and reconstruct the questionMedium-high2 minutesMulti-step problems, geometry
Solve a second wayDifferent method, same problemHighest, and slowest5+ minutesAnything that really matters, like test prep
Check unitsDo the units come out right?High for physics and chemistry; catches setup errors15 secondsScience problems
Ask a second AIPose the same question to a different modelLow-medium. Two models can share the same wrong reasoning1 minuteA tiebreaker, never a verdict
Ask the same AI “are you sure?”Request self-confirmationVery low. Models often agree with themselves or flip under pressure10 secondsNothing. Teach kids this is not verification

That bottom row is the one worth making a point of. When a model says “you’re right, I apologize, the answer is actually 12,” it hasn’t verified anything; it has responded to social pressure. Teaching a kid that “are you sure?” is not a check is one of the more durable things you can hand them.

How to Teach Your Kid About Verifying AI Math

Ages 5–8: The estimate game

Before any calculation, guess. “About how many?” Then compute. Then compare. Do it with grocery totals, minutes until dinner, how many blocks in the tower. You’re building the instinct that an answer should feel about the right size, which is the single most useful error-catching reflex a person can have. When they eventually use a calculator or a chatbot, the reflex fires automatically and catches the decimal-point disasters.

Ages 9–12: The substitution habit

Take any equation-based homework problem. After getting an answer, from anywhere, plug it back in and show that both sides match. Do it out loud: “if x is 5, then 3 times 5 is 15, plus 7 is 22. Checks.” Ten reps and it becomes automatic. Then introduce the twist: deliberately give them a wrong answer and let them catch it by substituting. The feeling of catching an error is what makes the habit stick.

Ages 13+: The 50% experiment

This one lands hard with teenagers. Have them ask an AI the same non-trivial math question in five separate fresh chats, then compare the five answers. On genuinely hard problems, they will not all agree. Log the results in a table: question, five answers, which matched, which was right. That is Bubeck’s finding, reproduced at your kitchen table with a free account, and it’s more convincing than any lecture about AI limitations.

The question to ask: “You got an answer. What would have to be true for it to be wrong?”

The weekly twenty-minute ritual

Same slot every week. Twenty minutes. Four steps.

Minutes 0–5: Pick one problem from the week. Their choice, from actual homework, ideally one they used AI on. No judgment about the use, that’s not what this is.

Minutes 5–10: Verify it with a named method. They pick from the table and say which one they’re using out loud. Substitute back, estimate, units, or solve a second way. The naming matters: it turns a vague “check your work” into a specific procedure.

Minutes 10–15: Find one error, anywhere. In their work, in the AI’s work, in the textbook. There’s almost always one. If the week produced no errors, take a harder problem next week. This step is what keeps the ritual from becoming a compliance exercise.

Minutes 15–20: One retrieval question, device down. Not about this problem. About the concept. “What’s the rule for when you can divide both sides?” Roediger and Karpicke (2006, Psychological Science) found repeated testing produced “substantially greater retention than studying,” and that students felt more confident after restudying even though testing worked better. Five minutes of retrieval is the best-supported study activity in the whole ritual.

Make it about the machine, not about them

The framing that keeps kids engaged is “let’s see if we can catch it.” Verification aimed at the AI feels like a game. Verification aimed at the kid feels like an audit. Same activity, completely different participation rate.

Use the oral question as the close

The University of Chicago Law School’s July 9, 2026 AI strategy replaced detection software with mandatory oral defenses, where students answer questions probing “the paper’s reasoning and argument implications.” The two-minute version, “walk me through why this step is allowed,” is the most reliable check that exists and it requires no tools.

What not to do

Don’t use this as a cheating investigation. The moment the ritual becomes an interrogation about whether they used AI, it’s over. The Wake County case is instructive on the cost of accusation-first approaches: a Green Hope High School freshman was accused of AI use after a teacher ran her work through three detection tools returning 62%, 75%, and 87%; another teacher reviewed the version history and confirmed she hadn’t. Weber-Wulff and colleagues (2023) tested 14 detection tools and found them “neither accurate nor reliable.”

What to Watch For Over the Next 3 Months

  • Week 4: Does your kid name the verification method before using it? Naming is the signal it’s becoming procedural rather than vague.
  • Month 2 red flags: Verification has collapsed into asking the AI “are you sure?” That’s the default failure mode, and it needs a direct correction.
  • Month 3 self-check: Hand them a wrong AI answer without telling them it’s wrong. Do they catch it? That’s the whole skill in one test.

Frequently Asked Questions

How often is AI wrong at math?

It depends entirely on difficulty. On routine arithmetic and algebra, modern models are usually right. At the research frontier, OpenAI researcher Sébastien Bubeck said the model that disproved an Erdős conjecture produced the correct solution in “50 percent” of repeated trials on the same prompt.

What’s the fastest way to check an AI math answer?

Substitute it back into the original problem. If the answer satisfies the equation or condition, it works; if not, it doesn’t. Thirty seconds, and it requires no additional tools or trust. Estimating the order of magnitude first is a close second.

Is asking the AI “are you sure?” a real check?

No, and this is worth teaching explicitly. Models often either agree with themselves or reverse under social pressure, neither of which is verification. Substituting back, checking units, or solving a second way are actual checks.

Can I just use a different AI to check the first one?

It’s a weak tiebreaker, not a verdict. Two models can share the same flawed reasoning, especially on problems where the common approach is wrong. Use it to flag disagreement, then verify properly.

Didn’t AI already solve a bunch of famous math problems?

That claim was retracted. In October 2025 an OpenAI executive said GPT-5 had solved 10 previously unsolved Erdős problems; mathematician Thomas Bloom called it “a dramatic misrepresentation” because the solutions already existed in the literature, and the post was deleted. The May 2026 unit-distance result was real.

Does verification take too long for daily homework?

The weekly ritual is twenty minutes total, on one problem. Daily verification should be the thirty-second version: substitute back, or estimate and compare. The point isn’t to verify everything; it’s to make verification automatic enough that it happens when it matters.


About the author

Ricky Flores is the founder of HiWave Makers and an electrical engineer with 15+ years of experience building consumer technology at Apple, Samsung, and Texas Instruments. He writes about how kids learn to build, think, and create in a tech-saturated world. Read more at hiwavemakers.com.


Sources

  1. Science News. (June 8, 2026). “AI guardrails and the Erdős math problem.” https://www.sciencenews.org/article/ai-guardrails-erdos-math-problem
  2. TechCrunch. (May 20, 2026). “OpenAI claims it solved an 80-year-old math problem, for real this time.” https://techcrunch.com/2026/05/20/openai-claims-it-solved-an-80-year-old-math-problem-for-real-this-time/
  3. OpenAI. (May 20, 2026). “A model disproves a discrete geometry conjecture.” https://openai.com/index/model-disproves-discrete-geometry-conjecture/
  4. Quanta Magazine. (Aug 3, 2026). “Why the legendary Erdős problems are falling to AI.” https://www.quantamagazine.org/why-the-legendary-erdos-problems-are-falling-to-ai-20260803/
  5. Roediger, H. L., III, & Karpicke, J. D. (2006). “Test-enhanced learning: Taking memory tests improves long-term retention.” Psychological Science, 17(3). https://doi.org/10.1111/j.1467-9280.2006.01693.x
  6. Weber-Wulff, D., Anohina-Naumeca, A., Bjelobaba, S., et al. (2023). “Testing of detection tools for AI-generated text.” International Journal for Educational Integrity, 19(26). https://doi.org/10.1007/s40979-023-00146-z
  7. WRAL. (May 5, 2026). “Wake County student says clear AI policies needed after being accused of cheating.” https://www.wral.com/news/education/wake-county-student-says-ai-policies-needed-after-cheating-accusation-may-2026/
  8. University of Chicago Law School. (July 9, 2026). “Rethinking Legal Education in the AI Era.” https://www.law.uchicago.edu/news/ai-strategy-statement
  9. Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). “Generative AI without guardrails can harm learning.” PNAS, 122(26). https://doi.org/10.1073/pnas.2422633122

Related reading: AI’s 50% math success rate and teaching kids reliability, how mathematicians verify an AI proof, and study mode vs. answer mode for kids.

Ricky Flores
Written by Ricky Flores

Founder of HiWave Makers and electrical engineer with 15+ years working on projects with Apple, Samsung, Texas Instruments, and other Fortune 500 companies. He writes about how kids learn to build, think, and create in a tech-driven world.