Table of Contents
When AI Is Wrong: Hallucinations, Bias, and What Kids Need to Know
AI hallucinations and bias aren't random glitches — they're predictable consequences of how these systems are built. Here's what parents and kids need to understand.
A seventh-grader cited a legal case in a school project that didn’t exist. She’d asked ChatGPT to find three landmark cases in environmental law, and it produced plausible-sounding citations — case names, years, courts — that were entirely fabricated. Her teacher caught it. But it raised a question the teacher hadn’t dealt with before: how do you teach a kid to use a tool that confidently makes things up?
The answer starts with understanding why it makes things up. Not “AI is stupid” or “AI is dangerous” — those framings don’t help. Understanding the mechanism gives kids the mental model they need to use these tools critically.
Key Takeaways
- AI hallucination happens because language models predict probable text, not factual statements — they can generate plausible-sounding falsehoods with complete confidence
- AI bias enters through training data (internet text reflects human prejudice), labeling (human raters have biases), and historical data used to train decision systems
- Hallucination rates vary by domain: medical and legal topics show particularly high error rates in current research
- Teaching kids to verify AI claims against primary sources is the key protective skill
- Different ages can practice different verification strategies — from simple cross-referencing to systematic source-checking
What Hallucination Actually Is
The word “hallucination” is borrowed from psychiatry, but the mechanism in AI is different from the human version. It’s not a perceptual error or a distorted reality. It’s a structural feature of how large language models produce text.
A language model generates tokens (word fragments) by predicting what word is most likely to follow the current sequence of words. It was trained on a massive corpus of human-written text, and it learned that certain kinds of sentences tend to follow certain other kinds of sentences. When asked for citations to legal cases, it generates text that looks like a citation — because citations have a consistent format and appear near discussions of legal topics in training data. Whether that specific citation corresponds to a real case is not something the model checks. It can’t check. It has no access to a truth database.
This is why hallucination is not “lying.” The model isn’t trying to deceive. It’s not aware that it’s wrong. It’s doing exactly what it was trained to do — generating probable text — and the probable text in the context of “list three landmark environmental cases” is text that looks like named cases, even if the specific names don’t correspond to real ones.
What Research Shows About Hallucination Rates
A 2023 benchmark study from Stanford’s HAI group evaluated 7 major LLMs on medical question answering and found hallucination rates ranging from 5% to 37% depending on the model and question type (Singhal et al., 2023). In domains with thin coverage in training data, rates go up.
Research specifically on legal AI found that when lawyers used AI tools to find case citations, the tools produced fabricated cases at rates that — in one documented lawsuit — led to sanctions from a federal judge (Mata v. Avianca, 2023). The lawyers cited cases that their AI had invented.
The rate varies by topic:
| Topic Domain | Approximate Hallucination Rate | Why It’s Higher/Lower |
|---|---|---|
| General knowledge (well-covered) | Low (3–10%) | High-frequency topic in training data |
| Medical (specific clinical details) | Medium-High (15–37%) | Nuanced facts not well-distributed in web text |
| Legal (specific case citations) | High (25–50%+) | Training data contains legal discussions, not full case databases |
| Current events (post-training cutoff) | Very High (50%+) | Model has no data about recent events |
| Creative writing (no “correct” answer) | Low (N/A) | No ground truth to hallucinate against |
Note: These are approximate ranges synthesized from multiple published studies. Individual model performance varies significantly.
What AI Bias Is (And Where It Enters)
Bias in AI is different from hallucination. It’s a systematic pattern of errors that correlates with identifiable characteristics — race, gender, age, disability status. It enters through several mechanisms.
Training Data Bias
Language models trained on internet text inherit the biases present in that text. If historical internet content associates certain professions predominantly with one gender (engineering = men, nursing = women), the model learns these associations. When asked to generate “describe an engineer,” it may produce a male-coded description by default.
A foundational 2016 paper by Bolukbasi et al. at Boston University demonstrated that word embedding models (a precursor to modern LLMs) had learned strong gender associations — “man is to doctor as woman is to nurse” — directly from training data (Bolukbasi et al., 2016).
Algorithmic Amplification
AI systems don’t just inherit biases — they can amplify them. If a résumé screening AI was trained on historical hiring data from a company that historically hired more men in technical roles, the AI learns that “successful candidates look like previous hires” — and perpetuates the same pattern. Amazon famously abandoned an internal hiring AI in 2018 after discovering it had learned to penalize résumés that included the word “women’s” (as in women’s chess club), because men had been disproportionately hired in the roles it was trained on (Dastin, 2018).
Healthcare and Criminal Justice
These two domains have the most documented cases of AI bias with real-world consequences.
In criminal justice, the COMPAS recidivism algorithm was analyzed in a 2016 ProPublica investigation that found the algorithm predicted higher recidivism risk for Black defendants at nearly twice the rate for white defendants who went on to commit no new crimes (Angwin et al., 2016). The algorithm’s designers contested the interpretation, and this case remains the subject of ongoing academic debate about how to define “fairness” mathematically — but the disparity in outcomes was real.
In healthcare, a 2019 study in Science found that a widely-used healthcare algorithm assigned lower risk scores to Black patients than to equally sick white patients, because it used historical healthcare spending as a proxy for healthcare need — and Black patients had historically received less care (Obermeyer et al., 2019).
How to Verify AI Claims: A Practical Framework
For kids who use AI tools, the skill is not “never trust AI” — that’s impractical. The skill is knowing when and how to verify.
High-verification topics (always verify):
- Specific facts: dates, names, statistics, quotes
- Medical or legal information
- Historical events or their interpretations
- Scientific claims, especially specific study results
- Anything your kid will cite in a school assignment
Lower-verification topics (AI is generally reliable):
- Writing assistance, grammar, rephrasing
- Explaining concepts you already understand well enough to evaluate
- Brainstorming where you’ll filter the results
- Generating creative content with no factual claims
The verification approach by age:
| Age Group | Verification Skill to Teach | Concrete Practice |
|---|---|---|
| 8–10 | Cross-reference with a second source | ”Google that fact to see if another website says the same thing” |
| 11–13 | Find the original source | ”Find the actual study, article, or law the AI is describing” |
| 14+ | Primary source evaluation | ”Who published this? Is it peer-reviewed? Can I read the abstract?” |
For a comprehensive look at how AI tools can assist kids’ learning while requiring critical evaluation, see our guide to AI tools for kids’ education in 2026.
How to Teach Your Kid About AI Errors
Ages 5–8: The “Sometimes Wrong” Story
Tell your child that even the smartest computers can be wrong sometimes — not because they’re bad, but because they’re guessing, like when you guess what word comes next in a sentence. Play the guessing game: “The dog ran toward the ___” and let them guess. Sometimes they’re right, sometimes not. Explain: AI does the same kind of guessing with words. That’s why grownups check important things.
Ages 9–12: The Fact-Check Hunt
Give your child a task: ask an AI for five facts about any topic they’re curious about. Then fact-check each one using Wikipedia or a trusted website. Keep a scorecard: correct, partially correct, wrong, can’t verify. Discuss the results. This is more effective than telling a child that AI can be wrong — experiencing it firsthand builds the skepticism.
Ages 13+: Bias Audit
Pick an AI image generation tool (Midjourney, DALL-E). Ask it to generate images of “a doctor,” “a nurse,” “a CEO,” “a criminal suspect” — without specifying any characteristics. Analyze the results: what patterns emerge in the generated images? Are the doctors predominantly one gender or race? This makes AI bias visible and discussable without requiring a deep dive into technical mechanisms.
The question to ask: “If the AI learned from the internet — and the internet was written by humans who have biases — what kinds of mistakes would you expect the AI to make?”
What to Watch For Over the Next 3 Months
Month 1: The next time your kid uses AI for a school project, ask them: “Did you check that? How?” The goal is to make verification a default habit, not an unusual step.
Month 2: Find and read one real news story about AI error together. The New York Times, MIT Technology Review, and Wired all cover documented AI failures. Understanding that these aren’t hypothetical makes the risk concrete.
Month 3: Introduce the concept of “high-stakes vs. low-stakes” AI use. Low stakes: brainstorming, entertainment, casual writing. High stakes: medical decisions, legal research, journalism, any citation. Kids who can sort tasks into these categories are meaningfully more capable AI users.
Frequently Asked Questions
Can AI hallucinations be fixed completely?
Not entirely, with current architectures. Retrieval-augmented generation (RAG) — where the model is given access to a specific database of verified documents — significantly reduces hallucination on the topics covered by that database. But general hallucination is structural to how language models work and can only be reduced, not eliminated.
Is AI bias getting better?
Yes, meaningfully. Most major AI companies now have fairness testing protocols. Models are evaluated on demographic parity before deployment. But improvement is uneven — some biases are easier to measure and correct than others. And as AI is deployed in higher-stakes contexts (medical, criminal justice, hiring), the bar for “good enough” rises.
Should I be worried about my kid getting medical information from AI?
Yes, selectively. AI can explain medical concepts well and is often more accessible than medical textbooks. But for specific symptoms, diagnoses, medication dosages, or treatment decisions, the hallucination rate and potential for confident misinformation is high enough that AI should not be the primary or only source. Peer-reviewed sources, CDC.gov, and actual medical consultation are the standard.
How do I teach my kid to fact-check without making it feel like punishment?
Frame it as a skill, not a punishment. “You found something interesting — let’s find out if it’s actually true, because that would make it even more interesting.” Kids who discover a factual error are often more engaged with the topic afterward, not less. The discovery of error is itself an educational event.
About the author Ricky Flores is the founder of HiWave Makers and an electrical engineer with 15+ years of experience building consumer technology at Apple, Samsung, and Texas Instruments. He writes about how kids learn to build, think, and create in a tech-saturated world. Read more at hiwavemakers.com.
Sources
- Singhal, K., Azizi, S., Tu, T., et al. (2023). “Large Language Models Encode Clinical Knowledge.” Nature, 620, pp. 172–180. https://doi.org/10.1038/s41586-023-06291-2
- Bolukbasi, T., Chang, K. W., Zou, J., et al. (2016). “Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings.” NIPS 2016. https://arxiv.org/abs/1607.06520
- Angwin, J., Larson, J., Mattu, S., & Kirchner, L. (2016). “Machine Bias.” ProPublica. https://www.propublica.org/article/machine-bias-risk-assessments-in-criminal-sentencing
- Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). “Dissecting Racial Bias in an Algorithm Used to Manage the Health of Populations.” Science, 366(6464), pp. 447–453. https://doi.org/10.1126/science.aax2342
- Dastin, J. (2018, October 10). “Amazon scraps secret AI recruiting tool that showed bias against women.” Reuters. https://www.reuters.com/article/us-amazon-com-jobs-automation-insight-idUSKCN1MK08G
- Mata v. Avianca, Inc. (2023). United States District Court, Southern District of New York. Case No. 22-cv-1461. https://www.courtlistener.com/docket/66732006/mata-v-avianca-inc/