Evaluating AI Safety Claims: A Checklist for Parents
Table of Contents

Evaluating AI Safety Claims: A Checklist for Parents

Evaluating AI safety claims takes about fifteen minutes if you know the six documents to look for. A checkable rubric, with what green and red look like.

Every AI company says it takes safety seriously. That sentence carries no information, which means evaluating AI safety claims has to be done on something else: the documents a company publishes, the thresholds it commits to in writing, and who besides the company itself has tested its models. All six of those things are public. Checking them takes about fifteen minutes, needs no technical background, and will tell you more than any press release. Here is the rubric I use, and what green and red actually look like.

Key Takeaways

  • A serious safety programme leaves paper. Six public artefacts tell you most of what an outsider can know, and a company missing four of them is telling you something.
  • The strongest single signal is a published threshold. The Seoul Frontier AI Safety Commitments ask signatories to “set out thresholds at which severe risks posed by a model or system, unless adequately mitigated, would be deemed intolerable.”
  • Independent scoring exists and is unflattering across the board. The Future of Life Institute graded seven developers in July 2025; the highest was a C+.
  • The useful question about a system card is not whether it exists but whether its known-limitations section got longer as the model got more capable.
  • Specific promises beat general ones by a wide margin. “Nothing publishes, sends, or spends without your approval” is checkable. “We prioritise user safety” is not.

Why the obvious test does not work

The intuitive approach is to read what a company says about safety. It fails for a structural reason: safety language is free to produce and carries no penalty when it turns out to be aspirational. Every company can write that it is committed to responsible development, and every company does.

Stanford’s 2025 AI Index Report puts the gap plainly: “AI-related incidents are rising sharply, yet standardized RAI evaluations remain rare among major industrial model developers,” and among companies “a gap persists between recognizing RAI risks and taking meaningful action.” Recognising a risk in a blog post and acting on it are different activities that produce different evidence.

So the test has to be about artefacts: things that exist or do not, that are dated, and that a company would find awkward to contradict later.

The six artefacts: evaluating AI safety claims in practice

1. A safety framework with actual thresholds

The distinguishing feature of a serious framework is a number, or at least a stated condition, at which the company says it will stop.

The benchmark here was set by the Frontier AI Safety Commitments, agreed at the Seoul summit in May 2024 by twenty organisations including Amazon, Anthropic, Cohere, Google, IBM, Meta, Microsoft, Mistral AI, NVIDIA, OpenAI, Samsung Electronics, xAI and Zhipu.ai. Signatories commit to “set out thresholds at which severe risks posed by a model or system, unless adequately mitigated, would be deemed intolerable,” and, crucially, to “not develop or deploy a model or system at all, if mitigations cannot be applied to keep risks below the thresholds.”

That second clause is the one to look for, because it is a commitment to forgo revenue under a stated condition. A framework that describes processes without ever naming a stopping point is a description of work, not a constraint on it.

Check it in two minutes: search the company’s site for “frontier safety framework,” “preparedness framework” or “responsible scaling policy.” Read only the section that names thresholds. If there are no thresholds, you have your answer.

2. System cards whose limitations section is growing

A system card is the document published alongside a model release describing what it was tested for, what it failed, and what remains unresolved. Our guide to reading a model card with your teen covers the structure.

Everyone publishes these now, so existence is no longer a signal. The signal is the trend. Compare the known-limitations section of the current release against the previous one. A more capable model should have a longer, more specific list of things it does badly, because better evaluation finds more failures. If the model got more powerful and the limitations list got shorter or vaguer, something changed in the writing process rather than in the model.

This is also the specific concern a departing OpenAI safety lead raised in October 2026, which we covered in the resignation that called the culture broken. When the person who writes the document says the process producing it is compromised, the document is still worth reading, with a thumb on the scale.

3. Third-party evaluation, published by the third party

“We conducted extensive internal testing” is the weakest form of this claim. “We commissioned an external red team” is better. “Here is the external team’s own report, which we did not edit” is the strong form.

The question to ask is not whether outside evaluation happened but whether the outside party published independently. A company can commission an evaluation and publish only a summary of it, which is marketing with extra steps.

4. Independent comparative scoring

Someone else has already done comparative work, and it is more useful than any single company’s self-description.

The Future of Life Institute’s AI Safety Index: Summer 2025, published July 17, 2025, graded seven developers across six domains: risk assessment, current harms, safety frameworks, existential safety, governance and accountability, and information sharing. Anthropic scored highest with a C+ (2.64), followed by OpenAI at C (2.10), Google DeepMind at C- (1.76), xAI at D (1.23), Meta at D (1.06), Zhipu AI at F (0.62) and DeepSeek at F (0.37). The panel’s summary judgment was blunt: “the industry is fundamentally unprepared for its own stated goals.”

Two honest notes. These are expert-panel grades, not measurements, and the panel’s methodology embeds a view about which risks matter. And the distribution matters more than the ranking: when nobody scores above C+, “which company is safest” is a less useful question than “what is missing across all of them.”

5. Incident disclosure you can find without the company’s help

A company that reports its own failures is doing something costly, which makes it informative. But you do not have to rely on self-reporting. The AI Incident Database, run by the Responsible AI Collaborative, indexes “the collective history of harms or near harms realized in the real world by the deployment of artificial intelligence systems.”

Search it for the product your family uses. Then check whether the company acknowledged the incidents you find. The gap between what is publicly documented and what the company has discussed is a direct measure of transparency, and it requires no insider access to compute.

6. Promises specific enough to fail

This is the test I would keep if I had to discard the other five.

Compare two real sentences. Meta, describing its personal AI agent in September 2026: “You’re in control: nothing publishes, sends, or spends without your approval.” And the generic form found on nearly every AI product page: some variation of “we are committed to building AI safely and responsibly.”

The first can be falsified by one counterexample. The second cannot be falsified at all. A company willing to write sentences that could embarrass it later is accepting a cost, and accepting costs is the only evidence of priority that exists.

SignalWhere to lookGreenRed
Safety frameworkSite search for “safety framework” or “scaling policy”Names thresholds and a condition for not deployingDescribes process only; no stopping point
System card trendCompare current and previous releasesLimitations section longer and more specificShorter or vaguer than last version
Third-party evaluationWho published the reportExternal team published its own findingsCompany summarised someone else’s work
Independent scoringFLI AI Safety Index and similarGraded, and the grade is discussed publiclyAbsent from comparative indexes
Incident disclosureAI Incident Database, plus company blogKnown incidents acknowledged with datesDocumented incidents never mentioned
Promise specificityProduct page and help docsSentences that one counterexample would breakOnly unfalsifiable commitments

The vocabulary that makes this easier

Two public documents give you the words to ask better questions.

The NIST AI Risk Management Framework, released January 26, 2023, organises risk work into four functions: govern, map, measure and manage. Its Generative AI Profile, NIST AI 600-1, followed on July 26, 2024. The framework is voluntary, which is its limitation, but it gives you a structure: a company that can describe what it does under all four functions is operating a programme, and one that can only talk about “measure” is running tests.

The EU AI Act implementation timeline tells you what stops being optional and when. Prohibitions and AI-literacy duties began applying on 2 February 2025. General-purpose AI model rules started on 2 August 2025, with existing models given until 2 August 2027. High-risk obligations under Annex III begin on 2 December 2027. A company already building to those requirements is making a different bet than one waiting for enforcement.

And regulators are asking their own questions. The US Federal Trade Commission’s 6(b) inquiry into AI chatbot companions, opened September 11, 2025, sent orders to Alphabet, Character Technologies, Instagram, Meta, OpenAI, Snap and X.AI asking, among other things, what steps they had taken “to evaluate the safety of their chatbots when acting as companions,” and how they “measure and monitor potential harms to minors before and after product launch.” Those are the right questions, and the answers will eventually be partly public.

What to do with the answer

Run the rubric once on the product you already use

Not on the most newsworthy company. On the assistant actually installed on your child’s phone. Fifteen minutes, six checks, and write the date on your notes.

Keep a dated file of specific promises

When you find a falsifiable sentence, save it with the date and the URL. Companies revise these pages quietly, and a dated quote is the only way to notice. Three dated quotes make you better informed than almost any other parent in the room at a school meeting. The reasoning behind this habit is in our piece on the difference between a guardrail and a promise.

Weight configuration over brand

Since no developer scores above a C+ on independent assessment, switching providers to escape poor safety governance has no clear destination. What you control is the configuration: age settings, permissions, payment access, logging. Those produce more safety per minute spent than brand selection does.

Ask the school the same six questions

When a district adopts an AI tool, the procurement conversation almost never includes any of this. A parent who asks “does the vendor publish a safety framework with thresholds, and has anyone here read its system card” changes the tenor of a meeting, usually productively.

What not to do

Do not treat a safety incident as proof of negligence, or its absence as proof of care. A company with a public incident may simply have better disclosure than a competitor with the same problems and a quieter blog. Compare disclosure practices, not incident counts.

What to Watch For Over the Next 3 Months

  • Week 4: Watch for new system cards from any model release. Read only the limitations section and compare against the previous version. Length and specificity are the whole signal.
  • Month 2 red flags: Watch for a company quietly removing a specific commitment from a help page. This is the most informative event available to an outsider and almost nobody notices it, because noticing requires having saved the old wording.
  • Month 3 self-check: Re-run the six checks on your primary provider. Did anything move from green to red, or the reverse? Three months is roughly the cadence at which these documents change.

Frequently Asked Questions

What is the fastest single check on whether a company is serious about AI safety?

Look for a published threshold: a stated condition under which the company says it will not deploy a model. The Seoul Frontier AI Safety Commitments asked signatories to set out thresholds at which severe risks would be deemed intolerable and to decline deployment if mitigations cannot keep risks below them. Frameworks without a stopping point describe work rather than constrain it.

Which AI company is safest for my kid?

Independent assessment does not support a clear answer. The Future of Life Institute’s Summer 2025 index graded seven developers, with the top score a C+ and two Fs. The practical implication is that configuration in your own house matters more than brand choice.

Are safety frameworks legally binding?

Generally no. The NIST AI Risk Management Framework is explicitly voluntary, and the Seoul commitments are commitments rather than law. The EU AI Act is binding within its scope on a published timeline, with general-purpose AI model rules applying from 2 August 2025 and high-risk obligations under Annex III from 2 December 2027.

How do I read a system card without a technical background?

Skip the benchmarks and read two sections: known limitations, and anything labelled unresolved or residual risk. Those are written in plainer language than the rest, and they are the only parts where the company is describing what it could not fix.

Does a public incident mean a company is careless?

Not by itself. Disclosure and incidence are different things, and a company with visible incidents may simply report more than a quieter competitor with the same problems. Compare how companies handle incidents rather than how many appear in the news.

What should I ask my child’s school before it adopts an AI tool?

Three questions cover most of it: does the vendor publish a safety framework with thresholds, has anyone at the school read the system card for the model being used, and what happens to student data. The third one is where most procurement conversations discover a problem.


About the author

Ricky Flores is the founder of HiWave Makers and an electrical engineer with 15+ years of experience building consumer technology at Apple, Samsung, and Texas Instruments. He writes about how kids learn to build, think, and create in a tech-saturated world. Read more at hiwavemakers.com.


Sources

  1. UK Department for Science, Innovation and Technology. (2024). “Frontier AI Safety Commitments, AI Seoul Summit 2024.” May 21, 2024. https://www.gov.uk/government/publications/frontier-ai-safety-commitments-ai-seoul-summit-2024
  2. Future of Life Institute. (2025). “AI Safety Index: Summer 2025.” July 17, 2025. https://futureoflife.org/ai-safety-index-summer-2025/
  3. Stanford Institute for Human-Centered AI. (2025). “The 2025 AI Index Report.” https://hai.stanford.edu/ai-index/2025-ai-index-report
  4. National Institute of Standards and Technology. (2023). “AI Risk Management Framework,” and the Generative AI Profile, NIST AI 600-1 (2024). https://www.nist.gov/itl/ai-risk-management-framework
  5. Responsible AI Collaborative. “AI Incident Database.” https://incidentdatabase.ai/
  6. Federal Trade Commission. (2025). “FTC Launches Inquiry into AI Chatbots Acting as Companions.” September 11, 2025. https://www.ftc.gov/news-events/news/press-releases/2025/09/ftc-launches-inquiry-ai-chatbots-acting-companions
  7. Future of Life Institute. “EU AI Act Implementation Timeline.” https://artificialintelligenceact.eu/implementation-timeline/
  8. Ha, Anthony. (2026). “OpenAI safety employee resigns, claiming the company’s ‘culture is broken’.” TechCrunch, October 3, 2026. https://techcrunch.com/2026/10/03/openai-safety-employee-resigns-claiming-the-companys-culture-is-broken/
Ricky Flores
Written by Ricky Flores

Founder of HiWave Makers and electrical engineer with 15+ years working on projects with Apple, Samsung, Texas Instruments, and other Fortune 500 companies. He writes about how kids learn to build, think, and create in a tech-driven world.