Standardized Testing Hasn't Made US Kids Smarter. 40 Years of Data.
Table of Contents

Standardized Testing Hasn't Made US Kids Smarter. 40 Years of Data.

40 years of NAEP data show standardized testing US kids hasn't worked. NCLB, Race to the Top, ESSA—what each promised vs. what scores actually show.

Standardized Testing Hasn’t Made US Kids Smarter. 40 Years of Data Says So.

Picture this scene: a fourth-grader comes home in October with a manila folder full of “test prep packets.” Her teacher explained that the state reading benchmark is eight weeks away. No new books this month—just practice passages and bubble sheets. The girl’s mother, a former teacher herself, emails the principal. The principal replies that the school’s Title I funding depends on the scores.

That’s not an edge case. That’s Tuesday in thousands of American schools. And the frustrating part isn’t that the tests are hard. It’s that four decades of evidence suggest they aren’t working—and the people who design education policy know it.

The US has been running a large, expensive, and increasingly coercive experiment in test-based accountability since the early 1980s. The results are in. If you want to understand why your kid comes home with test prep instead of a novel, this is where the data trail leads.

The Promise: What Standardized Tests Were Supposed to Fix

The modern era of standardized accountability begins with A Nation at Risk (1983), the federal report that declared American schools were drowning in “a rising tide of mediocrity.” The proposed cure: measure student performance consistently, hold schools accountable for results, and improvement would follow.

The logic had surface appeal. If you can’t measure it, you can’t fix it. Schools with low scores would identify weak instruction. Teachers would respond. Students would learn more.

Each subsequent reform used the same logic, escalated:

  • No Child Left Behind (2001): Required annual reading and math tests in grades 3–8 plus one high school grade. Threatened schools with sanctions—restructuring, staff replacement, closure—if they missed “Adequate Yearly Progress” targets. The goal was 100% proficiency in math and reading by 2014.
  • Race to the Top (2009): Competitive grants dangled before states that adopted Common Core standards, expanded charter schools, and tied teacher evaluations to student test scores.
  • Every Student Succeeds Act (2015): Replaced NCLB, kept annual testing, handed states more flexibility in how they used scores—while keeping the accountability architecture largely intact.

None of these reforms said “tests are the point.” They said tests are the instrument for driving better teaching. The distinction matters because you can defend the instrument forever by saying the problem is how it’s being used, not whether it works.

But after 40 years, the instrument has a track record. Let’s look at it.

NAEP “Nation’s Report Card”: What 40 Years of Scores Actually Show

The National Assessment of Educational Progress is the closest thing the US has to an honest, nationally consistent measure of student learning. Unlike state tests—which states can and do make easier to inflate scores—the NAEP is administered by a federally independent board, uses the same framework across all states, and has long-term trend data going back to the 1970s.

The NAEP Long-Term Trend Assessment (LTTF) tests 9-year-olds and 13-year-olds in reading and math. It’s the one national measure that isn’t politicized by individual reform cycles, which is exactly why it’s the right lens here.

What do 40 years of NAEP scores show?

The good news first: 9-year-olds made meaningful gains in both reading and math from the 1970s through the early 1990s—before NCLB existed. Between 1971 and 1990, average math scores for 9-year-olds rose about 11 points on the NAEP scale. Reading scores for the same group rose about 12 points.

Now the hard part: Those gains largely leveled off. The NAEP 2022 Long-Term Trend report—released after the COVID disruption—shows that 9-year-old math scores dropped to their lowest level since 1999. Reading scores dropped to their lowest level since 1990. Decades of test-driven accountability didn’t protect those gains. It didn’t produce new ones.

For 13-year-olds, the picture is worse. Math scores in 2022 were lower than in 2012. Reading scores were lower than in 2008. The NCLB era—the most intensive period of standardized-test-based accountability in US history—produced no statistically significant long-term improvement in this age group.

Education researcher Diane Ravitch, who served as Assistant Secretary of Education under George H.W. Bush and initially supported NCLB, reversed her position after reviewing the data. In The Death and Life of the Great American School System (2010), she documented how NCLB’s pressure created incentives for states to lower their proficiency standards rather than raise actual learning—a dynamic she called “lying to children.”

NCLB, Race to the Top, ESSA: What Each Did to Scores

Here’s the NAEP data at key reform milestones. Scores are average scale scores (0–500 for long-term trend; NAEP 4th/8th grade main assessments use a 0–500 scale):

Reform Era4th Grade Reading4th Grade Math8th Grade Reading8th Grade Math
Pre-NCLB baseline (2000)213226263273
End of NCLB era (2009)220240263282
Race to the Top peak (2013)221241267285
Post-ESSA (2019, pre-COVID)220241263282
2022 (post-COVID)217236260274

Sources: NAEP Main Assessments 2000–2022, National Center for Education Statistics.

The pattern: modest gains from 2000–2009, stagnation from 2009–2019, and COVID-era losses that erased much of the NCLB-era progress. NCLB showed real, if modest, gains in 4th grade math—a finding that accountability advocates cite. But 8th grade reading didn’t move in a decade of NCLB pressure. And the post-2009 plateau under Race to the Top and ESSA is hard to spin as success.

Researcher John Jennings, in a 2015 review published by Education Week, analyzed NCLB’s decade-long record and concluded that while early elementary gains were real, the law “never came close to its own goal of 100% proficiency,” and that the pressure had measurable negative effects on curriculum breadth. The gains that did occur may be partially explained by simultaneous investments in early childhood education—Head Start expansion, pre-K programs—rather than the testing regime itself.

Eric Hanushek at Stanford’s Hoover Institution, one of the more rigorous economists of education, has argued that accountability systems in principle can work—but that the US implementation failed to connect consequences to actual learning rather than score manipulation. His analysis of state NAEP data shows that states which lowered proficiency thresholds under NCLB showed better paper results but no NAEP improvement.

The Narrowing Curriculum: What Gets Cut When Tests Drive Teaching

Scores aren’t the only cost. There’s a second harm that doesn’t show up in NAEP tables: what disappeared from the school day to make room for test prep.

A 2007 report from the Center on Education Policy found that 62% of US school districts had increased time spent on math and reading since NCLB—by an average of 141 minutes per week in reading and 89 minutes per week in math. The time came from somewhere. Social studies instruction dropped by 76 minutes per week in many districts. Science, art, music, and physical education were cut in schools where scores were lowest and pressure was highest.

This matters for reasons beyond the obvious. The subjects that disappear under testing pressure—science, social studies, arts, project-based work—are the subjects where higher-order thinking develops. Reading comprehension improves when kids have broad domain knowledge to attach words to. Cutting social studies and science to prep for a reading test is self-defeating over time. But it produces short-term score bumps, which is what the accountability clock rewards.

Linda Darling-Hammond, who has studied accountability systems across countries for Stanford’s Learning Policy Institute, has documented how countries with strong outcomes—Finland, Canada, Singapore—use standardized assessments sparingly, primarily as diagnostic tools rather than as high-stakes accountability instruments. The tests inform instruction; they don’t drive it.

For a broader look at how US education spending compares to outcomes internationally, our analysis of the PISA paradox puts these NAEP numbers in a global frame.

What High-Performing Countries Use Instead (And Why)

The countries that consistently outperform the US on international assessments—Finland, Canada, Singapore, Estonia—share a counterintuitive feature: they test less, not more.

Finland administers no standardized tests until students are 16. National assessments are sample-based—they test a representative sample of students to get a system-level picture, rather than testing every child every year. Teachers use that data to adjust their own practice. Nobody’s job is tied to the results.

Canada’s PISA performance is driven largely by provinces like Alberta and Ontario, which invest heavily in teacher training and curriculum quality. Ontario’s turnaround in the early 2000s—under education researcher Michael Fullan’s framework—used data diagnostically, not punitively. Principals reviewed school-level trends to identify where to target professional development. The difference: data as a flashlight, not a weapon.

Singapore tests more than Finland, but the testing is embedded in a coherent curriculum system with highly trained, well-compensated teachers who use results as formative feedback. The high-stakes nature of Singaporean exams is real—but it operates alongside genuine teacher authority and curriculum depth. American high-stakes testing has the stakes without the rest of the architecture.

The pattern Linda Darling-Hammond summarized across multiple studies: accountability systems produce improvement when they build capacity (better teachers, better curriculum, better resources) rather than simply apply pressure. Testing without capacity-building is pressure without support. Schools and teachers respond to that combination predictably—they protect themselves from the consequences, not the students from ignorance.

The PISA 2022 results—covered in more depth in our PISA ranking breakdown for US parents—show the US sitting in the middle of developed nations on math, reading, and science, despite spending among the highest per pupil.

What Parents Can Do When Tests Drive Their Child’s Education

Knowing the research doesn’t change your school’s testing calendar. But it does change how you frame things for your kid and how you advocate within the system.

Separate your child’s self-concept from their score

A third-grader who scores “below proficient” on a state reading benchmark is not a below-proficient reader in any meaningful developmental sense. She may be a child whose school drilled comprehension passages but never read her a chapter book. Tell her the test measures one narrow thing. Her reading is bigger than that.

Find out what’s being cut

Ask your child’s teacher directly: “What subjects or activities have been reduced to make room for test preparation?” Most teachers are honest when parents ask without judgment. If science or social studies has been cut to 45 minutes per week, that’s worth advocating about at the school board level, where those scheduling decisions are made.

Push for diagnostic use, not shame use

Many schools post test scores publicly, rank teachers, and send home letters comparing kids against benchmarks in ways that feel punitive. The research on assessment is clear: feedback that labels students as failing produces less improvement than feedback that identifies specific gaps and offers a path forward (Hattie & Timperley, 2007, Review of Educational Research). Ask your school’s principal: “How do teachers use these scores to change their instruction?”

Supplement with actual learning

Test prep crowds out reading, writing, and thinking—the things that matter beyond third grade. The most protective thing a parent can do is independent of the school: read aloud to kids past the age when schools stop, have them explain things back to you, give them projects that require building and problem-solving. None of that appears on a bubble sheet. All of it compounds.

What to Watch for Over the Next School Year

If your school is in a high-testing environment, watch for these signs that the testing regime is doing concrete harm to your specific kid:

  • By October: Is your child coming home with test prep materials rather than books or projects? If test prep starts before Halloween for a spring exam, that’s a 6-month curriculum replacement.
  • By January: Has your child said anything like “I’m not good at reading” or “I’m bad at math”? These self-concepts form from test feedback. If you hear them, counter them with specifics: what can they read? what can they build? what math shows up in real life?
  • By April (pre-test month): Watch for anxiety symptoms—sleep disruption, stomachaches, reluctance to go to school—that spike in the weeks before a benchmark test. Some stress is normal. Sustained anxiety about an academic test in elementary school is a sign the pressure has been transmitted downward in ways the research doesn’t support.

Frequently Asked Questions

Did NCLB produce any real improvement at all?

Yes, narrowly. NCLB showed genuine gains in 4th grade math scores from 2000–2007—about 5 scale score points on the NAEP main assessment. The gain was real and not explained by demographic shifts alone. But 8th grade reading showed no gain over the same period, the improvements stalled after 2009, and many researchers attribute the 4th grade math gains partly to pre-K investments and early childhood programs that predated NCLB.

Why don’t politicians just admit standardized testing isn’t working?

Because the logic of accountability—test students, publish scores, apply consequences—has strong intuitive appeal and strong political constituencies. Testing companies, tutoring companies, curriculum companies, and accountability advocates all benefit from the current system. Diane Ravitch documented this ecosystem in detail. The data problem is that NAEP scores don’t offer an obvious culprit to punish, so the political incentive is to explain the scores away rather than restructure the system.

My child’s school says their scores are “above the state average.” Should I be reassured?

Only partially. State averages are set by state tests, which states control. Multiple states under NCLB gamed their proficiency thresholds downward to avoid sanctions—meaning “above state average” in some states meant less than it should have. The NAEP is the more honest benchmark; it’s worth asking whether your child’s school’s NAEP scores are available (many district-level data points are public through nces.ed.gov).

Is there any standardized test worth taking seriously?

The NAEP itself is well-designed—it’s not used to rank individual kids, just to track system-level trends. For parents, PISA and TIMSS international assessments are meaningful comparative tools. The problem isn’t that measurement is wrong—it’s that annual high-stakes testing of every child, tied to funding and teacher jobs, creates incentives that distort what gets taught. Diagnostic assessments used formatively by teachers are a different thing from accountability tests used punitively by administrators.

How do I talk to my child’s teacher about test prep concerns without seeming like a problem parent?

Frame it as curiosity, not criticism: “I’ve read that curriculum can get narrowed in testing seasons—what does your class look like in terms of subject variety right now?” Most teachers feel the squeeze acutely and are relieved when a parent acknowledges it rather than demanding higher scores. If you want to advocate more formally, school board meetings, where curriculum policy is actually set, are more effective than individual classroom conversations.

Are private and charter schools exempt from all this?

Mostly no. Federally, charter schools that receive Title I funding must administer state tests. Private schools are exempt from state testing mandates but often administer their own assessments (ISEE, SSAT, ERB) that carry their own pressures. The testing regime is heaviest in under-resourced public schools where federal funding accountability is tightest—which compounds the equity problem.


About the author

Ricky Flores is the founder of HiWave Makers and an electrical engineer with 15+ years of experience building consumer technology at Apple, Samsung, and Texas Instruments. He writes about how kids learn to build, think, and create in a tech-saturated world. Read more at hiwavemakers.com.

Sources

  1. National Center for Education Statistics. (2022). NAEP Long-Term Trend Assessment Results: Reading and Mathematics. U.S. Department of Education. https://nces.ed.gov/nationsreportcard/ltt/
  2. National Center for Education Statistics. (2022). NAEP Main Assessments: 4th and 8th Grade Reading and Mathematics, 2000–2022. U.S. Department of Education. https://nces.ed.gov/nationsreportcard/
  3. Ravitch, D. (2010). The Death and Life of the Great American School System: How Testing and Choice Are Undermining Education. Basic Books.
  4. Jennings, J. (2015). “NCLB: Ten Years After.” Education Week, 34(17), pp. 24–25.
  5. Hanushek, E. A., & Raymond, M. E. (2005). “Does school accountability lead to improved student performance?” Journal of Policy Analysis and Management, 24(2), pp. 297–327. https://doi.org/10.1002/pam.20091
  6. Center on Education Policy. (2007). Choices, Changes, and Challenges: Curriculum and Instruction in the NCLB Era. https://www.cep-dc.org
  7. Darling-Hammond, L. (2010). The Flat World and Education: How America’s Commitment to Equity Will Determine Our Future. Teachers College Press.
  8. Hattie, J., & Timperley, H. (2007). “The power of feedback.” Review of Educational Research, 77(1), pp. 81–112. https://doi.org/10.3102/003465430298487
Ricky Flores
Written by Ricky Flores

Founder of HiWave Makers and electrical engineer with 15+ years working on projects with Apple, Samsung, Texas Instruments, and other Fortune 500 companies. He writes about how kids learn to build, think, and create in a tech-driven world.