Assessment in the AI Era: Why Schools Are Rebuilding the Test
Table of Contents

Assessment in the AI Era: Why Schools Are Rebuilding the Test

Assessment in the AI era is the real problem: homework grades rise while exam scores fall. Here is how schools are rebuilding tests, and what to ask for.

Assessment in the AI era broke in a way that showed up as a number, not an argument. Faculty quoted by the Washington Post on September 22, 2026 described students who “turn in better homework but perform worse on exams.” They called it an illusion of learning. Strip the phrase of its drama and what remains is a measurement failure: the take-home task and the supervised task no longer agree about what the same student knows. When two instruments disagree, you do not argue about the student. You fix the instrument.

Key Takeaways

  • The homework-to-exam gap is a validity problem. An assignment only measures a student if the student is the only variable that changed.
  • The IES practice guide Organizing Instruction and Study to Improve Student Learning (Pashler et al., 2007) rates re-exposure quizzing and deep explanatory questions as having strong evidence. Both survive AI access because both require retrieval.
  • Four designs are in play: oral defence, supervised in-class production, process evidence from document version history, and performance tasks with a physical artefact.
  • Cornell’s classroom pilots, funded by a $2 million gift, will not report learning outcomes until spring or fall 2028. Nobody has the comparison data yet.
  • Teens already draw a line adults rarely notice: 54% told Pew it is acceptable to use ChatGPT to research a new topic, and only 18% said it is acceptable to write an essay with it.

What assessment in the AI era actually has to do

An assessment is an instrument for inferring what is inside a student’s head from something they produce. Its validity depends entirely on the gap between the two being small and predictable.

That link held for a century because producing a competent essay required having competent thoughts. It no longer does. The output can now be produced without the internal state, which means the inference breaks. This is not a moral problem about honesty. A perfectly honest student who asks an AI to tighten their paragraph structure has also made their essay a worse measurement of their paragraph structure.

So schools face a choice they have mostly avoided: either control the conditions of production, or measure something production cannot fake.

Why retrieval is the part that still works

The most useful research here predates AI by nearly two decades.

The IES practice guide Organizing Instruction and Study to Improve Student Learning (Pashler, Bain, Bottge, Graesser, Koedinger, McDaniel & Metcalfe, 2007) assigns evidence levels to specific instructional moves. Two carry strong evidence: using quizzes to re-expose students to key content, and asking deep explanatory questions. Spacing review across weeks or months earns moderate evidence, as does interleaving worked examples with problem-solving.

Notice what those have in common. Each requires the student to pull information out of memory rather than recognise it or produce it with help. Retrieval is both the mechanism of durable learning and, conveniently, the thing that is hardest to outsource. The guide’s framing is that retrieval “directly facilitates long-lasting memory traces,” which is why quizzing is instruction and not just measurement.

That gives schools a design rule with real research behind it: assessments built on retrieval and explanation keep working under tool access. Assessments built on production do not.

The four rebuilds schools are actually trying

DesignWhat it measuresCost to runWeakness
Oral defenceWhether the student can explain and justify their own submitted workHigh: minutes of teacher time per studentPenalises anxious and language-learning students unless scaffolded
Supervised in-class productionUnaided writing or problem-solving under time pressureMedium: class time, no techMeasures speed and nerves alongside knowledge; little revision practice
Process evidenceThe path from blank page to final draft, via version historyLow to set up, high to readEasy to game once students learn how; privacy questions about keystroke data
Performance task with artefactWhether the student can make a thing work in the physical worldHigh: materials, time, spaceHard to grade consistently; weak at measuring written reasoning

Most schools will blend these. The useful question for a parent is not which one is best but whether the school has changed anything, because an unchanged assessment regime in 2026 is reporting numbers it cannot defend.

The oral defence deserves a note. It is the oldest design on the list, it is what doctoral vivas have always done, and it is the only one that directly tests the link between the artefact and the person. It is also the most expensive, which is why it tends to appear as a short spot-check rather than a full examination. Five minutes per student across a class of thirty is two and a half hours. That is a real budget line, and when districts say they cannot afford to rebuild assessment, this is usually the item they mean.

Who is funding the rebuild, and when the answers arrive

Cornell announced classroom pilots on September 17, 2026, backed by a $2 million gift from the Dake family and the Office of the Provost, with embedded researchers helping identify and evaluate ways to incorporate AI. Deans were asked to identify a core aspect of teaching in their colleges to test. Teams may implement changes in spring 2027, fall 2027 or spring 2028, with results assessed and shared in spring or fall 2028. Vice Provost Steven Jackson described Cornell as “one of very few institutions that is taking such a deliberate and proactive approach.”

Read that timeline carefully. The most carefully designed assessment research announced this autumn produces its first findings two years from now. Everything happening in classrooms before then is improvisation, including the good improvisation.

Professional schools moved faster because their stakes are more immediate. Reuters reported on September 21, 2026 that at least 12 US law schools changed AI policies that autumn, with 36 of 180 surveyed now requiring AI instruction while others reinstated laptop bans. A laptop ban is an assessment decision disguised as a classroom rule: it controls the conditions of production.

Instruction Partners’ review of 20 classroom tools, published September 14, 2026, found that the tools the reviewers rated highest had “clear point of view about the parts of the lesson that teachers are best positioned to deliver.” Applied to assessment, that is the whole design principle in one sentence.

The assessment that never broke

There is one category of task AI access did not damage at all, and schools that teach it have a quiet advantage: anything where the work has to function.

A circuit either lights or it does not. A soldered joint either conducts or it does not. A program either compiles and produces the right output or it throws an error. A bridge made of spaghetti either holds the weight or snaps. You can ask an AI for the design, and the AI can be confidently wrong, and the artefact will tell you so in front of the whole class. The feedback is not a grade. It is physics.

That is why performance tasks appear on the rebuild list despite being the most expensive design in the table. They are the only assessments whose validity is enforced by the world rather than by the teacher. A student can get help at every stage and still have to understand enough to make the thing work, which is a different and more durable standard than producing text that reads well.

This also explains an asymmetry parents notice. Maths and lab subjects absorbed the AI shock with less disruption than history and English, not because those teachers are smarter, but because their assessments already had a verification layer built in. The subjects under the most pressure are the ones where the only evidence of thinking was a document.

What to do at home

Run the two-column check once a term

Write down your child’s grades on work completed at home and their grades on work completed in supervised conditions. Keep them in separate columns. A consistent gap is the signal the Washington Post reporting describes, and you will see it before any school does, because no gradebook separates those categories by default.

Ask your child to explain one submitted assignment out loud

Not as an interrogation. Pick something they were proud of and ask them to walk you through why they made two specific choices. A student who did the thinking enjoys this conversation. A student who did not will redirect to how long it took. That reaction is the data.

Ask the school one question about assessment design

“Has any assessment in this course changed in the last year in response to AI, and how?” If the answer is a detection tool, the school is policing production. If the answer is a change in task design, the school is rebuilding measurement. The second is harder and much more likely to help your child.

Push for retrieval, not more homework

Given the IES evidence, low-stakes quizzing spread across weeks does more for retention than longer assignments. If your child’s course has one midterm and one final, they are getting the least effective schedule the research describes. Asking for more frequent, lower-weight checks is a legitimate request with a citation behind it. Our pieces on portfolio assessment versus standardised tests and AI-proof homework cover the alternatives in more detail.

What not to do: do not rely on AI detectors as evidence

Detection tools produce probabilities, not findings, and a false positive lands on a child who has no way to prove a negative. If your school uses one, the question to ask is not about accuracy. It is about process: who reviews a flag, what the student is shown, and how an appeal works.

What to Watch For Over the Next 3 Months

  • Week 4: Watch for syllabus changes rather than policy announcements. A course that moves grade weight from take-home essays toward in-class work has made the real decision, usually without a press release.
  • Month 2 red flags: A school reporting improved outcomes using homework grades only. A new AI policy with no corresponding change to any assessment. An integrity process in which the detector’s score is treated as the finding.
  • Month 3 self-check: Ask your child to sit a past paper or an old quiz, closed-book, at the kitchen table. Compare it to their current grade. If the two disagree by a wide margin, you have found the gap in your own house and can act on it without waiting for 2028.

Frequently Asked Questions

Is the homework-to-exam gap proof that AI hurts learning?

No. The Washington Post reporting is observational, drawn from instructor accounts rather than a controlled comparison, and it concerns college students. It is strong enough to act on as a diagnostic and not strong enough to call a finding. The honest statement is that the two instruments disagree, and that is a problem regardless of cause.

Are oral defences fair to shy kids?

Only if designed for them. Done badly, an oral defence measures confidence. Done well, it is short, scheduled in advance, uses questions the student has seen the form of before, and allows notes. Ask whether the school offers a written alternative for students with anxiety accommodations.

Should my child avoid AI entirely to protect their grades?

That is not what the research supports, and teens already distinguish. In Pew’s survey of 1,391 US teens, 54% said researching a new topic with ChatGPT was acceptable while only 18% said the same about writing essays and 42% said it was not acceptable. Using a tool to find material is a different act from using it to produce the thing being graded.

Why do schools keep buying detection software instead of changing tests?

Because detection is cheap and procurable this quarter, while redesigning assessment costs teacher hours and training. Instruction Partners’ review of 20 tools found that purpose-built tools with a clear instructional role outperformed general-purpose chatbots, and the same logic applies to assessment: the cheap intervention is rarely the effective one.

When will we actually know what works?

Cornell’s embedded-research pilots report in spring or fall 2028. That is the nearest credible date for designed evidence from a US institution. Treat anything before then as a hypothesis, including the hypotheses in this article.


About the author

Ricky Flores is the founder of HiWave Makers and an electrical engineer with 15+ years of experience building consumer technology at Apple, Samsung, and Texas Instruments. He writes about how kids learn to build, think, and create in a tech-saturated world. Read more at hiwavemakers.com.


Sources

  1. The Washington Post. (2026, September 22). “‘It’s an illusion of learning’: How some educators say AI is hurting students.” https://www.washingtonpost.com/education/2026/09/22/its-illusion-learning-how-some-educators-say-ai-is-hurting-students/
  2. Pashler, H., Bain, P., Bottge, B., Graesser, A., Koedinger, K., McDaniel, M., & Metcalfe, J. (2007). “Organizing Instruction and Study to Improve Student Learning.” IES Practice Guide, National Center for Education Evaluation. https://ies.ed.gov/ncee/wwc/practiceguide/1
  3. Cornell Chronicle. (2026, September 17). “AI Education Initiatives Emphasize Experimentation, Balanced Use.” https://news.cornell.edu/stories/2026/09/ai-education-initiatives-emphasize-experimentation-balanced-use
  4. Reuters. (2026, September 21). “Laptop Bans, New Tech Courses: US Law Schools Grapple With AI.” https://www.reuters.com/legal/litigation/laptop-bans-new-tech-courses-us-law-schools-grapple-with-ai-2026-09-21/
  5. Education Week. (2026, September 14). “New Project Identifies Strengths and Weaknesses of a Collection of AI Learning Tools.” https://www.edweek.org/technology/new-project-identifies-strengths-and-weaknesses-of-a-collection-of-ai-learning-tools/2026/09
  6. Pew Research Center. (2025, January 15). “About a Quarter of U.S. Teens Have Used ChatGPT for Schoolwork.” https://www.pewresearch.org/short-reads/2025/01/15/about-a-quarter-of-us-teens-have-used-chatgpt-for-schoolwork-double-the-share-in-2023/
  7. Stanford Institute for Human-Centered AI. (2025). “AI Index Report 2025.” https://hai.stanford.edu/ai-index/2025-ai-index-report
  8. UNESCO. (2023). “Guidance for Generative AI in Education and Research.” https://www.unesco.org/en/articles/guidance-generative-ai-education-and-research
Ricky Flores
Written by Ricky Flores

Founder of HiWave Makers and electrical engineer with 15+ years working on projects with Apple, Samsung, Texas Instruments, and other Fortune 500 companies. He writes about how kids learn to build, think, and create in a tech-driven world.