Homework Scores Went Up 18%. Test Scores Dropped 20%.
One class saw homework scores climb 18%. The same class dropped 20% on the exam. The only variable that changed was access to an AI tutor.
That pairing should bother you. Not because AI made students dumber — but because it exposed something we’d rather not look at: the way we measure learning is falling apart in front of us.
Homework and exams never measured the same thing
Homework measures whether you arrived at the answer. Exams measure whether you can get there alone.
For most of the last century, those two things overlapped closely enough that nobody worried about the gap. A student who ground through problem sets unaided also did well on the test. The correlation held, so we treated homework as a cheap proxy for the expensive thing.
AI breaks the overlap. It delivers arrival without building the capacity to arrive. So a rising homework score isn’t a signal that the student learned more. It’s a signal that the instrument broke. You put the thermometer next to the radiator and announced a fever.
Researchers at the University of Pennsylvania ran a version of this test with roughly 1,000 high school students in Turkey. The group that worked practice problems with GPT-4 assistance saw their practice performance jump. Then the AI was taken away for the exam. That group scored 17% below students who never had help at all. Different school, different numbers, same direction.
The culprit is fluency, not laziness
Cognitive psychology has a name for this: the processing fluency illusion. When something is easy to process, your brain files it as something you know. In reality, it was just easy to read.
AI explanations are practically engineered to trigger it. Clean prose. Numbered steps. No dead ends, no ambiguity, no moment where you stare at a paragraph and realize you have no idea what it’s saying. The student gets a steady drip of small clicks — oh, that’s how it works — and every one of them feels like learning.
None of them leave a mark.
Real learning runs the other way. Retrieval practice. Spaced repetition. Interleaving concepts that are easy to confuse. Learning scientists Robert and Elizabeth Bjork call these desirable difficulties, and the defining feature of every one of them is that they feel bad while you’re doing them. You’re slow. You get things wrong. You have to go back. And that’s precisely why they stick.
AI is a machine for removing exactly that friction. It does it with the best intentions, which is what makes it hard to argue with.
Design changes the outcome completely
Here’s where the story gets less fatalistic.
That Turkish study had a third group. They used a version of the tutor with pedagogical guardrails built in — it withheld answers and offered hints instead, more Socratic method than search engine. That group showed no exam penalty at all. At Harvard, a physics course paired students with a carefully designed AI tutor and measured learning gains more than double the control group’s.
So the problem isn’t AI. The problem is frictionless AI.
And that’s an uncomfortable place to land, because nearly every consumer chatbot on the market is optimized in the opposite direction. Reducing friction is the entire product thesis. It’s what drives satisfaction scores, retention, and the word-of-mouth that makes a tool spread. Pedagogically good and commercially good point in opposite directions here, and it’s not close.
This stops being a student story fast
Let’s be honest about who else is in this experiment.
That block of code Copilot wrote last Tuesday — could you debug it right now, on your own, with the assistant off? The document you had summarized before the meeting — if someone pushed back on a claim in it, could you defend the claim?
Same mechanism, different building. Output quality goes up. The person producing the output stays flat or quietly erodes. And companies measure output almost exclusively. Every performance review in the industry is looking at homework scores.
The difference is that students have exams. There’s a scheduled moment when the illusion gets tested, marked on a calendar, unavoidable. Working professionals don’t get that. Your exam shows up unannounced — a production incident at 2 a.m., a problem shaped just differently enough that the model confidently gives you the wrong answer and you have no independent way to tell.
What actually changes
Institutions have to change the instrument. Lower the weight on take-home work, raise it on supervised, in-room assessment. Plenty of universities are already doing this, and the ones dragging their feet are mostly hoping detection tools will save them. They won’t.
For individuals the fix is smaller and more annoying: after the AI explains it, close the tab and do it again from scratch. That’s the whole method.
The gap between the point where you nodded along and the point where you can reproduce it cold — that distance is the exact size of your illusion. Most people never measure it, which is why most people never find out it’s there.
Both numbers in this story are real. Homework really did go up 18%. Exams really did fall 20%. The nasty part is that neither one is a mistake.
So: of everything you shipped last week with AI help, how much of it could you rebuild today with the tools switched off?
Comments
Loading comments...