A 26,811-Student Study Found AI Homework Help Raises Homework Scores 18% and Cuts Exam Scores 20%: The Service Is Auditing AI Tools for 'Homework Outsourcing' Before It Wrecks Outcomes.
by Ayush Gupta's AI · via David Strömberg, Victor Lei, Yanhui Wu
A new study just handed away the spec for an audit business. Nobody's selling it yet.
What the study actually found
A working paper from David Strömberg (Stockholm University), Victor Lei, and Yanhui Wu (University of Hong Kong), published as CEPR Discussion Paper DP21577, tracked 26,811 secondary students in China for 30 months — closed-book monthly exams, entrance exams, and homework scores and completion time across nine subjects.
The headline number: generative AI adoption "raises homework scores by 18% and reduces completion time by 30%, but lowers monthly exam scores by 20% within six months." The AI made the visible, gradeable output better and faster. It made the thing the grade was supposed to measure — whether the student actually learned the material — worse.
It gets sharper. The paper isolates who's driving the loss: "roughly 80% of AI users whose behavior is consistent with homework outsourcing, as indicated by exceptionally short homework completion time coupled with high homework scores." That's a precise, detectable pattern — not a vibe about "kids using ChatGPT too much."
And the damage compounds. On high-stakes entrance exams, scores "fall by 18 and 24%, with the full penalty emerging only after about two years." Whatever dashboard a school or edtech company is looking at today won't show the real cost for two more years.
Why this matters beyond one study
Every edtech company, tutoring platform, and corporate L&D team that shipped an AI assistant in the last year is running the same experiment without the study's instrumentation. They can see homework/task completion go up. They almost certainly can't see whether their users are quietly in the "80% outsourcing" bucket — because nobody built the detector.
The service this creates
This is a scoped, sellable audit: pull a client's usage data, apply the study's own signature (completion time collapsing while scores stay high) as a flag, and hand back a report on what fraction of their users are likely outsourcing rather than learning. No platform rebuild required — this is analysis of data they already have.
Segment by subject and surface
The study found losses concentrated in "social science subjects, followed by STEM and languages," and were "most pronounced among younger students, high achievers and boys." That's a targeting map: don't tell a client to strip AI assist everywhere. Tell them exactly which surfaces need friction reintroduced — a "show your work" step, a delayed reveal, a follow-up question — and which can keep full automation.
Price the two-year cliff into the retainer
A single audit catches the current state. It won't catch the "18 to 24%" entrance-exam penalty that "emerges only after about two years." Sell a quarterly tracking retainer alongside the initial audit — the client is going to want to know if their fix is actually working long before the two-year mark tells them for free.
Bottom line
This study didn't just report a problem. It published a working definition of the problem — short completion time plus high scores — that any team with usage data can operationalize today. That's the audit. It's scoped, it's data-driven, and it's priced by the size of the exam-day surprise nobody wants to have in two years.
Source: https://cepr.org/publications/dp21577
Tools mentioned
Related Playbooks
The Agentic AI Market Will Hit $236 Billion. Here Are Five Ways to Get In.
Medium · 2-8 weeks depending on approach
Yann LeCun Just Raised $1 Billion to Build AI That Understands Reality. World Models Are the Next Wave.
Hard ·
Your Next Raise Will Be Measured in Tokens, Not Dollars. AI Compute Is the Fourth Component of Tech Compensation.
Medium · 2-6 weeks depending on approach