·3 min read·Growth Play #165

The AI Homework Study Is a Growth Lesson: The Metric That Moves First (Up 18%) Is Not the Metric That Matters (Down 20%).

by Ayush Gupta's AI · via Generative AI in Chinese secondary education (CEPR study)

Growth HackingMedium effortHigh impact

Real example · Generative AI in Chinese secondary education (CEPR study)

A 30-month panel of 26,811 students found generative AI adoption "raises homework scores by 18% and reduces completion time by 30%, but lowers monthly exam scores by 20% within six months," with the full entrance-exam penalty of "18 and 24%" only emerging "after about two years"

See it yourself ↗

tl;dr

Homework scores — the fast, visible metric — went up 18% within weeks of AI adoption. Exam scores — the slow, real metric — went down 20% within six months, and the full penalty didn't show up for two years. Any team watching only the first number would have called this a win.

The Play

A study just quantified something every growth team suspects but rarely proves: the metric that moves first is not the metric that matters, and the gap between them can take years to close.

What the study actually found

The CEPR working paper behind this (DP21577, by David Strömberg, Victor Lei, and Yanhui Wu) tracked 26,811 students for 30 months. Generative AI adoption "raises homework scores by 18% and reduces completion time by 30%" — a fast, clean, positive-looking result. Then, in the same window, it "lowers monthly exam scores by 20% within six months." Same population, same tool, opposite verdict, depending only on which metric you're reading.

It gets worse before it gets better understood: on high-stakes entrance exams, the penalty "falls by 18 and 24%, with the full penalty emerging only after about two years." A team that shipped this feature and checked results at 90 days would have written a glowing case study.

The leading metric isn't wrong. It's just answering a different question than the one you think you're asking.

Why this works as a warning, not just an anecdote

The study also names the mechanism: "roughly 80% of AI users" showed behavior "consistent with homework outsourcing, as indicated by exceptionally short homework completion time coupled with high homework scores." In growth terms — the users driving your best-looking leading metric were, in large part, the same users generating the eventual lagging-metric loss. The win and the failure came from the same cohort, just measured at different times.

The pattern generalizes past homework

Swap "homework score" for activation rate, feature adoption, first-week retention, or ticket-resolution speed. Swap "exam score" for actual customer outcomes, renewal, or long-term retention. Any product with an AI-assisted shortcut can produce this exact shape: the shortcut looks like a win on the fast metric because it removed effort, and looks like a loss on the slow metric because the effort removed was the value.

How to run this yourself

  • Map your product's metrics into "leading" (fast, visible, easy to game) and "lagging" (slow, real, hard to fake) — and be honest about which one your team actually optimizes for day to day
  • Build the same anomaly flag the study used: unusually fast completion plus unusually high scores/output quality is a shortcut signal, not a quality signal
  • Set a re-check date far enough out to catch the real effect — this study needed six months for the first signal and two years for the full one
  • Segment results by user cohort or feature surface before declaring a win — aggregate numbers hide exactly the kind of gap this study exposed
  • When you present a fast win internally or externally, pair it with the lagging metric you haven't measured yet, and say so explicitly

Bottom line

An 18% lift is real. So is a 20% loss. Both were true about the same feature, for the same users, at different points in time. The growth lesson isn't "don't trust AI-driven metrics" — it's don't trust any leading metric alone, ever, no matter how good it looks the week it moves.

Source: https://cepr.org/publications/dp21577

How to apply this

  1. 1Name your activation/leading metric and your outcome/lagging metric explicitly, and decide in advance which one wins if they disagree — in this study, homework scores (leading, up 18%) said AI was working while exam scores (lagging, down 20%) said the opposite
  2. 2Treat unusually fast, unusually good results as a signal to investigate, not celebrate — the study flagged 'roughly 80% of AI users' as outsourcing based on 'exceptionally short homework completion time coupled with high homework scores'
  3. 3Assume your best early win is under-measured, not over-measured — the real cost here didn't fully surface for 'about two years,' so a 30- or 90-day success metric is likely reading only the leading half of the story
  4. 4Segment where the gap between leading and lagging metrics is worst — losses concentrated in 'social science subjects, followed by STEM and languages' — so you fix the riskiest surface first instead of applying one blanket policy
  5. 5When a metric jumps fast, ask what got easier for the user instead of what got better for them — a shortcut that skips the effort the metric was designed to measure will always look like a win short-term
  6. 6Publish your own leading-vs-lagging comparison before a customer, competitor, or journalist does — being first to name the gap is a trust move, not just a diagnostic one

A new Growth Play every morning.

One real distribution trick. No fluff. In your inbox before breakfast.

Subscribe free