AI Homework Help Cuts Exam Scores 20% in Chinese Schools
Administrative data from 26,000 secondary students reveal that self-directed generative AI use improves homework grades while undermining actual learning.
The homework paradox
Generative AI tools help students finish assignments faster and score higher on homework—but new research shows this productivity gain comes at a steep cost to actual learning. A study tracking 26,811 secondary school students in China over 30 months found that self-directed use of AI tools like ChatGPT and DeepSeek improved homework scores by 18% while reducing closed-book exam scores by 20% within six months of adoption.
The research, conducted by David Strömberg, Victor Lei, and Yanhui Wu, analyzed administrative records from grades 7 through 12 across nine subjects. Students completed homework 30% faster after adopting AI, dropping from 64 minutes to 45 minutes per assignment. Yet their performance on exams—which measure retained knowledge without AI assistance—declined significantly.
Why it matters
This study provides the first large-scale evidence of how students actually use generative AI in natural educational settings, rather than controlled experiments with prescribed tools. The findings suggest that short-term productivity gains mask cumulative learning losses that compound over time, with implications for college admissions, career trajectories, and the fundamental question of how to integrate AI into education without undermining its core purpose.
The penalty grows with time
The learning penalty starts small but intensifies as students become more proficient with AI tools and as more course material accumulates after adoption. For monthly exams testing recent material, the full effect appears around six months. For high-stakes entrance exams covering years of coursework, losses build over roughly two years but reach similar magnitude: 24% for high-school entrance exams and 18% for college entrance exams.
This time lag means existing short-duration studies likely underestimate the long-run costs of AI use in education, according to the researchers.
Two patterns of use
The data reveal a sharp divide in how students employ AI. Roughly 80% of AI users complete assignments in under 50 minutes—faster than any non-AI students—and receive exceptionally high homework scores matching AI capability. Their exam performance, however, drops markedly, consistent with outsourcing substantial homework to AI.
The remaining 20% of AI users spend as much time on homework as non-AI students and achieve similar exam scores, despite also using AI for homework as evidenced by higher homework grades. When AI students invest the same time, they learn as much.
Compression from the top
Unlike workplace studies showing AI disproportionately helps lower-skilled workers, the educational penalty falls hardest on initially high-performing students. Social science subjects saw the largest drops (27%), followed by STEM (22%), English (17%), and Chinese (9%). Junior students, male students, and top performers experienced disproportionately large losses.
By June 2025, roughly 80% of students in the study reported using generative AI, up from nearly zero in September 2022. The most common tools were Doubao and DeepSeek—general-purpose AI platforms without special educational features.
The incentive problem
The researchers note that well-designed tutoring tools already exist at near-zero cost, but students prefer general-purpose AI that makes homework outsourcing easy. This suggests the core challenge is not technological but organizational: how to incentivize productive AI use when tools that provide direct answers are readily available.
The study also found that homework scores no longer predict learning among AI users. Students with higher homework scores were more likely to perform worse on exams—inverting the traditional relationship and rendering homework grades unreliable as learning metrics.
These findings were first reported by Strömberg, Lei, and Wu in CEPR Discussion Paper 21577 and published on VoxEU.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call
