AI Grading Tools Inflate Student Essay Scores, Study Finds
Research shows ChatGPT and similar models award higher marks than human evaluators and cannot reliably assess extended written work.

Large language models consistently award higher grades than human evaluators when assessing student essays and cannot be trusted to accurately measure academic performance, according to new research that raises questions about the rush to automate university grading.
A study published in Assessment & Evaluation in Higher Education examined how ChatGPT evaluated 50 undergraduate bioscience essays across seven assessment criteria. Researchers tested two versions of the AI tool using four different prompting approaches, then compared results against human-assigned grades.
The grading gap
The findings revealed substantial misalignment between AI and human judgment. In nearly every test condition, the AI models returned higher average marks than human graders. The most extreme case showed a 40-point difference on a 100-point scale for a single essay.
The discrepancy followed a pattern: lower-performing essays received inflated marks from AI, while stronger essays were marked down compared to human assessment. Only mid-range essays showed reasonable alignment between AI and human evaluators.
Even when the AI tools produced consistent marks for the same essay using identical prompts, they failed to maintain that consistency across essays of varying quality. While overall composite scores showed some correlation with human marks, individual assessment criteria diverged considerably.
Why it matters
Universities face mounting pressure to reduce faculty workload, and AI-assisted grading has emerged as a potential solution. This research demonstrates that current technology cannot substitute for human judgment in evaluating subjective written work—a finding with immediate implications for institutions already piloting automated assessment tools. Beyond accuracy concerns, the study highlights unresolved ethical questions about submitting student work to commercial AI platforms without explicit consent.
Technical limitations persist
William Kay, senior lecturer in statistics at Cardiff University and study co-author, emphasized that the research never intended to explore full AI replacement of human graders. The team originally sought to evaluate whether generative AI could serve as a formative benchmarking tool to supplement faculty feedback.
"At present, GenAI is unable to reliably assign a mark to a subjective piece of written work comparable to that of humans—even with extensive training of the LLM," Kay said.
The researchers noted that more detailed grading rubrics using objectively distinct descriptions for each mark bracket might improve AI performance, rather than relying on subjective terms like "good, excellent, outstanding." However, they cautioned that inherent variation among human markers themselves may make perfect alignment impossible to achieve.
Kay acknowledged that future AI systems may develop better judgment-mimicking capabilities, but concluded that current models should not be relied upon for assigning grades to extended written work.
The findings were first reported by Inside Higher Ed.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call
