UK Tests AI-Generated Writing Samples for Student Assessments
Department for Education pilot uses ChatGPT to create reference materials for moderators grading 11-year-olds' literacy work.

UK education authorities turn to generative AI for assessment standards
The UK Department for Education has begun testing artificial intelligence to generate writing samples used in standardizing literacy assessments for students transitioning from primary to secondary school. The pilot program uses ChatGPT to create reference materials that moderators compare against actual student work when verifying teacher-assigned grades.
According to details first reported by New Scientist, the initiative targets Key Stage 2 assessments, which evaluate pupils around age 11 as they complete primary education in England and Wales. The DfE claims the approach could reduce annual costs by 95 percent compared to current methods.
How the current system works
Approximately 2,000 moderators currently participate in the KS2 writing assessment process. Their role is to ensure consistency in grading standards across different schools and teachers. Moderators cross-reference student submissions against benchmark writing samples that demonstrate expected performance levels.
Traditionally, these benchmark samples come from actual work produced by schoolchildren in previous years. The existing process of collecting, curating, and distributing these authentic student samples costs around £100,000 annually.
The AI alternative under evaluation
The DfE's pilot replaces human-written benchmark samples with text generated by ChatGPT. The department has not disclosed how it prompts the AI system to produce writing that reflects different proficiency levels, nor whether the generated samples have been validated against educational standards by literacy experts.
The cost savings stem from eliminating the labor-intensive work of gathering and preparing real student writing samples each assessment cycle. If the trial proves successful, the approach could be scaled across other educational assessments.
Why it matters
This pilot represents a significant shift in how educational authorities might use generative AI—not to assess students directly, but to create the reference materials human evaluators use to standardize their judgments. The approach raises questions about whether AI-generated text can authentically represent the developmental writing patterns, creativity, and errors typical of 11-year-old children. If moderators calibrate their standards against synthetic writing that doesn't reflect real student capabilities, it could systematically skew grading standards across England and Wales, affecting thousands of pupils at a critical educational transition point.
The initiative also highlights growing pressure on education budgets and the appeal of AI solutions that promise dramatic cost reductions. Whether the £95,000 in potential savings justifies using synthetic reference materials in high-stakes assessments remains an open question as the pilot continues.
New Scientist first reported the details of this Department for Education pilot program.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call
