The UK Department for Education is exploring an experimental—and controversial—plan to
supply AI-generated writing samples to the teachers who moderate Key Stage 2 (KS2) writing
assessments. Instead of reading real pupils’ work, moderators would receive anonymised
scripts produced by ChatGPT, a large-language model (LLM), in an effort to streamline the
process and reduce costs. Below, we unpack how the scheme would function, why it was
proposed, and what its wider implications might be for literacy standards, assessment
integrity, and classroom practice.
What Are Key Stage 2 Writing Assessments?
At the end of primary school (Year 6, age 11), pupils in England complete KS2 tests in
reading, maths, and spelling, punctuation & grammar (SPaG). Writing, however, is assessed
by teachers who use national teacher assessment frameworks. To ensure consistency,
external moderators visit a sample of schools and review pupils’ work, deciding whether
teachers have applied grading standards correctly.
Why Moderation Matters
KS2 outcomes influence secondary-school placement, local authority performance tables, and
targeted funding. An unreliable moderation system risks inflating or deflating pupils’
official attainment, undermining public confidence in national data and, ultimately,
educational equity.
The Government’s Cost-Saving Proposal
According to internal documents seen by educational unions, the Department for Education
spends several million pounds each year on moderation visits, largely because reviewers
must travel and examine multiple authentic writing portfolios. Officials believe that
replacing real scripts with ChatGPT-authored samples could:
- shorten the time moderators spend reading and grading;
- make it easier to generate writing at specific ability levels;
- eliminate the administrative burden of anonymising children’s work.
How Would ChatGPT Be Used?
The project team envisage prompting ChatGPT to produce short stories, reports, or
persuasive letters aligned with the national curriculum. Each script would be labelled as
“Working Towards”, “Expected”, or “Greater Depth” so moderators could demonstrate that they
can reliably identify the intended standard. Crucially, these AI texts would not
replace real pupils’ work in high-stakes grading—at least in the pilot phase—but would act
as calibration tools.
Potential Advantages
Advocates cite several possible benefits:
- Cost efficiency: fewer travel days and reduced printing or scanning costs.
- Consistency: AI can generate near-identical examples for every moderator
nationwide, reducing variability in training materials. - Rapid iteration: scripts can be updated instantly when frameworks change,
avoiding the lag involved in collecting new pupil work. - Privacy: no risk of doxxing or identification of individual children.
Key Concerns and Criticisms
The proposal has drawn sharp criticism from teacher unions, assessment experts, and
children’s-rights advocates:
- Authenticity: AI texts may lack the idiosyncrasies, vocabulary gaps,
spelling errors, and developmental quirks found in genuine Year 6 writing, leading to
unrealistic moderation conditions. - Bias & style homogenisation: LLMs often default to middle-class
cultural references and standardised syntax, potentially embedding bias into the
official benchmark of “good” writing. - Skill distortion: Moderators might over-focus on surface-level
features that ChatGPT reproduces well (e.g., grammar, cohesion) while under-valuing
creative voice and authentic pupil expression. - Transparency: The public may question whether AI use signals a wider
move toward automated marking, raising ethical and professional issues.
What Do Educators Say?
• The National Education Union argues that the scheme “risks turning moderation into a box-ticking exercise divorced from real classroom work.”
• Some moderators welcome the idea of having large banks of clearly tiered examples but
emphasise that these must be supplemented with actual children’s writing.
• Assessment researchers point out that reliability gains from AI exemplars could be
offset by validity losses if the tasks no longer mirror the complexities of
genuine pupil output.
Technical and Ethical Safeguards Under Discussion
The Department for Education is reportedly testing several safeguards:
- Fine-tuning the model on anonymised pupil scripts to retain realistic error patterns.
- Human vetting panels to approve or reject each AI script before release.
- Clear labelling that prevents moderators from mistaking AI exemplars for real
evidence. - Impact evaluations comparing moderation accuracy with and without AI materials.
Looking Ahead: Possible Scenarios
• Pilot adoption only: AI exemplars remain a niche training aid while genuine
writing continues to anchor the moderation process.
• Hybrid model: Both AI and pupil scripts are used, with LLM outputs covering
routine errors and genuine work illustrating creativity and diversity.
• Full automation drift: Success in moderation encourages policymakers to
explore AI-marked writing tests—an avenue already trialled in other jurisdictions.
The push to use ChatGPT in KS2 moderation highlights the tension between cost
efficiency and assessment authenticity. While AI-generated scripts might help
standardise training and cut administrative overheads, they raise substantive questions
about validity, bias, and the future role of human judgement in education. As pilots
progress, policymakers will need to demonstrate that any savings do not come at the
expense of the nuanced, child-centred understanding that effective writing assessment
demands.



