Accuracy

How accurate is Ariel?

Ariel marks GCSE English to within the same margin that AQA expects of two professional examiners. Here is what we measured, how we measured it, and how it compares to human marking.

1.70
MAE · English Literature Paper 1 (/30)
81%
Literature Paper 1 essays within 2 marks of the examiner
< 1
MAE · Language reading questions (Q2 & Q3)

MAE = mean absolute error: the average gap, in marks, between Ariel’s mark and a trained examiner’s. Lower is better. AQA’s own examiner tolerance band is 2–3 marks. Results for every paper and question are below.

Measured

What we measured

We marked GCSE English Literature Paper 1 essays that had already been marked by AQA examiners, using Ariel’s live marking pipeline — the exact system a teacher uses in production, one mark per essay, no averaging. The essays cover Macbeth, Romeo and Juliet, A Christmas Carol and Jekyll and Hyde, from 7 out of 30 to full marks. We then compared Ariel’s marks to the examiner marks.

Measure (English Literature, /30)Result
Mean absolute error (MAE)1.70 marks
Marks within 2 of the examiner81%
Marks within 3 of the examiner89%

We measure every paper the same way. The table below gives the latest result for each one. On English Languagereading questions (Questions 2 and 3), Ariel lands in the same level as the examiner 60–80% of the time and within one level 99–100% of the time. Every Literature figure sits inside AQA’s 2–3 mark examiner tolerance band — the margin within which two qualified examiners are themselves expected to agree.

AQA paper and questionOut ofMAE
Literature Paper 1 (Shakespeare & 19th-century novel)301.70
Literature Paper 2 Section A (modern texts)302.22
Literature Paper 2 Section B (poetry)301.81
Literature Paper 2 Section C (unseen poetry)241.67
Language Paper 1 Q1 (retrieval)40.02
Language Paper 1 Q2 (language)80.79
Language Paper 1 Q3 (structure)80.81
Language Paper 1 Q4 (evaluation)201.96
Language Paper 1 Q5 writing: AO5241.71
Language Paper 1 Q5 writing: AO6161.30
Language Paper 2 Q2 (summary)80.83
Language Paper 2 Q3 (language)120.70
Language Paper 2 Q4 (comparison)161.44
Language Paper 2 Q5 writing: AO5241.32
Language Paper 2 Q5 writing: AO6160.97

How that compares to human examiners

Marking is not a fixed answer key. When a script is re-marked by a senior examiner, the grade changes more often than most people expect. In Ofqual’s 2018 marking-consistency study, GCSE English Language & Literature came out as the least consistent subject pairing of all: senior examiners gave the same grade on re-mark roughly half the time, and agreed within one grade more than 95% of the time.

Against that background, Ariel’s mark is within two marks of the examiner’s on 81% of Literature Paper 1 essays, and on Language reading questions it lands within one level of the examiner 99–100% of the time.

How we measured it

  • Real examiner marks as ground truth. Every essay carried a mark from a trained AQA examiner; that mark is what we measured Ariel against.
  • Ariel’s live pipeline. Essays were marked by the same system that runs in production — not a tuned research model.
  • One mark per essay. No averaging across multiple runs. This is what a teacher actually receives.
  • Clean text. Marking was measured on transcribed text, which isolates marking accuracy from handwriting transcription.
  • Standard metrics. Mean absolute error, the share of essays within two and three marks of the examiner, and level agreement — the same measures used across the assessment-research literature.

What this does — and does not — show

  • Top-band scripts. Ariel is most cautious at the very top: a genuine top-band Literature essay — one an examiner would give 30 out of 30 — tends to come back at around 26–27. Always give top-band scripts an extra look.
  • Marking, not transcription. These figures measure marking on clean text. Accuracy on a scanned, handwritten upload also depends on transcription, which is a separate step.
  • A first marker, not a final one. Every Ariel mark is a draft with a full rationale for the teacher to review, adjust, or override. The accuracy above is what makes that draft worth starting from — not a reason to skip the review.

Source

Human-examiner consistency figures: Ofqual (2018), Marking Consistency Metrics — An update. Ariel figures: internal measurement of the live marking system against AQA examiner marks, June–October 2026.

See the accuracy on your own class set.

Apply for a free pilot and mark a real set of essays.

Enquire about a free pilot