IN BRIEF
AI-assisted evaluation of handwritten subjective answers digitises each script with OCR, scores it against the marking scheme, and routes it through a teacher who reviews and can override every score. CrazyGoldFish validated this in a government school in Srikakulam: 93% ICC total-score reliability against teachers, with the teacher as final authority.
IN SHORT
In December 2025 we ran a controlled validation of AI-assisted evaluation of handwritten subjective answers at a government Model School in Srikakulam, Andhra Pradesh. The school carries a dual designation: Andhra Pradesh Model School and PM Shri. The validation, the App-phase study, covered 23 answer copies, 801 questions, four teachers, and four assessments. It followed an earlier Web-App phase at the same school, published in the IndiaAI Impact Summit 2026 Compendium (MeitY, Case Study 8). The headline result of the App phase is that AI-assisted evaluation, used as a governed decision-support layer with teachers retaining final authority, was reliable enough to scale with safeguards. OCR succeeded on 95.8% of questions and question mapping on 99.6%. Total scores were highly reliable against teacher scores, with an intraclass correlation of 93% and a mean absolute error of 3.91 marks. On the cases where AI and teacher disagreed, an independent adjudication against the intended marking found AI was closer more often than the teacher: AI 89.7% versus teacher 82.8%. Teacher evaluation time fell by roughly 80% in this App phase. We report the 8.7% high-severity tail openly, because governing it is the point.
The problem: evaluation is the slow, uneven step
Of all the work a teacher does, evaluating handwritten subjective answers is among the most time-intensive. A teacher reads a student's reasoning line by line, weighs it against a marking scheme, and arrives at a number. Done carefully, it is slow. Done at volume, across hundreds of scripts, it is both slow and uneven, because two evaluators rarely score the same answer identically, and the same evaluator rarely scores identically across a long sitting.
The cost of that delay is not abstract. When evaluation takes a week or two, feedback arrives after the moment for remediation has passed. In low-connectivity government settings the problem compounds, because the tools that might speed evaluation often assume bandwidth and devices the classroom does not have.
This is the gap we work at. Evaluation is the one point in a learning system where a student's reasoning is actually interpreted. We treat it as infrastructure: the place where learning signals are captured and structured, rather than reduced to a score and discarded. The question this validation set out to answer was narrow and practical. Can AI-assisted evaluation be adopted safely and effectively in a real government classroom, with teachers retaining final authority?
The deployment: one school, two phases
The work took place at a single government Model School in Srikakulam, Andhra Pradesh, designated both as an Andhra Pradesh Model School and as a PM Shri school. It is one physical school, validated there in two phases.
The first, the Web-App phase, was the broader exploration: 240 students and 12 teachers over roughly one month, published in the IndiaAI Impact Summit 2026 Compendium produced by MeitY, as Case Study 8. It reported a reduction in teacher evaluation time of up to 60%, AI and teacher scores landing within three to five marks, and feedback cycles shortening from ten to fourteen days down to two or three.
The second, the App phase, is the study we describe here. It deliberately scaled down rather than up: 23 answer copies, 801 questions, four teachers, and four assessments, in December 2025. A controlled set, graded in parallel by both AI and teachers, is what makes rigorous reliability statistics possible. We ran the exploration first to learn the workflow, then narrowed to a validation set we could measure with confidence. Depth mattered more than headline size.
The method: adjudication as ground truth
The App phase ran the full evaluation pipeline and measured each stage. Handwritten answers were digitised through OCR, which succeeded on 95.8% of questions, leaving 34 of 801 with extraction issues. Each extracted answer was mapped to its question, which succeeded on 99.6%, leaving 3 of 801. Both failure rates are low and measurable, which is what lets us route flagged items to a teacher rather than let them pass silently.
The AI then suggested rubric-aligned scores and computed totals, compared against teacher totals on the same copies graded in parallel. When an AI score and a teacher score disagreed, we did not assume the teacher was right. Every disagreement was sent to an independent adjudication against the intended marking scheme. Adjudication, not the teacher's mark and not the AI's mark, was treated as ground truth. That single move is what turns a comparison into evidence.
Throughout, the teacher's overwrite and edit functions stayed active. The teacher remained the final authority on every score, and any override was captured as an improvement signal. The workflow was validated explicitly for offline and low-bandwidth use, because a tool that assumes connectivity is not one a Tier-3 government classroom can rely on.
The findings: reliable totals, and a counter-intuitive result on the hard cases
On total scores, agreement was strong: an intraclass correlation of 93% (95% CI 0.805 to 0.976) and a mean absolute error of 3.91 marks (CI 2.61 to 5.41). 87.0% of copies fell within six marks of the teacher's total, and 91.3% within 10% of total marks.
The more interesting result came from the disagreements. Across the cases where AI and teacher diverged, adjudication found AI closer to correct more often than the teacher: AI's adjudicated correctness was 89.7% against the teacher's 82.8%, a difference of 6.9 percentage points. Across the 215 mismatch cases, AI was closer in 60.9% (131), the teacher in 32.6% (70), and both acceptable in 6.5% (14).
We hold that finding carefully. It is always a statement about the mismatch set, never a claim that AI grades better than teachers in general, and the two numbers, 89.7% and 82.8%, always travel together. On time, the App-phase workflow reduced teacher evaluation time by roughly 80% — a figure that belongs to this App phase specifically, building on the up-to-60% reduction measured in the earlier Web-App phase.
Governance: we report the tail, because governing it is the work
Alongside the strong central results, 8.7% of copies showed high-severity deviations, defined as a gap of at least the greater of six marks or 10% of the total. The confidence interval is wide, from 0.0% to 21.7%, which is itself a reason to govern it rather than wave it away.
We report the 8.7% tail openly because it is the design point, not a footnote. The tail is managed through specific controls: an upload quality gate, hotspot routing so mismatch-prone question types go to teachers, outlier escalation, and a full review of every outlier copy. Above those sits a defined escalation ladder, from teacher to principal to block and district quality assurance, and a monthly audit pack with quarterly review.
The end-to-end workflow keeps a human in the loop at every consequential step: upload, pre-provisional, provisional publication, student review and AI re-evaluation, student query, teacher closure, and final publication, with reporting throughout. A frozen policy annex sets the rules: a 72-hour provisional window, a five-working-day query SLA, and escalation within two working days. The complete trail means scores are auditable and disputes traceable, which for a government deployment is as important as the accuracy itself.
What it means
The recommendation from the App phase was to proceed to scale in a hybrid mode with defined safeguards: total scores that are reliable, a consistent second reader more often closer to correct on the hard cases, a tail that is small and openly governed, and a teacher who remains the final authority throughout.
For the district, the gains are concrete. Teachers get time back, roughly 80% of evaluation time in this App phase. Students get faster, fairer feedback while the moment for remediation is still open. The community gains transparency, administrators gain auditability and fewer disputes. None of it depends on removing the teacher from the decision. This is what we mean when we describe evaluation as infrastructure rather than automation: not a machine that grades on its own, but a system that captures the reasoning in a student's answer, structures it, and gives the teacher a reliable, auditable second read.
KEY TAKEAWAYS
Key takeaways
- →One government Model School in Srikakulam, Andhra Pradesh (APMS and PM Shri), validated in two phases: a Web-App exploration and a controlled App-phase study (this one).
- →App-phase scope: 23 copies, 801 questions, 4 teachers, 4 assessments, December 2025.
- →Extraction and mapping were reliable: 95.8% OCR success, 99.6% mapping success.
- →Total scores highly reliable against teachers: 93% ICC, 3.91 marks mean absolute error, 87.0% within six marks.
- →On disagreements, adjudication found AI closer more often: AI 89.7% vs Teacher 82.8%; across 215 mismatches, AI better 60.9%, teacher better 32.6%, both acceptable 6.5%.
- →Teacher evaluation time fell roughly 80% in the App phase (up to 60% in the earlier Web-App phase).
- →The 8.7% high-severity tail is reported openly and governed through flagging, outlier review, escalation, and a monthly audit pack, with the teacher as final authority.
Read the full validation report
The complete App-phase methodology, evidence snapshot, mismatch analysis, governance workflow, and confidence intervals (anonymised). All values traceable to the study annexure with bootstrap confidence intervals at copy level. Free, drop us an email and we'll send it your way.
Get new CrazyGoldFish research as it ships
Future studies + the transparency package, straight to your inbox.
KEEP READING