RESEARCH · CASE STUDY

    AI-assisted evaluation in a government classroom: what a two-phase validation in Srikakulam actually found

    A controlled validation of AI-assisted handwritten subjective evaluation in one government school, run under teacher authority, with the full reliability and tail-risk evidence reported openly.

    Government Model School · Srikakulam, Andhra Pradesh · APMS + PM Shri

    TOTAL-SCORE RELIABILITY (ICC)

    93% vs teacher scores

    CrazyGoldFish Research Team · December 2025 · 9 min read

    95.8%

    OCR success on handwritten scripts (801 questions)

    93%

    Total-score reliability vs teachers (ICC · 95% CI 0.805–0.976)

    89.7% / 82.8%

    Adjudicated AI vs Teacher correct, on the mismatch cases

    ~80%

    Teacher evaluation time saved (App phase)

    App-phase results, Dec 2025. The ~80% time saving is the App phase; the earlier Web-App phase at the same school measured up to 60% (MeitY IndiaAI Compendium). One government school, two phases — figures are never blended.

    IN BRIEF

    AI-assisted evaluation of handwritten subjective answers digitises each script with OCR, scores it against the marking scheme, and routes it through a teacher who reviews and can override every score. CrazyGoldFish validated this in a government school in Srikakulam: 93% ICC total-score reliability against teachers, with the teacher as final authority.

    IN SHORT

    In December 2025 we ran a controlled validation of AI-assisted evaluation of handwritten subjective answers at a government Model School in Srikakulam, Andhra Pradesh. The school carries a dual designation: Andhra Pradesh Model School and PM Shri. The validation, the App-phase study, covered 23 answer copies, 801 questions, four teachers, and four assessments. It followed an earlier Web-App phase at the same school, published in the IndiaAI Impact Summit 2026 Compendium (MeitY, Case Study 8). The headline result of the App phase is that AI-assisted evaluation, used as a governed decision-support layer with teachers retaining final authority, was reliable enough to scale with safeguards. OCR succeeded on 95.8% of questions and question mapping on 99.6%. Total scores were highly reliable against teacher scores, with an intraclass correlation of 93% and a mean absolute error of 3.91 marks. On the cases where AI and teacher disagreed, an independent adjudication against the intended marking found AI was closer more often than the teacher: AI 89.7% versus teacher 82.8%. Teacher evaluation time fell by roughly 80% in this App phase. We report the 8.7% high-severity tail openly, because governing it is the point.

    The problem: evaluation is the slow, uneven step

    Of all the work a teacher does, evaluating handwritten subjective answers is among the most time-intensive. A teacher reads a student's reasoning line by line, weighs it against a marking scheme, and arrives at a number. Done carefully, it is slow. Done at volume, across hundreds of scripts, it is both slow and uneven, because two evaluators rarely score the same answer identically, and the same evaluator rarely scores identically across a long sitting.

    The cost of that delay is not abstract. When evaluation takes a week or two, feedback arrives after the moment for remediation has passed. In low-connectivity government settings the problem compounds, because the tools that might speed evaluation often assume bandwidth and devices the classroom does not have.

    This is the gap we work at. Evaluation is the one point in a learning system where a student's reasoning is actually interpreted. We treat it as infrastructure: the place where learning signals are captured and structured, rather than reduced to a score and discarded. The question this validation set out to answer was narrow and practical. Can AI-assisted evaluation be adopted safely and effectively in a real government classroom, with teachers retaining final authority?

    The deployment: one school, two phases

    The work took place at a single government Model School in Srikakulam, Andhra Pradesh, designated both as an Andhra Pradesh Model School and as a PM Shri school. It is one physical school, validated there in two phases.

    The first, the Web-App phase, was the broader exploration: 240 students and 12 teachers over roughly one month, published in the IndiaAI Impact Summit 2026 Compendium produced by MeitY, as Case Study 8. It reported a reduction in teacher evaluation time of up to 60%, AI and teacher scores landing within three to five marks, and feedback cycles shortening from ten to fourteen days down to two or three.

    The second, the App phase, is the study we describe here. It deliberately scaled down rather than up: 23 answer copies, 801 questions, four teachers, and four assessments, in December 2025. A controlled set, graded in parallel by both AI and teachers, is what makes rigorous reliability statistics possible. We ran the exploration first to learn the workflow, then narrowed to a validation set we could measure with confidence. Depth mattered more than headline size.

    The method: adjudication as ground truth

    The App phase ran the full evaluation pipeline and measured each stage. Handwritten answers were digitised through OCR, which succeeded on 95.8% of questions, leaving 34 of 801 with extraction issues. Each extracted answer was mapped to its question, which succeeded on 99.6%, leaving 3 of 801. Both failure rates are low and measurable, which is what lets us route flagged items to a teacher rather than let them pass silently.

    The AI then suggested rubric-aligned scores and computed totals, compared against teacher totals on the same copies graded in parallel. When an AI score and a teacher score disagreed, we did not assume the teacher was right. Every disagreement was sent to an independent adjudication against the intended marking scheme. Adjudication, not the teacher's mark and not the AI's mark, was treated as ground truth. That single move is what turns a comparison into evidence.

    Throughout, the teacher's overwrite and edit functions stayed active. The teacher remained the final authority on every score, and any override was captured as an improvement signal. The workflow was validated explicitly for offline and low-bandwidth use, because a tool that assumes connectivity is not one a Tier-3 government classroom can rely on.

    The findings: reliable totals, and a counter-intuitive result on the hard cases

    On total scores, agreement was strong: an intraclass correlation of 93% (95% CI 0.805 to 0.976) and a mean absolute error of 3.91 marks (CI 2.61 to 5.41). 87.0% of copies fell within six marks of the teacher's total, and 91.3% within 10% of total marks.

    The more interesting result came from the disagreements. Across the cases where AI and teacher diverged, adjudication found AI closer to correct more often than the teacher: AI's adjudicated correctness was 89.7% against the teacher's 82.8%, a difference of 6.9 percentage points. Across the 215 mismatch cases, AI was closer in 60.9% (131), the teacher in 32.6% (70), and both acceptable in 6.5% (14).

    We hold that finding carefully. It is always a statement about the mismatch set, never a claim that AI grades better than teachers in general, and the two numbers, 89.7% and 82.8%, always travel together. On time, the App-phase workflow reduced teacher evaluation time by roughly 80% — a figure that belongs to this App phase specifically, building on the up-to-60% reduction measured in the earlier Web-App phase.

    Governance: we report the tail, because governing it is the work

    Alongside the strong central results, 8.7% of copies showed high-severity deviations, defined as a gap of at least the greater of six marks or 10% of the total. The confidence interval is wide, from 0.0% to 21.7%, which is itself a reason to govern it rather than wave it away.

    We report the 8.7% tail openly because it is the design point, not a footnote. The tail is managed through specific controls: an upload quality gate, hotspot routing so mismatch-prone question types go to teachers, outlier escalation, and a full review of every outlier copy. Above those sits a defined escalation ladder, from teacher to principal to block and district quality assurance, and a monthly audit pack with quarterly review.

    The end-to-end workflow keeps a human in the loop at every consequential step: upload, pre-provisional, provisional publication, student review and AI re-evaluation, student query, teacher closure, and final publication, with reporting throughout. A frozen policy annex sets the rules: a 72-hour provisional window, a five-working-day query SLA, and escalation within two working days. The complete trail means scores are auditable and disputes traceable, which for a government deployment is as important as the accuracy itself.

    What it means

    The recommendation from the App phase was to proceed to scale in a hybrid mode with defined safeguards: total scores that are reliable, a consistent second reader more often closer to correct on the hard cases, a tail that is small and openly governed, and a teacher who remains the final authority throughout.

    For the district, the gains are concrete. Teachers get time back, roughly 80% of evaluation time in this App phase. Students get faster, fairer feedback while the moment for remediation is still open. The community gains transparency, administrators gain auditability and fewer disputes. None of it depends on removing the teacher from the decision. This is what we mean when we describe evaluation as infrastructure rather than automation: not a machine that grades on its own, but a system that captures the reasoning in a student's answer, structures it, and gives the teacher a reliable, auditable second read.

    KEY TAKEAWAYS

    Key takeaways

    • One government Model School in Srikakulam, Andhra Pradesh (APMS and PM Shri), validated in two phases: a Web-App exploration and a controlled App-phase study (this one).
    • App-phase scope: 23 copies, 801 questions, 4 teachers, 4 assessments, December 2025.
    • Extraction and mapping were reliable: 95.8% OCR success, 99.6% mapping success.
    • Total scores highly reliable against teachers: 93% ICC, 3.91 marks mean absolute error, 87.0% within six marks.
    • On disagreements, adjudication found AI closer more often: AI 89.7% vs Teacher 82.8%; across 215 mismatches, AI better 60.9%, teacher better 32.6%, both acceptable 6.5%.
    • Teacher evaluation time fell roughly 80% in the App phase (up to 60% in the earlier Web-App phase).
    • The 8.7% high-severity tail is reported openly and governed through flagging, outlier review, escalation, and a monthly audit pack, with the teacher as final authority.

    Read the full validation report

    The complete App-phase methodology, evidence snapshot, mismatch analysis, governance workflow, and confidence intervals (anonymised). All values traceable to the study annexure with bootstrap confidence intervals at copy level. Free, drop us an email and we'll send it your way.

    Get new CrazyGoldFish research as it ships

    Future studies + the transparency package, straight to your inbox.

    KEEP READING

    Keep reading

    // COMMON.QUESTIONS

    Common questions

    What was validated at the Srikakulam government school?

    Whether AI-assisted evaluation of handwritten subjective answers can be adopted safely in a government classroom with teachers retaining final authority. The App-phase study (December 2025) covered 23 copies, 801 questions, four teachers, and four assessments, measuring extraction accuracy, AI-versus-teacher reliability, and the cases where the two disagreed.

    How accurate was the AI compared with the teacher?

    On total scores, agreement was strong: a 93% intraclass correlation and a mean absolute error of 3.91 marks, with 87.0% of copies within six marks. On the cases where they disagreed, an independent adjudication against the intended marking found AI closer more often than the teacher, with adjudicated correctness of AI 89.7% versus teacher 82.8%.

    Does the AI replace the teacher?

    No. The teacher remains the final authority on every score. Overwrite and edit functions stay active, and any override is captured as an improvement signal. The AI is a governed decision-support layer that gives teachers a faster, more consistent second read.

    How much time did teachers save?

    In the App phase, teacher evaluation time fell by roughly 80%. The earlier Web-App phase at the same school, published in the IndiaAI Impact Summit 2026 Compendium (MeitY, Case Study 8), measured a reduction of up to 60%. The two figures belong to two different phases and cohorts.

    What about the cases where the AI got it wrong?

    We report the tail openly. 8.7% of copies showed high-severity deviations, governed through an upload quality gate, hotspot routing, outlier escalation, a full review of every outlier copy, and a monthly audit pack, under a defined escalation ladder from teacher to principal to district quality assurance.

    How do you evaluate handwritten subjective answers at scale with AI?

    Digitise the script with OCR, score each answer against the marking scheme, and keep a teacher reviewing and overriding every score. In CrazyGoldFish's government-school validation the AI and teacher total scores agreed at 93% ICC; on the cases where they disagreed, an independent adjudicator favoured the AI more often (89.7% versus the teacher's 82.8%, on the mismatch cases only). The teacher remains the final authority.

    How reliable is AI grading compared to human examiners?

    Measured with intraclass correlation, AI and teacher total scores agreed at 93% ICC (95% CI 0.805–0.976), with a mean absolute error of 3.91 marks. Reliability is reported; accuracy is claimed only against an independent reference.

    Want our research in your inbox?

    Future studies + the transparency package, straight to your inbox.

    See how this would run for your students

    This validation describes a governed workflow that keeps teachers in control while giving them their evaluation time back. If you run assessments at scale — a school system, a network, or a government programme — we can walk you through how the same governed evaluation layer fits your context.

    // ABOUT.THE.AUTHOR

    Rahul Khandelwal

    Rahul Khandelwal is the founder and CEO of CrazyGoldFish, where he is building the evaluation infrastructure — the memory layer — for education. His focus is the hardest part of the Indian assessment stack: scoring handwritten, subjective, multilingual answers reliably and at scale, with the teacher as the final authority. Under his direction, CrazyGoldFish has run production validations with the NGO Educate Girls (handwritten Hindi answers across two states) and a government model school in Srikakulam, and its work has been recognised in the MeitY IndiaAI Impact Summit compendium and listed on AIKosh. Before founding CrazyGoldFish, Rahul spent two years at Pratham, one of India's largest education non-profits. He writes on evaluation as infrastructure, human-in-the-loop AI, and why reliability must come before accuracy in education AI.