RESEARCH · CASE STUDY

    Reading Hindi handwriting at scale: what a national NGO found when it validated our AI evaluation

    Educate Girls tested CrazyGoldFish on real handwritten Hindi answer scripts across two states and seven subjects — then checked the results against standard psychometrics.

    Educate Girls · Hindi · 2 states · 7 subjects

    MEAN INTERNAL RELIABILITY

    α = 0.808 across 11 tests

    Rahul Khandelwal — Founder & CEO, CrazyGoldFish · June 2026 · 9 min read

    0.808

    Mean internal reliability (Cronbach's α), median 0.817

    10 of 11

    Tests clearing the α ≥ 0.70 acceptable threshold

    2,297

    Student-question evaluations · 162 questions · 7 subjects

    2,447

    Unique feedback variants · 94.4% grounded in the model answer

    α = 0.808 is internal reliability (scoring consistency) — not an accuracy figure. This study does not present one.

    IN BRIEF

    CrazyGoldFish evaluated handwritten Hindi-medium subjective answers in production with the NGO Educate Girls, across two states and seven subjects — 2,297 student-question evaluations. Scoring reached a mean internal reliability (Cronbach's alpha) of 0.808 across 11 tests, a measure of consistency, not accuracy, and generated 2,447 unique feedback explanations, 94.4% grounded in the model answer.

    IN SHORT

    Educate Girls ran handwritten Hindi-medium answer scripts from two states and seven subjects through our production evaluation system, then checked the results against standard psychometrics. Across 11 tests, AI scoring reached a mean internal reliability of α = 0.808, with 10 of 11 tests clearing the accepted threshold, and produced 2,447 unique feedback explanations for 2,297 student-question pairs — 94.4% grounded in the model answer. The validation makes no controlled accuracy claim; it reports what the dataset supports and ships the code so the numbers can be re-run.

    The problem

    Educate Girls works on out-of-school girls and learning outcomes in government schools, and it runs assessment cycles in Hindi, at a scale where there is not enough human evaluation capacity to give every student close attention on every script. That is the ordinary shape of the problem across most of Indian education. A teacher's read of a handwritten answer contains a great deal of signal: what concept the student grasped, where the reasoning broke, which mistake is a spelling slip and which is a real gap. In most systems that signal is collapsed into a single number and then discarded. The score survives. The learning event does not.

    Before scaling AI evaluation across its programme, the Educate Girls central team wanted to know whether the system could carry the weight of that read. Could it interpret Hindi handwriting across subjects, not only on language papers? Would per-question feedback hold up when a human reviewed it? Could one workflow handle every question format the programme actually uses, from multiple-choice up to explanatory essays? These are not questions a demo answers. They are questions a real dataset answers.

    The deployment

    The team uploaded real production answer scripts into our system. The validation covered two states — Madhya Pradesh and Rajasthan — across seven subjects, coming to 162 questions and roughly 140 students, producing 2,297 student-question evaluations in all. Our system is human-in-the-loop by design, with review, override and an audit trail built in. Here, the Educate Girls review process was itself the human layer. The point of the exercise was to see whether what the AI handed that layer was worth reviewing.

    The method

    We did not score the system against our own opinion of it. We computed the same statistics a psychometrician would apply to any new assessment instrument, working from the per-student per-question score matrices pulled directly from the production scoring API. Reliability was measured with Cronbach's alpha against the usual thresholds: 0.70 acceptable, 0.80 good, 0.90 excellent. Item quality was measured with corrected item-total correlation, where above 0.30 marks a strong discriminator. Difficulty was checked for monotonicity — the expectation that harder formats should, on average, score lower.

    The findings

    The scoring was reliable. Mean internal reliability across the 11 tests was α = 0.808, median 0.817. Ten of 11 cleared the 0.70 threshold; 7 reached good or excellent. The single test below threshold had only five students — a small-sample limitation, not a scoring failure.

    The difficulty curve held. Mean accuracy fell cleanly from objective and multiple-choice questions (66.2%) down to explanatory and essay questions (22.1%), with no inversions anywhere. A grader that did not understand what it was reading would have scored some essays higher than some multiple-choice items. This one did not.

    It caught a programme-level signal. On all three subjects shared between the two states, the Madhya Pradesh cohort outscored Rajasthan — by 23.3 points in Hindi, 16.4 in Home Science, 9.1 in Social Science — the kind of repeatable cross-cohort gap a programme director can act on.

    The feedback taught, rather than only scored. The system produced 2,447 unique feedback variants for 2,297 pairs — writing more distinct explanations than there were answers to grade, which templated feedback cannot do. 94.4% explicitly referenced the model answer, the strongest available check against feedback that drifts from the source.

    What it means

    Evaluation is the one point in the learning cycle where a student's reasoning is interpreted and validated. We build the layer that captures what happens there and turns it into structured, persistent learning memory — instead of letting it collapse into a score and vanish. The Educate Girls validation is evidence that the layer holds at the hardest part of the Indian assessment stack: handwritten, subjective answers, in Hindi, across every question format a real programme uses. Two things about how the evidence was produced matter as much as the numbers. We named what the study does not claim — no controlled accuracy figure, no inter-rater agreement study — and we shipped the code so the rest can be checked.

    KEY TAKEAWAYS

    Key takeaways

    • Mean internal reliability α = 0.808 across 11 tests; 10 of 11 cleared the accepted threshold
    • 73–100% of items per test were strong discriminators; clean difficulty curve, no inversions
    • 2,447 unique feedback variants for 2,297 pairs; 94.4% grounded in the model answer
    • Reliability — not accuracy. The study names its own limits and ships reproducible code

    Download the full 24-page study

    Complete methodology, per-test reliability tables and exemplars — prepared with Educate Girls. Free, drop us an email and we'll send it your way.

    Get new CrazyGoldFish research as it ships

    Future studies + the transparency package, straight to your inbox.

    KEEP READING

    Keep reading

    // COMMON.QUESTIONS

    Common questions

    Is α = 0.808 an accuracy score?

    No. α (Cronbach's alpha) measures internal reliability — how consistently the items in a test work together. It is not accuracy, and this validation does not present an accuracy figure.

    Can AI reliably grade handwritten Hindi subjective answers?

    In this production deployment, scoring across 11 tests reached a mean internal reliability of α = 0.808, with 10 of 11 tests clearing the accepted threshold and a clean difficulty curve from multiple-choice down to essays — grounded in real data across two states, seven subjects and 162 questions.

    How do you know the feedback is not templated?

    The system produced 2,447 unique feedback variants for 2,297 pairs — more distinct explanations than answers graded — and 94.4% explicitly referenced the model answer.

    Can the results be independently verified?

    Yes. A reproducibility package with the anonymised score matrices for all 11 tests, the analysis scripts and the data-quality filter is available on request, so the analysis can be re-run without API access.

    How do you grade Hindi or regional-language handwritten answers with AI?

    With OCR tuned for Indian-language handwriting, rubric-aligned scoring, and a teacher-override step. In CrazyGoldFish's Hindi deployment with Educate Girls, internal reliability was alpha = 0.808 across 11 tests — a consistency measure, not accuracy — with feedback grounded in the model answer 94.4% of the time.

    How do you evaluate handwritten subjective answers at scale with AI?

    Digitise with OCR, score against the marking scheme, and keep a teacher reviewing every score. This ran in production across two states and seven subjects (2,297 student-question evaluations), reaching a mean internal reliability of alpha = 0.808.

    Want our research in your inbox?

    Future studies + the transparency package, straight to your inbox.

    See it on your own students

    If you run Hindi-medium or subjective assessment at scale, the same validation can be run on your own scripts.

    // ABOUT.THE.AUTHOR

    Rahul Khandelwal

    Rahul Khandelwal is the founder and CEO of CrazyGoldFish, where he is building the evaluation infrastructure — the memory layer — for education. His focus is the hardest part of the Indian assessment stack: scoring handwritten, subjective, multilingual answers reliably and at scale, with the teacher as the final authority. Under his direction, CrazyGoldFish has run production validations with the NGO Educate Girls (handwritten Hindi answers across two states) and a government model school in Srikakulam, and its work has been recognised in the MeitY IndiaAI Impact Summit compendium and listed on AIKosh. Before founding CrazyGoldFish, Rahul spent two years at Pratham, one of India's largest education non-profits. He writes on evaluation as infrastructure, human-in-the-loop AI, and why reliability must come before accuracy in education AI.