IN BRIEF
CrazyGoldFish evaluated handwritten Hindi-medium subjective answers in production with the NGO Educate Girls, across two states and seven subjects — 2,297 student-question evaluations. Scoring reached a mean internal reliability (Cronbach's alpha) of 0.808 across 11 tests, a measure of consistency, not accuracy, and generated 2,447 unique feedback explanations, 94.4% grounded in the model answer.
IN SHORT
Educate Girls ran handwritten Hindi-medium answer scripts from two states and seven subjects through our production evaluation system, then checked the results against standard psychometrics. Across 11 tests, AI scoring reached a mean internal reliability of α = 0.808, with 10 of 11 tests clearing the accepted threshold, and produced 2,447 unique feedback explanations for 2,297 student-question pairs — 94.4% grounded in the model answer. The validation makes no controlled accuracy claim; it reports what the dataset supports and ships the code so the numbers can be re-run.
The problem
Educate Girls works on out-of-school girls and learning outcomes in government schools, and it runs assessment cycles in Hindi, at a scale where there is not enough human evaluation capacity to give every student close attention on every script. That is the ordinary shape of the problem across most of Indian education. A teacher's read of a handwritten answer contains a great deal of signal: what concept the student grasped, where the reasoning broke, which mistake is a spelling slip and which is a real gap. In most systems that signal is collapsed into a single number and then discarded. The score survives. The learning event does not.
Before scaling AI evaluation across its programme, the Educate Girls central team wanted to know whether the system could carry the weight of that read. Could it interpret Hindi handwriting across subjects, not only on language papers? Would per-question feedback hold up when a human reviewed it? Could one workflow handle every question format the programme actually uses, from multiple-choice up to explanatory essays? These are not questions a demo answers. They are questions a real dataset answers.
The deployment
The team uploaded real production answer scripts into our system. The validation covered two states — Madhya Pradesh and Rajasthan — across seven subjects, coming to 162 questions and roughly 140 students, producing 2,297 student-question evaluations in all. Our system is human-in-the-loop by design, with review, override and an audit trail built in. Here, the Educate Girls review process was itself the human layer. The point of the exercise was to see whether what the AI handed that layer was worth reviewing.
The method
We did not score the system against our own opinion of it. We computed the same statistics a psychometrician would apply to any new assessment instrument, working from the per-student per-question score matrices pulled directly from the production scoring API. Reliability was measured with Cronbach's alpha against the usual thresholds: 0.70 acceptable, 0.80 good, 0.90 excellent. Item quality was measured with corrected item-total correlation, where above 0.30 marks a strong discriminator. Difficulty was checked for monotonicity — the expectation that harder formats should, on average, score lower.
The findings
The scoring was reliable. Mean internal reliability across the 11 tests was α = 0.808, median 0.817. Ten of 11 cleared the 0.70 threshold; 7 reached good or excellent. The single test below threshold had only five students — a small-sample limitation, not a scoring failure.
The difficulty curve held. Mean accuracy fell cleanly from objective and multiple-choice questions (66.2%) down to explanatory and essay questions (22.1%), with no inversions anywhere. A grader that did not understand what it was reading would have scored some essays higher than some multiple-choice items. This one did not.
It caught a programme-level signal. On all three subjects shared between the two states, the Madhya Pradesh cohort outscored Rajasthan — by 23.3 points in Hindi, 16.4 in Home Science, 9.1 in Social Science — the kind of repeatable cross-cohort gap a programme director can act on.
The feedback taught, rather than only scored. The system produced 2,447 unique feedback variants for 2,297 pairs — writing more distinct explanations than there were answers to grade, which templated feedback cannot do. 94.4% explicitly referenced the model answer, the strongest available check against feedback that drifts from the source.
What it means
Evaluation is the one point in the learning cycle where a student's reasoning is interpreted and validated. We build the layer that captures what happens there and turns it into structured, persistent learning memory — instead of letting it collapse into a score and vanish. The Educate Girls validation is evidence that the layer holds at the hardest part of the Indian assessment stack: handwritten, subjective answers, in Hindi, across every question format a real programme uses. Two things about how the evidence was produced matter as much as the numbers. We named what the study does not claim — no controlled accuracy figure, no inter-rater agreement study — and we shipped the code so the rest can be checked.
KEY TAKEAWAYS
Key takeaways
- →Mean internal reliability α = 0.808 across 11 tests; 10 of 11 cleared the accepted threshold
- →73–100% of items per test were strong discriminators; clean difficulty curve, no inversions
- →2,447 unique feedback variants for 2,297 pairs; 94.4% grounded in the model answer
- →Reliability — not accuracy. The study names its own limits and ships reproducible code
Download the full 24-page study
Complete methodology, per-test reliability tables and exemplars — prepared with Educate Girls. Free, drop us an email and we'll send it your way.
Get new CrazyGoldFish research as it ships
Future studies + the transparency package, straight to your inbox.
KEEP READING