IN SHORT
Reliability and accuracy are different questions, and a serious validation answers reliability first. Reliability measures whether scoring holds together — it is computable from a score matrix with no reference needed. Accuracy measures correctness against an independent reference. CrazyGoldFish reports reliability — 93% ICC (intraclass correlation) in a government-school validation and an internal reliability of Cronbach's alpha = 0.808 in a Hindi handwritten deployment — and reports accuracy only when an independent reference exists, naming the limit each time.
Reliability vs accuracy at a glance
| Reliability | Accuracy | |
|---|---|---|
| Question it answers | Does the scoring hold together? | Is the score correct? |
| Needs a reference? | No (score matrix) | Yes (independent reference) |
| CrazyGoldFish metric | 93% ICC (APMS) · alpha 0.808 (Educate Girls) | Adjudication on mismatch cases |
| When CGF claims it | Always | Only with an independent reference |
Why accuracy needs a reference
Accuracy is a relative measurement: a score is accurate against a reference. Change the reference and the number changes. So an accuracy figure quoted without naming its reference has come loose from its mooring, and in practice it usually resolves to one of four things.
It might mean the AI agreed with itself on a re-run. That is consistency, not accuracy. It might mean the AI agreed with a single teacher on a single set of papers — one-grader agreement, which is the weakest available reference, because a single teacher is a noisy instrument and the same teacher reading the same script twice will not always produce the same score. It might mean the AI agreed with a model answer the system itself wrote, or a held-out set the same system curated — all circular. Or, the honest construction: the AI was tested against an independent reference — an adjudicator working from the intended marking scheme, with documented disagreements and known tail behaviour. We built that for one phase of one school deployment, and we return to it below.
This matters because education buyers are asked to make consequential decisions, sometimes at state scale, on these figures. A number that cannot survive the question "accurate against what?" is not yet a measurement. That is the discipline gap reliability is designed to close.
What reliability actually measures
Reliability is a narrower question, and it has the virtue of being answerable: does a scoring system hold together internally? Do the items in a test agree about who is stronger and weaker? When two graders mark the same scripts in parallel, how closely do their totals track? A small family of statistics does the computing.
Cronbach's alpha measures internal consistency across the items of a test, against the conventional thresholds 0.70 acceptable, 0.80 good, 0.90 excellent — the same yardstick a board examiner applies to a new question paper. Intraclass correlation (ICC) measures how reliably two graders or two scoring methods agree on totals across the same papers. Quadratic weighted kappa (QWK) measures agreement on ordered categories like a 1-to-5 rubric, penalising larger disagreements more — the statistic the largest open AI grading benchmarks have historically used.
None of these is accuracy. All are consistency, in different shapes, and that is the point: they tell you whether the scoring is coherent enough to be worth interpreting. Reliability is the floor; accuracy is the ceiling, and the ceiling is only worth discussing once the floor holds. There is a second reason to lead with reliability: it can be reproduced from a score matrix. We can publish the matrices, ship the analysis script, and let an external reviewer re-run the numbers without touching our production system. We report reliability where we can reproduce it, and accuracy only where we built the conditions to defend it.
Why we report reliability
In production validations we lead with reliability, name the threshold, publish the matrices and the script, and draw a hard line around what the study does and does not claim. The Educate Girls case study is the cleanest example. Across 11 tests in Hindi, two states, seven subjects, our scoring produced a mean Cronbach's alpha of 0.808, median 0.817, with 10 of 11 tests clearing the 0.70 threshold and seven reaching the 0.80 good-or-excellent band. The one test that did not clear had five students — a sample-size limitation, not a scoring failure. We did not run an independent accuracy adjudication in that study, and we said so on the page, in writing, twice.
That is method, not modesty. The same case study reports a difficulty curve falling cleanly from objective questions at 66% to essays at 22% with no inversions; 2,447 unique feedback variants for 2,297 student-question pairs, more distinct explanations than answers graded, which rules out templated repetition; and 94.4% of those variants explicitly referencing the model answer, the empirical anti-drift check. None of these is an accuracy number; all are evidence the scoring is doing something coherent. Evaluation earns trust the way an assessment instrument does — by being reliable, discriminating well, calibrating to difficulty, and being open to inspection. That is the standard we build to.
The teacher-as-ground-truth problem
There is a question buried here: if you want to make an accuracy claim, what counts as ground truth? The convenient answer is the classroom teacher. The honest answer is that the classroom teacher, at scale, is a noisy grader. Two teachers reading the same script will disagree; the same teacher in the second hour of a long sitting will not score as in the first. This is not a criticism of teachers — it is a property of grading as a cognitive task, and it is why psychometricians invented inter-rater statistics in the first place.
So "the AI agrees with the teacher 95% of the time" raises two questions at once: which teacher, and against what reference does the teacher itself score correctly? The APMS App-phase validation in Srikakulam took the hard path: every AI-teacher disagreement was sent to an independent adjudication against the intended marking scheme, and the adjudication — not the AI and not the teacher — was treated as the reference. The construction detail sits in the APMS case study. The conceptual point is only that an accuracy claim without an independent reference is not yet an accuracy claim.
How to read a reliability or accuracy claim
Four questions tell you whether a number is doing real work — questions worth asking of any AI grading claim, ours included.
- What is the reference? Not "what was the AI compared against," but specifically who or what adjudicated the disagreements, and against what document. If the answer is the teacher, the follow-up is how teacher noise was controlled.
- What is the reliability number, first? Cronbach's alpha, ICC, QWK — reportable across multiple tests, not a single hero figure, clearing conventional thresholds, with the cases that do not clear named and explained.
- Where are the score matrices? A claim you cannot share an anonymised matrix and analysis script for has not been built for an external check.
- What does the tail look like? Every grader misses some answers badly. The question is not whether the tail exists but how big it is, what items it concentrates on, and how the workflow handles them.
Reliability claims done right look like the two live validations on this site. Educate Girls reports mean alpha 0.808 across 11 tests with the matrices and script available on request. APMS Phase 2 reports a 93% intraclass correlation against teacher totals (95% CI 0.805–0.976); an independent adjudication on disagreements with AI at 89.7% against Teacher at 82.8%, always paired; the 8.7% high-severity tail reported openly with the governance workflow described; and the phase qualifier carried on every number so App-phase figures are never blurred into the earlier Web-App phase. That is the bar we hold. We would rather report a smaller defended number than a bigger decorative one — the memory layer for education is built on the defended ones.
KEY TAKEAWAYS
Key takeaways
- →Reliability and accuracy are different questions: reliability measures consistency; accuracy measures correctness against a reference. The reference is the part that is easy to skip.
- →An accuracy number without a named, defensible reference is not yet a measurement.
- →Reliability statistics (Cronbach's alpha, ICC, QWK) are computable from score matrices and reproducible by an external reviewer — the right floor for any AI grading claim.
- →Treating the classroom teacher as ground truth without controlling for grader noise makes an accuracy claim circular. An independent adjudicator against the intended marking is what makes accuracy defensible.
- →The two live validations show the discipline: Educate Girls reports internal reliability and explicitly presents no accuracy figure; APMS Phase 2 reports total-score reliability and adjudicated accuracy on disagreements, with the pairing and the tail named.
KEEP READING