RESEARCH · CONCEPT

    Reliability vs accuracy: the number that actually matters

    Why we lead with reliability, and claim accuracy only against a defensible reference.

    THE NUMBER THAT MATTERS

    Reliability against a defensible reference

    Rahul Khandelwal — Founder & CEO, CrazyGoldFish · June 2026 · 10 min read

    // AT.A.GLANCE

    TermWhat it means
    ReliabilityA measure of consistency. Does the grader produce coherent, repeatable scores across items, students, and sittings?
    AccuracyA measure of correctness against a reference. Without a defensible reference, accuracy has no meaning.
    What we reportReliability statistics in production, and accuracy only where an independent reference exists. We name the limit each time.
    "Reliability is what a psychometrician can compute from a score matrix alone. Accuracy needs an adjudicator that is not the AI and not the teacher. If the reference is not named, the accuracy claim cannot be checked."

    IN SHORT

    Reliability and accuracy are different questions, and a serious validation answers reliability first. Reliability measures whether scoring holds together — it is computable from a score matrix with no reference needed. Accuracy measures correctness against an independent reference. CrazyGoldFish reports reliability — 93% ICC (intraclass correlation) in a government-school validation and an internal reliability of Cronbach's alpha = 0.808 in a Hindi handwritten deployment — and reports accuracy only when an independent reference exists, naming the limit each time.

    Reliability vs accuracy at a glance

    ReliabilityAccuracy
    Question it answersDoes the scoring hold together?Is the score correct?
    Needs a reference?No (score matrix)Yes (independent reference)
    CrazyGoldFish metric93% ICC (APMS) · alpha 0.808 (Educate Girls)Adjudication on mismatch cases
    When CGF claims itAlwaysOnly with an independent reference

    Why accuracy needs a reference

    Accuracy is a relative measurement: a score is accurate against a reference. Change the reference and the number changes. So an accuracy figure quoted without naming its reference has come loose from its mooring, and in practice it usually resolves to one of four things.

    It might mean the AI agreed with itself on a re-run. That is consistency, not accuracy. It might mean the AI agreed with a single teacher on a single set of papers — one-grader agreement, which is the weakest available reference, because a single teacher is a noisy instrument and the same teacher reading the same script twice will not always produce the same score. It might mean the AI agreed with a model answer the system itself wrote, or a held-out set the same system curated — all circular. Or, the honest construction: the AI was tested against an independent reference — an adjudicator working from the intended marking scheme, with documented disagreements and known tail behaviour. We built that for one phase of one school deployment, and we return to it below.

    This matters because education buyers are asked to make consequential decisions, sometimes at state scale, on these figures. A number that cannot survive the question "accurate against what?" is not yet a measurement. That is the discipline gap reliability is designed to close.

    What reliability actually measures

    Reliability is a narrower question, and it has the virtue of being answerable: does a scoring system hold together internally? Do the items in a test agree about who is stronger and weaker? When two graders mark the same scripts in parallel, how closely do their totals track? A small family of statistics does the computing.

    Cronbach's alpha measures internal consistency across the items of a test, against the conventional thresholds 0.70 acceptable, 0.80 good, 0.90 excellent — the same yardstick a board examiner applies to a new question paper. Intraclass correlation (ICC) measures how reliably two graders or two scoring methods agree on totals across the same papers. Quadratic weighted kappa (QWK) measures agreement on ordered categories like a 1-to-5 rubric, penalising larger disagreements more — the statistic the largest open AI grading benchmarks have historically used.

    None of these is accuracy. All are consistency, in different shapes, and that is the point: they tell you whether the scoring is coherent enough to be worth interpreting. Reliability is the floor; accuracy is the ceiling, and the ceiling is only worth discussing once the floor holds. There is a second reason to lead with reliability: it can be reproduced from a score matrix. We can publish the matrices, ship the analysis script, and let an external reviewer re-run the numbers without touching our production system. We report reliability where we can reproduce it, and accuracy only where we built the conditions to defend it.

    Why we report reliability

    In production validations we lead with reliability, name the threshold, publish the matrices and the script, and draw a hard line around what the study does and does not claim. The Educate Girls case study is the cleanest example. Across 11 tests in Hindi, two states, seven subjects, our scoring produced a mean Cronbach's alpha of 0.808, median 0.817, with 10 of 11 tests clearing the 0.70 threshold and seven reaching the 0.80 good-or-excellent band. The one test that did not clear had five students — a sample-size limitation, not a scoring failure. We did not run an independent accuracy adjudication in that study, and we said so on the page, in writing, twice.

    That is method, not modesty. The same case study reports a difficulty curve falling cleanly from objective questions at 66% to essays at 22% with no inversions; 2,447 unique feedback variants for 2,297 student-question pairs, more distinct explanations than answers graded, which rules out templated repetition; and 94.4% of those variants explicitly referencing the model answer, the empirical anti-drift check. None of these is an accuracy number; all are evidence the scoring is doing something coherent. Evaluation earns trust the way an assessment instrument does — by being reliable, discriminating well, calibrating to difficulty, and being open to inspection. That is the standard we build to.

    The teacher-as-ground-truth problem

    There is a question buried here: if you want to make an accuracy claim, what counts as ground truth? The convenient answer is the classroom teacher. The honest answer is that the classroom teacher, at scale, is a noisy grader. Two teachers reading the same script will disagree; the same teacher in the second hour of a long sitting will not score as in the first. This is not a criticism of teachers — it is a property of grading as a cognitive task, and it is why psychometricians invented inter-rater statistics in the first place.

    So "the AI agrees with the teacher 95% of the time" raises two questions at once: which teacher, and against what reference does the teacher itself score correctly? The APMS App-phase validation in Srikakulam took the hard path: every AI-teacher disagreement was sent to an independent adjudication against the intended marking scheme, and the adjudication — not the AI and not the teacher — was treated as the reference. The construction detail sits in the APMS case study. The conceptual point is only that an accuracy claim without an independent reference is not yet an accuracy claim.

    How to read a reliability or accuracy claim

    Four questions tell you whether a number is doing real work — questions worth asking of any AI grading claim, ours included.

    • What is the reference? Not "what was the AI compared against," but specifically who or what adjudicated the disagreements, and against what document. If the answer is the teacher, the follow-up is how teacher noise was controlled.
    • What is the reliability number, first? Cronbach's alpha, ICC, QWK — reportable across multiple tests, not a single hero figure, clearing conventional thresholds, with the cases that do not clear named and explained.
    • Where are the score matrices? A claim you cannot share an anonymised matrix and analysis script for has not been built for an external check.
    • What does the tail look like? Every grader misses some answers badly. The question is not whether the tail exists but how big it is, what items it concentrates on, and how the workflow handles them.

    Reliability claims done right look like the two live validations on this site. Educate Girls reports mean alpha 0.808 across 11 tests with the matrices and script available on request. APMS Phase 2 reports a 93% intraclass correlation against teacher totals (95% CI 0.805–0.976); an independent adjudication on disagreements with AI at 89.7% against Teacher at 82.8%, always paired; the 8.7% high-severity tail reported openly with the governance workflow described; and the phase qualifier carried on every number so App-phase figures are never blurred into the earlier Web-App phase. That is the bar we hold. We would rather report a smaller defended number than a bigger decorative one — the memory layer for education is built on the defended ones.

    KEY TAKEAWAYS

    Key takeaways

    • Reliability and accuracy are different questions: reliability measures consistency; accuracy measures correctness against a reference. The reference is the part that is easy to skip.
    • An accuracy number without a named, defensible reference is not yet a measurement.
    • Reliability statistics (Cronbach's alpha, ICC, QWK) are computable from score matrices and reproducible by an external reviewer — the right floor for any AI grading claim.
    • Treating the classroom teacher as ground truth without controlling for grader noise makes an accuracy claim circular. An independent adjudicator against the intended marking is what makes accuracy defensible.
    • The two live validations show the discipline: Educate Girls reports internal reliability and explicitly presents no accuracy figure; APMS Phase 2 reports total-score reliability and adjudicated accuracy on disagreements, with the pairing and the tail named.

    KEEP READING

    Keep reading

    // COMMON.QUESTIONS

    Common questions

    Is reliability the same as accuracy?

    No. Reliability measures whether a scoring system is internally consistent and coherent across items, students, and graders. Accuracy measures whether scores are correct against a reference. A system can be reliable without being accurate; the two use different methods. Reliability is the floor; accuracy is the ceiling, measurable only once the floor holds and a defensible reference exists.

    What is Cronbach's alpha and why use it?

    It measures internal consistency across the items of a test, against conventional thresholds of 0.70 acceptable, 0.80 good, 0.90 excellent. We use it because it can be computed from a score matrix alone, so it can be reproduced and checked externally without access to our production system. The Educate Girls validation reports a mean alpha of 0.808 across 11 Hindi-medium tests.

    Why lead with reliability instead of a headline accuracy number?

    Because we report what the data supports, name the reference, draw the boundary of the claim, and ship the analysis script. A 0.808 alpha reproducible from a published matrix is stronger evidence than a high accuracy figure with no named reference. We would rather report a smaller defended number than a bigger decorative one.

    Have you ever reported an accuracy figure?

    Yes — in the APMS Phase 2 App-phase validation in Srikakulam, where we constructed an independent adjudication against the intended marking scheme as the reference. The result was AI 89.7% versus Teacher 82.8% on the disagreement set, with the pairing always shown together, the 215 mismatch cases broken down, and the 8.7% high-severity tail described openly. The detail is on the APMS case study page.

    Where can I see this discipline applied?

    On the two live case studies: Educate Girls for the Hindi handwriting reliability study (alpha 0.808 across 11 tests, reproducibility package on request), and APMS Phase 2 for the controlled reliability and adjudicated accuracy study (93% ICC, AI 89.7% and Teacher 82.8% adjudicated on mismatches, 8.7% tail reported and governed). Both name the boundary of what the study does and does not claim.

    How reliable is AI grading compared to human examiners?

    Reliability is measured with statistics like intraclass correlation (ICC) and Cronbach's alpha. In CrazyGoldFish's government-school validation, AI and teacher total scores agreed at 93% ICC; in a Hindi handwritten deployment, internal reliability was alpha = 0.808. On the cases where AI and teacher disagreed, an independent adjudicator favoured the AI more often — but the teacher remains the final authority on every score.

    What is the most accurate AI for grading descriptive answers?

    Ask for reliability before accuracy. A single accuracy percentage with no independent reference is not a trustworthy claim. CrazyGoldFish reports 93% ICC reliability from a government validation, and reports accuracy only against an independent adjudicator, naming the limit each time.

    See the discipline applied to your students

    If you run assessments at scale — a school system, a network, or a government programme — the same reliability discipline can be applied to your own scripts.

    // ABOUT.THE.AUTHOR

    Rahul Khandelwal

    Rahul Khandelwal is the founder and CEO of CrazyGoldFish, where he is building the evaluation infrastructure — the memory layer — for education. His focus is the hardest part of the Indian assessment stack: scoring handwritten, subjective, multilingual answers reliably and at scale, with the teacher as the final authority. Under his direction, CrazyGoldFish has run production validations with the NGO Educate Girls (handwritten Hindi answers across two states) and a government model school in Srikakulam, and its work has been recognised in the MeitY IndiaAI Impact Summit compendium and listed on AIKosh. Before founding CrazyGoldFish, Rahul spent two years at Pratham, one of India's largest education non-profits. He writes on evaluation as infrastructure, human-in-the-loop AI, and why reliability must come before accuracy in education AI.