RESEARCH · OVERVIEW

    Who checks the evaluator?

    When an AI grades your exam, who checks the AI? An invariant evaluator and a deterministic adjudicator, with a human as the final authority on every mark.

    CrazyGoldFish · 28 September 2026 · ~6 min read

    THE SHORT ANSWER

    When an AI grades a subjective answer, the honest question is not whether it is accurate, because there is no answer key, only a defensible range against the marking scheme. It is whether the mark is reproducible and checkable. CrazyGoldFish answers it with two things: an invariant evaluator that reads an answer the same way whichever model reads it, and a deterministic adjudicator that judges any mark against the board's own scheme through a fixed rule, with a human as the final authority.

    The Adjudicator

    Get the next study when it drops, plus the full PDF and transparency package.

    AT A GLANCE

    The numbers, up front

    Under 1%

    of score variance from which AI model marks it · four model families

    82% · 18%

    disputed marks resolved with an evidence-linked verdict · escalated to a human

    15.2% vs 0.9%

    marks that move with no scheme vs criteria merely reordered

    No self-preference

    direction, stated without a magnitude

    56 · 9 · 5

    question-cases · Class-XII copies · exams

    A second opinion is not enough

    Grade the same subjective answer twice and you can get two marks. Ask a second examiner and you can get a third.

    Two failure modes follow. First, agreement is not validity. Scoring an AI by how often it matches a teacher measures resemblance to a label that may itself be wrong. Second, the shared blind spot. When a teacher and an AI agree and both are wrong, no teacher-versus-AI comparison can catch it, because there is nothing between them to flag.

    This was studied in a government-school pilot featured in MeitY's IndiaAI Impact Summit 2026 Compendium.

    WHERE A SCORE'S VARIANCE COMES FROM

    • ~80%The answer itself
    • ~19%Question and interactions
    • <1%Which model (four families)
    Reproducibility across model families, not a claim of accuracy.

    The invariant evaluator: a reading that does not move

    The same answer is read the same way whichever model reads it. Under 1% of a score's variance comes from which AI model marks it, across four model families, while about 80% comes from the answer itself.

    This is a stable reading, not a claim of accuracy. It is the fixed picture a referee needs. Read more in reliability before accuracy.

    The deterministic adjudicator: a referee, not a smarter grader

    The adjudicator judges any mark, a teacher's or an AI's, against the board's own scheme through a fixed rule. It returns the same evidence-linked verdict every run, resolves 82% of disputed marks, and escalates the hardest 18% to a person with the evidence attached.

    It judges the mark, not the marker. Read more in the three-tier evaluation architecture.

    56 DISPUTED MARKS · WHAT THE ADJUDICATOR RETURNS

    • 46 (82%) Resolved with an evidence-linked verdict
    • 10 (18%) Escalated to a human

    THE 46 RESOLVED

    • 24 Both defensible
    • 9 AI closer
    • 5 Teacher closer
    • 8 Neither defensible

    Why it matters at scale

    Re-evaluation queues touch only challenged papers. Second-marker systems touch only disagreements. Both leave the quietly-agreed marks untouched, which is exactly where a shared mistake sits.

    The field evidence is in the APMS Phase 2 validation (Srikakulam), and in Educate Girls, the Hindi-medium study.

    An invariant, evidence-backed evaluator makes every mark checkable against a rule both sides already answer to, and turns "trust us" into "here is the test, run it yourself, including on us."

    For a board, a district, or a national programme, see evaluation for institutions and public programs.

    What we publish, and what we hold back

    We publish the result and the test, not the mechanism or the registry behind it. An institution does not need the recipe; it needs to run the test on any vendor, including us, and see who passes.

    Two things always hold: conformance to a marking scheme, not accuracy against an absolute truth, and a human as the final authority on every mark.

    KEY TAKEAWAYS

    Key takeaways

    • →For a subjective answer there is no answer key, only a defensible range against the scheme
    • →A stable reading must come first; a ruler that bends settles nothing
    • →The adjudicator judges the mark, not the marker, and returns the same verdict every run
    • →Jointly-wrong marks, where a teacher and an AI agree and both are wrong, are invisible to any teacher-vs-AI check
    • →A human is the final authority on every mark

    KEEP READING

    Keep reading

    The Adjudicator

    Get the next study when it drops, plus the full PDF and transparency package.

    // COMMON.QUESTIONS

    Who checks the evaluator?

    When an AI grades a subjective answer, a second human grader is not enough, because two readers can agree and both be wrong. CrazyGoldFish uses a deterministic adjudicator: a fixed rule that judges any mark against the board's own marking scheme, returns the same evidence-linked verdict every run, and escalates the hardest cases to a person who has the final say.

    Is AI grading reliable?

    Reliability is the right test for subjective answers, because there is no single correct score, only a defensible range against the scheme. In a Class-XII study, less than 1% of a score's variance came from which AI model marked it, across four model families. That is reproducibility, measured on a focused corpus, not a claim of absolute accuracy.

    What is an invariant evaluator?

    An invariant evaluator reads and scores an answer so the reading barely changes across different AI models. In this study under 1% of a score's variance came from which model marked it, while about 80% came from the answer itself. It is a stable reading, the fixed picture a referee needs before it can judge a disputed mark.

    Is AI more accurate than a teacher at grading?

    That is the wrong question. For a subjective answer there is no answer key to be accurate against, only a defensible range against the marking scheme. Both a teacher and an AI can be right, or both wrong in the same place. The useful test is conformance: does the mark follow the scheme the same way every time, and can it show you why.

    How do you catch a grading mistake two readers both miss?

    By reading the answer one value point at a time against the marking scheme instead of judging it whole. In one real Class-XII case a question needed eight value-point entries, two were never attempted, and both the teacher and the AI still awarded full marks. The value-point reading gave a defensible three out of four, catching what agreement had hidden.

    See the referee on your own scripts