RESEARCH · WHITE PAPER

    Every Mark Deserves a Referee

    The deterministic adjudicator: a referee that holds every mark, a teacher's or an AI's, to the same rule, every time.

    CrazyGoldFish · 28 September 2026 · ~9 min read

    SUMMARY

    Grade the same subjective answer twice and you can get two marks. Ask a second examiner and you can get a third. For a subjective answer there is no answer key, only a defensible range against the marking scheme. So the useful question is not whether a mark is accurate, but whether it follows the scheme the same way every time.

    This paper describes two components that make that possible. An invariant evaluator reads and scores an answer, and its reading barely changes across AI model families: less than 1% of a score's variance comes from which model marks it. A deterministic adjudicator then judges any mark, a teacher's or an AI's, against the board's own scheme through a fixed rule, and returns the same evidence-linked verdict on every run. It judges the mark, not the marker.

    We demonstrate both on real Class-XII answer scripts. We are precise about what that shows: conformance to a marking scheme, on a focused corpus, with a human as the final authority on every mark.

    The Adjudicator

    Get the next study when it drops, plus the full PDF and transparency package.

    AT A GLANCE

    The numbers, up front

    Under 1%

    of score variance from which AI model marks it · four model families

    82% / 18%

    disputed marks resolved with an evidence-linked verdict / escalated to a human

    15.2% vs 0.9%

    marks that move with no scheme vs criteria merely reordered

    No self-preference

    own model family removed · verdicts held (direction, not magnitude)

    56 · 9 · 5

    question-cases · Class-XII copies · exams

    Marked against a scheme derived from the board's published policy. This is conformance to a scheme, not accuracy against an absolute truth, and a human remains the final authority on every mark.

    1. There is no answer key

    For a subjective answer, a single correct score does not exist. There is a defensible range, judged against a marking scheme. Ask five fair examiners and you get a distribution, not a number. So "is the AI accurate?" has no reference point. Accurate against what?

    Two failure modes follow. First, agreement is not validity. Scoring an AI by how often it agrees with a teacher measures resemblance to a label that may itself be wrong. Second, and more dangerous, is the shared blind spot. When a teacher and an AI agree on a mark and both are wrong, no teacher-versus-AI comparison can see it. It shows up only when you read the answer one value point at a time against the scheme.

    The right question is not who is more accurate. It is who checks the evaluator. And the precondition for any answer is a reading that does not move. The same answer has to read the same way every time, before anyone can judge whether it is right.

    2. A stable reading comes first

    An invariant reading is not a correct reading, but it is a stable one, and you cannot settle a disagreement with a ruler that bends. If the reading moves with the model, there is nothing steady to judge against. The order is deliberate. A stable reading, then a referee, then a human with the final say.

    3. Two components: one we prove, one we teach

    We separate, on purpose, what we demonstrate from what we disclose.

    3a. The invariant evaluator: proven, not disclosed

    How our evaluator is built, the way the marking scheme is constructed and subject conventions are grounded so that independent models converge, is a confidential method. What we publish is not the recipe but the test, and the result. The test is one any institution can ask of any vendor. Run several independent AI model families on the same answer against the same scheme, and measure how much of the score's variance is explained by which model was used. For ours, that share is less than 1% (§4.1). An evaluator that passes this test is admissible as a reference. One that does not is another opinion. We prove ours passes. We do not publish how we made it pass, and a buyer does not need to know that, only to verify it.

    3b. The deterministic adjudicator: the method we teach

    On top of an admissible evaluator sits the referee, and this part we describe in full. The adjudicator is a fixed rule applied to the evaluator's frozen reading of a disputed mark, in five steps.

    1. Admissibility. Is there enough evidence to judge this mark at all?
    2. Defensibility. Does the mark fall within the band the marking scheme supports?
    3. Scheme-supported mark. Where it does not, what mark does the evidence earn?
    4. Evidence route. Every credit awarded points to a specific span in the student's answer.
    5. Escalation. Where the rule cannot decide, the case goes to a human, with the evidence attached.

    THE ADJUDICATOR · FIVE STEPS

    1. 01

      Admissibility

      Enough evidence to judge this mark?

    2. 02

      Defensibility

      Within the band the scheme supports?

    3. 03

      Scheme-supported mark

      What mark does the evidence earn?

    4. 04

      Evidence route

      Every credit points to a span

    5. 05

      Escalation

      Undecidable cases go to a human

    Same frozen reading in, same evidence-linked verdict out, on every run.

    The rule is deterministic. Given the same reading, it returns the same evidence-linked verdict on every run. It is a rule both sides already answer to, the board's scheme, not another opinion, and it judges the mark, not the marker. Anyone can build one. What they need first is an evaluator invariant enough to feed it.

    4. Results

    The adjudicator's job is to make a mark defensible, to say, on the evidence and against the scheme, what mark holds, the same way every time. The results below show it doing that, and show that the evaluator beneath it is stable enough to serve as the reference.

    Corpus: 56 question-cases from 9 Class-XII answer copies across 5 exams, and 35 value-point cells for the invariance analysis. English-medium. Marked against a scheme derived from the board's published marking policy.

    4.1 The reading barely depends on the model. Across four model families from four labs, less than 1% of a score's variance is attributable to which model marks it. About 80% comes from the answer itself. Panel generalizability is 0.95; a single model alone reaches about 0.80. The reliability figure reproduces three independent ways to within 0.01. Cross-family agreement (Gwet AC2) is 0.768, reported as a band of about 0.67 to 0.77. The 0.162 spread is between-cell variation, not a confidence interval. Per exam it ranges from 0.816 in Business Studies A to 0.699 in Accountancy B, weakest where the ledgers are hardest, holding across every subject. On the Landis and Koch scale, of 35 cells, 15 are almost-perfect, 12 substantial, 8 moderate, 0 slight. This is reproducibility, not a claim of correctness.

    4.2 The referee resolves 82%, and sends 18% to a human. Run on all 56 cases, the cascade returns a stated verdict on 46 of them, and refers the hardest 10 to a human with the evidence attached. Among the 46 resolved: 24 where both markers are defensible, 9 where the AI is closer, 5 where the teacher is closer, 8 where neither is defensible. The 82% reproduces on every run. A human's verdicts do not. An evidence-integrity guardrail fired zero times.

    VERDICTS ON 56 CASES

    • Both defensible24
    • AI closer9
    • Teacher closer5
    • Neither defensible8
    • Escalated to a human10
    46 resolved (82%) · 10 escalated with evidence (18%).

    4.3 It follows the scheme, not cosmetics. Shuffle the order of the marking criteria and 0.9% of value-point cells move by half a mark or more. Remove the marking scheme and let the model judge freely, and 15.2% move. The scheme structure is doing the work, not the model's taste. A wrong-curriculum control moved 1.1%, which we report as an honest null, not a causal claim.

    4.4 No self-preference. The fair concern is an AI panel favouring answers from its own model family. We removed our production model family from the panel and re-ran the study. The verdicts held. We report this as a direction, no self-preference, not as a magnitude.

    4.5 The shared blind spot, made visible. In one Accountancy case, a question required eight value-point entries, and two were never attempted. The teacher and the AI each awarded full marks, 4 out of 4. Reading value point by value point, the panel's expected credit was 2.69, a defensible 3 out of 4, catching two missing entries neither grader had docked. This is the error that is invisible to any teacher-versus-AI method, because both agreed.

    4.6 Evidence anchoring. About 82% of awarded credits are span-verified: each earned or partial credit points to a span in the student's answer. We state this as every credit being evidence-anchored, not as auditability being proven. A span existing is necessary but not sufficient for that span to justify the credit, and the definitive test is a human reading of the page image, which is a next step.

    5. The boundary we draw, deliberately

    Precision about scope is part of the work, not an apology for it.

    What this establishes: a reading that is reproducible across models, a referee that is deterministic and evidence-linked, the ability to catch jointly-wrong marks that pairwise checks miss, and no self-preference. What it is by design: a measure of conformance to a marking scheme, because for a subjective answer there is no absolute truth to be accurate against, and a demonstration of mechanisms on a focused corpus, built to scale rather than to assert a national rate.

    Larger-N validation and a human-expert study are the defined next steps. We state the roadmap plainly, because a claim you can check is worth more than one you cannot.

    6. Judgment stays human

    Reproducibility earns trust only if a person can still overrule it. The referee's most important behaviour is the escape hatch. When the rule cannot decide, it stops and escalates to a human, with the answer, the scheme, and the reasoning attached. Nothing changes a student's mark without a person signing it. This is the architecture, not an add-on: the AI scores first, and a human reviews before a result reaches the student. Review, override, audit trail, and approval states are built in. The companion white paper on the evaluation-governance layer covers this in full.

    7. Reproducibility

    The adjudicator's rule reproduces byte for byte on a frozen reading. The panel's reading reproduces to within the measured variance in §4.1. Stating both tiers precisely is part of the design, so a reader can see exactly what is fixed and what varies. The analysis is releasable as anonymised, DPDP-compliant score matrices with the analysis scripts, under our 60-day open-data transparency protocol.

    About CrazyGoldFish

    CrazyGoldFish builds evaluation infrastructure for education. It turns handwritten and typed student answers into structured, auditable learning signals, with human-in-the-loop control built in. Its evaluation work is featured in MeitY's IndiaAI Impact Summit 2026 Compendium and listed in the IndiaAI AIKosh use-case repository.

    The findings here are demonstrated on a focused, English-medium corpus of real Class-XII scripts. They establish conformance to a marking scheme, and a human remains the final authority on every mark. The underlying study is a working paper, being prepared for submission to a peer-reviewed journal. Correspondence: research@crazygoldfish.com.

    KEY TAKEAWAYS

    Key takeaways

    • →For a subjective answer there is no answer key, only a defensible range against the scheme
    • →A stable reading must come first; a ruler that bends settles nothing
    • →The adjudicator judges the mark, not the marker, and returns the same verdict every run
    • →Jointly-wrong marks (teacher and AI agree and both are wrong) are invisible to any teacher-vs-AI check
    • →A human is the final authority on every mark

    KEEP READING

    Keep reading

    The Adjudicator

    Get the next study when it drops, plus the full PDF and transparency package.

    // COMMON.QUESTIONS

    Who checks the evaluator in AI grading?

    A deterministic adjudicator: a fixed rule that judges any mark, a teacher's or an AI's, against the board's own marking scheme and returns the same evidence-linked verdict every run, with a human as the final authority.

    Is AI grading reproducible?

    In this study, less than 1% of a score's variance came from which AI model marked the answer, across four model families. This is reproducibility, a stable reading, not a claim of accuracy against an absolute truth.

    What is a deterministic adjudicator?

    A rule applied to a frozen reading of a disputed mark, in five steps (admissibility, defensibility, scheme-supported mark, evidence route, escalation to a human), that returns the same verdict on every run. It judges the mark, not the marker.

    Does AI replace teachers in evaluation?

    No. The AI scores first, the adjudicator judges against the scheme, and a human reviews with full evidence before any result reaches a student. A human is the final authority on every mark.

    See the referee on your own scripts