SUMMARY
Grade the same subjective answer twice and you can get two marks. Ask a second examiner and you can get a third. For a subjective answer there is no answer key, only a defensible range against the marking scheme. So the useful question is not whether a mark is accurate, but whether it follows the scheme the same way every time.
This paper describes two components that make that possible. An invariant evaluator reads and scores an answer, and its reading barely changes across AI model families: less than 1% of a score's variance comes from which model marks it. A deterministic adjudicator then judges any mark, a teacher's or an AI's, against the board's own scheme through a fixed rule, and returns the same evidence-linked verdict on every run. It judges the mark, not the marker.
We demonstrate both on real Class-XII answer scripts. We are precise about what that shows: conformance to a marking scheme, on a focused corpus, with a human as the final authority on every mark.
The Adjudicator
Get the next study when it drops, plus the full PDF and transparency package.
AT A GLANCE
The numbers, up front
Under 1%
of score variance from which AI model marks it · four model families
82% / 18%
disputed marks resolved with an evidence-linked verdict / escalated to a human
15.2% vs 0.9%
marks that move with no scheme vs criteria merely reordered
No self-preference
own model family removed · verdicts held (direction, not magnitude)
56 · 9 · 5
question-cases · Class-XII copies · exams
Marked against a scheme derived from the board's published policy. This is conformance to a scheme, not accuracy against an absolute truth, and a human remains the final authority on every mark.
1. There is no answer key
For a subjective answer, a single correct score does not exist. There is a defensible range, judged against a marking scheme. Ask five fair examiners and you get a distribution, not a number. So "is the AI accurate?" has no reference point. Accurate against what?
Two failure modes follow. First, agreement is not validity. Scoring an AI by how often it agrees with a teacher measures resemblance to a label that may itself be wrong. Second, and more dangerous, is the shared blind spot. When a teacher and an AI agree on a mark and both are wrong, no teacher-versus-AI comparison can see it. It shows up only when you read the answer one value point at a time against the scheme.
The right question is not who is more accurate. It is who checks the evaluator. And the precondition for any answer is a reading that does not move. The same answer has to read the same way every time, before anyone can judge whether it is right.
2. A stable reading comes first
An invariant reading is not a correct reading, but it is a stable one, and you cannot settle a disagreement with a ruler that bends. If the reading moves with the model, there is nothing steady to judge against. The order is deliberate. A stable reading, then a referee, then a human with the final say.
3. Two components: one we prove, one we teach
We separate, on purpose, what we demonstrate from what we disclose.
3a. The invariant evaluator: proven, not disclosed
How our evaluator is built, the way the marking scheme is constructed and subject conventions are grounded so that independent models converge, is a confidential method. What we publish is not the recipe but the test, and the result. The test is one any institution can ask of any vendor. Run several independent AI model families on the same answer against the same scheme, and measure how much of the score's variance is explained by which model was used. For ours, that share is less than 1% (§4.1). An evaluator that passes this test is admissible as a reference. One that does not is another opinion. We prove ours passes. We do not publish how we made it pass, and a buyer does not need to know that, only to verify it.
3b. The deterministic adjudicator: the method we teach
On top of an admissible evaluator sits the referee, and this part we describe in full. The adjudicator is a fixed rule applied to the evaluator's frozen reading of a disputed mark, in five steps.
- Admissibility. Is there enough evidence to judge this mark at all?
- Defensibility. Does the mark fall within the band the marking scheme supports?
- Scheme-supported mark. Where it does not, what mark does the evidence earn?
- Evidence route. Every credit awarded points to a specific span in the student's answer.
- Escalation. Where the rule cannot decide, the case goes to a human, with the evidence attached.
THE ADJUDICATOR · FIVE STEPS
01
Admissibility
Enough evidence to judge this mark?
02
Defensibility
Within the band the scheme supports?
03
Scheme-supported mark
What mark does the evidence earn?
04
Evidence route
Every credit points to a span
05
Escalation
Undecidable cases go to a human
The rule is deterministic. Given the same reading, it returns the same evidence-linked verdict on every run. It is a rule both sides already answer to, the board's scheme, not another opinion, and it judges the mark, not the marker. Anyone can build one. What they need first is an evaluator invariant enough to feed it.
4. Results
The adjudicator's job is to make a mark defensible, to say, on the evidence and against the scheme, what mark holds, the same way every time. The results below show it doing that, and show that the evaluator beneath it is stable enough to serve as the reference.
Corpus: 56 question-cases from 9 Class-XII answer copies across 5 exams, and 35 value-point cells for the invariance analysis. English-medium. Marked against a scheme derived from the board's published marking policy.
4.1 The reading barely depends on the model. Across four model families from four labs, less than 1% of a score's variance is attributable to which model marks it. About 80% comes from the answer itself. Panel generalizability is 0.95; a single model alone reaches about 0.80. The reliability figure reproduces three independent ways to within 0.01. Cross-family agreement (Gwet AC2) is 0.768, reported as a band of about 0.67 to 0.77. The 0.162 spread is between-cell variation, not a confidence interval. Per exam it ranges from 0.816 in Business Studies A to 0.699 in Accountancy B, weakest where the ledgers are hardest, holding across every subject. On the Landis and Koch scale, of 35 cells, 15 are almost-perfect, 12 substantial, 8 moderate, 0 slight. This is reproducibility, not a claim of correctness.
4.2 The referee resolves 82%, and sends 18% to a human. Run on all 56 cases, the cascade returns a stated verdict on 46 of them, and refers the hardest 10 to a human with the evidence attached. Among the 46 resolved: 24 where both markers are defensible, 9 where the AI is closer, 5 where the teacher is closer, 8 where neither is defensible. The 82% reproduces on every run. A human's verdicts do not. An evidence-integrity guardrail fired zero times.
VERDICTS ON 56 CASES
- Both defensible24
- AI closer9
- Teacher closer5
- Neither defensible8
- Escalated to a human10
4.3 It follows the scheme, not cosmetics. Shuffle the order of the marking criteria and 0.9% of value-point cells move by half a mark or more. Remove the marking scheme and let the model judge freely, and 15.2% move. The scheme structure is doing the work, not the model's taste. A wrong-curriculum control moved 1.1%, which we report as an honest null, not a causal claim.
4.4 No self-preference. The fair concern is an AI panel favouring answers from its own model family. We removed our production model family from the panel and re-ran the study. The verdicts held. We report this as a direction, no self-preference, not as a magnitude.
4.5 The shared blind spot, made visible. In one Accountancy case, a question required eight value-point entries, and two were never attempted. The teacher and the AI each awarded full marks, 4 out of 4. Reading value point by value point, the panel's expected credit was 2.69, a defensible 3 out of 4, catching two missing entries neither grader had docked. This is the error that is invisible to any teacher-versus-AI method, because both agreed.
4.6 Evidence anchoring. About 82% of awarded credits are span-verified: each earned or partial credit points to a span in the student's answer. We state this as every credit being evidence-anchored, not as auditability being proven. A span existing is necessary but not sufficient for that span to justify the credit, and the definitive test is a human reading of the page image, which is a next step.
5. The boundary we draw, deliberately
Precision about scope is part of the work, not an apology for it.
What this establishes: a reading that is reproducible across models, a referee that is deterministic and evidence-linked, the ability to catch jointly-wrong marks that pairwise checks miss, and no self-preference. What it is by design: a measure of conformance to a marking scheme, because for a subjective answer there is no absolute truth to be accurate against, and a demonstration of mechanisms on a focused corpus, built to scale rather than to assert a national rate.
Larger-N validation and a human-expert study are the defined next steps. We state the roadmap plainly, because a claim you can check is worth more than one you cannot.
6. Judgment stays human
Reproducibility earns trust only if a person can still overrule it. The referee's most important behaviour is the escape hatch. When the rule cannot decide, it stops and escalates to a human, with the answer, the scheme, and the reasoning attached. Nothing changes a student's mark without a person signing it. This is the architecture, not an add-on: the AI scores first, and a human reviews before a result reaches the student. Review, override, audit trail, and approval states are built in. The companion white paper on the evaluation-governance layer covers this in full.
7. Reproducibility
The adjudicator's rule reproduces byte for byte on a frozen reading. The panel's reading reproduces to within the measured variance in §4.1. Stating both tiers precisely is part of the design, so a reader can see exactly what is fixed and what varies. The analysis is releasable as anonymised, DPDP-compliant score matrices with the analysis scripts, under our 60-day open-data transparency protocol.
About CrazyGoldFish
CrazyGoldFish builds evaluation infrastructure for education. It turns handwritten and typed student answers into structured, auditable learning signals, with human-in-the-loop control built in. Its evaluation work is featured in MeitY's IndiaAI Impact Summit 2026 Compendium and listed in the IndiaAI AIKosh use-case repository.
The findings here are demonstrated on a focused, English-medium corpus of real Class-XII scripts. They establish conformance to a marking scheme, and a human remains the final authority on every mark. The underlying study is a working paper, being prepared for submission to a peer-reviewed journal. Correspondence: research@crazygoldfish.com.
KEY TAKEAWAYS
Key takeaways
- →For a subjective answer there is no answer key, only a defensible range against the scheme
- →A stable reading must come first; a ruler that bends settles nothing
- →The adjudicator judges the mark, not the marker, and returns the same verdict every run
- →Jointly-wrong marks (teacher and AI agree and both are wrong) are invisible to any teacher-vs-AI check
- →A human is the final authority on every mark
KEEP READING
Keep reading
The Adjudicator
Get the next study when it drops, plus the full PDF and transparency package.