RESEARCH · CONCEPT

    The three-tier evaluation architecture: production, audit, adjudication

    Three operational layers that turn AI-assisted evaluation from a vendor claim into a governed decision-support layer.

    ONE GOVERNED SYSTEM

    3 tiers · 1 teacher as final authority

    Rahul Khandelwal — Founder & CEO, CrazyGoldFish · June 2026 · 12 min read

    // AT.A.GLANCE

    TierWhat it isWhat anchors it
    Tier 1 · ProductionDaily-scale evaluation in every classroom — AI-assisted scoring, teacher as final authority, HITL built into the API, not a paid add-on.APMS App phase: 95.8% OCR success, 99.6% mapping success, ~80% teacher time saved.
    Tier 2 · AuditContinuous, sampled, governed review — hotspot routing, outlier escalation, monthly audit pack, quarterly review.APMS Policy Annex: monthly audit pack, 100% outlier-copy review, escalation ladder to district QA.
    Tier 3 · AdjudicationIndependent validation against the intended marking scheme — the methodology that makes a reliability claim defensible.APMS Phase 2: AI 89.7% vs Teacher 82.8% adjudicated correctness, paired, on the 215-case mismatch set.
    "The three tiers are not three products. They are three layers of one system, and a serious AI evaluation deployment runs all three. Tier 1 is what makes it work day to day. Tier 2 is what makes it infrastructure rather than a tool. Tier 3 is what makes a reliability claim citable rather than declared."

    IN SHORT

    CrazyGoldFish operates AI-assisted evaluation as three tiers, not one. Tier 1, Production, runs every day at classroom scale, with the teacher as final authority and HITL built into the API rather than sold as a paid feature. Tier 2, Audit, is the continuous sampled review that catches drift before it compounds: hotspot routing, outlier escalation, a monthly audit pack, a quarterly review, and a defined escalation ladder. Tier 3, Adjudication, is the periodic validation against an independent reference — the move that turns "the AI agrees with the teacher" into a defensible reliability claim. The two live case studies on this site — Educate Girls in Hindi and APMS Phase 2 in Srikakulam — are evidence of the tiers running. This page names the architecture underneath.

    The three layers

    AI-assisted evaluation is only trustworthy at scale when it is built as three coupled layers. The Production layer scores at daily volume. The Audit layer watches that scoring continuously and catches drift before it compounds. The Adjudication layer validates the whole system periodically against an independent reference. Run only the production layer and you have a tool: it grades papers, but you cannot trust what it graded six months later, at scale, when the cohort has changed, the question paper is new, and the model has been swapped twice. Add audit and adjudication, governed and on a cadence, and it becomes infrastructure. Our product claim is the three tiers together.

    Each layer carries a different kind of cost, which is why all three together is the standard we hold ourselves to. Production requires a working pipeline — hard, but bounded. Audit requires continuous instrumentation and ongoing review time; it never produces a hero number for a deck, but it is what proves the day-to-day grading is still doing what it did six months ago. Drift in a grading model looks small from inside the platform and large at the level of a student transcript. Adjudication requires an independent reference and a published methodology; run once it is a snapshot, run on a cadence it is a guarantee. We treat all three as non-negotiable, because that is what evaluation infrastructure means.

    Tier 1 · Production

    The Production tier runs every day, in every classroom that uses the platform. A copy enters through an upload. Handwritten answers are digitised through OCR. The extracted text is mapped to the question it belongs to. The AI suggests a rubric-aligned score and a per-question feedback note. The teacher reads, accepts, or overrides. The final score is published.

    Two design commitments sit underneath. First, the teacher remains the final authority. Overwrite and edit functions stay active on every score, every day, in every deployment — there is no mode where the AI publishes a grade the teacher cannot revisit, and any override is captured into the audit trail as an improvement signal. "AI-assisted" is exact: the AI assists, the teacher decides. Second, human-in-the-loop is built into the API, not sold as a paid feature. Review, override, audit trail, and approval states ship with the standard endpoints; a customer who buys the K12 API gets the governed workflow by default. We do not put governance behind a tier upgrade.

    In the APMS Phase 2 App-phase deployment in Srikakulam, OCR succeeded on 95.8% of 801 questions (34 routed for review) and question-mapping on 99.6% (3 routed). Teacher evaluation time fell by roughly 80% in the App phase, building on the up-to-60% reduction measured in the earlier Web-App phase at the same school, published in the IndiaAI Impact Summit 2026 Compendium (MeitY, Case Study 8). The two phase figures always carry their phase qualifier. In the Educate Girls Hindi-medium deployment — 11 tests across Madhya Pradesh and Rajasthan, seven subjects, 162 questions, 2,297 student-question pairs — the same Production pipeline produced a mean internal reliability (Cronbach's α) of 0.808, median 0.817, with 10 of 11 tests clearing the conventional acceptable threshold.

    Every Production decision creates an entry in the audit trail. The trail is not an export or an opt-in feature — it is the substrate that Tier 2 reads.

    Tier 2 · Audit

    The Audit tier is what continuous trust looks like, and it is the layer we treat as the difference between a tool and infrastructure. It samples, reviews, and routes on a cadence, watching the signals a careful examination authority watches when it standardises a board exam: mismatch-prone question types, outlier graders, drift across sittings, high-severity score deviations.

    In the APMS Phase 2 deployment, the audit layer runs through a specific set of controls. Hotspot routing: the system maintains a list of question types that have historically produced AI-teacher disagreement, and answers to those are routed automatically to teacher review; the list updates as new patterns surface. Outlier teacher detection: evaluator-to-evaluator deviation is tracked within a sitting, and a teacher whose scoring drifts from the cohort over a long sitting is flagged for principal review. Outlier copy review: every copy in the high-severity tail — defined as a gap of at least the greater of six marks or 10% of the total between AI and teacher scores — is routed for 100% review. The high-severity tail in the APMS App phase was 8.7% of copies (95% CI 0.0%–21.7%); we report the tail openly because governing it is the point. Monthly audit pack: the Policy Annex commits to a monthly pack — the month's mismatch rate, the question-level concentration of mismatches, the outlier-teacher signals, the tail size, the actions taken — and a quarterly review asking whether the marking scheme, teacher calibration, or AI model needs adjustment. Escalation ladder: a defined route up — teacher to principal to block QA to district QA — with a 72-hour provisional window, a five-working-day query SLA, and a two-working-day escalation clock, auditable end to end.

    This is not a logging system. A logging system records what happened; the audit tier responds to it, on a cadence, with a defined route. A tool grades. Infrastructure grades, watches, flags, routes, and writes the trail the next conversation rests on.

    Tier 3 · Adjudication

    The Adjudication tier is what runs when the reliability of the system has to be defended on its merits. When AI and teacher scores disagree, the easy assumption is that the teacher is right; the Adjudication tier does not make it. Every disagreement is routed to an independent adjudicator working from the intended marking scheme, and the adjudication — not the AI mark, not the teacher mark — is treated as the reference. We call this Adjudication-as-Ground-Truth (AGT).

    The APMS Phase 2 App-phase study is where the methodology was first reported externally. On the 215 cases where AI and teacher scores disagreed, adjudication found AI was closer 60.9% of the time, the teacher 32.6%, and both acceptable in 6.5%. Translated into adjudicated correctness, AI scored 89.7% on the mismatch set against the teacher's 82.8%. The two numbers always travel together — reading 89.7% without the 82.8% beside it is misreading the study.

    This is what makes a reliability claim defensible. Without an independent reference, "AI agrees with the teacher 95% of the time" is circular: the teacher is treated as ground truth to compute the agreement, with nothing left to check the teacher against. Adjudication breaks that circle. It runs on a cadence — per study, per major model release, per new deployment context. The APMS Phase 2 intraclass correlation of 93% (95% CI 0.805–0.976) and mean absolute error of 3.91 marks (CI 2.61–5.41) rest on the Adjudication tier having been run. The Educate Girls Hindi α of 0.808 is a reliability claim of a different shape — internal consistency, not adjudicated correctness — and that study is explicit that no controlled accuracy adjudication was run alongside it. The discipline shows up in what each study is permitted to claim.

    The Adjudication-as-Ground-Truth methodology is applied in the APMS Phase 2 validation, and the anonymised dataset it was computed from is published on AIKosh, so the analysis can be re-run independently. The APMS Phase 2 anonymised dataset — the score matrix and adjudication record from which the published numbers were computed — is hosted on AIKosh, the Government of India's national AI platform, with a 5-of-5 Data Quality Score (Beta). Any reader can re-run the analysis and arrive at the same numbers.

    How the tiers fit together

    The three tiers are a coupled loop, and the audit trail is the substrate that links them. Tier 1 produces evidence — every Production decision lands as a row in the trail. Tier 2 reads Tier 1 continuously, flags hotspots, surfaces outlier teachers, escalates high-severity copies, and feeds two streams back: a refreshed hotspot list and a calibration signal for teachers and the marking scheme. Tier 3 validates Tier 1 and Tier 2 periodically, recalibrates the reliability claims the lower tiers rest on, and feeds the mismatch concentration back into the next model release and the next Tier 2 sampling strategy. The loop closes there.

    The teacher is the final authority at every layer, not only Tier 1. A Tier 2 escalation adds a review; it does not remove the teacher's mark. A Tier 3 adjudication computes a reliability statement against a reference; it does not overwrite the teacher's mark. AI-assisted means AI-assisted at every tier, with the human decision protected throughout. This architecture sits inside the broader Evaluation Infrastructure category — the system layer where evaluation becomes a control point for capturing learning signals as persistent memory. The category framing lives at /evaluation-infrastructure; this page is the operational drill-down inside it.

    What this means for buyers

    If you are evaluating an AI grading vendor, the three tiers are a useful diligence frame — questions worth asking any vendor, ourselves included. A complete three-tier system answers all of them:

    • Continuous audit, or only one-time validation? A one-time report is a Tier 3 snapshot, not governance. Ask for the cadence, the audit pack, and the escalation ladder.
    • Where is the published adjudication methodology? A serious Tier 3 has a methodology a reader can find, with the reference defined. "We compared with teachers" invites the follow-up: which teachers, against what reference, and how was teacher noise controlled?
    • Where is the dataset, so the numbers can be re-run? A reliability claim is only as defensible as the matrix it rests on. The APMS Phase 2 anonymised dataset is on AIKosh; the Educate Girls reproducibility package is available on request with the analysis script.
    • Is HITL in the API, or a paid feature? Governance behind a paywall is governance you do not have at scale. Ask whether review, override, and audit trail ship with the standard endpoints.
    • What does the audit trail do end to end? A logging system is not an audit trail. Ask what feeds back, on what cadence, into what decision.
    • What is the high-severity tail, and how is it governed? Every grader has a tail. The APMS App phase reports an 8.7% tail openly, with the governance workflow described. That is the standard we hold.

    KEY TAKEAWAYS

    Key takeaways

    • AI evaluation runs as three tiers, not one: Tier 1 daily Production, Tier 2 continuous Audit, Tier 3 periodic Adjudication against an independent reference.
    • Production alone is a tool; three governed layers are infrastructure. Our product claim is the three together.
    • Production scales daily volume — APMS App phase: 95.8% OCR, 99.6% mapping, ~80% teacher time saved — with the teacher as final authority and HITL in the API.
    • Audit catches drift before it compounds: hotspot routing, outlier-teacher detection, 100% outlier-copy review, monthly audit pack, quarterly review, escalation ladder.
    • Adjudication makes reliability claims defensible: APMS Phase 2, AI 89.7% vs Teacher 82.8% on the 215-case mismatch set, paired, with the anonymised dataset on AIKosh for independent re-analysis.
    • The audit trail couples the three tiers; the teacher is the final authority at every tier.

    KEEP READING

    Keep reading

    // COMMON.QUESTIONS

    Common questions

    Is the three-tier architecture proprietary to CrazyGoldFish?

    It is our framing of how a governed AI evaluation system has to be built to be trustworthy at scale, and the structure we operate. The component disciplines — daily HITL, continuous audit, independent adjudication — are how serious psychometric work has been done for decades. What is specific to us is naming the three as a coupled architecture, building the audit trail as the substrate that links them, and publishing the dataset under each Tier 3 cycle so the claim is checkable.

    Do customers see all three tiers?

    Tier 1 is the daily customer surface — every customer on the K12 API operates it with HITL governance, override, and the audit trail by default. Tier 2 is enabled per deployment based on the audit cadence the customer commits to. Tier 3 runs as a validation cycle, per study or per major release, with the methodology and dataset published.

    How is Tier 2 different from a logging system?

    A logging system records what happened. The Audit tier responds to it — on a cadence, with a defined route, a monthly audit pack, a quarterly review, and an escalation ladder that puts every high-severity copy in front of a teacher and, where warranted, a principal or district QA reviewer.

    Is the AI better than the teacher according to Tier 3?

    On the cases where AI and teacher disagreed in the APMS Phase 2 App phase, independent adjudication found AI was closer more often: AI 89.7% and Teacher 82.8% on the 215-case mismatch set, with AI the closer reader in 60.9% of cases, the teacher in 32.6%, both acceptable in 6.5%. We hold the finding carefully — it is always a statement about the mismatch set, never that AI grades better than teachers in general. The teacher holds final authority across all three tiers.

    Where can I see the methodology in practice?

    The Adjudication-as-Ground-Truth methodology is applied in the APMS Phase 2 validation — see the case study at /research/case-studies/apms-phase-2-validation — and the underlying anonymised dataset is published on AIKosh, so the numbers can be re-run.

    See the three tiers running on your own assessments

    If you run assessments at scale — a school system, a network, or a government programme — the three-tier architecture can be configured to your context across Production, Audit, and Adjudication.