RESEARCH · CONCEPT

    Audit as infrastructure: the standard we hold for continuous evaluation

    What continuous evaluation audit actually does, what the monthly pack contains, and why it is the layer between a scoring pipeline and a governed system.

    THE AUDIT CADENCE

    Monthly pack · quarterly review · drift caught

    Rahul Khandelwal — Founder & CEO, CrazyGoldFish · June 2026 · 11 min read

    // AT.A.GLANCE

    8.7%APMS Phase 2 high-severity tail (CI 0.0%–21.7%)what the audit catches openly
    Monthly · Quarterlyaudit pack cadence with quarterly reviewfrom the APMS Policy Annex
    Teacher → Block QA → District QAthe escalation ladder in the deploymentgovernance, not just data
    "A scoring pipeline that produces a number is a tool. A pipeline that also watches its own outputs continuously, surfaces drift before it compounds, and publishes a monthly pack a quality team can read, is a governed system. That is the standard we hold, ourselves included."

    IN SHORT

    CrazyGoldFish treats audit as a continuous tier inside its three-tier evaluation architecture, sitting above Production and below Adjudication. The audit layer samples four surfaces (hotspot question routing, outlier teacher detection, 100% review of every high-severity outlier copy, and stratified continuous sampling for drift), publishes a monthly pack with six standard sections, escalates through a defined ladder, and feeds findings back into Production routing. The Policy Annex from the APMS Phase 2 validation in Srikakulam is where this framework first ran end to end. The methodology is shipped; the platform-wide pipeline is in build.

    Why continuous audit is part of the architecture

    A scoring pipeline on its own is one layer: scripts come in, an AI suggests scores, a teacher reviews, the day moves on. Call it Tier 1. Without continuous instrumentation alongside it, no one has to ask, three months in, whether the scoring is still doing what it did on launch day. Drift in a grading model looks small from inside the platform and large at the level of a student transcript. The layer that catches drift before it compounds is the one we treat as load-bearing.

    The discipline has names in adjacent fields. In financial audit it is the continuous controls testing between the daily ledger and the annual external opinion; in assessment it is the continuous psychometric monitoring a serious board examination runs across sittings. A daily-scale system produces evidence; a continuous-scale system reads that evidence, on a cadence, and responds. Two things are easy to mistake for an audit layer: compliance sampling (a random subset signed off — a useful snapshot, not drift detection) and a one-time pilot validation report (a periodic snapshot, not continuous monitoring). Neither substitutes for a continuous audit tier. The broader three-tier shape is at /research/concepts/three-tier-evaluation-architecture; this page is the deeper drill on the audit tier.

    What gets sampled, and how

    The audit layer samples four surfaces, each answering a different question.

    Hotspot question routing. A maintained list of question types that have historically produced AI-teacher disagreement; new answers to those questions are routed to teacher review before publication, and the list updates as new mismatch patterns surface.

    Outlier teacher detection. Evaluator-to-evaluator deviation tracked within a long marking sitting; a teacher whose scoring drifts from the cohort over the sitting is flagged for principal review — the fatigue-and-anchoring slip psychometricians have long known about.

    Outlier copy 100% review. Every copy in the high-severity tail (for APMS, a gap of at least the greater of six marks or 10% of the total between AI and teacher scores) is routed for full review. No high-severity copy leaves without a second pair of eyes.

    Stratified continuous sampling. A statistical sample across production volume on a continuous cadence, stratified by question type, subject, teacher, and class, so drift is monitored over time without waiting for a hotspot or an outlier. This is what makes it a layer rather than a set of alarms.

    The Tier 2 audit layer is not the Tier 3 adjudication study. In the APMS Phase 2 report, the 215 cases where AI and teacher disagreed were sent through an independent adjudicator against the intended marking scheme — that is the Tier 3 piece, the methodology behind the AI 89.7% versus Teacher 82.8% finding on the mismatch set. The monthly audit pack is the Tier 2 discipline that keeps Tier 1 honest between adjudication studies.

    What a monthly audit pack contains

    The monthly pack is the visible artefact of the audit layer — what a quality team reads on a cadence, and what a buyer should be able to ask to see before signing. Drawing on the APMS Phase 2 Policy Annex, a CGF audit pack carries six sections:

    1. Period summary — window dates, production volume, subject/teacher/class coverage, deployment context.

    2. Drift surfaces — three trend lines: per-question reliability (Cronbach's alpha) drift, total-score MAE drift teacher-by-teacher and at cohort level, and per-teacher consistency over time.

    3. Hotspot questions surfaced — a ranked list of the top mismatched question types this period, with anonymised sample copies. A type recurring at the top means the marking scheme, the teacher calibration, or the AI rubric enforcement on it needs attention; the pack surfaces the pattern, it does not pick the fix.

    4. High-severity tail report — count of high-severity outlier copies, the confidence interval around it, and the governance actions taken on each. For APMS Phase 2 the tail was 8.7% (CI 0.0%–21.7%); the width of that interval is itself a governance signal, reported openly.

    5. Escalation log — every dispute that walked the ladder this period: who escalated, what about, what level it resolved at, how long it took, what changed.

    6. Action items — what changed in Production routing as a result: new hotspot questions, marking-scheme revisions queued, teacher calibration scheduled, model releases flagged for the next adjudication cycle.

    Cadence holds it together: monthly pack as the working surface, quarterly review reading three packs as a trend, annual review feeding the next adjudication cycle's sampling.

    Drift and the high-severity tail

    Drift is a slow change in the relationship between inputs and scores. The pack's three trend lines instrument it: per-question alpha drift, total-score MAE drift across months (APMS reports the MAE at a point in time, 3.91 marks, CI 2.61–5.41; the audit layer reports the same number as a line), and per-teacher reliability over time. The high-severity tail is where drift becomes a per-student governance question: in APMS Phase 2 it was 8.7% of copies (CI 0.0%–21.7%), both numbers reported together, with 100% review of every tail copy and a defined escalation route. The standard we hold, ourselves included, is that a tail reported without an interval is not a measurement.

    The escalation ladder

    The APMS Phase 2 Policy Annex defines the escalation structure that runs alongside the audit layer. The teacher remains the final authority on every score; the ladder adds a defined route up when a decision is queried — Teacher, Principal, Block Educational Officer, Block QA, District QA. Three clocks hold it together: a 72-hour provisional window (scores are provisional when first published, open to query), a 5-working-day query SLA, and a 2-working-day escalation clock if a level does not resolve in time. The escalation log captures every step. The ladder is what the APMS deployment runs; the shape is configurable — an NGO programme, a K-12 chain, and a state rollout each have a different chain — and the discipline we hold is that the chain is named, the clocks defined, the log captured, and the audit pack reads the log each month.

    How audit closes the loop

    The audit layer does not just record; it writes back. A pack surfaces a drift signal in a question type; the action items add that type to hotspot routing; the next pack reports whether the mismatch rate fell or whether a marking-scheme revision is the better answer. A second loop runs at the teacher level: a consistency drift triggers a calibration session, and the next month rechecks the trend. The quarterly review takes three packs together — where a monthly trend looks like noise, three months can show whether drift is real — and hands a structured input back to the research team on where the next adjudication cycle should sample more heavily.

    The audit framework — the four sampling surfaces, the six-section pack, the escalation ladder — ran end to end in the APMS Phase 2 deployment under the Policy Annex. That is one deployment, one school system, one validation window. The platform-wide pipeline that would run daily across every production tenant is the architectural commitment we are building toward. The methodology is shipped; the pipeline is in build. We would rather say that openly than ship a page that implies more.

    What to expect from any serious evaluation system

    If you want a diligence ask-list specific to the audit layer, six questions describe what a serious evaluation system should answer — and we hold ourselves to the same list: a sample audit pack (the structured monthly artefact, not a deck screenshot); a drift-detection cadence (monthly/quarterly/annual, at what level, which surfaces); a high-severity tail governed openly (threshold, count with its interval, the workflow on tail copies); a named escalation ladder (defined clocks, captured log, configurable chain); a feedback path from audit findings back into routing; and audit packs from existing deployments — and where a pipeline is still in build, an honest answer that it is. Continuous audit is not a free byproduct of scoring; it is the architecture of trust. We do not put audit governance behind a tier upgrade or sell it as a separate compliance module — it is part of the infrastructure tier.

    KEY TAKEAWAYS

    Key takeaways

    • Audit is the continuous Tier 2 layer in the three-tier architecture. Production produces evidence; audit reads it on a cadence and responds; adjudication validates periodically.
    • It samples four surfaces: hotspot question routing, outlier teacher detection, 100% review of every high-severity outlier copy, and stratified continuous sampling for drift.
    • The monthly pack has six sections: period summary, drift surfaces, hotspot questions, high-severity tail report, escalation log, action items.
    • APMS Phase 2 high-severity tail: 8.7% (CI 0.0%–21.7%), reported openly with the response workflow.
    • Escalation ladder (APMS Policy Annex): Teacher → Principal → Block Educational Officer → Block QA → District QA, with a 72-hour provisional window, 5-working-day query SLA, 2-working-day escalation clock.
    • The framework ran end to end in the APMS Phase 2 deployment. The platform-wide pipeline is the architectural commitment we are building toward — shipped methodology, in-build pipeline.

    KEEP READING

    Keep reading

    // COMMON.QUESTIONS

    Common questions

    Is the audit pipeline generally available today across all CrazyGoldFish deployments?

    No. The framework, the four sampling surfaces, the six-section monthly pack, and the escalation ladder ran end to end in the APMS Phase 2 deployment in Srikakulam under the Policy Annex. The platform-wide pipeline that would run daily across every production tenant is the architectural commitment we are building toward — the methodology is shipped, the pipeline is in build.

    What does a monthly audit pack cost?

    Audit is part of the infrastructure tier, not a separate module priced on top — the same way human-in-the-loop review and the audit trail ship in the standard K12 API. Cost is scoped to the deployment: production volume, subject coverage, cadence, and the escalation ladder configuration.

    Can I run audit on data that has already been graded?

    Yes. The audit layer can read historical graded data, compute the drift surfaces, identify hotspot questions, and produce a retrospective pack that baselines the continuous packs that follow — also useful before a model release or a marking-scheme revision.

    Who reviews the audit pack?

    The customer's quality team — academic leadership at block or district level in a school system, the programme quality lead in an NGO, the assessment governance team in an enterprise. The audit layer is the data infrastructure; the reviewer is the human governance layer.

    How does Tier 3 adjudication relate to Tier 2 audit?

    Audit catches drift continuously on smaller samples as a normal part of operations; adjudication validates Tier 1 and Tier 2 periodically on larger samples against an independent reference, producing the citable reliability claim. The 215-case mismatch cohort from APMS Phase 2 is a Tier 3 adjudication — the source of the AI 89.7% versus Teacher 82.8% finding. They answer different questions on different clocks.

    Talk to us about an audit pack

    If you run assessments at scale — a school system, a network, or a government programme — the audit framework can be configured to your deployment: the sampling surfaces, the escalation ladder, and the routing loop scoped to your governance chain.