IN SHORT
A learning signal is the information in a student's answer that a careful teacher's read captures: which concept was grasped, where the reasoning broke, whether a mistake is a spelling slip or a real conceptual gap, what the next teachable step should be. The score is a one-number summary of that read; the signal is the read itself. The commitment we hold is that evaluation should capture the signal in structured form at the moment of grading, alongside the score, and let it persist. Three feedback-quality findings from the Educate Girls Hindi deployment describe what capture looks like in production: 2,447 unique feedback variants for 2,297 student-question pairs, 94.4% grounded in the model answer, and an average of 2.90 of 10 distinct teaching signals per variant (39% combining error identification with forward guidance). None of these is an accuracy number — all are feedback-quality evidence that the signal is captured, not collapsed. This page is the bridge between the Evaluation Infrastructure pillar, where capture happens, and the Memory Layer for Education pillar, where the captured signal persists and compounds.
What a learning signal is
Watch a careful teacher read a student's answer. The teacher sees the final number the answer earns, and the teacher sees everything that produced it: the concept the student understood, the step where the reasoning broke, the spelling slip that did not really mean the idea was missed, the piece of the question answered in full and the piece skipped without noticing, the misconception that looks like an arithmetic mistake until you look twice. That is the read. The score is a compression of the read into a number. The number survives the day; the read, in most systems, does not — the teacher carries it in their head until the next class, the next student, the next sitting, and eventually it is gone.
We call the information inside that read a learning signal: everything the read captured beyond the number — concept grasped, reasoning step where the answer broke, type of error, distance between the student's answer and the model answer, the quality of the partial credit the answer deserves, the next teachable step. Two students can earn the same wrong final answer for completely different reasons: one misread the question and answered a different one in full; the other set up the reasoning correctly and made an arithmetic slip in the last line. The score is identical. The signal is not, and the next teaching action for each is not the same. A score-only record cannot tell them apart; a signal-capturing record can, because it captured the difference at the moment of grading and stored it. The score is the floor. The signal is the work above it.
Why scores hide signals
The score and the signal both exist at the moment of evaluation. The score is the headline; the signal is everything the headline summarised. The collapse — keeping one, discarding the other — is a design choice, not a property of grading, and it happens for an operational reason. A row that holds a number and a date fits in any system. A row that also holds a list of teaching signals, a reference to the rubric, a pointer to the model answer, and the verbatim explanation the grader wrote is a much larger row: it needs a schema, a storage layer that knows the shape of an evaluation rather than just the shape of a number, and a workflow that captures the signal when the grade is published rather than as a separate exercise afterward.
The reason it matters is that the signal cannot be reconstructed from the score. A 4 of 10 carries nothing about why; the signal that would explain it was there at the moment of grading, and if it was not captured in structured form then, no later analysis recovers it. The standard we hold, ourselves included, is that capture must happen at the moment of evaluation, in structured form, alongside the score. We treated the collapse-into-number as the design problem and built around it: the pipeline produces a score per question, a model-answer-grounded feedback variant, a set of tagged teaching signals, a record of the rubric path that produced the partial credit, and a teacher review state — every piece captured the moment the grade is published. The score is one field of many on the row. The wider architecture (Production, Audit, Adjudication) is at /research/concepts/three-tier-evaluation-architecture; this page is the deeper drill on what those layers capture.
What CrazyGoldFish captures
For every student-question pair the system grades, the captured record carries: the student's actual answer as recognised from the script; the model answer it was checked against; the rubric path that produced the score and partial credit; the per-question feedback variant in natural language, in the language of instruction; a set of tagged teaching signals on that variant; a reference to the question; a reference to the teacher's review state; and the timestamp and provenance that make the row auditable. The score is one field on that row; the signal is the rest.
The Educate Girls Hindi-medium deployment in Madhya Pradesh and Rajasthan is where this capture has been published externally as feedback-quality evidence — across 11 tests, two states, seven subjects, 162 questions, and 2,297 student-question pairs, three findings describe what the captured signal looks like at scale.
INLINE CALLOUT
More distinct explanations than answers graded: 2,447 unique feedback variants for 2,297 pairs. The ratio rules out templated feedback assigned to a small list of score buckets — templates cannot produce more unique outputs than inputs. Feedback-quality evidence, not an accuracy figure: the signal was captured per response, not assigned per bucket.
INLINE CALLOUT
Variants grounded in the model answer: 94.4% of those variants explicitly reference the model answer — the strongest empirical anchor against feedback that drifts from the source the student was checked against. Also a feedback-quality number, not an accuracy figure.
INLINE CALLOUT
Multiple teaching signals per variant: every variant was tagged for 10 distinct teaching signals (error identification, forward guidance, references to the student's response, references to the model answer, format-error identification, pedagogical explanation, no-answer identification, question convention, partial-credit explanation, praise, spelling-error identification). The average variant carried 2.90 of those 10 — and as a paired finding, 39% of variants combine error identification with forward guidance, the highest-value combination: diagnose what went wrong and point to the next teachable step in one feedback, on more than a third of records, on handwritten Hindi.
The labelling discipline is load-bearing: these are feedback-quality numbers, not accuracy figures. Whether the score itself is correct is a different question, reported against under the reliability discipline at /research/concepts/reliability-vs-accuracy and the adjudication tier at /research/concepts/three-tier-evaluation-architecture. The signal-capture finding is that the read was captured in structured form alongside the score, anchored to the model answer, multi-dimensional per record. That is what one captured signal looks like in production.
How signals become persistent
A captured signal is not yet a persisted one. Capture happens at grading; persistence is the commitment that the captured row is kept, linked, and queryable so a reader weeks or terms later can act on it. Four properties separate the two. Structured — the signal is a schema, not a free-text field; the 10 teaching tags, the model-answer grounding, the rubric path, and the feedback variant (with language tagged) are fields, and a field can be queried where a blob can only be searched. Linked — the row carries references: to the student (aggregate over time for one learner), to the question (aggregate over one item across students), to the rubric (recover the path), to the teacher review state (the human record travels with the AI record). Machine-readable — the structure and references can be re-read downstream without a human translating the row; memory is built on machine-readable signal, not archived PDFs. Persisted as a layer, not an export — the captured signal lives in a memory layer the next evaluation, analytics surface, and research read can query without re-fetching from the production grading log.
The memory layer is the second product pillar; the deeper drill on the schema is at /memory-layer-for-education. The point here is the bridge: evaluation captures the signal, the memory layer persists it, and the commitment is to write the structured, linked, machine-readable row into the layer at the moment of grading. A captured-but-not-persisted signal is a missed opportunity — the read happened, the signal was named, the row was written, then discarded at the boundary of the production grading log. We treat that boundary as the architectural seam, and the commitment is that the seam is closed.
What signals enable when stored
Five operations become available once the signal layer is in place, none of them available from a score-only record. Longitudinal learning trajectories — a learner's record becomes a sequence of captured signals, not just grades, so a teacher at the start of a term picks up where the last left off rather than re-deriving it. Per-question reliability over time — the same item across cohorts and teachers generates a reliability trend from the score matrix; the signal layer makes the reliability discipline (at /research/concepts/reliability-vs-accuracy) computable continuously, not only at a study's end. Drift detection — the audit layer (at /research/concepts/audit-pipeline) reads the signal stream continuously for per-question reliability, MAE, and per-teacher consistency trends; without persistence, drift detection is a one-time report, not a continuous reading. Cohort-level intelligence — a teacher, school leader, programme director, or state authority can ask cohort-shaped questions of the layer; the Educate Girls deployment surfaced one such gap, students in Madhya Pradesh outscoring Rajasthan by 23.3 percentage points on Hindi, readable because the signal was persisted in an aggregable form. Memory for the system itself — a system that persists signals has a memory of its own evaluations: the next reading is informed by the previous one, the next teacher has the previous teacher's feedback as context, the next adjudication study samples where drift concentrates. The system stops being a sequence of independent gradings and becomes a coupled record that compounds across time, learners, items, and teachers. The captured signal is what evaluation produces; the persisted signal is what memory holds; this page is the part of the chain between them.
What to expect from any serious evaluation system
If you want a diligence ask-list specific to signal capture and persistence, six questions describe what a serious evaluation system should answer — and we hold ourselves to the same list. Signal captured at the moment of evaluation, not reconstructed afterward (ask to see one grading record end to end, with score, feedback variant, rubric path, model-answer grounding, and teaching-signal tags as discrete fields). Structured, machine-readable signal (ask for the schema; a free-text blob is not a captured signal). Signal grounded in the rubric and model answer (ask for the percentage grounded — the Educate Girls deployment reports 94.4%). Multi-dimensional signal per record, not a single tag (ask how many distinct signal types are tagged and the average per record — Educate Girls tagged 10, averaging 2.90, with 39% carrying the error-ID-plus-forward-guidance combination). Persistent signal that compounds across time, learners, and assessments (ask where the signal lives, how it links to the learner record, what queries the layer supports). The same signal stream feeding audit, longitudinal analytics, and personalisation downstream (ask whether all three read from one captured layer; three separate extracts signal an incoherent substrate). The standard we hold, ourselves included, is signal at the capture point, in structured form, grounded in the rubric, multi-dimensional per record, persisted in a layer downstream reads share.
KEY TAKEAWAYS
Key takeaways
- →A learning signal is the information a careful teacher's read captures: concept grasped, reasoning step broken, type of error, partial-credit reasoning, next teachable step. The score is the compression of that read into a number.
- →The collapse from signal to score is an architectural choice, not a property of grading. The commitment we hold is to capture the signal at the moment of evaluation, in structured form, alongside the score.
- →The Educate Girls Hindi deployment is the published evidence: 2,447 unique feedback variants for 2,297 pairs, 94.4% grounded in the model answer, an average of 2.90 of 10 teaching signals per variant, and 39% combining error identification with forward guidance. These are feedback-quality findings, not accuracy figures.
- →Captured signals become persistent when written as structured, linked, machine-readable rows into a memory layer downstream reads can query — rather than discarded at the production grading boundary.
- →A persisted signal layer enables longitudinal trajectories, per-question reliability over time, drift detection for audit, cohort-level intelligence, and memory for the system itself.
- →Learning signals are the bridge between the Evaluation Infrastructure pillar (where capture happens) and the Memory Layer for Education pillar (where the captured signal persists and compounds).
KEEP READING