Never stub what you test
The design rule of the whole package, in one sentence:An eval never stubs the thing it is testing.It exists because the harness this one replaced broke it. The earlier data-entry evals injected a scripted extractor, so the model was never exercised - and a classification or extraction regression could not fail the eval. A green board meant the script still matched itself.
What is real, per agent
Every target runs the agent’s real compiled graph.The Typewriter’s row is the strongest one, and it is arithmetic rather than a promise.That graph has exactly one injected boundary - the model - and no storage of any kind.
There is nothing left to substitute, which is why its target is the shortest file in the package.
The distinction that makes the rule usable
“Never stub anything” is not a workable rule; every eval has to stop somewhere. The line this package draws:Legitimate to substitute
A storage side effect.
The Clerk’s queue and file sink become in-memory doubles behind the same interface, so no eval run can write into the live intake queue.Every governance decision still executes the real code.
Never substitute
The reasoning under test.
The extractor, the classifier, the posting rules, the check families, the scope gate.Substituting any of those produces a board that grades the substitute.
Three design rules, each with the guard test that enforces it
1. A missing prerequisite reports SKIP, never a pass
No Postgres, no model credentials, an agent branch that has not landed: the example reports SKIP. A skip is never counted as a pass, so a green board always means green code. And--strict makes skips fatal, so an environment that is supposed to have everything wired can demand it.
This was verified rather than assumed: pointed at an unreachable database host, the ledger cases report [SKIP] ... no Postgres reachable, score nothing, and the process exits 0 normally but 1 under --strict.
2. No agent is parked as un-gradeable
SKIPPED_AGENTS is empty, and a test keeps it empty.
3. An expectation nobody reads is worse than no expectation
test_every_reference_key_is_graded_by_some_evaluator asserts that every key in every example’s expected outputs is read by at least one of that agent’s evaluators.
An ungraded expectation still shows PASS, which is the most quietly dangerous state a case can be in.
That test caught a dead expectation key during the build.
The judge boundary
A model judges in exactly one place, and it is not the accounting.The single judged score is the Clerk’s free-form in-scope chat reply, graded against that example’s written rubric.
It fails closed: an unavailable judge scores 0, never a pass.Account codes, balance, entry counts, hash-chain findings, flag fields, severities, scope verdicts, schema types and row counts are all exact comparisons.
Recording a known defect
A permanently red case hides a regression exactly as well as a permanently skipped one: the next real failure lands on an already-red board and nobody looks twice. So a case that fails for a known, accepted reason declares it - by evaluator key, and by the situation that produces it.
Three things about this mechanism are load-bearing.
Naming the keys.
A bare “this case may fail” flag would swallow a new regression inside a case that was already red.
only_when, which is the same rule one level finer.
The Typewriter’s intermittent classify miss fails every payload-derived key at once, but only by ending the turn with zero proposals.
Excusing doc_type_correct unconditionally would also excuse a turn that did propose and named the wrong one of the 146 types - which is the mis-proposal the dataset exists to catch.
Conditioned on {"n_proposals": 0}, the same key is an accepted defect on the turn that reproduces it and a hard failure on every other.
XPASS is a failure.
When a recorded defect is fixed, the run says so and goes red, so a stale excuse cannot outlive the bug.
That is not theoretical: the first run against the number-coercion fix reported three XPASS - “the defect this baseline recorded appears fixed” - which is how the fix announced itself.
The structure tests check both directions of it: that a marked case still goes red on a wrong doc type, that the anti-fabrication bound is never excused, that a recorded intermittent miss never turns the board red, and that no marker names a key nothing scores - because a marker naming an impossible key would excuse a case forever for a reason that cannot occur.
How a score is computed
Every evaluator returns zero or more scores. Returning nothing means “this evaluator does not apply to this example”, so a mixed dataset never manufactures a 1.0 for a case nothing graded. An example passes when every score it received is 1.0. One deliberate wrinkle: extraction and field accuracy are reported as a raw ratio alongside a pass/fail floor. The floor is the gate; the ratio is a trend line. A case carrying a floor of 0.85 therefore still fails at 0.90 on the raw key, which is a real and recorded cause of cases sitting below green.Layout
The in-repo JSON is the source of truth.
The upload rebuilds the LangSmith copy rather than merging into it, so an edit made in the dashboard is overwritten on the next sync.
A dataset is code, and it lives in the repo with the code it grades.
Where it is enforced
Related
- The scoreboard - what this harness produced
- What an eval is, and is not - why the engine is deliberately not on this board
- How this gates a release - which half of this runs on every PR
- Reading a run - the trace a graded run leaves behind