The evaluation model
Every evaluation is a pure function of one fixture-run: a versioned
envelope pairing a fixture's expected checks with the system-under-test's
actual output. The evaluator runs the declared checks, assigns each a
verdict, and aggregates them. Same input, byte-identical report — no randomness, no
wall-clock, no network, no datastore.
The input envelope is validated first. An unsupported contractVersion,
an unknown check name, or a malformed envelope fails closed as a contract
error (exit code 3), kept distinct from a fixture whose scenario
kind is malformed.
The three failure modes
Forgetting — a stale or lost fact returned in place of the current one
Scored with containment checks over the answer: mustContain asserts the
current/retained value is present, and mustNotContain asserts a known-stale
value has not leaked back in. answerEquals can pin an exact
answer. A leaked stale value or a missing retained value is a fail.
Chronology — facts returned in the wrong time order
Scored with the order check: the actual.order sequence must
deep-equal the expected.order. Any reordering — even of correct items — is a
fail, because downstream reasoning over mis-ordered history is unsafe.
Entity confusion — a fact attributed to the wrong entity
Scored with the attributedTo check: actual.entity must match
the expected entity (case- and whitespace-insensitive). Attribution to the wrong entity
is a fail.
The tri-state verdict
Checks aggregate with a fixed precedence: fail > unknown > pass.
- fail — at least one check failed.
- unknown — no check failed, but the evidence needed to assign a pass
was missing (for example, no
actual.answerwas supplied). Missing evidence is reported with a reason; it never fabricates a pass or a fail. - pass — at least one check passed and none were failed or unknown.
The CLI maps verdicts to exit codes — 0 pass, 1 fail,
2 unknown, 3 contract error — so a suite gates CI directly.
Receipts and reproducibility
Every report carries a receipt with a sha256 digest of the canonicalized
input (recursively key-sorted JSON) plus the contract and evaluator versions. Two runs of
the same input on the same evaluator version produce the same digest, so any result is
auditable and replayable.
Try it
# Forgetting: a stale value leaked back in → exit 1 (fail)
node src/cli.mjs fixtures/forgetting/negative-stale-value.json --human
# Missing evidence → exit 2 (unknown), never a false pass/fail
node src/cli.mjs fixtures/edge/missing-data.json --human
Contract version 1, evaluator
1.0.0.