clay-borg/research/CB-RES-0004-replay-and-kernel-coverage.md
tegwick 923c2038be CB-RES-0004: the instruments count names, not assertions
Tier-L survey plus the full review trail. v1 was REJECTED by adversarial
review with four blocking findings; this is v2, and the finding below is
the reviewer's as much as mine.

The root cause is neither v1's story ("coverage scope is too narrow" —
true but shallow) nor the reviewer's cheaper one ("the link check points
at one file" — true but catches 3 of 7). It is:

  Every coverage instrument in this project counts NAMES. None counts
  ASSERTIONS. M-D1-COV counts covers: tags; M-D1-LNK counts rule-ID
  strings. Both answer "is this rule mentioned?" and neither answers
  "does anything fail if this rule is violated?"

Four of seven defects found this pass are named in the source and inert,
so no name-based check finds them. This subsumes the oldest error class:
all five prior harness-does-nothing instances have the same shape.

Two are mutation-proven, because the reviewer set that bar for findings
they had only read:

  AM-7 — folding a 100k-event log from fresh(999) instead of fresh(42),
  an unrelated genesis state, leaves the test green. The hash reaches
  only a println!; the sole assertion is on elapsed time. And the log is
  built across games seeded 42,43,44... so it could not be asserted as
  written. evidence/CB-EV-0001 reports "AM-7 replay met, 2,290x" —
  unearned on the hash-identical clause. HDN instance #6.

  K9 — making Snapshot::take discard its EventSeq and store 0, which
  destroys the half of K9 that says "+ the EventId it includes", leaves
  the test green. On the acceptance criterion GameKernel §2.4 is zero of
  three, not two of three.

Also found: K10, K14, K18 named nowhere in source; AM-10 vacuous (a K6
determinism lint reported under a D4 leak row, over a cb-*-api crate that
does not exist); AM-11's conformance suite does not exist; AM-1b measured
and omitted from the evidence scoreboard.

All twelve challenge findings conceded, none defended on its facts. v1's
prediction was rejected as unfalsifiable AND already determined; the
replacement names its refuting outcome and is beatable in both
directions.

Chaos d4=1, no override. Declaration 1 of the 12-declaration window.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 14:19:30 +02:00

14 KiB
Raw Blame History

CB-RES-0004: the instruments count names, not assertions

capability: meta.coverage.assertions + kernel.replay status: approved (v2, after adversarial review — v1 was rejected with four blocking findings) tier: L (structural L — creates a new capability port (durable log storage); changes what several headline acceptance numbers mean; chaos d4=1 → no override) instrument: extended make coverage (spec→code over every numbered spec) plus a new make mutation-check and make replay-test baseline: cited — boardgame.io, Rune, rr/Pernosco, Playwright trace, event-sourcing log framing

Survey opened 2026-07-31. v1 was rejected; the review trail is history/260731-replay-kernel-coverage-{research,challenge,response}.md. The finding below is the reviewer's as much as mine.


1. The finding

Every coverage instrument in this project counts names. None counts assertions. M-D1-COV counts covers: tags in scenario YAML. M-D1-LNK counts rule-ID strings in one source file. Both answer "is this rule mentioned?"; neither answers "does anything fail if this rule is violated?"

Seven defects were found this pass. Four of them are named in the source and unasserted — invisible to any name-based check, including the cheaper one the reviewer proposed:

defect named in code? a name-based check finds it?
K10 replay bundle — no writer, reader, or --replay no yes
K14 commit window — CommitWindow has zero non-test users no yes
K18 benches from scenarios — bench hardcodes commands no yes
K9 snapshot + the EventId it includes yes no
K11 durable format, truncated-tail detection yes no
AM-7 hash-identical yes no
AM-11 passing the same conformance suite yes no

This subsumes the project's oldest error class. All five prior harness-does-nothing instances have the same shape as AM-7 and K9 today: the name is there, the assertion is not.

1.1 The two that are mutation-proven

AM-7's hash-identical clause detects nothing. games/ground/src/lib.rs computes state_hash_hex(&state) and passes it to a println!. The only assertion is on elapsed time. Mutating the fold's genesis from fresh(42) to fresh(999) — replaying a 100k-event log into an unrelated starting state — leaves the test green:

test replay_probe::replay_100k_events_is_linear_and_fast ... ok

The hash could not be asserted as written anyway: the log is built across games seeded 42, 43, 44… (source = fresh(42 + games)) and folded into a state from fresh(42). It is a fold-rate microbenchmark over a semantically invalid stream, reported in evidence/CB-EV-0001 as AM-7 replay | met, 2,290×. That verdict is unearned on the hash clause and needs a correction note. HDN instance #6.

K9's through field is unverified. The whole evidence for K9 is one test round-tripping a BTreeMap<String, u8> with EventSeq(17) as a literal — no game aggregate, no events applied, no from-genesis comparison. Snapshot has no callers outside that test. Mutating Snapshot::take to discard its EventSeq and store EventSeq(0) leaves the test green.

On the acceptance criterion, GameKernel §2.4 is zero of three, not two of three.

1.2 The three that are named nowhere

spec:  K1..K18       code:  K1..K9 K11 K12 K13 K15 K16 K17
diff:  K10, K14, K18
  • K10 — no .cbreplay writer, reader, or replays/ directory; the only fs:: call in the workspace is read_to_string for scenario YAML. cb-sim has no flag parsing at all, so --replay has no place to go. MetricsAndScenarios §2 promises a failing run writes a bundle. It does not. INTENT names this design decision 8 of 10.
  • K14 — "one commit window per Select". CommitWindow has zero non-test users and games/ground does not import it; GROUND collects selections in its own aggregate. A dead abstraction certified by a test that exercises only itself.
  • K18 — "Criterion benches driving the same scenario format". The bench hardcodes commands in Rust and never touches ScenarioFile. benchmarks/ contains only baselines/model-prices.toml.

1.3 The scoped-denominator sweep

v1 claimed this problem was general and swept in one place. Swept properly, in descending order of severity:

AM-10 is vacuous — the population is empty. The row reads M-D4-LEAK: foreign types in cb-*-api-visible signatures | 0. There is no cb-*-api crate. What is actually measured is a clippy.toml deny of HashMap/HashSet whose stated reason cites K6 (determinism), reported under a D4 leak row. A determinism lint standing in for a leak metric, over a crate that does not exist.

AM-1's denominator is one spec, and the metric says otherwise. M-D1-COV is defined as "numbered spec rules covered by ≥1 passing scenario", and covers is documented as "numbered rules from the capability spec" with a non-GR example. RULE_RE matches GR- only, against GroundRules.md only. So 58/58 (100%) is not merely narrowly scoped — the instrument is non-conformant with its own metric definition, and 100% is a wrong answer to the question M-D1-COV asks.

AM-1b is measured and dropped from the scoreboard. make coverage prints 49/58 and MetricsAndScenarios records it Unmet; grep -c 'AM-1b' evidence/CB-EV-0001 returns 0. The artifact carrying the headline reports the flattering half of the gate and omits the unflattering half the same command prints two lines below it.

AM-11's suite does not exist. M-D4-SWAP requires impls "passing the same conformance suite". grep -rn conformance over all .rs returns one doc comment describing future work. The RNG pair is exercised by two separate, non-shared tests. met, narrow is unearned on the suite clause.

AM-2, AM-3, AM-5, AM-9 are unreported and were unmentioned by v1.

The history file's stated weakness — "the evidence for this is n=1" — is retired. It is n≥5 in this repo, found without leaving the tree.


2. Baseline — what the state of the art does

Cited, not locally measured; per MetricsAndScenarios §3 this caps comparison verdicts at parity.

system durable log exportable failure bundle machine re-executable
boardgame.io in-memory, optional server persistence debug panel, in-session time travel no
Rune rollback ring buffer, transient no no
rr / Pernosco full process recording yes — the gold standard yes, bit-exact
Playwright trace per-run trace zip yes, one file partly (UI, not logic)
Kafka / EventStore append-only, length-prefixed, versioned, checksummed n/a yes, by re-consumption

Two lessons:

  1. The valuable property is machine re-execution, not inspection. boardgame.io's debug panel is more visible than a .cbreplay file and far less useful to an agent, because a human must interpret it. rr and Playwright are the precedents worth taking.
  2. The log format is solved engineering. Length prefix, version byte, per-record framing, tail detection: Kafka and EventStore converged on the same shape decades apart. It is ~80120 lines.

Provenance. Log framing: adapted:event-sourcing-log-framing — length-prefixed versioned records with tail detection, dropping the distributed concerns (partitions, offsets, compaction) a single-game log does not have. Bundle: adapted:playwright-trace — one artifact per failed run containing everything needed to reproduce, adapted so the consumer is an agent, not a human viewer. Assertion coverage: adapted:mutation-testing — mutate the property, expect a red test; applied only to acceptance rows rather than exhaustively.


3. The four dimensions

D1 — ease of specification. Strongly positive. A .cbreplay bundle turns "the scenario failed" into an artifact with a stated contract. More importantly, mutation coverage makes "this rule is enforced" a claim with a test rather than a tag.

D2 — efficiency of implementation. The log format adds no dependency (serde is already shipped). The bundle writer is not "a directory of four files" — v1 said that and it was wrong. scenario.rs creates an EventLog, appends to it, and never reads it; it is dropped at the end of execute. Pass carries the end state, not an initial snapshot, and failures are a formatted String, not structured expected-vs-actual. The bundle requires changing the runner's data flow: plumb the log out, capture an initial snapshot, restructure RunOutcome::Failed.

D3 — speed of execution. Bundle writes happen only on failure, so the hot path is untouched. Log append gains a length-prefix write. Mutation checks run in CI, not in the loop.

D4 — optionality. Where the pass earns its L, on a corrected premise. v1 argued AM-11 "claims" a storage port; it does not — that is the target column, and the evidence honestly reports 1 of 1, met, narrow. The real argument is stronger: M-D4-SWAP is a bool over a conformance suite that does not exist. Building a durable-log port with a genuine shared suite discharges AM-11 properly and creates the seam stage 3 (networked sessions) and stage 4 (game-creation framework) both need.


4. Options

A — implement K10/K11, leave the instruments alone. Rejected. Fixes what was noticed; leaves the mechanism that hid it.

B — extend coverage to K-rules, defer the rest. Rejected as a final state. Adopted as ordering.

C — adopt an embedded store (sled, redb). Rejected — and v1 rejected it against the wrong budget. The durable log (K11) is shipped-runtime, so AM-4a applies with 1.5% headroom; the bundle writer (K10) is dev-only behind the scenarios feature, so AM-4b applies with 9.4%. v1 wanted it both ways without saying so. With the split stated, C is rejected on simplicity for the log — a length prefix is not a hard implementation, and "assimilate the implementation" is for hard ones — and the headroom argument is dropped for the bundle, where it never applied.

E — extend AM-1b's spec→code link to every numbered spec and every crate. ~10 lines against an instrument that already exists. Catches K10, K14, K18 immediately. Adopted as step one — and demonstrably insufficient, since it catches 3 of the 7 defects and none of the four mutation-class ones. AGGREGATE must become a list.

F — E, plus mutation checks on acceptance rows, plus K10/K11. Recommended. The name-based check is cheap and catches the absent; the mutation check is the only thing that catches the present-but-inert; the capability work discharges the rules both instruments expose.


5. What this pass must not do

  • Do not tag K-rules as covered to make a number go up. The overclaim CB-WP-0001 committed once and M-D1-LNK exists to catch. A K-rule is covered when breaking it breaks a test.
  • Do not let the widened coverage gate fail the build on day one. rule-coverage.py returns 2 when missing is non-empty and coverage is in all, so 18 uncovered K-rules fail make all immediately. The kernel denominator must report outside the exit code until the pass closes the gap, with a stated date for when it binds.
  • Replicate the zero-rules positive control on the new denominator. The existing arm refuses to report over zero rules — a defect it was fixed for. A kernel regex matching nothing must abort, not print 0/0 (100%).
  • make replay-test must not be able to pass vacuously. It reports ok under at least four silent failures: no scenario fails so zero bundles are round-tripped; the writer emits nothing and identical error strings satisfy "same failure"; the replayed hash is compared to one recomputed in the same process (assert_eq!(h, h) — the AM-7 defect exactly); the truncation path is never exercised. Required controls: a committed deliberately-failing fixture plus bundles > 0; the comparison hash read out of the bundle, written by the first process; truncate-by-one-byte and corrupt-length-prefix negative controls; and a mutated-seed negative control so the round-trip is capable of failing.

6. Prediction

Stated as a mechanism per CB-WP-0004 T06, and with the refuting outcome named — v1's prediction was rejected for being unfalsifiable and already determined (any tag-based kernel count is 0 of 18 today, computable before writing a line).

claim mechanism after the fix refuted if
absent rules become visible spec→code runs over every numbered spec × every crate; no manual path any numbered rule in any spec is unnamed in source after the pass
present-but-inert rules become visible mutation: each acceptance row's property is inverted and the suite must go red fewer than 9 of the 12 acceptance rows have a mutation that turns a test red
replay is real acceptance test re-executes a bundle in a fresh process against a hash read from the bundle the replayed hash is recomputed in-process, or the round-trip cannot be made to fail
K9/K11 are enforced snapshot-at-N + events N+1..M ≡ genesis fold on GroundState; truncation rejected either passes when mutated

The middle row is the one that carries information. 9 of 12 is a guess, and it is meant to be beatable in both directions — if 12 of 12 mutate red, the instruments were better than this survey claims and the finding shrinks to the three absent rules; if 3 of 12 do, the problem is larger than the pass is scoped for and the pass should stop and re-plan.