clay-borg/research/CB-RES-0004-replay-and-kernel-coverage.md
tegwick 923c2038be CB-RES-0004: the instruments count names, not assertions
Tier-L survey plus the full review trail. v1 was REJECTED by adversarial
review with four blocking findings; this is v2, and the finding below is
the reviewer's as much as mine.

The root cause is neither v1's story ("coverage scope is too narrow" —
true but shallow) nor the reviewer's cheaper one ("the link check points
at one file" — true but catches 3 of 7). It is:

  Every coverage instrument in this project counts NAMES. None counts
  ASSERTIONS. M-D1-COV counts covers: tags; M-D1-LNK counts rule-ID
  strings. Both answer "is this rule mentioned?" and neither answers
  "does anything fail if this rule is violated?"

Four of seven defects found this pass are named in the source and inert,
so no name-based check finds them. This subsumes the oldest error class:
all five prior harness-does-nothing instances have the same shape.

Two are mutation-proven, because the reviewer set that bar for findings
they had only read:

  AM-7 — folding a 100k-event log from fresh(999) instead of fresh(42),
  an unrelated genesis state, leaves the test green. The hash reaches
  only a println!; the sole assertion is on elapsed time. And the log is
  built across games seeded 42,43,44... so it could not be asserted as
  written. evidence/CB-EV-0001 reports "AM-7 replay met, 2,290x" —
  unearned on the hash-identical clause. HDN instance #6.

  K9 — making Snapshot::take discard its EventSeq and store 0, which
  destroys the half of K9 that says "+ the EventId it includes", leaves
  the test green. On the acceptance criterion GameKernel §2.4 is zero of
  three, not two of three.

Also found: K10, K14, K18 named nowhere in source; AM-10 vacuous (a K6
determinism lint reported under a D4 leak row, over a cb-*-api crate that
does not exist); AM-11's conformance suite does not exist; AM-1b measured
and omitted from the evidence scoreboard.

All twelve challenge findings conceded, none defended on its facts. v1's
prediction was rejected as unfalsifiable AND already determined; the
replacement names its refuting outcome and is beatable in both
directions.

Chaos d4=1, no override. Declaration 1 of the 12-declaration window.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 14:19:30 +02:00

272 lines
14 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# CB-RES-0004: the instruments count names, not assertions
capability: meta.coverage.assertions + kernel.replay
status: approved (v2, after adversarial review — v1 was rejected with four
blocking findings)
tier: L (structural L — creates a new capability port (durable log
storage); changes what several headline acceptance numbers mean; chaos
d4=1 → no override)
instrument: extended `make coverage` (spec→code over every numbered spec)
plus a new `make mutation-check` and `make replay-test`
baseline: cited — boardgame.io, Rune, rr/Pernosco, Playwright trace,
event-sourcing log framing
Survey opened 2026-07-31. **v1 was rejected**; the review trail is
`history/260731-replay-kernel-coverage-{research,challenge,response}.md`.
The finding below is the reviewer's as much as mine.
---
## 1. The finding
> **Every coverage instrument in this project counts names. None counts
> assertions.** M-D1-COV counts `covers:` tags in scenario YAML. M-D1-LNK
> counts rule-ID strings in one source file. Both answer *"is this rule
> mentioned?"*; neither answers *"does anything fail if this rule is
> violated?"*
Seven defects were found this pass. Four of them are **named in the source
and unasserted** — invisible to any name-based check, including the
cheaper one the reviewer proposed:
| defect | named in code? | a name-based check finds it? |
|---|---|---|
| K10 replay bundle — no writer, reader, or `--replay` | no | **yes** |
| K14 commit window — `CommitWindow` has zero non-test users | no | **yes** |
| K18 benches from scenarios — bench hardcodes commands | no | **yes** |
| K9 snapshot `+ the EventId it includes` | yes | no |
| K11 durable format, truncated-tail detection | yes | no |
| AM-7 `hash-identical` | yes | no |
| AM-11 `passing the same conformance suite` | yes | no |
This subsumes the project's oldest error class. All five prior
harness-does-nothing instances have the same shape as AM-7 and K9 today:
**the name is there, the assertion is not.**
### 1.1 The two that are mutation-proven
**AM-7's `hash-identical` clause detects nothing.**
`games/ground/src/lib.rs` computes `state_hash_hex(&state)` and passes it
to a `println!`. The only assertion is on elapsed time. Mutating the
fold's genesis from `fresh(42)` to `fresh(999)` — replaying a 100k-event
log into an unrelated starting state — leaves the test green:
```text
test replay_probe::replay_100k_events_is_linear_and_fast ... ok
```
The hash could not be asserted as written anyway: the log is built across
games seeded 42, 43, 44… (`source = fresh(42 + games)`) and folded into a
state from `fresh(42)`. It is a fold-rate microbenchmark over a
semantically invalid stream, reported in `evidence/CB-EV-0001` as
**`AM-7 replay | met, 2,290×`**. That verdict is unearned on the
hash clause and needs a correction note. **HDN instance #6.**
**K9's `through` field is unverified.** The whole evidence for K9 is one
test round-tripping a `BTreeMap<String, u8>` with `EventSeq(17)` as a
literal — no game aggregate, no events applied, no from-genesis
comparison. `Snapshot` has no callers outside that test. Mutating
`Snapshot::take` to discard its `EventSeq` and store `EventSeq(0)` leaves
the test green.
On the acceptance criterion, **GameKernel §2.4 is zero of three, not two
of three.**
### 1.2 The three that are named nowhere
```text
spec: K1..K18 code: K1..K9 K11 K12 K13 K15 K16 K17
diff: K10, K14, K18
```
- **K10** — no `.cbreplay` writer, reader, or `replays/` directory; the
only `fs::` call in the workspace is `read_to_string` for scenario YAML.
`cb-sim` has no flag parsing at all, so `--replay` has no place to go.
`MetricsAndScenarios` §2 promises a failing run writes a bundle. It does
not. INTENT names this design decision **8 of 10**.
- **K14** — "one commit window per Select". `CommitWindow` has zero
non-test users and `games/ground` does not import it; GROUND collects
selections in its own aggregate. A dead abstraction certified by a test
that exercises only itself.
- **K18** — "Criterion benches driving the same scenario format". The
bench hardcodes commands in Rust and never touches `ScenarioFile`.
`benchmarks/` contains only `baselines/model-prices.toml`.
### 1.3 The scoped-denominator sweep
v1 claimed this problem was general and swept in one place. Swept properly,
in descending order of severity:
**AM-10 is vacuous — the population is empty.** The row reads
*M-D4-LEAK: foreign types in `cb-*-api`-visible signatures | 0*. There is
no `cb-*-api` crate. What is actually measured is a `clippy.toml` deny of
`HashMap`/`HashSet` whose stated reason cites **K6 (determinism)**,
reported under a **D4 leak** row. A determinism lint standing in for a
leak metric, over a crate that does not exist.
**AM-1's denominator is one spec, and the metric says otherwise.**
M-D1-COV is defined as *"numbered spec rules covered by ≥1 passing
scenario"*, and `covers` is documented as "numbered rules from the
capability spec" with a non-GR example. `RULE_RE` matches `GR-` only,
against `GroundRules.md` only. So `58/58 (100%)` is not merely narrowly
scoped — the instrument is **non-conformant with its own metric
definition**, and 100% is a wrong answer to the question M-D1-COV asks.
**AM-1b is measured and dropped from the scoreboard.** `make coverage`
prints `49/58` and `MetricsAndScenarios` records it **Unmet**;
`grep -c 'AM-1b' evidence/CB-EV-0001` returns **0**. The artifact carrying
the headline reports the flattering half of the gate and omits the
unflattering half the same command prints two lines below it.
**AM-11's suite does not exist.** M-D4-SWAP requires impls "passing **the
same conformance suite**". `grep -rn conformance` over all `.rs` returns
one doc comment describing future work. The RNG pair is exercised by two
separate, non-shared tests. `met, narrow` is unearned on the suite clause.
**AM-2, AM-3, AM-5, AM-9 are unreported** and were unmentioned by v1.
The history file's stated weakness — *"the evidence for this is n=1"* — is
retired. It is n≥5 in this repo, found without leaving the tree.
---
## 2. Baseline — what the state of the art does
Cited, not locally measured; per MetricsAndScenarios §3 this caps
comparison verdicts at `parity`.
| system | durable log | exportable failure bundle | machine re-executable |
|---|---|---|---|
| **boardgame.io** | in-memory, optional server persistence | debug panel, in-session time travel | no |
| **Rune** | rollback ring buffer, transient | no | no |
| **rr / Pernosco** | full process recording | yes — the gold standard | **yes**, bit-exact |
| **Playwright trace** | per-run trace zip | **yes**, one file | partly (UI, not logic) |
| **Kafka / EventStore** | append-only, length-prefixed, versioned, checksummed | n/a | yes, by re-consumption |
Two lessons:
1. **The valuable property is machine re-execution, not inspection.**
boardgame.io's debug panel is more visible than a `.cbreplay` file and
far less useful to an agent, because a human must interpret it. rr and
Playwright are the precedents worth taking.
2. **The log format is solved engineering.** Length prefix, version byte,
per-record framing, tail detection: Kafka and EventStore converged on
the same shape decades apart. It is ~80120 lines.
**Provenance.** Log framing: `adapted:event-sourcing-log-framing` —
length-prefixed versioned records with tail detection, dropping the
distributed concerns (partitions, offsets, compaction) a single-game log
does not have. Bundle: `adapted:playwright-trace` — one artifact per
failed run containing everything needed to reproduce, adapted so the
consumer is an agent, not a human viewer. Assertion coverage:
`adapted:mutation-testing` — mutate the property, expect a red test;
applied only to acceptance rows rather than exhaustively.
---
## 3. The four dimensions
**D1 — ease of specification.** Strongly positive. A `.cbreplay` bundle
turns "the scenario failed" into an artifact with a stated contract. More
importantly, mutation coverage makes "this rule is enforced" a *claim with
a test* rather than a tag.
**D2 — efficiency of implementation.** The log format adds no dependency
(serde is already shipped). **The bundle writer is not "a directory of
four files"** — v1 said that and it was wrong. `scenario.rs` creates an
`EventLog`, appends to it, and never reads it; it is dropped at the end of
`execute`. `Pass` carries the *end* state, not an initial snapshot, and
failures are a formatted `String`, not structured expected-vs-actual. The
bundle requires changing the runner's data flow: plumb the log out, capture
an initial snapshot, restructure `RunOutcome::Failed`.
**D3 — speed of execution.** Bundle writes happen only on failure, so the
hot path is untouched. Log append gains a length-prefix write. Mutation
checks run in CI, not in the loop.
**D4 — optionality.** Where the pass earns its L, on a corrected premise.
v1 argued AM-11 "claims" a storage port; it does not — that is the target
column, and the evidence honestly reports `1 of 1, met, narrow`. The real
argument is stronger: **M-D4-SWAP is a bool over a conformance suite that
does not exist.** Building a durable-log port with a genuine shared suite
discharges AM-11 properly and creates the seam stage 3 (networked
sessions) and stage 4 (game-creation framework) both need.
---
## 4. Options
**A — implement K10/K11, leave the instruments alone.** Rejected. Fixes
what was noticed; leaves the mechanism that hid it.
**B — extend coverage to K-rules, defer the rest.** Rejected as a final
state. Adopted as *ordering*.
**C — adopt an embedded store (sled, redb).** Rejected — and v1 rejected
it against the wrong budget. The **durable log (K11) is shipped-runtime**,
so AM-4a applies with 1.5% headroom; the **bundle writer (K10) is dev-only**
behind the `scenarios` feature, so AM-4b applies with 9.4%. v1 wanted it
both ways without saying so. With the split stated, C is rejected on
simplicity for the log — a length prefix is not a hard implementation, and
"assimilate the implementation" is for hard ones — and the headroom
argument is dropped for the bundle, where it never applied.
**E — extend AM-1b's spec→code link to every numbered spec and every
crate.** ~10 lines against an instrument that already exists. Catches
K10, K14, K18 immediately. **Adopted as step one — and demonstrably
insufficient**, since it catches 3 of the 7 defects and none of the four
mutation-class ones. `AGGREGATE` must become a list.
**F — E, plus mutation checks on acceptance rows, plus K10/K11.**
**Recommended.** The name-based check is cheap and catches the absent; the
mutation check is the only thing that catches the present-but-inert; the
capability work discharges the rules both instruments expose.
---
## 5. What this pass must not do
- **Do not tag K-rules as covered to make a number go up.** The overclaim
CB-WP-0001 committed once and M-D1-LNK exists to catch. A K-rule is
covered when breaking it breaks a test.
- **Do not let the widened coverage gate fail the build on day one.**
`rule-coverage.py` returns `2` when `missing` is non-empty and
`coverage` is in `all`, so 18 uncovered K-rules fail `make all`
immediately. The kernel denominator must report outside the exit code
until the pass closes the gap, with a stated date for when it binds.
- **Replicate the zero-rules positive control on the new denominator.**
The existing arm refuses to report over zero rules — a defect it was
fixed for. A kernel regex matching nothing must abort, not print
`0/0 (100%)`.
- **`make replay-test` must not be able to pass vacuously.** It reports
`ok` under at least four silent failures: no scenario fails so zero
bundles are round-tripped; the writer emits nothing and identical error
strings satisfy "same failure"; the replayed hash is compared to one
recomputed in the same process (`assert_eq!(h, h)` — the AM-7 defect
exactly); the truncation path is never exercised. **Required controls:**
a committed deliberately-failing fixture plus `bundles > 0`; the
comparison hash read *out of the bundle*, written by the first process;
truncate-by-one-byte and corrupt-length-prefix negative controls; and a
mutated-seed negative control so the round-trip is capable of failing.
---
## 6. Prediction
Stated as a mechanism per CB-WP-0004 T06, and with the refuting outcome
named — v1's prediction was rejected for being unfalsifiable *and* already
determined (any tag-based kernel count is **0 of 18** today, computable
before writing a line).
| claim | mechanism after the fix | refuted if |
|---|---|---|
| absent rules become visible | spec→code runs over every numbered spec × every crate; no manual path | any numbered rule in any spec is unnamed in source after the pass |
| **present-but-inert rules become visible** | mutation: each acceptance row's property is inverted and the suite must go red | **fewer than 9 of the 12 acceptance rows have a mutation that turns a test red** |
| replay is real | acceptance test re-executes a bundle in a **fresh process** against a hash read from the bundle | the replayed hash is recomputed in-process, or the round-trip cannot be made to fail |
| K9/K11 are enforced | snapshot-at-N + events N+1..M ≡ genesis fold on `GroundState`; truncation rejected | either passes when mutated |
The middle row is the one that carries information. **9 of 12 is a
guess, and it is meant to be beatable in both directions** — if 12 of 12
mutate red, the instruments were better than this survey claims and the
finding shrinks to the three absent rules; if 3 of 12 do, the problem is
larger than the pass is scoped for and the pass should stop and re-plan.