clay-borg/history/260731-replay-kernel-coverage-response.md
tegwick 923c2038be CB-RES-0004: the instruments count names, not assertions
Tier-L survey plus the full review trail. v1 was REJECTED by adversarial
review with four blocking findings; this is v2, and the finding below is
the reviewer's as much as mine.

The root cause is neither v1's story ("coverage scope is too narrow" —
true but shallow) nor the reviewer's cheaper one ("the link check points
at one file" — true but catches 3 of 7). It is:

  Every coverage instrument in this project counts NAMES. None counts
  ASSERTIONS. M-D1-COV counts covers: tags; M-D1-LNK counts rule-ID
  strings. Both answer "is this rule mentioned?" and neither answers
  "does anything fail if this rule is violated?"

Four of seven defects found this pass are named in the source and inert,
so no name-based check finds them. This subsumes the oldest error class:
all five prior harness-does-nothing instances have the same shape.

Two are mutation-proven, because the reviewer set that bar for findings
they had only read:

  AM-7 — folding a 100k-event log from fresh(999) instead of fresh(42),
  an unrelated genesis state, leaves the test green. The hash reaches
  only a println!; the sole assertion is on elapsed time. And the log is
  built across games seeded 42,43,44... so it could not be asserted as
  written. evidence/CB-EV-0001 reports "AM-7 replay met, 2,290x" —
  unearned on the hash-identical clause. HDN instance #6.

  K9 — making Snapshot::take discard its EventSeq and store 0, which
  destroys the half of K9 that says "+ the EventId it includes", leaves
  the test green. On the acceptance criterion GameKernel §2.4 is zero of
  three, not two of three.

Also found: K10, K14, K18 named nowhere in source; AM-10 vacuous (a K6
determinism lint reported under a D4 leak row, over a cb-*-api crate that
does not exist); AM-11's conformance suite does not exist; AM-1b measured
and omitted from the evidence scoreboard.

All twelve challenge findings conceded, none defended on its facts. v1's
prediction was rejected as unfalsifiable AND already determined; the
replacement names its refuting outcome and is beatable in both
directions.

Chaos d4=1, no override. Declaration 1 of the 12-declaration window.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 14:19:30 +02:00

8.5 KiB

2026-07-31 — response to the CB-RES-0004 challenge

All twelve findings conceded. Nothing was defended on its facts, because I re-derived every load-bearing claim myself and the reviewer was right on each one. Two findings the reviewer explicitly marked as read-only, not mutation-tested are now mutation-proven below — that was the right bar to set and it deserved to be met rather than argued around.

The survey is withdrawn and rewritten. What follows is what changed and, more importantly, the one thing the round produced that neither the survey nor the challenge had.


The mutation tests the challenge asked for

"A response that disputes findings 1 or 2 should be answered with a mutation test, not with re-reading." — challenge, sampling limit 2

I am not disputing them. I ran the mutations anyway, because a conceded finding backed by evidence is worth more to the next reader than a conceded finding backed by agreement.

Finding 1 — AM-7. Changed the fold's genesis from fresh(42) to fresh(999), so the 100k-event log is replayed into a completely different starting state:

MUTATION: fold starts from fresh(999) — a different genesis state.
test replay_probe::replay_100k_events_is_linear_and_fast ... ok
test result: ok. 1 passed; 0 failed

The test passes. AM-7's hash-identical clause detects nothing at all. The reviewer read this correctly from source; it is now demonstrated.

Finding 2 — K9. Changed Snapshot::take to discard the EventSeq it is handed and store EventSeq(0), destroying the half of K9 that says a snapshot is "the full aggregate + the EventId it includes":

MUTATION: Snapshot::take DISCARDS its EventSeq and stores 0.
test tests::snapshot_roundtrip ... ok
test result: ok. 3 passed; 0 failed

The test passes. K9's through field is unverified.

What the round actually produced

The challenge's finding 4 says the root cause is cheaper than the survey claimed: extend AM-1b's spec→code link to every numbered spec and every crate, and K10/K14/K18 fall out in one line. That is correct and it is adopted as Option E, first step of the pass.

But findings 1 and 2 prove Option E is not sufficient, and the reviewer did not connect their own findings. Sort the seven defects this round surfaced by whether a name-based check would catch them:

defect named in code? Option E catches it?
K10 replay bundle no yes
K14 commit window no yes
K18 bench-from-scenarios no yes
K9 snapshot through yes no
K11 durable format yes no
AM-7 hash-identical yes no
AM-11 conformance suite yes no

Four of seven are named in the source and unasserted. So:

Every coverage instrument in this project counts names. None counts assertions. M-D1-COV counts covers: tags. M-D1-LNK counts rule-ID strings in source. Both answer "is this rule mentioned?" and neither answers "does anything fail if this rule is violated?"

That is the correct root cause. It is neither the survey's story ("coverage scope is too narrow" — true but shallow) nor the challenge's ("the link check is pointed at one file" — true but only covers 3 of 7). It also subsumes the project's oldest error class: all five prior harness-does-nothing instances, plus AM-7 and K9 today, are the same shape — the name is there, the assertion is not.

The instrument this implies is mutation, not another grep. A rule is covered when breaking it breaks a test. That is expensive to do exhaustively and cheap to do on the acceptance rows, which is exactly where every serious error in this project has been.

The loop change this earns

The challenge's sampling limit 6 is the sharpest paragraph in the document:

"I found finding 1 only because a test had 'replay' in its name and I opened it out of curiosity — not because any systematic step pointed there. 'Reproduce the number independently' is satisfiable by re-running make coverage, which is exactly what does not find this class."

Conceded and adopted. InnerLoop §Step 2's "numbers" row currently says the reviewer must reproduce the number independently. Re-running the command that prints it satisfies that wording and finds nothing. The row must require reading the assertion behind the number, and mutating it — state what the harness would report if the property were false, then make it false and check.

The reviewer notes this is n=1 evidence for a process claim. It is n=2: CB-WP-0002's dedup blind spot had the same structure — a verification step that re-ran the author's method and inherited the author's blindness. The fix there was an assertion over all data at execution time; the fix here is mutation. Both replace re-derivation with adversarial execution.

Finding-by-finding disposition

# severity disposition
1 BLOCKING conceded, mutation-proven. AM-7 correction enters the pass scope; evidence/CB-EV-0001 needs a correction note. HDN instance #6.
2 BLOCKING conceded, mutation-proven. §2.4 is zero of three on the acceptance criterion, not two of three.
3 BLOCKING conceded fully. The prediction was unfalsifiable and already determined (0 of 18). Replaced with a prediction that names its refuting outcome.
4 BLOCKING conceded. Rescoped to K10/K14/K18; Option E adopted as step one — and shown above to be insufficient, which is the survey's real justification, not the one it gave.
5 MATERIAL conceded. "A parameterization, not a rewrite" withdrawn. The K-rule numerator must come from named executable assertions, not covers: tags. AGGREGATE becomes a list.
6 MATERIAL conceded. The AM-11 misreading is deleted. The real overclaim — M-D4-SWAP requires "the same conformance suite" and no suite exists — replaces it, and is a stronger basis for the port argument than what I wrote.
7 MATERIAL conceded with a clarification the survey owed: the durable log (K11) is shipped-runtime and AM-4a applies; the bundle writer (K10) is dev-only behind the scenarios feature and AM-4b applies. The survey wanted it both ways without saying so. Option C is now rejected on simplicity for the log, and the headroom claim is dropped for the bundle.
8 MATERIAL conceded. All four controls adopted verbatim, including the two negative controls I had not thought of (truncate-by-one-byte, mutated seed).
9 MATERIAL conceded. The exit-code path is named, and the new denominator gets the zero-rules positive control the existing arm already has.
10 MATERIAL conceded. The sweep is folded in with AM-10 as lead example. AM-10 is the better instance than AM-1: a K6 determinism lint reported under a D4 leak row, over a crate that does not exist.
11 MINOR conceded. D2's estimate corrected: Pass extended, RunOutcome::Failed restructured, initial snapshot captured, log plumbed out of execute.
12 MINOR conceded. "The number is not false" deleted. M-D1-COV is defined over "numbered spec rules"; the instrument is non-conformant with its own metric definition and 58/58 (100%) is a wrong answer to the question the metric asks.

One thing the challenge could not have found

Running its mutation tests surfaced a limit in CB-WP-0004 T01, closed two hours earlier:

$ cargo test -p games-ground ...
/bin/bash: line 22: cargo: command not found

make env-test proves every tool and every make target runs from any directory with no cargo on PATH. It does not put cargo on the agent's interactive PATH, so an ad-hoc cargo invocation still fails. T01's evidence claimed environment friction went from 85 turns to 1; that number stands for the measured window, but the fix is narrower than "no environment setup", and CB-EV-0003 should say so.

Small, and recorded rather than quietly patched, because CB-WP-0004 T05's whole finding was that tooling only recovers capacity when the manual path stops existing — and here it did not stop existing.

Approval status

Not approved. The survey is withdrawn and rewritten as research/CB-RES-0004 v2 incorporating all twelve findings, the corrected root cause, and a rescoped pass that now includes AM-7, K9, K14, K18, AM-10 and AM-11 — none of which were in the original scope.

The review cost is recorded in the pass evidence. On the two prior passes this step cost $0.66 and $1.11 and each time found an approval-blocking defect. This time it found four, one of which invalidates a committed met verdict in evidence/CB-EV-0001.