Tier-L survey plus the full review trail. v1 was REJECTED by adversarial
review with four blocking findings; this is v2, and the finding below is
the reviewer's as much as mine.
The root cause is neither v1's story ("coverage scope is too narrow" —
true but shallow) nor the reviewer's cheaper one ("the link check points
at one file" — true but catches 3 of 7). It is:
Every coverage instrument in this project counts NAMES. None counts
ASSERTIONS. M-D1-COV counts covers: tags; M-D1-LNK counts rule-ID
strings. Both answer "is this rule mentioned?" and neither answers
"does anything fail if this rule is violated?"
Four of seven defects found this pass are named in the source and inert,
so no name-based check finds them. This subsumes the oldest error class:
all five prior harness-does-nothing instances have the same shape.
Two are mutation-proven, because the reviewer set that bar for findings
they had only read:
AM-7 — folding a 100k-event log from fresh(999) instead of fresh(42),
an unrelated genesis state, leaves the test green. The hash reaches
only a println!; the sole assertion is on elapsed time. And the log is
built across games seeded 42,43,44... so it could not be asserted as
written. evidence/CB-EV-0001 reports "AM-7 replay met, 2,290x" —
unearned on the hash-identical clause. HDN instance #6.
K9 — making Snapshot::take discard its EventSeq and store 0, which
destroys the half of K9 that says "+ the EventId it includes", leaves
the test green. On the acceptance criterion GameKernel §2.4 is zero of
three, not two of three.
Also found: K10, K14, K18 named nowhere in source; AM-10 vacuous (a K6
determinism lint reported under a D4 leak row, over a cb-*-api crate that
does not exist); AM-11's conformance suite does not exist; AM-1b measured
and omitted from the evidence scoreboard.
All twelve challenge findings conceded, none defended on its facts. v1's
prediction was rejected as unfalsifiable AND already determined; the
replacement names its refuting outcome and is beatable in both
directions.
Chaos d4=1, no override. Declaration 1 of the 12-declaration window.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
13 KiB
Challenge — CB-RES-0004 (replay bundles, durable log, and the rules nobody measures)
Reviewer: adversarial pass, InnerLoop §Step 2, one round. Separate agent
session, given the survey and the harness and the evidence file per
the "numbers" row of the review-target table.
Tree reviewed: 59b73c3.
Verdict: not approvable as written. Four BLOCKING, six MATERIAL, two
MINOR.
The survey is right that something is unmeasured. It is wrong about what, wrong about how much, and its own baseline contains an instance of the exact error class it is diagnosing. The headline finding is under-scoped by at least three rules, and the instrument it proposes is heavier and less effective than a one-line extension of an instrument that already exists in this repo.
1. BLOCKING — the AM-7 replay row does not measure the half of AM-7 the survey defends as "honest". Instance #6 of the harness-does-nothing class is live in the tree right now.
The survey §1.3 calls the AM-7 row "honest but easily misread". It is not honest. AM-7's target is "replay of 100k events ≤ 5 s, hash-identical".
games/ground/src/lib.rs:2241-2256:
let hash = state_hash_hex(&state);
let elapsed = start.elapsed();
println!("replay {} events in {:?} ({:.0} events/s), hash {}", ...);
assert!(elapsed.as_secs_f64() < 5.0, "AM-7: 100k replay under 5s");
hash reaches exactly one place: a println!. There is no assertion on
it. The Criterion bench is the same shape.
Worse, the hash could not be asserted as written. The log is built
across multiple games (lib.rs:2231-2235):
while log.len() < target {
if source.outcome.is_some() { games += 1; source = fresh(42 + games); }
so the "replay" folds events from games seeded 42, 43, 44… into a single
state built from fresh(42). It is not a replay of anything. It is a
fold-rate microbenchmark over a semantically invalid event stream,
reported as AM-7 replay | met, 2,290×.
Must change: §1.3 rewritten — AM-7's hash-identical clause is
unmeasured, the fold is over a multi-game log, and the evidence verdict is
unearned on that clause. This is a correction to evidence/CB-EV-0001,
and it puts AM-7 in the pass scope.
2. BLOCKING — the survey grants K9 as implemented. K9's acceptance property is asserted nowhere.
K9 requires "snapshot + remaining events → state must be hash-identical
to a from-genesis fold". The entire evidence is one unit test
(cb-events/src/lib.rs:158-167) that round-trips a
BTreeMap<String, u8> with EventSeq(17) as a literal. No game
aggregate, no events applied, no from-genesis comparison. And Snapshot
has no callers outside its own test.
The survey established which rules are implemented by reading pub fn
signatures — presence-of-API, which is exactly the tag-counting it
condemns in §5. On the acceptance criterion §2.4 is not "two of three",
it is zero of three.
3. BLOCKING — the "below 100%" prediction is arithmetic, not a prediction, and §6 is constructed so no outcome can refute it.
Below 100% → thesis confirmed. Near 100% → "the instrument measures tags", thesis confirmed. No stated result counts against the survey — the unfalsifiable framing the survey claims to guard against, dressed in the vocabulary of falsifiability.
It is also not a prediction: every covers: id in the corpus is GR-*
(58 ids, zero K*), and ls scenarios/ returns one directory. Any
tag-counting extension must report 0 of 18 (0%) on first run.
Computable today, before writing a line.
Must change: a prediction that can fail, with the refuting outcome named. Delete the §6 escape clause.
4. BLOCKING — K10 is not the only unimplemented kernel rule. K14 and K18 are too, and the one-line check that finds all three already exists in this repo.
spec: K1..K18
code: K1..K9 K11 K12 K13 K15 K16 K17
diff: K10, K14, K18
K14 ("one commit window per Select") — CommitWindow has zero
non-test users and games/ground does not import it. GROUND collects
selections in its own aggregate. A dead abstraction certified by a test
that exercises only itself.
K18 ("Criterion benches driving the same scenario format") — the
bench hardcodes commands in Rust and never touches ScenarioFile. And
ls -R benchmarks contains only baselines/model-prices.toml.
Why this is BLOCKING. The check that found these is a two-line
generalization of M-D1-LNK / AM-1b, which this repo already runs every
make coverage. Pointing parse_code_ids at specs/GameKernel.md ×
crates/** would have surfaced all three the day AM-1b shipped. The
survey diagnoses a gap a cheap existing instrument covers and prescribes a
new expensive capability as the remedy — and, having chosen the expensive
framing, found 1 of the 3 gaps the cheap check finds in one line.
Must change: rescope to K10/K14/K18 (K9, K11 partial). Add Option E: extend AM-1b's spec→code link to every numbered spec and every crate, and argue why the tier-L work is warranted on top of it.
5. MATERIAL — K-rules cannot be covered by the mechanism that covers GR-rules. "A parameterization, not a rewrite" is false.
The M-D1-COV numerator comes from covers: lists consumed by a runner
requiring ScenarioGame: Aggregate + Serialize + DeserializeOwned, with
dispatch by game prefix and ground the only arm. There is no kernel
aggregate, setup preset, or command vocabulary. scenarios/kernel/*.yaml
would be a tag in a directory the runner cannot dispatch — the inflation
§5 forbids. Second blocker: AGGREGATE = "games/ground/src/lib.rs" is
one file, so every K-rule reports "unlinked" forever.
6. MATERIAL — the AM-11 argument, on which tier L rests, misreads a target as a claim and misses the actual overclaim next to it.
Nobody claims what the survey says. GameKernel.md:175 is the target
column, and the evidence reports AM-11 impl pairs | 1 of 1 (KernelRng) | met, narrow, adding "will mean more once storage has one". The evidence
already says what the survey accuses it of hiding.
The real overclaim is one clause away: M-D4-SWAP requires impls
"passing the same conformance suite", and no conformance suite exists
— one doc comment describing future work. The RNG pair is exercised by two
separate non-shared tests. The met, narrow verdict is unearned on the
suite clause.
7. MATERIAL — Option C is rejected against the wrong budget line.
AM-4a (1.5% headroom) is the shipped-runtime budget. The bundle is
written by the scenario runner on failure, and scenario tooling is
explicitly dev-only. A store behind the scenarios feature is charged to
AM-4b, headroom 32,979 lines (9.4%) — nine times larger, and
comfortably above redb-class crates. The conclusion may survive on
simplicity, but ADR-0004 exists precisely to stop AM-4 numbers being used
as post-hoc justification, and this is that.
8. MATERIAL — what make replay-test prints if it does nothing: ok.
Four vacuous passes, none controlled for:
| silent failure | what it prints |
|---|---|
| no scenario fails (all 21 pass today) → zero bundles round-tripped | 0 bundles verified — ok |
| writer emits nothing, reader tolerates; identical error strings satisfy "same failure" | PASS — failure reproduced |
replayed hash compared to one recomputed in the same process — assert_eq!(h, h), the AM-7 defect exactly |
PASS — bit-identical |
| truncation path never exercised; K11's operative clause is dead code | PASS |
Required controls: a deliberately-failing fixture scenario plus
bundles > 0; the comparison hash read out of the bundle, not
recomputed; truncate-by-one-byte and corrupt-length-prefix negative
controls; and a mutated-seed negative control so the round-trip can fail.
9. MATERIAL — the extended make coverage contradicts §5 on day one.
rule-coverage.py:195 returns 2 when missing is non-empty, and
coverage is in all with no - prefix. Adding 18 uncovered K-rules
fails make all immediately, which §5 bullet 2 forbids. Requires a
separate kernel denominator outside the exit code, plus a decision on when
it starts binding.
For the record: the existing arm cannot silently do nothing —
rule-coverage.py:121-129 refuses to report over zero rules, and that
control is self-tested. The new denominator must replicate it, or a
kernel regex matching nothing reports 0/0 (100%).
10. MATERIAL — the survey says the scoped-denominator problem is general, then sweeps in one place. I swept.
(a) AM-10 is vacuous. M-D4-LEAK: foreign types in cb-*-api-visible signatures | 0. There is no cb-*-api crate — true over an empty set.
What is measured is a clippy.toml deny of HashMap/HashSet citing
K6 (determinism), reported under a D4 leak row. Stronger than
AM-1 as an instance of the survey's own thesis.
(b) AM-1b is measured and dropped from the scoreboard. make coverage
prints 49/58 and MetricsAndScenarios records it Unmet;
grep -c 'AM-1b' evidence/CB-EV-0001 returns 0. The artifact carrying
the headline reports the flattering half and omits the unflattering half
the same command prints two lines below.
(c) M-D1-LNK's own denominator is one file.
(d) AM-2, AM-3, AM-5, AM-9 are unreported and unmentioned by the survey.
Must change: generalize §1.2 to a swept table, lead with (a). The history file's "n=1" weakness can be retired: it is n≥4 in this repo.
11. MINOR — D2 understates the bundle writer.
scenario.rs:139 creates EventLog, appends at :157, and never reads
it — not returned in Pass, not iterated, dropped at end of execute.
The one production instantiation of the K11 log is a write-only sink.
Pass carries end state, not an initial snapshot. Failures are a
formatted String, not structured expected-vs-actual. The bundle needs a
change to the runner's data flow, not "a directory of four files".
12. MINOR — "the number is not false" is itself false.
M-D1-COV is defined as "numbered spec rules covered by ≥1 passing
scenario", and covers is documented as "numbered rules from the
capability spec" with a non-GR example. The instrument is non-conformant
with its own metric definition; 58/58 (100%) is a wrong answer to the
question M-D1-COV asks. The concession "the number is not false" gives
away the finding — it lets a future reader conclude only the reading was
careless.
What I could not break
- No write path in the workspace. Wider idiom set than the author's
(
OpenOptions,create_dir,tempfile,to_writer,std::io::Write) over the whole tree including#[cfg(test)]and benches. Every hit is aDisplay::fmtwrite!orread_to_string. Holds. - No
--replayflag and no flag parsing. Read all 88 lines ofcb-sim/src/main.rs. The survey's stronger phrasing is accurate. - No
.cbreplayartifact, reader, writer, or directory. Confirmed across.rs,.py,Makefile, CI. - 18 K-rules, zero in coverage scope. Verified two ways;
RULE_REand input read from source, not inferred from output. - AM-4a headroom is 1.5%. Number right; applicability disputed (7).
make coveragereports 58/58 (100%) over 21 scenarios. Reproduced.- No counter-example to the absence claims. Hunted for a K-rule in a scenario, a replay entry behind a feature flag, a bundle writer in CI, a replay round-trip test. None exists.
Sampling limits of this review
- The "no write path" claim has one sample and my check does not fix
that. The workspace contains ten
.rsfiles total; the author's sample is the population. My independence is in the idiom list and in reading whole files, not in file selection. This is the CB-WP-0002 failure shape — two checks, one population — and I cannot make it otherwise from inside this repo. - I did not execute the test suite.
cargo test -- --listenumerates 19 tests; I ran none. Findings 1 and 2 come from reading source, not from mutation testing. A response disputing them should answer with a mutation test, not re-reading. - Finding 4 rests on grep for rule IDs — the same tag-counting I criticize. K14 and K18 have independent structural confirmation. "K10/K14/K18 are untracked" is solid; "they are the only further unimplemented K-rules" is not a claim I make. I did not audit K1–K9, K12–K17 against their acceptance text one by one. Given K9 and K14 both failed on close reading, a full audit should find more.
- I did not verify the §2 baseline table. Still cited, by both of us.
- AM-12 / CostAccounting were out of scope. I did not audit
cb-cost.pyor the price sheet. - One process limit. I found finding 1 only because a test had
"replay" in its name and I opened it out of curiosity — not because
any systematic step pointed there. "Reproduce the number independently"
is satisfiable by re-running
make coverage, which is exactly what does not find this class. The procedure should require reading the assertion behind every quoted acceptance number, not re-running the command that prints it. Offered as a loop change; n=1 evidence for a process claim, the same weakness I charged the survey with in 10.