The pattern is now measured over three rounds: 5 fatal, then 3 (2 from the previous round's corrections), then 4 (3 from them). The corrections are not getting safer. FATAL 1: round 2's short-cell assertion went into regulation.rs only. attack-value.rs — which produced every number in CB-EV-0030's DARVO table — still just warned, and the gate registered to close the finding claimed the property for both. FATAL 2, the sharpest of the three rounds: counting games proves they STARTED. Stopping the engine after one round gives 200 games, all-zero columns and exit 0 — byte for byte the signature CB-EV-0030 says the instrumentation distinguishes from a real result. Both harnesses now require every counted game to have reached an outcome over five rounds. FATAL 3: round 2's `.csv` filter was applied to all three loops, so catalog.yaml and rules_delta.yaml — whose missing digests were round 1's finding — were recorded and then never compared, and never checked against upstream at all. Only the parser loop filters now. FATAL 4: five of six tiebreak comparators had no coverage. GR-E04's tiebreak never executes in any scenario. All four are now covered and mutation-verified; the Blame key needed compensating claims to be reachable at all, since Blame also lowers the coalition score. SERIOUS: "peak held" computed the same number as "peak assigned" for every possible input — the real gap was that START_STRESS was an unchecked constant, now read off the dealt state; cadence="none" was a pure loophole, removed; sibling discovery swapped a hand-written list for hand-written globs and missed metadata.json and VARIANT.md, both named in the package's own changed_files — now walked, and it found them immediately; and "~72,000 games" was unsourced, make panels runs 17,600. Also separated two kinds of number that were presented alike: seats×games is invariant, 363 and 29 vary 7.1%-11.5% across samples. Round 4 owed. The conclusion is not that the work is nearly right — it is that author-made corrections to measurement work should be assumed defective until a fresh reader has attacked them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
6.3 KiB
CB-REV-0003 — round 3
Fresh agent again. Run 2026-08-08 against the corrections made after
CB-REV-0002.
Verdict: not approvable. Four FATAL, five SERIOUS, two MINOR — and three of the four FATAL were created or left by round 2's corrections.
The pattern is now established over three rounds and is the most useful thing this review has produced.
| round | challenges | FATAL | of which introduced by the previous round's fix |
|---|---|---|---|
| 1 | 13 | 5 | — |
| 2 | 8 | 3 | 2 |
| 3 | 11 | 4 | 3 |
The four fatal ones
1. The fix for round 2's #1 was applied to one of the two harnesses
regulation.rs got assert_eq!(games, GAMES). attack-value.rs — which
produced every number in CB-EV-0030's DARVO table — kept if games != 200 { eprintln!(..) }. Injected setup refusals gave exit 0 and a full table
over 195-game columns.
And the gate registered to close the finding asserted the property for
both. Its checks line was false of half of what it gates.
Conceded. Both harnesses assert now; the checks line says what is
actually checked.
2. Counting games proves they STARTED, not that they were PLAYED
Stopping the engine after one round (round >= 5 → >= 1) gives 200
games in every cell, every value zero, exit 0, make panels green —
byte for byte the signature CB-EV-0030 §3 claims the instrumentation
distinguishes from a real result.
Conceded, and it is the sharpest finding of the three rounds. The
counter answered "did play return Ok" while the claim made of it was
"the zeros are real". Both harnesses now require every counted game to
have reached an outcome over five rounds; the one-round mutation fails
with the right message.
3. Round 2's .csv filter disabled the checks it was added beside
The filter was applied to all three loops — digest comparison, the CSV
parser check, and upstream freshness. So catalog.yaml and
rules_delta.yaml, whose missing digests were round 1's finding, were
recorded and then never compared, and never checked against upstream
at all. Tampering with both produced [ok] × 12, exit 0.
Conceded. Only the parser loop filters; tampering with either YAML now fails on digest and freshness.
4. Five of six tiebreak comparators had no coverage
The strengthened oracle covered GR-E03's Stress key. GR-E03's Bond key
and both of GR-E04's keys could be reversed with 179 tests and 26
scenarios green — GR-E04's tiebreak never executes in any scenario,
because the only GR-E04 scenario sets group_success: false.
Conceded. All four are now covered and each was verified red under
mutation. The Blame key took care to reach: a coalition's score is
sum(claimed − blame), so a Blame token lowers the score and key 1
decides first — the key is only reachable when claims compensate, and the
test asserts the scores tie before relying on it.
The serious ones
| # | challenge | outcome |
|---|---|---|
| 5 | "peak Stress held, not assigned" — peak_held ≡ max(start, peak_payload) for every possible input, so the correction computed the same number |
conceded. The fix was cosmetic and the comment was false of the new code too. The real coupling was missing, which is #6 |
| 6 | START_STRESS was an unchecked constant — setting it to 7 printed peak 7 everywhere |
conceded. Now read off the dealt state. Verified: changing setup's Stress to 3 moves the reported peak to 3 |
| 7 | cadence = "none" was a pure loophole — a real target that runs nowhere and passes |
conceded, removed. It had one user, the empty-target CHAOS entry, which the loop already skips. manual remains self-declared and unchecked, and that is now stated rather than implied |
| 8 | sibling discovery swapped a hand-written list for three hand-written globs; metadata.json and VARIANT.md were invisible, and both are named in the package's own changed_files |
conceded. Walked, not globbed. It found both immediately, plus nested files the globs could never reach |
| 9 | "~72,000 games" is not producible; make panels runs 17,600 |
conceded. Adopted from a reviewer's message and never re-derived — in the file whose correction history is about exactly that. The conclusion (zero divergence) reproduces |
The minor ones
- 10 — two kinds of number were presented alike.
600/800/1200is exactlyseats × gameson every sample;363and29vary — 7.1%–11.5% inert across four windows. "92% of arms are live" is a fact about seeds0..200. Now separated. - 11 — the 29/363 citation, wrong twice, is now stated as one sample
of a varying quantity. Both "independent confirmations" used the same
seeds, so what was confirmed was the definition. Whether round 1's
StrictReactivewas the corrected policy cannot be determined: that reviewer's code was never kept, which is itself worth fixing.
What held, after being attacked
Every one of the 24 cells in CB-EV-0030's DARVO table reproduces
exactly — the first of three attempts at that table to produce no wrong
number. regulation's assert_eq!(c.games, GAMES) catches both the setup
and the play path. The #13 fix is genuinely controlled. H1-B's mutation
coverage is real. make all runs panels and fails on it. The
1-arm-per-seat-per-game invariant holds on every sample — and is still
unexplained.
What could not be checked
Upstream freshness (../ground-game not checked out, so that half of
edition-check has never run here); whether metadata.json and
VARIANT.md are load-bearing to anything; round 1's StrictReactive;
felt-play; cost.
The finding that outlasts H1
Three rounds, each correcting the last, each introducing defects of the same class. The corrections are not getting safer.
A correction is written under the belief that the error is now understood. That belief is the condition under which this class of error is produced — so the correction inherits it, and the next round finds the same shape one level in.
Round 4 is owed by the same argument. The honest conclusion is not that the work is nearly right; it is that author-made corrections to measurement work should be assumed defective until a fresh reader has attacked them, and this project should stop treating "corrected" as a state closer to done than "found wrong".