Three FATAL, five SERIOUS. The substance of round 1's corrections held — Reactive is genuinely one arm different, the five replacement controls are non-inert, the inert metric is right, the numbers reproduce. What failed were the CLAIMS about them, and two defects the corrections introduced. FATAL 1: the fix for round 1's #11 did not fix it. The assertion was `games + setup_fails == 200`, and a refused setup increments setup_fails while skipping games — so the sum is invariant under exactly the failure it claimed to catch. Injecting setup failures gave exit 0 over 196-game columns. Now asserts games == GAMES, verified to exit 101. FATAL 2: the correction to the selective-column FATAL was itself selective. "81-1000 per cell, baseline AND H1" and "31-1000" twelve lines apart, both taken from the baseline row; under H1 rank-75 arms are 59/0/0/0. Every cell is now printed rather than summarised, and the corrected verdict is the opposite of the one it replaced: under rank-75, H1 REDUCES DARVO arms to zero at 3p and above. FATAL 3: "DARVO arms 2 per seat per game" is 1 per seat per game, exactly, at every band. SERIOUS: the tiebreak oracle asserted only that the winner set CHANGED, so reversing the tiebreak left it green; the #13 defect's impact was claimed and never measured (72,000 games: zero divergences — real in principle, witnessed only by a constructed board); a 29-of-363 citation pointed at a file that did not contain it (round 1's reviewer did report it, and it was never transcribed — the record was wrong, not the number); the harnesses were run by NO GATE, so every published figure came from a manual run of an ungated binary, including the assertion added for #1; and edition-check's sibling handling — added by the last correction — was self-certifying, crashed instead of failing, and counted Markdown lines as coverage. Now discovered on disk, and it found a real gap on its first run: Rules_Text.csv vendored with no digest. Also: "peak Stress held" was dead code kept quiet by `let _ = held;` — the numbers were right by coincidence. make panels is now a registered gate. Round 3 is owed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
6.5 KiB
CB-REV-0002 — re-review of the H1 corrections
Round 2 of the tier-L review. A fresh agent, not the round-1 reviewer,
because a re-review by the same reviewer inherits its own sampling. Run
2026-08-08 against the corrections made after
CB-REV-0001.
Verdict: not approvable. Three FATAL, five SERIOUS.
The substance of round 1's corrections held. The claims about them did not, and two of the defects were introduced by the corrections.
The reviewer was briefed on the specific risk — the author of the corrections is the author of the errors — and re-derived on five independent seed windows.
What it could not break, after trying
Stated first, because it is what round 1's corrections were for.
Reactiveis genuinely one arm different. The reviewer built aPureDelegatewith the same loop shape and no override, hashed every command of every game, and got identical output toGreedyPolicyin 8/8 cells over 500 seeds, both variants. Round 1's worst finding is properly closed. Residual: nothing asserts it; it is true by construction, not by test.- All five replacement controls are non-inert, re-verified by mutation, and the SOLVE control is red on nine alternate seeds.
h1b_soothes_the_darvo_stage_attack_tooreally covers that path — disabling H1-B only for the DARVO extra Attack reddens that test alone.- The inert-arm metric is right, agreeing exactly with an independent definition in every cell of every window.
- Every number in CB-EV-0031 §1 reproduces and is stable across five windows.
score()fromworkis correct per the delta and changes no baseline path.
The three fatal ones
1. The fix for round 1's #11 did not fix it
The assertion was games + setup_fails == 200. A refused setup
increments setup_fails and skips games, so the sum is invariant under
exactly the failure it claimed to catch. Injecting setup failures gave
exit 0 and a full table over 196-game columns, under a banner reading
"200 games per cell".
Conceded. Now assert_eq!(c.games, GAMES), verified to exit 101 under
injected failures, with GAMES named once so the banner and the assertion
cannot disagree.
2. The correction to the selective-column FATAL was itself selective
CB-EV-0030's correction quoted "81–1000 per cell … baseline and H1"
and, twelve lines away, "31–1000". Both came from the baseline
rank-75 row. Under H1, rank-75 arms are 2p 59, 3p 0, 4p 0, 6p 0.
Conceded, and it is the same error twice. The remedy for choosing a favourable column chose a favourable column. Every cell is now printed rather than summarised, because two attempts at a range produced three wrong numbers — and the corrected verdict is the opposite of the one it replaced: under rank-75, H1 reduces DARVO arms to zero at 3p and above.
3. "DARVO arms 2 per seat per game" is off by a factor of two
It is 1 per seat per game, exactly, at every band: 600/(200×3), 800/(200×4), 1200/(200×6). Deterministic across five windows.
Conceded. A headline number, wrong, in a corrected section.
The serious ones
| # | challenge | outcome |
|---|---|---|
| 4 | the tiebreak test's oracle was assert_ne! — "the set changed", not "changed correctly". Reversing GR-E03's tiebreak left it green |
conceded, fixed. The oracle now names the expected winners before and after; the reviewer's M7 goes red |
| 5 | the #13 "real defect" was claimed to affect the reported modes and never measured | conceded. ~72,000 games across three modes: zero divergences, attack-value byte-identical. Real in principle, and its only witness is a constructed board. Now stated |
| 6 | "matches the reviewer's independent 29-of-363" cites a file that does not contain it | conceded, with a correction to the correction. Round 1's reviewer did report 29 of 363 — in its message, never transcribed into the review file. So the figure is real and the record was not. Recorded in CB-REV-0001 now |
| 7 | regulation.rs and attack-value.rs are run by no gate |
conceded, and it is the sharpest finding. Every number in both evidence files came from a manual run of an ungated binary — including the assertion added for #1, unreachable from make. Now make panels, registered in gates.toml |
| 8 | edition-check's sibling handling was self-certifying: it read the file list out of PROVENANCE, crashed instead of failing on an absent file, and "coverage" counted Markdown lines |
conceded — a defect introduced by the correction. Siblings are now discovered on disk; an absent recorded file FAILs; undocumented siblings FAIL. It found a real gap on its first run: Rules_Text.csv was vendored with no digest |
The minor ones
- 9 — the CSV falsifier ran over YAML. Fixed; it now filters to
.csv. - 10 — "peak Stress held" was dead code.
heldwas built, never read, andlet _ = held;kept clippy quiet; the metric was still max over payloads plus a floor. The published numbers were right by coincidence. Now actually read fromheld, and the figures are unchanged — which is how it stayed hidden. - 11 —
Reactive's doc comment still argued for the copy that round 1 ruled fatal, giving cover to restore it. Rewritten. - 12 — §3's withdrawal ground was wrong. "There is no constant to
explain" — there is:
darvo == seats × games, exactly. The constant moved from 400 ton·gamesrather than dissolving, and what forces exactly one arm per seat per game is still unexplained. - 13 — assorted, including that the delta's DARVO clause is conditional ("if they share the same path") where the test's comment says "outright".
What the reviewer could not check
Digest freshness against upstream (edition-check reports
upstream not checked out when the sibling repo is absent — it did not
have it); several untouched modules; ground-game's design note; felt-play.
The pattern, now visible across two rounds
Round 1 found five fatal errors of the form correct computation, wrong subject. Round 2 found three more, in the corrections to them — an assertion invariant under its own target, a selective fix for a selective report, and a rate wrong by a factor of two.
A correction is new work and deserves the scrutiny of new work. It arrives feeling like a conclusion, and that is exactly when it is least examined.
Round 3 is owed. Same reason as round 2.