clay-borg/reviews/CB-REV-0003-h1.md
tegwick c4a8a227c0
Some checks failed
ci / check (push) Failing after 3s
CB-REV-0003: round 3, and three of four FATAL came from round 2's fixes
The pattern is now measured over three rounds: 5 fatal, then 3 (2 from the
previous round's corrections), then 4 (3 from them). The corrections are
not getting safer.

FATAL 1: round 2's short-cell assertion went into regulation.rs only.
attack-value.rs — which produced every number in CB-EV-0030's DARVO table
— still just warned, and the gate registered to close the finding claimed
the property for both.

FATAL 2, the sharpest of the three rounds: counting games proves they
STARTED. Stopping the engine after one round gives 200 games, all-zero
columns and exit 0 — byte for byte the signature CB-EV-0030 says the
instrumentation distinguishes from a real result. Both harnesses now
require every counted game to have reached an outcome over five rounds.

FATAL 3: round 2's `.csv` filter was applied to all three loops, so
catalog.yaml and rules_delta.yaml — whose missing digests were round 1's
finding — were recorded and then never compared, and never checked against
upstream at all. Only the parser loop filters now.

FATAL 4: five of six tiebreak comparators had no coverage. GR-E04's
tiebreak never executes in any scenario. All four are now covered and
mutation-verified; the Blame key needed compensating claims to be
reachable at all, since Blame also lowers the coalition score.

SERIOUS: "peak held" computed the same number as "peak assigned" for every
possible input — the real gap was that START_STRESS was an unchecked
constant, now read off the dealt state; cadence="none" was a pure
loophole, removed; sibling discovery swapped a hand-written list for
hand-written globs and missed metadata.json and VARIANT.md, both named in
the package's own changed_files — now walked, and it found them
immediately; and "~72,000 games" was unsourced, make panels runs 17,600.

Also separated two kinds of number that were presented alike: seats×games
is invariant, 363 and 29 vary 7.1%-11.5% across samples.

Round 4 owed. The conclusion is not that the work is nearly right — it is
that author-made corrections to measurement work should be assumed
defective until a fresh reader has attacked them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 10:26:25 +02:00

6.3 KiB
Raw Blame History

CB-REV-0003 — round 3

Fresh agent again. Run 2026-08-08 against the corrections made after CB-REV-0002.

Verdict: not approvable. Four FATAL, five SERIOUS, two MINOR — and three of the four FATAL were created or left by round 2's corrections.

The pattern is now established over three rounds and is the most useful thing this review has produced.

round challenges FATAL of which introduced by the previous round's fix
1 13 5
2 8 3 2
3 11 4 3

The four fatal ones

1. The fix for round 2's #1 was applied to one of the two harnesses

regulation.rs got assert_eq!(games, GAMES). attack-value.rs — which produced every number in CB-EV-0030's DARVO table — kept if games != 200 { eprintln!(..) }. Injected setup refusals gave exit 0 and a full table over 195-game columns.

And the gate registered to close the finding asserted the property for both. Its checks line was false of half of what it gates.

Conceded. Both harnesses assert now; the checks line says what is actually checked.

2. Counting games proves they STARTED, not that they were PLAYED

Stopping the engine after one round (round >= 5>= 1) gives 200 games in every cell, every value zero, exit 0, make panels green — byte for byte the signature CB-EV-0030 §3 claims the instrumentation distinguishes from a real result.

Conceded, and it is the sharpest finding of the three rounds. The counter answered "did play return Ok" while the claim made of it was "the zeros are real". Both harnesses now require every counted game to have reached an outcome over five rounds; the one-round mutation fails with the right message.

3. Round 2's .csv filter disabled the checks it was added beside

The filter was applied to all three loops — digest comparison, the CSV parser check, and upstream freshness. So catalog.yaml and rules_delta.yaml, whose missing digests were round 1's finding, were recorded and then never compared, and never checked against upstream at all. Tampering with both produced [ok] × 12, exit 0.

Conceded. Only the parser loop filters; tampering with either YAML now fails on digest and freshness.

4. Five of six tiebreak comparators had no coverage

The strengthened oracle covered GR-E03's Stress key. GR-E03's Bond key and both of GR-E04's keys could be reversed with 179 tests and 26 scenarios green — GR-E04's tiebreak never executes in any scenario, because the only GR-E04 scenario sets group_success: false.

Conceded. All four are now covered and each was verified red under mutation. The Blame key took care to reach: a coalition's score is sum(claimed blame), so a Blame token lowers the score and key 1 decides first — the key is only reachable when claims compensate, and the test asserts the scores tie before relying on it.

The serious ones

# challenge outcome
5 "peak Stress held, not assigned" — peak_held ≡ max(start, peak_payload) for every possible input, so the correction computed the same number conceded. The fix was cosmetic and the comment was false of the new code too. The real coupling was missing, which is #6
6 START_STRESS was an unchecked constant — setting it to 7 printed peak 7 everywhere conceded. Now read off the dealt state. Verified: changing setup's Stress to 3 moves the reported peak to 3
7 cadence = "none" was a pure loophole — a real target that runs nowhere and passes conceded, removed. It had one user, the empty-target CHAOS entry, which the loop already skips. manual remains self-declared and unchecked, and that is now stated rather than implied
8 sibling discovery swapped a hand-written list for three hand-written globs; metadata.json and VARIANT.md were invisible, and both are named in the package's own changed_files conceded. Walked, not globbed. It found both immediately, plus nested files the globs could never reach
9 "~72,000 games" is not producible; make panels runs 17,600 conceded. Adopted from a reviewer's message and never re-derived — in the file whose correction history is about exactly that. The conclusion (zero divergence) reproduces

The minor ones

  • 10 — two kinds of number were presented alike. 600/800/1200 is exactly seats × games on every sample; 363 and 29 vary — 7.1%11.5% inert across four windows. "92% of arms are live" is a fact about seeds 0..200. Now separated.
  • 11 — the 29/363 citation, wrong twice, is now stated as one sample of a varying quantity. Both "independent confirmations" used the same seeds, so what was confirmed was the definition. Whether round 1's StrictReactive was the corrected policy cannot be determined: that reviewer's code was never kept, which is itself worth fixing.

What held, after being attacked

Every one of the 24 cells in CB-EV-0030's DARVO table reproduces exactly — the first of three attempts at that table to produce no wrong number. regulation's assert_eq!(c.games, GAMES) catches both the setup and the play path. The #13 fix is genuinely controlled. H1-B's mutation coverage is real. make all runs panels and fails on it. The 1-arm-per-seat-per-game invariant holds on every sample — and is still unexplained.

What could not be checked

Upstream freshness (../ground-game not checked out, so that half of edition-check has never run here); whether metadata.json and VARIANT.md are load-bearing to anything; round 1's StrictReactive; felt-play; cost.


The finding that outlasts H1

Three rounds, each correcting the last, each introducing defects of the same class. The corrections are not getting safer.

A correction is written under the belief that the error is now understood. That belief is the condition under which this class of error is produced — so the correction inherits it, and the next round finds the same shape one level in.

Round 4 is owed by the same argument. The honest conclusion is not that the work is nearly right; it is that author-made corrections to measurement work should be assumed defective until a fresh reader has attacked them, and this project should stop treating "corrected" as a state closer to done than "found wrong".