Some checks failed
ci / check (push) Failing after 3s
The pattern is now measured over three rounds: 5 fatal, then 3 (2 from the previous round's corrections), then 4 (3 from them). The corrections are not getting safer. FATAL 1: round 2's short-cell assertion went into regulation.rs only. attack-value.rs — which produced every number in CB-EV-0030's DARVO table — still just warned, and the gate registered to close the finding claimed the property for both. FATAL 2, the sharpest of the three rounds: counting games proves they STARTED. Stopping the engine after one round gives 200 games, all-zero columns and exit 0 — byte for byte the signature CB-EV-0030 says the instrumentation distinguishes from a real result. Both harnesses now require every counted game to have reached an outcome over five rounds. FATAL 3: round 2's `.csv` filter was applied to all three loops, so catalog.yaml and rules_delta.yaml — whose missing digests were round 1's finding — were recorded and then never compared, and never checked against upstream at all. Only the parser loop filters now. FATAL 4: five of six tiebreak comparators had no coverage. GR-E04's tiebreak never executes in any scenario. All four are now covered and mutation-verified; the Blame key needed compensating claims to be reachable at all, since Blame also lowers the coalition score. SERIOUS: "peak held" computed the same number as "peak assigned" for every possible input — the real gap was that START_STRESS was an unchecked constant, now read off the dealt state; cadence="none" was a pure loophole, removed; sibling discovery swapped a hand-written list for hand-written globs and missed metadata.json and VARIANT.md, both named in the package's own changed_files — now walked, and it found them immediately; and "~72,000 games" was unsourced, make panels runs 17,600. Also separated two kinds of number that were presented alike: seats×games is invariant, 363 and 29 vary 7.1%-11.5% across samples. Round 4 owed. The conclusion is not that the work is nearly right — it is that author-made corrections to measurement work should be assumed defective until a fresh reader has attacked them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
127 lines
6.3 KiB
Markdown
127 lines
6.3 KiB
Markdown
# CB-REV-0003 — round 3
|
||
|
||
Fresh agent again. Run 2026-08-08 against the corrections made after
|
||
[`CB-REV-0002`](CB-REV-0002-h1.md).
|
||
|
||
> **Verdict: not approvable. Four FATAL, five SERIOUS, two MINOR — and
|
||
> three of the four FATAL were created or left by round 2's corrections.**
|
||
>
|
||
> **The pattern is now established over three rounds** and is the most
|
||
> useful thing this review has produced.
|
||
|
||
| round | challenges | FATAL | of which introduced by the previous round's fix |
|
||
|---|---:|---:|---:|
|
||
| 1 | 13 | 5 | — |
|
||
| 2 | 8 | 3 | 2 |
|
||
| 3 | 11 | 4 | 3 |
|
||
|
||
---
|
||
|
||
## The four fatal ones
|
||
|
||
### 1. The fix for round 2's #1 was applied to one of the two harnesses
|
||
|
||
`regulation.rs` got `assert_eq!(games, GAMES)`. **`attack-value.rs` — which
|
||
produced every number in CB-EV-0030's DARVO table — kept `if games != 200
|
||
{ eprintln!(..) }`.** Injected setup refusals gave exit 0 and a full table
|
||
over 195-game columns.
|
||
|
||
**And the gate registered to close the finding asserted the property for
|
||
both.** Its `checks` line was false of half of what it gates.
|
||
|
||
**Conceded.** Both harnesses assert now; the `checks` line says what is
|
||
actually checked.
|
||
|
||
### 2. Counting games proves they STARTED, not that they were PLAYED
|
||
|
||
Stopping the engine after one round (`round >= 5` → `>= 1`) gives **200
|
||
games in every cell, every value zero, exit 0, `make panels` green** —
|
||
byte for byte the signature CB-EV-0030 §3 claims the instrumentation
|
||
distinguishes from a real result.
|
||
|
||
**Conceded, and it is the sharpest finding of the three rounds.** The
|
||
counter answered *"did `play` return `Ok`"* while the claim made of it was
|
||
*"the zeros are real"*. Both harnesses now require every counted game to
|
||
have reached an outcome over five rounds; the one-round mutation fails
|
||
with the right message.
|
||
|
||
### 3. Round 2's `.csv` filter disabled the checks it was added beside
|
||
|
||
The filter was applied to **all three** loops — digest comparison, the CSV
|
||
parser check, and upstream freshness. So `catalog.yaml` and
|
||
`rules_delta.yaml`, whose missing digests were round 1's finding, were
|
||
recorded and then **never compared**, and never checked against upstream
|
||
at all. Tampering with both produced `[ok] × 12, exit 0`.
|
||
|
||
**Conceded.** Only the parser loop filters; tampering with either YAML now
|
||
fails on digest *and* freshness.
|
||
|
||
### 4. Five of six tiebreak comparators had no coverage
|
||
|
||
The strengthened oracle covered GR-E03's Stress key. **GR-E03's Bond key
|
||
and both of GR-E04's keys could be reversed with 179 tests and 26
|
||
scenarios green** — GR-E04's tiebreak never executes in any scenario,
|
||
because the only GR-E04 scenario sets `group_success: false`.
|
||
|
||
**Conceded.** All four are now covered and each was verified red under
|
||
mutation. **The Blame key took care to reach**: a coalition's score is
|
||
`sum(claimed − blame)`, so a Blame token lowers the score and key 1
|
||
decides first — the key is only reachable when claims compensate, and the
|
||
test asserts the scores tie before relying on it.
|
||
|
||
## The serious ones
|
||
|
||
| # | challenge | outcome |
|
||
|---|---|---|
|
||
| 5 | "peak Stress **held**, not assigned" — `peak_held ≡ max(start, peak_payload)` for every possible input, so the correction computed the same number | **conceded.** The fix was cosmetic and the comment was false of the new code too. The real coupling was missing, which is #6 |
|
||
| 6 | `START_STRESS` was an unchecked constant — setting it to 7 printed `peak 7` everywhere | **conceded.** Now read off the dealt state. Verified: changing setup's Stress to 3 moves the reported peak to 3 |
|
||
| 7 | `cadence = "none"` was a pure loophole — a real target that runs nowhere and passes | **conceded, removed.** It had one user, the empty-target CHAOS entry, which the loop already skips. `manual` remains self-declared and unchecked, and that is now stated rather than implied |
|
||
| 8 | sibling discovery swapped a hand-written list for three hand-written globs; `metadata.json` and `VARIANT.md` were invisible, and both are named in the package's own `changed_files` | **conceded.** Walked, not globbed. **It found both immediately**, plus nested files the globs could never reach |
|
||
| 9 | "~72,000 games" is not producible; `make panels` runs **17,600** | **conceded.** Adopted from a reviewer's message and never re-derived — in the file whose correction history is about exactly that. The conclusion (zero divergence) reproduces |
|
||
|
||
## The minor ones
|
||
|
||
- **10 — two kinds of number were presented alike.** `600/800/1200` is
|
||
exactly `seats × games` on every sample; **`363` and `29` vary** —
|
||
7.1%–11.5% inert across four windows. "92% of arms are live" is a fact
|
||
about seeds `0..200`. Now separated.
|
||
- **11 — the 29/363 citation, wrong twice, is now stated as one sample**
|
||
of a varying quantity. Both "independent confirmations" used the same
|
||
seeds, so what was confirmed was the *definition*. Whether round 1's
|
||
`StrictReactive` was the corrected policy cannot be determined: **that
|
||
reviewer's code was never kept**, which is itself worth fixing.
|
||
|
||
## What held, after being attacked
|
||
|
||
**Every one of the 24 cells in CB-EV-0030's DARVO table reproduces
|
||
exactly** — the first of three attempts at that table to produce no wrong
|
||
number. `regulation`'s `assert_eq!(c.games, GAMES)` catches both the setup
|
||
and the `play` path. The #13 fix is genuinely controlled. H1-B's mutation
|
||
coverage is real. `make all` runs `panels` and fails on it. The
|
||
1-arm-per-seat-per-game invariant holds on every sample — **and is still
|
||
unexplained.**
|
||
|
||
## What could not be checked
|
||
|
||
Upstream freshness (`../ground-game` not checked out, so that half of
|
||
`edition-check` has never run here); whether `metadata.json` and
|
||
`VARIANT.md` are load-bearing to anything; round 1's `StrictReactive`;
|
||
felt-play; cost.
|
||
|
||
---
|
||
|
||
## The finding that outlasts H1
|
||
|
||
Three rounds, each correcting the last, each introducing defects of the
|
||
same class. **The corrections are not getting safer.**
|
||
|
||
> A correction is written under the belief that the error is now
|
||
> understood. That belief is the condition under which this class of
|
||
> error is produced — so the correction inherits it, and the next round
|
||
> finds the same shape one level in.
|
||
|
||
**Round 4 is owed by the same argument.** The honest conclusion is not
|
||
that the work is nearly right; it is that **author-made corrections to
|
||
measurement work should be assumed defective until a fresh reader has
|
||
attacked them**, and this project should stop treating "corrected" as a
|
||
state closer to done than "found wrong".
|