clay-borg/reviews/CB-REV-0002-h1.md
tegwick da58e78e4a
Some checks failed
ci / check (push) Failing after 4s
CB-REV-0002: round 2, and the corrections were not approvable either
Three FATAL, five SERIOUS. The substance of round 1's corrections held —
Reactive is genuinely one arm different, the five replacement controls are
non-inert, the inert metric is right, the numbers reproduce. What failed
were the CLAIMS about them, and two defects the corrections introduced.

FATAL 1: the fix for round 1's #11 did not fix it. The assertion was
`games + setup_fails == 200`, and a refused setup increments setup_fails
while skipping games — so the sum is invariant under exactly the failure
it claimed to catch. Injecting setup failures gave exit 0 over 196-game
columns. Now asserts games == GAMES, verified to exit 101.

FATAL 2: the correction to the selective-column FATAL was itself
selective. "81-1000 per cell, baseline AND H1" and "31-1000" twelve lines
apart, both taken from the baseline row; under H1 rank-75 arms are
59/0/0/0. Every cell is now printed rather than summarised, and the
corrected verdict is the opposite of the one it replaced: under rank-75,
H1 REDUCES DARVO arms to zero at 3p and above.

FATAL 3: "DARVO arms 2 per seat per game" is 1 per seat per game, exactly,
at every band.

SERIOUS: the tiebreak oracle asserted only that the winner set CHANGED, so
reversing the tiebreak left it green; the #13 defect's impact was claimed
and never measured (72,000 games: zero divergences — real in principle,
witnessed only by a constructed board); a 29-of-363 citation pointed at a
file that did not contain it (round 1's reviewer did report it, and it was
never transcribed — the record was wrong, not the number); the harnesses
were run by NO GATE, so every published figure came from a manual run of
an ungated binary, including the assertion added for #1; and edition-check's
sibling handling — added by the last correction — was self-certifying,
crashed instead of failing, and counted Markdown lines as coverage. Now
discovered on disk, and it found a real gap on its first run: Rules_Text.csv
vendored with no digest.

Also: "peak Stress held" was dead code kept quiet by `let _ = held;` — the
numbers were right by coincidence.

make panels is now a registered gate. Round 3 is owed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 09:44:00 +02:00

120 lines
6.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# CB-REV-0002 — re-review of the H1 corrections
Round 2 of the tier-L review. A **fresh** agent, not the round-1 reviewer,
because a re-review by the same reviewer inherits its own sampling. Run
2026-08-08 against the corrections made after
[`CB-REV-0001`](CB-REV-0001-h1.md).
> **Verdict: not approvable.** Three FATAL, five SERIOUS.
>
> **The substance of round 1's corrections held. The claims about them did
> not**, and two of the defects were introduced *by* the corrections.
The reviewer was briefed on the specific risk — **the author of the
corrections is the author of the errors** — and re-derived on five
independent seed windows.
---
## What it could not break, after trying
Stated first, because it is what round 1's corrections were for.
- **`Reactive` is genuinely one arm different.** The reviewer built a
`PureDelegate` with the same loop shape and *no* override, hashed every
command of every game, and got **identical output to `GreedyPolicy` in
8/8 cells over 500 seeds**, both variants. Round 1's worst finding is
properly closed. *Residual: nothing asserts it; it is true by
construction, not by test.*
- **All five replacement controls are non-inert**, re-verified by
mutation, and the SOLVE control is red on nine alternate seeds.
- **`h1b_soothes_the_darvo_stage_attack_too` really covers that path** —
disabling H1-B *only* for the DARVO extra Attack reddens that test alone.
- **The inert-arm metric is right**, agreeing exactly with an independent
definition in every cell of every window.
- **Every number in CB-EV-0031 §1 reproduces** and is stable across five
windows.
- **`score()` from `work` is correct per the delta and changes no baseline
path.**
## The three fatal ones
### 1. The fix for round 1's #11 did not fix it
The assertion was `games + setup_fails == 200`. **A refused setup
increments `setup_fails` and skips `games`, so the sum is invariant under
exactly the failure it claimed to catch.** Injecting setup failures gave
exit 0 and a full table over 196-game columns, under a banner reading
"200 games per cell".
**Conceded.** Now `assert_eq!(c.games, GAMES)`, verified to exit 101 under
injected failures, with `GAMES` named once so the banner and the assertion
cannot disagree.
### 2. The correction to the selective-column FATAL was itself selective
CB-EV-0030's correction quoted *"811000 per cell … baseline **and** H1"*
and, twelve lines away, *"311000"*. **Both came from the baseline
rank-75 row.** Under H1, rank-75 arms are `2p 59, 3p 0, 4p 0, 6p 0`.
**Conceded, and it is the same error twice.** The remedy for choosing a
favourable column chose a favourable column. **Every cell is now printed**
rather than summarised, because two attempts at a range produced three
wrong numbers — and the corrected verdict is the opposite of the one it
replaced: under rank-75, **H1 reduces DARVO arms to zero** at 3p and above.
### 3. "DARVO arms 2 per seat per game" is off by a factor of two
It is **1 per seat per game**, exactly, at every band: 600/(200×3),
800/(200×4), 1200/(200×6). Deterministic across five windows.
**Conceded.** A headline number, wrong, in a corrected section.
## The serious ones
| # | challenge | outcome |
|---|---|---|
| 4 | the tiebreak test's oracle was `assert_ne!` — "the set changed", not "changed correctly". Reversing GR-E03's tiebreak left it green | **conceded, fixed.** The oracle now names the expected winners before and after; the reviewer's M7 goes red |
| 5 | the #13 "real defect" was claimed to affect the reported modes and never measured | **conceded.** ~72,000 games across three modes: **zero divergences**, `attack-value` byte-identical. Real in principle, and its only witness is a constructed board. Now stated |
| 6 | "matches the reviewer's independent 29-of-363" cites a file that does not contain it | **conceded, with a correction to the correction.** Round 1's reviewer *did* report 29 of 363 — in its message, never transcribed into the review file. So the figure is real and the **record** was not. Recorded in CB-REV-0001 now |
| 7 | `regulation.rs` and `attack-value.rs` are run by **no gate** | **conceded, and it is the sharpest finding.** Every number in both evidence files came from a manual run of an ungated binary — including the assertion added for #1, unreachable from `make`. Now `make panels`, registered in `gates.toml` |
| 8 | `edition-check`'s sibling handling was self-certifying: it read the file list *out of PROVENANCE*, crashed instead of failing on an absent file, and "coverage" counted Markdown lines | **conceded — a defect introduced by the correction.** Siblings are now discovered on disk; an absent recorded file FAILs; undocumented siblings FAIL. **It found a real gap on its first run**: `Rules_Text.csv` was vendored with no digest |
## The minor ones
- **9 — the CSV falsifier ran over YAML.** Fixed; it now filters to `.csv`.
- **10 — "peak Stress held" was dead code.** `held` was built, never read,
and `let _ = held;` kept clippy quiet; the metric was still max over
payloads plus a floor. **The published numbers were right by
coincidence.** Now actually read from `held`, and the figures are
unchanged — which is how it stayed hidden.
- **11 — `Reactive`'s doc comment still argued for the copy** that round 1
ruled fatal, giving cover to restore it. Rewritten.
- **12 — §3's withdrawal ground was wrong.** *"There is no constant to
explain"* — there is: `darvo == seats × games`, exactly. The constant
moved from 400 to `n·games` rather than dissolving, and **what forces
exactly one arm per seat per game is still unexplained.**
- **13 — assorted**, including that the delta's DARVO clause is
conditional (*"if they share the same path"*) where the test's comment
says "outright".
## What the reviewer could not check
Digest freshness against upstream (`edition-check` reports
`upstream not checked out` when the sibling repo is absent — it did not
have it); several untouched modules; ground-game's design note; felt-play.
---
## The pattern, now visible across two rounds
Round 1 found five fatal errors of the form *correct computation, wrong
subject*. **Round 2 found three more, in the corrections to them** — an
assertion invariant under its own target, a selective fix for a selective
report, and a rate wrong by a factor of two.
> **A correction is new work and deserves the scrutiny of new work.** It
> arrives feeling like a conclusion, and that is exactly when it is least
> examined.
**Round 3 is owed.** Same reason as round 2.