Some checks failed
ci / check (push) Failing after 4s
Three FATAL, five SERIOUS. The substance of round 1's corrections held — Reactive is genuinely one arm different, the five replacement controls are non-inert, the inert metric is right, the numbers reproduce. What failed were the CLAIMS about them, and two defects the corrections introduced. FATAL 1: the fix for round 1's #11 did not fix it. The assertion was `games + setup_fails == 200`, and a refused setup increments setup_fails while skipping games — so the sum is invariant under exactly the failure it claimed to catch. Injecting setup failures gave exit 0 over 196-game columns. Now asserts games == GAMES, verified to exit 101. FATAL 2: the correction to the selective-column FATAL was itself selective. "81-1000 per cell, baseline AND H1" and "31-1000" twelve lines apart, both taken from the baseline row; under H1 rank-75 arms are 59/0/0/0. Every cell is now printed rather than summarised, and the corrected verdict is the opposite of the one it replaced: under rank-75, H1 REDUCES DARVO arms to zero at 3p and above. FATAL 3: "DARVO arms 2 per seat per game" is 1 per seat per game, exactly, at every band. SERIOUS: the tiebreak oracle asserted only that the winner set CHANGED, so reversing the tiebreak left it green; the #13 defect's impact was claimed and never measured (72,000 games: zero divergences — real in principle, witnessed only by a constructed board); a 29-of-363 citation pointed at a file that did not contain it (round 1's reviewer did report it, and it was never transcribed — the record was wrong, not the number); the harnesses were run by NO GATE, so every published figure came from a manual run of an ungated binary, including the assertion added for #1; and edition-check's sibling handling — added by the last correction — was self-certifying, crashed instead of failing, and counted Markdown lines as coverage. Now discovered on disk, and it found a real gap on its first run: Rules_Text.csv vendored with no digest. Also: "peak Stress held" was dead code kept quiet by `let _ = held;` — the numbers were right by coincidence. make panels is now a registered gate. Round 3 is owed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
143 lines
7.8 KiB
Markdown
143 lines
7.8 KiB
Markdown
# CB-REV-0001 — adversarial review of the H1 measurement
|
||
|
||
The tier-L review CB-WP-0038 owed (InnerLoop Step 2). One round: challenge,
|
||
then response. Run 2026-08-08 by a separate agent against
|
||
CB-WP-0038/CB-EV-0030 and CB-WP-0039/CB-EV-0031.
|
||
|
||
> **Verdict: not approvable as submitted. All thirteen are now closed** —
|
||
> see the response under each. **Re-review is owed** before any of this
|
||
> travels: the corrections were made by the author of the errors.
|
||
>
|
||
> **Original verdict:** Thirteen challenges, five
|
||
> rated FATAL. **Every FATAL is conceded.** Nothing from either evidence
|
||
> file had reached `ground-game`, which is the only reason this is a
|
||
> correction rather than a retraction.
|
||
|
||
The reviewer reproduced every number, re-derived on **three** samples
|
||
(seeds 0..200, 1000..1200, 5000..5500) and ran **14 kernel mutations**.
|
||
|
||
---
|
||
|
||
## The five that were fatal
|
||
|
||
### 1. `Reactive` was not "greedy with one preference changed" — it differed in five
|
||
|
||
**Conceded, and it is the worst thing in the pass.** The policy re-typed
|
||
an abridged copy of greedy's ranking and changed one line *of the copy*,
|
||
under a comment reading `// THE ONE LINE`. It also differed in
|
||
`SpendFreedom` (95 unconditional against greedy's `95 if gated else 0`),
|
||
`Solve` on a claimed Problem, `ChooseGroundMode` and `RespondToSupport`.
|
||
|
||
**The `SpendFreedom` difference is a second change to the exact mechanism
|
||
under study**: the seat burned its Freedom token in round one of every
|
||
game, ungated. Measured by the reviewer: 1200 spends against greedy's 0.
|
||
|
||
**This pass claimed ADR-0018's one-varying-parameter discipline in its own
|
||
workplan while violating it.** That is worse than not claiming it.
|
||
|
||
**Fixed structurally rather than by testing:** `GreedyPolicy::rank` is now
|
||
`pub`, and `Reactive` delegates to it and overrides a single match arm. A
|
||
caller that delegates cannot drift. The reviewer's `StrictReactive`
|
||
numbers reproduce exactly.
|
||
|
||
### 2. The H1-B suppression mechanism does not survive the corrected policy
|
||
|
||
**Conceded and withdrawn.** The reviewer disabled H1-B and re-ran:
|
||
`darvo` is **identical in every cell**. The effect CB-EV-0031 §3 attributed
|
||
to "H1-B's arithmetic" was an interaction with challenge 1's bugs.
|
||
|
||
**The hedge was on the wrong variable.** §3 disclaimed *"the number 2"*
|
||
and defended *"the direction"*. The direction is what failed.
|
||
|
||
### 3. `[5,5,4,4,4,4]` does not show that two seats armed
|
||
|
||
**Conceded.** `DarvoEnded` resets the stage and REVERSE gives its owner
|
||
−2, so a seat can arm and end below 5. The reviewer exhibited **six**
|
||
seats arming behind the same signature. A terminal snapshot was used to
|
||
prove a path property.
|
||
|
||
### 4. Criterion 1 was failed on the greedy column alone
|
||
|
||
**Conceded.** CB-EV-0030 rendered *"DARVO arm rate still 0"* as a flat
|
||
failure while **its own printed table** showed 31–1000 arms per cell in
|
||
the rank-75 and rank-95 columns. **The selective-column move, in the file
|
||
that names selective reporting as the thing to avoid.** Restated per
|
||
column.
|
||
|
||
### 5. "Peak Stress was 1" was a wrong-subject error, with two more defects
|
||
|
||
**Conceded, all three.** `peak` maximised over `StressSet` event
|
||
*payloads*; starting Stress is written by `setup` and never by an event,
|
||
so a table sitting at 2 reported 1. True peak held: **2**. The baseline
|
||
game count was **1,600**, not 3,200. And the generalisation — *"a reckless
|
||
policy plays identically to a careful one"* — is refuted by this repo's
|
||
own rank-95 policy, which under **baseline** drives Stress to 5.
|
||
|
||
**What survives is narrower and is now stated that way**: under the two
|
||
policies measured, neither of which selects ATTACK under baseline, no
|
||
Stress is ever added. ATTACK is the sole inbound pressure.
|
||
|
||
---
|
||
|
||
## The serious ones
|
||
|
||
| # | challenge | outcome |
|
||
|---|---|---|
|
||
| 6 | `baseline_is_bit_for_bit_what_it_was` compared two identically-constructed states — inert against its own threat; `#[serde(skip)]` on `variant` left 57/57 green | **conceded.** Replaced by `the_variant_reaches_the_hash_and_baseline_is_not_a_change`, which asserts baseline and H1 hash **differently**. The reviewer's M6 now goes red |
|
||
| 7 | the `unchanged:` test checked **3 of 7** entries; making SOLVE illegal under H1 left it green | **conceded.** Now compares `legal_commands` per seat (catches M9), checks relation slots at the **boundary** rather than on an empty seat, and asserts the DARVO arm is at 5 (catches M7). Both verified red |
|
||
| 8 | "criteria met" rests on a forced move — `reactive` ranks ATTACK at 10 and picks it only when the gate leaves nothing else | **conceded as a limitation, recorded, not fixed.** It is true that the instrument cannot show a null result once H1-A reaches Stress 4. That is a real weakness of the measurement and is now stated in CB-EV-0031 §4 |
|
||
| 9 | three claimed properties had no failing test: H1-A's ordering, H1-B on the DARVO extra Attack, H1-B on an OU-cancelled Attack | **all three fixed and mutation-verified**: `h1a_pressure_arms_darvo_in_the_same_round_end`, `h1b_does_not_soothe_an_ou_cancelled_attack`, `h1b_soothes_the_darvo_stage_attack_too` |
|
||
|
||
## The rest
|
||
|
||
- **10 — H1-B's `after_target_and_relation_effects` ordering is vacuous
|
||
here.** Conceded: nothing between the two positions touches the
|
||
attacker's Stress, so the clause cannot be checked in this kernel. The
|
||
honest statement replaces the claim that it was got right.
|
||
- **11 — `regulation.rs` still skips setup failures silently.** Conceded
|
||
and **fixed**: counted, and a short cell now fails an assertion rather
|
||
than printing a number a reader must notice.
|
||
- **13 — round-5 arms, and scoring from `self`.** *(Round 1's reviewer
|
||
reported the arm figure as **29 of 363 at 2p, 8%**, under its
|
||
`StrictReactive`. That number was omitted when this file was written and
|
||
later cited to it — see CB-REV-0002 #6. Recorded here now.)* Both **fixed**, and the
|
||
second was a real defect rather than a reporting one: `score` reads
|
||
Stress for the GR-E03/GR-E04 tiebreaks, so the final round's pressure
|
||
was invisible to the two modes CB-EV-0030 reports on. Inert arms are now
|
||
reported separately — **29 of 363 at 2p, none above** — which matches the
|
||
reviewer's independent figure and leaves criterion 1 met.
|
||
|
||
**Fixing it produced one more wrong-subject error**, caught before it
|
||
was reported: the first inert-arm metric tested `g.rounds >= 5`, a
|
||
property of the *game* rather than the *event*, so it marked every arm
|
||
in every completed game inert and briefly read as "criterion 1 fails
|
||
after all".
|
||
- **12 — reproduction and sample robustness: no problem found.** Every
|
||
number reproduced; no conclusion was seed-specific. **The failures were
|
||
of construction and interpretation, not sampling.**
|
||
|
||
## What the reviewer could not check, and it is recorded rather than glossed
|
||
|
||
The design note itself (criteria quoted from CB-EV-0030, not read from
|
||
source); the catalog digest, which the reviewer could not find and
|
||
reported as **unverified rather than absent**; `--variant` on the CLI;
|
||
`make cost`; and the two non-SHARED modes under `reactive`.
|
||
|
||
**The catalog digest is a real gap.** CB-WP-0038 T01's control said *"the
|
||
catalog is vendored with a digest, like every other borrowed file"*, and
|
||
`PROVENANCE.md` records digests for the ten CSVs and **not** for
|
||
`catalog.yaml` or the H1 package. The control was claimed and not met.
|
||
|
||
---
|
||
|
||
## What this cost, and what it bought
|
||
|
||
The review found **five fatal defects in numbers that were one step from
|
||
another repository's design decision**, and four of the five were errors
|
||
of the exact class this project has been cataloguing since ADR-0018 —
|
||
correct computation, wrong subject.
|
||
|
||
**The instrument caught the instrument's author.** CB-WP-0039 was itself
|
||
written to correct CB-EV-0030, and it introduced worse errors than the
|
||
ones it fixed. **The lesson is not "review works" — it is that a pass
|
||
written to correct a previous pass inherits none of its caution.**
|