clay-borg/reviews/CB-REV-0001-h1.md
tegwick da58e78e4a
Some checks failed
ci / check (push) Failing after 4s
CB-REV-0002: round 2, and the corrections were not approvable either
Three FATAL, five SERIOUS. The substance of round 1's corrections held —
Reactive is genuinely one arm different, the five replacement controls are
non-inert, the inert metric is right, the numbers reproduce. What failed
were the CLAIMS about them, and two defects the corrections introduced.

FATAL 1: the fix for round 1's #11 did not fix it. The assertion was
`games + setup_fails == 200`, and a refused setup increments setup_fails
while skipping games — so the sum is invariant under exactly the failure
it claimed to catch. Injecting setup failures gave exit 0 over 196-game
columns. Now asserts games == GAMES, verified to exit 101.

FATAL 2: the correction to the selective-column FATAL was itself
selective. "81-1000 per cell, baseline AND H1" and "31-1000" twelve lines
apart, both taken from the baseline row; under H1 rank-75 arms are
59/0/0/0. Every cell is now printed rather than summarised, and the
corrected verdict is the opposite of the one it replaced: under rank-75,
H1 REDUCES DARVO arms to zero at 3p and above.

FATAL 3: "DARVO arms 2 per seat per game" is 1 per seat per game, exactly,
at every band.

SERIOUS: the tiebreak oracle asserted only that the winner set CHANGED, so
reversing the tiebreak left it green; the #13 defect's impact was claimed
and never measured (72,000 games: zero divergences — real in principle,
witnessed only by a constructed board); a 29-of-363 citation pointed at a
file that did not contain it (round 1's reviewer did report it, and it was
never transcribed — the record was wrong, not the number); the harnesses
were run by NO GATE, so every published figure came from a manual run of
an ungated binary, including the assertion added for #1; and edition-check's
sibling handling — added by the last correction — was self-certifying,
crashed instead of failing, and counted Markdown lines as coverage. Now
discovered on disk, and it found a real gap on its first run: Rules_Text.csv
vendored with no digest.

Also: "peak Stress held" was dead code kept quiet by `let _ = held;` — the
numbers were right by coincidence.

make panels is now a registered gate. Round 3 is owed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 09:44:00 +02:00

7.8 KiB
Raw Blame History

CB-REV-0001 — adversarial review of the H1 measurement

The tier-L review CB-WP-0038 owed (InnerLoop Step 2). One round: challenge, then response. Run 2026-08-08 by a separate agent against CB-WP-0038/CB-EV-0030 and CB-WP-0039/CB-EV-0031.

Verdict: not approvable as submitted. All thirteen are now closed — see the response under each. Re-review is owed before any of this travels: the corrections were made by the author of the errors.

Original verdict: Thirteen challenges, five rated FATAL. Every FATAL is conceded. Nothing from either evidence file had reached ground-game, which is the only reason this is a correction rather than a retraction.

The reviewer reproduced every number, re-derived on three samples (seeds 0..200, 1000..1200, 5000..5500) and ran 14 kernel mutations.


The five that were fatal

1. Reactive was not "greedy with one preference changed" — it differed in five

Conceded, and it is the worst thing in the pass. The policy re-typed an abridged copy of greedy's ranking and changed one line of the copy, under a comment reading // THE ONE LINE. It also differed in SpendFreedom (95 unconditional against greedy's 95 if gated else 0), Solve on a claimed Problem, ChooseGroundMode and RespondToSupport.

The SpendFreedom difference is a second change to the exact mechanism under study: the seat burned its Freedom token in round one of every game, ungated. Measured by the reviewer: 1200 spends against greedy's 0.

This pass claimed ADR-0018's one-varying-parameter discipline in its own workplan while violating it. That is worse than not claiming it.

Fixed structurally rather than by testing: GreedyPolicy::rank is now pub, and Reactive delegates to it and overrides a single match arm. A caller that delegates cannot drift. The reviewer's StrictReactive numbers reproduce exactly.

2. The H1-B suppression mechanism does not survive the corrected policy

Conceded and withdrawn. The reviewer disabled H1-B and re-ran: darvo is identical in every cell. The effect CB-EV-0031 §3 attributed to "H1-B's arithmetic" was an interaction with challenge 1's bugs.

The hedge was on the wrong variable. §3 disclaimed "the number 2" and defended "the direction". The direction is what failed.

3. [5,5,4,4,4,4] does not show that two seats armed

Conceded. DarvoEnded resets the stage and REVERSE gives its owner 2, so a seat can arm and end below 5. The reviewer exhibited six seats arming behind the same signature. A terminal snapshot was used to prove a path property.

4. Criterion 1 was failed on the greedy column alone

Conceded. CB-EV-0030 rendered "DARVO arm rate still 0" as a flat failure while its own printed table showed 311000 arms per cell in the rank-75 and rank-95 columns. The selective-column move, in the file that names selective reporting as the thing to avoid. Restated per column.

5. "Peak Stress was 1" was a wrong-subject error, with two more defects

Conceded, all three. peak maximised over StressSet event payloads; starting Stress is written by setup and never by an event, so a table sitting at 2 reported 1. True peak held: 2. The baseline game count was 1,600, not 3,200. And the generalisation — "a reckless policy plays identically to a careful one" — is refuted by this repo's own rank-95 policy, which under baseline drives Stress to 5.

What survives is narrower and is now stated that way: under the two policies measured, neither of which selects ATTACK under baseline, no Stress is ever added. ATTACK is the sole inbound pressure.


The serious ones

# challenge outcome
6 baseline_is_bit_for_bit_what_it_was compared two identically-constructed states — inert against its own threat; #[serde(skip)] on variant left 57/57 green conceded. Replaced by the_variant_reaches_the_hash_and_baseline_is_not_a_change, which asserts baseline and H1 hash differently. The reviewer's M6 now goes red
7 the unchanged: test checked 3 of 7 entries; making SOLVE illegal under H1 left it green conceded. Now compares legal_commands per seat (catches M9), checks relation slots at the boundary rather than on an empty seat, and asserts the DARVO arm is at 5 (catches M7). Both verified red
8 "criteria met" rests on a forced move — reactive ranks ATTACK at 10 and picks it only when the gate leaves nothing else conceded as a limitation, recorded, not fixed. It is true that the instrument cannot show a null result once H1-A reaches Stress 4. That is a real weakness of the measurement and is now stated in CB-EV-0031 §4
9 three claimed properties had no failing test: H1-A's ordering, H1-B on the DARVO extra Attack, H1-B on an OU-cancelled Attack all three fixed and mutation-verified: h1a_pressure_arms_darvo_in_the_same_round_end, h1b_does_not_soothe_an_ou_cancelled_attack, h1b_soothes_the_darvo_stage_attack_too

The rest

  • 10 — H1-B's after_target_and_relation_effects ordering is vacuous here. Conceded: nothing between the two positions touches the attacker's Stress, so the clause cannot be checked in this kernel. The honest statement replaces the claim that it was got right.

  • 11 — regulation.rs still skips setup failures silently. Conceded and fixed: counted, and a short cell now fails an assertion rather than printing a number a reader must notice.

  • 13 — round-5 arms, and scoring from self. (Round 1's reviewer reported the arm figure as 29 of 363 at 2p, 8%, under its StrictReactive. That number was omitted when this file was written and later cited to it — see CB-REV-0002 #6. Recorded here now.) Both fixed, and the second was a real defect rather than a reporting one: score reads Stress for the GR-E03/GR-E04 tiebreaks, so the final round's pressure was invisible to the two modes CB-EV-0030 reports on. Inert arms are now reported separately — 29 of 363 at 2p, none above — which matches the reviewer's independent figure and leaves criterion 1 met.

    Fixing it produced one more wrong-subject error, caught before it was reported: the first inert-arm metric tested g.rounds >= 5, a property of the game rather than the event, so it marked every arm in every completed game inert and briefly read as "criterion 1 fails after all".

  • 12 — reproduction and sample robustness: no problem found. Every number reproduced; no conclusion was seed-specific. The failures were of construction and interpretation, not sampling.

What the reviewer could not check, and it is recorded rather than glossed

The design note itself (criteria quoted from CB-EV-0030, not read from source); the catalog digest, which the reviewer could not find and reported as unverified rather than absent; --variant on the CLI; make cost; and the two non-SHARED modes under reactive.

The catalog digest is a real gap. CB-WP-0038 T01's control said "the catalog is vendored with a digest, like every other borrowed file", and PROVENANCE.md records digests for the ten CSVs and not for catalog.yaml or the H1 package. The control was claimed and not met.


What this cost, and what it bought

The review found five fatal defects in numbers that were one step from another repository's design decision, and four of the five were errors of the exact class this project has been cataloguing since ADR-0018 — correct computation, wrong subject.

The instrument caught the instrument's author. CB-WP-0039 was itself written to correct CB-EV-0030, and it introduced worse errors than the ones it fixed. The lesson is not "review works" — it is that a pass written to correct a previous pass inherits none of its caution.