clay-borg/evidence/CB-EV-0030-h1-measured.md
tegwick da58e78e4a
Some checks failed
ci / check (push) Failing after 4s
CB-REV-0002: round 2, and the corrections were not approvable either
Three FATAL, five SERIOUS. The substance of round 1's corrections held —
Reactive is genuinely one arm different, the five replacement controls are
non-inert, the inert metric is right, the numbers reproduce. What failed
were the CLAIMS about them, and two defects the corrections introduced.

FATAL 1: the fix for round 1's #11 did not fix it. The assertion was
`games + setup_fails == 200`, and a refused setup increments setup_fails
while skipping games — so the sum is invariant under exactly the failure
it claimed to catch. Injecting setup failures gave exit 0 over 196-game
columns. Now asserts games == GAMES, verified to exit 101.

FATAL 2: the correction to the selective-column FATAL was itself
selective. "81-1000 per cell, baseline AND H1" and "31-1000" twelve lines
apart, both taken from the baseline row; under H1 rank-75 arms are
59/0/0/0. Every cell is now printed rather than summarised, and the
corrected verdict is the opposite of the one it replaced: under rank-75,
H1 REDUCES DARVO arms to zero at 3p and above.

FATAL 3: "DARVO arms 2 per seat per game" is 1 per seat per game, exactly,
at every band.

SERIOUS: the tiebreak oracle asserted only that the winner set CHANGED, so
reversing the tiebreak left it green; the #13 defect's impact was claimed
and never measured (72,000 games: zero divergences — real in principle,
witnessed only by a constructed board); a 29-of-363 citation pointed at a
file that did not contain it (round 1's reviewer did report it, and it was
never transcribed — the record was wrong, not the number); the harnesses
were run by NO GATE, so every published figure came from a manual run of
an ungated binary, including the assertion added for #1; and edition-check's
sibling handling — added by the last correction — was self-certifying,
crashed instead of failing, and counted Markdown lines as coverage. Now
discovered on disk, and it found a real gap on its first run: Rules_Text.csv
vendored with no digest.

Also: "peak Stress held" was dead code kept quiet by `let _ = held;` — the
numbers were right by coincidence.

make panels is now a registered gate. Round 3 is owed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 09:44:00 +02:00

6.6 KiB
Raw Blame History

CB-EV-0030 — H1 measured

CB-WP-0038. Tier L (structural L; chaos d8 = 8 → override, redraw = L — the override changed nothing). Declaration 10 of chaos window 3, and the window's first 8. Closed 2026-08-08.

This evidence has not had its tier-L adversarial review (InnerLoop Step 2). It may not go to ground-game until it has. It is a measurement of another repo's design hypothesis, which is precisely the case the review exists for.

Delivered: variant selection in the kernel, H1-A and H1-B implemented from rules_delta.yaml, ground-game's unchanged: list asserted rather than trusted, and the ATTACK panel re-run for both rule sets.


1. The result, against their own criteria

ground-game's success criteria (design note §3.2), answered one by one. Three of four fail.

CORRECTED TWICE. Round 1 found criteria 1 and 2 rendered as flat failures on the greedy column alone, while this pass's own harness printed non-zero DARVO arms in the rank-75 and rank-95 columns — the selective-column move, in the file that names it.

Round 2 found the correction was itself selective, and its numbers wrong (CB-REV-0002 #2): it quoted "811000 per cell … baseline and H1" and, twelve lines away, "311000". Both ranges came from the baseline rank-75 row. Under H1, rank-75 arms are 2p 59, 3p 0, 4p 0, 6p 0 — three cells at zero and the fourth below the stated floor.

Every cell is now given, because two attempts at summarising this into a range have now produced three wrong numbers.

DARVO arms per 200-game cell, SHARED GROUND (all three modes agree on arms; they differ only in who wins):

greedy rank-75 rank-95
baseline 2p / 3p / 4p / 6p 0 / 0 / 0 / 0 81 / 125 / 218 / 31 400 / 805 / 1000 / 1000
H1 2p / 3p / 4p / 6p 0 / 0 / 0 / 0 59 / 0 / 0 / 0 600 / 870 / 941 / 970
# criterion verdict
1 DARVO arm rate non-trivial, "not 0/500" depends entirely on the policy, and H1 does not improve it for the middle one. Zero under greedy in both. Under rank-75, H1 reduces arms to zero at 3p and above. Only rank-95 — a seat that always attacks — arms freely, and it does so under the baseline too
2 ATTACK selection rises for some subpopulation no, for the policy the criterion is about. Greedy attacks 0 times under both. rank-75's attacks fall 604 → 180 at 2p and 592/923/389 → 0 at 3p/4p/6p. Only the deliberately unregulated seat of CB-EV-0031 rises, and it never wins
3 group success does not collapse fails hardest — greedy's wins go 165 → 0 (3p), 190 → 0 (4p), 200 → 0 (6p)
4 Bond/GROUND stay better than DARVO holds, but vacuously: DARVO never fires

SHARED GROUND, greedy column, 200 games per cell:

seats baseline won H1 won
2p 132 68
3p 165 0
4p 190 0
6p 200 0

2. Why — and the mechanism is not the one H1 assumed

H1-A was meant to raise Stress toward the DARVO arm at 5. It does not get there under competent play. One 3-player game, seed 1:

rounds final Stress Problems unclaimed DARVO armed outcome
baseline 5 0, 1, 2 0 0 9 of 7 — success
H1 5 3, 3, 3 2 0 4 of 7 — failure

Greedy spends its actions on Stress instead of Problems. The pressure is answered by regulating, so Stress plateaus around 3 — below the gate at 4 and well below the arm at 5 — while two Problems go unclaimed and the total lands under the threshold.

So H1-A behaves as a solve-rate tax, not as DARVO pressure. And since Stress never reaches 4, H1-B never fires at all under competent play: the self-soothe is unreachable by the policy it was designed to tempt.

The seat-count pattern is the tell. At 2p — three Problems, so a table can clear the board — H1 helps in SHARED (132 → 68 is a fall, but 2p is the only band that still wins at all). At 3p and above the board is bigger than the pressure allows a table to clear, and success goes to zero. The tax scales with the number of Problems; the intended effect does not.

3. The harness defect caught before the claim

The first run showed won, atk and darvo all exactly 0 at 3+ seats. That is the signature of games being discarded, not lost.

sweep had let Ok(g) = play(st, &mut ps) else { continue }; and never reported its game count, so a run in which nothing played and a run in which nothing won were the same output.

Reporting "H1 makes the game unwinnable" when the truth might have been "our harness threw every H1 game away" would have been the ADR-0018 family — aimed, this time, at another repository's design decision.

Instrumented: setup failures counted, play errors captured and printed, games != 200 reported. All 200 ran in every cell. The zeros are real, and now they are known to be real.

4. What this evidence does not say

  • It does not say H1 is a bad idea. It says these two deltas, at these magnitudes, measured by this panel, do not do what §3.2 asks. Their §6 already has the branch for it: "H1 measured | DARVO still ~0 | add Idea 2 last-place stress (semi only) as H2".
  • The panel is greedy-family throughout — greedy, and greedy with ATTACK reranked to 75 and 95. Plural, per RetrospectiveAnalysis §4.1, but not diverse. A policy that manages Stress badly would reach 5 and arm DARVO, and none of these do. That is the obvious next measurement and it has not been run.
  • No felt-play. GROUND-WP-0003 T01T05 runs human tables separately, and 200-game aggregates cannot say what a table feels like.

5. A tuning direction, offered as a question and not a finding

If the aim is Stress reaching 5 sometimes, the tax must outrun regulation. Candidates, none measured: pressure scaling with the number of unclaimed Problems rather than a flat +1; applying it only above a Round threshold; or lowering the DARVO arm below 5. Each is a rules change and therefore ground-game's to make, not ours to propose as a fix.

6. Cost

make cost, not inlined. One vendored catalog and one experiment package; no new dependency.

Open after this pass

  • The tier-L adversarial review, before any of this travels.
  • A policy that regulates badly, which is the measurement §4 names.
  • H1-B has never been observed firing in a real game — only in a unit test. A delta that competent play cannot reach is untested where it matters.