CB-WP-0042 T05: H2 measured — it largely succeeds where H1 failed
Some checks failed
ci / check (push) Failing after 4s

Measured against ground-game's own §3 criteria, read from their design
note rather than reused from H1.

Criterion 1 met with room: greedy SHARED wins 120 at 3p and 175 at 4p,
against H1's 0 and 0, restoring 73% and 92% of baseline.

Criterion 3 met, and it was H1's clearest failure. Under H1 the
unregulated seat armed DARVO constantly and never won; under H2 it arms
and wins 13/13/57. "Non-zero for some policy that still sometimes wins" is
exactly the shape H1 could not produce.

Criterion 2 met at 2p/4p/6p and missed at 3p — 1.50 against baseline's
1.57 — reported as a miss because that is what this sample says. The
mechanism is visible: H1-greedy's spread is 0.00 at 3p+, because a flat
tax on every seat creates no variance at all. That is the clearest
statement of why scoping was the right correction.

Criterion 5 is the best evidence in the pass. Forcing every scope to
global and changing nothing else reproduces H1's collapse exactly — 120 to
0 at 3p, 175 to 0 at 4p — so the scoping is what saves it, not any other
difference between the packages.

Criterion 4 came out backwards and the prediction held. The workplan said
this panel might be unable to test it, because no policy here models
another seat or knows what a scope is; bond claim rates are LOWER than
personal at 3p and 4p, driven by suit availability rather than incentive.
Reported as untested with an incidental figure pointing the wrong way,
not as a refutation.

Wired into make panels. First pass declared after ADR-0021, so no chaos
roll is recorded.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
tegwick 2026-08-08 16:09:16 +02:00
parent 04e26b3077
commit ff89288581
4 changed files with 352 additions and 2 deletions

View file

@ -2,7 +2,7 @@
id: CB-WP-0042
kind: product
title: "H2 — scoped problem stress"
status: active
status: done
state_hub_workstream_id: "8f11d55b-3452-49f3-9857-d0bf06752f68"
---
@ -176,7 +176,7 @@ The test must fail if it does.
```task
id: CB-WP-0042-T05
status: todo
status: done
priority: high
state_hub_task_id: "0bcd97da-60b8-46b2-ae86-7a94f6af8dd5"
```
@ -196,6 +196,34 @@ Same instrument as H1, same panel, both variants in one run.
- **`make panels` runs it**, or the figures come from an ungated binary
again (CB-REV-0002 #7).
**Done 2026-08-08.**
[CB-EV-0032](../evidence/CB-EV-0032-h2-measured.md). **H2 largely
succeeds where H1 failed.**
| # | their criterion | verdict |
|---|---|---|
| 1 | greedy SHARED 34p above H1's 0 | **met** — 120 and 175 against H1's 0 and 0 |
| 2 | variance above baseline and H1-greedy | **met at 2p/4p/6p, missed at 3p** (1.50 vs 1.57) |
| 3 | DARVO for a policy that still sometimes wins | **met** — H1's clearest failure, now 13/13/57 wins with arms |
| 4 | bond SOLVE rate above personal | **not met, and untestable here** |
| 5 | control: all scopes global → collapse | **met decisively** — 120→0 and 175→0 |
**Criterion 5 is the best evidence in the pass.** Forcing every scope to
`global` and changing nothing else reproduces H1's collapse exactly, which
says the **scoping** is what saves it rather than any other difference
between the packages.
**Criterion 2's mechanism is visible in the numbers**: H1-greedy's spread
is **0.00** at 3p+. A flat tax on every seat creates no variance at all,
which is the clearest statement of why scoping was the right correction.
**Criterion 4 came out backwards, and the prediction held.** The workplan
said our panel might not be able to test it, because no policy here models
another seat or knows what a scope is. Bond claim rates are **lower** than
personal at 3p and 4p — driven by suit availability, not by incentive.
**Reported as untested with an incidental figure pointing the wrong way**,
not as a refutation.
## Not in this workplan
- **No stacking with H1.** The package forbids it explicitly.