ground-game packages hypotheses as selectable rules variants — a catalog, a rules_delta.yaml, and prose — and their note is explicit that CSV text alone is not executable here. So the kernel gains a Variant in game state: in the state, therefore in the hash, therefore in the recording, because a scenario replayed under a different variant would diverge silently. Baseline is bit-for-bit what it was, asserted across seat counts and seeds. A variant system that perturbs the baseline invalidates every measurement this repo has. H1-A and H1-B implemented from rules_delta.yaml and mutation-proven on their own defects: "unclaimed" misread as face-up-and-unsolved, and the attacker's Stress read after the attack's effects. Their `unchanged:` list is asserted rather than trusted — that list is their claim about their own experiment. Measured, and three of their four criteria fail. DARVO arm rate is still 0 under greedy; ATTACK selection does not rise and falls for the rank-75 policy; group success collapses from 165/190/200 to 0 at 3/4/6 seats. The mechanism is not the assumed one: greedy answers the pressure by regulating, Stress plateaus at 3, so it never reaches the gate at 4 or the arm at 5 — H1-A acts as a solve-rate tax and H1-B is unreachable under competent play. A harness defect was caught before the claim: sweep discarded refused games silently and never reported its count, so "nobody won" and "nothing played" printed identically. Reporting H1 as unwinnable on that basis would have been the ADR-0018 family aimed at another repo's design. All 200 games ran in every cell; the zeros are real. Chaos d8 = 8 — the window's first override, redrew L against a structural L, so it changed nothing. Window 3 recorded in ChaosRollHistory. NOT REVIEWED: tier L owes a separate-agent adversarial review, and no H1 result may reach ground-game until it has run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
3.6 KiB
Chaos roll — window records
One entry per chaos window. Split out of InnerLoopReference.md when that
file crossed the loadability limit: this is a log that grows, and a log
inside a reference eventually crowds out the reference.
The rule itself lives in InnerLoop.md §Loop tiers; the
current window's terms are there. This is the history the verdicts rest on.
Window 1 (d4, closed 2026-08-02): 12 declarations, 2 overrides, one each way, and both changed the outcome. The mechanism was kept and the rate dropped d4 → d8 (CB-EV-0015 §5, CB-EV-0016 §4).
Window 2 (d8, 2026-08-03 → 2026-08-07): 12 declarations, of which 11 rolled at d8 — declaration 1 (CB-WP-0018) opened the window and rolled at the old d4. Expected 8s: 1.375. Observed: 1.
| decl | pass | roll | effect |
|---|---|---|---|
| 1 | CB-WP-0018 | d4 = 3 | opened the window at the old rate |
| 2 | CB-WP-0019 | d8 = 5 | — |
| 3 | CB-WP-0020 | d8 = 8 | override — drew S over a structural S: changed nothing |
| 4 | CB-WP-0021 | d8 = 7 | — |
| 5–9 | CB-WP-0022…0026 | d8 = 6 ×5 | — |
| 10 | CB-WP-0027 | d8 = 7 | — |
| 11 | CB-WP-0028 | d8 = 1 | — |
| 12 | CB-WP-0029 | d8 = 4 | — |
Verdict (ADR-0017): the rate is behaving as designed. One override against 1.375 expected is not a shortage of evidence.
Why the retirement condition changed. "An override changes nothing twice running" requires a consecutive pair, each with P = 1/3, so ~12 overrides are expected before one occurs — at ~1.4 overrides per window, ~9 windows or roughly 100 declarations. A gate that cannot cash out on any realistic horizon is decoration (ADR-0006 D3). It is now evaluated per window, needing two consecutive qualifying windows: ~24 declarations rather than ~100.
Four evidence files claimed window 2 produced zero overrides. They were wrong, and each cited the one before rather than counting. See F23 — the failure is a claim with no source, asserted once and repeated, which no gate here detects.
A post-hoc observation, deliberately not acted on. Declarations 5–9
rolled six five times running (~1 in 370 for some run of five in eleven
rolls). shuf was tested over 200 rapid successive calls and looks
uniform, longest run three. Recorded so a future window can check whether
it recurs; not evidence of anything on its own.
Window 3 — opened 2026-08-07 at d8, running to 12 declarations
| # | pass | roll | override |
|---|---|---|---|
| 1 | CB-WP-0030 | d8 = 7 | — |
| 2 | CB-WP-0031 | d8 = 2 | — |
| 3 | CB-WP-0032 | d8 = 6 | — |
| 4 | CB-WP-0033 | d8 = 7 | — |
| 5 | CB-WP-0034 | d8 = 4 | — |
| 6 | CB-WP-0035 | d8 = 4 | — |
| 7 | CB-WP-0036 (first declaration) | d8 = 7 | — |
| 8 | CB-WP-0036 (re-declared L→M) | d8 = 7 | — |
| 9 | CB-WP-0037 | d8 = 1 | — |
| 10 | CB-WP-0038 | d8 = 8 | yes — redraw L, structural was L, so it changed nothing |
The window's first 8, at declaration 10. Expectation over ten rolls at d8 is 1.25; one is exactly on rate.
The override changed nothing, which is the observation ADR-0017 D2's retirement condition is built from — it needs a full window whose overrides all change nothing, in two consecutive windows. This window now has one qualifying override and two declarations left to run.
Declaration 8 is a re-declaration of the same workplan, counted separately because it was a materially different pass: CB-WP-0036 was re-scoped from L to M after the maintainer moved the animation work out of the repo, and a changed declaration is a new declaration or the roll is not binding on what was actually built.