clay-borg/specs/ChaosRollHistory.md
tegwick a928b5925c
Some checks failed
ci / check (push) Failing after 3s
CB-WP-0038: variant selection, H1 implemented, and H1 measured
ground-game packages hypotheses as selectable rules variants — a catalog,
a rules_delta.yaml, and prose — and their note is explicit that CSV text
alone is not executable here. So the kernel gains a Variant in game state:
in the state, therefore in the hash, therefore in the recording, because a
scenario replayed under a different variant would diverge silently.

Baseline is bit-for-bit what it was, asserted across seat counts and
seeds. A variant system that perturbs the baseline invalidates every
measurement this repo has.

H1-A and H1-B implemented from rules_delta.yaml and mutation-proven on
their own defects: "unclaimed" misread as face-up-and-unsolved, and the
attacker's Stress read after the attack's effects. Their `unchanged:` list
is asserted rather than trusted — that list is their claim about their own
experiment.

Measured, and three of their four criteria fail. DARVO arm rate is still
0 under greedy; ATTACK selection does not rise and falls for the rank-75
policy; group success collapses from 165/190/200 to 0 at 3/4/6 seats.
The mechanism is not the assumed one: greedy answers the pressure by
regulating, Stress plateaus at 3, so it never reaches the gate at 4 or the
arm at 5 — H1-A acts as a solve-rate tax and H1-B is unreachable under
competent play.

A harness defect was caught before the claim: sweep discarded refused
games silently and never reported its count, so "nobody won" and "nothing
played" printed identically. Reporting H1 as unwinnable on that basis
would have been the ADR-0018 family aimed at another repo's design. All
200 games ran in every cell; the zeros are real.

Chaos d8 = 8 — the window's first override, redrew L against a structural
L, so it changed nothing. Window 3 recorded in ChaosRollHistory.

NOT REVIEWED: tier L owes a separate-agent adversarial review, and no H1
result may reach ground-game until it has run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 00:50:08 +02:00

3.6 KiB
Raw Blame History

Chaos roll — window records

One entry per chaos window. Split out of InnerLoopReference.md when that file crossed the loadability limit: this is a log that grows, and a log inside a reference eventually crowds out the reference.

The rule itself lives in InnerLoop.md §Loop tiers; the current window's terms are there. This is the history the verdicts rest on.

Window 1 (d4, closed 2026-08-02): 12 declarations, 2 overrides, one each way, and both changed the outcome. The mechanism was kept and the rate dropped d4 → d8 (CB-EV-0015 §5, CB-EV-0016 §4).

Window 2 (d8, 2026-08-03 → 2026-08-07): 12 declarations, of which 11 rolled at d8 — declaration 1 (CB-WP-0018) opened the window and rolled at the old d4. Expected 8s: 1.375. Observed: 1.

decl pass roll effect
1 CB-WP-0018 d4 = 3 opened the window at the old rate
2 CB-WP-0019 d8 = 5
3 CB-WP-0020 d8 = 8 override — drew S over a structural S: changed nothing
4 CB-WP-0021 d8 = 7
59 CB-WP-0022…0026 d8 = 6 ×5
10 CB-WP-0027 d8 = 7
11 CB-WP-0028 d8 = 1
12 CB-WP-0029 d8 = 4

Verdict (ADR-0017): the rate is behaving as designed. One override against 1.375 expected is not a shortage of evidence.

Why the retirement condition changed. "An override changes nothing twice running" requires a consecutive pair, each with P = 1/3, so ~12 overrides are expected before one occurs — at ~1.4 overrides per window, ~9 windows or roughly 100 declarations. A gate that cannot cash out on any realistic horizon is decoration (ADR-0006 D3). It is now evaluated per window, needing two consecutive qualifying windows: ~24 declarations rather than ~100.

Four evidence files claimed window 2 produced zero overrides. They were wrong, and each cited the one before rather than counting. See F23 — the failure is a claim with no source, asserted once and repeated, which no gate here detects.

A post-hoc observation, deliberately not acted on. Declarations 59 rolled six five times running (~1 in 370 for some run of five in eleven rolls). shuf was tested over 200 rapid successive calls and looks uniform, longest run three. Recorded so a future window can check whether it recurs; not evidence of anything on its own.


Window 3 — opened 2026-08-07 at d8, running to 12 declarations

# pass roll override
1 CB-WP-0030 d8 = 7
2 CB-WP-0031 d8 = 2
3 CB-WP-0032 d8 = 6
4 CB-WP-0033 d8 = 7
5 CB-WP-0034 d8 = 4
6 CB-WP-0035 d8 = 4
7 CB-WP-0036 (first declaration) d8 = 7
8 CB-WP-0036 (re-declared L→M) d8 = 7
9 CB-WP-0037 d8 = 1
10 CB-WP-0038 d8 = 8 yes — redraw L, structural was L, so it changed nothing

The window's first 8, at declaration 10. Expectation over ten rolls at d8 is 1.25; one is exactly on rate.

The override changed nothing, which is the observation ADR-0017 D2's retirement condition is built from — it needs a full window whose overrides all change nothing, in two consecutive windows. This window now has one qualifying override and two declarations left to run.

Declaration 8 is a re-declaration of the same workplan, counted separately because it was a materially different pass: CB-WP-0036 was re-scoped from L to M after the maintainer moved the animation work out of the repo, and a changed declaration is a new declaration or the roll is not binding on what was actually built.