clay-borg/specs/ChaosRollHistory.md
tegwick 713a9df7fd
Some checks failed
ci / check (push) Failing after 3s
CB-WP-0040: name the stratum before naming the defect
The maintainer could not tell whether "error", "failure", "finding" or
"correction" referred to the game's design, our formalisation of it, the
code, the measuring apparatus, or the sentences we wrote. Three review
rounds produced twenty-odd defect statements spanning five systems, all
called errors. The confusion was ours.

specs/Taxonomy.md, grounded in named canon rather than invented here: six
strata from Sargent's problem entity / conceptual model / computerized
model, extended where a simulation-V&V frame stops — we also own an
instrument and an account. The two relations are what was missing:
GAME<->MODEL is validation, MODEL<->ENGINE is verification, and nearly
every argument about "our bug or their gap" was that distinction going
unnamed.

Fault/error/failure from Avizienis et al., applied within a stratum, plus
the rule that explains the review history: a failure in one stratum is a
fault in the next. And it finally defines the family ADR-0018 could only
point at — a wrong-subject error is an ACCOUNT failure with no INSTRUMENT
fault, which is why tests never catch them.

MDA supplies the game-facing layers and one hard limit: our panels measure
dynamics, our trial logs sample aesthetics, and a win rate does not answer
"is it fun".

specs/Positioning.md names the field fairly — Ludii is the closest
relative and the right benchmark — and the four differentiators, each
already built rather than aspired to. Clay-borg is a design-evidence
instrument; anyone can produce the number. Three tracks named and none
started: a second game, game theory as the lens on dynamics, and
assimilated knowledge about why games work.

Track A is the falsifier for the whole positioning: every abstraction here
has exactly one instance, which by our own rule may mean invented rather
than observed.

Chaos window 3 closes at 12 declarations with one override that changed
nothing. Its verdict is due and is deliberately not written here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 11:41:16 +02:00

4.2 KiB
Raw Blame History

Chaos roll — window records

One entry per chaos window. Split out of InnerLoopReference.md when that file crossed the loadability limit: this is a log that grows, and a log inside a reference eventually crowds out the reference.

The rule itself lives in InnerLoop.md §Loop tiers; the current window's terms are there. This is the history the verdicts rest on.

Window 1 (d4, closed 2026-08-02): 12 declarations, 2 overrides, one each way, and both changed the outcome. The mechanism was kept and the rate dropped d4 → d8 (CB-EV-0015 §5, CB-EV-0016 §4).

Window 2 (d8, 2026-08-03 → 2026-08-07): 12 declarations, of which 11 rolled at d8 — declaration 1 (CB-WP-0018) opened the window and rolled at the old d4. Expected 8s: 1.375. Observed: 1.

decl pass roll effect
1 CB-WP-0018 d4 = 3 opened the window at the old rate
2 CB-WP-0019 d8 = 5
3 CB-WP-0020 d8 = 8 override — drew S over a structural S: changed nothing
4 CB-WP-0021 d8 = 7
59 CB-WP-0022…0026 d8 = 6 ×5
10 CB-WP-0027 d8 = 7
11 CB-WP-0028 d8 = 1
12 CB-WP-0029 d8 = 4

Verdict (ADR-0017): the rate is behaving as designed. One override against 1.375 expected is not a shortage of evidence.

Why the retirement condition changed. "An override changes nothing twice running" requires a consecutive pair, each with P = 1/3, so ~12 overrides are expected before one occurs — at ~1.4 overrides per window, ~9 windows or roughly 100 declarations. A gate that cannot cash out on any realistic horizon is decoration (ADR-0006 D3). It is now evaluated per window, needing two consecutive qualifying windows: ~24 declarations rather than ~100.

Four evidence files claimed window 2 produced zero overrides. They were wrong, and each cited the one before rather than counting. See F23 — the failure is a claim with no source, asserted once and repeated, which no gate here detects.

A post-hoc observation, deliberately not acted on. Declarations 59 rolled six five times running (~1 in 370 for some run of five in eleven rolls). shuf was tested over 200 rapid successive calls and looks uniform, longest run three. Recorded so a future window can check whether it recurs; not evidence of anything on its own.


Window 3 — opened 2026-08-07 at d8, running to 12 declarations

# pass roll override
1 CB-WP-0030 d8 = 7
2 CB-WP-0031 d8 = 2
3 CB-WP-0032 d8 = 6
4 CB-WP-0033 d8 = 7
5 CB-WP-0034 d8 = 4
6 CB-WP-0035 d8 = 4
7 CB-WP-0036 (first declaration) d8 = 7
8 CB-WP-0036 (re-declared L→M) d8 = 7
9 CB-WP-0037 d8 = 1
10 CB-WP-0038 d8 = 8 yes — redraw L, structural was L, so it changed nothing
11 CB-WP-0039 d8 = 2
12 CB-WP-0040 d8 = 5

The window's first 8, at declaration 10. Expectation over ten rolls at d8 is 1.25; one is exactly on rate.

The override changed nothing, which is the observation ADR-0017 D2's retirement condition is built from — it needs a full window whose overrides all change nothing, in two consecutive windows. This window now had one qualifying override, and it changed nothing.

Window 3 is closed at 12 declarations. Its verdict is due and is not written here — recording it is a change to how the loop constrains its own operation, which is a tier-M trigger in its own right, and window 2's verdict was delayed the same way. One override in twelve, changing nothing, is the second consecutive window to produce no override that changed an outcome; ADR-0017 D2's retirement condition asks for exactly that in two consecutive windows and should now be evaluated rather than restated.

Declaration 8 is a re-declaration of the same workplan, counted separately because it was a materially different pass: CB-WP-0036 was re-scoped from L to M after the maintainer moved the animation work out of the repo, and a changed declaration is a new declaration or the roll is not binding on what was actually built.