clay-borg/specs/ChaosRollHistory.md

126 lines
5.6 KiB
Markdown
Raw Normal View History

ADR-0017: window 2's verdict — the mechanism worked, my account of it did not Tier M (changes how the loop constrains its own operation), declared at d8 because the rate for window 3 is what this document decides and declaring at a rate it invents would be circular. chaos d8 = 7, no override. I CLAIMED WINDOW 2 PRODUCED ZERO OVERRIDES, FIVE TIMES, AND IT IS FALSE. Declaration 3 (CB-WP-0020) rolled d8 = 8, overrode, drew S against a structural S, and changed nothing -- and CB-WP-0020 recorded it correctly at the time, in those words: "the first override at d8... It changed nothing... One." Counting the workplans takes one command and I never ran it. CB-EV-0024 asserted "zero" without checking; CB-EV-0025, 0026, 0027 and CB-WP-0029 each cited the one before. A claim propagated five times by citation rather than by measurement, in files whose subject was that exact failure. facts-check catches a copied number that disagrees with its source; nothing catches a number with NO source, asserted once and repeated. Registered F23, and all four evidence files carry an in-place correction rather than a silent edit (ADR-0012 D5). THE ACTUAL VERDICT: THE RATE IS WORKING. Eleven rolls at d8 -- declaration 1 opened the window at the old d4 -- against 1.375 eights expected, 1 observed. Not a shortage of evidence; the design. BUT THE RETIREMENT CONDITION GENUINELY CANNOT FIRE, and that took computing to see. "An override changes nothing twice running" needs a consecutive pair at P=1/3 each, so ~12 overrides expected, at ~1.4 per window: ~9 windows, roughly 100 declarations. A gate that cannot cash out on any realistic horizon is decoration, which ADR-0006 D3 forbids. Restated to be evaluated PER WINDOW: retire if a full window's overrides all change nothing, met in two consecutive windows. A window with no overrides is inconclusive and advances nothing. ~24 declarations rather than ~100. Window 2 counts as the first; window 3 opens at d8 and decides. Recorded and deliberately not acted on: declarations 5-9 rolled six five times running, ~1 in 370 for some run of five in eleven rolls. shuf tested over 200 rapid successive calls looks uniform, longest run three. Found post hoc, which is how coincidences become findings, so it is logged for a future window to check rather than treated as evidence. InnerLoop.md then crossed the loadability limit, and so did InnerLoopReference.md. The window log moved to specs/ChaosRollHistory.md: it grows by one entry per window, and a log inside a reference eventually crowds out the reference. make all: exit 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 10:55:54 +02:00
# Chaos roll — window records
ADR-0021: the chaos roll is retired, and the condition that retired it was wrong The condition is met and the tally was verified rather than recalled. ADR-0017 D2 named window 3 as the decider; windows 2 and 3 each produced exactly one override and each changed nothing. Every roll in window 3 was cross-checked against the workplan that made it, because the last time this project tallied chaos rolls from memory it was wrong and asserted the wrong figure four times (F23). All twelve agree. But meeting the condition is not evidence. P(an override changes nothing) is 1/3, and each window had exactly one override, so the condition fires on a 1/9 coincidence. ADR-0017 restated it to be REACHABLE and made it weak in the process; reachability was checked and discriminating power was not. Worse, it measures the wrong subject. It asks whether an override changed the tier; the question is whether changing the tier helped. Under it, a die that always changed the tier could never be retired however useless its changes were. The real ground is stronger. Four overrides across roughly forty declarations, and the mechanism's value has never once been demonstrated. The one substantive intervention dropped CB-WP-0011 from a structural L to S, and that work then needed CB-WP-0016 and CB-WP-0017 to fix defects a human found by playing. Not offered as causation — a tier is process weight, not a guarantee — but it is the only evidence we have about an override's consequences and it points the wrong way. And the purpose has no live evidence of need: 17 M, 11 S, 4 L across every workplan, with the only two structural/declared mismatches being the window-1 overrides themselves. Tier declaration has not ossified. So: retired, with nothing replacing it. Adding a successor to guard against ossification that has not occurred would invent a gate for an instance we do not have. The revival trigger is stated: tiers collapsing toward one value, or a pass declaring below its structural tier to dodge a review. InnerLoop loses the chaos paragraph, loop-lint loses the chaos-recorded check — a check that outlives its rule becomes an obstruction — and its self-test now asserts the opposite: a note with no roll must pass. ChaosRollHistory is closed. Window 4 ends incomplete at four declarations, the last of which is this one, rolled because the rule was still in force. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 15:51:52 +02:00
> **CLOSED 2026-08-08.** The mechanism is retired
> ([ADR-0021](../decisions/ADR-0021-the-chaos-roll-is-retired.md)). This
> file is history and takes no new entries.
>
> **Four windows, four overrides, no demonstrated benefit.** The
> retirement condition was met — and was itself wrong: it fires on a 1/9
> coincidence, and it measured whether an override *changed the tier*
> rather than whether the change *helped*. The real ground was D4's
> tally, not the condition.
ADR-0017: window 2's verdict — the mechanism worked, my account of it did not Tier M (changes how the loop constrains its own operation), declared at d8 because the rate for window 3 is what this document decides and declaring at a rate it invents would be circular. chaos d8 = 7, no override. I CLAIMED WINDOW 2 PRODUCED ZERO OVERRIDES, FIVE TIMES, AND IT IS FALSE. Declaration 3 (CB-WP-0020) rolled d8 = 8, overrode, drew S against a structural S, and changed nothing -- and CB-WP-0020 recorded it correctly at the time, in those words: "the first override at d8... It changed nothing... One." Counting the workplans takes one command and I never ran it. CB-EV-0024 asserted "zero" without checking; CB-EV-0025, 0026, 0027 and CB-WP-0029 each cited the one before. A claim propagated five times by citation rather than by measurement, in files whose subject was that exact failure. facts-check catches a copied number that disagrees with its source; nothing catches a number with NO source, asserted once and repeated. Registered F23, and all four evidence files carry an in-place correction rather than a silent edit (ADR-0012 D5). THE ACTUAL VERDICT: THE RATE IS WORKING. Eleven rolls at d8 -- declaration 1 opened the window at the old d4 -- against 1.375 eights expected, 1 observed. Not a shortage of evidence; the design. BUT THE RETIREMENT CONDITION GENUINELY CANNOT FIRE, and that took computing to see. "An override changes nothing twice running" needs a consecutive pair at P=1/3 each, so ~12 overrides expected, at ~1.4 per window: ~9 windows, roughly 100 declarations. A gate that cannot cash out on any realistic horizon is decoration, which ADR-0006 D3 forbids. Restated to be evaluated PER WINDOW: retire if a full window's overrides all change nothing, met in two consecutive windows. A window with no overrides is inconclusive and advances nothing. ~24 declarations rather than ~100. Window 2 counts as the first; window 3 opens at d8 and decides. Recorded and deliberately not acted on: declarations 5-9 rolled six five times running, ~1 in 370 for some run of five in eleven rolls. shuf tested over 200 rapid successive calls looks uniform, longest run three. Found post hoc, which is how coincidences become findings, so it is logged for a future window to check rather than treated as evidence. InnerLoop.md then crossed the loadability limit, and so did InnerLoopReference.md. The window log moved to specs/ChaosRollHistory.md: it grows by one entry per window, and a log inside a reference eventually crowds out the reference. make all: exit 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 10:55:54 +02:00
One entry per chaos window. Split out of `InnerLoopReference.md` when that
file crossed the loadability limit: this is a **log that grows**, and a log
inside a reference eventually crowds out the reference.
The rule itself lives in [`InnerLoop.md`](InnerLoop.md) §Loop tiers; the
current window's terms are there. This is the history the verdicts rest on.
**Window 1** (d4, closed 2026-08-02): 12 declarations, **2 overrides**, one
each way, and **both changed the outcome**. The mechanism was kept and the
rate dropped d4 → d8 (CB-EV-0015 §5, CB-EV-0016 §4).
**Window 2** (d8, 2026-08-03 → 2026-08-07): 12 declarations, of which
**11 rolled at d8** — declaration 1 (CB-WP-0018) opened the window and
rolled at the old d4. Expected 8s: **1.375**. Observed: **1**.
| decl | pass | roll | effect |
|---:|---|---|---|
| 1 | CB-WP-0018 | d4 = 3 | opened the window at the old rate |
| 2 | CB-WP-0019 | d8 = 5 | — |
| **3** | **CB-WP-0020** | **d8 = 8** | **override — drew S over a structural S: changed nothing** |
| 4 | CB-WP-0021 | d8 = 7 | — |
| 59 | CB-WP-0022…0026 | d8 = 6 ×5 | — |
| 10 | CB-WP-0027 | d8 = 7 | — |
| 11 | CB-WP-0028 | d8 = 1 | — |
| 12 | CB-WP-0029 | d8 = 4 | — |
**Verdict (ADR-0017): the rate is behaving as designed.** One override
against 1.375 expected is not a shortage of evidence.
**Why the retirement condition changed.** *"An override changes nothing
twice running"* requires a consecutive pair, each with P = 1/3, so ~12
overrides are expected before one occurs — at ~1.4 overrides per window,
**~9 windows or roughly 100 declarations**. A gate that cannot cash out on
any realistic horizon is decoration (ADR-0006 D3). It is now evaluated
**per window**, needing two consecutive qualifying windows: ~24
declarations rather than ~100.
**Four evidence files claimed window 2 produced zero overrides.** They were
wrong, and each cited the one before rather than counting. See F23 —
the failure is a claim with no source, asserted once and repeated, which
no gate here detects.
**A post-hoc observation, deliberately not acted on.** Declarations 59
rolled six five times running (~1 in 370 for some run of five in eleven
rolls). `shuf` was tested over 200 rapid successive calls and looks
uniform, longest run three. Recorded so a future window can check whether
it recurs; not evidence of anything on its own.
CB-WP-0038: variant selection, H1 implemented, and H1 measured ground-game packages hypotheses as selectable rules variants — a catalog, a rules_delta.yaml, and prose — and their note is explicit that CSV text alone is not executable here. So the kernel gains a Variant in game state: in the state, therefore in the hash, therefore in the recording, because a scenario replayed under a different variant would diverge silently. Baseline is bit-for-bit what it was, asserted across seat counts and seeds. A variant system that perturbs the baseline invalidates every measurement this repo has. H1-A and H1-B implemented from rules_delta.yaml and mutation-proven on their own defects: "unclaimed" misread as face-up-and-unsolved, and the attacker's Stress read after the attack's effects. Their `unchanged:` list is asserted rather than trusted — that list is their claim about their own experiment. Measured, and three of their four criteria fail. DARVO arm rate is still 0 under greedy; ATTACK selection does not rise and falls for the rank-75 policy; group success collapses from 165/190/200 to 0 at 3/4/6 seats. The mechanism is not the assumed one: greedy answers the pressure by regulating, Stress plateaus at 3, so it never reaches the gate at 4 or the arm at 5 — H1-A acts as a solve-rate tax and H1-B is unreachable under competent play. A harness defect was caught before the claim: sweep discarded refused games silently and never reported its count, so "nobody won" and "nothing played" printed identically. Reporting H1 as unwinnable on that basis would have been the ADR-0018 family aimed at another repo's design. All 200 games ran in every cell; the zeros are real. Chaos d8 = 8 — the window's first override, redrew L against a structural L, so it changed nothing. Window 3 recorded in ChaosRollHistory. NOT REVIEWED: tier L owes a separate-agent adversarial review, and no H1 result may reach ground-game until it has run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 00:50:08 +02:00
---
## Window 3 — opened 2026-08-07 at d8, running to 12 declarations
| # | pass | roll | override |
|---|---|---|---|
| 1 | CB-WP-0030 | d8 = 7 | — |
| 2 | CB-WP-0031 | d8 = 2 | — |
| 3 | CB-WP-0032 | d8 = 6 | — |
| 4 | CB-WP-0033 | d8 = 7 | — |
| 5 | CB-WP-0034 | d8 = 4 | — |
| 6 | CB-WP-0035 | d8 = 4 | — |
| 7 | CB-WP-0036 (first declaration) | d8 = 7 | — |
| 8 | CB-WP-0036 (re-declared L→M) | d8 = 7 | — |
| 9 | CB-WP-0037 | d8 = 1 | — |
| 10 | CB-WP-0038 | **d8 = 8** | **yes — redraw L, structural was L, so it changed nothing** |
CB-WP-0039: a seat that does not regulate — and it changes H1's verdict CB-EV-0030 concluded H1's DARVO arm rate was still 0. That was true of the panel, and the panel was greedy-family throughout. GreedyPolicy ranks `Ground if gated => 100`, so it grounds the instant the stress gate bites, Stress plateaus at 3, and the arm at 5 is unreachable by construction. "H1 does nothing" was really "H1 does nothing to a seat that already manages its Stress" — and H1 was written for the seat that does not. `reactive` is greedy with exactly one preference changed: GROUND demoted below ATTACK. Under it, H1's criteria 1 and 2 are MET — DARVO arms 400 times per cell, ATTACK is chosen 3 times per seat per game. Criterion 3 fails harder: reactive wins nothing at any seat count. The larger finding is about the baseline. Greedy and reactive play IDENTICALLY under baseline, and peak Stress across 3,200 baseline games was 1 — against a starting value of 2. The gate at 4, the DARVO arm at 5 and the Freedom token are all unreachable, and a policy built to be reckless with Stress is indistinguishable from one built to husband it. That is a deeper account of F17 than F17 has. Not raised as a finding yet: it wants the plural panel first. A constant was investigated rather than reported: darvo was exactly 400 in every cell while atk scaled with seats. Six-player final Stress is [5,5,4,4,4,4] every seed — H1-B holds the attacker at 4, below the arm, and pushes its targets to 5. The self-soothe suppresses DARVO in the aggressor and concentrates it in the attacked. The direction follows from H1-B's arithmetic; the number 2 is partly an artifact of reactive's first-legal targeting, and is labelled as such. Still unreviewed: tier L review outstanding on CB-WP-0038, and nothing here reaches ground-game until it runs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 01:08:40 +02:00
| 11 | CB-WP-0039 | d8 = 2 | — |
CB-WP-0040: name the stratum before naming the defect The maintainer could not tell whether "error", "failure", "finding" or "correction" referred to the game's design, our formalisation of it, the code, the measuring apparatus, or the sentences we wrote. Three review rounds produced twenty-odd defect statements spanning five systems, all called errors. The confusion was ours. specs/Taxonomy.md, grounded in named canon rather than invented here: six strata from Sargent's problem entity / conceptual model / computerized model, extended where a simulation-V&V frame stops — we also own an instrument and an account. The two relations are what was missing: GAME<->MODEL is validation, MODEL<->ENGINE is verification, and nearly every argument about "our bug or their gap" was that distinction going unnamed. Fault/error/failure from Avizienis et al., applied within a stratum, plus the rule that explains the review history: a failure in one stratum is a fault in the next. And it finally defines the family ADR-0018 could only point at — a wrong-subject error is an ACCOUNT failure with no INSTRUMENT fault, which is why tests never catch them. MDA supplies the game-facing layers and one hard limit: our panels measure dynamics, our trial logs sample aesthetics, and a win rate does not answer "is it fun". specs/Positioning.md names the field fairly — Ludii is the closest relative and the right benchmark — and the four differentiators, each already built rather than aspired to. Clay-borg is a design-evidence instrument; anyone can produce the number. Three tracks named and none started: a second game, game theory as the lens on dynamics, and assimilated knowledge about why games work. Track A is the falsifier for the whole positioning: every abstraction here has exactly one instance, which by our own rule may mean invented rather than observed. Chaos window 3 closes at 12 declarations with one override that changed nothing. Its verdict is due and is deliberately not written here. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 11:41:16 +02:00
| 12 | CB-WP-0040 | d8 = 5 | — |
CB-WP-0038: variant selection, H1 implemented, and H1 measured ground-game packages hypotheses as selectable rules variants — a catalog, a rules_delta.yaml, and prose — and their note is explicit that CSV text alone is not executable here. So the kernel gains a Variant in game state: in the state, therefore in the hash, therefore in the recording, because a scenario replayed under a different variant would diverge silently. Baseline is bit-for-bit what it was, asserted across seat counts and seeds. A variant system that perturbs the baseline invalidates every measurement this repo has. H1-A and H1-B implemented from rules_delta.yaml and mutation-proven on their own defects: "unclaimed" misread as face-up-and-unsolved, and the attacker's Stress read after the attack's effects. Their `unchanged:` list is asserted rather than trusted — that list is their claim about their own experiment. Measured, and three of their four criteria fail. DARVO arm rate is still 0 under greedy; ATTACK selection does not rise and falls for the rank-75 policy; group success collapses from 165/190/200 to 0 at 3/4/6 seats. The mechanism is not the assumed one: greedy answers the pressure by regulating, Stress plateaus at 3, so it never reaches the gate at 4 or the arm at 5 — H1-A acts as a solve-rate tax and H1-B is unreachable under competent play. A harness defect was caught before the claim: sweep discarded refused games silently and never reported its count, so "nobody won" and "nothing played" printed identically. Reporting H1 as unwinnable on that basis would have been the ADR-0018 family aimed at another repo's design. All 200 games ran in every cell; the zeros are real. Chaos d8 = 8 — the window's first override, redrew L against a structural L, so it changed nothing. Window 3 recorded in ChaosRollHistory. NOT REVIEWED: tier L owes a separate-agent adversarial review, and no H1 result may reach ground-game until it has run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 00:50:08 +02:00
**The window's first 8, at declaration 10.** Expectation over ten rolls at
d8 is 1.25; one is exactly on rate.
**The override changed nothing**, which is the observation ADR-0017 D2's
retirement condition is built from — it needs *a full window whose
overrides all change nothing*, in two consecutive windows. This window now
CB-WP-0040: name the stratum before naming the defect The maintainer could not tell whether "error", "failure", "finding" or "correction" referred to the game's design, our formalisation of it, the code, the measuring apparatus, or the sentences we wrote. Three review rounds produced twenty-odd defect statements spanning five systems, all called errors. The confusion was ours. specs/Taxonomy.md, grounded in named canon rather than invented here: six strata from Sargent's problem entity / conceptual model / computerized model, extended where a simulation-V&V frame stops — we also own an instrument and an account. The two relations are what was missing: GAME<->MODEL is validation, MODEL<->ENGINE is verification, and nearly every argument about "our bug or their gap" was that distinction going unnamed. Fault/error/failure from Avizienis et al., applied within a stratum, plus the rule that explains the review history: a failure in one stratum is a fault in the next. And it finally defines the family ADR-0018 could only point at — a wrong-subject error is an ACCOUNT failure with no INSTRUMENT fault, which is why tests never catch them. MDA supplies the game-facing layers and one hard limit: our panels measure dynamics, our trial logs sample aesthetics, and a win rate does not answer "is it fun". specs/Positioning.md names the field fairly — Ludii is the closest relative and the right benchmark — and the four differentiators, each already built rather than aspired to. Clay-borg is a design-evidence instrument; anyone can produce the number. Three tracks named and none started: a second game, game theory as the lens on dynamics, and assimilated knowledge about why games work. Track A is the falsifier for the whole positioning: every abstraction here has exactly one instance, which by our own rule may mean invented rather than observed. Chaos window 3 closes at 12 declarations with one override that changed nothing. Its verdict is due and is deliberately not written here. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 11:41:16 +02:00
had one qualifying override, and it changed nothing.
**Window 3 is closed at 12 declarations. Its verdict is due and is not
written here** — recording it is a change to how the loop constrains its
own operation, which is a tier-M trigger in its own right, and window 2's
verdict was delayed the same way. **One override in twelve, changing
nothing**, is the second consecutive window to produce no override that
changed an outcome; ADR-0017 D2's retirement condition asks for exactly
that in two consecutive windows and should now be evaluated rather than
restated.
CB-WP-0038: variant selection, H1 implemented, and H1 measured ground-game packages hypotheses as selectable rules variants — a catalog, a rules_delta.yaml, and prose — and their note is explicit that CSV text alone is not executable here. So the kernel gains a Variant in game state: in the state, therefore in the hash, therefore in the recording, because a scenario replayed under a different variant would diverge silently. Baseline is bit-for-bit what it was, asserted across seat counts and seeds. A variant system that perturbs the baseline invalidates every measurement this repo has. H1-A and H1-B implemented from rules_delta.yaml and mutation-proven on their own defects: "unclaimed" misread as face-up-and-unsolved, and the attacker's Stress read after the attack's effects. Their `unchanged:` list is asserted rather than trusted — that list is their claim about their own experiment. Measured, and three of their four criteria fail. DARVO arm rate is still 0 under greedy; ATTACK selection does not rise and falls for the rank-75 policy; group success collapses from 165/190/200 to 0 at 3/4/6 seats. The mechanism is not the assumed one: greedy answers the pressure by regulating, Stress plateaus at 3, so it never reaches the gate at 4 or the arm at 5 — H1-A acts as a solve-rate tax and H1-B is unreachable under competent play. A harness defect was caught before the claim: sweep discarded refused games silently and never reported its count, so "nobody won" and "nothing played" printed identically. Reporting H1 as unwinnable on that basis would have been the ADR-0018 family aimed at another repo's design. All 200 games ran in every cell; the zeros are real. Chaos d8 = 8 — the window's first override, redrew L against a structural L, so it changed nothing. Window 3 recorded in ChaosRollHistory. NOT REVIEWED: tier L owes a separate-agent adversarial review, and no H1 result may reach ground-game until it has run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 00:50:08 +02:00
**Declaration 8 is a re-declaration of the same workplan**, counted
separately because it was a materially different pass: CB-WP-0036 was
re-scoped from L to M after the maintainer moved the animation work out of
the repo, and a changed declaration is a new declaration or the roll is
not binding on what was actually built.
CB-WP-0041 done: ADR-0020 refuses the port, and T02 is why T02 — all chance derives from one root seed. Three chance points, all reading it: the setup deck shuffle, the setup Lead draw, and the reshuffle permutation. The Problems deal is not chance at all. So in extensive-form terms the tree has a single chance node at the root. That test was wrong first, and the mutation caught it. It compared state hashes — and GroundState carries `seed` as a field, so "different seeds differ" was true by construction. Mutating the shuffle away left it green. It now compares the dealt configuration, and the same mutation fails it: a wrong-subject error inside the control written for T02. The reshuffle is a pure function of (seed, round) because K5 requires deterministic replay, where a real table reshuffles independently. That is a modelling restriction, not a defect, and it is now pinned. T03 — commit/reveal checked in both directions: before Reveal each seat sees its own selection and no other; after Reveal the information sets merge, because an encoding that hides forever is not commit/reveal either. T04 — ADR-0020 refuses the EFG port, and the blocker is T02 rather than T01, which inverts what the workplan expected. Perfect recall looked like the risk and is a constraint with a known answer: key on observation histories. Making chance explicit is the expensive one — the reshuffle would become a real chance node and break the K5 purity that every recording, replay bundle and trial-note hash depends on. A port would trade the property this project is built on for one it has never needed. Track B's first move is therefore a question, not a build: take "is exploitability meaningful for a co-operative game with a shared threshold" to OpenSpiel on a toy model, where answering it costs nothing. D4 states what being wrong looks like — OpenSpiel settling on a toy what three rounds of policy sweeps could not — and makes watching for it the next action. Taxonomy §4.1 records the EFG correspondence with the test that checks each row, so a later pass starts from a specification rather than a memory. Chaos window 4 at three declarations. Window 3's verdict is now two windows behind and should be evaluated rather than restated. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 15:35:03 +02:00
---
## Window 4 — opened 2026-08-08 at d8
| # | pass | roll | override |
|---|---|---|---|
| 1 | CB-WP-0041 | d8 = 5 | — |
| 2 | CB-WP-0042 | d8 = 4 | — |
| 3 | ADR-0020 | d8 = 7 | — |
ADR-0021: the chaos roll is retired, and the condition that retired it was wrong The condition is met and the tally was verified rather than recalled. ADR-0017 D2 named window 3 as the decider; windows 2 and 3 each produced exactly one override and each changed nothing. Every roll in window 3 was cross-checked against the workplan that made it, because the last time this project tallied chaos rolls from memory it was wrong and asserted the wrong figure four times (F23). All twelve agree. But meeting the condition is not evidence. P(an override changes nothing) is 1/3, and each window had exactly one override, so the condition fires on a 1/9 coincidence. ADR-0017 restated it to be REACHABLE and made it weak in the process; reachability was checked and discriminating power was not. Worse, it measures the wrong subject. It asks whether an override changed the tier; the question is whether changing the tier helped. Under it, a die that always changed the tier could never be retired however useless its changes were. The real ground is stronger. Four overrides across roughly forty declarations, and the mechanism's value has never once been demonstrated. The one substantive intervention dropped CB-WP-0011 from a structural L to S, and that work then needed CB-WP-0016 and CB-WP-0017 to fix defects a human found by playing. Not offered as causation — a tier is process weight, not a guarantee — but it is the only evidence we have about an override's consequences and it points the wrong way. And the purpose has no live evidence of need: 17 M, 11 S, 4 L across every workplan, with the only two structural/declared mismatches being the window-1 overrides themselves. Tier declaration has not ossified. So: retired, with nothing replacing it. Adding a successor to guard against ossification that has not occurred would invent a gate for an instance we do not have. The revival trigger is stated: tiers collapsing toward one value, or a pass declaring below its structural tier to dodge a review. InnerLoop loses the chaos paragraph, loop-lint loses the chaos-recorded check — a check that outlives its rule becomes an obstruction — and its self-test now asserts the opposite: a note with no roll must pass. ChaosRollHistory is closed. Window 4 ends incomplete at four declarations, the last of which is this one, rolled because the rule was still in force. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 15:51:52 +02:00
| 4 | ADR-0021 | d8 = 7 | — |
**Window 4 ends here, incomplete, at four declarations.** The condition
was met at the end of window 3; waiting out a fourth window to satisfy a
symmetry nobody requires would have been ceremony. **Declaration 4 is the
last chaos roll this project made** — the one that declared its
retirement, which was rolled because the rule was still in force.
**Window 3's verdict, delivered 2026-08-08** in
[ADR-0021](../decisions/ADR-0021-the-chaos-roll-is-retired.md): the
condition was met, and **the condition was defective**. Every roll in this
window was cross-checked against the workplan that made it before the
verdict was written, because the last tally taken from memory was wrong
four times over (F23).