Some checks failed
ci / check (push) Failing after 3s
ground-game ruled all ten U-items on 2026-08-03, every one CONFIRMED as the default clay-borg simulates, and confirmed five of six provisional scenarios. clay-borg never collected the answers: CB-RES-0007 reported "0 of 10 ruled" the same day, and CB-WP-0022 built the finding register two days later still recording them as `reported`. make design's first run is what noticed -- not a human, not the adversarial review that found four other things. That is the unread-inbox failure running in the opposite direction, and it appears nowhere in the declaration, survey, ADR or spec of the pass that was built entirely around the forward version. It is arguably worse: an unread message is visible as silence, while a collected-but-unapplied ruling looks exactly like work in progress. Ten rulings quoted into §Underdetermined (the three conditional ones verbatim -- U1's designer note, U2's End-only trigger, U8's consume-only-if-it-cancels). Five provisional flags lifted, replaced by ruled/ruled_by/ruled_note so the flag went and the provenance stayed. Register queue 9 -> 0. T02's control came back clean: make sim is 26 passed, 59 rules covered, nothing red. Had a scenario gone red it would have meant we described our own behaviour incorrectly to ground-game. I wrote two U-item mappings and both were wrong. gr-a04 -> U1 (it asserts consent is REQUIRED; U1 asks WHEN the target accepts) and gr-d05 -> U5 (it exercises the UNREJECTED Reverse; U5 is the rejected one). Both plausible from covers:, neither survived reading the description. Third and fourth instance of this defect; the first two reached ground-game. So encodes_u_item is now a declaration and design.py asserts the file names what it claims -- and that check's own first version grepped for mentions and went red when two files recorded why they do NOT encode U1 and U5. A mention is not a claim, which is exactly the looseness that let "six of the ten have provisional scenarios" stand. Two positive controls went red for the best possible reason, both broken the same way -- asserting against live repo data instead of constructing their condition. rule-coverage.py required at least one provisional item to EXIST; it now builds a fixture and reports the live count as a diagnostic, because there is no number of provisional items this project should have. design-baseline.py pinned "2 of 6" while recomputing one row from a live glob, so the dated snapshot was never a snapshot; frozen to its 2026-08-03 list and unwired from self-tests, since per ADR-0012 D8 it is no longer a reporting tool. ScenarioFile is deny_unknown_fields and refused the four new fields until declared -- correct: a corpus accepting unknown metadata would let a typo'd encodes_u_iem sit there claiming nothing. DEVIATION: ADR-0012 D2 said "no new file". GroundRules.md crossed the loadability limit, so the register moved to specs/FindingRegister.md. D2's substance holds -- one register, same machinery, nothing competing -- but the literal instruction did not, and it resolves an awkwardness D2 named itself. make all: exit 0. loop-lint clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
209 lines
9.2 KiB
Markdown
209 lines
9.2 KiB
Markdown
# CB-EV-0022 — collect the rulings
|
|
|
|
CB-WP-0026 T05. Tier S (structural S — applies rulings inside an existing
|
|
capability; chaos d8=6 → no override). Declaration 9 of chaos window 2.
|
|
Closed 2026-08-05.
|
|
|
|
**Delivered:** ten rulings recorded in `§Underdetermined`, five
|
|
`provisional: true` flags lifted, four scenario schema fields, the
|
|
`encodes_u_item` declaration, and an empty finding queue.
|
|
|
|
---
|
|
|
|
## 1. How long the answers sat, and what noticed them
|
|
|
|
**Two days**, and the thing that noticed was **the register's first run** —
|
|
not a human, not the adversarial review.
|
|
|
|
`ground-game` ruled all ten U-items on **2026-08-03** (GROUND-WP-0002 T05),
|
|
every one confirmed. On the same day CB-RES-0007 reported *"0 of 10
|
|
ruled."* Two days later CB-WP-0022 built the register recording them as
|
|
`reported`, and the adversarial review — which found four other things —
|
|
did not catch it either. `make design` did, on its first execution.
|
|
|
|
**This is the symmetric failure nobody designed for.** CB-WP-0022 was
|
|
shaped end to end around *we send findings and nobody reads them*: the
|
|
four-day unread inbox is quoted in the declaration, the survey, the ADR
|
|
and the spec. The mirror case — *they answer and we do not collect it* —
|
|
appears in none of them.
|
|
|
|
It is arguably the worse of the two. An unread message is visible as
|
|
silence; a collected-but-unapplied ruling looks exactly like work in
|
|
progress.
|
|
|
|
## 2. Did any confirmed default fail to match the kernel? No.
|
|
|
|
This was the control that mattered. Every ruling was a *confirmation* of
|
|
what we told `ground-game` we simulate — so a red scenario would have
|
|
meant **we described our own behaviour incorrectly to them**, a defect in
|
|
our report rather than in their ruling, and the most serious class
|
|
available since it would be about our own code.
|
|
|
|
```
|
|
make sim → 26 passed, 59 rules covered
|
|
```
|
|
|
|
**No scenario went red.** The five confirmed defaults are the behaviour
|
|
implemented. That is the strongest single result here and it is a
|
|
negative: nothing was wrong.
|
|
|
|
## 3. The mapping gap was not what the survey said, and I reproduced the defect writing it
|
|
|
|
CB-RES-0007: *"six of the ten already have provisional scenarios."*
|
|
|
|
**Measured: one.** And the interesting part is how the other nine were
|
|
lost.
|
|
|
|
I wrote two mappings from the `covers:` lists and **both were wrong**:
|
|
|
|
| claimed | why it was withdrawn |
|
|
|---|---|
|
|
| `gr-a04-bond-support` → U1 | asserts consent is **required**; U1 asks **when** the target accepts |
|
|
| `gr-d05-darvo-reverse` → U5 | exercises the **unrejected** REVERSE; U5 is the **rejected** one (GROUND—ND) |
|
|
|
|
Both were plausible from `covers:`. Neither survived reading the
|
|
description. **These are the third and fourth instances of this exact
|
|
defect** — a link that looks right from metadata, asserted without
|
|
checking what the artifact exercises — and the first two reached
|
|
`ground-game`.
|
|
|
|
So the answer to *"do the four unlinked scenarios encode U-item defaults
|
|
at all?"* is: **at least two of the four do not**, and the survey's "six of
|
|
ten" was not an under-documented truth. It was wrong.
|
|
|
|
`encodes_u_item` is now a declaration a scenario makes or omits, and
|
|
`design.py` asserts the file names what it claims.
|
|
|
|
### The check's own first version was the same looseness
|
|
|
|
Written as `grep -lE "\bU<n>\b"`, it went red the moment two scenarios
|
|
recorded *why they do not* encode U1 and U5 — reporting `['U1','U2','U5']`.
|
|
|
|
**A mention is not a claim.** That is precisely the imprecision that let
|
|
*"six of the ten have provisional scenarios"* stand unchallenged for five
|
|
days: someone grepped for U-item strings and counted hits. The check now
|
|
asserts on the declaration.
|
|
|
|
## 4. The queue reached 0
|
|
|
|
```
|
|
QUEUE (open findings) (none)
|
|
open, lacking a reproduction 0 target 0
|
|
closed (log) 12 [U1..U10, F11, F13]
|
|
with a resolving reproduction 3/12 = 25%
|
|
notes 2 F12, F14
|
|
```
|
|
|
|
**`9 → 0`**, the number CB-WP-0026 T04 named. First evidence that
|
|
ADR-0012 D5's lifecycle is real rather than drawn: findings entered a
|
|
state, moved through it, and left the queue.
|
|
|
|
**25% reproduced must not be read as a failure.** Nine U-items closed by a
|
|
**ruling**, and a ruling is not an artifact. The metric is now honest
|
|
about something the survey's 33% concealed: most of our findings close
|
|
because someone answered them, not because anything demonstrates them.
|
|
|
|
**What the lifecycle could not express** — the honest gap: `applied` is
|
|
defined as *the source changed and the provisional default was deleted*.
|
|
Here the rulings **confirmed** our defaults, so nothing in the rules moved;
|
|
what changed is that the flags came off. The state fits, but the
|
|
definition had to be read generously. If a future ruling *overturns* a
|
|
default, `applied` will mean something materially different from what it
|
|
meant today, and D5 does not distinguish them.
|
|
|
|
## 5. Two things the schema caught
|
|
|
|
**`ScenarioFile` is `deny_unknown_fields`**, so five scenarios failed to
|
|
parse until `ruled`, `ruled_by`, `ruled_note` and `encodes_u_item` were
|
|
declared in the Rust struct. A corpus that accepted unknown metadata would
|
|
let a typo'd `encodes_u_iem` sit forever claiming nothing — and this
|
|
pass's whole subject is claims nobody checks.
|
|
|
|
**`record.rs` had to set them explicitly.** A recorded game is evidence of
|
|
what happened, not a claim about an undetermined rule. Filling the fields
|
|
via `..Default::default()` would have been shorter and would let a
|
|
recording silently inherit a U-item claim, pointing a reproduction at a
|
|
finding it has nothing to do with.
|
|
|
|
## 5b. Two gates went red for the best possible reason
|
|
|
|
Lifting the last five provisional flags broke two positive controls, and
|
|
**both were broken in the same way**: they asserted against live repo data
|
|
instead of constructing the condition they test.
|
|
|
|
**`rule-coverage.py`** required `bool(prov)` — *at least one provisional
|
|
item must exist*. That guard was the right instinct (a control that passes
|
|
vacuously is worthless) wired the wrong way. With nothing provisional, it
|
|
went red. It now builds a fixture, asserts the missing-owner case is
|
|
**caught**, and reports the live count as a diagnostic — because **there is
|
|
no number of provisional items this project should have.**
|
|
|
|
**`design-baseline.py`** pinned *"the measured baseline is 2 of 6"* and
|
|
reported `1/6`. Its "six provisional defaults" row **globbed
|
|
`provisional: true` at run time**, so the dated snapshot was never a
|
|
snapshot — it drifted with the repo. C2 dismantled this tool four hours
|
|
earlier and missed this: a hand-maintained dict with one dynamically
|
|
computed row is worse than a fully hand-maintained one, because the
|
|
recomputed row silently disagrees with the date in the header.
|
|
|
|
Frozen to the literal list it measured on 2026-08-03, and **removed from
|
|
`make self-tests`** — that target is *"a positive control for every
|
|
reporting tool"*, and per ADR-0012 D8 this is no longer a reporting tool.
|
|
Leaving it wired in meant a superseded instrument could fail the build.
|
|
|
|
**The pattern across both**: a control that reads the world it is meant to
|
|
audit will eventually audit a world that has moved. Neither was caught by
|
|
review; both were caught by the world moving.
|
|
|
|
## 6. The register moved, and ADR-0012 D2 gave way to loadability
|
|
|
|
`specs/GroundRules.md` crossed the ~400-line limit and `loop-lint`
|
|
required a split. ADR-0012 D2 said **"no new file"**, so this is a
|
|
deviation and is recorded as one.
|
|
|
|
**D2's substance holds.** Its argument was *one register, not a second
|
|
mechanism competing with the first* — and `specs/FindingRegister.md` is
|
|
that same register, moved, still driving off the `provisional`/ruling
|
|
machinery, still the only one. What was traded away is the literal "no new
|
|
file", which was D2's *implementation*, not its reason.
|
|
|
|
It also resolves an awkwardness D2 named itself: *"a finding about the
|
|
engine's behaviour sits in a document about the game."* Now it does not.
|
|
|
|
## 7. Chaos window 2
|
|
|
|
**Declaration 9 of 12.** Structural S, d8 = 6, no override. Recorded per
|
|
§Loop tiers even though it changed nothing.
|
|
|
|
**Window 2 still has no override to evaluate** — nine declarations, zero
|
|
8s. The retirement condition (*retire if an override changes nothing twice
|
|
running*) cannot be assessed, and at d8 the expected count over twelve
|
|
declarations is 1.5, so this is unremarkable rather than evidence of
|
|
anything.
|
|
|
|
## 8. Cost
|
|
|
|
CB-WP-0022's cost, by re-running the instrument:
|
|
|
|
```
|
|
CB-WP-0022-T05 $ 6.43
|
|
CB-WP-0022-T01 $ 4.17
|
|
CB-WP-0022-T02 $ 3.56 (the separate reviewer's own spend is in the subagent tree)
|
|
CB-WP-0022-T04 $ 1.92
|
|
CB-WP-0022-T03 $ 1.01
|
|
```
|
|
|
|
**The chain did not snap this time** — CB-EV-0019 §4 predicted it might,
|
|
having found `cb-cost.py --slug CB-WP-0020` aborts for want of retained
|
|
transcripts. CB-WP-0022 is recent enough to still be in the window. **The
|
|
bound CB-EV-0019 asked for is still owed**; this pass is evidence that the
|
|
rule works for a pass one step back, not that it works generally.
|
|
|
|
## Open after this pass
|
|
|
|
- **Nine U-items are `applied` with no reproduction.** That is recorded,
|
|
not hidden, but it means nine rules rest on a ruling nobody can re-run.
|
|
Cheap to fix incrementally: each needs one scenario naming its item.
|
|
- **`applied` conflates *confirmed* with *overturned*** (§4). It will
|
|
matter the first time a ruling goes against us.
|
|
- **The cost-chain bound** (CB-EV-0019 §4) is still unwritten.
|