ground-game ruled all ten U-items on 2026-08-03, every one CONFIRMED as the default clay-borg simulates, and confirmed five of six provisional scenarios. clay-borg never collected the answers: CB-RES-0007 reported "0 of 10 ruled" the same day, and CB-WP-0022 built the finding register two days later still recording them as `reported`. make design's first run is what noticed -- not a human, not the adversarial review that found four other things. That is the unread-inbox failure running in the opposite direction, and it appears nowhere in the declaration, survey, ADR or spec of the pass that was built entirely around the forward version. It is arguably worse: an unread message is visible as silence, while a collected-but-unapplied ruling looks exactly like work in progress. Ten rulings quoted into §Underdetermined (the three conditional ones verbatim -- U1's designer note, U2's End-only trigger, U8's consume-only-if-it-cancels). Five provisional flags lifted, replaced by ruled/ruled_by/ruled_note so the flag went and the provenance stayed. Register queue 9 -> 0. T02's control came back clean: make sim is 26 passed, 59 rules covered, nothing red. Had a scenario gone red it would have meant we described our own behaviour incorrectly to ground-game. I wrote two U-item mappings and both were wrong. gr-a04 -> U1 (it asserts consent is REQUIRED; U1 asks WHEN the target accepts) and gr-d05 -> U5 (it exercises the UNREJECTED Reverse; U5 is the rejected one). Both plausible from covers:, neither survived reading the description. Third and fourth instance of this defect; the first two reached ground-game. So encodes_u_item is now a declaration and design.py asserts the file names what it claims -- and that check's own first version grepped for mentions and went red when two files recorded why they do NOT encode U1 and U5. A mention is not a claim, which is exactly the looseness that let "six of the ten have provisional scenarios" stand. Two positive controls went red for the best possible reason, both broken the same way -- asserting against live repo data instead of constructing their condition. rule-coverage.py required at least one provisional item to EXIST; it now builds a fixture and reports the live count as a diagnostic, because there is no number of provisional items this project should have. design-baseline.py pinned "2 of 6" while recomputing one row from a live glob, so the dated snapshot was never a snapshot; frozen to its 2026-08-03 list and unwired from self-tests, since per ADR-0012 D8 it is no longer a reporting tool. ScenarioFile is deny_unknown_fields and refused the four new fields until declared -- correct: a corpus accepting unknown metadata would let a typo'd encodes_u_iem sit there claiming nothing. DEVIATION: ADR-0012 D2 said "no new file". GroundRules.md crossed the loadability limit, so the register moved to specs/FindingRegister.md. D2's substance holds -- one register, same machinery, nothing competing -- but the literal instruction did not, and it resolves an awkwardness D2 named itself. make all: exit 0. loop-lint clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
9.2 KiB
CB-EV-0022 — collect the rulings
CB-WP-0026 T05. Tier S (structural S — applies rulings inside an existing capability; chaos d8=6 → no override). Declaration 9 of chaos window 2. Closed 2026-08-05.
Delivered: ten rulings recorded in §Underdetermined, five
provisional: true flags lifted, four scenario schema fields, the
encodes_u_item declaration, and an empty finding queue.
1. How long the answers sat, and what noticed them
Two days, and the thing that noticed was the register's first run — not a human, not the adversarial review.
ground-game ruled all ten U-items on 2026-08-03 (GROUND-WP-0002 T05),
every one confirmed. On the same day CB-RES-0007 reported "0 of 10
ruled." Two days later CB-WP-0022 built the register recording them as
reported, and the adversarial review — which found four other things —
did not catch it either. make design did, on its first execution.
This is the symmetric failure nobody designed for. CB-WP-0022 was shaped end to end around we send findings and nobody reads them: the four-day unread inbox is quoted in the declaration, the survey, the ADR and the spec. The mirror case — they answer and we do not collect it — appears in none of them.
It is arguably the worse of the two. An unread message is visible as silence; a collected-but-unapplied ruling looks exactly like work in progress.
2. Did any confirmed default fail to match the kernel? No.
This was the control that mattered. Every ruling was a confirmation of
what we told ground-game we simulate — so a red scenario would have
meant we described our own behaviour incorrectly to them, a defect in
our report rather than in their ruling, and the most serious class
available since it would be about our own code.
make sim → 26 passed, 59 rules covered
No scenario went red. The five confirmed defaults are the behaviour implemented. That is the strongest single result here and it is a negative: nothing was wrong.
3. The mapping gap was not what the survey said, and I reproduced the defect writing it
CB-RES-0007: "six of the ten already have provisional scenarios."
Measured: one. And the interesting part is how the other nine were lost.
I wrote two mappings from the covers: lists and both were wrong:
| claimed | why it was withdrawn |
|---|---|
gr-a04-bond-support → U1 |
asserts consent is required; U1 asks when the target accepts |
gr-d05-darvo-reverse → U5 |
exercises the unrejected REVERSE; U5 is the rejected one (GROUND—ND) |
Both were plausible from covers:. Neither survived reading the
description. These are the third and fourth instances of this exact
defect — a link that looks right from metadata, asserted without
checking what the artifact exercises — and the first two reached
ground-game.
So the answer to "do the four unlinked scenarios encode U-item defaults at all?" is: at least two of the four do not, and the survey's "six of ten" was not an under-documented truth. It was wrong.
encodes_u_item is now a declaration a scenario makes or omits, and
design.py asserts the file names what it claims.
The check's own first version was the same looseness
Written as grep -lE "\bU<n>\b", it went red the moment two scenarios
recorded why they do not encode U1 and U5 — reporting ['U1','U2','U5'].
A mention is not a claim. That is precisely the imprecision that let "six of the ten have provisional scenarios" stand unchallenged for five days: someone grepped for U-item strings and counted hits. The check now asserts on the declaration.
4. The queue reached 0
QUEUE (open findings) (none)
open, lacking a reproduction 0 target 0
closed (log) 12 [U1..U10, F11, F13]
with a resolving reproduction 3/12 = 25%
notes 2 F12, F14
9 → 0, the number CB-WP-0026 T04 named. First evidence that
ADR-0012 D5's lifecycle is real rather than drawn: findings entered a
state, moved through it, and left the queue.
25% reproduced must not be read as a failure. Nine U-items closed by a ruling, and a ruling is not an artifact. The metric is now honest about something the survey's 33% concealed: most of our findings close because someone answered them, not because anything demonstrates them.
What the lifecycle could not express — the honest gap: applied is
defined as the source changed and the provisional default was deleted.
Here the rulings confirmed our defaults, so nothing in the rules moved;
what changed is that the flags came off. The state fits, but the
definition had to be read generously. If a future ruling overturns a
default, applied will mean something materially different from what it
meant today, and D5 does not distinguish them.
5. Two things the schema caught
ScenarioFile is deny_unknown_fields, so five scenarios failed to
parse until ruled, ruled_by, ruled_note and encodes_u_item were
declared in the Rust struct. A corpus that accepted unknown metadata would
let a typo'd encodes_u_iem sit forever claiming nothing — and this
pass's whole subject is claims nobody checks.
record.rs had to set them explicitly. A recorded game is evidence of
what happened, not a claim about an undetermined rule. Filling the fields
via ..Default::default() would have been shorter and would let a
recording silently inherit a U-item claim, pointing a reproduction at a
finding it has nothing to do with.
5b. Two gates went red for the best possible reason
Lifting the last five provisional flags broke two positive controls, and both were broken in the same way: they asserted against live repo data instead of constructing the condition they test.
rule-coverage.py required bool(prov) — at least one provisional
item must exist. That guard was the right instinct (a control that passes
vacuously is worthless) wired the wrong way. With nothing provisional, it
went red. It now builds a fixture, asserts the missing-owner case is
caught, and reports the live count as a diagnostic — because there is
no number of provisional items this project should have.
design-baseline.py pinned "the measured baseline is 2 of 6" and
reported 1/6. Its "six provisional defaults" row globbed
provisional: true at run time, so the dated snapshot was never a
snapshot — it drifted with the repo. C2 dismantled this tool four hours
earlier and missed this: a hand-maintained dict with one dynamically
computed row is worse than a fully hand-maintained one, because the
recomputed row silently disagrees with the date in the header.
Frozen to the literal list it measured on 2026-08-03, and removed from
make self-tests — that target is "a positive control for every
reporting tool", and per ADR-0012 D8 this is no longer a reporting tool.
Leaving it wired in meant a superseded instrument could fail the build.
The pattern across both: a control that reads the world it is meant to audit will eventually audit a world that has moved. Neither was caught by review; both were caught by the world moving.
6. The register moved, and ADR-0012 D2 gave way to loadability
specs/GroundRules.md crossed the ~400-line limit and loop-lint
required a split. ADR-0012 D2 said "no new file", so this is a
deviation and is recorded as one.
D2's substance holds. Its argument was one register, not a second
mechanism competing with the first — and specs/FindingRegister.md is
that same register, moved, still driving off the provisional/ruling
machinery, still the only one. What was traded away is the literal "no new
file", which was D2's implementation, not its reason.
It also resolves an awkwardness D2 named itself: "a finding about the engine's behaviour sits in a document about the game." Now it does not.
7. Chaos window 2
Declaration 9 of 12. Structural S, d8 = 6, no override. Recorded per §Loop tiers even though it changed nothing.
Window 2 still has no override to evaluate — nine declarations, zero 8s. The retirement condition (retire if an override changes nothing twice running) cannot be assessed, and at d8 the expected count over twelve declarations is 1.5, so this is unremarkable rather than evidence of anything.
8. Cost
CB-WP-0022's cost, by re-running the instrument:
CB-WP-0022-T05 $ 6.43
CB-WP-0022-T01 $ 4.17
CB-WP-0022-T02 $ 3.56 (the separate reviewer's own spend is in the subagent tree)
CB-WP-0022-T04 $ 1.92
CB-WP-0022-T03 $ 1.01
The chain did not snap this time — CB-EV-0019 §4 predicted it might,
having found cb-cost.py --slug CB-WP-0020 aborts for want of retained
transcripts. CB-WP-0022 is recent enough to still be in the window. The
bound CB-EV-0019 asked for is still owed; this pass is evidence that the
rule works for a pass one step back, not that it works generally.
Open after this pass
- Nine U-items are
appliedwith no reproduction. That is recorded, not hidden, but it means nine rules rest on a ruling nobody can re-run. Cheap to fix incrementally: each needs one scenario naming its item. appliedconflates confirmed with overturned (§4). It will matter the first time a ruling goes against us.- The cost-chain bound (CB-EV-0019 §4) is still unwritten.