clay-borg/evidence/CB-EV-0022-collect-the-rulings.md
tegwick 6be9fbc9af
Some checks failed
ci / check (push) Failing after 3s
CB-WP-0026: collect the rulings -- ten answers that arrived and were never applied
ground-game ruled all ten U-items on 2026-08-03, every one CONFIRMED as
the default clay-borg simulates, and confirmed five of six provisional
scenarios. clay-borg never collected the answers: CB-RES-0007 reported "0
of 10 ruled" the same day, and CB-WP-0022 built the finding register two
days later still recording them as `reported`. make design's first run is
what noticed -- not a human, not the adversarial review that found four
other things.

That is the unread-inbox failure running in the opposite direction, and it
appears nowhere in the declaration, survey, ADR or spec of the pass that
was built entirely around the forward version. It is arguably worse: an
unread message is visible as silence, while a collected-but-unapplied
ruling looks exactly like work in progress.

Ten rulings quoted into §Underdetermined (the three conditional ones
verbatim -- U1's designer note, U2's End-only trigger, U8's
consume-only-if-it-cancels). Five provisional flags lifted, replaced by
ruled/ruled_by/ruled_note so the flag went and the provenance stayed.
Register queue 9 -> 0.

T02's control came back clean: make sim is 26 passed, 59 rules covered,
nothing red. Had a scenario gone red it would have meant we described our
own behaviour incorrectly to ground-game.

I wrote two U-item mappings and both were wrong. gr-a04 -> U1 (it asserts
consent is REQUIRED; U1 asks WHEN the target accepts) and gr-d05 -> U5 (it
exercises the UNREJECTED Reverse; U5 is the rejected one). Both plausible
from covers:, neither survived reading the description. Third and fourth
instance of this defect; the first two reached ground-game. So encodes_u_item
is now a declaration and design.py asserts the file names what it claims --
and that check's own first version grepped for mentions and went red when
two files recorded why they do NOT encode U1 and U5. A mention is not a
claim, which is exactly the looseness that let "six of the ten have
provisional scenarios" stand.

Two positive controls went red for the best possible reason, both broken
the same way -- asserting against live repo data instead of constructing
their condition. rule-coverage.py required at least one provisional item
to EXIST; it now builds a fixture and reports the live count as a
diagnostic, because there is no number of provisional items this project
should have. design-baseline.py pinned "2 of 6" while recomputing one row
from a live glob, so the dated snapshot was never a snapshot; frozen to
its 2026-08-03 list and unwired from self-tests, since per ADR-0012 D8 it
is no longer a reporting tool.

ScenarioFile is deny_unknown_fields and refused the four new fields until
declared -- correct: a corpus accepting unknown metadata would let a typo'd
encodes_u_iem sit there claiming nothing.

DEVIATION: ADR-0012 D2 said "no new file". GroundRules.md crossed the
loadability limit, so the register moved to specs/FindingRegister.md. D2's
substance holds -- one register, same machinery, nothing competing -- but
the literal instruction did not, and it resolves an awkwardness D2 named
itself.

make all: exit 0. loop-lint clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 16:13:37 +02:00

209 lines
9.2 KiB
Markdown

# CB-EV-0022 — collect the rulings
CB-WP-0026 T05. Tier S (structural S — applies rulings inside an existing
capability; chaos d8=6 → no override). Declaration 9 of chaos window 2.
Closed 2026-08-05.
**Delivered:** ten rulings recorded in `§Underdetermined`, five
`provisional: true` flags lifted, four scenario schema fields, the
`encodes_u_item` declaration, and an empty finding queue.
---
## 1. How long the answers sat, and what noticed them
**Two days**, and the thing that noticed was **the register's first run**
not a human, not the adversarial review.
`ground-game` ruled all ten U-items on **2026-08-03** (GROUND-WP-0002 T05),
every one confirmed. On the same day CB-RES-0007 reported *"0 of 10
ruled."* Two days later CB-WP-0022 built the register recording them as
`reported`, and the adversarial review — which found four other things —
did not catch it either. `make design` did, on its first execution.
**This is the symmetric failure nobody designed for.** CB-WP-0022 was
shaped end to end around *we send findings and nobody reads them*: the
four-day unread inbox is quoted in the declaration, the survey, the ADR
and the spec. The mirror case — *they answer and we do not collect it*
appears in none of them.
It is arguably the worse of the two. An unread message is visible as
silence; a collected-but-unapplied ruling looks exactly like work in
progress.
## 2. Did any confirmed default fail to match the kernel? No.
This was the control that mattered. Every ruling was a *confirmation* of
what we told `ground-game` we simulate — so a red scenario would have
meant **we described our own behaviour incorrectly to them**, a defect in
our report rather than in their ruling, and the most serious class
available since it would be about our own code.
```
make sim → 26 passed, 59 rules covered
```
**No scenario went red.** The five confirmed defaults are the behaviour
implemented. That is the strongest single result here and it is a
negative: nothing was wrong.
## 3. The mapping gap was not what the survey said, and I reproduced the defect writing it
CB-RES-0007: *"six of the ten already have provisional scenarios."*
**Measured: one.** And the interesting part is how the other nine were
lost.
I wrote two mappings from the `covers:` lists and **both were wrong**:
| claimed | why it was withdrawn |
|---|---|
| `gr-a04-bond-support` → U1 | asserts consent is **required**; U1 asks **when** the target accepts |
| `gr-d05-darvo-reverse` → U5 | exercises the **unrejected** REVERSE; U5 is the **rejected** one (GROUND—ND) |
Both were plausible from `covers:`. Neither survived reading the
description. **These are the third and fourth instances of this exact
defect** — a link that looks right from metadata, asserted without
checking what the artifact exercises — and the first two reached
`ground-game`.
So the answer to *"do the four unlinked scenarios encode U-item defaults
at all?"* is: **at least two of the four do not**, and the survey's "six of
ten" was not an under-documented truth. It was wrong.
`encodes_u_item` is now a declaration a scenario makes or omits, and
`design.py` asserts the file names what it claims.
### The check's own first version was the same looseness
Written as `grep -lE "\bU<n>\b"`, it went red the moment two scenarios
recorded *why they do not* encode U1 and U5 — reporting `['U1','U2','U5']`.
**A mention is not a claim.** That is precisely the imprecision that let
*"six of the ten have provisional scenarios"* stand unchallenged for five
days: someone grepped for U-item strings and counted hits. The check now
asserts on the declaration.
## 4. The queue reached 0
```
QUEUE (open findings) (none)
open, lacking a reproduction 0 target 0
closed (log) 12 [U1..U10, F11, F13]
with a resolving reproduction 3/12 = 25%
notes 2 F12, F14
```
**`9 → 0`**, the number CB-WP-0026 T04 named. First evidence that
ADR-0012 D5's lifecycle is real rather than drawn: findings entered a
state, moved through it, and left the queue.
**25% reproduced must not be read as a failure.** Nine U-items closed by a
**ruling**, and a ruling is not an artifact. The metric is now honest
about something the survey's 33% concealed: most of our findings close
because someone answered them, not because anything demonstrates them.
**What the lifecycle could not express** — the honest gap: `applied` is
defined as *the source changed and the provisional default was deleted*.
Here the rulings **confirmed** our defaults, so nothing in the rules moved;
what changed is that the flags came off. The state fits, but the
definition had to be read generously. If a future ruling *overturns* a
default, `applied` will mean something materially different from what it
meant today, and D5 does not distinguish them.
## 5. Two things the schema caught
**`ScenarioFile` is `deny_unknown_fields`**, so five scenarios failed to
parse until `ruled`, `ruled_by`, `ruled_note` and `encodes_u_item` were
declared in the Rust struct. A corpus that accepted unknown metadata would
let a typo'd `encodes_u_iem` sit forever claiming nothing — and this
pass's whole subject is claims nobody checks.
**`record.rs` had to set them explicitly.** A recorded game is evidence of
what happened, not a claim about an undetermined rule. Filling the fields
via `..Default::default()` would have been shorter and would let a
recording silently inherit a U-item claim, pointing a reproduction at a
finding it has nothing to do with.
## 5b. Two gates went red for the best possible reason
Lifting the last five provisional flags broke two positive controls, and
**both were broken in the same way**: they asserted against live repo data
instead of constructing the condition they test.
**`rule-coverage.py`** required `bool(prov)` — *at least one provisional
item must exist*. That guard was the right instinct (a control that passes
vacuously is worthless) wired the wrong way. With nothing provisional, it
went red. It now builds a fixture, asserts the missing-owner case is
**caught**, and reports the live count as a diagnostic — because **there is
no number of provisional items this project should have.**
**`design-baseline.py`** pinned *"the measured baseline is 2 of 6"* and
reported `1/6`. Its "six provisional defaults" row **globbed
`provisional: true` at run time**, so the dated snapshot was never a
snapshot — it drifted with the repo. C2 dismantled this tool four hours
earlier and missed this: a hand-maintained dict with one dynamically
computed row is worse than a fully hand-maintained one, because the
recomputed row silently disagrees with the date in the header.
Frozen to the literal list it measured on 2026-08-03, and **removed from
`make self-tests`** — that target is *"a positive control for every
reporting tool"*, and per ADR-0012 D8 this is no longer a reporting tool.
Leaving it wired in meant a superseded instrument could fail the build.
**The pattern across both**: a control that reads the world it is meant to
audit will eventually audit a world that has moved. Neither was caught by
review; both were caught by the world moving.
## 6. The register moved, and ADR-0012 D2 gave way to loadability
`specs/GroundRules.md` crossed the ~400-line limit and `loop-lint`
required a split. ADR-0012 D2 said **"no new file"**, so this is a
deviation and is recorded as one.
**D2's substance holds.** Its argument was *one register, not a second
mechanism competing with the first* — and `specs/FindingRegister.md` is
that same register, moved, still driving off the `provisional`/ruling
machinery, still the only one. What was traded away is the literal "no new
file", which was D2's *implementation*, not its reason.
It also resolves an awkwardness D2 named itself: *"a finding about the
engine's behaviour sits in a document about the game."* Now it does not.
## 7. Chaos window 2
**Declaration 9 of 12.** Structural S, d8 = 6, no override. Recorded per
§Loop tiers even though it changed nothing.
**Window 2 still has no override to evaluate** — nine declarations, zero
8s. The retirement condition (*retire if an override changes nothing twice
running*) cannot be assessed, and at d8 the expected count over twelve
declarations is 1.5, so this is unremarkable rather than evidence of
anything.
## 8. Cost
CB-WP-0022's cost, by re-running the instrument:
```
CB-WP-0022-T05 $ 6.43
CB-WP-0022-T01 $ 4.17
CB-WP-0022-T02 $ 3.56 (the separate reviewer's own spend is in the subagent tree)
CB-WP-0022-T04 $ 1.92
CB-WP-0022-T03 $ 1.01
```
**The chain did not snap this time** — CB-EV-0019 §4 predicted it might,
having found `cb-cost.py --slug CB-WP-0020` aborts for want of retained
transcripts. CB-WP-0022 is recent enough to still be in the window. **The
bound CB-EV-0019 asked for is still owed**; this pass is evidence that the
rule works for a pass one step back, not that it works generally.
## Open after this pass
- **Nine U-items are `applied` with no reproduction.** That is recorded,
not hidden, but it means nine rules rest on a ruling nobody can re-run.
Cheap to fix incrementally: each needs one scenario naming its item.
- **`applied` conflates *confirmed* with *overturned*** (§4). It will
matter the first time a ruling goes against us.
- **The cost-chain bound** (CB-EV-0019 §4) is still unwritten.