clay-borg/workplans/CB-WP-0026-collect-the-rulings.md
tegwick 5f9c2839bb
Some checks failed
ci / check (push) Failing after 4s
Sync hub state for CB-WP-0026
Workplan and task ids written back by fix-consistency C-06; WORK-RECORDS
regenerated.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 16:15:22 +02:00

292 lines
12 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
id: CB-WP-0026
kind: product
title: "Collect the rulings: ten answers that arrived and were never applied"
status: done
state_hub_workstream_id: "289bcc55-0c1b-4e10-a40b-cb7ecdab46be"
---
# Purpose
```
structural tier S (applies rulings inside an existing capability —
scenario data and a spec section, no new port, no
canonical interface, no dependency)
chaos d8 = 6 → no override
declared tier S
```
Declaration 9 of chaos window 2.
## The register's first run found this, which is the only reason it exists
`ground-game` **ruled all ten U-items on 2026-08-03** (GROUND-WP-0002 T05)
and confirmed five of six provisional scenarios (T03). Every U-item was
**confirmed** as the default clay-borg already simulates. The workplan is
`finished`.
**clay-borg never collected the answers.** CB-RES-0007 reported *"0 of 10
ruled"* on 2026-08-03 — the same day — and CB-WP-0022 built the finding
register two days later still recording them as `reported`. It took
`make design`'s first run to notice.
**This is the unread-inbox failure running in the opposite direction.**
The whole design-instrument pass was shaped around *we send findings and
nobody reads them*. The symmetric failure — *they answer and we do not
collect it* — was not anticipated by the declaration, the survey, or the
adversarial review.
**Nothing here is a rules decision.** Every ruling is already made. This is
clay-borg applying them, which is the `ruled``applied` transition
ADR-0012 D5 defines: *the source changed and the provisional default was
deleted.*
## Why the rulings being "confirm" makes this more subtle, not less
Every U-item was confirmed, so **no kernel behaviour changes**. It is
tempting to conclude nothing needs doing.
That is wrong, and the register is what makes it visible: the scenarios
still carry `provisional: true` for choices that are now **settled**. A
provisional flag means *we picked this and the owner has not ruled*. After
2026-08-03 that statement is false in five files, and `make coverage`
keeps ageing them toward a 30-day warning for a question that has been
answered.
**A flag that outlives its question is the same defect as a hash that
outlives its content** — ADR-0011's *"a lie with a timestamp."*
## Task: record the rulings where the defaults live
```task
id: CB-WP-0026-T01
status: done
priority: high
state_hub_task_id: "ef4defb1-5d89-4d33-b5c1-6d7f6a12fb6b"
```
`specs/GroundRules.md §Underdetermined` states ten defaults as *"must be
confirmed or corrected by ground-game."* All ten now are. Each U-item
gains its ruling and the date, from GROUND-WP-0002 T05.
**Quote the ruling, do not paraphrase it.** Three carry notes that are
more than "confirm" and a paraphrase would lose them:
- **U1** — the Bond formed by *this* Support still does not cancel DARVO
(a printed designer note, and it is the interaction GR-L05 encodes);
- **U2** — clamp on every application, but **DARVO still arms only at End**
when at Stress 5 (GR-R08). The clamping rule and the trigger rule are
separate, and conflating them is the obvious misreading;
- **U8** — GROUND—OU applies first and **Protection is consumed only if it
actually cancels**.
**Controls:**
- the §Underdetermined preamble no longer says the items await a ruling —
it is the sentence that will otherwise be true-looking and wrong;
- **U2, U6 and U8 are the ones ground-game flagged as changing play.**
Each must be checkable against the kernel behaviour it describes, or the
record says a rule is implemented that may not be.
**Done 2026-08-05.** All ten rulings quoted into `§Underdetermined`, with
the three conditional ones (U1's designer note, U2's End-only trigger,
U8's consume-only-if-it-cancels) verbatim. The section heading and
preamble no longer say the items await a ruling.
## Task: lift the provisional flags the rulings settle
```task
id: CB-WP-0026-T02
status: done
priority: high
state_hub_task_id: "ab08dc38-46c0-460e-8042-7600542d3908"
```
Five scenarios carry `provisional: true` with `provisional_owner:
ground-game`: `gr-a04-bond-support`, `gr-d01-darvo-trigger`,
`gr-d05-darvo-reverse`, `gr-e02-shared-ground`, `gr-e04-coalitions`. All
five were **confirmed** on 2026-08-03. The flag comes off, and each
records which ruling settled it.
**Do not delete the provenance with the flag.** CB-WP-0021 T03 is the
precedent: `gr-e01` was *rewritten, not retired*, because a reader needs
to learn the question was answered rather than that a tag vanished.
**Controls:**
- `make coverage` reports **0 provisional items** afterwards, and that is
the measurable outcome of this task;
- `make sim` still passes — a confirmed default is the behaviour already
implemented, so a red scenario here means the default was **not** what
we told ground-game it was, which is a finding and must be raised, not
patched;
- the ruling date survives in each file.
**Done 2026-08-05.** All five flags lifted; `make coverage` reports **0
provisional items**. Each file gained `ruled` / `ruled_by` / `ruled_note`
instead — the flag came off, the provenance did not.
**The schema refused the new fields and was right to.** `ScenarioFile` is
`deny_unknown_fields`, so five scenarios failed to parse until `ruled`,
`ruled_by`, `ruled_note` and `encodes_u_item` were declared. A scenario
corpus that silently accepted unknown metadata would let a typo'd
`encodes_u_iem` sit there claiming nothing.
**T02's control held: `make sim` is 26 passed, 59 rules covered, with no
scenario red.** That is the answer to the serious version of this task —
had one gone red, it would have meant we described our own behaviour
incorrectly to ground-game, which would be a defect in *our* report rather
than in their ruling.
## Task: make the U-item ↔ scenario mapping checkable
```task
id: CB-WP-0026-T03
status: done
priority: high
state_hub_task_id: "2a4b9ef4-707a-44d1-a61a-0d1abbb8a321"
```
**Only U2 names its U-item in a scenario.** Measured in CB-WP-0022 T05:
```
for u in U1..U10; do grep -lE "\b$u\b" scenarios/ground/*.yaml; done
```
CB-RES-0007 asserted *"six of the ten already have provisional
scenarios."* Four probably do — but **the link is not written down, so it
is not checkable**, and an uncheckable link asserted as a fact is the
defect class that produced three wrong premises.
Each scenario that encodes a U-item default names it. Where no scenario
covers a U-item, **say so** rather than inventing a mapping — that is the
honest half, and guessing here would reproduce exactly the error being
fixed.
**Controls:**
- a check that every U-item claiming a reproduction resolves to a file
that **names it**`design.py`'s self-test already asserts the current
state (`cited == ["U2"]`) and must be updated to the new truth, not
deleted;
- the register's reproduction column matches the mapping, since a register
that disagrees with the files is worse than none.
**Done 2026-08-05, and it caught me writing the same defect a third and
fourth time.**
I wrote two mappings from the `covers:` lists — `gr-a04` → U1, `gr-d05`
U5 — and both are wrong on reading what the scenarios actually exercise:
| claimed | why withdrawn |
|---|---|
| `gr-a04` → U1 | it asserts consent is **required**; U1 asks **when** the target accepts |
| `gr-d05` → U5 | it exercises the **unrejected** REVERSE; U5 is the **rejected** one (GROUND—ND) |
Both were plausible from `covers:` and neither survived reading the
description. **The measurement stands: one U-item, U2.** Nine have no
reproduction and are recorded as having none.
`encodes_u_item` is now a declaration, and `design.py` asserts a scenario
naming a U-item actually contains it. **The check's first version was
itself the loose proxy**: it grepped for `\bU<n>\b` and went red when two
files recorded *why they do not* encode U1 and U5. A mention is not a
claim — and that is exactly the looseness that let *"six of the ten have
provisional scenarios"* stand unchallenged.
## Task: close the register, and see whether it closes
```task
id: CB-WP-0026-T04
status: done
priority: high
state_hub_task_id: "335457ff-61b4-4991-a787-a7eec55ac899"
```
Move U1U10 from `ruled` to `applied` and re-run `make design`.
**The number to beat is `open, lacking a reproduction: 9 → 0`**, and the
queue emptying is the first evidence that ADR-0012 D5's lifecycle is real
rather than drawn.
**If it does not reach 0, do not adjust the metric.** Whatever remains is
either a U-item with no scenario (T03's honest half, which should then be
a `note`, not a finding) or a state the lifecycle cannot express — and the
second is a finding about ADR-0012 that belongs in evidence.
**Controls:**
- `make design` shows an empty queue for the U-items, with the log
carrying all ten;
- `make self-tests` and `make sim` pass;
- **the closure is checkable from the register alone** — a reader should
not need this workplan to see that ten findings closed on 2026-08-03 and
were applied on 2026-08-05.
**Done 2026-08-05. The queue is empty.**
```
QUEUE (open findings) (none)
open, lacking a reproduction 0 target 0
closed (log) 12 [U1..U10, F11, F13]
with a resolving reproduction 3/12 = 25%
```
**`open, lacking a reproduction: 9 → 0`** — the number this task named.
It is the first evidence that ADR-0012 D5's lifecycle is real rather than
drawn: findings entered a state, moved through it, and left the queue.
**25% reproduced is not a failure and must not be read as one.** Nine
U-items closed by a *ruling*, and a ruling is not an artifact. The metric
is now honest about a thing the survey's 33% concealed: most of our
findings closed because someone answered them, not because anything
demonstrated them.
## Task: evidence
```task
id: CB-WP-0026-T05
status: done
priority: medium
state_hub_task_id: "28197e69-1326-4181-b673-61b3a20fb720"
```
`evidence/CB-EV-0022-*.md`. Short — this is an S pass.
- **How long the answers sat uncollected**, and what noticed them. Two
days, and the thing that noticed was the register's first run, not a
human and not the review.
- **Whether any confirmed default turned out not to match the kernel** —
T02's control. A red scenario here would mean we described our own
behaviour incorrectly to ground-game, which would be a fourth premise
defect and the most serious kind, since it was *our* behaviour.
- **What the mapping gap really was**: whether the four unlinked scenarios
encode U-item defaults at all, or whether CB-RES-0007's "six of ten" was
simply wrong.
- **Whether the queue reached 0**, and if not, what the lifecycle could not
express.
- **Quote CB-WP-0022's cost by re-running the instrument** — and note that
CB-EV-0019 §4 found this chain snaps at roughly four passes, so if the
instrument aborts, that is the second instance and the rule needs its
bound.
**Done 2026-08-05.**
[CB-EV-0022](../evidence/CB-EV-0022-collect-the-rulings.md).
- **Two days uncollected, and the register's first run is what noticed** —
not a human, not the adversarial review that found four other things.
- **No confirmed default failed to match the kernel**: `make sim` 26
passed, 59 rules covered, nothing red. The serious version of this task
came back clean.
- **The mapping gap was not under-documentation, it was error.** At least
two of the four unlinked scenarios do **not** encode the U-items the
survey implied, and I wrote both wrong mappings before checking them.
- **The queue reached 0.** `applied` conflates *confirmed* with
*overturned*, which is the lifecycle gap this pass found and did not fix.
- **The cost chain held this time** — CB-WP-0022 is one pass back and
still in the window. The bound CB-EV-0019 asked for is still owed.
**A gate went red for the best possible reason.** `rule-coverage.py`'s
positive control required `bool(prov)` — at least one provisional item to
exist. Lifting the last five turned it red. The guard was the right
instinct (a vacuous pass is worthless) wired the wrong way: it bound a
control to live repo data. It now builds its own fixture, asserts the
missing-owner case is *caught*, and reports the live count as a diagnostic
— because there is no number of provisional items this project should
have.