clay-borg/workplans/CB-WP-0005-assertion-coverage.md
tegwick e2a2957b3a
Some checks failed
ci / check (push) Failing after 3s
Fix loadability, and put the deferred analysis where the work is
loop-lint failed: CB-WP-0005 was 401 lines against a ~400 limit, one over,
after the status-vocabulary note. Pushed before checking — my error; the
gate caught it on the next run.

The fix is not a trim. The gate exposed a circular reference I had
created: CB-WP-0005's cancelled tasks held the full analysis while
CB-WP-0006 pointed back at them for detail, so the live workplan deferred
to a cancelled one. The analysis now lives in CB-WP-0006 Phase B next to
the work, and CB-WP-0005 keeps a forward pointer per task. One copy,
single source of fact, and the reference points forward.

CB-WP-0006 T05 and T06 gain the detail that moved: K9's mutation proof and
what its acceptance property actually is, K11's detection clause and
budget attribution, and the D2 correction the reviewer forced — the bundle
is not "a directory of four files" but a change to the runner's data flow,
because scenario.rs creates an EventLog, appends to it and never reads it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 17:58:36 +02:00

345 lines
15 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
id: CB-WP-0005
title: "Make the instruments count assertions, then fix what they expose"
status: proposed
state_hub_workstream_id: "0b95a1e3-7780-43d0-81e9-072ef7978734"
---
# Purpose
`research/CB-RES-0004-replay-and-kernel-coverage.md` (v2, after an
adversarial review that rejected v1 with four blocking findings):
> **Every coverage instrument in this project counts names. None counts
> assertions.** Four of the seven defects found this pass are named in the
> source and inert, so no name-based check finds them.
Two are mutation-proven: AM-7's `hash-identical` clause survives folding a
100k-event log from an unrelated genesis state, and K9's `through` field
survives `Snapshot::take` discarding it entirely. `evidence/CB-EV-0001`
carries **`AM-7 replay | met, 2,290×`** for the first of those.
[ADR-0005](../decisions/ADR-0005-assertion-coverage-and-replay.md) decides
the response. This workplan executes it in four phases: teach the
instruments to count assertions, correct the record they falsified,
implement what they expose, then measure whether any of it worked.
**Headline numbers are expected to get worse before they get better.** A
pass that widened every denominator and reported the same percentages
would have proved the instruments still count names. Per InnerLoop §Step 4
no target moves in the commit that measures it.
## Phase A — teach the instruments to count assertions
## Task: spec→code link over every numbered spec and every crate
```task
id: CB-WP-0005-T01
status: done
priority: high
state_hub_task_id: "d47338c6-eb2e-4443-aad2-0b08c2d91895"
```
`tools/rule-coverage.py` hardcodes `AGGREGATE = "games/ground/src/lib.rs"`
and `RULE_RE = \*\*(GR-[A-Z]+\d+)` against `GroundRules.md` alone. The
reviewer's finding: a ~10-line generalization would have surfaced K10, K14
and K18 the day AM-1b shipped.
Deliver: `AGGREGATE` becomes a list of source roots; rule patterns become
per-spec; the link runs over **every numbered spec × every crate**.
Three constraints from ADR-0005 §5, all of which must be visible in the
output:
1. The kernel denominator **reports and does not feed the exit code until
2026-08-31**, after which it binds. The date lives in the tool and the
tool prints the days remaining — an open-ended "gate it later" is how
AM-4's targets went unratified for four workplans.
2. **Replicate the zero-rules positive control** on every new denominator.
The existing arm refuses to report over zero rules, a defect it was
fixed for; a kernel regex matching nothing must abort, not print
`0/0 (100%)`.
3. State the limit in the output, as the GR arm already does: this counts
names. It is the cheap half.
**Predicted:** K10, K14, K18 reported unlinked on the first run.
**Refuted if** any numbered rule in any spec is still unnamed in source
after the pass and the tool does not say so.
**Delivered, and the prediction held on the first run:**
```text
AM-1b kernel spec->code link: 15/18 (83%) K-rules named across 10 source files
NOTE: link only — K-rules are kernel invariants with no scenario
mechanism; and this counts names, not assertions
gate: reporting only for 31 more day(s), binds 2026-08-31 (ADR-0005 §5)
unlinked (declared in the spec, named nowhere in source):
K10 K14 K18
```
`AGGREGATE` is now a list of source roots, rule patterns are per-spec, and
the link runs over every numbered spec × every crate. The binding date
lives in the tool and the days remaining are printed every run. The
kernel figures (`k_rules`, `k_linked`, `k_unlinked`) are registered facts,
so they are under `make facts-check` from the day they first exist rather
than after they drift.
**The self-test passed while the tool was completely broken.** A
`print(` inside `say()` was rewritten to `say(`, so every real
`make coverage` died with `RecursionError` — while `--self-test` reported
all-ok, because it only ever called `kernel_arm(quiet=True)` and never
executed the reporting path.
That is this task's own thesis in miniature: **the control named the
behaviour and did not assert it.** Fixed by exercising the loud path under
`redirect_stdout` and asserting it prints, and verified by re-breaking
`say()` and confirming the two new checks go red. Seventh instance of the
harness-does-nothing shape, found in the tool written to find that shape.
**A limit of `facts-check` surfaced here and is recorded, not patched:**
the check is line-based, so a tagged value that prose-wraps onto the next
line fails. It cost three edits to place two tags. Reported for T07 —
either the checker spans a paragraph, or the rule is stated as
"tagged values must not wrap".
## Task: M-D1-MUT — one mutation per acceptance row
```task
id: CB-WP-0005-T02
status: done
priority: high
state_hub_task_id: "c88f2696-ca12-44df-add3-0db3d2819c08"
```
The task this workplan exists for. Name-based checks catch 3 of the 7
defects; this catches the other 4.
Deliver `make mutation-check`: for each of the twelve AM-* acceptance
rows, a committed fixture that **inverts the row's stated property** and
an assertion that the suite goes **red**. A row whose mutation leaves the
suite green is a row backed by nothing.
Per ADR-0005 §1, `adapted:mutation-testing` — the denominator is
acceptance rows, not source lines, because the failure mode here is not an
untested branch but a headline number backed by nothing.
**The escape hatch is closed in advance:** a row for which no mutation can
be written is recorded `unmutatable` **with the reason** and **counts
against** the metric. A row nobody can invert asserts nothing.
Carries `--self-test`. Its own positive control is the one that matters:
a mutation harness that fails to apply its mutation reports every row as
`unmutatable` and looks thorough — this project has six harness-does-
nothing instances and one of them is being fixed in T03.
**Predicted: 9 of 12 rows turn red.** Deliberately beatable in both
directions — **12 of 12 refutes CB-RES-0004** (the instruments were better
than claimed, the finding shrinks to three absent rules, and M-D1-MUT was
not worth its CI cost); **3 of 12 means the pass is under-scoped and must
stop and re-plan** rather than proceed to Phase C.
**Delivered. Measured: 4 of 14 rows enforced. The prediction is badly
unmet.**
```text
M-D1-MUT: 4/14 rows enforced
PARTIAL 2 (AM-7, AM-8 — some clauses live, some inert)
unmutatable 8 (no property to invert, reason stated per row)
SURVIVED 0
```
**First correction: there are 14 acceptance rows, not 12.** ADR-0005 and
this workplan both said twelve; AM-4 splits into a/b/c. The 9-of-12 (75%)
prediction is evaluated as ≥10 of 14 on the same basis. Measured 4 (29%).
**Second correction, and the one that matters: my first run reported two
SURVIVED rows, and both were my own bad mutations.**
- AM-8: `pub struct NullRng;``pub struct NullRng {}` — semantically
identical, a no-op.
- AM-12: renaming `max_age_days` in the price sheet — CA-17 reads it with
`.get("max_age_days", 90)`, so removing it changes nothing.
Both would have been published as *"this row asserts nothing"* — a false
accusation against code that is in fact fine. Replaced with real
inversions (inject a per-construction counter into the ChaCha seed;
revert the AC-9 output resolution to the first-wins bug it was fixed for),
after which both go red. **T08's question — "is writing a weak mutation
the new grep?" — is answered on the first attempt: yes, demonstrably.**
**The finding is larger than the workplan assumed.** 8 of 14 rows are
`unmutatable`: AM-2, AM-3, AM-4c, AM-5, AM-6, AM-9, AM-10, AM-11 have no
instrument at all. AM-6 is the sharpest — **nothing in the workspace
compares any number to 100,000 events/s**, the headline throughput claim.
The problem is not three unimplemented rules; it is that **more than half
the acceptance table has nothing behind it.**
Harness controls that earned their place: a stale find-string reports
`HARNESS-BROKEN` rather than silently scoring the baseline as the mutant;
a red baseline reports `inconclusive` rather than `red`; the tree is
restored in a `finally` and the restoration is verified.
`make mutation-check` is deliberately **not** in `make all` — it rebuilds
the workspace once per mutated row. `--self-test` is in `make self-tests`.
**Stop condition: see the note in T07 and the decision recorded there.**
## Phase B — correct the record
## Task: correct three committed verdicts and restore the fourth
```task
id: CB-WP-0005-T03
status: done
priority: high
state_hub_task_id: "87fa2245-5f27-48af-906e-18affa41efcf"
```
Per ADR-0005 §4. Corrections are made **in place with a dated note**, not
silently edited — the correction trail is the artifact.
| row | committed | corrected to |
|---|---|---|
| AM-7 replay | `met, 2,290×` | timing clause **met**; `hash-identical` clause **withdrawn**, mutation-proven unmeasured, and the fold spans multiple game seeds so it is not a replay |
| AM-10 foreign types | `met` | **withdrawn as written** — no `cb-*-api` crate, so the population is empty; re-stated as a K6 determinism row, M-D4-LEAK marked **unmeasured** |
| AM-11 impl pairs | `met, narrow` | **unmet** until a shared conformance suite exists; the pair exists, the suite does not, and the metric is a bool over the suite |
| AM-1b | *absent from the scoreboard* | **added**, at its real value |
The last one is the point of the phase. `make coverage` prints `49/58` two
lines below the `100%` the scoreboard carried, and the scoreboard kept the
flattering half. Update `specs/GameKernel.md` §5 and
`specs/MetricsAndScenarios.md` §1 to match, and tag the new figures under
`make facts-check` so they cannot drift.
**Delivered.** `evidence/CB-EV-0001` §1 now carries a dated correction
note and a new **Enforced** column carrying M-D1-MUT — because a row can
be *measured* and still enforce nothing, and the scoreboard had no way to
say so. §1a records what changed and how each was found.
**A fifth correction surfaced that ADR-0005 did not list:** AM-12 still
read **$248.46**, the figure CB-WP-0002 disproved and corrected to
$93.15. It had been stale in the evidence file ever since — **untagged,
and therefore invisible to `facts-check`**. It is now tagged
`fact:pinned_total`. A DFD instance that survived the gate built to catch
DFD, because that gate only checks copies that opted in.
`specs/GameKernel.md` §5 carries the AM-10 withdrawal and the AM-11
downgrade inline, so a reader of the spec cannot reach the old claim.
## Phase C — DEFERRED before starting (see T02)
> **Deferred 2026-07-31 by the stop condition in T02, maintainer decision.**
> M-D1-MUT measured **4 of 14** (28.6%) against a 25% floor and a 71%
> prediction. Phase C is scoped to K9/K10/K11/K14/K18 — five rules — but
> the measurement says **8 of 14 acceptance rows have no instrument at
> all**, which Phase C does not address. Building it as written would mean
> proceeding on a diagnosis the instrument had just contradicted.
>
> T04, T05 and T06 stay in this file, unstarted, with their analysis
> intact; they move to **CB-WP-0006**, which is scoped to the finding that
> was actually measured rather than the one that was predicted.
>
> They carry `status: cancel` — the hub's status vocabulary is
> `wait|todo|progress|done|cancel`, and of those `cancel` is the accurate
> one: *these* task records are superseded, and equivalents live in
> CB-WP-0006 T05T07. The work is deferred, not abandoned.
## Phase C — implement what the instruments expose (cancelled here)
The three tasks below were written in full, then cancelled unstarted. Their
analysis now lives where the work will happen — **[CB-WP-0006](CB-WP-0006-instrument-the-table.md)
Phase B** — rather than here, so there is one copy and it is next to the
code it describes. What each was, and what it became:
## Task: K9 and K11 — durable log, LogStore port, conformance suite
```task
id: CB-WP-0005-T04
status: cancel
state_hub_task_id: "35876b28-97ac-4379-bb71-72e74c4116d4"
```
K9's acceptance property is mutation-proven unasserted; K11 has no durable
format; AM-11's conformance suite does not exist. → **CB-WP-0006 T05**.
## Task: K10 — the replay bundle and `--replay`
```task
id: CB-WP-0005-T05
status: cancel
state_hub_task_id: "8592b6e7-8b83-4b47-b310-90e16a35eb6a"
```
INTENT design decision 8 of 10, unimplemented; four controls required
verbatim from ADR-0005 §6. → **CB-WP-0006 T06**.
## Task: K14 and K18 — implement, or amend the spec
```task
id: CB-WP-0005-T06
status: cancel
state_hub_task_id: "774b1c8a-3b26-4c4f-a2c8-46cb0b7ba5a7"
```
`CommitWindow` has zero non-test users; the bench never touches
`ScenarioFile`. Implement, or amend with a recorded argument. →
**CB-WP-0006 T07**.
## Phase D — measure it, then say what the loop should change
## Task: control loop — did the numbers get worse, then better?
```task
id: CB-WP-0005-T07
status: todo
priority: high
state_hub_task_id: "a34baa46-fbfd-473b-9e15-811e17697791"
```
Commit `evidence/CB-EV-0004-assertion-coverage.md`. Three tests, all
reported:
1. **Did the denominators widen?** Name coverage before and after, per
spec. A widened denominator that reports the same percentage is the
result to publish — it would mean the instruments still count names.
2. **M-D1-MUT against the 9-of-12 prediction.** Report unmet if unmet; no
target moves in the commit that measures it. State which rows are
`unmutatable` and why.
3. **Did quality hold?** `make all` green with the new gates; and the
honest question this pass raises — did widening the instruments surface
*new* defects, or only the seven already known? Finding none would be
evidence the sweep was as complete as CB-RES-0004 claims, which is a
stronger claim than it sounds and should be stated as such.
Normalize per unit of work, as CB-WP-0004 T05 established: report share of
pass alongside absolute figures, and use `cb-cost --since` to window this
pass against the last.
## Task: retrospective and InnerLoop v1.4
```task
id: CB-WP-0005-T08
status: todo
priority: medium
state_hub_task_id: "3add7e1e-2082-44ef-9287-742c90e5c958"
```
One loop change is already earned and should be written whatever else this
pass finds. InnerLoop §Step 2's "numbers" row says the reviewer must
*reproduce the number independently* — satisfiable by re-running the
command that prints it, which is exactly what does **not** find this class.
The reviewer found AM-7 by opening a test out of curiosity, and said so.
> **v1.4:** the reviewer must read the assertion behind every quoted
> acceptance number and **mutate it**, not re-run the command that prints
> it.
This is the **second** instance of a verification step inheriting the
author's blindness — CB-WP-0002's dedup blind spot was the first. Both
fixes replace re-derivation with adversarial execution. Record whether
that generalizes.
The question to answer honestly: **CB-WP-0004 concluded that tooling
recovers capacity only where it removes the manual path. Does the same
test predict which gates work?** `env-test` and `task-done` removed the
manual path and held. Does a mutation gate remove the manual path — or is
writing a weak mutation the new `grep`?