AGGREGATE becomes a list of source roots and rule patterns become
per-spec, so the link runs over every numbered spec x every crate rather
than GroundRules.md x games/ground/src/lib.rs.
The prediction held on the first run:
AM-1b kernel spec->code link: 15/18 (83%) across 10 source files
unlinked: K10 K14 K18
Kernel rules are link-only by design, and the output says so: they are
kernel invariants with no aggregate, setup preset or command vocabulary,
so scenarios/kernel/*.yaml with covers: [K11] would be a tag in a
directory the runner cannot dispatch. Claiming scenario coverage for them
is the inflation this gate exists to prevent.
Per ADR-0005 §5 the kernel arm reports without feeding the exit code
until 2026-08-31, then binds — the date in the tool, not in prose, with
days remaining printed every run, because open-ended "gate it later" is
how AM-4's targets went unratified for four workplans. The self-test
asserts the gate returns 0 before that date and 2 after.
The zero-rules positive control is replicated on the new denominator: a
kernel regex that stops matching aborts rather than printing 0/0 as
though it were 100%.
The self-test passed while the tool was completely broken. A print(
inside say() became say(), so every real `make coverage` died with
RecursionError while --self-test reported all-ok — it only ever called
kernel_arm(quiet=True) and never executed the reporting path. The control
named the behaviour and did not assert it, which is precisely what this
workplan is about. Fixed by exercising the loud path and asserting it
prints, then verified by re-breaking say() and confirming both new checks
go red. Seventh instance of the harness-does-nothing shape, in the tool
written to find that shape.
Also caught by its own gate: a self-test label that printed "0 K-ids"
beside a passing ">5" assertion, because the detail string rebuilt the
pattern with different escaping. A label that contradicts its own check
is worse than no label.
k_rules, k_linked and k_unlinked are registered facts under facts-check.
A limit of that checker is recorded rather than patched: it is
line-based, so a tagged value that prose-wraps fails.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
322 lines
14 KiB
Markdown
322 lines
14 KiB
Markdown
---
|
||
id: CB-WP-0005
|
||
title: "Make the instruments count assertions, then fix what they expose"
|
||
status: proposed
|
||
state_hub_workstream_id: "0b95a1e3-7780-43d0-81e9-072ef7978734"
|
||
---
|
||
|
||
# Purpose
|
||
|
||
`research/CB-RES-0004-replay-and-kernel-coverage.md` (v2, after an
|
||
adversarial review that rejected v1 with four blocking findings):
|
||
|
||
> **Every coverage instrument in this project counts names. None counts
|
||
> assertions.** Four of the seven defects found this pass are named in the
|
||
> source and inert, so no name-based check finds them.
|
||
|
||
Two are mutation-proven: AM-7's `hash-identical` clause survives folding a
|
||
100k-event log from an unrelated genesis state, and K9's `through` field
|
||
survives `Snapshot::take` discarding it entirely. `evidence/CB-EV-0001`
|
||
carries **`AM-7 replay | met, 2,290×`** for the first of those.
|
||
|
||
[ADR-0005](../decisions/ADR-0005-assertion-coverage-and-replay.md) decides
|
||
the response. This workplan executes it in four phases: teach the
|
||
instruments to count assertions, correct the record they falsified,
|
||
implement what they expose, then measure whether any of it worked.
|
||
|
||
**Headline numbers are expected to get worse before they get better.** A
|
||
pass that widened every denominator and reported the same percentages
|
||
would have proved the instruments still count names. Per InnerLoop §Step 4
|
||
no target moves in the commit that measures it.
|
||
|
||
## Phase A — teach the instruments to count assertions
|
||
|
||
## Task: spec→code link over every numbered spec and every crate
|
||
|
||
```task
|
||
id: CB-WP-0005-T01
|
||
status: todo
|
||
priority: high
|
||
state_hub_task_id: "d47338c6-eb2e-4443-aad2-0b08c2d91895"
|
||
```
|
||
|
||
`tools/rule-coverage.py` hardcodes `AGGREGATE = "games/ground/src/lib.rs"`
|
||
and `RULE_RE = \*\*(GR-[A-Z]+\d+)` against `GroundRules.md` alone. The
|
||
reviewer's finding: a ~10-line generalization would have surfaced K10, K14
|
||
and K18 the day AM-1b shipped.
|
||
|
||
Deliver: `AGGREGATE` becomes a list of source roots; rule patterns become
|
||
per-spec; the link runs over **every numbered spec × every crate**.
|
||
|
||
Three constraints from ADR-0005 §5, all of which must be visible in the
|
||
output:
|
||
|
||
1. The kernel denominator **reports and does not feed the exit code until
|
||
2026-08-31**, after which it binds. The date lives in the tool and the
|
||
tool prints the days remaining — an open-ended "gate it later" is how
|
||
AM-4's targets went unratified for four workplans.
|
||
2. **Replicate the zero-rules positive control** on every new denominator.
|
||
The existing arm refuses to report over zero rules, a defect it was
|
||
fixed for; a kernel regex matching nothing must abort, not print
|
||
`0/0 (100%)`.
|
||
3. State the limit in the output, as the GR arm already does: this counts
|
||
names. It is the cheap half.
|
||
|
||
**Predicted:** K10, K14, K18 reported unlinked on the first run.
|
||
**Refuted if** any numbered rule in any spec is still unnamed in source
|
||
after the pass and the tool does not say so.
|
||
|
||
**Delivered, and the prediction held on the first run:**
|
||
|
||
```text
|
||
AM-1b kernel spec->code link: 15/18 (83%) K-rules named across 10 source files
|
||
NOTE: link only — K-rules are kernel invariants with no scenario
|
||
mechanism; and this counts names, not assertions
|
||
gate: reporting only for 31 more day(s), binds 2026-08-31 (ADR-0005 §5)
|
||
unlinked (declared in the spec, named nowhere in source):
|
||
K10 K14 K18
|
||
```
|
||
|
||
`AGGREGATE` is now a list of source roots, rule patterns are per-spec, and
|
||
the link runs over every numbered spec × every crate. The binding date
|
||
lives in the tool and the days remaining are printed every run. The
|
||
kernel figures (`k_rules`, `k_linked`, `k_unlinked`) are registered facts,
|
||
so they are under `make facts-check` from the day they first exist rather
|
||
than after they drift.
|
||
|
||
**The self-test passed while the tool was completely broken.** A
|
||
`print(` inside `say()` was rewritten to `say(`, so every real
|
||
`make coverage` died with `RecursionError` — while `--self-test` reported
|
||
all-ok, because it only ever called `kernel_arm(quiet=True)` and never
|
||
executed the reporting path.
|
||
|
||
That is this task's own thesis in miniature: **the control named the
|
||
behaviour and did not assert it.** Fixed by exercising the loud path under
|
||
`redirect_stdout` and asserting it prints, and verified by re-breaking
|
||
`say()` and confirming the two new checks go red. Seventh instance of the
|
||
harness-does-nothing shape, found in the tool written to find that shape.
|
||
|
||
**A limit of `facts-check` surfaced here and is recorded, not patched:**
|
||
the check is line-based, so a tagged value that prose-wraps onto the next
|
||
line fails. It cost three edits to place two tags. Reported for T07 —
|
||
either the checker spans a paragraph, or the rule is stated as
|
||
"tagged values must not wrap".
|
||
|
||
## Task: M-D1-MUT — one mutation per acceptance row
|
||
|
||
```task
|
||
id: CB-WP-0005-T02
|
||
status: todo
|
||
priority: high
|
||
state_hub_task_id: "c88f2696-ca12-44df-add3-0db3d2819c08"
|
||
```
|
||
|
||
The task this workplan exists for. Name-based checks catch 3 of the 7
|
||
defects; this catches the other 4.
|
||
|
||
Deliver `make mutation-check`: for each of the twelve AM-* acceptance
|
||
rows, a committed fixture that **inverts the row's stated property** and
|
||
an assertion that the suite goes **red**. A row whose mutation leaves the
|
||
suite green is a row backed by nothing.
|
||
|
||
Per ADR-0005 §1, `adapted:mutation-testing` — the denominator is
|
||
acceptance rows, not source lines, because the failure mode here is not an
|
||
untested branch but a headline number backed by nothing.
|
||
|
||
**The escape hatch is closed in advance:** a row for which no mutation can
|
||
be written is recorded `unmutatable` **with the reason** and **counts
|
||
against** the metric. A row nobody can invert asserts nothing.
|
||
|
||
Carries `--self-test`. Its own positive control is the one that matters:
|
||
a mutation harness that fails to apply its mutation reports every row as
|
||
`unmutatable` and looks thorough — this project has six harness-does-
|
||
nothing instances and one of them is being fixed in T03.
|
||
|
||
**Predicted: 9 of 12 rows turn red.** Deliberately beatable in both
|
||
directions — **12 of 12 refutes CB-RES-0004** (the instruments were better
|
||
than claimed, the finding shrinks to three absent rules, and M-D1-MUT was
|
||
not worth its CI cost); **3 of 12 means the pass is under-scoped and must
|
||
stop and re-plan** rather than proceed to Phase C.
|
||
|
||
## Phase B — correct the record
|
||
|
||
## Task: correct three committed verdicts and restore the fourth
|
||
|
||
```task
|
||
id: CB-WP-0005-T03
|
||
status: todo
|
||
priority: high
|
||
state_hub_task_id: "87fa2245-5f27-48af-906e-18affa41efcf"
|
||
```
|
||
|
||
Per ADR-0005 §4. Corrections are made **in place with a dated note**, not
|
||
silently edited — the correction trail is the artifact.
|
||
|
||
| row | committed | corrected to |
|
||
|---|---|---|
|
||
| AM-7 replay | `met, 2,290×` | timing clause **met**; `hash-identical` clause **withdrawn**, mutation-proven unmeasured, and the fold spans multiple game seeds so it is not a replay |
|
||
| AM-10 foreign types | `met` | **withdrawn as written** — no `cb-*-api` crate, so the population is empty; re-stated as a K6 determinism row, M-D4-LEAK marked **unmeasured** |
|
||
| AM-11 impl pairs | `met, narrow` | **unmet** until a shared conformance suite exists; the pair exists, the suite does not, and the metric is a bool over the suite |
|
||
| AM-1b | *absent from the scoreboard* | **added**, at its real value |
|
||
|
||
The last one is the point of the phase. `make coverage` prints `49/58` two
|
||
lines below the `100%` the scoreboard carried, and the scoreboard kept the
|
||
flattering half. Update `specs/GameKernel.md` §5 and
|
||
`specs/MetricsAndScenarios.md` §1 to match, and tag the new figures under
|
||
`make facts-check` so they cannot drift.
|
||
|
||
## Phase C — implement what the instruments expose
|
||
|
||
## Task: K9 and K11 — durable log, LogStore port, real conformance suite
|
||
|
||
```task
|
||
id: CB-WP-0005-T04
|
||
status: todo
|
||
priority: high
|
||
state_hub_task_id: "35876b28-97ac-4379-bb71-72e74c4116d4"
|
||
```
|
||
|
||
**K9** currently round-trips a `BTreeMap<String, u8>` with `EventSeq(17)`
|
||
as a literal. Deliver the real property: *snapshot at seq N + events
|
||
N+1..M ≡ genesis fold*, on `GroundState`, hash-compared.
|
||
|
||
**K11** — append-only, length-prefixed, versioned framing, with a
|
||
truncated tail **detected**. Reimplemented, not assimilated (ADR-0005 §2):
|
||
~100 lines against a format Kafka and EventStore converged on
|
||
independently. Charged to **AM-4a** (shipped runtime, 1.5% headroom) and
|
||
adds no dependency.
|
||
|
||
**The port** (ADR-0005 §3): a `LogStore` trait with a shared
|
||
`fn conformance<S: LogStore>(…)` driven by an in-memory impl **and** a
|
||
file-backed impl. Retro-fit the same shape to `KernelRng`, whose pair is
|
||
currently exercised by two separate non-shared tests — which is why AM-11
|
||
was never earned.
|
||
|
||
Negative controls, not optional: truncate by one byte → reject; corrupt
|
||
the length prefix → reject. K11's operative clause is *detection*.
|
||
|
||
## Task: K10 — the replay bundle and `--replay`
|
||
|
||
```task
|
||
id: CB-WP-0005-T05
|
||
status: todo
|
||
priority: high
|
||
state_hub_task_id: "8592b6e7-8b83-4b47-b310-90e16a35eb6a"
|
||
```
|
||
|
||
INTENT design decision **8 of 10**, unimplemented. `cb-sim` has no flag
|
||
parsing at all, so `--replay` has nowhere to go yet.
|
||
|
||
D2 note carried from the review: this is **not** "a directory of four
|
||
files". `scenario.rs` creates an `EventLog`, appends to it and never reads
|
||
it — the one production instantiation of the K11 log is a write-only sink.
|
||
`Pass` carries the *end* state, not an initial snapshot, and failures are a
|
||
formatted `String`, not structured expected-vs-actual. Plumb the log out of
|
||
`execute`, capture an initial snapshot, restructure `RunOutcome::Failed`.
|
||
|
||
The bundle writer is **dev-only** behind the `scenarios` feature and is
|
||
charged to **AM-4b** (9.4% headroom), not AM-4a.
|
||
|
||
**Four controls, from ADR-0005 §6, verbatim and all required:**
|
||
|
||
1. a committed **deliberately-failing fixture** scenario, plus an
|
||
assertion that bundles produced `> 0` and equals expected failures —
|
||
all 21 scenarios pass today, so a corpus sweep verifies zero bundles
|
||
and prints `ok`;
|
||
2. the comparison hash **read out of the bundle**, written by the first
|
||
process — never recomputed in the second, which is `assert_eq!(h, h)`
|
||
and the AM-7 defect exactly;
|
||
3. truncate-by-one-byte and corrupt-length-prefix rejection;
|
||
4. a **mutated-seed** control: change the manifest seed and the replay
|
||
must fail to reproduce. A round-trip that cannot fail proves nothing.
|
||
|
||
## Task: K14 and K18 — implement, or amend the spec and say why
|
||
|
||
```task
|
||
id: CB-WP-0005-T06
|
||
status: todo
|
||
priority: medium
|
||
state_hub_task_id: "774b1c8a-71c1-4e0a-91b1-0247a3e397cf"
|
||
```
|
||
|
||
Both are named nowhere in source, and both have an honest second option
|
||
that must be considered rather than assumed away.
|
||
|
||
**K14** — "one commit window per Select". `CommitWindow` exists, has zero
|
||
non-test users, and `games/ground` does not import it; GROUND collects
|
||
selections in its own aggregate. Either wire GROUND through the runtime
|
||
primitive as K14 requires, **or** amend K14 to describe what the kernel
|
||
actually guarantees and record why the primitive stays. A dead abstraction
|
||
certified by a test that exercises only itself is worse than no
|
||
abstraction — it is an AM-11-shaped claim waiting to be quoted.
|
||
|
||
**K18** — "Criterion benches driving the same scenario format". The bench
|
||
hardcodes commands in Rust and never touches `ScenarioFile`, and
|
||
`benchmarks/` contains only `baselines/model-prices.toml` against a spec
|
||
that says benches live there. Either drive the bench from scenario files,
|
||
**or** amend K18 and MetricsAndScenarios §3 to match reality.
|
||
|
||
Whichever way each goes, the outcome is a rule and an implementation that
|
||
agree. Deleting a rule to make a gate green is forbidden; deleting a rule
|
||
*with an argument*, recorded, is a legitimate outcome and the more likely
|
||
one for K18.
|
||
|
||
## Phase D — measure it, then say what the loop should change
|
||
|
||
## Task: control loop — did the numbers get worse, then better?
|
||
|
||
```task
|
||
id: CB-WP-0005-T07
|
||
status: todo
|
||
priority: high
|
||
state_hub_task_id: "a34baa46-fbfd-473b-9e15-811e17697791"
|
||
```
|
||
|
||
Commit `evidence/CB-EV-0004-assertion-coverage.md`. Three tests, all
|
||
reported:
|
||
|
||
1. **Did the denominators widen?** Name coverage before and after, per
|
||
spec. A widened denominator that reports the same percentage is the
|
||
result to publish — it would mean the instruments still count names.
|
||
2. **M-D1-MUT against the 9-of-12 prediction.** Report unmet if unmet; no
|
||
target moves in the commit that measures it. State which rows are
|
||
`unmutatable` and why.
|
||
3. **Did quality hold?** `make all` green with the new gates; and the
|
||
honest question this pass raises — did widening the instruments surface
|
||
*new* defects, or only the seven already known? Finding none would be
|
||
evidence the sweep was as complete as CB-RES-0004 claims, which is a
|
||
stronger claim than it sounds and should be stated as such.
|
||
|
||
Normalize per unit of work, as CB-WP-0004 T05 established: report share of
|
||
pass alongside absolute figures, and use `cb-cost --since` to window this
|
||
pass against the last.
|
||
|
||
## Task: retrospective and InnerLoop v1.4
|
||
|
||
```task
|
||
id: CB-WP-0005-T08
|
||
status: todo
|
||
priority: medium
|
||
state_hub_task_id: "3add7e1e-2082-44ef-9287-742c90e5c958"
|
||
```
|
||
|
||
One loop change is already earned and should be written whatever else this
|
||
pass finds. InnerLoop §Step 2's "numbers" row says the reviewer must
|
||
*reproduce the number independently* — satisfiable by re-running the
|
||
command that prints it, which is exactly what does **not** find this class.
|
||
The reviewer found AM-7 by opening a test out of curiosity, and said so.
|
||
|
||
> **v1.4:** the reviewer must read the assertion behind every quoted
|
||
> acceptance number and **mutate it**, not re-run the command that prints
|
||
> it.
|
||
|
||
This is the **second** instance of a verification step inheriting the
|
||
author's blindness — CB-WP-0002's dedup blind spot was the first. Both
|
||
fixes replace re-derivation with adversarial execution. Record whether
|
||
that generalizes.
|
||
|
||
The question to answer honestly: **CB-WP-0004 concluded that tooling
|
||
recovers capacity only where it removes the manual path. Does the same
|
||
test predict which gates work?** `env-test` and `task-done` removed the
|
||
manual path and held. Does a mutation gate remove the manual path — or is
|
||
writing a weak mutation the new `grep`?
|