Some checks failed
ci / check (push) Failing after 3s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
368 lines
16 KiB
Markdown
368 lines
16 KiB
Markdown
---
|
||
id: CB-WP-0005
|
||
title: "Make the instruments count assertions, then fix what they expose"
|
||
status: proposed
|
||
state_hub_workstream_id: "0b95a1e3-7780-43d0-81e9-072ef7978734"
|
||
---
|
||
|
||
# Purpose
|
||
|
||
`research/CB-RES-0004-replay-and-kernel-coverage.md` (v2, after an
|
||
adversarial review that rejected v1 with four blocking findings):
|
||
|
||
> **Every coverage instrument in this project counts names. None counts
|
||
> assertions.** Four of the seven defects found this pass are named in the
|
||
> source and inert, so no name-based check finds them.
|
||
|
||
Two are mutation-proven: AM-7's `hash-identical` clause survives folding a
|
||
100k-event log from an unrelated genesis state, and K9's `through` field
|
||
survives `Snapshot::take` discarding it entirely. `evidence/CB-EV-0001`
|
||
carries **`AM-7 replay | met, 2,290×`** for the first of those.
|
||
|
||
[ADR-0005](../decisions/ADR-0005-assertion-coverage-and-replay.md) decides
|
||
the response. This workplan executes it in four phases: teach the
|
||
instruments to count assertions, correct the record they falsified,
|
||
implement what they expose, then measure whether any of it worked.
|
||
|
||
**Headline numbers are expected to get worse before they get better.** A
|
||
pass that widened every denominator and reported the same percentages
|
||
would have proved the instruments still count names. Per InnerLoop §Step 4
|
||
no target moves in the commit that measures it.
|
||
|
||
## Phase A — teach the instruments to count assertions
|
||
|
||
## Task: spec→code link over every numbered spec and every crate
|
||
|
||
```task
|
||
id: CB-WP-0005-T01
|
||
status: done
|
||
priority: high
|
||
state_hub_task_id: "d47338c6-eb2e-4443-aad2-0b08c2d91895"
|
||
```
|
||
|
||
`tools/rule-coverage.py` hardcodes `AGGREGATE = "games/ground/src/lib.rs"`
|
||
and `RULE_RE = \*\*(GR-[A-Z]+\d+)` against `GroundRules.md` alone. The
|
||
reviewer's finding: a ~10-line generalization would have surfaced K10, K14
|
||
and K18 the day AM-1b shipped.
|
||
|
||
Deliver: `AGGREGATE` becomes a list of source roots; rule patterns become
|
||
per-spec; the link runs over **every numbered spec × every crate**.
|
||
|
||
Three constraints from ADR-0005 §5, all of which must be visible in the
|
||
output:
|
||
|
||
1. The kernel denominator **reports and does not feed the exit code until
|
||
2026-08-31**, after which it binds. The date lives in the tool and the
|
||
tool prints the days remaining — an open-ended "gate it later" is how
|
||
AM-4's targets went unratified for four workplans.
|
||
2. **Replicate the zero-rules positive control** on every new denominator.
|
||
The existing arm refuses to report over zero rules, a defect it was
|
||
fixed for; a kernel regex matching nothing must abort, not print
|
||
`0/0 (100%)`.
|
||
3. State the limit in the output, as the GR arm already does: this counts
|
||
names. It is the cheap half.
|
||
|
||
**Predicted:** K10, K14, K18 reported unlinked on the first run.
|
||
**Refuted if** any numbered rule in any spec is still unnamed in source
|
||
after the pass and the tool does not say so.
|
||
|
||
**Delivered, and the prediction held on the first run:**
|
||
|
||
```text
|
||
AM-1b kernel spec->code link: 15/18 (83%) K-rules named across 10 source files
|
||
NOTE: link only — K-rules are kernel invariants with no scenario
|
||
mechanism; and this counts names, not assertions
|
||
gate: reporting only for 31 more day(s), binds 2026-08-31 (ADR-0005 §5)
|
||
unlinked (declared in the spec, named nowhere in source):
|
||
K10 K14 K18
|
||
```
|
||
|
||
`AGGREGATE` is now a list of source roots, rule patterns are per-spec, and
|
||
the link runs over every numbered spec × every crate. The binding date
|
||
lives in the tool and the days remaining are printed every run. The
|
||
kernel figures (`k_rules`, `k_linked`, `k_unlinked`) are registered facts,
|
||
so they are under `make facts-check` from the day they first exist rather
|
||
than after they drift.
|
||
|
||
**The self-test passed while the tool was completely broken.** A
|
||
`print(` inside `say()` was rewritten to `say(`, so every real
|
||
`make coverage` died with `RecursionError` — while `--self-test` reported
|
||
all-ok, because it only ever called `kernel_arm(quiet=True)` and never
|
||
executed the reporting path.
|
||
|
||
That is this task's own thesis in miniature: **the control named the
|
||
behaviour and did not assert it.** Fixed by exercising the loud path under
|
||
`redirect_stdout` and asserting it prints, and verified by re-breaking
|
||
`say()` and confirming the two new checks go red. Seventh instance of the
|
||
harness-does-nothing shape, found in the tool written to find that shape.
|
||
|
||
**A limit of `facts-check` surfaced here and is recorded, not patched:**
|
||
the check is line-based, so a tagged value that prose-wraps onto the next
|
||
line fails. It cost three edits to place two tags. Reported for T07 —
|
||
either the checker spans a paragraph, or the rule is stated as
|
||
"tagged values must not wrap".
|
||
|
||
## Task: M-D1-MUT — one mutation per acceptance row
|
||
|
||
```task
|
||
id: CB-WP-0005-T02
|
||
status: done
|
||
priority: high
|
||
state_hub_task_id: "c88f2696-ca12-44df-add3-0db3d2819c08"
|
||
```
|
||
|
||
The task this workplan exists for. Name-based checks catch 3 of the 7
|
||
defects; this catches the other 4.
|
||
|
||
Deliver `make mutation-check`: for each of the twelve AM-* acceptance
|
||
rows, a committed fixture that **inverts the row's stated property** and
|
||
an assertion that the suite goes **red**. A row whose mutation leaves the
|
||
suite green is a row backed by nothing.
|
||
|
||
Per ADR-0005 §1, `adapted:mutation-testing` — the denominator is
|
||
acceptance rows, not source lines, because the failure mode here is not an
|
||
untested branch but a headline number backed by nothing.
|
||
|
||
**The escape hatch is closed in advance:** a row for which no mutation can
|
||
be written is recorded `unmutatable` **with the reason** and **counts
|
||
against** the metric. A row nobody can invert asserts nothing.
|
||
|
||
Carries `--self-test`. Its own positive control is the one that matters:
|
||
a mutation harness that fails to apply its mutation reports every row as
|
||
`unmutatable` and looks thorough — this project has six harness-does-
|
||
nothing instances and one of them is being fixed in T03.
|
||
|
||
**Predicted: 9 of 12 rows turn red.** Deliberately beatable in both
|
||
directions — **12 of 12 refutes CB-RES-0004** (the instruments were better
|
||
than claimed, the finding shrinks to three absent rules, and M-D1-MUT was
|
||
not worth its CI cost); **3 of 12 means the pass is under-scoped and must
|
||
stop and re-plan** rather than proceed to Phase C.
|
||
|
||
**Delivered. Measured: 4 of 14 rows enforced. The prediction is badly
|
||
unmet.**
|
||
|
||
```text
|
||
M-D1-MUT: 4/14 rows enforced
|
||
PARTIAL 2 (AM-7, AM-8 — some clauses live, some inert)
|
||
unmutatable 8 (no property to invert, reason stated per row)
|
||
SURVIVED 0
|
||
```
|
||
|
||
**First correction: there are 14 acceptance rows, not 12.** ADR-0005 and
|
||
this workplan both said twelve; AM-4 splits into a/b/c. The 9-of-12 (75%)
|
||
prediction is evaluated as ≥10 of 14 on the same basis. Measured 4 (29%).
|
||
|
||
**Second correction, and the one that matters: my first run reported two
|
||
SURVIVED rows, and both were my own bad mutations.**
|
||
|
||
- AM-8: `pub struct NullRng;` → `pub struct NullRng {}` — semantically
|
||
identical, a no-op.
|
||
- AM-12: renaming `max_age_days` in the price sheet — CA-17 reads it with
|
||
`.get("max_age_days", 90)`, so removing it changes nothing.
|
||
|
||
Both would have been published as *"this row asserts nothing"* — a false
|
||
accusation against code that is in fact fine. Replaced with real
|
||
inversions (inject a per-construction counter into the ChaCha seed;
|
||
revert the AC-9 output resolution to the first-wins bug it was fixed for),
|
||
after which both go red. **T08's question — "is writing a weak mutation
|
||
the new grep?" — is answered on the first attempt: yes, demonstrably.**
|
||
|
||
**The finding is larger than the workplan assumed.** 8 of 14 rows are
|
||
`unmutatable`: AM-2, AM-3, AM-4c, AM-5, AM-6, AM-9, AM-10, AM-11 have no
|
||
instrument at all. AM-6 is the sharpest — **nothing in the workspace
|
||
compares any number to 100,000 events/s**, the headline throughput claim.
|
||
The problem is not three unimplemented rules; it is that **more than half
|
||
the acceptance table has nothing behind it.**
|
||
|
||
Harness controls that earned their place: a stale find-string reports
|
||
`HARNESS-BROKEN` rather than silently scoring the baseline as the mutant;
|
||
a red baseline reports `inconclusive` rather than `red`; the tree is
|
||
restored in a `finally` and the restoration is verified.
|
||
|
||
`make mutation-check` is deliberately **not** in `make all` — it rebuilds
|
||
the workspace once per mutated row. `--self-test` is in `make self-tests`.
|
||
|
||
**Stop condition: see the note in T07 and the decision recorded there.**
|
||
|
||
## Phase B — correct the record
|
||
|
||
## Task: correct three committed verdicts and restore the fourth
|
||
|
||
```task
|
||
id: CB-WP-0005-T03
|
||
status: todo
|
||
priority: high
|
||
state_hub_task_id: "87fa2245-5f27-48af-906e-18affa41efcf"
|
||
```
|
||
|
||
Per ADR-0005 §4. Corrections are made **in place with a dated note**, not
|
||
silently edited — the correction trail is the artifact.
|
||
|
||
| row | committed | corrected to |
|
||
|---|---|---|
|
||
| AM-7 replay | `met, 2,290×` | timing clause **met**; `hash-identical` clause **withdrawn**, mutation-proven unmeasured, and the fold spans multiple game seeds so it is not a replay |
|
||
| AM-10 foreign types | `met` | **withdrawn as written** — no `cb-*-api` crate, so the population is empty; re-stated as a K6 determinism row, M-D4-LEAK marked **unmeasured** |
|
||
| AM-11 impl pairs | `met, narrow` | **unmet** until a shared conformance suite exists; the pair exists, the suite does not, and the metric is a bool over the suite |
|
||
| AM-1b | *absent from the scoreboard* | **added**, at its real value |
|
||
|
||
The last one is the point of the phase. `make coverage` prints `49/58` two
|
||
lines below the `100%` the scoreboard carried, and the scoreboard kept the
|
||
flattering half. Update `specs/GameKernel.md` §5 and
|
||
`specs/MetricsAndScenarios.md` §1 to match, and tag the new figures under
|
||
`make facts-check` so they cannot drift.
|
||
|
||
## Phase C — implement what the instruments expose
|
||
|
||
## Task: K9 and K11 — durable log, LogStore port, real conformance suite
|
||
|
||
```task
|
||
id: CB-WP-0005-T04
|
||
status: todo
|
||
priority: high
|
||
state_hub_task_id: "35876b28-97ac-4379-bb71-72e74c4116d4"
|
||
```
|
||
|
||
**K9** currently round-trips a `BTreeMap<String, u8>` with `EventSeq(17)`
|
||
as a literal. Deliver the real property: *snapshot at seq N + events
|
||
N+1..M ≡ genesis fold*, on `GroundState`, hash-compared.
|
||
|
||
**K11** — append-only, length-prefixed, versioned framing, with a
|
||
truncated tail **detected**. Reimplemented, not assimilated (ADR-0005 §2):
|
||
~100 lines against a format Kafka and EventStore converged on
|
||
independently. Charged to **AM-4a** (shipped runtime, 1.5% headroom) and
|
||
adds no dependency.
|
||
|
||
**The port** (ADR-0005 §3): a `LogStore` trait with a shared
|
||
`fn conformance<S: LogStore>(…)` driven by an in-memory impl **and** a
|
||
file-backed impl. Retro-fit the same shape to `KernelRng`, whose pair is
|
||
currently exercised by two separate non-shared tests — which is why AM-11
|
||
was never earned.
|
||
|
||
Negative controls, not optional: truncate by one byte → reject; corrupt
|
||
the length prefix → reject. K11's operative clause is *detection*.
|
||
|
||
## Task: K10 — the replay bundle and `--replay`
|
||
|
||
```task
|
||
id: CB-WP-0005-T05
|
||
status: todo
|
||
priority: high
|
||
state_hub_task_id: "8592b6e7-8b83-4b47-b310-90e16a35eb6a"
|
||
```
|
||
|
||
INTENT design decision **8 of 10**, unimplemented. `cb-sim` has no flag
|
||
parsing at all, so `--replay` has nowhere to go yet.
|
||
|
||
D2 note carried from the review: this is **not** "a directory of four
|
||
files". `scenario.rs` creates an `EventLog`, appends to it and never reads
|
||
it — the one production instantiation of the K11 log is a write-only sink.
|
||
`Pass` carries the *end* state, not an initial snapshot, and failures are a
|
||
formatted `String`, not structured expected-vs-actual. Plumb the log out of
|
||
`execute`, capture an initial snapshot, restructure `RunOutcome::Failed`.
|
||
|
||
The bundle writer is **dev-only** behind the `scenarios` feature and is
|
||
charged to **AM-4b** (9.4% headroom), not AM-4a.
|
||
|
||
**Four controls, from ADR-0005 §6, verbatim and all required:**
|
||
|
||
1. a committed **deliberately-failing fixture** scenario, plus an
|
||
assertion that bundles produced `> 0` and equals expected failures —
|
||
all 21 scenarios pass today, so a corpus sweep verifies zero bundles
|
||
and prints `ok`;
|
||
2. the comparison hash **read out of the bundle**, written by the first
|
||
process — never recomputed in the second, which is `assert_eq!(h, h)`
|
||
and the AM-7 defect exactly;
|
||
3. truncate-by-one-byte and corrupt-length-prefix rejection;
|
||
4. a **mutated-seed** control: change the manifest seed and the replay
|
||
must fail to reproduce. A round-trip that cannot fail proves nothing.
|
||
|
||
## Task: K14 and K18 — implement, or amend the spec and say why
|
||
|
||
```task
|
||
id: CB-WP-0005-T06
|
||
status: todo
|
||
priority: medium
|
||
state_hub_task_id: "774b1c8a-71c1-4e0a-91b1-0247a3e397cf"
|
||
```
|
||
|
||
Both are named nowhere in source, and both have an honest second option
|
||
that must be considered rather than assumed away.
|
||
|
||
**K14** — "one commit window per Select". `CommitWindow` exists, has zero
|
||
non-test users, and `games/ground` does not import it; GROUND collects
|
||
selections in its own aggregate. Either wire GROUND through the runtime
|
||
primitive as K14 requires, **or** amend K14 to describe what the kernel
|
||
actually guarantees and record why the primitive stays. A dead abstraction
|
||
certified by a test that exercises only itself is worse than no
|
||
abstraction — it is an AM-11-shaped claim waiting to be quoted.
|
||
|
||
**K18** — "Criterion benches driving the same scenario format". The bench
|
||
hardcodes commands in Rust and never touches `ScenarioFile`, and
|
||
`benchmarks/` contains only `baselines/model-prices.toml` against a spec
|
||
that says benches live there. Either drive the bench from scenario files,
|
||
**or** amend K18 and MetricsAndScenarios §3 to match reality.
|
||
|
||
Whichever way each goes, the outcome is a rule and an implementation that
|
||
agree. Deleting a rule to make a gate green is forbidden; deleting a rule
|
||
*with an argument*, recorded, is a legitimate outcome and the more likely
|
||
one for K18.
|
||
|
||
## Phase D — measure it, then say what the loop should change
|
||
|
||
## Task: control loop — did the numbers get worse, then better?
|
||
|
||
```task
|
||
id: CB-WP-0005-T07
|
||
status: todo
|
||
priority: high
|
||
state_hub_task_id: "a34baa46-fbfd-473b-9e15-811e17697791"
|
||
```
|
||
|
||
Commit `evidence/CB-EV-0004-assertion-coverage.md`. Three tests, all
|
||
reported:
|
||
|
||
1. **Did the denominators widen?** Name coverage before and after, per
|
||
spec. A widened denominator that reports the same percentage is the
|
||
result to publish — it would mean the instruments still count names.
|
||
2. **M-D1-MUT against the 9-of-12 prediction.** Report unmet if unmet; no
|
||
target moves in the commit that measures it. State which rows are
|
||
`unmutatable` and why.
|
||
3. **Did quality hold?** `make all` green with the new gates; and the
|
||
honest question this pass raises — did widening the instruments surface
|
||
*new* defects, or only the seven already known? Finding none would be
|
||
evidence the sweep was as complete as CB-RES-0004 claims, which is a
|
||
stronger claim than it sounds and should be stated as such.
|
||
|
||
Normalize per unit of work, as CB-WP-0004 T05 established: report share of
|
||
pass alongside absolute figures, and use `cb-cost --since` to window this
|
||
pass against the last.
|
||
|
||
## Task: retrospective and InnerLoop v1.4
|
||
|
||
```task
|
||
id: CB-WP-0005-T08
|
||
status: todo
|
||
priority: medium
|
||
state_hub_task_id: "3add7e1e-2082-44ef-9287-742c90e5c958"
|
||
```
|
||
|
||
One loop change is already earned and should be written whatever else this
|
||
pass finds. InnerLoop §Step 2's "numbers" row says the reviewer must
|
||
*reproduce the number independently* — satisfiable by re-running the
|
||
command that prints it, which is exactly what does **not** find this class.
|
||
The reviewer found AM-7 by opening a test out of curiosity, and said so.
|
||
|
||
> **v1.4:** the reviewer must read the assertion behind every quoted
|
||
> acceptance number and **mutate it**, not re-run the command that prints
|
||
> it.
|
||
|
||
This is the **second** instance of a verification step inheriting the
|
||
author's blindness — CB-WP-0002's dedup blind spot was the first. Both
|
||
fixes replace re-derivation with adversarial execution. Record whether
|
||
that generalizes.
|
||
|
||
The question to answer honestly: **CB-WP-0004 concluded that tooling
|
||
recovers capacity only where it removes the manual path. Does the same
|
||
test predict which gates work?** `env-test` and `task-done` removed the
|
||
manual path and held. Does a mutation gate remove the manual path — or is
|
||
writing a weak mutation the new `grep`?
|