278 lines
12 KiB
Markdown
278 lines
12 KiB
Markdown
|
|
---
|
|||
|
|
id: CB-WP-0005
|
|||
|
|
title: "Make the instruments count assertions, then fix what they expose"
|
|||
|
|
status: proposed
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
# Purpose
|
|||
|
|
|
|||
|
|
`research/CB-RES-0004-replay-and-kernel-coverage.md` (v2, after an
|
|||
|
|
adversarial review that rejected v1 with four blocking findings):
|
|||
|
|
|
|||
|
|
> **Every coverage instrument in this project counts names. None counts
|
|||
|
|
> assertions.** Four of the seven defects found this pass are named in the
|
|||
|
|
> source and inert, so no name-based check finds them.
|
|||
|
|
|
|||
|
|
Two are mutation-proven: AM-7's `hash-identical` clause survives folding a
|
|||
|
|
100k-event log from an unrelated genesis state, and K9's `through` field
|
|||
|
|
survives `Snapshot::take` discarding it entirely. `evidence/CB-EV-0001`
|
|||
|
|
carries **`AM-7 replay | met, 2,290×`** for the first of those.
|
|||
|
|
|
|||
|
|
[ADR-0005](../decisions/ADR-0005-assertion-coverage-and-replay.md) decides
|
|||
|
|
the response. This workplan executes it in four phases: teach the
|
|||
|
|
instruments to count assertions, correct the record they falsified,
|
|||
|
|
implement what they expose, then measure whether any of it worked.
|
|||
|
|
|
|||
|
|
**Headline numbers are expected to get worse before they get better.** A
|
|||
|
|
pass that widened every denominator and reported the same percentages
|
|||
|
|
would have proved the instruments still count names. Per InnerLoop §Step 4
|
|||
|
|
no target moves in the commit that measures it.
|
|||
|
|
|
|||
|
|
## Phase A — teach the instruments to count assertions
|
|||
|
|
|
|||
|
|
## Task: spec→code link over every numbered spec and every crate
|
|||
|
|
|
|||
|
|
```task
|
|||
|
|
id: CB-WP-0005-T01
|
|||
|
|
status: todo
|
|||
|
|
priority: high
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
`tools/rule-coverage.py` hardcodes `AGGREGATE = "games/ground/src/lib.rs"`
|
|||
|
|
and `RULE_RE = \*\*(GR-[A-Z]+\d+)` against `GroundRules.md` alone. The
|
|||
|
|
reviewer's finding: a ~10-line generalization would have surfaced K10, K14
|
|||
|
|
and K18 the day AM-1b shipped.
|
|||
|
|
|
|||
|
|
Deliver: `AGGREGATE` becomes a list of source roots; rule patterns become
|
|||
|
|
per-spec; the link runs over **every numbered spec × every crate**.
|
|||
|
|
|
|||
|
|
Three constraints from ADR-0005 §5, all of which must be visible in the
|
|||
|
|
output:
|
|||
|
|
|
|||
|
|
1. The kernel denominator **reports and does not feed the exit code until
|
|||
|
|
2026-08-31**, after which it binds. The date lives in the tool and the
|
|||
|
|
tool prints the days remaining — an open-ended "gate it later" is how
|
|||
|
|
AM-4's targets went unratified for four workplans.
|
|||
|
|
2. **Replicate the zero-rules positive control** on every new denominator.
|
|||
|
|
The existing arm refuses to report over zero rules, a defect it was
|
|||
|
|
fixed for; a kernel regex matching nothing must abort, not print
|
|||
|
|
`0/0 (100%)`.
|
|||
|
|
3. State the limit in the output, as the GR arm already does: this counts
|
|||
|
|
names. It is the cheap half.
|
|||
|
|
|
|||
|
|
**Predicted:** K10, K14, K18 reported unlinked on the first run.
|
|||
|
|
**Refuted if** any numbered rule in any spec is still unnamed in source
|
|||
|
|
after the pass and the tool does not say so.
|
|||
|
|
|
|||
|
|
## Task: M-D1-MUT — one mutation per acceptance row
|
|||
|
|
|
|||
|
|
```task
|
|||
|
|
id: CB-WP-0005-T02
|
|||
|
|
status: todo
|
|||
|
|
priority: high
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
The task this workplan exists for. Name-based checks catch 3 of the 7
|
|||
|
|
defects; this catches the other 4.
|
|||
|
|
|
|||
|
|
Deliver `make mutation-check`: for each of the twelve AM-* acceptance
|
|||
|
|
rows, a committed fixture that **inverts the row's stated property** and
|
|||
|
|
an assertion that the suite goes **red**. A row whose mutation leaves the
|
|||
|
|
suite green is a row backed by nothing.
|
|||
|
|
|
|||
|
|
Per ADR-0005 §1, `adapted:mutation-testing` — the denominator is
|
|||
|
|
acceptance rows, not source lines, because the failure mode here is not an
|
|||
|
|
untested branch but a headline number backed by nothing.
|
|||
|
|
|
|||
|
|
**The escape hatch is closed in advance:** a row for which no mutation can
|
|||
|
|
be written is recorded `unmutatable` **with the reason** and **counts
|
|||
|
|
against** the metric. A row nobody can invert asserts nothing.
|
|||
|
|
|
|||
|
|
Carries `--self-test`. Its own positive control is the one that matters:
|
|||
|
|
a mutation harness that fails to apply its mutation reports every row as
|
|||
|
|
`unmutatable` and looks thorough — this project has six harness-does-
|
|||
|
|
nothing instances and one of them is being fixed in T03.
|
|||
|
|
|
|||
|
|
**Predicted: 9 of 12 rows turn red.** Deliberately beatable in both
|
|||
|
|
directions — **12 of 12 refutes CB-RES-0004** (the instruments were better
|
|||
|
|
than claimed, the finding shrinks to three absent rules, and M-D1-MUT was
|
|||
|
|
not worth its CI cost); **3 of 12 means the pass is under-scoped and must
|
|||
|
|
stop and re-plan** rather than proceed to Phase C.
|
|||
|
|
|
|||
|
|
## Phase B — correct the record
|
|||
|
|
|
|||
|
|
## Task: correct three committed verdicts and restore the fourth
|
|||
|
|
|
|||
|
|
```task
|
|||
|
|
id: CB-WP-0005-T03
|
|||
|
|
status: todo
|
|||
|
|
priority: high
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Per ADR-0005 §4. Corrections are made **in place with a dated note**, not
|
|||
|
|
silently edited — the correction trail is the artifact.
|
|||
|
|
|
|||
|
|
| row | committed | corrected to |
|
|||
|
|
|---|---|---|
|
|||
|
|
| AM-7 replay | `met, 2,290×` | timing clause **met**; `hash-identical` clause **withdrawn**, mutation-proven unmeasured, and the fold spans multiple game seeds so it is not a replay |
|
|||
|
|
| AM-10 foreign types | `met` | **withdrawn as written** — no `cb-*-api` crate, so the population is empty; re-stated as a K6 determinism row, M-D4-LEAK marked **unmeasured** |
|
|||
|
|
| AM-11 impl pairs | `met, narrow` | **unmet** until a shared conformance suite exists; the pair exists, the suite does not, and the metric is a bool over the suite |
|
|||
|
|
| AM-1b | *absent from the scoreboard* | **added**, at its real value |
|
|||
|
|
|
|||
|
|
The last one is the point of the phase. `make coverage` prints `49/58` two
|
|||
|
|
lines below the `100%` the scoreboard carried, and the scoreboard kept the
|
|||
|
|
flattering half. Update `specs/GameKernel.md` §5 and
|
|||
|
|
`specs/MetricsAndScenarios.md` §1 to match, and tag the new figures under
|
|||
|
|
`make facts-check` so they cannot drift.
|
|||
|
|
|
|||
|
|
## Phase C — implement what the instruments expose
|
|||
|
|
|
|||
|
|
## Task: K9 and K11 — durable log, LogStore port, real conformance suite
|
|||
|
|
|
|||
|
|
```task
|
|||
|
|
id: CB-WP-0005-T04
|
|||
|
|
status: todo
|
|||
|
|
priority: high
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
**K9** currently round-trips a `BTreeMap<String, u8>` with `EventSeq(17)`
|
|||
|
|
as a literal. Deliver the real property: *snapshot at seq N + events
|
|||
|
|
N+1..M ≡ genesis fold*, on `GroundState`, hash-compared.
|
|||
|
|
|
|||
|
|
**K11** — append-only, length-prefixed, versioned framing, with a
|
|||
|
|
truncated tail **detected**. Reimplemented, not assimilated (ADR-0005 §2):
|
|||
|
|
~100 lines against a format Kafka and EventStore converged on
|
|||
|
|
independently. Charged to **AM-4a** (shipped runtime, 1.5% headroom) and
|
|||
|
|
adds no dependency.
|
|||
|
|
|
|||
|
|
**The port** (ADR-0005 §3): a `LogStore` trait with a shared
|
|||
|
|
`fn conformance<S: LogStore>(…)` driven by an in-memory impl **and** a
|
|||
|
|
file-backed impl. Retro-fit the same shape to `KernelRng`, whose pair is
|
|||
|
|
currently exercised by two separate non-shared tests — which is why AM-11
|
|||
|
|
was never earned.
|
|||
|
|
|
|||
|
|
Negative controls, not optional: truncate by one byte → reject; corrupt
|
|||
|
|
the length prefix → reject. K11's operative clause is *detection*.
|
|||
|
|
|
|||
|
|
## Task: K10 — the replay bundle and `--replay`
|
|||
|
|
|
|||
|
|
```task
|
|||
|
|
id: CB-WP-0005-T05
|
|||
|
|
status: todo
|
|||
|
|
priority: high
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
INTENT design decision **8 of 10**, unimplemented. `cb-sim` has no flag
|
|||
|
|
parsing at all, so `--replay` has nowhere to go yet.
|
|||
|
|
|
|||
|
|
D2 note carried from the review: this is **not** "a directory of four
|
|||
|
|
files". `scenario.rs` creates an `EventLog`, appends to it and never reads
|
|||
|
|
it — the one production instantiation of the K11 log is a write-only sink.
|
|||
|
|
`Pass` carries the *end* state, not an initial snapshot, and failures are a
|
|||
|
|
formatted `String`, not structured expected-vs-actual. Plumb the log out of
|
|||
|
|
`execute`, capture an initial snapshot, restructure `RunOutcome::Failed`.
|
|||
|
|
|
|||
|
|
The bundle writer is **dev-only** behind the `scenarios` feature and is
|
|||
|
|
charged to **AM-4b** (9.4% headroom), not AM-4a.
|
|||
|
|
|
|||
|
|
**Four controls, from ADR-0005 §6, verbatim and all required:**
|
|||
|
|
|
|||
|
|
1. a committed **deliberately-failing fixture** scenario, plus an
|
|||
|
|
assertion that bundles produced `> 0` and equals expected failures —
|
|||
|
|
all 21 scenarios pass today, so a corpus sweep verifies zero bundles
|
|||
|
|
and prints `ok`;
|
|||
|
|
2. the comparison hash **read out of the bundle**, written by the first
|
|||
|
|
process — never recomputed in the second, which is `assert_eq!(h, h)`
|
|||
|
|
and the AM-7 defect exactly;
|
|||
|
|
3. truncate-by-one-byte and corrupt-length-prefix rejection;
|
|||
|
|
4. a **mutated-seed** control: change the manifest seed and the replay
|
|||
|
|
must fail to reproduce. A round-trip that cannot fail proves nothing.
|
|||
|
|
|
|||
|
|
## Task: K14 and K18 — implement, or amend the spec and say why
|
|||
|
|
|
|||
|
|
```task
|
|||
|
|
id: CB-WP-0005-T06
|
|||
|
|
status: todo
|
|||
|
|
priority: medium
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Both are named nowhere in source, and both have an honest second option
|
|||
|
|
that must be considered rather than assumed away.
|
|||
|
|
|
|||
|
|
**K14** — "one commit window per Select". `CommitWindow` exists, has zero
|
|||
|
|
non-test users, and `games/ground` does not import it; GROUND collects
|
|||
|
|
selections in its own aggregate. Either wire GROUND through the runtime
|
|||
|
|
primitive as K14 requires, **or** amend K14 to describe what the kernel
|
|||
|
|
actually guarantees and record why the primitive stays. A dead abstraction
|
|||
|
|
certified by a test that exercises only itself is worse than no
|
|||
|
|
abstraction — it is an AM-11-shaped claim waiting to be quoted.
|
|||
|
|
|
|||
|
|
**K18** — "Criterion benches driving the same scenario format". The bench
|
|||
|
|
hardcodes commands in Rust and never touches `ScenarioFile`, and
|
|||
|
|
`benchmarks/` contains only `baselines/model-prices.toml` against a spec
|
|||
|
|
that says benches live there. Either drive the bench from scenario files,
|
|||
|
|
**or** amend K18 and MetricsAndScenarios §3 to match reality.
|
|||
|
|
|
|||
|
|
Whichever way each goes, the outcome is a rule and an implementation that
|
|||
|
|
agree. Deleting a rule to make a gate green is forbidden; deleting a rule
|
|||
|
|
*with an argument*, recorded, is a legitimate outcome and the more likely
|
|||
|
|
one for K18.
|
|||
|
|
|
|||
|
|
## Phase D — measure it, then say what the loop should change
|
|||
|
|
|
|||
|
|
## Task: control loop — did the numbers get worse, then better?
|
|||
|
|
|
|||
|
|
```task
|
|||
|
|
id: CB-WP-0005-T07
|
|||
|
|
status: todo
|
|||
|
|
priority: high
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Commit `evidence/CB-EV-0004-assertion-coverage.md`. Three tests, all
|
|||
|
|
reported:
|
|||
|
|
|
|||
|
|
1. **Did the denominators widen?** Name coverage before and after, per
|
|||
|
|
spec. A widened denominator that reports the same percentage is the
|
|||
|
|
result to publish — it would mean the instruments still count names.
|
|||
|
|
2. **M-D1-MUT against the 9-of-12 prediction.** Report unmet if unmet; no
|
|||
|
|
target moves in the commit that measures it. State which rows are
|
|||
|
|
`unmutatable` and why.
|
|||
|
|
3. **Did quality hold?** `make all` green with the new gates; and the
|
|||
|
|
honest question this pass raises — did widening the instruments surface
|
|||
|
|
*new* defects, or only the seven already known? Finding none would be
|
|||
|
|
evidence the sweep was as complete as CB-RES-0004 claims, which is a
|
|||
|
|
stronger claim than it sounds and should be stated as such.
|
|||
|
|
|
|||
|
|
Normalize per unit of work, as CB-WP-0004 T05 established: report share of
|
|||
|
|
pass alongside absolute figures, and use `cb-cost --since` to window this
|
|||
|
|
pass against the last.
|
|||
|
|
|
|||
|
|
## Task: retrospective and InnerLoop v1.4
|
|||
|
|
|
|||
|
|
```task
|
|||
|
|
id: CB-WP-0005-T08
|
|||
|
|
status: todo
|
|||
|
|
priority: medium
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
One loop change is already earned and should be written whatever else this
|
|||
|
|
pass finds. InnerLoop §Step 2's "numbers" row says the reviewer must
|
|||
|
|
*reproduce the number independently* — satisfiable by re-running the
|
|||
|
|
command that prints it, which is exactly what does **not** find this class.
|
|||
|
|
The reviewer found AM-7 by opening a test out of curiosity, and said so.
|
|||
|
|
|
|||
|
|
> **v1.4:** the reviewer must read the assertion behind every quoted
|
|||
|
|
> acceptance number and **mutate it**, not re-run the command that prints
|
|||
|
|
> it.
|
|||
|
|
|
|||
|
|
This is the **second** instance of a verification step inheriting the
|
|||
|
|
author's blindness — CB-WP-0002's dedup blind spot was the first. Both
|
|||
|
|
fixes replace re-derivation with adversarial execution. Record whether
|
|||
|
|
that generalizes.
|
|||
|
|
|
|||
|
|
The question to answer honestly: **CB-WP-0004 concluded that tooling
|
|||
|
|
recovers capacity only where it removes the manual path. Does the same
|
|||
|
|
test predict which gates work?** `env-test` and `task-done` removed the
|
|||
|
|
manual path and held. Does a mutation gate remove the manual path — or is
|
|||
|
|
writing a weak mutation the new `grep`?
|