# 2026-08-01 — retrospective: is a weak mutation the new grep? CB-WP-0006 T09. The pass instrumented AM-2, AM-6, AM-9 and AM-11, implemented K10, K11 and K18, amended K14, withdrew AM-4c, and measured itself in [CB-EV-0005](../evidence/CB-EV-0005-instrument-the-table.md). ## The question this task inherited > **M-D1-MUT does not remove a manual path — writing a weak mutation is > exactly as easy as writing a strong one, and the harness cannot tell the > difference. Is it a real instrument, or a name-counter with extra > steps?** **It is a real instrument — but only because it was hardened three times in one pass, and it would have been a name-counter without that.** ## What it took to make one mutation trustworthy The naive idea — *invert the property, require red* — is not enough. Five controls now stand between a mutation and a `red` verdict, and **every one of them exists because its failure actually occurred**: | control | verdict | the failure that motivated it | |---|---|---| | the mutation must apply | `HARNESS-BROKEN` | a stale find-string would score the baseline as the mutant | | the baseline must be green first | `inconclusive` | "mutant red" proves nothing if it was already red | | the tree must be restored, and verified | — | every later row runs against a corrupted tree | | the failure must match a **stated reason** | `WRONG-REASON` | a mutation that merely fails to compile would credit the row | | the stated reason must be **absent from passing output** | `EXPECT-VACUOUS` | my first AM-2 `expect` was `"AM-2"`, which the passing report contains | That last one is the sharpest. It means the FA guard needed a guard, and the failure it prevents is one I committed on the guard's second use. > **The generalizable finding: an instrument that measures whether other > instruments work needs more controls than the instruments it measures.** > M-D1-MUT carries five; `dep-weight` and `rule-coverage` carry one each. > That asymmetry is not overhead, it is the cost of a meta-instrument, and > a project that adds one should budget for it. ## The failure mode that is still open CB-WP-0005 predicted a weak mutation would be *written* weak. This pass found a worse case: **a mutation can become weak without anyone touching it.** AM-6's 4,000 `black_box` iterations were calibrated against debug's 3.4× headroom. When T04 moved the gate to release — correctly, because the debug measurement was invalid — the same mutation became invisible against 20× headroom and the row went `SURVIVED`. Nothing about the row, the mutation, or the code changed. The harness caught it, so the loop is not blind here. But it means **mutation strength is coupled to measurement conditions**, and a mutation is not a write-once artifact. That is a standing maintenance cost, and it is the honest answer to "is this cheaper than it looks": no. **CB-WP-0005 T08's stronger remedy — mutations written by someone other than the author of the assertion — was not tested.** The adversarial reviewer wrote none of these. `EXPECT-VACUOUS` covers the cheap failure (a guard that matches everything); it does **not** cover the expensive one (a mutation too weak to trip a real threshold), which is exactly what AM-6 demonstrated. That remedy remains untested and should not be assumed unnecessary. ## Did the "removes the manual path" test predict which gates work? No, and this pass is the counter-example that settles it. `mutation-check` fails that test outright — `grep` is still one keystroke away, and a weak mutation is as easy to write as a strong one. It also produced **six defects nothing else would have found**: AM-6 measuring contention, AM-5's 61% measurement error, RSS over-reported 3×, K10's first round trip not reproducing, the bench workload existing twice, and two spec copies going stale. CB-WP-0004 T06 said tooling recovers capacity only where it removes the manual path. **That is a predictor of whether a gate saves money, not of whether it is worth having** — a correction first stated in CB-WP-0005 T08 and now supported by a full pass. `mutation-check` costs money and earns its place on findings. ## The cost result refutes the pass **Mechanical share rose to 50%** — the highest recorded, above the 38% baseline that motivated CB-WP-0004. Mean context **493,486** against a 200,000 target, up from 315,170. This is not a tooling regression, and the distinction matters: environment setup and task closes are still at **zero**, two passes on. It is the *other* half of T06's mechanism arriving in force. Text patching (45 turns, $13.79) and orientation (19 turns, $10.97) never had their manual path removed, and a code-heavy pass is precisely where that spends. So the mechanism holds in both directions, which is the strongest evidence it is real: what was removed stayed removed; what was merely improved stayed expensive, and got worse under load. ## Prediction discipline, third outing | pass | prediction | measured | error | |---|---|---|---| | CB-WP-0004 | 25–30 points recovered | 6 | 4–5× | | CB-WP-0005 | ≥10 of 14 rows | 4 | 2.5× | | **CB-WP-0006** | **9 per-task outcomes** | **7 met, 2 restated with an argument** | **small** | The first two were aggregate predictions with a confidence label. This pass predicted **per task, as a mechanism, with the alternative named** — "instrument it, *or* restate it with an argument" — and the error collapsed. That is CB-WP-0004 T06's recommendation working. It also shows why: a per-task prediction with a named alternative cannot be dodged, because *both* branches are outcomes someone has to defend. AM-3 and AM-4c both took the second branch, and both are better resolved than if a number had been forced. ## What the loop should change **Nothing in InnerLoop this pass.** v1.4's mutation requirement was added by CB-WP-0005 and this pass is its first full exercise; changing it again before a second use would be the invention-in-isolation INTENT warns about, and the same argument used to amend K14 four hours ago. **The named next candidate is `specs/SessionShape.md`.** SS-01…SS-05 have been stated since CB-WP-0003 and **none has ever been enforced** — they are the only acceptance-adjacent numbers in this project with no gate at all. Measured this pass: mean context 493,486 (target 200,000), p90 594,888 (target 300,000), batching well under target. A pass whose entire subject was ungated numbers ran at 2.5× the context ceiling its own spec sets, and nothing said a word. ## Open, not closed - **AM-3 blocked** on an artifact that has never been built — a minimal synthetic game on our kernel. Needs a deliverable, not a metric tweak. - **AM-7 and AM-8 remain PARTIAL**: no code computes AM-7's throughput ratio, and nothing enforces AM-8's N=10. - **`CommitWindow` is provisional with a delete-by date of 2026-12-31.** - **The kernel gate binds 2026-08-31** — 18/18 today, so it binds green. - **SessionShape is ungated**, and is now the largest measured gap. - **The chaos roll is at declaration 3 of 12** (T04 rolled d4=2).