The question was whether M-D1-MUT is a real instrument or a name-counter with extra steps, given that writing a weak mutation is as easy as writing a strong one. It is real, but only because it was hardened three times in one pass. Five controls now stand between a mutation and a red verdict — the mutation must apply, the baseline must be green, the tree must be restored and verified, the failure must match a stated reason, and that stated reason must be absent from passing output — and every one of them exists because its failure actually occurred. The last is the sharpest: the FA guard needed a guard, because my first AM-2 expect was "AM-2", which the passing report contains. Generalizable: an instrument that measures whether other instruments work needs more controls than the instruments it measures. M-D1-MUT carries five; dep-weight and rule-coverage carry one each. That asymmetry is the cost of a meta-instrument, and a project adding one should budget for it. A worse failure mode than CB-WP-0005 predicted: a mutation can become weak without anyone touching it. AM-6's went SURVIVED when T04 moved its gate from debug to release — nothing about the row, the mutation or the code changed, only the headroom. Mutation strength is coupled to measurement conditions, so a mutation is not a write-once artifact. CB-WP-0005 T08's stronger remedy — mutations written by someone other than the author — was NOT tested and should not be assumed unnecessary: EXPECT-VACUOUS covers the cheap failure, not the expensive one AM-6 demonstrated. The "removes the manual path" test is settled as a predictor of cost, not of worth. mutation-check fails it outright and produced six defects nothing else would have found. Prediction error collapsed: 4-5x, then 2.5x, now small — because this pass predicted per task, as a mechanism, with the alternative named. Both branches are outcomes someone must defend, so the prediction cannot be dodged. AM-3 and AM-4c took the second branch and are better resolved for it than if a number had been forced. No InnerLoop change. v1.4's mutation requirement is one pass old and changing it before a second use would be the invention-in-isolation INTENT warns about — the same argument used to amend K14 four hours earlier. Named next candidate: specs/SessionShape.md. SS-01..SS-05 have been stated since CB-WP-0003 and none has ever been enforced. This pass ran at 2.5x the context ceiling its own spec sets and nothing said a word. CB-WP-0006 status -> done, 9/9. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
7 KiB
2026-08-01 — retrospective: is a weak mutation the new grep?
CB-WP-0006 T09. The pass instrumented AM-2, AM-6, AM-9 and AM-11, implemented K10, K11 and K18, amended K14, withdrew AM-4c, and measured itself in CB-EV-0005.
The question this task inherited
M-D1-MUT does not remove a manual path — writing a weak mutation is exactly as easy as writing a strong one, and the harness cannot tell the difference. Is it a real instrument, or a name-counter with extra steps?
It is a real instrument — but only because it was hardened three times in one pass, and it would have been a name-counter without that.
What it took to make one mutation trustworthy
The naive idea — invert the property, require red — is not enough. Five
controls now stand between a mutation and a red verdict, and every one
of them exists because its failure actually occurred:
| control | verdict | the failure that motivated it |
|---|---|---|
| the mutation must apply | HARNESS-BROKEN |
a stale find-string would score the baseline as the mutant |
| the baseline must be green first | inconclusive |
"mutant red" proves nothing if it was already red |
| the tree must be restored, and verified | — | every later row runs against a corrupted tree |
| the failure must match a stated reason | WRONG-REASON |
a mutation that merely fails to compile would credit the row |
| the stated reason must be absent from passing output | EXPECT-VACUOUS |
my first AM-2 expect was "AM-2", which the passing report contains |
That last one is the sharpest. It means the FA guard needed a guard, and the failure it prevents is one I committed on the guard's second use.
The generalizable finding: an instrument that measures whether other instruments work needs more controls than the instruments it measures. M-D1-MUT carries five;
dep-weightandrule-coveragecarry one each. That asymmetry is not overhead, it is the cost of a meta-instrument, and a project that adds one should budget for it.
The failure mode that is still open
CB-WP-0005 predicted a weak mutation would be written weak. This pass found a worse case: a mutation can become weak without anyone touching it.
AM-6's 4,000 black_box iterations were calibrated against debug's 3.4×
headroom. When T04 moved the gate to release — correctly, because the
debug measurement was invalid — the same mutation became invisible against
20× headroom and the row went SURVIVED. Nothing about the row, the
mutation, or the code changed.
The harness caught it, so the loop is not blind here. But it means mutation strength is coupled to measurement conditions, and a mutation is not a write-once artifact. That is a standing maintenance cost, and it is the honest answer to "is this cheaper than it looks": no.
CB-WP-0005 T08's stronger remedy — mutations written by someone other
than the author of the assertion — was not tested. The adversarial
reviewer wrote none of these. EXPECT-VACUOUS covers the cheap failure
(a guard that matches everything); it does not cover the expensive one
(a mutation too weak to trip a real threshold), which is exactly what
AM-6 demonstrated. That remedy remains untested and should not be assumed
unnecessary.
Did the "removes the manual path" test predict which gates work?
No, and this pass is the counter-example that settles it.
mutation-check fails that test outright — grep is still one keystroke
away, and a weak mutation is as easy to write as a strong one. It also
produced six defects nothing else would have found: AM-6 measuring
contention, AM-5's 61% measurement error, RSS over-reported 3×, K10's
first round trip not reproducing, the bench workload existing twice, and
two spec copies going stale.
CB-WP-0004 T06 said tooling recovers capacity only where it removes the
manual path. That is a predictor of whether a gate saves money, not of
whether it is worth having — a correction first stated in CB-WP-0005 T08
and now supported by a full pass. mutation-check costs money and earns
its place on findings.
The cost result refutes the pass
Mechanical share rose to 50% — the highest recorded, above the 38% baseline that motivated CB-WP-0004. Mean context 493,486 against a 200,000 target, up from 315,170.
This is not a tooling regression, and the distinction matters: environment setup and task closes are still at zero, two passes on. It is the other half of T06's mechanism arriving in force. Text patching (45 turns, $13.79) and orientation (19 turns, $10.97) never had their manual path removed, and a code-heavy pass is precisely where that spends.
So the mechanism holds in both directions, which is the strongest evidence it is real: what was removed stayed removed; what was merely improved stayed expensive, and got worse under load.
Prediction discipline, third outing
| pass | prediction | measured | error |
|---|---|---|---|
| CB-WP-0004 | 25–30 points recovered | 6 | 4–5× |
| CB-WP-0005 | ≥10 of 14 rows | 4 | 2.5× |
| CB-WP-0006 | 9 per-task outcomes | 7 met, 2 restated with an argument | small |
The first two were aggregate predictions with a confidence label. This pass predicted per task, as a mechanism, with the alternative named — "instrument it, or restate it with an argument" — and the error collapsed.
That is CB-WP-0004 T06's recommendation working. It also shows why: a per-task prediction with a named alternative cannot be dodged, because both branches are outcomes someone has to defend. AM-3 and AM-4c both took the second branch, and both are better resolved than if a number had been forced.
What the loop should change
Nothing in InnerLoop this pass. v1.4's mutation requirement was added by CB-WP-0005 and this pass is its first full exercise; changing it again before a second use would be the invention-in-isolation INTENT warns about, and the same argument used to amend K14 four hours ago.
The named next candidate is specs/SessionShape.md. SS-01…SS-05 have
been stated since CB-WP-0003 and none has ever been enforced — they
are the only acceptance-adjacent numbers in this project with no gate at
all. Measured this pass: mean context 493,486 (target 200,000), p90
594,888 (target 300,000), batching well under target. A pass whose entire
subject was ungated numbers ran at 2.5× the context ceiling its own spec
sets, and nothing said a word.
Open, not closed
- AM-3 blocked on an artifact that has never been built — a minimal synthetic game on our kernel. Needs a deliverable, not a metric tweak.
- AM-7 and AM-8 remain PARTIAL: no code computes AM-7's throughput ratio, and nothing enforces AM-8's N=10.
CommitWindowis provisional with a delete-by date of 2026-12-31.- The kernel gate binds 2026-08-31 — 18/18 today, so it binds green.
- SessionShape is ungated, and is now the largest measured gap.
- The chaos roll is at declaration 3 of 12 (T04 rolled d4=2).