CB-WP-0006 T09: retrospective — a real instrument, hardened three times

The question was whether M-D1-MUT is a real instrument or a name-counter
with extra steps, given that writing a weak mutation is as easy as writing
a strong one.

It is real, but only because it was hardened three times in one pass. Five
controls now stand between a mutation and a red verdict — the mutation
must apply, the baseline must be green, the tree must be restored and
verified, the failure must match a stated reason, and that stated reason
must be absent from passing output — and every one of them exists because
its failure actually occurred. The last is the sharpest: the FA guard
needed a guard, because my first AM-2 expect was "AM-2", which the passing
report contains.

Generalizable: an instrument that measures whether other instruments work
needs more controls than the instruments it measures. M-D1-MUT carries
five; dep-weight and rule-coverage carry one each. That asymmetry is the
cost of a meta-instrument, and a project adding one should budget for it.

A worse failure mode than CB-WP-0005 predicted: a mutation can become weak
without anyone touching it. AM-6's went SURVIVED when T04 moved its gate
from debug to release — nothing about the row, the mutation or the code
changed, only the headroom. Mutation strength is coupled to measurement
conditions, so a mutation is not a write-once artifact. CB-WP-0005 T08's
stronger remedy — mutations written by someone other than the author — was
NOT tested and should not be assumed unnecessary: EXPECT-VACUOUS covers
the cheap failure, not the expensive one AM-6 demonstrated.

The "removes the manual path" test is settled as a predictor of cost, not
of worth. mutation-check fails it outright and produced six defects
nothing else would have found.

Prediction error collapsed: 4-5x, then 2.5x, now small — because this pass
predicted per task, as a mechanism, with the alternative named. Both
branches are outcomes someone must defend, so the prediction cannot be
dodged. AM-3 and AM-4c took the second branch and are better resolved for
it than if a number had been forced.

No InnerLoop change. v1.4's mutation requirement is one pass old and
changing it before a second use would be the invention-in-isolation INTENT
warns about — the same argument used to amend K14 four hours earlier.

Named next candidate: specs/SessionShape.md. SS-01..SS-05 have been stated
since CB-WP-0003 and none has ever been enforced. This pass ran at 2.5x
the context ceiling its own spec sets and nothing said a word.

CB-WP-0006 status -> done, 9/9.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
tegwick 2026-08-01 13:40:58 +02:00
parent ce353adde8
commit 039e4191dd
2 changed files with 183 additions and 3 deletions

View file

@ -0,0 +1,141 @@
# 2026-08-01 — retrospective: is a weak mutation the new grep?
CB-WP-0006 T09. The pass instrumented AM-2, AM-6, AM-9 and AM-11,
implemented K10, K11 and K18, amended K14, withdrew AM-4c, and measured
itself in [CB-EV-0005](../evidence/CB-EV-0005-instrument-the-table.md).
## The question this task inherited
> **M-D1-MUT does not remove a manual path — writing a weak mutation is
> exactly as easy as writing a strong one, and the harness cannot tell the
> difference. Is it a real instrument, or a name-counter with extra
> steps?**
**It is a real instrument — but only because it was hardened three times
in one pass, and it would have been a name-counter without that.**
## What it took to make one mutation trustworthy
The naive idea — *invert the property, require red* — is not enough. Five
controls now stand between a mutation and a `red` verdict, and **every one
of them exists because its failure actually occurred**:
| control | verdict | the failure that motivated it |
|---|---|---|
| the mutation must apply | `HARNESS-BROKEN` | a stale find-string would score the baseline as the mutant |
| the baseline must be green first | `inconclusive` | "mutant red" proves nothing if it was already red |
| the tree must be restored, and verified | — | every later row runs against a corrupted tree |
| the failure must match a **stated reason** | `WRONG-REASON` | a mutation that merely fails to compile would credit the row |
| the stated reason must be **absent from passing output** | `EXPECT-VACUOUS` | my first AM-2 `expect` was `"AM-2"`, which the passing report contains |
That last one is the sharpest. It means the FA guard needed a guard, and
the failure it prevents is one I committed on the guard's second use.
> **The generalizable finding: an instrument that measures whether other
> instruments work needs more controls than the instruments it measures.**
> M-D1-MUT carries five; `dep-weight` and `rule-coverage` carry one each.
> That asymmetry is not overhead, it is the cost of a meta-instrument, and
> a project that adds one should budget for it.
## The failure mode that is still open
CB-WP-0005 predicted a weak mutation would be *written* weak. This pass
found a worse case: **a mutation can become weak without anyone touching
it.**
AM-6's 4,000 `black_box` iterations were calibrated against debug's 3.4×
headroom. When T04 moved the gate to release — correctly, because the
debug measurement was invalid — the same mutation became invisible against
20× headroom and the row went `SURVIVED`. Nothing about the row, the
mutation, or the code changed.
The harness caught it, so the loop is not blind here. But it means
**mutation strength is coupled to measurement conditions**, and a mutation
is not a write-once artifact. That is a standing maintenance cost, and it
is the honest answer to "is this cheaper than it looks": no.
**CB-WP-0005 T08's stronger remedy — mutations written by someone other
than the author of the assertion — was not tested.** The adversarial
reviewer wrote none of these. `EXPECT-VACUOUS` covers the cheap failure
(a guard that matches everything); it does **not** cover the expensive one
(a mutation too weak to trip a real threshold), which is exactly what
AM-6 demonstrated. That remedy remains untested and should not be assumed
unnecessary.
## Did the "removes the manual path" test predict which gates work?
No, and this pass is the counter-example that settles it.
`mutation-check` fails that test outright — `grep` is still one keystroke
away, and a weak mutation is as easy to write as a strong one. It also
produced **six defects nothing else would have found**: AM-6 measuring
contention, AM-5's 61% measurement error, RSS over-reported 3×, K10's
first round trip not reproducing, the bench workload existing twice, and
two spec copies going stale.
CB-WP-0004 T06 said tooling recovers capacity only where it removes the
manual path. **That is a predictor of whether a gate saves money, not of
whether it is worth having** — a correction first stated in CB-WP-0005 T08
and now supported by a full pass. `mutation-check` costs money and earns
its place on findings.
## The cost result refutes the pass
**Mechanical share rose to 50%** — the highest recorded, above the 38%
baseline that motivated CB-WP-0004. Mean context **493,486** against a
200,000 target, up from 315,170.
This is not a tooling regression, and the distinction matters: environment
setup and task closes are still at **zero**, two passes on. It is the
*other* half of T06's mechanism arriving in force. Text patching (45
turns, $13.79) and orientation (19 turns, $10.97) never had their manual
path removed, and a code-heavy pass is precisely where that spends.
So the mechanism holds in both directions, which is the strongest evidence
it is real: what was removed stayed removed; what was merely improved
stayed expensive, and got worse under load.
## Prediction discipline, third outing
| pass | prediction | measured | error |
|---|---|---|---|
| CB-WP-0004 | 2530 points recovered | 6 | 45× |
| CB-WP-0005 | ≥10 of 14 rows | 4 | 2.5× |
| **CB-WP-0006** | **9 per-task outcomes** | **7 met, 2 restated with an argument** | **small** |
The first two were aggregate predictions with a confidence label. This
pass predicted **per task, as a mechanism, with the alternative named**
"instrument it, *or* restate it with an argument" — and the error
collapsed.
That is CB-WP-0004 T06's recommendation working. It also shows why:
a per-task prediction with a named alternative cannot be dodged, because
*both* branches are outcomes someone has to defend. AM-3 and AM-4c both
took the second branch, and both are better resolved than if a number had
been forced.
## What the loop should change
**Nothing in InnerLoop this pass.** v1.4's mutation requirement was added
by CB-WP-0005 and this pass is its first full exercise; changing it again
before a second use would be the invention-in-isolation INTENT warns
about, and the same argument used to amend K14 four hours ago.
**The named next candidate is `specs/SessionShape.md`.** SS-01…SS-05 have
been stated since CB-WP-0003 and **none has ever been enforced** — they
are the only acceptance-adjacent numbers in this project with no gate at
all. Measured this pass: mean context 493,486 (target 200,000), p90
594,888 (target 300,000), batching well under target. A pass whose entire
subject was ungated numbers ran at 2.5× the context ceiling its own spec
sets, and nothing said a word.
## Open, not closed
- **AM-3 blocked** on an artifact that has never been built — a minimal
synthetic game on our kernel. Needs a deliverable, not a metric tweak.
- **AM-7 and AM-8 remain PARTIAL**: no code computes AM-7's throughput
ratio, and nothing enforces AM-8's N=10.
- **`CommitWindow` is provisional with a delete-by date of 2026-12-31.**
- **The kernel gate binds 2026-08-31** — 18/18 today, so it binds green.
- **SessionShape is ungated**, and is now the largest measured gap.
- **The chaos roll is at declaration 3 of 12** (T04 rolled d4=2).

View file

@ -1,7 +1,7 @@
--- ---
id: CB-WP-0006 id: CB-WP-0006
title: "Instrument the acceptance table, then implement what it exposes" title: "Instrument the acceptance table, then implement what it exposes"
status: in_progress status: done
state_hub_workstream_id: "8a6327cc-fd5c-4e2c-a29b-b437c27d1e71" state_hub_workstream_id: "8a6327cc-fd5c-4e2c-a29b-b437c27d1e71"
--- ---
@ -222,14 +222,23 @@ right for K18.
```task ```task
id: CB-WP-0006-T08 id: CB-WP-0006-T08
status: todo status: done
priority: high priority: high
state_hub_task_id: "1d4455c3-c642-47d8-916b-ab35bc512207" state_hub_task_id: "1d4455c3-c642-47d8-916b-ab35bc512207"
``` ```
Commit `evidence/CB-EV-0004`. The baseline is **4 of 14**, committed and Commit `evidence/CB-EV-0005`**numbering corrected**, CB-EV-0004 was
already CB-WP-0005's. The baseline is **4 of 14**, committed and
reproducible via `make mutation-check`. reproducible via `make mutation-check`.
**Measured — [CB-EV-0005](../evidence/CB-EV-0005-instrument-the-table.md).
4 of 14 → 8 of 14** (8 of 10 enforceable; the 14 stays the headline and
AM-4c stays in it on purpose). Kernel link 15/18 → **18/18**, names only.
One row regressed to `SURVIVED` and was caught. **0 vacuous expects**
across 14 rows — test 3 is now mechanical, via a new `EXPECT-VACUOUS`
verdict. **The cost result refutes the pass:** mechanical share rose to
**50%**, the highest recorded.
Three tests, all reported: Three tests, all reported:
1. **Did the enforced count rise, per row?** Against the per-task 1. **Did the enforced count rise, per row?** Against the per-task
@ -268,3 +277,33 @@ the latter, say so and propose what would actually bind — the candidate
being that the mutation must be written by someone other than the author being that the mutation must be written by someone other than the author
of the assertion, which is the adversarial-review principle applied one of the assertion, which is the adversarial-review principle applied one
level down. level down.
**Answered — [260801-instrument-the-table-retrospective.md](../history/260801-instrument-the-table-retrospective.md).**
**A real instrument — but only because it was hardened three times in one
pass.** Five controls now stand between a mutation and a `red` verdict,
and every one exists because its failure actually occurred. The
generalizable finding: **an instrument that measures whether other
instruments work needs more controls than the instruments it measures** —
M-D1-MUT carries five, `dep-weight` and `rule-coverage` carry one each.
**A worse failure mode than predicted:** a mutation can become weak
**without anyone touching it**. AM-6's went `SURVIVED` when T04 moved its
gate from debug to release. Mutation strength is coupled to measurement
conditions, so a mutation is not a write-once artifact. CB-WP-0005 T08's
stronger remedy — mutations written by someone other than the author —
was **not tested** and should not be assumed unnecessary: `EXPECT-VACUOUS`
covers the cheap failure, not the expensive one AM-6 demonstrated.
**Prediction error collapsed** — 45×, then 2.5×, now small — because this
pass predicted **per task, as a mechanism, with the alternative named**.
Both branches are outcomes someone must defend, so the prediction cannot
be dodged; AM-3 and AM-4c took the second branch and are better resolved
for it.
**No InnerLoop change.** v1.4's mutation requirement is one pass old and
changing it before a second use would be the invention-in-isolation INTENT
warns about — the same argument used to amend K14. **The named next
candidate is `specs/SessionShape.md`**: SS-01…SS-05 have never been
enforced, and this pass ran at **2.5× the context ceiling its own spec
sets** without anything saying a word.