CB-WP-0006 T09: retrospective — a real instrument, hardened three times

The question was whether M-D1-MUT is a real instrument or a name-counter
with extra steps, given that writing a weak mutation is as easy as writing
a strong one.

It is real, but only because it was hardened three times in one pass. Five
controls now stand between a mutation and a red verdict — the mutation
must apply, the baseline must be green, the tree must be restored and
verified, the failure must match a stated reason, and that stated reason
must be absent from passing output — and every one of them exists because
its failure actually occurred. The last is the sharpest: the FA guard
needed a guard, because my first AM-2 expect was "AM-2", which the passing
report contains.

Generalizable: an instrument that measures whether other instruments work
needs more controls than the instruments it measures. M-D1-MUT carries
five; dep-weight and rule-coverage carry one each. That asymmetry is the
cost of a meta-instrument, and a project adding one should budget for it.

A worse failure mode than CB-WP-0005 predicted: a mutation can become weak
without anyone touching it. AM-6's went SURVIVED when T04 moved its gate
from debug to release — nothing about the row, the mutation or the code
changed, only the headroom. Mutation strength is coupled to measurement
conditions, so a mutation is not a write-once artifact. CB-WP-0005 T08's
stronger remedy — mutations written by someone other than the author — was
NOT tested and should not be assumed unnecessary: EXPECT-VACUOUS covers
the cheap failure, not the expensive one AM-6 demonstrated.

The "removes the manual path" test is settled as a predictor of cost, not
of worth. mutation-check fails it outright and produced six defects
nothing else would have found.

Prediction error collapsed: 4-5x, then 2.5x, now small — because this pass
predicted per task, as a mechanism, with the alternative named. Both
branches are outcomes someone must defend, so the prediction cannot be
dodged. AM-3 and AM-4c took the second branch and are better resolved for
it than if a number had been forced.

No InnerLoop change. v1.4's mutation requirement is one pass old and
changing it before a second use would be the invention-in-isolation INTENT
warns about — the same argument used to amend K14 four hours earlier.

Named next candidate: specs/SessionShape.md. SS-01..SS-05 have been stated
since CB-WP-0003 and none has ever been enforced. This pass ran at 2.5x
the context ceiling its own spec sets and nothing said a word.

CB-WP-0006 status -> done, 9/9.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
tegwick 2026-08-01 13:40:58 +02:00
parent ce353adde8
commit 039e4191dd
2 changed files with 183 additions and 3 deletions

View file

@ -1,7 +1,7 @@
---
id: CB-WP-0006
title: "Instrument the acceptance table, then implement what it exposes"
status: in_progress
status: done
state_hub_workstream_id: "8a6327cc-fd5c-4e2c-a29b-b437c27d1e71"
---
@ -222,14 +222,23 @@ right for K18.
```task
id: CB-WP-0006-T08
status: todo
status: done
priority: high
state_hub_task_id: "1d4455c3-c642-47d8-916b-ab35bc512207"
```
Commit `evidence/CB-EV-0004`. The baseline is **4 of 14**, committed and
Commit `evidence/CB-EV-0005`**numbering corrected**, CB-EV-0004 was
already CB-WP-0005's. The baseline is **4 of 14**, committed and
reproducible via `make mutation-check`.
**Measured — [CB-EV-0005](../evidence/CB-EV-0005-instrument-the-table.md).
4 of 14 → 8 of 14** (8 of 10 enforceable; the 14 stays the headline and
AM-4c stays in it on purpose). Kernel link 15/18 → **18/18**, names only.
One row regressed to `SURVIVED` and was caught. **0 vacuous expects**
across 14 rows — test 3 is now mechanical, via a new `EXPECT-VACUOUS`
verdict. **The cost result refutes the pass:** mechanical share rose to
**50%**, the highest recorded.
Three tests, all reported:
1. **Did the enforced count rise, per row?** Against the per-task
@ -268,3 +277,33 @@ the latter, say so and propose what would actually bind — the candidate
being that the mutation must be written by someone other than the author
of the assertion, which is the adversarial-review principle applied one
level down.
**Answered — [260801-instrument-the-table-retrospective.md](../history/260801-instrument-the-table-retrospective.md).**
**A real instrument — but only because it was hardened three times in one
pass.** Five controls now stand between a mutation and a `red` verdict,
and every one exists because its failure actually occurred. The
generalizable finding: **an instrument that measures whether other
instruments work needs more controls than the instruments it measures** —
M-D1-MUT carries five, `dep-weight` and `rule-coverage` carry one each.
**A worse failure mode than predicted:** a mutation can become weak
**without anyone touching it**. AM-6's went `SURVIVED` when T04 moved its
gate from debug to release. Mutation strength is coupled to measurement
conditions, so a mutation is not a write-once artifact. CB-WP-0005 T08's
stronger remedy — mutations written by someone other than the author —
was **not tested** and should not be assumed unnecessary: `EXPECT-VACUOUS`
covers the cheap failure, not the expensive one AM-6 demonstrated.
**Prediction error collapsed** — 45×, then 2.5×, now small — because this
pass predicted **per task, as a mechanism, with the alternative named**.
Both branches are outcomes someone must defend, so the prediction cannot
be dodged; AM-3 and AM-4c took the second branch and are better resolved
for it.
**No InnerLoop change.** v1.4's mutation requirement is one pass old and
changing it before a second use would be the invention-in-isolation INTENT
warns about — the same argument used to amend K14. **The named next
candidate is `specs/SessionShape.md`**: SS-01…SS-05 have never been
enforced, and this pass ran at **2.5× the context ceiling its own spec
sets** without anything saying a word.