clay-borg/history/260801-instrument-the-table-retrospective.md
tegwick 039e4191dd CB-WP-0006 T09: retrospective — a real instrument, hardened three times
The question was whether M-D1-MUT is a real instrument or a name-counter
with extra steps, given that writing a weak mutation is as easy as writing
a strong one.

It is real, but only because it was hardened three times in one pass. Five
controls now stand between a mutation and a red verdict — the mutation
must apply, the baseline must be green, the tree must be restored and
verified, the failure must match a stated reason, and that stated reason
must be absent from passing output — and every one of them exists because
its failure actually occurred. The last is the sharpest: the FA guard
needed a guard, because my first AM-2 expect was "AM-2", which the passing
report contains.

Generalizable: an instrument that measures whether other instruments work
needs more controls than the instruments it measures. M-D1-MUT carries
five; dep-weight and rule-coverage carry one each. That asymmetry is the
cost of a meta-instrument, and a project adding one should budget for it.

A worse failure mode than CB-WP-0005 predicted: a mutation can become weak
without anyone touching it. AM-6's went SURVIVED when T04 moved its gate
from debug to release — nothing about the row, the mutation or the code
changed, only the headroom. Mutation strength is coupled to measurement
conditions, so a mutation is not a write-once artifact. CB-WP-0005 T08's
stronger remedy — mutations written by someone other than the author — was
NOT tested and should not be assumed unnecessary: EXPECT-VACUOUS covers
the cheap failure, not the expensive one AM-6 demonstrated.

The "removes the manual path" test is settled as a predictor of cost, not
of worth. mutation-check fails it outright and produced six defects
nothing else would have found.

Prediction error collapsed: 4-5x, then 2.5x, now small — because this pass
predicted per task, as a mechanism, with the alternative named. Both
branches are outcomes someone must defend, so the prediction cannot be
dodged. AM-3 and AM-4c took the second branch and are better resolved for
it than if a number had been forced.

No InnerLoop change. v1.4's mutation requirement is one pass old and
changing it before a second use would be the invention-in-isolation INTENT
warns about — the same argument used to amend K14 four hours earlier.

Named next candidate: specs/SessionShape.md. SS-01..SS-05 have been stated
since CB-WP-0003 and none has ever been enforced. This pass ran at 2.5x
the context ceiling its own spec sets and nothing said a word.

CB-WP-0006 status -> done, 9/9.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 13:40:58 +02:00

7 KiB
Raw Blame History

2026-08-01 — retrospective: is a weak mutation the new grep?

CB-WP-0006 T09. The pass instrumented AM-2, AM-6, AM-9 and AM-11, implemented K10, K11 and K18, amended K14, withdrew AM-4c, and measured itself in CB-EV-0005.

The question this task inherited

M-D1-MUT does not remove a manual path — writing a weak mutation is exactly as easy as writing a strong one, and the harness cannot tell the difference. Is it a real instrument, or a name-counter with extra steps?

It is a real instrument — but only because it was hardened three times in one pass, and it would have been a name-counter without that.

What it took to make one mutation trustworthy

The naive idea — invert the property, require red — is not enough. Five controls now stand between a mutation and a red verdict, and every one of them exists because its failure actually occurred:

control verdict the failure that motivated it
the mutation must apply HARNESS-BROKEN a stale find-string would score the baseline as the mutant
the baseline must be green first inconclusive "mutant red" proves nothing if it was already red
the tree must be restored, and verified every later row runs against a corrupted tree
the failure must match a stated reason WRONG-REASON a mutation that merely fails to compile would credit the row
the stated reason must be absent from passing output EXPECT-VACUOUS my first AM-2 expect was "AM-2", which the passing report contains

That last one is the sharpest. It means the FA guard needed a guard, and the failure it prevents is one I committed on the guard's second use.

The generalizable finding: an instrument that measures whether other instruments work needs more controls than the instruments it measures. M-D1-MUT carries five; dep-weight and rule-coverage carry one each. That asymmetry is not overhead, it is the cost of a meta-instrument, and a project that adds one should budget for it.

The failure mode that is still open

CB-WP-0005 predicted a weak mutation would be written weak. This pass found a worse case: a mutation can become weak without anyone touching it.

AM-6's 4,000 black_box iterations were calibrated against debug's 3.4× headroom. When T04 moved the gate to release — correctly, because the debug measurement was invalid — the same mutation became invisible against 20× headroom and the row went SURVIVED. Nothing about the row, the mutation, or the code changed.

The harness caught it, so the loop is not blind here. But it means mutation strength is coupled to measurement conditions, and a mutation is not a write-once artifact. That is a standing maintenance cost, and it is the honest answer to "is this cheaper than it looks": no.

CB-WP-0005 T08's stronger remedy — mutations written by someone other than the author of the assertion — was not tested. The adversarial reviewer wrote none of these. EXPECT-VACUOUS covers the cheap failure (a guard that matches everything); it does not cover the expensive one (a mutation too weak to trip a real threshold), which is exactly what AM-6 demonstrated. That remedy remains untested and should not be assumed unnecessary.

Did the "removes the manual path" test predict which gates work?

No, and this pass is the counter-example that settles it.

mutation-check fails that test outright — grep is still one keystroke away, and a weak mutation is as easy to write as a strong one. It also produced six defects nothing else would have found: AM-6 measuring contention, AM-5's 61% measurement error, RSS over-reported 3×, K10's first round trip not reproducing, the bench workload existing twice, and two spec copies going stale.

CB-WP-0004 T06 said tooling recovers capacity only where it removes the manual path. That is a predictor of whether a gate saves money, not of whether it is worth having — a correction first stated in CB-WP-0005 T08 and now supported by a full pass. mutation-check costs money and earns its place on findings.

The cost result refutes the pass

Mechanical share rose to 50% — the highest recorded, above the 38% baseline that motivated CB-WP-0004. Mean context 493,486 against a 200,000 target, up from 315,170.

This is not a tooling regression, and the distinction matters: environment setup and task closes are still at zero, two passes on. It is the other half of T06's mechanism arriving in force. Text patching (45 turns, $13.79) and orientation (19 turns, $10.97) never had their manual path removed, and a code-heavy pass is precisely where that spends.

So the mechanism holds in both directions, which is the strongest evidence it is real: what was removed stayed removed; what was merely improved stayed expensive, and got worse under load.

Prediction discipline, third outing

pass prediction measured error
CB-WP-0004 2530 points recovered 6 45×
CB-WP-0005 ≥10 of 14 rows 4 2.5×
CB-WP-0006 9 per-task outcomes 7 met, 2 restated with an argument small

The first two were aggregate predictions with a confidence label. This pass predicted per task, as a mechanism, with the alternative named — "instrument it, or restate it with an argument" — and the error collapsed.

That is CB-WP-0004 T06's recommendation working. It also shows why: a per-task prediction with a named alternative cannot be dodged, because both branches are outcomes someone has to defend. AM-3 and AM-4c both took the second branch, and both are better resolved than if a number had been forced.

What the loop should change

Nothing in InnerLoop this pass. v1.4's mutation requirement was added by CB-WP-0005 and this pass is its first full exercise; changing it again before a second use would be the invention-in-isolation INTENT warns about, and the same argument used to amend K14 four hours ago.

The named next candidate is specs/SessionShape.md. SS-01…SS-05 have been stated since CB-WP-0003 and none has ever been enforced — they are the only acceptance-adjacent numbers in this project with no gate at all. Measured this pass: mean context 493,486 (target 200,000), p90 594,888 (target 300,000), batching well under target. A pass whose entire subject was ungated numbers ran at 2.5× the context ceiling its own spec sets, and nothing said a word.

Open, not closed

  • AM-3 blocked on an artifact that has never been built — a minimal synthetic game on our kernel. Needs a deliverable, not a metric tweak.
  • AM-7 and AM-8 remain PARTIAL: no code computes AM-7's throughput ratio, and nothing enforces AM-8's N=10.
  • CommitWindow is provisional with a delete-by date of 2026-12-31.
  • The kernel gate binds 2026-08-31 — 18/18 today, so it binds green.
  • SessionShape is ungated, and is now the largest measured gap.
  • The chaos roll is at declaration 3 of 12 (T04 rolled d4=2).