CB-WP-0006 T09: retrospective — a real instrument, hardened three times
The question was whether M-D1-MUT is a real instrument or a name-counter with extra steps, given that writing a weak mutation is as easy as writing a strong one. It is real, but only because it was hardened three times in one pass. Five controls now stand between a mutation and a red verdict — the mutation must apply, the baseline must be green, the tree must be restored and verified, the failure must match a stated reason, and that stated reason must be absent from passing output — and every one of them exists because its failure actually occurred. The last is the sharpest: the FA guard needed a guard, because my first AM-2 expect was "AM-2", which the passing report contains. Generalizable: an instrument that measures whether other instruments work needs more controls than the instruments it measures. M-D1-MUT carries five; dep-weight and rule-coverage carry one each. That asymmetry is the cost of a meta-instrument, and a project adding one should budget for it. A worse failure mode than CB-WP-0005 predicted: a mutation can become weak without anyone touching it. AM-6's went SURVIVED when T04 moved its gate from debug to release — nothing about the row, the mutation or the code changed, only the headroom. Mutation strength is coupled to measurement conditions, so a mutation is not a write-once artifact. CB-WP-0005 T08's stronger remedy — mutations written by someone other than the author — was NOT tested and should not be assumed unnecessary: EXPECT-VACUOUS covers the cheap failure, not the expensive one AM-6 demonstrated. The "removes the manual path" test is settled as a predictor of cost, not of worth. mutation-check fails it outright and produced six defects nothing else would have found. Prediction error collapsed: 4-5x, then 2.5x, now small — because this pass predicted per task, as a mechanism, with the alternative named. Both branches are outcomes someone must defend, so the prediction cannot be dodged. AM-3 and AM-4c took the second branch and are better resolved for it than if a number had been forced. No InnerLoop change. v1.4's mutation requirement is one pass old and changing it before a second use would be the invention-in-isolation INTENT warns about — the same argument used to amend K14 four hours earlier. Named next candidate: specs/SessionShape.md. SS-01..SS-05 have been stated since CB-WP-0003 and none has ever been enforced. This pass ran at 2.5x the context ceiling its own spec sets and nothing said a word. CB-WP-0006 status -> done, 9/9. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
ce353adde8
commit
039e4191dd
2 changed files with 183 additions and 3 deletions
141
history/260801-instrument-the-table-retrospective.md
Normal file
141
history/260801-instrument-the-table-retrospective.md
Normal file
|
|
@ -0,0 +1,141 @@
|
|||
# 2026-08-01 — retrospective: is a weak mutation the new grep?
|
||||
|
||||
CB-WP-0006 T09. The pass instrumented AM-2, AM-6, AM-9 and AM-11,
|
||||
implemented K10, K11 and K18, amended K14, withdrew AM-4c, and measured
|
||||
itself in [CB-EV-0005](../evidence/CB-EV-0005-instrument-the-table.md).
|
||||
|
||||
## The question this task inherited
|
||||
|
||||
> **M-D1-MUT does not remove a manual path — writing a weak mutation is
|
||||
> exactly as easy as writing a strong one, and the harness cannot tell the
|
||||
> difference. Is it a real instrument, or a name-counter with extra
|
||||
> steps?**
|
||||
|
||||
**It is a real instrument — but only because it was hardened three times
|
||||
in one pass, and it would have been a name-counter without that.**
|
||||
|
||||
## What it took to make one mutation trustworthy
|
||||
|
||||
The naive idea — *invert the property, require red* — is not enough. Five
|
||||
controls now stand between a mutation and a `red` verdict, and **every one
|
||||
of them exists because its failure actually occurred**:
|
||||
|
||||
| control | verdict | the failure that motivated it |
|
||||
|---|---|---|
|
||||
| the mutation must apply | `HARNESS-BROKEN` | a stale find-string would score the baseline as the mutant |
|
||||
| the baseline must be green first | `inconclusive` | "mutant red" proves nothing if it was already red |
|
||||
| the tree must be restored, and verified | — | every later row runs against a corrupted tree |
|
||||
| the failure must match a **stated reason** | `WRONG-REASON` | a mutation that merely fails to compile would credit the row |
|
||||
| the stated reason must be **absent from passing output** | `EXPECT-VACUOUS` | my first AM-2 `expect` was `"AM-2"`, which the passing report contains |
|
||||
|
||||
That last one is the sharpest. It means the FA guard needed a guard, and
|
||||
the failure it prevents is one I committed on the guard's second use.
|
||||
|
||||
> **The generalizable finding: an instrument that measures whether other
|
||||
> instruments work needs more controls than the instruments it measures.**
|
||||
> M-D1-MUT carries five; `dep-weight` and `rule-coverage` carry one each.
|
||||
> That asymmetry is not overhead, it is the cost of a meta-instrument, and
|
||||
> a project that adds one should budget for it.
|
||||
|
||||
## The failure mode that is still open
|
||||
|
||||
CB-WP-0005 predicted a weak mutation would be *written* weak. This pass
|
||||
found a worse case: **a mutation can become weak without anyone touching
|
||||
it.**
|
||||
|
||||
AM-6's 4,000 `black_box` iterations were calibrated against debug's 3.4×
|
||||
headroom. When T04 moved the gate to release — correctly, because the
|
||||
debug measurement was invalid — the same mutation became invisible against
|
||||
20× headroom and the row went `SURVIVED`. Nothing about the row, the
|
||||
mutation, or the code changed.
|
||||
|
||||
The harness caught it, so the loop is not blind here. But it means
|
||||
**mutation strength is coupled to measurement conditions**, and a mutation
|
||||
is not a write-once artifact. That is a standing maintenance cost, and it
|
||||
is the honest answer to "is this cheaper than it looks": no.
|
||||
|
||||
**CB-WP-0005 T08's stronger remedy — mutations written by someone other
|
||||
than the author of the assertion — was not tested.** The adversarial
|
||||
reviewer wrote none of these. `EXPECT-VACUOUS` covers the cheap failure
|
||||
(a guard that matches everything); it does **not** cover the expensive one
|
||||
(a mutation too weak to trip a real threshold), which is exactly what
|
||||
AM-6 demonstrated. That remedy remains untested and should not be assumed
|
||||
unnecessary.
|
||||
|
||||
## Did the "removes the manual path" test predict which gates work?
|
||||
|
||||
No, and this pass is the counter-example that settles it.
|
||||
|
||||
`mutation-check` fails that test outright — `grep` is still one keystroke
|
||||
away, and a weak mutation is as easy to write as a strong one. It also
|
||||
produced **six defects nothing else would have found**: AM-6 measuring
|
||||
contention, AM-5's 61% measurement error, RSS over-reported 3×, K10's
|
||||
first round trip not reproducing, the bench workload existing twice, and
|
||||
two spec copies going stale.
|
||||
|
||||
CB-WP-0004 T06 said tooling recovers capacity only where it removes the
|
||||
manual path. **That is a predictor of whether a gate saves money, not of
|
||||
whether it is worth having** — a correction first stated in CB-WP-0005 T08
|
||||
and now supported by a full pass. `mutation-check` costs money and earns
|
||||
its place on findings.
|
||||
|
||||
## The cost result refutes the pass
|
||||
|
||||
**Mechanical share rose to 50%** — the highest recorded, above the 38%
|
||||
baseline that motivated CB-WP-0004. Mean context **493,486** against a
|
||||
200,000 target, up from 315,170.
|
||||
|
||||
This is not a tooling regression, and the distinction matters: environment
|
||||
setup and task closes are still at **zero**, two passes on. It is the
|
||||
*other* half of T06's mechanism arriving in force. Text patching (45
|
||||
turns, $13.79) and orientation (19 turns, $10.97) never had their manual
|
||||
path removed, and a code-heavy pass is precisely where that spends.
|
||||
|
||||
So the mechanism holds in both directions, which is the strongest evidence
|
||||
it is real: what was removed stayed removed; what was merely improved
|
||||
stayed expensive, and got worse under load.
|
||||
|
||||
## Prediction discipline, third outing
|
||||
|
||||
| pass | prediction | measured | error |
|
||||
|---|---|---|---|
|
||||
| CB-WP-0004 | 25–30 points recovered | 6 | 4–5× |
|
||||
| CB-WP-0005 | ≥10 of 14 rows | 4 | 2.5× |
|
||||
| **CB-WP-0006** | **9 per-task outcomes** | **7 met, 2 restated with an argument** | **small** |
|
||||
|
||||
The first two were aggregate predictions with a confidence label. This
|
||||
pass predicted **per task, as a mechanism, with the alternative named** —
|
||||
"instrument it, *or* restate it with an argument" — and the error
|
||||
collapsed.
|
||||
|
||||
That is CB-WP-0004 T06's recommendation working. It also shows why:
|
||||
a per-task prediction with a named alternative cannot be dodged, because
|
||||
*both* branches are outcomes someone has to defend. AM-3 and AM-4c both
|
||||
took the second branch, and both are better resolved than if a number had
|
||||
been forced.
|
||||
|
||||
## What the loop should change
|
||||
|
||||
**Nothing in InnerLoop this pass.** v1.4's mutation requirement was added
|
||||
by CB-WP-0005 and this pass is its first full exercise; changing it again
|
||||
before a second use would be the invention-in-isolation INTENT warns
|
||||
about, and the same argument used to amend K14 four hours ago.
|
||||
|
||||
**The named next candidate is `specs/SessionShape.md`.** SS-01…SS-05 have
|
||||
been stated since CB-WP-0003 and **none has ever been enforced** — they
|
||||
are the only acceptance-adjacent numbers in this project with no gate at
|
||||
all. Measured this pass: mean context 493,486 (target 200,000), p90
|
||||
594,888 (target 300,000), batching well under target. A pass whose entire
|
||||
subject was ungated numbers ran at 2.5× the context ceiling its own spec
|
||||
sets, and nothing said a word.
|
||||
|
||||
## Open, not closed
|
||||
|
||||
- **AM-3 blocked** on an artifact that has never been built — a minimal
|
||||
synthetic game on our kernel. Needs a deliverable, not a metric tweak.
|
||||
- **AM-7 and AM-8 remain PARTIAL**: no code computes AM-7's throughput
|
||||
ratio, and nothing enforces AM-8's N=10.
|
||||
- **`CommitWindow` is provisional with a delete-by date of 2026-12-31.**
|
||||
- **The kernel gate binds 2026-08-31** — 18/18 today, so it binds green.
|
||||
- **SessionShape is ungated**, and is now the largest measured gap.
|
||||
- **The chaos roll is at declaration 3 of 12** (T04 rolled d4=2).
|
||||
|
|
@ -1,7 +1,7 @@
|
|||
---
|
||||
id: CB-WP-0006
|
||||
title: "Instrument the acceptance table, then implement what it exposes"
|
||||
status: in_progress
|
||||
status: done
|
||||
state_hub_workstream_id: "8a6327cc-fd5c-4e2c-a29b-b437c27d1e71"
|
||||
---
|
||||
|
||||
|
|
@ -222,14 +222,23 @@ right for K18.
|
|||
|
||||
```task
|
||||
id: CB-WP-0006-T08
|
||||
status: todo
|
||||
status: done
|
||||
priority: high
|
||||
state_hub_task_id: "1d4455c3-c642-47d8-916b-ab35bc512207"
|
||||
```
|
||||
|
||||
Commit `evidence/CB-EV-0004`. The baseline is **4 of 14**, committed and
|
||||
Commit `evidence/CB-EV-0005` — **numbering corrected**, CB-EV-0004 was
|
||||
already CB-WP-0005's. The baseline is **4 of 14**, committed and
|
||||
reproducible via `make mutation-check`.
|
||||
|
||||
**Measured — [CB-EV-0005](../evidence/CB-EV-0005-instrument-the-table.md).
|
||||
4 of 14 → 8 of 14** (8 of 10 enforceable; the 14 stays the headline and
|
||||
AM-4c stays in it on purpose). Kernel link 15/18 → **18/18**, names only.
|
||||
One row regressed to `SURVIVED` and was caught. **0 vacuous expects**
|
||||
across 14 rows — test 3 is now mechanical, via a new `EXPECT-VACUOUS`
|
||||
verdict. **The cost result refutes the pass:** mechanical share rose to
|
||||
**50%**, the highest recorded.
|
||||
|
||||
Three tests, all reported:
|
||||
|
||||
1. **Did the enforced count rise, per row?** Against the per-task
|
||||
|
|
@ -268,3 +277,33 @@ the latter, say so and propose what would actually bind — the candidate
|
|||
being that the mutation must be written by someone other than the author
|
||||
of the assertion, which is the adversarial-review principle applied one
|
||||
level down.
|
||||
|
||||
**Answered — [260801-instrument-the-table-retrospective.md](../history/260801-instrument-the-table-retrospective.md).**
|
||||
|
||||
**A real instrument — but only because it was hardened three times in one
|
||||
pass.** Five controls now stand between a mutation and a `red` verdict,
|
||||
and every one exists because its failure actually occurred. The
|
||||
generalizable finding: **an instrument that measures whether other
|
||||
instruments work needs more controls than the instruments it measures** —
|
||||
M-D1-MUT carries five, `dep-weight` and `rule-coverage` carry one each.
|
||||
|
||||
**A worse failure mode than predicted:** a mutation can become weak
|
||||
**without anyone touching it**. AM-6's went `SURVIVED` when T04 moved its
|
||||
gate from debug to release. Mutation strength is coupled to measurement
|
||||
conditions, so a mutation is not a write-once artifact. CB-WP-0005 T08's
|
||||
stronger remedy — mutations written by someone other than the author —
|
||||
was **not tested** and should not be assumed unnecessary: `EXPECT-VACUOUS`
|
||||
covers the cheap failure, not the expensive one AM-6 demonstrated.
|
||||
|
||||
**Prediction error collapsed** — 4–5×, then 2.5×, now small — because this
|
||||
pass predicted **per task, as a mechanism, with the alternative named**.
|
||||
Both branches are outcomes someone must defend, so the prediction cannot
|
||||
be dodged; AM-3 and AM-4c took the second branch and are better resolved
|
||||
for it.
|
||||
|
||||
**No InnerLoop change.** v1.4's mutation requirement is one pass old and
|
||||
changing it before a second use would be the invention-in-isolation INTENT
|
||||
warns about — the same argument used to amend K14. **The named next
|
||||
candidate is `specs/SessionShape.md`**: SS-01…SS-05 have never been
|
||||
enforced, and this pass ran at **2.5× the context ceiling its own spec
|
||||
sets** without anything saying a word.
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue