diff --git a/evidence/CB-EV-0005-instrument-the-table.md b/evidence/CB-EV-0005-instrument-the-table.md deleted file mode 100644 index e532523..0000000 --- a/evidence/CB-EV-0005-instrument-the-table.md +++ /dev/null @@ -1,163 +0,0 @@ -# CB-EV-0005: did instrumenting the acceptance table change what it enforces? - -research: [CB-RES-0004](../research/CB-RES-0004-replay-and-kernel-coverage.md) -adr: [ADR-0005](../decisions/ADR-0005-assertion-coverage-and-replay.md) -workplan: [CB-WP-0006](../workplans/CB-WP-0006-instrument-the-table.md) -log: [260801-cb-wp-0006-log.md](../history/260801-cb-wp-0006-log.md) -instruments: `make mutation-check`, `make coverage`, `make cost-mix` -window: `cb-cost --since dfd0d6d` — 144 responses, **$51.52** - -> **Numbering correction.** CB-WP-0006 T08 says "commit -> `evidence/CB-EV-0004`". That number was already taken by CB-WP-0005's -> evidence. This is **CB-EV-0005**. - ---- - -## Test 1 — did the enforced count rise, per row? - -**M-D1-MUT: 4 of 14 → 8 of 14.** - -| row | before | after | task | -|---|---|---|---| -| AM-1 | red | red | — | -| **AM-2** | unmutatable | **red** | T02 | -| AM-3 | unmutatable | unmutatable — **blocked**, artifact never built | T02 | -| AM-4a / AM-4b | red | red | — | -| AM-4c | unmutatable | unmutatable — **withdrawn**, retained in denominator | T04 | -| AM-5 | unmutatable | unmutatable — **measured and met**, spec declares it ungated | T03 | -| **AM-6** | unmutatable | **red** | T01 | -| AM-7 | PARTIAL 1/3 | PARTIAL **2/3** — `hash-identical` re-earned | T06 | -| AM-8 | PARTIAL 1/2 | PARTIAL 1/2 | — | -| **AM-9** | unmutatable | **red** | T03 | -| AM-10 | unmutatable | unmutatable — withdrawn | — | -| **AM-11** | unmutatable | **red** | T05 | -| AM-12 | red | red | — | - -Per-task predictions, reported against what happened: - -| task | predicted | measured | -|---|---|---| -| T01 AM-6 | `unmutatable` → `red` | **met** | -| T02 AM-2 | instrumented | **met** | -| T02 AM-3 | instrumented **or** restated | **restated** — blocked on an artifact, not a metric | -| T03 AM-9 | measured | **met** | -| T03 AM-5 | measured **or** withdrawn | **measured**, and the first reading was wrong (below) | -| T04 AM-4c | targeted **or** dropped | **dropped**, with the argument | -| T05 AM-11 | `unmutatable` → `red` | **met** | -| T06 K10 | replay real, AM-7 re-earnable | **met** — hash clause re-earned | -| T07 K14/K18 | implement **or** amend | **one of each** | - -**The honest denominator, both ways.** 8 of 14 is 57%. But four rows -*cannot* be enforced — AM-3 is blocked on an artifact, AM-4c is withdrawn, -AM-5 is declared ungated by the spec, AM-10 is withdrawn — and two are -partial. Of the **10 enforceable** rows, **8 are enforced**. - -The 14 stays the headline. AM-4c is deliberately retained in that -denominator after withdrawal, and the tool says so on every run: *a score -improved by deleting the question is not an improvement.* - -**Kernel spec→code link: 15/18 → 18/18.** With the caveat the gate prints -every run: that is names, not assertions. - -## Test 2 — did any row regress from `red` to `SURVIVED`? - -**Yes, once, and it was caught.** - -In T04, moving AM-6's gate from debug to release turned its mutation -`SURVIVED`. The 4,000 `black_box` iterations were calibrated against -debug's 3.4× headroom and are invisible against release's 20×. Raised to -100,000; back to `red`. - -The generalizable finding: **a weak mutation is not a fixed property of a -row. It can become weak when the row's measurement conditions change** — -without the row, the mutation, or the code being touched. That is a -maintenance burden M-D1-MUT carries and T09 should weigh. - -No other row regressed. `SURVIVED` count across the final run: **0**. - -## Test 3 — were the new mutations strong? - -The workplan required every new mutation to be *shown to fail for the -stated reason*. That is now **mechanical, not asserted**. - -`mutation-check` gained an `EXPECT-VACUOUS` verdict: if a row's `expect` -string appears in **passing** output, the FA guard would accept any -failure at all, and the row is reported broken rather than `red`. - -**Final run: 0 vacuous expects across all 14 rows.** - -This control exists because the failure it detects happened. My first -`expect` for AM-2 was `"AM-2"` — which appears in the passing report, so -the guard would have accepted a compile error as proof the row was -enforced. Caught by hand in T02; caught mechanically from T08 onward. - -Every mutation added this pass is a **property** mutation, not a threshold -tweak: - -| row | mutation | expect | -|---|---|---| -| AM-6 | 100,000 `black_box` iterations in `GroundState::fold` | `AM-6 UNMET` | -| AM-2 | ~800 filler lines into the impl | `FAIL target <= 40` | -| AM-9 | a 300 MB allocation in the workload | `FAIL target <= 64 MB` | -| AM-11 | `NullRng::draw` returns its bound | `outside 0..` | - -## What the instruments caught that nothing else would - -Six defects, all found by a gate rather than by review: - -1. **AM-6's gate measured contention, not throughput** — 38,753 ev/s - inside `make all` against 341,280 in isolation, because `cargo test` - runs concurrently. -2. **AM-5's first reading was wrong by 61%** — 87.0 s under load against - 37.3 s quiet. Reported as a 45% breach; it is met with 1.6× headroom. - *Same error class as (1), committed two tasks later by the same author - in the row immediately after diagnosing it.* -3. **RSS over-reported 3×** — `getrusage(RUSAGE_CHILDREN)` is a - high-water mark across every reaped child, so `cargo`'s memory was - charged to the workload. Found only by cross-checking `/usr/bin/time`. -4. **K10's first round trip did not reproduce** — hashing a - `serde_json::Value` is a different canonical form than hashing the - typed aggregate. A round trip that recomputed its own comparison value - would have *passed* this. -5. **The bench workload existed twice** — `bench_shape` kept passing - while `bench-test` broke on the same edit. -6. **`facts-check` caught two spec copies going stale** within the hour - the numbers moved. - -## Cost, and a regression that matters - -| | CB-WP-0004 window | CB-WP-0005 | **CB-WP-0006** | -|---|---|---|---| -| pass cost | — | $16.03 | **$51.52** | -| responses | — | 83 | **144** | -| mechanical share | 26% | 38% | **50%** | -| environment setup | 3 turns | 0 | **0** | -| hub + workplan edits | 1 turn | 0 | **0** | -| ad-hoc text patching | 8 turns | 11 | **45 turns, $13.79** | -| orientation / inspect | 41 turns | 7 | **19 turns, $10.97** | -| mean context | — | 315,170 | **493,486** | - -**Mechanical share rose to 50% — the highest ever recorded, above the 38% -baseline that motivated CB-WP-0004.** This is not a tooling regression: -environment setup and task closes are still at **zero**, two passes on. -It is the *other* half of CB-WP-0004 T06's finding arriving in force — -text patching and orientation never had their manual path removed, and a -code-heavy pass is exactly where that spends. - -**Context is the sharper problem.** Mean 493,486 against a 200,000 target, -up from 315,170 last pass. `specs/SessionShape.md` has stated SS-01…SS-05 -since CB-WP-0003 and **none of them has ever been enforced** — the only -acceptance-adjacent numbers in this project with no gate at all, in a pass -whose entire subject was ungated numbers. - -## Verdict - -| claim | status | -|---|---| -| enforced count rose | **confirmed** — 4 → 8 of 14, 8 of 10 enforceable | -| kernel link complete | **confirmed** — 18/18, names only | -| no row regressed to SURVIVED | **confirmed** — one near-miss, caught | -| new mutations are strong | **confirmed mechanically** — 0 vacuous expects | -| AM-3 instrumented | **not met** — blocked on an artifact, stated not hidden | -| AM-7, AM-8 fully enforced | **not met** — partial, reported as partial | -| the pass got cheaper | **refuted** — mechanical share rose to 50% | diff --git a/history/260801-instrument-the-table-retrospective.md b/history/260801-instrument-the-table-retrospective.md deleted file mode 100644 index 8717da1..0000000 --- a/history/260801-instrument-the-table-retrospective.md +++ /dev/null @@ -1,141 +0,0 @@ -# 2026-08-01 — retrospective: is a weak mutation the new grep? - -CB-WP-0006 T09. The pass instrumented AM-2, AM-6, AM-9 and AM-11, -implemented K10, K11 and K18, amended K14, withdrew AM-4c, and measured -itself in [CB-EV-0005](../evidence/CB-EV-0005-instrument-the-table.md). - -## The question this task inherited - -> **M-D1-MUT does not remove a manual path — writing a weak mutation is -> exactly as easy as writing a strong one, and the harness cannot tell the -> difference. Is it a real instrument, or a name-counter with extra -> steps?** - -**It is a real instrument — but only because it was hardened three times -in one pass, and it would have been a name-counter without that.** - -## What it took to make one mutation trustworthy - -The naive idea — *invert the property, require red* — is not enough. Five -controls now stand between a mutation and a `red` verdict, and **every one -of them exists because its failure actually occurred**: - -| control | verdict | the failure that motivated it | -|---|---|---| -| the mutation must apply | `HARNESS-BROKEN` | a stale find-string would score the baseline as the mutant | -| the baseline must be green first | `inconclusive` | "mutant red" proves nothing if it was already red | -| the tree must be restored, and verified | — | every later row runs against a corrupted tree | -| the failure must match a **stated reason** | `WRONG-REASON` | a mutation that merely fails to compile would credit the row | -| the stated reason must be **absent from passing output** | `EXPECT-VACUOUS` | my first AM-2 `expect` was `"AM-2"`, which the passing report contains | - -That last one is the sharpest. It means the FA guard needed a guard, and -the failure it prevents is one I committed on the guard's second use. - -> **The generalizable finding: an instrument that measures whether other -> instruments work needs more controls than the instruments it measures.** -> M-D1-MUT carries five; `dep-weight` and `rule-coverage` carry one each. -> That asymmetry is not overhead, it is the cost of a meta-instrument, and -> a project that adds one should budget for it. - -## The failure mode that is still open - -CB-WP-0005 predicted a weak mutation would be *written* weak. This pass -found a worse case: **a mutation can become weak without anyone touching -it.** - -AM-6's 4,000 `black_box` iterations were calibrated against debug's 3.4× -headroom. When T04 moved the gate to release — correctly, because the -debug measurement was invalid — the same mutation became invisible against -20× headroom and the row went `SURVIVED`. Nothing about the row, the -mutation, or the code changed. - -The harness caught it, so the loop is not blind here. But it means -**mutation strength is coupled to measurement conditions**, and a mutation -is not a write-once artifact. That is a standing maintenance cost, and it -is the honest answer to "is this cheaper than it looks": no. - -**CB-WP-0005 T08's stronger remedy — mutations written by someone other -than the author of the assertion — was not tested.** The adversarial -reviewer wrote none of these. `EXPECT-VACUOUS` covers the cheap failure -(a guard that matches everything); it does **not** cover the expensive one -(a mutation too weak to trip a real threshold), which is exactly what -AM-6 demonstrated. That remedy remains untested and should not be assumed -unnecessary. - -## Did the "removes the manual path" test predict which gates work? - -No, and this pass is the counter-example that settles it. - -`mutation-check` fails that test outright — `grep` is still one keystroke -away, and a weak mutation is as easy to write as a strong one. It also -produced **six defects nothing else would have found**: AM-6 measuring -contention, AM-5's 61% measurement error, RSS over-reported 3×, K10's -first round trip not reproducing, the bench workload existing twice, and -two spec copies going stale. - -CB-WP-0004 T06 said tooling recovers capacity only where it removes the -manual path. **That is a predictor of whether a gate saves money, not of -whether it is worth having** — a correction first stated in CB-WP-0005 T08 -and now supported by a full pass. `mutation-check` costs money and earns -its place on findings. - -## The cost result refutes the pass - -**Mechanical share rose to 50%** — the highest recorded, above the 38% -baseline that motivated CB-WP-0004. Mean context **493,486** against a -200,000 target, up from 315,170. - -This is not a tooling regression, and the distinction matters: environment -setup and task closes are still at **zero**, two passes on. It is the -*other* half of T06's mechanism arriving in force. Text patching (45 -turns, $13.79) and orientation (19 turns, $10.97) never had their manual -path removed, and a code-heavy pass is precisely where that spends. - -So the mechanism holds in both directions, which is the strongest evidence -it is real: what was removed stayed removed; what was merely improved -stayed expensive, and got worse under load. - -## Prediction discipline, third outing - -| pass | prediction | measured | error | -|---|---|---|---| -| CB-WP-0004 | 25–30 points recovered | 6 | 4–5× | -| CB-WP-0005 | ≥10 of 14 rows | 4 | 2.5× | -| **CB-WP-0006** | **9 per-task outcomes** | **7 met, 2 restated with an argument** | **small** | - -The first two were aggregate predictions with a confidence label. This -pass predicted **per task, as a mechanism, with the alternative named** — -"instrument it, *or* restate it with an argument" — and the error -collapsed. - -That is CB-WP-0004 T06's recommendation working. It also shows why: -a per-task prediction with a named alternative cannot be dodged, because -*both* branches are outcomes someone has to defend. AM-3 and AM-4c both -took the second branch, and both are better resolved than if a number had -been forced. - -## What the loop should change - -**Nothing in InnerLoop this pass.** v1.4's mutation requirement was added -by CB-WP-0005 and this pass is its first full exercise; changing it again -before a second use would be the invention-in-isolation INTENT warns -about, and the same argument used to amend K14 four hours ago. - -**The named next candidate is `specs/SessionShape.md`.** SS-01…SS-05 have -been stated since CB-WP-0003 and **none has ever been enforced** — they -are the only acceptance-adjacent numbers in this project with no gate at -all. Measured this pass: mean context 493,486 (target 200,000), p90 -594,888 (target 300,000), batching well under target. A pass whose entire -subject was ungated numbers ran at 2.5× the context ceiling its own spec -sets, and nothing said a word. - -## Open, not closed - -- **AM-3 blocked** on an artifact that has never been built — a minimal - synthetic game on our kernel. Needs a deliverable, not a metric tweak. -- **AM-7 and AM-8 remain PARTIAL**: no code computes AM-7's throughput - ratio, and nothing enforces AM-8's N=10. -- **`CommitWindow` is provisional with a delete-by date of 2026-12-31.** -- **The kernel gate binds 2026-08-31** — 18/18 today, so it binds green. -- **SessionShape is ungated**, and is now the largest measured gap. -- **The chaos roll is at declaration 3 of 12** (T04 rolled d4=2). diff --git a/tools/mutation-check.py b/tools/mutation-check.py index 06682fd..c3a578e 100644 --- a/tools/mutation-check.py +++ b/tools/mutation-check.py @@ -282,20 +282,10 @@ def check_row(row): # Positive control 2: the baseline must be green, or "mutant red" # proves nothing. - base_ok, base_tail, base_out = run(row.verify) + base_ok, base_tail, _ = run(row.verify) if not base_ok: return "inconclusive", f"baseline already red: {base_tail}" - # CB-WP-0006 T08: the FA guard is only a guard if its `expect` string - # cannot appear in PASSING output. An expect of "AM-2" would match the - # normal report and accept any failure at all — which is how the guard - # goes vacuous without anyone noticing. My first attempt on AM-2 did - # exactly that. - if row.expect and row.expect in base_out: - return "EXPECT-VACUOUS", ( - f"expect string {row.expect!r} appears in PASSING output, so it " - f"would accept any failure — the FA guard is inert for this row") - try: mutated = original.replace(old, new, 1) # Positive control 3: the file content must actually differ. @@ -347,13 +337,12 @@ def report(only=None): detail = (f"{len(r.clauses) - len(unmet)}/{len(r.clauses)} " f"clauses enforced") tally[verdict] = tally.get(verdict, 0) + 1 - if verdict in ("HARNESS-BROKEN", "EXPECT-VACUOUS"): + if verdict == "HARNESS-BROKEN": broken.append(r.id) mark = {"red": "red ", "SURVIVED": "SURVIVED ", "unmutatable": "unmutatable", "inconclusive": "inconclusive", "PARTIAL": "PARTIAL ", "HARNESS-BROKEN": "BROKEN ", - "EXPECT-VACUOUS": "EXPECT-VOID", "WRONG-REASON": "WRONG-REASON"}[verdict] print(f" [{mark}] {r.id:<6} {r.claim[:52]}") if detail: @@ -377,8 +366,8 @@ def report(only=None): print(f" ({withdrawn} withdrawn row(s) retained in the denominator " f"on purpose — a score\n improved by deleting the question " f"is not an improvement)") - for k in ("PARTIAL", "SURVIVED", "WRONG-REASON", "EXPECT-VACUOUS", - "unmutatable", "inconclusive"): + for k in ("PARTIAL", "SURVIVED", "WRONG-REASON", "unmutatable", + "inconclusive"): if tally.get(k): print(f" {k:<13} {tally[k]}") if only: diff --git a/workplans/CB-WP-0006-instrument-the-table.md b/workplans/CB-WP-0006-instrument-the-table.md index 9ec2513..0de81d9 100644 --- a/workplans/CB-WP-0006-instrument-the-table.md +++ b/workplans/CB-WP-0006-instrument-the-table.md @@ -1,7 +1,7 @@ --- id: CB-WP-0006 title: "Instrument the acceptance table, then implement what it exposes" -status: done +status: in_progress state_hub_workstream_id: "8a6327cc-fd5c-4e2c-a29b-b437c27d1e71" --- @@ -222,23 +222,14 @@ right for K18. ```task id: CB-WP-0006-T08 -status: done +status: todo priority: high state_hub_task_id: "1d4455c3-c642-47d8-916b-ab35bc512207" ``` -Commit `evidence/CB-EV-0005` — **numbering corrected**, CB-EV-0004 was -already CB-WP-0005's. The baseline is **4 of 14**, committed and +Commit `evidence/CB-EV-0004`. The baseline is **4 of 14**, committed and reproducible via `make mutation-check`. -**Measured — [CB-EV-0005](../evidence/CB-EV-0005-instrument-the-table.md). -4 of 14 → 8 of 14** (8 of 10 enforceable; the 14 stays the headline and -AM-4c stays in it on purpose). Kernel link 15/18 → **18/18**, names only. -One row regressed to `SURVIVED` and was caught. **0 vacuous expects** -across 14 rows — test 3 is now mechanical, via a new `EXPECT-VACUOUS` -verdict. **The cost result refutes the pass:** mechanical share rose to -**50%**, the highest recorded. - Three tests, all reported: 1. **Did the enforced count rise, per row?** Against the per-task @@ -255,7 +246,7 @@ Three tests, all reported: ```task id: CB-WP-0006-T09 -status: done +status: todo priority: medium state_hub_task_id: "4f291d5a-90f0-4f1f-bcaf-7ec28856bfcf" ``` @@ -277,33 +268,3 @@ the latter, say so and propose what would actually bind — the candidate being that the mutation must be written by someone other than the author of the assertion, which is the adversarial-review principle applied one level down. - -**Answered — [260801-instrument-the-table-retrospective.md](../history/260801-instrument-the-table-retrospective.md).** - -**A real instrument — but only because it was hardened three times in one -pass.** Five controls now stand between a mutation and a `red` verdict, -and every one exists because its failure actually occurred. The -generalizable finding: **an instrument that measures whether other -instruments work needs more controls than the instruments it measures** — -M-D1-MUT carries five, `dep-weight` and `rule-coverage` carry one each. - -**A worse failure mode than predicted:** a mutation can become weak -**without anyone touching it**. AM-6's went `SURVIVED` when T04 moved its -gate from debug to release. Mutation strength is coupled to measurement -conditions, so a mutation is not a write-once artifact. CB-WP-0005 T08's -stronger remedy — mutations written by someone other than the author — -was **not tested** and should not be assumed unnecessary: `EXPECT-VACUOUS` -covers the cheap failure, not the expensive one AM-6 demonstrated. - -**Prediction error collapsed** — 4–5×, then 2.5×, now small — because this -pass predicted **per task, as a mechanism, with the alternative named**. -Both branches are outcomes someone must defend, so the prediction cannot -be dodged; AM-3 and AM-4c took the second branch and are better resolved -for it. - -**No InnerLoop change.** v1.4's mutation requirement is one pass old and -changing it before a second use would be the invention-in-isolation INTENT -warns about — the same argument used to amend K14. **The named next -candidate is `specs/SessionShape.md`**: SS-01…SS-05 have never been -enforced, and this pass ran at **2.5× the context ceiling its own spec -sets** without anything saying a word.