Compare commits

..

No commits in common. "1a09c4c57f7145397a77a9e5b1fdc60a56db09f1" and "c51c7c9b47c00831f9044d0f554669561802f0d9" have entirely different histories.

4 changed files with 8 additions and 362 deletions

View file

@ -1,163 +0,0 @@
# CB-EV-0005: did instrumenting the acceptance table change what it enforces?
research: [CB-RES-0004](../research/CB-RES-0004-replay-and-kernel-coverage.md)
adr: [ADR-0005](../decisions/ADR-0005-assertion-coverage-and-replay.md)
workplan: [CB-WP-0006](../workplans/CB-WP-0006-instrument-the-table.md)
log: [260801-cb-wp-0006-log.md](../history/260801-cb-wp-0006-log.md)
instruments: `make mutation-check`, `make coverage`, `make cost-mix`
window: `cb-cost --since dfd0d6d` — 144 responses, **$51.52**
> **Numbering correction.** CB-WP-0006 T08 says "commit
> `evidence/CB-EV-0004`". That number was already taken by CB-WP-0005's
> evidence. This is **CB-EV-0005**.
---
## Test 1 — did the enforced count rise, per row?
**M-D1-MUT: 4 of 14 → 8 of 14.**
| row | before | after | task |
|---|---|---|---|
| AM-1 | red | red | — |
| **AM-2** | unmutatable | **red** | T02 |
| AM-3 | unmutatable | unmutatable — **blocked**, artifact never built | T02 |
| AM-4a / AM-4b | red | red | — |
| AM-4c | unmutatable | unmutatable — **withdrawn**, retained in denominator | T04 |
| AM-5 | unmutatable | unmutatable — **measured and met**, spec declares it ungated | T03 |
| **AM-6** | unmutatable | **red** | T01 |
| AM-7 | PARTIAL 1/3 | PARTIAL **2/3** — `hash-identical` re-earned | T06 |
| AM-8 | PARTIAL 1/2 | PARTIAL 1/2 | — |
| **AM-9** | unmutatable | **red** | T03 |
| AM-10 | unmutatable | unmutatable — withdrawn | — |
| **AM-11** | unmutatable | **red** | T05 |
| AM-12 | red | red | — |
Per-task predictions, reported against what happened:
| task | predicted | measured |
|---|---|---|
| T01 AM-6 | `unmutatable` → `red` | **met** |
| T02 AM-2 | instrumented | **met** |
| T02 AM-3 | instrumented **or** restated | **restated** — blocked on an artifact, not a metric |
| T03 AM-9 | measured | **met** |
| T03 AM-5 | measured **or** withdrawn | **measured**, and the first reading was wrong (below) |
| T04 AM-4c | targeted **or** dropped | **dropped**, with the argument |
| T05 AM-11 | `unmutatable` → `red` | **met** |
| T06 K10 | replay real, AM-7 re-earnable | **met** — hash clause re-earned |
| T07 K14/K18 | implement **or** amend | **one of each** |
**The honest denominator, both ways.** 8 of 14 is 57%. But four rows
*cannot* be enforced — AM-3 is blocked on an artifact, AM-4c is withdrawn,
AM-5 is declared ungated by the spec, AM-10 is withdrawn — and two are
partial. Of the **10 enforceable** rows, **8 are enforced**.
The 14 stays the headline. AM-4c is deliberately retained in that
denominator after withdrawal, and the tool says so on every run: *a score
improved by deleting the question is not an improvement.*
**Kernel spec→code link: 15/18 → 18/18.** With the caveat the gate prints
every run: that is names, not assertions.
## Test 2 — did any row regress from `red` to `SURVIVED`?
**Yes, once, and it was caught.**
In T04, moving AM-6's gate from debug to release turned its mutation
`SURVIVED`. The 4,000 `black_box` iterations were calibrated against
debug's 3.4× headroom and are invisible against release's 20×. Raised to
100,000; back to `red`.
The generalizable finding: **a weak mutation is not a fixed property of a
row. It can become weak when the row's measurement conditions change** —
without the row, the mutation, or the code being touched. That is a
maintenance burden M-D1-MUT carries and T09 should weigh.
No other row regressed. `SURVIVED` count across the final run: **0**.
## Test 3 — were the new mutations strong?
The workplan required every new mutation to be *shown to fail for the
stated reason*. That is now **mechanical, not asserted**.
`mutation-check` gained an `EXPECT-VACUOUS` verdict: if a row's `expect`
string appears in **passing** output, the FA guard would accept any
failure at all, and the row is reported broken rather than `red`.
**Final run: 0 vacuous expects across all 14 rows.**
This control exists because the failure it detects happened. My first
`expect` for AM-2 was `"AM-2"` — which appears in the passing report, so
the guard would have accepted a compile error as proof the row was
enforced. Caught by hand in T02; caught mechanically from T08 onward.
Every mutation added this pass is a **property** mutation, not a threshold
tweak:
| row | mutation | expect |
|---|---|---|
| AM-6 | 100,000 `black_box` iterations in `GroundState::fold` | `AM-6 UNMET` |
| AM-2 | ~800 filler lines into the impl | `FAIL target <= 40` |
| AM-9 | a 300 MB allocation in the workload | `FAIL target <= 64 MB` |
| AM-11 | `NullRng::draw` returns its bound | `outside 0..` |
## What the instruments caught that nothing else would
Six defects, all found by a gate rather than by review:
1. **AM-6's gate measured contention, not throughput** — 38,753 ev/s
inside `make all` against 341,280 in isolation, because `cargo test`
runs concurrently.
2. **AM-5's first reading was wrong by 61%** — 87.0 s under load against
37.3 s quiet. Reported as a 45% breach; it is met with 1.6× headroom.
*Same error class as (1), committed two tasks later by the same author
in the row immediately after diagnosing it.*
3. **RSS over-reported 3×** — `getrusage(RUSAGE_CHILDREN)` is a
high-water mark across every reaped child, so `cargo`'s memory was
charged to the workload. Found only by cross-checking `/usr/bin/time`.
4. **K10's first round trip did not reproduce** — hashing a
`serde_json::Value` is a different canonical form than hashing the
typed aggregate. A round trip that recomputed its own comparison value
would have *passed* this.
5. **The bench workload existed twice** — `bench_shape` kept passing
while `bench-test` broke on the same edit.
6. **`facts-check` caught two spec copies going stale** within the hour
the numbers moved.
## Cost, and a regression that matters
| | CB-WP-0004 window | CB-WP-0005 | **CB-WP-0006** |
|---|---|---|---|
| pass cost | — | $16.03 | **$51.52** |
| responses | — | 83 | **144** |
| mechanical share | 26% | 38% | **50%** |
| environment setup | 3 turns | 0 | **0** |
| hub + workplan edits | 1 turn | 0 | **0** |
| ad-hoc text patching | 8 turns | 11 | **45 turns, $13.79** |
| orientation / inspect | 41 turns | 7 | **19 turns, $10.97** |
| mean context | — | 315,170 | **493,486** |
**Mechanical share rose to 50% — the highest ever recorded, above the 38%
baseline that motivated CB-WP-0004.** This is not a tooling regression:
environment setup and task closes are still at **zero**, two passes on.
It is the *other* half of CB-WP-0004 T06's finding arriving in force —
text patching and orientation never had their manual path removed, and a
code-heavy pass is exactly where that spends.
**Context is the sharper problem.** Mean 493,486 against a 200,000 target,
up from 315,170 last pass. `specs/SessionShape.md` has stated SS-01…SS-05
since CB-WP-0003 and **none of them has ever been enforced** — the only
acceptance-adjacent numbers in this project with no gate at all, in a pass
whose entire subject was ungated numbers.
## Verdict
| claim | status |
|---|---|
| enforced count rose | **confirmed** — 4 → 8 of 14, 8 of 10 enforceable |
| kernel link complete | **confirmed** — 18/18, names only |
| no row regressed to SURVIVED | **confirmed** — one near-miss, caught |
| new mutations are strong | **confirmed mechanically** — 0 vacuous expects |
| AM-3 instrumented | **not met** — blocked on an artifact, stated not hidden |
| AM-7, AM-8 fully enforced | **not met** — partial, reported as partial |
| the pass got cheaper | **refuted** — mechanical share rose to 50% |

View file

@ -1,141 +0,0 @@
# 2026-08-01 — retrospective: is a weak mutation the new grep?
CB-WP-0006 T09. The pass instrumented AM-2, AM-6, AM-9 and AM-11,
implemented K10, K11 and K18, amended K14, withdrew AM-4c, and measured
itself in [CB-EV-0005](../evidence/CB-EV-0005-instrument-the-table.md).
## The question this task inherited
> **M-D1-MUT does not remove a manual path — writing a weak mutation is
> exactly as easy as writing a strong one, and the harness cannot tell the
> difference. Is it a real instrument, or a name-counter with extra
> steps?**
**It is a real instrument — but only because it was hardened three times
in one pass, and it would have been a name-counter without that.**
## What it took to make one mutation trustworthy
The naive idea — *invert the property, require red* — is not enough. Five
controls now stand between a mutation and a `red` verdict, and **every one
of them exists because its failure actually occurred**:
| control | verdict | the failure that motivated it |
|---|---|---|
| the mutation must apply | `HARNESS-BROKEN` | a stale find-string would score the baseline as the mutant |
| the baseline must be green first | `inconclusive` | "mutant red" proves nothing if it was already red |
| the tree must be restored, and verified | — | every later row runs against a corrupted tree |
| the failure must match a **stated reason** | `WRONG-REASON` | a mutation that merely fails to compile would credit the row |
| the stated reason must be **absent from passing output** | `EXPECT-VACUOUS` | my first AM-2 `expect` was `"AM-2"`, which the passing report contains |
That last one is the sharpest. It means the FA guard needed a guard, and
the failure it prevents is one I committed on the guard's second use.
> **The generalizable finding: an instrument that measures whether other
> instruments work needs more controls than the instruments it measures.**
> M-D1-MUT carries five; `dep-weight` and `rule-coverage` carry one each.
> That asymmetry is not overhead, it is the cost of a meta-instrument, and
> a project that adds one should budget for it.
## The failure mode that is still open
CB-WP-0005 predicted a weak mutation would be *written* weak. This pass
found a worse case: **a mutation can become weak without anyone touching
it.**
AM-6's 4,000 `black_box` iterations were calibrated against debug's 3.4×
headroom. When T04 moved the gate to release — correctly, because the
debug measurement was invalid — the same mutation became invisible against
20× headroom and the row went `SURVIVED`. Nothing about the row, the
mutation, or the code changed.
The harness caught it, so the loop is not blind here. But it means
**mutation strength is coupled to measurement conditions**, and a mutation
is not a write-once artifact. That is a standing maintenance cost, and it
is the honest answer to "is this cheaper than it looks": no.
**CB-WP-0005 T08's stronger remedy — mutations written by someone other
than the author of the assertion — was not tested.** The adversarial
reviewer wrote none of these. `EXPECT-VACUOUS` covers the cheap failure
(a guard that matches everything); it does **not** cover the expensive one
(a mutation too weak to trip a real threshold), which is exactly what
AM-6 demonstrated. That remedy remains untested and should not be assumed
unnecessary.
## Did the "removes the manual path" test predict which gates work?
No, and this pass is the counter-example that settles it.
`mutation-check` fails that test outright — `grep` is still one keystroke
away, and a weak mutation is as easy to write as a strong one. It also
produced **six defects nothing else would have found**: AM-6 measuring
contention, AM-5's 61% measurement error, RSS over-reported 3×, K10's
first round trip not reproducing, the bench workload existing twice, and
two spec copies going stale.
CB-WP-0004 T06 said tooling recovers capacity only where it removes the
manual path. **That is a predictor of whether a gate saves money, not of
whether it is worth having** — a correction first stated in CB-WP-0005 T08
and now supported by a full pass. `mutation-check` costs money and earns
its place on findings.
## The cost result refutes the pass
**Mechanical share rose to 50%** — the highest recorded, above the 38%
baseline that motivated CB-WP-0004. Mean context **493,486** against a
200,000 target, up from 315,170.
This is not a tooling regression, and the distinction matters: environment
setup and task closes are still at **zero**, two passes on. It is the
*other* half of T06's mechanism arriving in force. Text patching (45
turns, $13.79) and orientation (19 turns, $10.97) never had their manual
path removed, and a code-heavy pass is precisely where that spends.
So the mechanism holds in both directions, which is the strongest evidence
it is real: what was removed stayed removed; what was merely improved
stayed expensive, and got worse under load.
## Prediction discipline, third outing
| pass | prediction | measured | error |
|---|---|---|---|
| CB-WP-0004 | 25–30 points recovered | 6 | 4–5× |
| CB-WP-0005 | ≥10 of 14 rows | 4 | 2.5× |
| **CB-WP-0006** | **9 per-task outcomes** | **7 met, 2 restated with an argument** | **small** |
The first two were aggregate predictions with a confidence label. This
pass predicted **per task, as a mechanism, with the alternative named** —
"instrument it, *or* restate it with an argument" — and the error
collapsed.
That is CB-WP-0004 T06's recommendation working. It also shows why:
a per-task prediction with a named alternative cannot be dodged, because
*both* branches are outcomes someone has to defend. AM-3 and AM-4c both
took the second branch, and both are better resolved than if a number had
been forced.
## What the loop should change
**Nothing in InnerLoop this pass.** v1.4's mutation requirement was added
by CB-WP-0005 and this pass is its first full exercise; changing it again
before a second use would be the invention-in-isolation INTENT warns
about, and the same argument used to amend K14 four hours ago.
**The named next candidate is `specs/SessionShape.md`.** SS-01…SS-05 have
been stated since CB-WP-0003 and **none has ever been enforced** — they
are the only acceptance-adjacent numbers in this project with no gate at
all. Measured this pass: mean context 493,486 (target 200,000), p90
594,888 (target 300,000), batching well under target. A pass whose entire
subject was ungated numbers ran at 2.5× the context ceiling its own spec
sets, and nothing said a word.
## Open, not closed
- **AM-3 blocked** on an artifact that has never been built — a minimal
synthetic game on our kernel. Needs a deliverable, not a metric tweak.
- **AM-7 and AM-8 remain PARTIAL**: no code computes AM-7's throughput
ratio, and nothing enforces AM-8's N=10.
- **`CommitWindow` is provisional with a delete-by date of 2026-12-31.**
- **The kernel gate binds 2026-08-31** — 18/18 today, so it binds green.
- **SessionShape is ungated**, and is now the largest measured gap.
- **The chaos roll is at declaration 3 of 12** (T04 rolled d4=2).

View file

@ -282,20 +282,10 @@ def check_row(row):
# Positive control 2: the baseline must be green, or "mutant red" # Positive control 2: the baseline must be green, or "mutant red"
# proves nothing. # proves nothing.
base_ok, base_tail, base_out = run(row.verify) base_ok, base_tail, _ = run(row.verify)
if not base_ok: if not base_ok:
return "inconclusive", f"baseline already red: {base_tail}" return "inconclusive", f"baseline already red: {base_tail}"
# CB-WP-0006 T08: the FA guard is only a guard if its `expect` string
# cannot appear in PASSING output. An expect of "AM-2" would match the
# normal report and accept any failure at all — which is how the guard
# goes vacuous without anyone noticing. My first attempt on AM-2 did
# exactly that.
if row.expect and row.expect in base_out:
return "EXPECT-VACUOUS", (
f"expect string {row.expect!r} appears in PASSING output, so it "
f"would accept any failure — the FA guard is inert for this row")
try: try:
mutated = original.replace(old, new, 1) mutated = original.replace(old, new, 1)
# Positive control 3: the file content must actually differ. # Positive control 3: the file content must actually differ.
@ -347,13 +337,12 @@ def report(only=None):
detail = (f"{len(r.clauses) - len(unmet)}/{len(r.clauses)} " detail = (f"{len(r.clauses) - len(unmet)}/{len(r.clauses)} "
f"clauses enforced") f"clauses enforced")
tally[verdict] = tally.get(verdict, 0) + 1 tally[verdict] = tally.get(verdict, 0) + 1
if verdict in ("HARNESS-BROKEN", "EXPECT-VACUOUS"): if verdict == "HARNESS-BROKEN":
broken.append(r.id) broken.append(r.id)
mark = {"red": "red ", "SURVIVED": "SURVIVED ", mark = {"red": "red ", "SURVIVED": "SURVIVED ",
"unmutatable": "unmutatable", "inconclusive": "inconclusive", "unmutatable": "unmutatable", "inconclusive": "inconclusive",
"PARTIAL": "PARTIAL ", "PARTIAL": "PARTIAL ",
"HARNESS-BROKEN": "BROKEN ", "HARNESS-BROKEN": "BROKEN ",
"EXPECT-VACUOUS": "EXPECT-VOID",
"WRONG-REASON": "WRONG-REASON"}[verdict] "WRONG-REASON": "WRONG-REASON"}[verdict]
print(f" [{mark}] {r.id:<6} {r.claim[:52]}") print(f" [{mark}] {r.id:<6} {r.claim[:52]}")
if detail: if detail:
@ -377,8 +366,8 @@ def report(only=None):
print(f" ({withdrawn} withdrawn row(s) retained in the denominator " print(f" ({withdrawn} withdrawn row(s) retained in the denominator "
f"on purpose — a score\n improved by deleting the question " f"on purpose — a score\n improved by deleting the question "
f"is not an improvement)") f"is not an improvement)")
for k in ("PARTIAL", "SURVIVED", "WRONG-REASON", "EXPECT-VACUOUS", for k in ("PARTIAL", "SURVIVED", "WRONG-REASON", "unmutatable",
"unmutatable", "inconclusive"): "inconclusive"):
if tally.get(k): if tally.get(k):
print(f" {k:<13} {tally[k]}") print(f" {k:<13} {tally[k]}")
if only: if only:

View file

@ -1,7 +1,7 @@
--- ---
id: CB-WP-0006 id: CB-WP-0006
title: "Instrument the acceptance table, then implement what it exposes" title: "Instrument the acceptance table, then implement what it exposes"
status: done status: in_progress
state_hub_workstream_id: "8a6327cc-fd5c-4e2c-a29b-b437c27d1e71" state_hub_workstream_id: "8a6327cc-fd5c-4e2c-a29b-b437c27d1e71"
--- ---
@ -222,23 +222,14 @@ right for K18.
```task ```task
id: CB-WP-0006-T08 id: CB-WP-0006-T08
status: done status: todo
priority: high priority: high
state_hub_task_id: "1d4455c3-c642-47d8-916b-ab35bc512207" state_hub_task_id: "1d4455c3-c642-47d8-916b-ab35bc512207"
``` ```
Commit `evidence/CB-EV-0005` — **numbering corrected**, CB-EV-0004 was Commit `evidence/CB-EV-0004`. The baseline is **4 of 14**, committed and
already CB-WP-0005's. The baseline is **4 of 14**, committed and
reproducible via `make mutation-check`. reproducible via `make mutation-check`.
**Measured — [CB-EV-0005](../evidence/CB-EV-0005-instrument-the-table.md).
4 of 14 → 8 of 14** (8 of 10 enforceable; the 14 stays the headline and
AM-4c stays in it on purpose). Kernel link 15/18 → **18/18**, names only.
One row regressed to `SURVIVED` and was caught. **0 vacuous expects**
across 14 rows — test 3 is now mechanical, via a new `EXPECT-VACUOUS`
verdict. **The cost result refutes the pass:** mechanical share rose to
**50%**, the highest recorded.
Three tests, all reported: Three tests, all reported:
1. **Did the enforced count rise, per row?** Against the per-task 1. **Did the enforced count rise, per row?** Against the per-task
@ -255,7 +246,7 @@ Three tests, all reported:
```task ```task
id: CB-WP-0006-T09 id: CB-WP-0006-T09
status: done status: todo
priority: medium priority: medium
state_hub_task_id: "4f291d5a-90f0-4f1f-bcaf-7ec28856bfcf" state_hub_task_id: "4f291d5a-90f0-4f1f-bcaf-7ec28856bfcf"
``` ```
@ -277,33 +268,3 @@ the latter, say so and propose what would actually bind — the candidate
being that the mutation must be written by someone other than the author being that the mutation must be written by someone other than the author
of the assertion, which is the adversarial-review principle applied one of the assertion, which is the adversarial-review principle applied one
level down. level down.
**Answered — [260801-instrument-the-table-retrospective.md](../history/260801-instrument-the-table-retrospective.md).**
**A real instrument — but only because it was hardened three times in one
pass.** Five controls now stand between a mutation and a `red` verdict,
and every one exists because its failure actually occurred. The
generalizable finding: **an instrument that measures whether other
instruments work needs more controls than the instruments it measures** —
M-D1-MUT carries five, `dep-weight` and `rule-coverage` carry one each.
**A worse failure mode than predicted:** a mutation can become weak
**without anyone touching it**. AM-6's went `SURVIVED` when T04 moved its
gate from debug to release. Mutation strength is coupled to measurement
conditions, so a mutation is not a write-once artifact. CB-WP-0005 T08's
stronger remedy — mutations written by someone other than the author —
was **not tested** and should not be assumed unnecessary: `EXPECT-VACUOUS`
covers the cheap failure, not the expensive one AM-6 demonstrated.
**Prediction error collapsed** — 4–5×, then 2.5×, now small — because this
pass predicted **per task, as a mechanism, with the alternative named**.
Both branches are outcomes someone must defend, so the prediction cannot
be dodged; AM-3 and AM-4c took the second branch and are better resolved
for it.
**No InnerLoop change.** v1.4's mutation requirement is one pass old and
changing it before a second use would be the invention-in-isolation INTENT
warns about — the same argument used to amend K14. **The named next
candidate is `specs/SessionShape.md`**: SS-01…SS-05 have never been
enforced, and this pass ran at **2.5× the context ceiling its own spec
sets** without anything saying a word.