clay-borg/evidence/CB-EV-0004-assertion-coverage.md
tegwick fd19f4e878 CB-WP-0005 T07/T08: control loop, retrospective, InnerLoop v1.4
T07 — CB-EV-0004. 70 responses, $16.03.

Test 1 MET: 1 spec -> 2, 58 rules -> 76, 1 source file -> 10, and the new
denominator came in at 83% naming K10/K14/K18. The stated failure mode — a
widened denominator reporting the same percentage — did not occur.

Test 2 UNMET: M-D1-MUT 4 of 14 against >=10, and wrong about what as well
as how much. The diagnosis was three absent kernel rules; the measurement
found eight rows with no instrument at all.

Test 3: quality held, and widening surfaced far more than the seven known
defects — six further uninstrumented rows, HDN #7 (rule-coverage's
self-test green while the tool was broken), a stale $248.46 invisible to
facts-check because it was untagged, and a fifth error class.

The clean test CB-WP-0004 was owed is now run, on a pass that used the
tools without building them. The two categories whose tools removed the
manual path are at 0 turns two passes on; the two that merely offered a
better option are now the entire mechanical cost of a pass. Absolute
mechanical cost fell $51.76 -> $6.13. The mechanism holds.

Reported because nothing else would: SH-3 batching is 0.0% this pass — 67
tool calls across 67 responses — against a 20% target, and mean context
315,170 against 200,000. SessionShape has stated these since CB-WP-0003
and none has ever been enforced.

T08 — the retrospective answers its question: yes, a weak mutation is the
new grep, and it is worse, because it fails in the opposite direction.
mutation-check's first run produced two SURVIVED verdicts and both were
the author's own no-op mutations.

The fifth error class: false accusation. HDN, TA, SSB and DFD all
under-report — a real problem passes. FA over-reports: it publishes the
claim that working code is broken, sends the next pass to fix something
that is not broken, and is more credible than the truth because it arrives
with a measurement attached. Thirteen instances, five classes, six passes,
and the newest class is one that hardening created.

One correction to CB-WP-0004 T06: "a gate only pays if it removes the
manual path" is a predictor of whether a gate saves money, not a criterion
for whether it is worth having. mutation-check fails that test and
produced the most valuable findings of the pass.

InnerLoop v1.4: where a claim rests on numbers, the adversarial reviewer
must read the assertion behind each quoted number and mutate it.
Re-running the command that prints a number is not verification of that
number. Second verification step to inherit the author's blindness; both
fixes replace re-derivation with adversarial execution.

CB-WP-0005 status -> done, 5 of 8 tasks, 3 cancelled into CB-WP-0006.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 18:05:31 +02:00

168 lines
7.9 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# CB-EV-0004: did widening the instruments change what they report?
research: [CB-RES-0004](../research/CB-RES-0004-replay-and-kernel-coverage.md)
adr: [ADR-0005](../decisions/ADR-0005-assertion-coverage-and-replay.md)
workplan: [CB-WP-0005](../workplans/CB-WP-0005-assertion-coverage.md)
instruments: `make coverage`, `make mutation-check`, `make cost-mix`
window: `cb-cost --since bd4423a` — 70 responses, **$16.03**
Phase C was deferred unstarted, so this closes T01T03 only. That is
itself the headline result: **the pass stopped because its own instrument
contradicted the plan it was executing.**
---
## Test 1 — did the denominators widen?
| | before | after |
|---|---|---|
| specs in scope | 1 (`GroundRules.md`) | **2** (+ `GameKernel.md`) |
| numbered rules in scope | 58 | **76** |
| source files searched for links | 1 | **10** |
| denominators reported | 1 | **3** |
```text
AM-1 rule coverage: 58/58 (100%) over 21 scenarios
AM-1b spec->code link: 49/58 claimed rules named in the aggregate
AM-1b kernel spec->code link: 15/18 (83%) across 10 source files
unlinked: K10 K14 K18
```
**Met.** The workplan's stated failure mode — *a widened denominator that
reports the same percentage would prove the instruments still count
names* — did not occur. The new denominator came in at **83%**, and named
three rules that four workplans of green gates never mentioned.
The gate reports without feeding the exit code **until 2026-08-31**, then
binds. The date is in `tools/rule-coverage.py`, the days remaining print
on every run, and the self-test asserts the arm returns 0 before that date
and 2 after — so it cannot quietly become never, which is how AM-4's
targets went unratified for four workplans.
## Test 2 — M-D1-MUT against the prediction
**Prediction ≥10 of 14 (the ADR's 9-of-12, 75%). Measured 4. UNMET.**
No target moved in the commit that measured it.
```text
M-D1-MUT: 4/14 rows enforced
PARTIAL 2 AM-7, AM-8 — some clauses live, some inert
unmutatable 8 no property to invert
SURVIVED 0
```
**Two corrections to our own numbers**, both recorded rather than quietly
absorbed:
1. **There are 14 acceptance rows, not 12.** ADR-0005 and CB-WP-0005 both
said twelve; AM-4 splits into a/b/c. The prediction is evaluated as
≥10 of 14 on the same 75% basis.
2. **The first run reported two `SURVIVED` rows and both were my own
no-op mutations.** `pub struct NullRng;``pub struct NullRng {}` is
semantically identical; renaming `max_age_days` changes nothing
because CA-17 reads it with `.get(..., 90)`. Replaced with real
inversions, after which both go red.
The eight `unmutatable` rows, each with the reason the harness records:
| row | why nothing can be inverted |
|---|---|
| **AM-6** | **nothing compares any number to 100,000 events/s** — the headline throughput claim |
| AM-2 | no instrument divides LOC by rule count or compares to 40 |
| AM-3 | the synthetic workload's LOC is never measured |
| AM-4c | `reported, not targeted` — no threshold, so nothing can fail |
| AM-5 | `recorded not gated`, and not recorded either |
| AM-9 | nothing measures resident memory |
| AM-10 | population empty — no `cb-*-api` crate exists |
| AM-11 | the conformance suite the metric is a bool over does not exist |
**The prediction was wrong about *what*, not only *how much*.**
CB-RES-0004 diagnosed three absent kernel rules. The measurement found
that **more than half the acceptance table has no instrument at all**
which Phase C, scoped to five rules, would not have touched.
## Test 3 — did quality hold, and did widening surface anything new?
`make all` green throughout, now including two gates that did not exist:
the kernel coverage arm and `mutation-check --self-test`.
**It surfaced substantially more than the seven defects CB-RES-0004
named.** Six further acceptance rows with no instrument (above), plus
three new defect instances:
- **HDN #7** — `rule-coverage.py`'s `--self-test` printed all-ok while
every real `make coverage` died with `RecursionError`. A `print(` inside
`say()` had become `say(`; the control only ever called the quiet path.
*The control named the behaviour and did not assert it* — the pass's own
thesis, inside the tool written to prove it.
- **A new error class (below)** — the two no-op mutations.
- **DFD** — `evidence/CB-EV-0001` still carried **$248.46** for AM-12, the
figure CB-WP-0002 disproved four workplans ago. Untagged, and therefore
invisible to `facts-check`. The gate built to catch duplicated facts
only checks copies that opted in.
### The fifth error class: false accusation (FA)
Every class on record under-reports: a harness that does nothing, trusted
arithmetic, a blind sample, a stale copy. All four let a real problem pass.
The weak mutation is the first that **over-reports**. A no-op mutation
yields `SURVIVED`, which reads as *"this row asserts nothing"* — a
published claim that working code is broken. It sends the next pass to fix
something that is not broken, and it is *more* credible than the truth
because it arrives with a measurement attached.
| class | direction | caught by |
|---|---|---|
| HDN, TA, SSB, DFD | under-report — a real defect passes | assertions, re-derivation, all-data checks, reading the copy |
| **FA** *(new)* | **over-report — working code is indicted** | checking that the mutation fails *for its stated reason* |
Ten instances across four classes became **thirteen across five**.
## Cost, and the clean test CB-WP-0004 was owed
CB-WP-0004 T05 said its measurement was confounded because the pass built
the tools it measured, and that *"the clean test is the next pass, which
uses the tools without building them."* T01T03 are largely that pass.
| category | baseline (662 resp) | CB-WP-0004 build | **this pass (70 resp)** |
|---|---|---|---|
| environment setup | 85 turns, $15.58 | 3 turns, $0.20 | **0 turns** |
| hub task status + workplan edit | 46 turns, $11.52 | 1 turn, $0.13 | **0 turns** |
| orientation / inspect | 49 turns, $6.87 | 41 turns, $4.10 | 7 turns, **$3.85** |
| ad-hoc text patching | 76 turns, $14.11 | 8 turns, $0.75 | 11 turns, **$2.28** |
| **mechanical share of pass** | **38.2%** | 26% | **38%** |
**The verdict is clean, and it confirms CB-WP-0004 T06's mechanism
exactly.** The two categories whose tools *removed the manual path* went
to **zero and stayed there** with no further work. The two whose tools
merely *offered a better option*`make status` and `facts-check` — are
now the entire mechanical cost of a pass.
Mechanical share returned to 38% not because the wins reversed, but
because the denominator shrank while the misses did not. **Absolute
mechanical cost per pass fell from $51.76 to $6.13; mechanical cost per
response rose slightly, $0.078 → $0.088.** Share is noisy at n=70 against
a $16 denominator and should not be read as a trend.
**A regression to report:** `SH-3` batching is **0.0%** this pass — 67
tool calls across 67 responses, none batched, against a 20% target. Mean
context 315,170 against a 200,000 target. Both worse than the pass before.
Nothing in this workplan addressed session shape, and nothing gates it.
## Verdict
| claim | status |
|---|---|
| denominators widened, and the number moved | **confirmed** |
| kernel rules became visible (K10, K14, K18) | **confirmed** |
| the record was corrected, five rows | **confirmed** |
| M-D1-MUT ≥10 of 14 | **unmet — 4 of 14** |
| the diagnosis in CB-RES-0004 was complete | **refuted** — 6 more rows, 3 new defects, 1 new class |
| CB-WP-0004's "remove the manual path" mechanism | **confirmed on a clean window** |
| session shape | **regressed, ungated** |
The pass did not finish what it planned. It stopped because the
instrument it built contradicted the plan — which is the outcome the stop
condition existed to produce, and the first time this loop has spent money
to be told it was wrong and then acted on it.