T07 — CB-EV-0004. 70 responses, $16.03. Test 1 MET: 1 spec -> 2, 58 rules -> 76, 1 source file -> 10, and the new denominator came in at 83% naming K10/K14/K18. The stated failure mode — a widened denominator reporting the same percentage — did not occur. Test 2 UNMET: M-D1-MUT 4 of 14 against >=10, and wrong about what as well as how much. The diagnosis was three absent kernel rules; the measurement found eight rows with no instrument at all. Test 3: quality held, and widening surfaced far more than the seven known defects — six further uninstrumented rows, HDN #7 (rule-coverage's self-test green while the tool was broken), a stale $248.46 invisible to facts-check because it was untagged, and a fifth error class. The clean test CB-WP-0004 was owed is now run, on a pass that used the tools without building them. The two categories whose tools removed the manual path are at 0 turns two passes on; the two that merely offered a better option are now the entire mechanical cost of a pass. Absolute mechanical cost fell $51.76 -> $6.13. The mechanism holds. Reported because nothing else would: SH-3 batching is 0.0% this pass — 67 tool calls across 67 responses — against a 20% target, and mean context 315,170 against 200,000. SessionShape has stated these since CB-WP-0003 and none has ever been enforced. T08 — the retrospective answers its question: yes, a weak mutation is the new grep, and it is worse, because it fails in the opposite direction. mutation-check's first run produced two SURVIVED verdicts and both were the author's own no-op mutations. The fifth error class: false accusation. HDN, TA, SSB and DFD all under-report — a real problem passes. FA over-reports: it publishes the claim that working code is broken, sends the next pass to fix something that is not broken, and is more credible than the truth because it arrives with a measurement attached. Thirteen instances, five classes, six passes, and the newest class is one that hardening created. One correction to CB-WP-0004 T06: "a gate only pays if it removes the manual path" is a predictor of whether a gate saves money, not a criterion for whether it is worth having. mutation-check fails that test and produced the most valuable findings of the pass. InnerLoop v1.4: where a claim rests on numbers, the adversarial reviewer must read the assertion behind each quoted number and mutate it. Re-running the command that prints a number is not verification of that number. Second verification step to inherit the author's blindness; both fixes replace re-derivation with adversarial execution. CB-WP-0005 status -> done, 5 of 8 tasks, 3 cancelled into CB-WP-0006. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
7.9 KiB
CB-EV-0004: did widening the instruments change what they report?
research: CB-RES-0004
adr: ADR-0005
workplan: CB-WP-0005
instruments: make coverage, make mutation-check, make cost-mix
window: cb-cost --since bd4423a — 70 responses, $16.03
Phase C was deferred unstarted, so this closes T01–T03 only. That is itself the headline result: the pass stopped because its own instrument contradicted the plan it was executing.
Test 1 — did the denominators widen?
| before | after | |
|---|---|---|
| specs in scope | 1 (GroundRules.md) |
2 (+ GameKernel.md) |
| numbered rules in scope | 58 | 76 |
| source files searched for links | 1 | 10 |
| denominators reported | 1 | 3 |
AM-1 rule coverage: 58/58 (100%) over 21 scenarios
AM-1b spec->code link: 49/58 claimed rules named in the aggregate
AM-1b kernel spec->code link: 15/18 (83%) across 10 source files
unlinked: K10 K14 K18
Met. The workplan's stated failure mode — a widened denominator that reports the same percentage would prove the instruments still count names — did not occur. The new denominator came in at 83%, and named three rules that four workplans of green gates never mentioned.
The gate reports without feeding the exit code until 2026-08-31, then
binds. The date is in tools/rule-coverage.py, the days remaining print
on every run, and the self-test asserts the arm returns 0 before that date
and 2 after — so it cannot quietly become never, which is how AM-4's
targets went unratified for four workplans.
Test 2 — M-D1-MUT against the prediction
Prediction ≥10 of 14 (the ADR's 9-of-12, 75%). Measured 4. UNMET. No target moved in the commit that measured it.
M-D1-MUT: 4/14 rows enforced
PARTIAL 2 AM-7, AM-8 — some clauses live, some inert
unmutatable 8 no property to invert
SURVIVED 0
Two corrections to our own numbers, both recorded rather than quietly absorbed:
- There are 14 acceptance rows, not 12. ADR-0005 and CB-WP-0005 both said twelve; AM-4 splits into a/b/c. The prediction is evaluated as ≥10 of 14 on the same 75% basis.
- The first run reported two
SURVIVEDrows and both were my own no-op mutations.pub struct NullRng;→pub struct NullRng {}is semantically identical; renamingmax_age_dayschanges nothing because CA-17 reads it with.get(..., 90). Replaced with real inversions, after which both go red.
The eight unmutatable rows, each with the reason the harness records:
| row | why nothing can be inverted |
|---|---|
| AM-6 | nothing compares any number to 100,000 events/s — the headline throughput claim |
| AM-2 | no instrument divides LOC by rule count or compares to 40 |
| AM-3 | the synthetic workload's LOC is never measured |
| AM-4c | reported, not targeted — no threshold, so nothing can fail |
| AM-5 | recorded not gated, and not recorded either |
| AM-9 | nothing measures resident memory |
| AM-10 | population empty — no cb-*-api crate exists |
| AM-11 | the conformance suite the metric is a bool over does not exist |
The prediction was wrong about what, not only how much. CB-RES-0004 diagnosed three absent kernel rules. The measurement found that more than half the acceptance table has no instrument at all — which Phase C, scoped to five rules, would not have touched.
Test 3 — did quality hold, and did widening surface anything new?
make all green throughout, now including two gates that did not exist:
the kernel coverage arm and mutation-check --self-test.
It surfaced substantially more than the seven defects CB-RES-0004 named. Six further acceptance rows with no instrument (above), plus three new defect instances:
- HDN #7 —
rule-coverage.py's--self-testprinted all-ok while every realmake coveragedied withRecursionError. Aprint(insidesay()had becomesay(; the control only ever called the quiet path. The control named the behaviour and did not assert it — the pass's own thesis, inside the tool written to prove it. - A new error class (below) — the two no-op mutations.
- DFD —
evidence/CB-EV-0001still carried $248.46 for AM-12, the figure CB-WP-0002 disproved four workplans ago. Untagged, and therefore invisible tofacts-check. The gate built to catch duplicated facts only checks copies that opted in.
The fifth error class: false accusation (FA)
Every class on record under-reports: a harness that does nothing, trusted arithmetic, a blind sample, a stale copy. All four let a real problem pass.
The weak mutation is the first that over-reports. A no-op mutation
yields SURVIVED, which reads as "this row asserts nothing" — a
published claim that working code is broken. It sends the next pass to fix
something that is not broken, and it is more credible than the truth
because it arrives with a measurement attached.
| class | direction | caught by |
|---|---|---|
| HDN, TA, SSB, DFD | under-report — a real defect passes | assertions, re-derivation, all-data checks, reading the copy |
| FA (new) | over-report — working code is indicted | checking that the mutation fails for its stated reason |
Ten instances across four classes became thirteen across five.
Cost, and the clean test CB-WP-0004 was owed
CB-WP-0004 T05 said its measurement was confounded because the pass built the tools it measured, and that "the clean test is the next pass, which uses the tools without building them." T01–T03 are largely that pass.
| category | baseline (662 resp) | CB-WP-0004 build | this pass (70 resp) |
|---|---|---|---|
| environment setup | 85 turns, $15.58 | 3 turns, $0.20 | 0 turns |
| hub task status + workplan edit | 46 turns, $11.52 | 1 turn, $0.13 | 0 turns |
| orientation / inspect | 49 turns, $6.87 | 41 turns, $4.10 | 7 turns, $3.85 |
| ad-hoc text patching | 76 turns, $14.11 | 8 turns, $0.75 | 11 turns, $2.28 |
| mechanical share of pass | 38.2% | 26% | 38% |
The verdict is clean, and it confirms CB-WP-0004 T06's mechanism
exactly. The two categories whose tools removed the manual path went
to zero and stayed there with no further work. The two whose tools
merely offered a better option — make status and facts-check — are
now the entire mechanical cost of a pass.
Mechanical share returned to 38% not because the wins reversed, but because the denominator shrank while the misses did not. Absolute mechanical cost per pass fell from $51.76 to $6.13; mechanical cost per response rose slightly, $0.078 → $0.088. Share is noisy at n=70 against a $16 denominator and should not be read as a trend.
A regression to report: SH-3 batching is 0.0% this pass — 67
tool calls across 67 responses, none batched, against a 20% target. Mean
context 315,170 against a 200,000 target. Both worse than the pass before.
Nothing in this workplan addressed session shape, and nothing gates it.
Verdict
| claim | status |
|---|---|
| denominators widened, and the number moved | confirmed |
| kernel rules became visible (K10, K14, K18) | confirmed |
| the record was corrected, five rows | confirmed |
| M-D1-MUT ≥10 of 14 | unmet — 4 of 14 |
| the diagnosis in CB-RES-0004 was complete | refuted — 6 more rows, 3 new defects, 1 new class |
| CB-WP-0004's "remove the manual path" mechanism | confirmed on a clean window |
| session shape | regressed, ungated |
The pass did not finish what it planned. It stopped because the instrument it built contradicted the plan — which is the outcome the stop condition existed to produce, and the first time this loop has spent money to be told it was wrong and then acted on it.