clay-borg/evidence/CB-EV-0004-assertion-coverage.md
tegwick fd19f4e878 CB-WP-0005 T07/T08: control loop, retrospective, InnerLoop v1.4
T07 — CB-EV-0004. 70 responses, $16.03.

Test 1 MET: 1 spec -> 2, 58 rules -> 76, 1 source file -> 10, and the new
denominator came in at 83% naming K10/K14/K18. The stated failure mode — a
widened denominator reporting the same percentage — did not occur.

Test 2 UNMET: M-D1-MUT 4 of 14 against >=10, and wrong about what as well
as how much. The diagnosis was three absent kernel rules; the measurement
found eight rows with no instrument at all.

Test 3: quality held, and widening surfaced far more than the seven known
defects — six further uninstrumented rows, HDN #7 (rule-coverage's
self-test green while the tool was broken), a stale $248.46 invisible to
facts-check because it was untagged, and a fifth error class.

The clean test CB-WP-0004 was owed is now run, on a pass that used the
tools without building them. The two categories whose tools removed the
manual path are at 0 turns two passes on; the two that merely offered a
better option are now the entire mechanical cost of a pass. Absolute
mechanical cost fell $51.76 -> $6.13. The mechanism holds.

Reported because nothing else would: SH-3 batching is 0.0% this pass — 67
tool calls across 67 responses — against a 20% target, and mean context
315,170 against 200,000. SessionShape has stated these since CB-WP-0003
and none has ever been enforced.

T08 — the retrospective answers its question: yes, a weak mutation is the
new grep, and it is worse, because it fails in the opposite direction.
mutation-check's first run produced two SURVIVED verdicts and both were
the author's own no-op mutations.

The fifth error class: false accusation. HDN, TA, SSB and DFD all
under-report — a real problem passes. FA over-reports: it publishes the
claim that working code is broken, sends the next pass to fix something
that is not broken, and is more credible than the truth because it arrives
with a measurement attached. Thirteen instances, five classes, six passes,
and the newest class is one that hardening created.

One correction to CB-WP-0004 T06: "a gate only pays if it removes the
manual path" is a predictor of whether a gate saves money, not a criterion
for whether it is worth having. mutation-check fails that test and
produced the most valuable findings of the pass.

InnerLoop v1.4: where a claim rests on numbers, the adversarial reviewer
must read the assertion behind each quoted number and mutate it.
Re-running the command that prints a number is not verification of that
number. Second verification step to inherit the author's blindness; both
fixes replace re-derivation with adversarial execution.

CB-WP-0005 status -> done, 5 of 8 tasks, 3 cancelled into CB-WP-0006.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 18:05:31 +02:00

7.9 KiB
Raw Blame History

CB-EV-0004: did widening the instruments change what they report?

research: CB-RES-0004 adr: ADR-0005 workplan: CB-WP-0005 instruments: make coverage, make mutation-check, make cost-mix window: cb-cost --since bd4423a — 70 responses, $16.03

Phase C was deferred unstarted, so this closes T01T03 only. That is itself the headline result: the pass stopped because its own instrument contradicted the plan it was executing.


Test 1 — did the denominators widen?

before after
specs in scope 1 (GroundRules.md) 2 (+ GameKernel.md)
numbered rules in scope 58 76
source files searched for links 1 10
denominators reported 1 3
AM-1  rule coverage:              58/58 (100%) over 21 scenarios
AM-1b spec->code link:            49/58 claimed rules named in the aggregate
AM-1b kernel spec->code link:     15/18 (83%) across 10 source files
      unlinked: K10 K14 K18

Met. The workplan's stated failure mode — a widened denominator that reports the same percentage would prove the instruments still count names — did not occur. The new denominator came in at 83%, and named three rules that four workplans of green gates never mentioned.

The gate reports without feeding the exit code until 2026-08-31, then binds. The date is in tools/rule-coverage.py, the days remaining print on every run, and the self-test asserts the arm returns 0 before that date and 2 after — so it cannot quietly become never, which is how AM-4's targets went unratified for four workplans.

Test 2 — M-D1-MUT against the prediction

Prediction ≥10 of 14 (the ADR's 9-of-12, 75%). Measured 4. UNMET. No target moved in the commit that measured it.

  M-D1-MUT: 4/14 rows enforced
    PARTIAL       2      AM-7, AM-8 — some clauses live, some inert
    unmutatable   8      no property to invert
    SURVIVED      0

Two corrections to our own numbers, both recorded rather than quietly absorbed:

  1. There are 14 acceptance rows, not 12. ADR-0005 and CB-WP-0005 both said twelve; AM-4 splits into a/b/c. The prediction is evaluated as ≥10 of 14 on the same 75% basis.
  2. The first run reported two SURVIVED rows and both were my own no-op mutations. pub struct NullRng;pub struct NullRng {} is semantically identical; renaming max_age_days changes nothing because CA-17 reads it with .get(..., 90). Replaced with real inversions, after which both go red.

The eight unmutatable rows, each with the reason the harness records:

row why nothing can be inverted
AM-6 nothing compares any number to 100,000 events/s — the headline throughput claim
AM-2 no instrument divides LOC by rule count or compares to 40
AM-3 the synthetic workload's LOC is never measured
AM-4c reported, not targeted — no threshold, so nothing can fail
AM-5 recorded not gated, and not recorded either
AM-9 nothing measures resident memory
AM-10 population empty — no cb-*-api crate exists
AM-11 the conformance suite the metric is a bool over does not exist

The prediction was wrong about what, not only how much. CB-RES-0004 diagnosed three absent kernel rules. The measurement found that more than half the acceptance table has no instrument at all — which Phase C, scoped to five rules, would not have touched.

Test 3 — did quality hold, and did widening surface anything new?

make all green throughout, now including two gates that did not exist: the kernel coverage arm and mutation-check --self-test.

It surfaced substantially more than the seven defects CB-RES-0004 named. Six further acceptance rows with no instrument (above), plus three new defect instances:

  • HDN #7rule-coverage.py's --self-test printed all-ok while every real make coverage died with RecursionError. A print( inside say() had become say(; the control only ever called the quiet path. The control named the behaviour and did not assert it — the pass's own thesis, inside the tool written to prove it.
  • A new error class (below) — the two no-op mutations.
  • DFDevidence/CB-EV-0001 still carried $248.46 for AM-12, the figure CB-WP-0002 disproved four workplans ago. Untagged, and therefore invisible to facts-check. The gate built to catch duplicated facts only checks copies that opted in.

The fifth error class: false accusation (FA)

Every class on record under-reports: a harness that does nothing, trusted arithmetic, a blind sample, a stale copy. All four let a real problem pass.

The weak mutation is the first that over-reports. A no-op mutation yields SURVIVED, which reads as "this row asserts nothing" — a published claim that working code is broken. It sends the next pass to fix something that is not broken, and it is more credible than the truth because it arrives with a measurement attached.

class direction caught by
HDN, TA, SSB, DFD under-report — a real defect passes assertions, re-derivation, all-data checks, reading the copy
FA (new) over-report — working code is indicted checking that the mutation fails for its stated reason

Ten instances across four classes became thirteen across five.

Cost, and the clean test CB-WP-0004 was owed

CB-WP-0004 T05 said its measurement was confounded because the pass built the tools it measured, and that "the clean test is the next pass, which uses the tools without building them." T01T03 are largely that pass.

category baseline (662 resp) CB-WP-0004 build this pass (70 resp)
environment setup 85 turns, $15.58 3 turns, $0.20 0 turns
hub task status + workplan edit 46 turns, $11.52 1 turn, $0.13 0 turns
orientation / inspect 49 turns, $6.87 41 turns, $4.10 7 turns, $3.85
ad-hoc text patching 76 turns, $14.11 8 turns, $0.75 11 turns, $2.28
mechanical share of pass 38.2% 26% 38%

The verdict is clean, and it confirms CB-WP-0004 T06's mechanism exactly. The two categories whose tools removed the manual path went to zero and stayed there with no further work. The two whose tools merely offered a better optionmake status and facts-check — are now the entire mechanical cost of a pass.

Mechanical share returned to 38% not because the wins reversed, but because the denominator shrank while the misses did not. Absolute mechanical cost per pass fell from $51.76 to $6.13; mechanical cost per response rose slightly, $0.078 → $0.088. Share is noisy at n=70 against a $16 denominator and should not be read as a trend.

A regression to report: SH-3 batching is 0.0% this pass — 67 tool calls across 67 responses, none batched, against a 20% target. Mean context 315,170 against a 200,000 target. Both worse than the pass before. Nothing in this workplan addressed session shape, and nothing gates it.

Verdict

claim status
denominators widened, and the number moved confirmed
kernel rules became visible (K10, K14, K18) confirmed
the record was corrected, five rows confirmed
M-D1-MUT ≥10 of 14 unmet — 4 of 14
the diagnosis in CB-RES-0004 was complete refuted — 6 more rows, 3 new defects, 1 new class
CB-WP-0004's "remove the manual path" mechanism confirmed on a clean window
session shape regressed, ungated

The pass did not finish what it planned. It stopped because the instrument it built contradicted the plan — which is the outcome the stop condition existed to produce, and the first time this loop has spent money to be told it was wrong and then acted on it.