clay-borg/tools
tegwick e7312b3b8a CB-WP-0006 T03: measure AM-5 and AM-9; AM-5 breaches
AM-9 is met and gated: 13.4 MB peak RSS against a 64 MB target, 4.8x
headroom, in `make all` via --fast. CB-EV-0001's "very unlikely to bind"
was right, but it is now measured rather than assumed, and verified red by
a property mutation (a 300 MB allocation in the workload).

AM-5 is BREACHED on both readings, on the machine the spec names:

  dev toolchain (default features)        87.0 s  [FAIL target <= 60 s]
  shipped runtime (--no-default-features) 61.3 s  [FAIL target <= 60 s]

bnt-lap001, 8 cores — a direct comparison, not a directional one. A row
declared "recorded not gated" and never recorded fails its own target by
45% on first measurement.

The tool reports and exits 0 because the spec says the row is ungated.
Gating it is a spec change needing an ADR; a tool that promotes itself is
how a target starts binding without anyone deciding it should. So AM-5
stays unmutatable — for the accurate reason now — and the breach is raised
as a maintainer decision: speed the build, move the target by ADR (arguing
why 60 s was wrong rather than why 87 s is convenient), or withdraw the
row.

The measurement itself had a real bug, found only by cross-validation.
getrusage(RUSAGE_CHILDREN) is a high-water mark across every reaped child,
so it attributed cargo's memory to the workload and reported 38.2 MB for a
run that used 12.3 MB — a 3x over-report that was plausible, passed its
target, and would have been published. Fixed with os.wait4, which returns
that specific child's rusage, and the self-test now cross-checks against
/usr/bin/time -v.

That is the false-accusation shape in the measurement layer rather than
the mutation layer: an instrument confidently reporting a number it had
not earned.

Also: the clean build measures into a throwaway CARGO_TARGET_DIR rather
than running `cargo clean`, so measuring the metric does not cost several
minutes of rebuild afterwards. A metric that punishes its own measurement
gets measured once and never again.

M-D1-MUT: 6 -> 7 of 14.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 18:59:32 +02:00
..
__pycache__ CB-WP-0006 T03: measure AM-5 and AM-9; AM-5 breaches 2026-07-31 18:59:32 +02:00
cb-sim CI: enforce every gate; close the silent-skip holes 2026-07-31 04:02:33 +02:00
cb-cost.py CB-WP-0004 T05: control loop — 6 points recovered, not 25-30 2026-07-31 10:27:09 +02:00
dep-weight.py CB-WP-0004 T01: fix environment friction at the root 2026-07-31 10:13:52 +02:00
facts.py CB-WP-0005 T02: M-D1-MUT — 4 of 14 acceptance rows are enforced 2026-07-31 17:27:23 +02:00
loop-lint.py CB-WP-0004 T01: fix environment friction at the root 2026-07-31 10:13:52 +02:00
mutation-check.py CB-WP-0006 T03: measure AM-5 and AM-9; AM-5 breaches 2026-07-31 18:59:32 +02:00
repo.py CB-WP-0004 T01: fix environment friction at the root 2026-07-31 10:13:52 +02:00
rule-coverage.py CB-WP-0005 T01: spec->code link over every numbered spec and every crate 2026-07-31 16:54:07 +02:00
runtime-metrics.py CB-WP-0006 T03: measure AM-5 and AM-9; AM-5 breaches 2026-07-31 18:59:32 +02:00
size-metrics.py CB-WP-0006 T02: instrument AM-2; report AM-3 blocked, with the argument 2026-07-31 18:38:15 +02:00
status.py CB-WP-0004 T03: make status — one-shot orientation 2026-07-31 10:20:27 +02:00
task-done.py chore: mark T01/T02 done (measured: $2.33 + $1.68, 46 responses, opus-5) 2026-07-31 10:18:34 +02:00