clay-borg/evidence
tegwick ee37b82675
Some checks failed
ci / check (push) Failing after 3s
CB-WP-0015: the two inert clauses, AM-7 scaling and AM-8 N=10
Provenance (tier S, one paragraph in lieu of survey and ADR): the two
clauses mutation-check has reported inert since CB-WP-0005. AM-7's
scaling ratio was held up by a test literally named
replay_100k_events_is_linear_and_fast that computed both throughputs,
printed both, and never divided one by the other. AM-8's N=10 was held
up by a runner that does two.

Both are now red. AM-7 3/3, AM-8 2/2, M-D1-MUT 10/14, and ADR-0005's
>=10-of-14 prediction MET for the first time. Neither was closed by
amending the question away, which was the live risk: the denominator is
unchanged and the four unenforced rows are the four already
unenforceable.

AM-7 needed three estimators. Best-of-N per leg then divide (AM-6's,
correct for a floor on one number) gave 0.581-1.085 on an unchanged
binary; legs back-to-back gave medians 0.931-1.004; legs interleaved at
fold granularity give 0.987/0.991/0.989, and 0.989 under 8-way CPU
contention while absolute throughput fell 4x. The INDETERMINATE guard
demanded unanimity and failed a good measurement over one sample
0.001 under the floor; it now requires a two-thirds majority. The
control that matters: AM-6's constant-cost mutation halves throughput
and leaves this ratio at 0.999x green, so AM-7 is not a second AM-6.

AM-8 kept N=10 because the measurement said so. Perturbing the RNG only
from its fourth construction on: --runs 2 PASSES, --runs 10 fails. A
late-onset divergence is deterministic, not flaky, so it is a control
rather than a coin flip. Ten runs live on one scenario (make am8, ~2s)
rather than all 25 (47s a build). GameKernel 5b records it.

The full run also found AM-4a's own mutation stale since ADR-0008 D3
moved the target 250,000 -> 161,000 in CB-WP-0013 -- reported
HARNESS-BROKEN, no score published. The build-free half of that check
is now a --self-test assertion, so make all catches the next one.

mutation-check clauses may now carry their own verify and mutation, and
then the enforced flag is measured rather than declared; a declaration
disagreeing with its measurement is refused.

make all exits 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 14:07:08 +02:00
..
CB-EV-0001-game-kernel.md CB-WP-0013-T02/T03: retire SH-3 as a gate; correct AM-4a and its target 2026-08-02 07:30:14 +02:00
CB-EV-0002-cost-accounting.md CB-WP-0004 T04: fact registry and make facts-check — DFD gets a gate 2026-07-31 10:24:39 +02:00
CB-EV-0003-mechanical-work.md CB-WP-0004 T05: control loop — 6 points recovered, not 25-30 2026-07-31 10:27:09 +02:00
CB-EV-0004-assertion-coverage.md CB-WP-0005 T07/T08: control loop, retrospective, InnerLoop v1.4 2026-07-31 18:05:31 +02:00
CB-EV-0005-instrument-the-table.md CB-WP-0006 T08: control loop — 4 of 14 to 8 of 14, and a cost regression 2026-08-01 13:39:12 +02:00
CB-EV-0007-stage-0.md CB-WP-0008-T04: CB-EV-0007 — stage 0 is shipped, and the spend curve is a V 2026-08-01 15:25:19 +02:00
CB-EV-0008-adaptive-gates.md CB-WP-0009-T04: CB-EV-0008 — the gate changes, measured 2026-08-01 15:45:08 +02:00
CB-EV-0009-inspectable-table.md CB-WP-0011-T03: evidence — the chaos roll fired, and it was right 2026-08-02 02:48:51 +02:00
CB-EV-0010-render-port.md CB-WP-0012-T05: evidence — tier L deleted its own deliverable 2026-08-02 04:30:10 +02:00
CB-EV-0011-instrument-corrections.md CB-WP-0013-T04: evidence — and a rule about quoting your own cost 2026-08-02 07:32:07 +02:00
CB-EV-0012-execute-the-javascript.md CB-WP-0014-T03: hot-seat evidenced; stage 1 open on one human check 2026-08-02 07:58:59 +02:00
CB-EV-0013-the-inert-clauses.md CB-WP-0015: the two inert clauses, AM-7 scaling and AM-8 N=10 2026-08-02 14:07:08 +02:00