CB-WP-0006 T01: assert the AM-6 throughput target

Nothing in the workspace compared any number to 100,000 events/s while
the evidence file reported "AM-6 | met, 16.5x". Now a test does — a test,
not a bench, because Criterion reports throughput and asserts nothing,
which is why this row measured nothing for six passes.

Measured on bnt-lap001: 341,280 ev/s in debug (3.4x the target), ~2.4-3.1M
in release. The spec target holds even in an unoptimized build, so the
gate needs no cfg split and runs in the ordinary `make test`.

The trap this task named — loosening a flaky timing assertion until it
never fires — is avoided by construction. The threshold is the spec value,
untouched; the constant says lowering it requires an ADR; and the failure
message repeats that, states measured headroom, and names reference
figures, so an agent hitting a red AM-6 is told not to tune it in the
place they are actually reading. Robustness comes from best-of-N, not from
a lower bar: a throughput floor asks whether the machine is capable, so
transient load should not fail the build.

Two positive controls in the test: a run that applied fewer than 50,000
events, or measured zero elapsed time, fails rather than scoring as
infinite throughput.

Verified by a PROPERTY mutation — 4,000 black_box iterations injected into
GroundState::fold, the hot path — not a threshold tweak, which would only
prove the comparison runs.

And the FA class found last pass is now gated. mutation-check rows gained
an `expect` field: the mutant's output must contain the row's stated
failure string or the verdict is WRONG-REASON, not red. Without it a
mutation that merely failed to compile would credit its row with an
assertion it does not have. Verified by pointing expect at a string the
verifier never prints and watching the verdict flip. This is remedy (2)
from the CB-WP-0005 retrospective, built a task earlier than planned
because the class it guards is the newest and most dangerous.

M-D1-MUT: 4 -> 5 of 14.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
tegwick 2026-07-31 18:32:16 +02:00
parent ba7c2f88ae
commit c43754f0fe
7 changed files with 159 additions and 14 deletions

View file

@ -63,6 +63,45 @@ names.
**Verified by:** `make mutation-check --row AM-6` goes from `unmutatable`
to `red`.
**Delivered.** `am6_throughput_clears_the_spec_target` in
`games/ground/src/lib.rs` — a test, not a bench. Best of 3 samples of
50,000 applied events each.
Measured on bnt-lap001 2026-07-31: **341,280 ev/s in debug (3.4× the
target)**, ~2.43.1M in release (~2430×). **The spec target holds even in
an unoptimized build**, so the gate needed no `cfg` split and runs in the
ordinary `make test`.
**The trap was avoided by construction, not by intention.** The threshold
is the spec value `100_000`, untouched; the constant carries a comment
saying lowering it requires an ADR; and the failure message repeats that,
states the measured headroom, and names the reference figures — so a
future agent hitting a red AM-6 is told not to tune it, in the place they
will actually be reading. Robustness comes from **best-of-N**, not from a
lower bar: a throughput *floor* asks "is this machine capable", so
transient load should not fail the build.
Two positive controls in the test itself: a run that applied fewer than
50,000 events, or measured zero elapsed time, fails rather than scoring as
infinite throughput.
**Verified:** `make mutation-check --row AM-6`**red**, via a *property*
mutation (4,000 `black_box` iterations injected into `GroundState::fold`,
the hot path) rather than a threshold tweak — raising the target would only
prove the comparison runs.
**And the FA class is now gated.** `mutation-check` rows gained an
`expect` field: the mutant's output must contain the row's stated failure
string, or the verdict is **`WRONG-REASON`**, not `red`. Without it, a
mutation that merely failed to compile would credit its row with an
assertion it does not have. Verified by pointing `expect` at a string the
verifier never prints and confirming the verdict flips. This is
remedy (2) from the CB-WP-0005 retrospective, built one task earlier than
T08 planned because the class it guards is the newest and the most
dangerous.
**M-D1-MUT: 4 → 5 of 14.**
## Task: AM-2, AM-3 — instrument the size metrics
```task