CB-WP-0006 T01: assert the AM-6 throughput target
Nothing in the workspace compared any number to 100,000 events/s while the evidence file reported "AM-6 | met, 16.5x". Now a test does — a test, not a bench, because Criterion reports throughput and asserts nothing, which is why this row measured nothing for six passes. Measured on bnt-lap001: 341,280 ev/s in debug (3.4x the target), ~2.4-3.1M in release. The spec target holds even in an unoptimized build, so the gate needs no cfg split and runs in the ordinary `make test`. The trap this task named — loosening a flaky timing assertion until it never fires — is avoided by construction. The threshold is the spec value, untouched; the constant says lowering it requires an ADR; and the failure message repeats that, states measured headroom, and names reference figures, so an agent hitting a red AM-6 is told not to tune it in the place they are actually reading. Robustness comes from best-of-N, not from a lower bar: a throughput floor asks whether the machine is capable, so transient load should not fail the build. Two positive controls in the test: a run that applied fewer than 50,000 events, or measured zero elapsed time, fails rather than scoring as infinite throughput. Verified by a PROPERTY mutation — 4,000 black_box iterations injected into GroundState::fold, the hot path — not a threshold tweak, which would only prove the comparison runs. And the FA class found last pass is now gated. mutation-check rows gained an `expect` field: the mutant's output must contain the row's stated failure string or the verdict is WRONG-REASON, not red. Without it a mutation that merely failed to compile would credit its row with an assertion it does not have. Verified by pointing expect at a string the verifier never prints and watching the verdict flip. This is remedy (2) from the CB-WP-0005 retrospective, built a task earlier than planned because the class it guards is the newest and most dangerous. M-D1-MUT: 4 -> 5 of 14. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
ba7c2f88ae
commit
c43754f0fe
7 changed files with 159 additions and 14 deletions
|
|
@ -63,6 +63,45 @@ names.
|
|||
**Verified by:** `make mutation-check --row AM-6` goes from `unmutatable`
|
||||
to `red`.
|
||||
|
||||
**Delivered.** `am6_throughput_clears_the_spec_target` in
|
||||
`games/ground/src/lib.rs` — a test, not a bench. Best of 3 samples of
|
||||
50,000 applied events each.
|
||||
|
||||
Measured on bnt-lap001 2026-07-31: **341,280 ev/s in debug (3.4× the
|
||||
target)**, ~2.4–3.1M in release (~24–30×). **The spec target holds even in
|
||||
an unoptimized build**, so the gate needed no `cfg` split and runs in the
|
||||
ordinary `make test`.
|
||||
|
||||
**The trap was avoided by construction, not by intention.** The threshold
|
||||
is the spec value `100_000`, untouched; the constant carries a comment
|
||||
saying lowering it requires an ADR; and the failure message repeats that,
|
||||
states the measured headroom, and names the reference figures — so a
|
||||
future agent hitting a red AM-6 is told not to tune it, in the place they
|
||||
will actually be reading. Robustness comes from **best-of-N**, not from a
|
||||
lower bar: a throughput *floor* asks "is this machine capable", so
|
||||
transient load should not fail the build.
|
||||
|
||||
Two positive controls in the test itself: a run that applied fewer than
|
||||
50,000 events, or measured zero elapsed time, fails rather than scoring as
|
||||
infinite throughput.
|
||||
|
||||
**Verified:** `make mutation-check --row AM-6` → **red**, via a *property*
|
||||
mutation (4,000 `black_box` iterations injected into `GroundState::fold`,
|
||||
the hot path) rather than a threshold tweak — raising the target would only
|
||||
prove the comparison runs.
|
||||
|
||||
**And the FA class is now gated.** `mutation-check` rows gained an
|
||||
`expect` field: the mutant's output must contain the row's stated failure
|
||||
string, or the verdict is **`WRONG-REASON`**, not `red`. Without it, a
|
||||
mutation that merely failed to compile would credit its row with an
|
||||
assertion it does not have. Verified by pointing `expect` at a string the
|
||||
verifier never prints and confirming the verdict flips. This is
|
||||
remedy (2) from the CB-WP-0005 retrospective, built one task earlier than
|
||||
T08 planned because the class it guards is the newest and the most
|
||||
dangerous.
|
||||
|
||||
**M-D1-MUT: 4 → 5 of 14.**
|
||||
|
||||
## Task: AM-2, AM-3 — instrument the size metrics
|
||||
|
||||
```task
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue