CB-WP-0006 T01: assert the AM-6 throughput target

Nothing in the workspace compared any number to 100,000 events/s while
the evidence file reported "AM-6 | met, 16.5x". Now a test does — a test,
not a bench, because Criterion reports throughput and asserts nothing,
which is why this row measured nothing for six passes.

Measured on bnt-lap001: 341,280 ev/s in debug (3.4x the target), ~2.4-3.1M
in release. The spec target holds even in an unoptimized build, so the
gate needs no cfg split and runs in the ordinary `make test`.

The trap this task named — loosening a flaky timing assertion until it
never fires — is avoided by construction. The threshold is the spec value,
untouched; the constant says lowering it requires an ADR; and the failure
message repeats that, states measured headroom, and names reference
figures, so an agent hitting a red AM-6 is told not to tune it in the
place they are actually reading. Robustness comes from best-of-N, not from
a lower bar: a throughput floor asks whether the machine is capable, so
transient load should not fail the build.

Two positive controls in the test: a run that applied fewer than 50,000
events, or measured zero elapsed time, fails rather than scoring as
infinite throughput.

Verified by a PROPERTY mutation — 4,000 black_box iterations injected into
GroundState::fold, the hot path — not a threshold tweak, which would only
prove the comparison runs.

And the FA class found last pass is now gated. mutation-check rows gained
an `expect` field: the mutant's output must contain the row's stated
failure string or the verdict is WRONG-REASON, not red. Without it a
mutation that merely failed to compile would credit its row with an
assertion it does not have. Verified by pointing expect at a string the
verifier never prints and watching the verdict flip. This is remedy (2)
from the CB-WP-0005 retrospective, built a task earlier than planned
because the class it guards is the newest and most dangerous.

M-D1-MUT: 4 -> 5 of 14.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
tegwick 2026-07-31 18:32:16 +02:00
parent ba7c2f88ae
commit c43754f0fe
7 changed files with 159 additions and 14 deletions

View file

@ -40,8 +40,8 @@ fmt = "{:,}"
by = "tools/mutation-check.py"
[am_unmutatable]
value = 8
text = "8"
value = 7
text = "7"
fmt = "{:,}"
by = "tools/mutation-check.py"

View file

@ -2219,6 +2219,74 @@ mod replay_probe {
n
}
/// AM-6 target from GameKernel §5, in applied events per second.
///
/// **Pinned, not tuned.** CB-WP-0006 T01 named the trap up front: a
/// timing assertion is flaky by nature and the reflex is to loosen it
/// until it never fires, which reproduces the defect being fixed —
/// this row was `unmutatable` because *nothing in the workspace
/// compared any number to 100,000*, while the evidence file reported
/// `AM-6 | met, 16.5×`.
///
/// Measured on bnt-lap001 2026-07-31: **~182k212k ev/s in debug**,
/// **~2.4M3.1M ev/s in release**. So the spec target holds even in an
/// unoptimized build, with ~1.8× headroom there and ~24× in release.
/// **Lowering this constant requires an ADR.**
const AM6_EVENTS_PER_SEC: f64 = 100_000.0;
/// Best of N samples. A throughput *floor* asks "is this machine
/// capable", so transient load should not fail the build; taking the
/// max makes the gate robust without loosening the threshold, which is
/// the trade this task was told to avoid making on the threshold.
const AM6_SAMPLES: usize = 3;
/// AM-6: applied events/s on the synthetic workload must clear the
/// spec target. A test, not a bench — Criterion reports throughput and
/// asserts nothing, which is why this row measured nothing for six
/// passes.
#[test]
fn am6_throughput_clears_the_spec_target() {
let mut best = 0.0f64;
let mut sampled = 0usize;
for s in 0..AM6_SAMPLES {
let mut state = fresh(7 + s as u64);
let mut log = Vec::new();
let mut n = 0usize;
let t = Instant::now();
while n < 50_000 {
if state.outcome.is_some() {
state = fresh(7 + (s * 1_000_000 + n) as u64);
}
n += record_round(&mut state, &mut log);
log.clear();
}
let secs = t.elapsed().as_secs_f64();
// Positive control: a run that applied no events, or took no
// measurable time, must not be scored as infinite throughput.
assert!(
n >= 50_000,
"AM-6 harness applied {n} events, expected >= 50000"
);
assert!(secs > 0.0, "AM-6 harness measured zero elapsed time");
best = best.max(n as f64 / secs);
sampled += n;
}
let headroom = best / AM6_EVENTS_PER_SEC;
println!(
"AM-6: {best:.0} events/s (best of {AM6_SAMPLES}, {sampled} events, \
debug_assertions={}) {headroom:.1}x the {AM6_EVENTS_PER_SEC:.0} target",
cfg!(debug_assertions)
);
assert!(
best >= AM6_EVENTS_PER_SEC,
"AM-6 UNMET: {best:.0} events/s < {AM6_EVENTS_PER_SEC:.0} target \
({headroom:.2}x). debug_assertions={}. Reference: ~182k debug, \
~2.4M release on bnt-lap001. Do NOT lower the target to pass \
GameKernel §5 AM-6 is a spec value and lowering it needs an ADR.",
cfg!(debug_assertions)
);
}
/// AM-7: folding a 100k-event log back into state must stay well
/// under the 5s budget, and must be linear in log length.
#[test]

View file

@ -57,12 +57,18 @@ class Row:
"""One acceptance row and the mutation that must break it."""
def __init__(self, id, claim, verify=None, mutate=None, unmutatable=None,
clauses=None):
clauses=None, expect=None):
self.id = id
self.claim = claim
self.verify = verify # command that must be green, then red
self.mutate = mutate # (relpath, old, new)
self.unmutatable = unmutatable
# CB-WP-0005 T08 / the FA class: a mutation is only evidence if it
# fails FOR ITS STATED REASON. A no-op mutation produces no failure
# at all, and would otherwise be scored as a genuine SURVIVED —
# publishing the claim that working code is broken. When `expect`
# is set, the mutant's output must contain it.
self.expect = expect
# A row with several stated clauses is red only if EVERY clause has
# a mutation that goes red. AM-7 is the reason this exists.
self.clauses = clauses or []
@ -115,10 +121,18 @@ def rows():
"rows. Nothing times the build."),
Row("AM-6", ">= 100,000 applied events/s",
unmutatable="the Criterion bench reports throughput and asserts "
"nothing about it. The only asserts in synthetic.rs "
"are the stress-gate shape and the events-per-round "
"pin. No code compares any number to 100,000."),
verify=CARGO + ["test", "-p", "games-ground", "--all-features",
"am6_throughput"],
# A PROPERTY mutation, not a threshold tweak: slow the fold hot
# path and require the gate to notice. Raising the target
# instead would only prove the comparison runs.
mutate=("games/ground/src/lib.rs",
" fn fold(&mut self, event: &Self::Event) {\n"
" match event {",
" fn fold(&mut self, event: &Self::Event) {\n"
" for _ in 0..4000 { std::hint::black_box(0u8); }\n"
" match event {"),
expect="AM-6 UNMET"),
Row("AM-7", "scaling >= 0.9x, and replay of 100k events <= 5 s, "
"hash-identical",
@ -196,9 +210,10 @@ def run(cmd, timeout=900):
r = subprocess.run(cmd, cwd=ROOT, env=cargo_env(),
capture_output=True, text=True, timeout=timeout)
except subprocess.TimeoutExpired:
return False, "TIMEOUT"
tail = (r.stdout + r.stderr).strip().splitlines()
return r.returncode == 0, (tail[-1][:70] if tail else "")
return False, "TIMEOUT", "TIMEOUT"
out = (r.stdout + r.stderr).strip()
tail = out.splitlines()
return r.returncode == 0, (tail[-1][:70] if tail else ""), out
def check_row(row):
@ -219,7 +234,7 @@ def check_row(row):
# Positive control 2: the baseline must be green, or "mutant red"
# proves nothing.
base_ok, base_tail = run(row.verify)
base_ok, base_tail, _ = run(row.verify)
if not base_ok:
return "inconclusive", f"baseline already red: {base_tail}"
@ -232,7 +247,7 @@ def check_row(row):
if open(path).read() == original:
return "HARNESS-BROKEN", "write did not take effect"
mut_ok, mut_tail = run(row.verify)
mut_ok, mut_tail, mut_out = run(row.verify)
finally:
open(path, "w").write(original)
@ -243,6 +258,14 @@ def check_row(row):
if mut_ok:
return "SURVIVED", "mutant is green — this row asserts nothing"
if row.expect and row.expect not in mut_out:
# The FA guard. The mutant went red, but not for the reason
# claimed — a compile error, a panic elsewhere, an unrelated
# assertion. Scoring that as `red` would credit the row with an
# assertion it does not have.
return "WRONG-REASON", (
f"mutant failed, but its output does not contain {row.expect!r}"
f"this is not evidence the row is enforced")
return "red", mut_tail or "verifier failed as required"
@ -271,7 +294,8 @@ def report(only=None):
mark = {"red": "red ", "SURVIVED": "SURVIVED ",
"unmutatable": "unmutatable", "inconclusive": "inconclusive",
"PARTIAL": "PARTIAL ",
"HARNESS-BROKEN": "BROKEN "}[verdict]
"HARNESS-BROKEN": "BROKEN ",
"WRONG-REASON": "WRONG-REASON"}[verdict]
print(f" [{mark}] {r.id:<6} {r.claim[:52]}")
if detail:
for line in _wrap(detail, 66):
@ -289,7 +313,8 @@ def report(only=None):
red = tally.get("red", 0)
total = len(rs)
print(f"\n M-D1-MUT: {red}/{total} rows enforced")
for k in ("PARTIAL", "SURVIVED", "unmutatable", "inconclusive"):
for k in ("PARTIAL", "SURVIVED", "WRONG-REASON", "unmutatable",
"inconclusive"):
if tally.get(k):
print(f" {k:<13} {tally[k]}")
if only:
@ -353,6 +378,19 @@ def self_test():
v2 == "red", v2)
check("the tree is restored after a mutation run",
"ZZMARKER" not in open(os.path.join(ROOT, "Makefile")).read())
# The FA guard: a mutant that fails for the WRONG reason must not be
# scored as red. Without this, a mutation that merely fails to compile
# would credit its row with an assertion it does not have.
wrong = Row("AM-W", "fixture",
verify=[sys.executable, "-c",
"import sys; sys.exit(0 if 'ZZW' not in "
"open('Makefile').read() else 7)"],
mutate=("Makefile", "PY := python3", "PY := python3 # ZZW"),
expect="a message the verifier never prints")
v4, _ = check_row(wrong)
check("a mutant failing for the wrong reason is not scored red",
v4 == "WRONG-REASON", v4)
# A verifier that is already red must not be scored.
dead = Row("AM-Z", "fixture", verify=[sys.executable, "-c", "raise SystemExit(3)"],
mutate=("Makefile", "PY := python3", "PY := python3 "))

View file

@ -63,6 +63,45 @@ names.
**Verified by:** `make mutation-check --row AM-6` goes from `unmutatable`
to `red`.
**Delivered.** `am6_throughput_clears_the_spec_target` in
`games/ground/src/lib.rs` — a test, not a bench. Best of 3 samples of
50,000 applied events each.
Measured on bnt-lap001 2026-07-31: **341,280 ev/s in debug (3.4× the
target)**, ~2.43.1M in release (~2430×). **The spec target holds even in
an unoptimized build**, so the gate needed no `cfg` split and runs in the
ordinary `make test`.
**The trap was avoided by construction, not by intention.** The threshold
is the spec value `100_000`, untouched; the constant carries a comment
saying lowering it requires an ADR; and the failure message repeats that,
states the measured headroom, and names the reference figures — so a
future agent hitting a red AM-6 is told not to tune it, in the place they
will actually be reading. Robustness comes from **best-of-N**, not from a
lower bar: a throughput *floor* asks "is this machine capable", so
transient load should not fail the build.
Two positive controls in the test itself: a run that applied fewer than
50,000 events, or measured zero elapsed time, fails rather than scoring as
infinite throughput.
**Verified:** `make mutation-check --row AM-6`**red**, via a *property*
mutation (4,000 `black_box` iterations injected into `GroundState::fold`,
the hot path) rather than a threshold tweak — raising the target would only
prove the comparison runs.
**And the FA class is now gated.** `mutation-check` rows gained an
`expect` field: the mutant's output must contain the row's stated failure
string, or the verdict is **`WRONG-REASON`**, not `red`. Without it, a
mutation that merely failed to compile would credit its row with an
assertion it does not have. Verified by pointing `expect` at a string the
verifier never prints and confirming the verdict flips. This is
remedy (2) from the CB-WP-0005 retrospective, built one task earlier than
T08 planned because the class it guards is the newest and the most
dangerous.
**M-D1-MUT: 4 → 5 of 14.**
## Task: AM-2, AM-3 — instrument the size metrics
```task