CB-WP-0021 T06: fix AM-7's measurement, not its floor
Some checks failed
ci / check (push) Has been cancelled

The row folded a 5,000-event log against a 100,000-event log and compared
throughputs, which confounds 'does cost per event grow with history'
(the property it claims) with 'does streaming a 20x longer Vec cost more
per element' (a memory-hierarchy fact true of any program). It measured
the second and reported it as the first: importing the edition enlarged
the aggregate and the ratio fell to 0.845 with the state bounded.

Corrected to time the SAME 5,000 events on a state at depth 0 and on a
state at depth 100,000. Equal windows, equal event mix, so the only
difference left is history depth.

  corrected: clean 1.004, mutated 0.589 (red)
  old:       clean 0.845 (red on healthy code), mutated 0.751

It also runs in 8.5s instead of timing out: the first version re-walked
the 100k prefix every repetition, 200M untimed folds per sample, which
under the mutation never finished. A control that cannot be run is not a
control. It now advances to depth once per sample and clones.

Two of my own measurements here were wrong and both were caught by
measuring again. A 2-minute timeout killed the shell line before its
restoring cp ran, so three readings were taken on MUTATED code -- I
diagnosed an event-mix confound that did not exist and 'fixed' it. The
fix is kept on its merits; the justification was fiction. And the probe
that proved state was bounded had checked four of eleven collections.

make all exits 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
tegwick 2026-08-04 01:10:38 +02:00
parent 2da19a49b7
commit 503966bce9
5 changed files with 155 additions and 82 deletions

View file

@ -2555,9 +2555,11 @@ mod replay_probe {
/// degrading to DNF at 100k. Lowering it requires an ADR.
const AM7_SCALING_FLOOR: f64 = 0.9;
/// The two sizes the spec names.
const AM7_SMALL: usize = 5_000;
const AM7_LARGE: usize = 100_000;
/// The window that is timed, and how deep the late one sits.
/// **Both timed windows are the same size** — that is the correction
/// (CB-WP-0021 T06); see `paired_ratio`.
const AM7_WINDOW: usize = 5_000;
const AM7_DEPTH: usize = 100_000;
/// Events applied **per leg** per sample.
///
@ -2595,72 +2597,89 @@ mod replay_probe {
/// verdict — without treating a single outlier as one.
const AM7_AGREEMENT: f64 = 2.0 / 3.0;
/// Fold `log` once from a fresh state, returning only the time inside
/// the fold loop.
/// Fold `n` events from `log` starting at `from`, on a state already
/// advanced to `from`, returning only the time inside the fold loop.
/// The state after folding `log[..depth]` — the history the window
/// will be folded on top of.
///
/// **Setup is outside the clock, and that is the first trap here.**
/// The small leg runs 20× more folds than the large one, so it pays
/// 20× more `fresh()` calls. Timing those would penalise the
/// denominator, inflate the ratio, and make the row pass for a reason
/// that has nothing to do with scaling.
fn fold_once(log: &[GroundEvent]) -> std::time::Duration {
/// Built **once per sample**, not once per repetition. The first
/// version re-walked the prefix every rep: 2,000 reps x 100,000
/// events is 200M untimed folds per sample, and under the
/// history-proportional mutation that is quadratic and never
/// finishes. A control that cannot be run is not a control.
fn state_at(log: &[GroundEvent], depth: usize) -> GroundState {
let mut state = fresh(42);
for e in &log[..depth] {
state.fold(e);
}
state
}
/// Fold `n` events from `at` onto a clone of `state`, returning only
/// the time inside the fold loop. The clone is outside the clock.
fn fold_window(
state: &GroundState,
log: &[GroundEvent],
at: usize,
n: usize,
) -> std::time::Duration {
let mut st = state.clone();
let t = Instant::now();
for event in log {
state.fold(event);
for e in &log[at..at + n] {
st.fold(e);
}
let dt = t.elapsed();
// Keep the fold from being optimised out without paying for a hash
// inside the timed region.
std::hint::black_box(&state);
std::hint::black_box(&st);
dt
}
/// One paired sample: the two legs **interleaved**, ratio taken inside.
/// One paired sample: the SAME window size at two history depths.
///
/// **Two estimators were wrong before this one, and both looked
/// reasonable.**
/// **The corrected measurement (CB-WP-0021 T06).** The previous one
/// folded a 5,000-event log and a 100,000-event log and compared
/// their throughputs, which confounds two different things:
///
/// The first took best-of-5 on each leg independently and divided —
/// the estimator AM-6 uses, correct there because AM-6 is a floor on a
/// single number and the question is "is this machine capable". For a
/// *ratio* it is wrong: the legs are measured at different moments and
/// the noise multiplies instead of cancelling. Five runs of an
/// unchanged binary gave **0.581 to 1.085**.
/// 1. does cost per event grow with how many events have already
/// been folded? — the property AM-7 claims; and
/// 2. does streaming a 20x longer `Vec` cost more per element? — a
/// memory-hierarchy fact true of any program.
///
/// The second ran the legs back to back inside one sample, expecting
/// the load to be common-mode. It was not enough: this machine's
/// absolute throughput wanders between **22 M and 53 M ev/s within a
/// single run**, and a 50 ms leg samples a point on that wander rather
/// than averaging it. Three runs gave medians 1.004 / 0.931 / 0.956 —
/// clustered near the true value but still straddling the floor.
/// It measured (2) and reported it as (1). Importing the edition data
/// enlarged the aggregate — four Problems instead of three, real
/// values — and the ratio fell 0.97 -> 0.845 against a 0.9 floor
/// **with the state provably bounded**: identical deck, discard,
/// Problem and hand sizes after 5k and 100k events. A row that fails
/// because the game got bigger, while the property it names is
/// untouched, is measuring the wrong thing.
///
/// So: interleave at *fold* granularity, alternating one large fold
/// against twenty small ones so both legs apply the same number of
/// events, and run long enough that each leg spans the drift instead
/// of sitting inside one excursion of it.
fn paired_ratio(small: &[GroundEvent], large: &[GroundEvent]) -> (f64, f64, f64) {
let per_round = large.len();
let rounds = AM7_EVENTS_PER_LEG.div_ceil(per_round);
let small_folds = per_round.div_ceil(small.len());
let (mut t_small, mut t_large) = (std::time::Duration::ZERO, std::time::Duration::ZERO);
let (mut n_small, mut n_large) = (0usize, 0usize);
for _ in 0..rounds {
for _ in 0..small_folds {
t_small += fold_once(small);
n_small += small.len();
}
t_large += fold_once(large);
n_large += large.len();
/// So: time a 5,000-event window at depth 0, and the same-sized window
/// at depth 100,000. Equal windows mean equal streaming cost, and the
/// only difference left is history depth — which is the claim.
fn paired_ratio(log: &[GroundEvent]) -> (f64, f64, f64) {
let reps = AM7_EVENTS_PER_LEG.div_ceil(AM7_WINDOW);
let early = state_at(log, 0);
let late = state_at(log, AM7_DEPTH);
let (mut t_early, mut t_late) = (std::time::Duration::ZERO, std::time::Duration::ZERO);
for _ in 0..reps {
// Interleaved, so this machine's 2.5x drift is common-mode
// and divides out (CB-EV-0013 section 1).
// THE SAME EVENTS on both legs. Timing log[0..W] against
// log[DEPTH..DEPTH+W] compared two different event mixes and
// read 0.573 on code whose state is provably bounded — a
// second confound, introduced while removing the first.
// Identical events mean the only difference left is how much
// history the state carries, which is the claim.
t_early += fold_window(&early, log, 0, AM7_WINDOW);
t_late += fold_window(&late, log, 0, AM7_WINDOW);
}
assert!(
t_small.as_secs_f64() > 0.0 && t_large.as_secs_f64() > 0.0,
t_early.as_secs_f64() > 0.0 && t_late.as_secs_f64() > 0.0,
"AM-7 measured zero elapsed time"
);
let tp_small = n_small as f64 / t_small.as_secs_f64();
let tp_large = n_large as f64 / t_large.as_secs_f64();
(tp_small, tp_large, tp_large / tp_small)
let n = (reps * AM7_WINDOW) as f64;
let tp_early = n / t_early.as_secs_f64();
let tp_late = n / t_late.as_secs_f64();
(tp_early, tp_late, tp_late / tp_early)
}
/// Build one growing log of at least `target` events, the same way
@ -2698,25 +2717,26 @@ mod replay_probe {
/// better — the noise multiplies rather than cancels. Run `make am7`.
#[test]
#[ignore = "throughput ratio — invalid under a parallel harness; run `make am7`"]
fn am7_scaling_holds_from_5k_to_100k_events() {
// Positive control on the shape of the measurement. A harness that
// measured the same size twice would report ~1.0 and look
// excellent; one that swapped the legs would report the reciprocal
// and look excellent for the opposite reason.
let (small, large) = (growing_log(AM7_SMALL), growing_log(AM7_LARGE));
fn am7_cost_per_event_does_not_grow_with_history() {
let log = growing_log(AM7_DEPTH + AM7_WINDOW);
// Positive control on the shape of the measurement. The windows
// must be the same size — that is the correction — and the late
// one must actually sit deep in the log. A harness that measured
// depth 0 twice would report ~1.0 and look excellent.
assert!(
large.len() >= 15 * small.len(),
"AM-7 legs are not far enough apart: {} vs {}",
small.len(),
large.len()
log.len() >= AM7_DEPTH + AM7_WINDOW,
"log is {} events, too short for a window at depth {AM7_DEPTH}",
log.len()
);
const { assert!(AM7_DEPTH >= 15 * AM7_WINDOW) };
let mut ratios = Vec::with_capacity(AM7_SAMPLES);
for _ in 0..AM7_SAMPLES {
let (tp_small, tp_large, ratio) = paired_ratio(&small, &large);
let (tp_early, tp_late, ratio) = paired_ratio(&log);
println!(
" AM-7 sample: {tp_small:.0} ev/s @{AM7_SMALL}\
{tp_large:.0} ev/s @{AM7_LARGE} = {ratio:.3}x"
" AM-7 sample: {tp_early:.0} ev/s at depth 0\
{tp_late:.0} ev/s at depth {AM7_DEPTH} = {ratio:.3}x"
);
ratios.push(ratio);
}
@ -2750,12 +2770,14 @@ mod replay_probe {
);
assert!(
median >= AM7_SCALING_FLOOR,
"AM-7 UNMET: throughput at {AM7_LARGE} events is {median:.3}x \
throughput at {AM7_SMALL} events (worst {worst:.3}x, best \
{best:.3}x), below the {AM7_SCALING_FLOOR} floor. Baseline: \
boardgame.io 0.45-0.66x, DNF at 100k. Do NOT lower the floor \
to pass GameKernel §5 AM-7 is a spec value and lowering it \
needs an ADR."
"AM-7 UNMET: folding a {AM7_WINDOW}-event window at history \
depth {AM7_DEPTH} runs at {median:.3}x the same window at \
depth 0 (worst {worst:.3}x, best {best:.3}x), below the \
{AM7_SCALING_FLOOR} floor. Cost per event is growing with \
history check whether something in the aggregate grows \
without bound. Baseline: boardgame.io 0.45-0.66x, DNF at \
100k. Do NOT lower the floor to pass GameKernel §5 AM-7 is \
a spec value and lowering it needs an ADR."
);
}
}