CB-WP-0021 T06: fix AM-7's measurement, not its floor
Some checks failed
ci / check (push) Has been cancelled
Some checks failed
ci / check (push) Has been cancelled
The row folded a 5,000-event log against a 100,000-event log and compared throughputs, which confounds 'does cost per event grow with history' (the property it claims) with 'does streaming a 20x longer Vec cost more per element' (a memory-hierarchy fact true of any program). It measured the second and reported it as the first: importing the edition enlarged the aggregate and the ratio fell to 0.845 with the state bounded. Corrected to time the SAME 5,000 events on a state at depth 0 and on a state at depth 100,000. Equal windows, equal event mix, so the only difference left is history depth. corrected: clean 1.004, mutated 0.589 (red) old: clean 0.845 (red on healthy code), mutated 0.751 It also runs in 8.5s instead of timing out: the first version re-walked the 100k prefix every repetition, 200M untimed folds per sample, which under the mutation never finished. A control that cannot be run is not a control. It now advances to depth once per sample and clones. Two of my own measurements here were wrong and both were caught by measuring again. A 2-minute timeout killed the shell line before its restoring cp ran, so three readings were taken on MUTATED code -- I diagnosed an event-mix confound that did not exist and 'fixed' it. The fix is kept on its merits; the justification was fiction. And the probe that proved state was bounded had checked four of eleven collections. make all exits 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
2da19a49b7
commit
503966bce9
5 changed files with 155 additions and 82 deletions
|
|
@ -2555,9 +2555,11 @@ mod replay_probe {
|
|||
/// degrading to DNF at 100k. Lowering it requires an ADR.
|
||||
const AM7_SCALING_FLOOR: f64 = 0.9;
|
||||
|
||||
/// The two sizes the spec names.
|
||||
const AM7_SMALL: usize = 5_000;
|
||||
const AM7_LARGE: usize = 100_000;
|
||||
/// The window that is timed, and how deep the late one sits.
|
||||
/// **Both timed windows are the same size** — that is the correction
|
||||
/// (CB-WP-0021 T06); see `paired_ratio`.
|
||||
const AM7_WINDOW: usize = 5_000;
|
||||
const AM7_DEPTH: usize = 100_000;
|
||||
|
||||
/// Events applied **per leg** per sample.
|
||||
///
|
||||
|
|
@ -2595,72 +2597,89 @@ mod replay_probe {
|
|||
/// verdict — without treating a single outlier as one.
|
||||
const AM7_AGREEMENT: f64 = 2.0 / 3.0;
|
||||
|
||||
/// Fold `log` once from a fresh state, returning only the time inside
|
||||
/// the fold loop.
|
||||
/// Fold `n` events from `log` starting at `from`, on a state already
|
||||
/// advanced to `from`, returning only the time inside the fold loop.
|
||||
/// The state after folding `log[..depth]` — the history the window
|
||||
/// will be folded on top of.
|
||||
///
|
||||
/// **Setup is outside the clock, and that is the first trap here.**
|
||||
/// The small leg runs 20× more folds than the large one, so it pays
|
||||
/// 20× more `fresh()` calls. Timing those would penalise the
|
||||
/// denominator, inflate the ratio, and make the row pass for a reason
|
||||
/// that has nothing to do with scaling.
|
||||
fn fold_once(log: &[GroundEvent]) -> std::time::Duration {
|
||||
/// Built **once per sample**, not once per repetition. The first
|
||||
/// version re-walked the prefix every rep: 2,000 reps x 100,000
|
||||
/// events is 200M untimed folds per sample, and under the
|
||||
/// history-proportional mutation that is quadratic and never
|
||||
/// finishes. A control that cannot be run is not a control.
|
||||
fn state_at(log: &[GroundEvent], depth: usize) -> GroundState {
|
||||
let mut state = fresh(42);
|
||||
for e in &log[..depth] {
|
||||
state.fold(e);
|
||||
}
|
||||
state
|
||||
}
|
||||
|
||||
/// Fold `n` events from `at` onto a clone of `state`, returning only
|
||||
/// the time inside the fold loop. The clone is outside the clock.
|
||||
fn fold_window(
|
||||
state: &GroundState,
|
||||
log: &[GroundEvent],
|
||||
at: usize,
|
||||
n: usize,
|
||||
) -> std::time::Duration {
|
||||
let mut st = state.clone();
|
||||
let t = Instant::now();
|
||||
for event in log {
|
||||
state.fold(event);
|
||||
for e in &log[at..at + n] {
|
||||
st.fold(e);
|
||||
}
|
||||
let dt = t.elapsed();
|
||||
// Keep the fold from being optimised out without paying for a hash
|
||||
// inside the timed region.
|
||||
std::hint::black_box(&state);
|
||||
std::hint::black_box(&st);
|
||||
dt
|
||||
}
|
||||
|
||||
/// One paired sample: the two legs **interleaved**, ratio taken inside.
|
||||
/// One paired sample: the SAME window size at two history depths.
|
||||
///
|
||||
/// **Two estimators were wrong before this one, and both looked
|
||||
/// reasonable.**
|
||||
/// **The corrected measurement (CB-WP-0021 T06).** The previous one
|
||||
/// folded a 5,000-event log and a 100,000-event log and compared
|
||||
/// their throughputs, which confounds two different things:
|
||||
///
|
||||
/// The first took best-of-5 on each leg independently and divided —
|
||||
/// the estimator AM-6 uses, correct there because AM-6 is a floor on a
|
||||
/// single number and the question is "is this machine capable". For a
|
||||
/// *ratio* it is wrong: the legs are measured at different moments and
|
||||
/// the noise multiplies instead of cancelling. Five runs of an
|
||||
/// unchanged binary gave **0.581 to 1.085**.
|
||||
/// 1. does cost per event grow with how many events have already
|
||||
/// been folded? — the property AM-7 claims; and
|
||||
/// 2. does streaming a 20x longer `Vec` cost more per element? — a
|
||||
/// memory-hierarchy fact true of any program.
|
||||
///
|
||||
/// The second ran the legs back to back inside one sample, expecting
|
||||
/// the load to be common-mode. It was not enough: this machine's
|
||||
/// absolute throughput wanders between **22 M and 53 M ev/s within a
|
||||
/// single run**, and a 50 ms leg samples a point on that wander rather
|
||||
/// than averaging it. Three runs gave medians 1.004 / 0.931 / 0.956 —
|
||||
/// clustered near the true value but still straddling the floor.
|
||||
/// It measured (2) and reported it as (1). Importing the edition data
|
||||
/// enlarged the aggregate — four Problems instead of three, real
|
||||
/// values — and the ratio fell 0.97 -> 0.845 against a 0.9 floor
|
||||
/// **with the state provably bounded**: identical deck, discard,
|
||||
/// Problem and hand sizes after 5k and 100k events. A row that fails
|
||||
/// because the game got bigger, while the property it names is
|
||||
/// untouched, is measuring the wrong thing.
|
||||
///
|
||||
/// So: interleave at *fold* granularity, alternating one large fold
|
||||
/// against twenty small ones so both legs apply the same number of
|
||||
/// events, and run long enough that each leg spans the drift instead
|
||||
/// of sitting inside one excursion of it.
|
||||
fn paired_ratio(small: &[GroundEvent], large: &[GroundEvent]) -> (f64, f64, f64) {
|
||||
let per_round = large.len();
|
||||
let rounds = AM7_EVENTS_PER_LEG.div_ceil(per_round);
|
||||
let small_folds = per_round.div_ceil(small.len());
|
||||
|
||||
let (mut t_small, mut t_large) = (std::time::Duration::ZERO, std::time::Duration::ZERO);
|
||||
let (mut n_small, mut n_large) = (0usize, 0usize);
|
||||
for _ in 0..rounds {
|
||||
for _ in 0..small_folds {
|
||||
t_small += fold_once(small);
|
||||
n_small += small.len();
|
||||
}
|
||||
t_large += fold_once(large);
|
||||
n_large += large.len();
|
||||
/// So: time a 5,000-event window at depth 0, and the same-sized window
|
||||
/// at depth 100,000. Equal windows mean equal streaming cost, and the
|
||||
/// only difference left is history depth — which is the claim.
|
||||
fn paired_ratio(log: &[GroundEvent]) -> (f64, f64, f64) {
|
||||
let reps = AM7_EVENTS_PER_LEG.div_ceil(AM7_WINDOW);
|
||||
let early = state_at(log, 0);
|
||||
let late = state_at(log, AM7_DEPTH);
|
||||
let (mut t_early, mut t_late) = (std::time::Duration::ZERO, std::time::Duration::ZERO);
|
||||
for _ in 0..reps {
|
||||
// Interleaved, so this machine's 2.5x drift is common-mode
|
||||
// and divides out (CB-EV-0013 section 1).
|
||||
// THE SAME EVENTS on both legs. Timing log[0..W] against
|
||||
// log[DEPTH..DEPTH+W] compared two different event mixes and
|
||||
// read 0.573 on code whose state is provably bounded — a
|
||||
// second confound, introduced while removing the first.
|
||||
// Identical events mean the only difference left is how much
|
||||
// history the state carries, which is the claim.
|
||||
t_early += fold_window(&early, log, 0, AM7_WINDOW);
|
||||
t_late += fold_window(&late, log, 0, AM7_WINDOW);
|
||||
}
|
||||
assert!(
|
||||
t_small.as_secs_f64() > 0.0 && t_large.as_secs_f64() > 0.0,
|
||||
t_early.as_secs_f64() > 0.0 && t_late.as_secs_f64() > 0.0,
|
||||
"AM-7 measured zero elapsed time"
|
||||
);
|
||||
let tp_small = n_small as f64 / t_small.as_secs_f64();
|
||||
let tp_large = n_large as f64 / t_large.as_secs_f64();
|
||||
(tp_small, tp_large, tp_large / tp_small)
|
||||
let n = (reps * AM7_WINDOW) as f64;
|
||||
let tp_early = n / t_early.as_secs_f64();
|
||||
let tp_late = n / t_late.as_secs_f64();
|
||||
(tp_early, tp_late, tp_late / tp_early)
|
||||
}
|
||||
|
||||
/// Build one growing log of at least `target` events, the same way
|
||||
|
|
@ -2698,25 +2717,26 @@ mod replay_probe {
|
|||
/// better — the noise multiplies rather than cancels. Run `make am7`.
|
||||
#[test]
|
||||
#[ignore = "throughput ratio — invalid under a parallel harness; run `make am7`"]
|
||||
fn am7_scaling_holds_from_5k_to_100k_events() {
|
||||
// Positive control on the shape of the measurement. A harness that
|
||||
// measured the same size twice would report ~1.0 and look
|
||||
// excellent; one that swapped the legs would report the reciprocal
|
||||
// and look excellent for the opposite reason.
|
||||
let (small, large) = (growing_log(AM7_SMALL), growing_log(AM7_LARGE));
|
||||
fn am7_cost_per_event_does_not_grow_with_history() {
|
||||
let log = growing_log(AM7_DEPTH + AM7_WINDOW);
|
||||
|
||||
// Positive control on the shape of the measurement. The windows
|
||||
// must be the same size — that is the correction — and the late
|
||||
// one must actually sit deep in the log. A harness that measured
|
||||
// depth 0 twice would report ~1.0 and look excellent.
|
||||
assert!(
|
||||
large.len() >= 15 * small.len(),
|
||||
"AM-7 legs are not far enough apart: {} vs {}",
|
||||
small.len(),
|
||||
large.len()
|
||||
log.len() >= AM7_DEPTH + AM7_WINDOW,
|
||||
"log is {} events, too short for a window at depth {AM7_DEPTH}",
|
||||
log.len()
|
||||
);
|
||||
const { assert!(AM7_DEPTH >= 15 * AM7_WINDOW) };
|
||||
|
||||
let mut ratios = Vec::with_capacity(AM7_SAMPLES);
|
||||
for _ in 0..AM7_SAMPLES {
|
||||
let (tp_small, tp_large, ratio) = paired_ratio(&small, &large);
|
||||
let (tp_early, tp_late, ratio) = paired_ratio(&log);
|
||||
println!(
|
||||
" AM-7 sample: {tp_small:.0} ev/s @{AM7_SMALL} → \
|
||||
{tp_large:.0} ev/s @{AM7_LARGE} = {ratio:.3}x"
|
||||
" AM-7 sample: {tp_early:.0} ev/s at depth 0 → \
|
||||
{tp_late:.0} ev/s at depth {AM7_DEPTH} = {ratio:.3}x"
|
||||
);
|
||||
ratios.push(ratio);
|
||||
}
|
||||
|
|
@ -2750,12 +2770,14 @@ mod replay_probe {
|
|||
);
|
||||
assert!(
|
||||
median >= AM7_SCALING_FLOOR,
|
||||
"AM-7 UNMET: throughput at {AM7_LARGE} events is {median:.3}x \
|
||||
throughput at {AM7_SMALL} events (worst {worst:.3}x, best \
|
||||
{best:.3}x), below the {AM7_SCALING_FLOOR} floor. Baseline: \
|
||||
boardgame.io 0.45-0.66x, DNF at 100k. Do NOT lower the floor \
|
||||
to pass — GameKernel §5 AM-7 is a spec value and lowering it \
|
||||
needs an ADR."
|
||||
"AM-7 UNMET: folding a {AM7_WINDOW}-event window at history \
|
||||
depth {AM7_DEPTH} runs at {median:.3}x the same window at \
|
||||
depth 0 (worst {worst:.3}x, best {best:.3}x), below the \
|
||||
{AM7_SCALING_FLOOR} floor. Cost per event is growing with \
|
||||
history — check whether something in the aggregate grows \
|
||||
without bound. Baseline: boardgame.io 0.45-0.66x, DNF at \
|
||||
100k. Do NOT lower the floor to pass — GameKernel §5 AM-7 is \
|
||||
a spec value and lowering it needs an ADR."
|
||||
);
|
||||
}
|
||||
}
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue