caught me repeating C1 specs/RetrospectiveAnalysis.md v1.0 plus games/ground/benches/search.rs, which exists because ADR-0013 D7 refused to let the spec quote either disputed figure. THE BENCHMARK'S OWN FIRST FIXTURE WAS DEFECTIVE, and it is the same defect class the review caught one layer up. Stopping at a fixed step 20 put 2p and 4p in states where seat 0 had NO legal commands, so it timed an empty Vec (~120 ns) and silently skipped validate_fold because there was nothing to validate. It now advances until the seat has a real branch and ASSERTS it. A clone benchmark was added too: a search must copy state per branch, and iter_batched excludes setup from timing, so without it the budget would again rest on an unmeasured span. Measured at real decision points: legal_commands 4.06-4.76 us, clone 378-639 ns, validate+fold 0.5-3.8 us. Per-child cost is NOT uniform -- some commands resolve cascades -- so budgets use the upper end (~5 us/child). That settles D3 with real numbers. Joint branching over the last two rounds is ~5x10^2 / 1.6x10^5 / 5.7x10^5 at 2/3/4 seats, so K=2 costs negligible / 0.8 s / 2.9 s and holds at two to four seats. It does NOT hold at five or six, where the tool must reduce K and say that it did rather than silently searching less. §4.1 is a normative prohibition, not a preference: a single policy's win rate MAY NOT be reported as a difficulty. The spec carries the measured reason -- greedy 100% against first-legal 0% on identical deals -- because this project already made that error and nearly exported it to a repo that is blocked waiting on the number. §2.3 makes the empty-result wording normative: "no winning line found in the last K rounds", never "unwinnable". A bounded search cannot establish unwinnability and that sentence is what a player who just lost reads. Also corrected: the T01 completion record still asserted all three withdrawn claims as fact. It now carries claimed / withdrawn / survives explicitly rather than being rewritten -- a retraction that does not propagate to every place the claim lives is how the earlier ones survived. make all: exit 0. loop-lint clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
128 lines
5.2 KiB
Rust
128 lines
5.2 KiB
Rust
//! CB-WP-0025 T04 — what one search node actually costs.
|
||
//!
|
||
//! **ADR-0013 D7 exists because two measurements disagreed by 5×.** The
|
||
//! survey published 112–161 µs/node from a timer that bracketed two
|
||
//! `setup`s and a whole greedy game (C1). The author's re-measurement said
|
||
//! 3.0–4.1 µs with `Instant::now()` around each call; the adversarial
|
||
//! reviewer's isolation said 15.6–20.4 µs. Both agreed the published
|
||
//! figure was wrong by 1–2 orders and neither established which
|
||
//! replacement was right.
|
||
//!
|
||
//! So the spec quotes **this** and nothing else. `criterion` handles the
|
||
//! things hand-rolled timing gets wrong here: per-call clock overhead
|
||
//! against a ~microsecond subject, warm-up, and run-to-run variance —
|
||
//! which is what let the survey's figure move 161 → 112 between two runs
|
||
//! of the same unmodified binary.
|
||
//!
|
||
//! Two subjects, because a search node is not one call:
|
||
//!
|
||
//! * `legal_commands` — enumerating a seat's options;
|
||
//! * `validate + fold` — taking one branch, which any search does per
|
||
//! child and which the survey never separated out.
|
||
|
||
use cb_game_runtime::{ScenarioGame, Setup};
|
||
use cb_kernel::{Actor, Aggregate, PlayerId};
|
||
use criterion::{criterion_group, criterion_main, BatchSize, Criterion};
|
||
use games_ground::bot::{legal_commands, play, GreedyPolicy, Policy};
|
||
use games_ground::GroundState;
|
||
use std::collections::BTreeMap;
|
||
|
||
fn setup(players: u8, seed: u64) -> GroundState {
|
||
GroundState::setup(
|
||
&Setup {
|
||
players,
|
||
preset: format!("standard-{players}p"),
|
||
patch: BTreeMap::new(),
|
||
},
|
||
seed,
|
||
)
|
||
.expect("preset")
|
||
}
|
||
|
||
/// A **mid-game state at a real decision point** for `seat`.
|
||
///
|
||
/// Not a fresh deal: at deal time most branches do not exist yet, and a
|
||
/// node cost taken there would flatter any search proposal.
|
||
///
|
||
/// **And not a fixed step count either.** The first version stopped at
|
||
/// step 20 for every seat count, which put 2p and 4p in a state where
|
||
/// seat 0 had *no* legal commands at all — so the benchmark reported
|
||
/// ~120 ns (the cost of returning an empty `Vec`) and silently skipped
|
||
/// `validate_fold` because there was nothing to validate. A fixture that
|
||
/// measures the empty case and calls it a node cost is the same defect
|
||
/// class this whole pass exists to correct, one layer down.
|
||
///
|
||
/// So: advance until the seat genuinely has a choice, and assert it.
|
||
fn midgame(players: u8, seed: u64, seat: PlayerId) -> GroundState {
|
||
let mut ps: Vec<Box<dyn Policy>> = (0..players)
|
||
.map(|_| Box::new(GreedyPolicy) as Box<dyn Policy>)
|
||
.collect();
|
||
let game = play(setup(players, seed), &mut ps).expect("a complete game");
|
||
let mut state = setup(players, seed);
|
||
let mut best: Option<GroundState> = None;
|
||
for (i, (actor, cmd)) in game.steps.iter().enumerate() {
|
||
// Past the opening, take the first state where the seat has a real
|
||
// branch. `> 1` rather than `> 0`: a forced move is not a node.
|
||
if i >= 8 && legal_commands(&state, seat).len() > 1 {
|
||
best = Some(state.clone());
|
||
break;
|
||
}
|
||
if let Ok(events) = state.validate(*actor, cmd) {
|
||
for e in &events {
|
||
state.fold(e);
|
||
}
|
||
}
|
||
}
|
||
let state = best.expect("a mid-game state where the seat has a choice");
|
||
assert!(
|
||
legal_commands(&state, seat).len() > 1,
|
||
"benchmark fixture has no branch to measure — it would time the empty case"
|
||
);
|
||
state
|
||
}
|
||
|
||
fn bench(c: &mut Criterion) {
|
||
for players in [2u8, 3, 4] {
|
||
let seat = PlayerId(0);
|
||
let state = midgame(players, 7, seat);
|
||
let width = legal_commands(&state, seat).len();
|
||
println!(" fixture {players}p: {width} legal commands at the measured node");
|
||
|
||
c.bench_function(&format!("legal_commands/{players}p"), |b| {
|
||
b.iter(|| std::hint::black_box(legal_commands(&state, seat)))
|
||
});
|
||
|
||
// A search must COPY the state per branch (or undo, which we do
|
||
// not have). `iter_batched` excludes setup from the timing, so
|
||
// without this the budget would rest on an unmeasured span —
|
||
// which is the exact mistake C1 caught in the survey.
|
||
c.bench_function(&format!("clone/{players}p"), |b| {
|
||
b.iter(|| std::hint::black_box(state.clone()))
|
||
});
|
||
|
||
// One branch taken: what a search pays per CHILD, on top of
|
||
// enumeration. The survey folded this into "us/node" without
|
||
// separating it, and a search's real cost is enumeration once plus
|
||
// this per child.
|
||
let legal = legal_commands(&state, seat);
|
||
if let Some(cmd) = legal.first() {
|
||
c.bench_function(&format!("validate_fold/{players}p"), |b| {
|
||
b.iter_batched(
|
||
|| state.clone(),
|
||
|mut s| {
|
||
if let Ok(events) = s.validate(Actor::Player(seat), cmd) {
|
||
for e in &events {
|
||
s.fold(e);
|
||
}
|
||
}
|
||
std::hint::black_box(s)
|
||
},
|
||
BatchSize::SmallInput,
|
||
)
|
||
});
|
||
}
|
||
}
|
||
}
|
||
|
||
criterion_group!(benches, bench);
|
||
criterion_main!(benches);
|