clay-borg/games/ground/examples/scenario-panel.rs

182 lines
7 KiB
Rust
Raw Normal View History

CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
//! Every board, in every mode, at every seat band (CB-WP-0047).
//!
//! The engine dealt `SCN_01` and only `SCN_01` for the whole life of this
//! repo: `deal` took a scenario id, `setup` passed a literal, and **15 of
//! the 20 Problem cards had never been dealt by anything**. The three
//! scoring modes were reachable from the driver but no measurement had
//! ever compared them.
//!
//! This plays all of it. It is a **characterisation panel**, not a
//! hypothesis test: nothing here predicts an outcome, and the numbers
//! exist so that a later change to a deck or a mode is visible as a
//! change rather than discovered by a player.
//!
//! **SCN_01 and SCN_02 are the same board.** Identical suits and values
//! at every priority; they differ only in prose. They are both played
//! anyway — a panel that dropped one would hide the day they diverge —
//! and the column is expected to match, which is itself a control.
use cb_game_runtime::{ScenarioGame, Setup};
CB-WP-0049 T02/T03: a seat that plays its objective, and F27 splits in two objective() reads GroundState::score (now public) rather than restating what winning is; a copy in the bot would disagree with the kernel the first time ground-game rules on F28. Working out WHERE the modes can differ was most of the task and it bounds the result: SOLVE always claims for the actor, so own-score and group-score want the same SOLVE nearly everywhere. That is a fact about GROUND's action set, not a shortcoming of the bot. Two real divergences, both readable off the table: SUPPORT regulates someone else (worth less against a rival, worth MORE under coalitions where a Bond merges them into my side), and SOLVE's value is the card's value, which greedy ignores entirely. THE RESULT — F27 splits in two: group success UNCHANGED in 34 of 36 cells who wins MOVES: BONDED COALITIONS at 4p goes 2.04 -> 2.98, 2.12 -> 3.29, 2.05 -> 3.01 winning seats per game So "the competitive modes are scoring lenses over cooperative play" was too strong and is withdrawn. The sharper claim: GROUND's scoring modes change WHO WINS, not WHETHER THE GROUP SUCCEEDS. And the effect is seat-band dependent -- 2p none, 4p largest, 6p none under coalitions; two relation slots capping network growth is a candidate explanation and is untested. The panel now prints BOTH policies side by side. That was a correction mid-task: the first version printed only the new one and I compared it against a figure remembered from CB-WP-0047 -- a comparison against a board nobody re-ran. Control that makes the numbers mean anything: under SHARED GROUND the two policies agree at all but <=2 decision points across 12 boards, so a moving column is mode-awareness and not simply a different bot. Also: two T01 tests keyed on `status: proposed`, which ground-game renamed to `ready-for-implement` mid-session. They now find the module by asking resolve() -- the structural property is ours and does not move when another repo edits its vocabulary. Also: `make vendor` replaces three hand re-vendors with a tool that regenerates digests by walking editions/, and reports one-sided files rather than resolving them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 23:31:11 +02:00
use games_ground::bot::{play, GreedyPolicy, ObjectivePolicy, Policy};
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
use games_ground::{GroundState, ScoringMode};
/// Games per cell. Named once so the banner and the assertion cannot
/// disagree — a previous panel printed "200 games per cell" over
/// 196-game columns (CB-REV-0002 #1).
const GAMES: u32 = 100;
struct Cell {
games: u32,
/// Games that reached an outcome over the full five rounds. A game
/// that RAN is not a game that was PLAYED (CB-REV-0003 #2).
played: u32,
group_success: u32,
/// Seats named a winner. Under SHARED GROUND that is everyone or
/// nobody; under the other two it is the point of the mode.
winners: u32,
/// Claimed value summed over games, so the threshold has something
/// to be compared against.
total: u32,
threshold: u32,
setup_fails: u32,
}
CB-WP-0049 T02/T03: a seat that plays its objective, and F27 splits in two objective() reads GroundState::score (now public) rather than restating what winning is; a copy in the bot would disagree with the kernel the first time ground-game rules on F28. Working out WHERE the modes can differ was most of the task and it bounds the result: SOLVE always claims for the actor, so own-score and group-score want the same SOLVE nearly everywhere. That is a fact about GROUND's action set, not a shortcoming of the bot. Two real divergences, both readable off the table: SUPPORT regulates someone else (worth less against a rival, worth MORE under coalitions where a Bond merges them into my side), and SOLVE's value is the card's value, which greedy ignores entirely. THE RESULT — F27 splits in two: group success UNCHANGED in 34 of 36 cells who wins MOVES: BONDED COALITIONS at 4p goes 2.04 -> 2.98, 2.12 -> 3.29, 2.05 -> 3.01 winning seats per game So "the competitive modes are scoring lenses over cooperative play" was too strong and is withdrawn. The sharper claim: GROUND's scoring modes change WHO WINS, not WHETHER THE GROUP SUCCEEDS. And the effect is seat-band dependent -- 2p none, 4p largest, 6p none under coalitions; two relation slots capping network growth is a candidate explanation and is untested. The panel now prints BOTH policies side by side. That was a correction mid-task: the first version printed only the new one and I compared it against a figure remembered from CB-WP-0047 -- a comparison against a board nobody re-ran. Control that makes the numbers mean anything: under SHARED GROUND the two policies agree at all but <=2 decision points across 12 boards, so a moving column is mode-awareness and not simply a different bot. Also: two T01 tests keyed on `status: proposed`, which ground-game renamed to `ready-for-implement` mid-session. They now find the module by asking resolve() -- the structural property is ours and does not move when another repo edits its vocabulary. Also: `make vendor` replaces three hand re-vendors with a tool that regenerates digests by walking editions/, and reports one-sided files rather than resolving them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 23:31:11 +02:00
/// Which bot fills every seat. CB-WP-0049 T03: F27 says the three modes
/// produce identical play because greedy never reads `state.mode`; the
/// only way to know whether that is a fact about GROUND or a fact about
/// our bot is to run the same panel with a seat that does read it.
#[derive(Clone, Copy, PartialEq)]
enum Seat {
Greedy,
Objective,
}
fn sweep(scenario: &str, mode: ScoringMode, players: u8, who: Seat) -> Cell {
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
let mut c = Cell {
games: 0,
played: 0,
group_success: 0,
winners: 0,
total: 0,
threshold: 0,
setup_fails: 0,
};
let preset = if scenario == "SCN_01" {
format!("standard-{players}p")
} else {
format!("scn-{}-{players}p", scenario.rsplit('_').next().unwrap())
};
for seed in 0..GAMES as u64 {
let Ok(mut st) = GroundState::setup(
&Setup {
players,
preset: preset.clone(),
patch: Default::default(),
},
seed,
) else {
// Counted, not skipped (CB-REV-0001 #11).
c.setup_fails += 1;
continue;
};
st.mode = mode;
// The board actually dealt, so the panel cannot report a
// scenario it did not play.
assert_eq!(st.scenario, scenario, "{preset} dealt {}", st.scenario);
let mut ps: Vec<Box<dyn Policy>> = (0..players)
CB-WP-0049 T02/T03: a seat that plays its objective, and F27 splits in two objective() reads GroundState::score (now public) rather than restating what winning is; a copy in the bot would disagree with the kernel the first time ground-game rules on F28. Working out WHERE the modes can differ was most of the task and it bounds the result: SOLVE always claims for the actor, so own-score and group-score want the same SOLVE nearly everywhere. That is a fact about GROUND's action set, not a shortcoming of the bot. Two real divergences, both readable off the table: SUPPORT regulates someone else (worth less against a rival, worth MORE under coalitions where a Bond merges them into my side), and SOLVE's value is the card's value, which greedy ignores entirely. THE RESULT — F27 splits in two: group success UNCHANGED in 34 of 36 cells who wins MOVES: BONDED COALITIONS at 4p goes 2.04 -> 2.98, 2.12 -> 3.29, 2.05 -> 3.01 winning seats per game So "the competitive modes are scoring lenses over cooperative play" was too strong and is withdrawn. The sharper claim: GROUND's scoring modes change WHO WINS, not WHETHER THE GROUP SUCCEEDS. And the effect is seat-band dependent -- 2p none, 4p largest, 6p none under coalitions; two relation slots capping network growth is a candidate explanation and is untested. The panel now prints BOTH policies side by side. That was a correction mid-task: the first version printed only the new one and I compared it against a figure remembered from CB-WP-0047 -- a comparison against a board nobody re-ran. Control that makes the numbers mean anything: under SHARED GROUND the two policies agree at all but <=2 decision points across 12 boards, so a moving column is mode-awareness and not simply a different bot. Also: two T01 tests keyed on `status: proposed`, which ground-game renamed to `ready-for-implement` mid-session. They now find the module by asking resolve() -- the structural property is ours and does not move when another repo edits its vocabulary. Also: `make vendor` replaces three hand re-vendors with a tool that regenerates digests by walking editions/, and reports one-sided files rather than resolving them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 23:31:11 +02:00
.map(|_| match who {
Seat::Greedy => Box::new(GreedyPolicy) as Box<dyn Policy>,
Seat::Objective => Box::new(ObjectivePolicy) as Box<dyn Policy>,
})
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
.collect();
let g = match play(st, &mut ps) {
Ok(g) => g,
Err(e) => {
eprintln!(" !! {scenario} {mode:?} {players}p seed {seed}: {e:?}");
continue;
}
};
c.games += 1;
if g.state.outcome.is_some() && g.rounds == 5 {
c.played += 1;
}
if let Some(o) = &g.state.outcome {
if o.group_success {
c.group_success += 1;
}
c.winners += o.winners.len() as u32;
c.total += o.total;
c.threshold = o.threshold;
}
}
assert_eq!(
c.games, GAMES,
"{scenario} {mode:?} {players}p: only {} of {GAMES} games ran \
({} setups refused) every number in the cell is over a sample \
nobody chose",
c.games, c.setup_fails
);
assert_eq!(
c.played, GAMES,
"{scenario} {mode:?} {players}p: {} of {} games reached an outcome \
over five rounds",
c.played, c.games
);
c
}
fn main() {
println!("CB-WP-0047 — the four Scenarios, the three Modes\n");
println!("Greedy throughout; {GAMES} games per cell; seeds 0..{GAMES}.");
println!("`won` is group success; `win/g` is winning SEATS per game,");
println!("which is what separates the modes — SHARED GROUND names all");
println!("or none, the other two name a subset.\n");
let scenarios = games_ground::edition::scenarios().expect("Scenarios.csv");
// Every mode the kernel has. Adding a fourth without extending the
// panel is the omission this list exists to make loud.
let modes = [
("SHARED GROUND ", ScoringMode::SharedGround),
("COMMON PROBLEM ", ScoringMode::CommonProblem),
("BONDED COALITIONS ", ScoringMode::BondedCoalitions),
];
for players in [2u8, 4, 6] {
println!("{players} players");
println!(
CB-WP-0049 T02/T03: a seat that plays its objective, and F27 splits in two objective() reads GroundState::score (now public) rather than restating what winning is; a copy in the bot would disagree with the kernel the first time ground-game rules on F28. Working out WHERE the modes can differ was most of the task and it bounds the result: SOLVE always claims for the actor, so own-score and group-score want the same SOLVE nearly everywhere. That is a fact about GROUND's action set, not a shortcoming of the bot. Two real divergences, both readable off the table: SUPPORT regulates someone else (worth less against a rival, worth MORE under coalitions where a Bond merges them into my side), and SOLVE's value is the card's value, which greedy ignores entirely. THE RESULT — F27 splits in two: group success UNCHANGED in 34 of 36 cells who wins MOVES: BONDED COALITIONS at 4p goes 2.04 -> 2.98, 2.12 -> 3.29, 2.05 -> 3.01 winning seats per game So "the competitive modes are scoring lenses over cooperative play" was too strong and is withdrawn. The sharper claim: GROUND's scoring modes change WHO WINS, not WHETHER THE GROUP SUCCEEDS. And the effect is seat-band dependent -- 2p none, 4p largest, 6p none under coalitions; two relation slots capping network growth is a candidate explanation and is untested. The panel now prints BOTH policies side by side. That was a correction mid-task: the first version printed only the new one and I compared it against a figure remembered from CB-WP-0047 -- a comparison against a board nobody re-ran. Control that makes the numbers mean anything: under SHARED GROUND the two policies agree at all but <=2 decision points across 12 boards, so a moving column is mode-awareness and not simply a different bot. Also: two T01 tests keyed on `status: proposed`, which ground-game renamed to `ready-for-implement` mid-session. They now find the module by asking resolve() -- the structural property is ours and does not move when another repo edits its vocabulary. Also: `make vendor` replaces three hand re-vendors with a tool that regenerates digests by walking editions/, and reports one-sided files rather than resolving them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 23:31:11 +02:00
" {:<32} {:<18} {:>5} {:>5} {:>7} {:>7} {:>6} {:>9}",
"scenario", "mode", "grdy", "obj", "win/g g", "win/g o", "pts", "threshold"
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
);
for s in &scenarios {
for (label, mode) in &modes {
CB-WP-0049 T02/T03: a seat that plays its objective, and F27 splits in two objective() reads GroundState::score (now public) rather than restating what winning is; a copy in the bot would disagree with the kernel the first time ground-game rules on F28. Working out WHERE the modes can differ was most of the task and it bounds the result: SOLVE always claims for the actor, so own-score and group-score want the same SOLVE nearly everywhere. That is a fact about GROUND's action set, not a shortcoming of the bot. Two real divergences, both readable off the table: SUPPORT regulates someone else (worth less against a rival, worth MORE under coalitions where a Bond merges them into my side), and SOLVE's value is the card's value, which greedy ignores entirely. THE RESULT — F27 splits in two: group success UNCHANGED in 34 of 36 cells who wins MOVES: BONDED COALITIONS at 4p goes 2.04 -> 2.98, 2.12 -> 3.29, 2.05 -> 3.01 winning seats per game So "the competitive modes are scoring lenses over cooperative play" was too strong and is withdrawn. The sharper claim: GROUND's scoring modes change WHO WINS, not WHETHER THE GROUP SUCCEEDS. And the effect is seat-band dependent -- 2p none, 4p largest, 6p none under coalitions; two relation slots capping network growth is a candidate explanation and is untested. The panel now prints BOTH policies side by side. That was a correction mid-task: the first version printed only the new one and I compared it against a figure remembered from CB-WP-0047 -- a comparison against a board nobody re-ran. Control that makes the numbers mean anything: under SHARED GROUND the two policies agree at all but <=2 decision points across 12 boards, so a moving column is mode-awareness and not simply a different bot. Also: two T01 tests keyed on `status: proposed`, which ground-game renamed to `ready-for-implement` mid-session. They now find the module by asking resolve() -- the structural property is ours and does not move when another repo edits its vocabulary. Also: `make vendor` replaces three hand re-vendors with a tool that regenerates digests by walking editions/, and reports one-sided files rather than resolving them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 23:31:11 +02:00
let g = sweep(&s.id, *mode, players, Seat::Greedy);
let c = sweep(&s.id, *mode, players, Seat::Objective);
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
println!(
CB-WP-0049 T02/T03: a seat that plays its objective, and F27 splits in two objective() reads GroundState::score (now public) rather than restating what winning is; a copy in the bot would disagree with the kernel the first time ground-game rules on F28. Working out WHERE the modes can differ was most of the task and it bounds the result: SOLVE always claims for the actor, so own-score and group-score want the same SOLVE nearly everywhere. That is a fact about GROUND's action set, not a shortcoming of the bot. Two real divergences, both readable off the table: SUPPORT regulates someone else (worth less against a rival, worth MORE under coalitions where a Bond merges them into my side), and SOLVE's value is the card's value, which greedy ignores entirely. THE RESULT — F27 splits in two: group success UNCHANGED in 34 of 36 cells who wins MOVES: BONDED COALITIONS at 4p goes 2.04 -> 2.98, 2.12 -> 3.29, 2.05 -> 3.01 winning seats per game So "the competitive modes are scoring lenses over cooperative play" was too strong and is withdrawn. The sharper claim: GROUND's scoring modes change WHO WINS, not WHETHER THE GROUP SUCCEEDS. And the effect is seat-band dependent -- 2p none, 4p largest, 6p none under coalitions; two relation slots capping network growth is a candidate explanation and is untested. The panel now prints BOTH policies side by side. That was a correction mid-task: the first version printed only the new one and I compared it against a figure remembered from CB-WP-0047 -- a comparison against a board nobody re-ran. Control that makes the numbers mean anything: under SHARED GROUND the two policies agree at all but <=2 decision points across 12 boards, so a moving column is mode-awareness and not simply a different bot. Also: two T01 tests keyed on `status: proposed`, which ground-game renamed to `ready-for-implement` mid-session. They now find the module by asking resolve() -- the structural property is ours and does not move when another repo edits its vocabulary. Also: `make vendor` replaces three hand re-vendors with a tool that regenerates digests by walking editions/, and reports one-sided files rather than resolving them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 23:31:11 +02:00
" {:<32} {label} {:>5} {:>5} {:>7.2} {:>7.2} {:>6.1} {:>9}",
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
format!("{} {}", s.id, s.title),
CB-WP-0049 T02/T03: a seat that plays its objective, and F27 splits in two objective() reads GroundState::score (now public) rather than restating what winning is; a copy in the bot would disagree with the kernel the first time ground-game rules on F28. Working out WHERE the modes can differ was most of the task and it bounds the result: SOLVE always claims for the actor, so own-score and group-score want the same SOLVE nearly everywhere. That is a fact about GROUND's action set, not a shortcoming of the bot. Two real divergences, both readable off the table: SUPPORT regulates someone else (worth less against a rival, worth MORE under coalitions where a Bond merges them into my side), and SOLVE's value is the card's value, which greedy ignores entirely. THE RESULT — F27 splits in two: group success UNCHANGED in 34 of 36 cells who wins MOVES: BONDED COALITIONS at 4p goes 2.04 -> 2.98, 2.12 -> 3.29, 2.05 -> 3.01 winning seats per game So "the competitive modes are scoring lenses over cooperative play" was too strong and is withdrawn. The sharper claim: GROUND's scoring modes change WHO WINS, not WHETHER THE GROUP SUCCEEDS. And the effect is seat-band dependent -- 2p none, 4p largest, 6p none under coalitions; two relation slots capping network growth is a candidate explanation and is untested. The panel now prints BOTH policies side by side. That was a correction mid-task: the first version printed only the new one and I compared it against a figure remembered from CB-WP-0047 -- a comparison against a board nobody re-ran. Control that makes the numbers mean anything: under SHARED GROUND the two policies agree at all but <=2 decision points across 12 boards, so a moving column is mode-awareness and not simply a different bot. Also: two T01 tests keyed on `status: proposed`, which ground-game renamed to `ready-for-implement` mid-session. They now find the module by asking resolve() -- the structural property is ours and does not move when another repo edits its vocabulary. Also: `make vendor` replaces three hand re-vendors with a tool that regenerates digests by walking editions/, and reports one-sided files rather than resolving them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 23:31:11 +02:00
g.group_success,
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
c.group_success,
CB-WP-0049 T02/T03: a seat that plays its objective, and F27 splits in two objective() reads GroundState::score (now public) rather than restating what winning is; a copy in the bot would disagree with the kernel the first time ground-game rules on F28. Working out WHERE the modes can differ was most of the task and it bounds the result: SOLVE always claims for the actor, so own-score and group-score want the same SOLVE nearly everywhere. That is a fact about GROUND's action set, not a shortcoming of the bot. Two real divergences, both readable off the table: SUPPORT regulates someone else (worth less against a rival, worth MORE under coalitions where a Bond merges them into my side), and SOLVE's value is the card's value, which greedy ignores entirely. THE RESULT — F27 splits in two: group success UNCHANGED in 34 of 36 cells who wins MOVES: BONDED COALITIONS at 4p goes 2.04 -> 2.98, 2.12 -> 3.29, 2.05 -> 3.01 winning seats per game So "the competitive modes are scoring lenses over cooperative play" was too strong and is withdrawn. The sharper claim: GROUND's scoring modes change WHO WINS, not WHETHER THE GROUP SUCCEEDS. And the effect is seat-band dependent -- 2p none, 4p largest, 6p none under coalitions; two relation slots capping network growth is a candidate explanation and is untested. The panel now prints BOTH policies side by side. That was a correction mid-task: the first version printed only the new one and I compared it against a figure remembered from CB-WP-0047 -- a comparison against a board nobody re-ran. Control that makes the numbers mean anything: under SHARED GROUND the two policies agree at all but <=2 decision points across 12 boards, so a moving column is mode-awareness and not simply a different bot. Also: two T01 tests keyed on `status: proposed`, which ground-game renamed to `ready-for-implement` mid-session. They now find the module by asking resolve() -- the structural property is ours and does not move when another repo edits its vocabulary. Also: `make vendor` replaces three hand re-vendors with a tool that regenerates digests by walking editions/, and reports one-sided files rather than resolving them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 23:31:11 +02:00
// **Both, side by side.** Reading one policy's number
// against a figure remembered from an earlier pass is
// how a comparison gets made against a board nobody
// re-ran.
f64::from(g.winners) / f64::from(g.games),
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
f64::from(c.winners) / f64::from(c.games),
f64::from(c.total) / f64::from(c.games),
c.threshold,
);
}
}
println!();
}
CB-WP-0049 T02/T03: a seat that plays its objective, and F27 splits in two objective() reads GroundState::score (now public) rather than restating what winning is; a copy in the bot would disagree with the kernel the first time ground-game rules on F28. Working out WHERE the modes can differ was most of the task and it bounds the result: SOLVE always claims for the actor, so own-score and group-score want the same SOLVE nearly everywhere. That is a fact about GROUND's action set, not a shortcoming of the bot. Two real divergences, both readable off the table: SUPPORT regulates someone else (worth less against a rival, worth MORE under coalitions where a Bond merges them into my side), and SOLVE's value is the card's value, which greedy ignores entirely. THE RESULT — F27 splits in two: group success UNCHANGED in 34 of 36 cells who wins MOVES: BONDED COALITIONS at 4p goes 2.04 -> 2.98, 2.12 -> 3.29, 2.05 -> 3.01 winning seats per game So "the competitive modes are scoring lenses over cooperative play" was too strong and is withdrawn. The sharper claim: GROUND's scoring modes change WHO WINS, not WHETHER THE GROUP SUCCEEDS. And the effect is seat-band dependent -- 2p none, 4p largest, 6p none under coalitions; two relation slots capping network growth is a candidate explanation and is untested. The panel now prints BOTH policies side by side. That was a correction mid-task: the first version printed only the new one and I compared it against a figure remembered from CB-WP-0047 -- a comparison against a board nobody re-ran. Control that makes the numbers mean anything: under SHARED GROUND the two policies agree at all but <=2 decision points across 12 boards, so a moving column is mode-awareness and not simply a different bot. Also: two T01 tests keyed on `status: proposed`, which ground-game renamed to `ready-for-implement` mid-session. They now find the module by asking resolve() -- the structural property is ours and does not move when another repo edits its vocabulary. Also: `make vendor` replaces three hand re-vendors with a tool that regenerates digests by walking editions/, and reports one-sided files rather than resolving them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 23:31:11 +02:00
println!("`grdy` is GreedyPolicy, which never reads state.mode; `obj`");
println!("is ObjectivePolicy, which plays its seat's own objective.");
println!("If the obj column varies ACROSS modes where grdy does not,");
println!("F27 was a fact about our bot. If it does not, it is a fact");
println!("about GROUND.\n");
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
println!("SCN_01 and SCN_02 are the same board — identical suits and");
println!("values at every priority. Their rows are expected to MATCH,");
println!("and a divergence means one of the two decks was edited.");
}