25 test games, no faults — and F17 gets the artifact that changes what it
Some checks failed
ci / check (push) Failing after 4s
Some checks failed
ci / check (push) Failing after 4s
says Five games per seat count, 2-6 players. NO ANOMALIES: every game reaches 5 rounds with an outcome, no stalls, no stress above the cap, no over-claimed Problems. But the series showed something a crash never would. DARVO NEVER FIRED IN 25 GAMES and stress never exceeded 2. Measured wider: GreedyPolicy plays ATTACK exactly ZERO times in 10,000 selections across 500 games. THAT NUMBER IS ABOUT OUR BOT, NOT THE GAME. bot.rs ranks `Action::Attack => 10`, below everything. Reporting "the game gives no incentive to attack" from a policy we programmed to rank attack last would have been CB-WP-0025's C4 error committed again -- a single policy's behaviour presented as the game's. So the artifact varies exactly one number: ATTACK's rank in an otherwise identical policy, 200 games per cell. rank 10 (below all): 132/165/190/200/200 wins, 0 attacks, 0 DARVO rank 75 (above SUPPORT): 132/165/190/200/200 wins, 315-923, 13-218 rank 95 (above SOLVE): 0/0/0/0/0 wins, 1400-5170, 400-1000 THE MIDDLE ROW IS THE FINDING. Identical win counts at every seat count, while attacking hundreds of times and arming DARVO repeatedly. Attacking is not punished -- it is INERT with respect to the goal. Group success is a function of SOLVE alone, and ATTACK costs anything only when it ranks above SOLVE and displaces it. The maintainer was right and the reason is sharper than his phrasing: there is no incentive because there is no PATH. ATTACK's effects (Stress, Rivalry, DARVO) feed nothing that decides group_success. Bounded honestly to SHARED GROUND. Blame costs PERSONAL score, so ATTACK may earn its place in GR-E03 and GR-E04 -- which have never been played to the end (F14), and that is where to ask next. And this is NOT a claim the game is broken: DARVO is the pattern the game is about not falling into, so a self-destructive ATTACK may be the design. The question for ground-game is whether the namesake mechanic being unreachable in competent co-op play is intended. F17 promoted from note to raised, with games/ground/examples/attack-value.rs as its reproduction. Register: 18 findings, 8 with a resolving reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
d5b3d27b43
commit
f4eeddd726
2 changed files with 176 additions and 11 deletions
143
games/ground/examples/attack-value.rs
Normal file
143
games/ground/examples/attack-value.rs
Normal file
|
|
@ -0,0 +1,143 @@
|
|||
//! F17's reproduction: what is ATTACK worth? (CB-WP-0029 follow-up.)
|
||||
//!
|
||||
//! The maintainer reported *"there is no incentive to play attacks as long
|
||||
//! as I have positive cards"*. F17 sat as a note because nothing
|
||||
//! demonstrated it. This is the artifact.
|
||||
//!
|
||||
//! GreedyPolicy ranks Attack at 10, below everything (bot.rs). So its zero
|
||||
//! attacks measure OUR HEURISTIC, not the game. The test is a policy
|
||||
//! identical in every other respect with Attack promoted: if it wins as
|
||||
//! often, attacking is neutral; much less, avoiding it is correct play;
|
||||
//! more, and greedy is simply wrong.
|
||||
use cb_game_runtime::{ScenarioGame, Setup};
|
||||
use cb_kernel::PlayerId;
|
||||
use games_ground::bot::{play, Choice, GreedyPolicy, Policy};
|
||||
use games_ground::{Action, GroundCommand, GroundState};
|
||||
|
||||
/// Greedy's ordering with ATTACK promoted to the top of the action list.
|
||||
/// A copy of the ranking, not a call into it: `GreedyPolicy::rank` is
|
||||
/// private, and the point is to vary exactly one number.
|
||||
struct Attacker(i32);
|
||||
impl Policy for Attacker {
|
||||
fn name(&self) -> &'static str {
|
||||
"attacker"
|
||||
}
|
||||
fn choose(
|
||||
&mut self,
|
||||
state: &GroundState,
|
||||
seat: PlayerId,
|
||||
legal: &[GroundCommand],
|
||||
_p: bool,
|
||||
) -> Choice {
|
||||
let gated = state
|
||||
.players
|
||||
.get(&seat)
|
||||
.is_some_and(|p| p.stress >= 4 && !p.freedom_gate_lifted);
|
||||
let rank = |c: &GroundCommand| -> i32 {
|
||||
match c {
|
||||
GroundCommand::SelectAction {
|
||||
action, problem, ..
|
||||
} => match action {
|
||||
Action::Ground if gated => 100,
|
||||
Action::Attack => self.0, // <-- the ONE varying number (greedy: 10)
|
||||
Action::Solve
|
||||
if problem
|
||||
.and_then(|p| state.problems.get(&p))
|
||||
.is_some_and(|p| p.claimed_by.is_some()) =>
|
||||
{
|
||||
5
|
||||
}
|
||||
Action::Solve => 90,
|
||||
Action::Investigate => 80,
|
||||
Action::Support => 70,
|
||||
Action::Ground => 60,
|
||||
},
|
||||
GroundCommand::SpendFreedom => {
|
||||
if gated {
|
||||
96
|
||||
} else {
|
||||
0
|
||||
}
|
||||
}
|
||||
GroundCommand::Reveal | GroundCommand::Resolve | GroundCommand::EndRound => -1,
|
||||
_ => 50,
|
||||
}
|
||||
};
|
||||
let mut best = 0usize;
|
||||
for (i, c) in legal.iter().enumerate() {
|
||||
if rank(c) > rank(&legal[best]) {
|
||||
best = i;
|
||||
}
|
||||
}
|
||||
Choice::Command(best)
|
||||
}
|
||||
}
|
||||
|
||||
fn sweep(
|
||||
name: &str,
|
||||
players: u8,
|
||||
mk: &dyn Fn(u64, u8) -> Vec<Box<dyn Policy>>,
|
||||
) -> (u32, u32, u32, u32) {
|
||||
let (mut games, mut won, mut atk, mut darvo) = (0, 0, 0, 0);
|
||||
for seed in 0..200u64 {
|
||||
let Ok(st) = GroundState::setup(
|
||||
&Setup {
|
||||
players,
|
||||
preset: format!("standard-{players}p"),
|
||||
patch: Default::default(),
|
||||
},
|
||||
seed,
|
||||
) else {
|
||||
continue;
|
||||
};
|
||||
let mut ps = mk(seed, players);
|
||||
let Ok(g) = play(st, &mut ps) else { continue };
|
||||
games += 1;
|
||||
if g.state.outcome.as_ref().is_some_and(|o| o.group_success) {
|
||||
won += 1;
|
||||
}
|
||||
for (_, c) in &g.steps {
|
||||
if let GroundCommand::SelectAction {
|
||||
action: Action::Attack,
|
||||
..
|
||||
} = c
|
||||
{
|
||||
atk += 1;
|
||||
}
|
||||
}
|
||||
for e in &g.events {
|
||||
if matches!(e, games_ground::GroundEvent::DarvoTriggered { .. }) {
|
||||
darvo += 1;
|
||||
}
|
||||
}
|
||||
}
|
||||
let _ = name;
|
||||
(games, won, atk, darvo)
|
||||
}
|
||||
|
||||
fn main() {
|
||||
println!("F17: does ATTACK cost you the game, or does our bot just avoid it?\n");
|
||||
// The gradient, not just the extremes: greedy ranks Attack at 10,
|
||||
// below everything. 75 puts it above SUPPORT and below INVESTIGATE --
|
||||
// "attack when convenient". 95 is "attack whenever legal".
|
||||
println!(" rank=10 (greedy) rank=75 (sometimes) rank=95 (always)");
|
||||
println!("seats won atk darvo won atk darvo won atk darvo");
|
||||
for players in [2u8, 3, 4, 5, 6] {
|
||||
let cell = |r: i32, p: u8| {
|
||||
sweep("x", p, &move |_s, n| {
|
||||
(0..n)
|
||||
.map(|_| Box::new(Attacker(r)) as Box<dyn Policy>)
|
||||
.collect()
|
||||
})
|
||||
};
|
||||
let (_, gw, ga, gd) = sweep("greedy", players, &|_s, p| {
|
||||
(0..p)
|
||||
.map(|_| Box::new(GreedyPolicy) as Box<dyn Policy>)
|
||||
.collect()
|
||||
});
|
||||
let (_, mw, ma, md) = cell(75, players);
|
||||
let (_, aw, aa, ad) = cell(95, players);
|
||||
println!(" {players}p {gw:>4} {ga:>4} {gd:>5} {mw:>4} {ma:>4} {md:>5} {aw:>4} {aa:>4} {ad:>5}");
|
||||
}
|
||||
println!("\n(200 games per cell; greedy shown as the rank=10 column)");
|
||||
}
|
||||
Loading…
Add table
Add a link
Reference in a new issue