CB-REV-0003: round 3, and three of four FATAL came from round 2's fixes
Some checks failed
ci / check (push) Failing after 3s
Some checks failed
ci / check (push) Failing after 3s
The pattern is now measured over three rounds: 5 fatal, then 3 (2 from the previous round's corrections), then 4 (3 from them). The corrections are not getting safer. FATAL 1: round 2's short-cell assertion went into regulation.rs only. attack-value.rs — which produced every number in CB-EV-0030's DARVO table — still just warned, and the gate registered to close the finding claimed the property for both. FATAL 2, the sharpest of the three rounds: counting games proves they STARTED. Stopping the engine after one round gives 200 games, all-zero columns and exit 0 — byte for byte the signature CB-EV-0030 says the instrumentation distinguishes from a real result. Both harnesses now require every counted game to have reached an outcome over five rounds. FATAL 3: round 2's `.csv` filter was applied to all three loops, so catalog.yaml and rules_delta.yaml — whose missing digests were round 1's finding — were recorded and then never compared, and never checked against upstream at all. Only the parser loop filters now. FATAL 4: five of six tiebreak comparators had no coverage. GR-E04's tiebreak never executes in any scenario. All four are now covered and mutation-verified; the Blame key needed compensating claims to be reachable at all, since Blame also lowers the coalition score. SERIOUS: "peak held" computed the same number as "peak assigned" for every possible input — the real gap was that START_STRESS was an unchecked constant, now read off the dealt state; cadence="none" was a pure loophole, removed; sibling discovery swapped a hand-written list for hand-written globs and missed metadata.json and VARIANT.md, both named in the package's own changed_files — now walked, and it found them immediately; and "~72,000 games" was unsourced, make panels runs 17,600. Also separated two kinds of number that were presented alike: seats×games is invariant, 363 and 29 vary 7.1%-11.5% across samples. Round 4 owed. The conclusion is not that the work is nearly right — it is that author-made corrections to measurement work should be assumed defective until a fresh reader has attacked them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
feb68027f9
commit
c4a8a227c0
9 changed files with 386 additions and 57 deletions
|
|
@ -62,6 +62,8 @@ sha256 eb21fa3237637790fef601fe6668a190a549715e9a78b7dfe47b40cd069b648e ../cat
|
||||||
sha256 f58e81f84ea2b0d16e39932261eb3f3d9890345cdf37ad6f0b3abc00636840be ../experiments/h1-problem-stress/rules_delta.yaml
|
sha256 f58e81f84ea2b0d16e39932261eb3f3d9890345cdf37ad6f0b3abc00636840be ../experiments/h1-problem-stress/rules_delta.yaml
|
||||||
sha256 7b1cc0149122b855e827bc930576ed165bf7dd8d62707e845a9e514ce3521f8e ../experiments/h1-problem-stress/Actions.csv
|
sha256 7b1cc0149122b855e827bc930576ed165bf7dd8d62707e845a9e514ce3521f8e ../experiments/h1-problem-stress/Actions.csv
|
||||||
sha256 62785f5e7e245c60171624d15de2f40187a44fec54f93c7d9705cf52584b1078 ../experiments/h1-problem-stress/Rules_Text.csv
|
sha256 62785f5e7e245c60171624d15de2f40187a44fec54f93c7d9705cf52584b1078 ../experiments/h1-problem-stress/Rules_Text.csv
|
||||||
|
sha256 002e3adf2d038579f2803975d8c6344e2b77ee9da9eaaa1b52c8bd1666bb5201 ../experiments/h1-problem-stress/VARIANT.md
|
||||||
|
sha256 443199db94601cc889557e5e86823f374dbfdf875962fc84f95c01865605101c ../experiments/h1-problem-stress/metadata.json
|
||||||
```
|
```
|
||||||
|
|
||||||
**`rules_delta.yaml` is the load-bearing one**: it is the executable
|
**`rules_delta.yaml` is the load-bearing one**: it is the executable
|
||||||
|
|
|
||||||
|
|
@ -143,12 +143,18 @@ takes the **clamp** at 5 to collapse a gap and change who wins.
|
||||||
|
|
||||||
**Its impact was claimed and never measured, and the measurement is
|
**Its impact was claimed and never measured, and the measurement is
|
||||||
zero** (round 2, #5). This file said the pressure was invisible to *"the
|
zero** (round 2, #5). This file said the pressure was invisible to *"the
|
||||||
two modes CB-EV-0030 reports on"* without checking. The reviewer ran
|
two modes CB-EV-0030 reports on"* without checking. Reverting the fix
|
||||||
~72,000 games across all three modes with a divergence detector and found
|
leaves `attack-value`'s output **byte-identical across all 24 cells and
|
||||||
**no divergence at all**; `attack-value`'s output is byte-identical with
|
all three modes**, including the two whose `won` column reads through the
|
||||||
the fix and without it. **The defect is real in principle and its only
|
tiebreak. **The defect is real in principle and its only witness is a
|
||||||
witness is a hand-constructed board.** It is worth fixing and it changed
|
hand-constructed board.**
|
||||||
nothing that has been reported.
|
|
||||||
|
**The sample size quoted here was "~72,000 games" and was not
|
||||||
|
derivable from anything in this repo** (round 3, #9). `make panels` runs
|
||||||
|
**17,600**: `attack-value` is 2 variants × 3 modes × 3 policies × 4 seat
|
||||||
|
counts × 200 = 14,400, and `regulation` is 2 × 4 × 2 × 200 = 3,200. The
|
||||||
|
figure was adopted from a reviewer's message and never re-derived — in
|
||||||
|
the file whose whole correction history is about exactly that.
|
||||||
|
|
||||||
**Inert arms, and a second wrong-subject error caught on the way.** A
|
**Inert arms, and a second wrong-subject error caught on the way.** A
|
||||||
DARVO arm at the End of Round 5 can never advance a stage — `GameEnded`
|
DARVO arm at the End of Round 5 can never advance a stage — `GameEnded`
|
||||||
|
|
@ -162,19 +168,32 @@ follows immediately — and ground-game's criterion 1 is about DARVO
|
||||||
| 4p | 800 | 0 |
|
| 4p | 800 | 0 |
|
||||||
| 6p | 1200 | 0 |
|
| 6p | 1200 | 0 |
|
||||||
|
|
||||||
|
**Two kinds of number are in that table and this file did not distinguish
|
||||||
|
them** (round 3, #10). `600 / 800 / 1200` is **exactly `seats × games`**
|
||||||
|
and holds on every sample tried. `363` and `29` are **sample-specific**:
|
||||||
|
across four disjoint windows the 2p arms run 354–366 and the inert share
|
||||||
|
runs **7.1%–11.5%**. So "92% of arms at 2p are live" is a fact about seeds
|
||||||
|
`0..200`, not about the game — it reads 88.5% elsewhere.
|
||||||
|
|
||||||
|
The same applies to the baseline win counts quoted throughout
|
||||||
|
(`132/165/190/200`): the *identity* of greedy and reactive under baseline
|
||||||
|
holds on every sample, the **digits do not**.
|
||||||
|
|
||||||
**The first version of that metric was wrong and equalled `darvo` in
|
**The first version of that metric was wrong and equalled `darvo` in
|
||||||
every cell**, because it tested `g.rounds >= 5` — a property of the
|
every cell**, because it tested `g.rounds >= 5` — a property of the
|
||||||
*game*, not of the *event*. Every arm in every completed game was marked
|
*game*, not of the *event*. Every arm in every completed game was marked
|
||||||
inert, which briefly looked like "criterion 1 is not met after all". An
|
inert, which briefly looked like "criterion 1 is not met after all". An
|
||||||
arm is inert when no `RoundEnded` follows it.
|
arm is inert when no `RoundEnded` follows it.
|
||||||
|
|
||||||
**29-of-363 was independently confirmed twice**, and the citation for it
|
**The 29-of-363 citation has now been wrong twice and is stated plainly
|
||||||
was wrong the first time (round 2, #6): this file credited
|
here.** Round 2 found it credited to `CB-REV-0001`, which did not contain
|
||||||
`CB-REV-0001`, which does not contain the figure. Round 1's reviewer *did*
|
the figure; round 1's reviewer *had* reported it, in its message only.
|
||||||
report it — 29 of 363 at 2p, 8% — but only in its message, and it was
|
Round 3 found the remaining problem (#10, #11): both "independent
|
||||||
never transcribed into the review file. **The record is now corrected
|
confirmations" used **seeds 0..200**, so what was confirmed twice is the
|
||||||
there.** Round 2 re-derived it independently, by a different definition
|
**definition**, not the figure — and whether round 1's `StrictReactive`
|
||||||
(no later DARVO event names that player), and agrees in every cell.
|
was the corrected one-arm policy or the five-difference one **cannot be
|
||||||
|
determined from anything in this repo**, because that reviewer's code was
|
||||||
|
never kept. Treat 29/363 as one sample of a quantity that varies.
|
||||||
|
|
||||||
**Criterion 1 stands as met**: 92% of arms at 2p and all of them above are
|
**Criterion 1 stands as met**: 92% of arms at 2p and all of them above are
|
||||||
live.
|
live.
|
||||||
|
|
|
||||||
|
|
@ -24,6 +24,10 @@ use cb_kernel::PlayerId;
|
||||||
use games_ground::bot::{play, Choice, GreedyPolicy, Policy};
|
use games_ground::bot::{play, Choice, GreedyPolicy, Policy};
|
||||||
use games_ground::{Action, GroundCommand, GroundState, ScoringMode};
|
use games_ground::{Action, GroundCommand, GroundState, ScoringMode};
|
||||||
|
|
||||||
|
/// Games per cell. Named once so the footer and the assertion cannot
|
||||||
|
/// disagree — the footer said "200 games per cell" over 195-game columns.
|
||||||
|
const GAMES: u32 = 200;
|
||||||
|
|
||||||
/// Greedy's ordering with ATTACK's rank as a parameter.
|
/// Greedy's ordering with ATTACK's rank as a parameter.
|
||||||
///
|
///
|
||||||
/// A copy of the ranking rather than a call into it: `GreedyPolicy::rank`
|
/// A copy of the ranking rather than a call into it: `GreedyPolicy::rank`
|
||||||
|
|
@ -98,9 +102,10 @@ fn sweep(
|
||||||
mk: &dyn Fn(u8) -> Vec<Box<dyn Policy>>,
|
mk: &dyn Fn(u8) -> Vec<Box<dyn Policy>>,
|
||||||
) -> (u32, u32, u32, u32) {
|
) -> (u32, u32, u32, u32) {
|
||||||
let (mut games, mut won, mut atk, mut darvo) = (0, 0, 0, 0);
|
let (mut games, mut won, mut atk, mut darvo) = (0, 0, 0, 0);
|
||||||
|
let mut played = 0u32;
|
||||||
let mut errs: Vec<String> = Vec::new();
|
let mut errs: Vec<String> = Vec::new();
|
||||||
let mut setup_fails = 0u32;
|
let mut setup_fails = 0u32;
|
||||||
for seed in 0..200u64 {
|
for seed in 0..GAMES as u64 {
|
||||||
let Ok(mut st) = GroundState::setup(
|
let Ok(mut st) = GroundState::setup(
|
||||||
&Setup {
|
&Setup {
|
||||||
players,
|
players,
|
||||||
|
|
@ -130,6 +135,9 @@ fn sweep(
|
||||||
}
|
}
|
||||||
};
|
};
|
||||||
games += 1;
|
games += 1;
|
||||||
|
if g.state.outcome.is_some() && g.rounds == 5 {
|
||||||
|
played += 1;
|
||||||
|
}
|
||||||
|
|
||||||
// In the co-op mode the table wins together. In the other two the
|
// In the co-op mode the table wins together. In the other two the
|
||||||
// question is whether SEAT 0 is among the winners — because that
|
// question is whether SEAT 0 is among the winners — because that
|
||||||
|
|
@ -161,20 +169,30 @@ fn sweep(
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
if setup_fails > 0 {
|
// **Asserted, not printed** (CB-REV-0003 #1). This harness produced
|
||||||
eprintln!(" !! {setup_fails} of 200 SETUPS failed ({players}p)");
|
// every number in CB-EV-0030's DARVO table, and it still only warned
|
||||||
}
|
// on a short cell — the very defect fixed in `regulation.rs` and left
|
||||||
if games != 200 {
|
// here, while the gate registered to close it claimed the property
|
||||||
eprintln!(" !! only {games} of 200 games ran ({players}p, {mode:?})");
|
// for both.
|
||||||
}
|
assert_eq!(
|
||||||
if !errs.is_empty() {
|
games,
|
||||||
eprintln!(
|
GAMES,
|
||||||
" !! {} of 200 games did not run ({}p): first = {}",
|
"{players}p {mode:?} {variant:?}: only {games} of {GAMES} games ran \
|
||||||
errs.len(),
|
({setup_fails} setups refused, {} play errors) — the cell is short, so \
|
||||||
players,
|
every number in it is over a sample nobody chose",
|
||||||
errs[0]
|
errs.len()
|
||||||
);
|
);
|
||||||
}
|
// **And that a game RAN is not that it was PLAYED** (CB-REV-0003 #2).
|
||||||
|
// Stopping the engine after one round gave 200 games, all-zero
|
||||||
|
// columns and exit 0 — byte for byte the signature CB-EV-0030 §3
|
||||||
|
// claims to distinguish from a real result. A counted game must have
|
||||||
|
// reached an outcome.
|
||||||
|
assert_eq!(
|
||||||
|
played, GAMES,
|
||||||
|
"{players}p {mode:?} {variant:?}: {played} of {games} games reached an \
|
||||||
|
outcome — the rest stopped early, and all-zero columns would read as \
|
||||||
|
a result rather than as nothing having happened"
|
||||||
|
);
|
||||||
(games, won, atk, darvo)
|
(games, won, atk, darvo)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -71,9 +71,15 @@ impl Policy for Reactive {
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
/// `Scenarios.csv`: "All players start at Stress 2." Named rather than
|
/// The starting Stress, **read off the dealt state** rather than declared.
|
||||||
/// inlined so the peak metric cannot silently disagree with setup.
|
///
|
||||||
const START_STRESS: u8 = 2;
|
/// It was a `const 2` "named so the peak metric cannot silently disagree
|
||||||
|
/// with setup" — and nothing compared it to setup, so raising it to 7
|
||||||
|
/// printed `peak 7` in every cell with no warning (CB-REV-0003 #6).
|
||||||
|
/// Naming a constant is not checking it.
|
||||||
|
fn start_stress(state: &GroundState) -> u8 {
|
||||||
|
state.players.values().map(|p| p.stress).max().unwrap_or(0)
|
||||||
|
}
|
||||||
|
|
||||||
/// Games per cell. Named once so the assertion and the banner cannot
|
/// Games per cell. Named once so the assertion and the banner cannot
|
||||||
/// disagree — the banner said "200 games per cell" over 196-game columns.
|
/// disagree — the banner said "200 games per cell" over 196-game columns.
|
||||||
|
|
@ -94,6 +100,13 @@ struct Cell {
|
||||||
/// arm is real and can do nothing. Reported separately, because
|
/// arm is real and can do nothing. Reported separately, because
|
||||||
/// ground-game's criterion 1 is asking about DARVO *mattering*.
|
/// ground-game's criterion 1 is asking about DARVO *mattering*.
|
||||||
inert_arms: u32,
|
inert_arms: u32,
|
||||||
|
/// Games that reached an outcome over the full five rounds.
|
||||||
|
///
|
||||||
|
/// **`games` counts `play` returning `Ok`, which is not the claim.**
|
||||||
|
/// Stopping the engine after one round produced 200 games, all-zero
|
||||||
|
/// columns and exit 0 — the exact signature CB-EV-0030 §3 says the
|
||||||
|
/// instrumentation distinguishes from a real result.
|
||||||
|
played: u32,
|
||||||
}
|
}
|
||||||
|
|
||||||
fn sweep(mode: ScoringMode, variant: Variant, players: u8, reactive: bool) -> Cell {
|
fn sweep(mode: ScoringMode, variant: Variant, players: u8, reactive: bool) -> Cell {
|
||||||
|
|
@ -105,6 +118,7 @@ fn sweep(mode: ScoringMode, variant: Variant, players: u8, reactive: bool) -> Ce
|
||||||
peak_stress: 0,
|
peak_stress: 0,
|
||||||
setup_fails: 0,
|
setup_fails: 0,
|
||||||
inert_arms: 0,
|
inert_arms: 0,
|
||||||
|
played: 0,
|
||||||
};
|
};
|
||||||
for seed in 0..GAMES as u64 {
|
for seed in 0..GAMES as u64 {
|
||||||
let Ok(mut st) = GroundState::setup(
|
let Ok(mut st) = GroundState::setup(
|
||||||
|
|
@ -120,6 +134,9 @@ fn sweep(mode: ScoringMode, variant: Variant, players: u8, reactive: bool) -> Ce
|
||||||
};
|
};
|
||||||
st.mode = mode;
|
st.mode = mode;
|
||||||
st.variant = variant;
|
st.variant = variant;
|
||||||
|
// The dealt state, kept so the peak metric can read the real
|
||||||
|
// starting Stress rather than trust a constant.
|
||||||
|
let started = st.clone();
|
||||||
let mut ps: Vec<Box<dyn Policy>> = (0..players)
|
let mut ps: Vec<Box<dyn Policy>> = (0..players)
|
||||||
.map(|_| {
|
.map(|_| {
|
||||||
if reactive {
|
if reactive {
|
||||||
|
|
@ -138,6 +155,10 @@ fn sweep(mode: ScoringMode, variant: Variant, players: u8, reactive: bool) -> Ce
|
||||||
}
|
}
|
||||||
};
|
};
|
||||||
c.games += 1;
|
c.games += 1;
|
||||||
|
// A game that RAN is not a game that was PLAYED (CB-REV-0003 #2).
|
||||||
|
if g.state.outcome.is_some() && g.rounds == 5 {
|
||||||
|
c.played += 1;
|
||||||
|
}
|
||||||
if g.state.outcome.as_ref().is_some_and(|o| o.group_success) {
|
if g.state.outcome.as_ref().is_some_and(|o| o.group_success) {
|
||||||
c.won += 1;
|
c.won += 1;
|
||||||
}
|
}
|
||||||
|
|
@ -150,15 +171,19 @@ fn sweep(mode: ScoringMode, variant: Variant, players: u8, reactive: bool) -> Ce
|
||||||
c.atk += 1;
|
c.atk += 1;
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
// Peak Stress **held**, not peak Stress *assigned*.
|
// Peak Stress **held**.
|
||||||
//
|
//
|
||||||
// The first version took a maximum over `StressSet` PAYLOADS.
|
// The original took a maximum over `StressSet` PAYLOADS, and the
|
||||||
// Starting Stress is 2 and is written by `setup`, never by an
|
// starting value is written by `setup` and never by an event — so
|
||||||
// event, so a table that sat at 2 all game reported **1**, and a
|
// a table that sat at 2 all game reported 1.
|
||||||
// table with no `StressSet` at all would report 0. CB-EV-0031 §2
|
//
|
||||||
// built its headline claim on that number. Wrong subject: the
|
// **The first correction did not change the number it computes**
|
||||||
// metric answered "highest value ever assigned", the prose said
|
// (CB-REV-0003 #5): `held`'s values are exactly
|
||||||
// "highest Stress reached".
|
// `{start} ∪ {payloads}`, so with the floor applied
|
||||||
|
// `peak_held ≡ max(start, peak_payload)` for every possible input,
|
||||||
|
// and the comment claiming a change of subject was false of the
|
||||||
|
// new code too. The real fix is that `start` is now read off the
|
||||||
|
// dealt state instead of asserted by a constant.
|
||||||
// An arm is inert when NO `RoundEnded` follows it: the game
|
// An arm is inert when NO `RoundEnded` follows it: the game
|
||||||
// ended in the same End step, so the sequence never advances a
|
// ended in the same End step, so the sequence never advances a
|
||||||
// stage. `g.rounds >= 5` is a property of the GAME, not of the
|
// stage. `g.rounds >= 5` is a property of the GAME, not of the
|
||||||
|
|
@ -169,9 +194,10 @@ fn sweep(mode: ScoringMode, variant: Variant, players: u8, reactive: bool) -> Ce
|
||||||
.events
|
.events
|
||||||
.iter()
|
.iter()
|
||||||
.rposition(|e| matches!(e, games_ground::GroundEvent::RoundEnded { .. }));
|
.rposition(|e| matches!(e, games_ground::GroundEvent::RoundEnded { .. }));
|
||||||
|
let start = start_stress(&started);
|
||||||
let mut held: std::collections::BTreeMap<PlayerId, u8> =
|
let mut held: std::collections::BTreeMap<PlayerId, u8> =
|
||||||
g.state.players.keys().map(|s| (*s, START_STRESS)).collect();
|
g.state.players.keys().map(|s| (*s, start)).collect();
|
||||||
c.peak_stress = c.peak_stress.max(START_STRESS);
|
c.peak_stress = c.peak_stress.max(start);
|
||||||
for (i, e) in g.events.iter().enumerate() {
|
for (i, e) in g.events.iter().enumerate() {
|
||||||
if matches!(e, games_ground::GroundEvent::DarvoTriggered { .. }) {
|
if matches!(e, games_ground::GroundEvent::DarvoTriggered { .. }) {
|
||||||
c.darvo += 1;
|
c.darvo += 1;
|
||||||
|
|
@ -190,7 +216,7 @@ fn sweep(mode: ScoringMode, variant: Variant, players: u8, reactive: bool) -> Ce
|
||||||
// START_STRESS floor made the two agree.
|
// START_STRESS floor made the two agree.
|
||||||
c.peak_stress = c
|
c.peak_stress = c
|
||||||
.peak_stress
|
.peak_stress
|
||||||
.max(held.values().copied().max().unwrap_or(START_STRESS));
|
.max(held.values().copied().max().unwrap_or(start));
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
// **`games`, not `games + setup_fails`** (CB-REV-0002 #1).
|
// **`games`, not `games + setup_fails`** (CB-REV-0002 #1).
|
||||||
|
|
@ -207,6 +233,11 @@ fn sweep(mode: ScoringMode, variant: Variant, players: u8, reactive: bool) -> Ce
|
||||||
the cell is short, so every number in it is over a sample nobody chose",
|
the cell is short, so every number in it is over a sample nobody chose",
|
||||||
c.games, c.setup_fails
|
c.games, c.setup_fails
|
||||||
);
|
);
|
||||||
|
assert_eq!(
|
||||||
|
c.played, GAMES,
|
||||||
|
"{players}p {variant:?}: {} of {} games reached an outcome over five rounds",
|
||||||
|
c.played, c.games
|
||||||
|
);
|
||||||
if c.setup_fails > 0 {
|
if c.setup_fails > 0 {
|
||||||
eprintln!(
|
eprintln!(
|
||||||
" !! {players}p {variant:?}: {} setups refused",
|
" !! {players}p {variant:?}: {} setups refused",
|
||||||
|
|
|
||||||
|
|
@ -2427,6 +2427,108 @@ mod tests {
|
||||||
);
|
);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// **Every tiebreak comparator, both modes** (CB-REV-0003 #4).
|
||||||
|
///
|
||||||
|
/// The strengthened oracle covered **1 of 6**: GR-E03's Stress
|
||||||
|
/// key. Reversing GR-E03's *Bond* key, or **either** of GR-E04's
|
||||||
|
/// two keys, left all 179 tests and all 26 scenarios green —
|
||||||
|
/// GR-E04's tiebreak had no coverage at all, because the only
|
||||||
|
/// GR-E04 scenario sets `group_success: false`, so the winners
|
||||||
|
/// branch returns empty and the comparator never runs.
|
||||||
|
///
|
||||||
|
/// Each case is built so exactly one key decides, and the
|
||||||
|
/// expected winner is named — asserting "the set changed" is what
|
||||||
|
/// let a reversed comparator pass in the first place.
|
||||||
|
#[test]
|
||||||
|
fn every_tiebreak_key_decides_the_way_the_rules_say() {
|
||||||
|
let board = |mode: ScoringMode| {
|
||||||
|
let mut s = setup(4, Variant::Baseline, 5);
|
||||||
|
s.mode = mode;
|
||||||
|
let seats: Vec<PlayerId> = s.players.keys().copied().collect();
|
||||||
|
let ids: Vec<u32> = s.problems.keys().copied().collect();
|
||||||
|
for id in &ids {
|
||||||
|
s.problems.remove(id);
|
||||||
|
}
|
||||||
|
(s, seats)
|
||||||
|
};
|
||||||
|
let mk = |value: u8, claimed_by: Option<PlayerId>| ProblemState {
|
||||||
|
suit: Suit::Repair,
|
||||||
|
value,
|
||||||
|
face_up: true,
|
||||||
|
denied: false,
|
||||||
|
claimed_by,
|
||||||
|
protected_this_round: false,
|
||||||
|
};
|
||||||
|
|
||||||
|
// GR-E03 key 2, Stress: equal claims, lower Stress wins.
|
||||||
|
let (mut s, seats) = board(ScoringMode::CommonProblem);
|
||||||
|
s.problems.insert(1, mk(4, Some(seats[0])));
|
||||||
|
s.problems.insert(2, mk(4, Some(seats[1])));
|
||||||
|
s.players.get_mut(&seats[0]).expect("a").stress = 5;
|
||||||
|
s.players.get_mut(&seats[1]).expect("b").stress = 1;
|
||||||
|
assert_eq!(
|
||||||
|
s.score().winners,
|
||||||
|
vec![seats[1]],
|
||||||
|
"GR-E03: lower Stress must win the tiebreak"
|
||||||
|
);
|
||||||
|
|
||||||
|
// GR-E03 key 3, Bonds: equal claims AND equal Stress, more
|
||||||
|
// Bonds wins.
|
||||||
|
s.players.get_mut(&seats[0]).expect("a").stress = 1;
|
||||||
|
s.relations
|
||||||
|
.insert(Pair::new(seats[0], seats[2]), Relation::Bond);
|
||||||
|
assert_eq!(
|
||||||
|
s.score().winners,
|
||||||
|
vec![seats[0]],
|
||||||
|
"GR-E03: with claims and Stress level, more Bonds must win"
|
||||||
|
);
|
||||||
|
|
||||||
|
// GR-E04 key 2, combined Stress: two coalitions of equal
|
||||||
|
// score, the calmer one wins.
|
||||||
|
let (mut s, seats) = board(ScoringMode::BondedCoalitions);
|
||||||
|
s.problems.insert(1, mk(4, Some(seats[0])));
|
||||||
|
s.problems.insert(2, mk(4, Some(seats[2])));
|
||||||
|
s.relations
|
||||||
|
.insert(Pair::new(seats[0], seats[1]), Relation::Bond);
|
||||||
|
s.relations
|
||||||
|
.insert(Pair::new(seats[2], seats[3]), Relation::Bond);
|
||||||
|
for seat in &seats {
|
||||||
|
s.players.get_mut(seat).expect("p").stress = 1;
|
||||||
|
}
|
||||||
|
s.players.get_mut(&seats[0]).expect("p").stress = 5;
|
||||||
|
let won = s.score().winners;
|
||||||
|
assert!(
|
||||||
|
won.contains(&seats[2]) && !won.contains(&seats[0]),
|
||||||
|
"GR-E04: the coalition with lower combined Stress must win, got {won:?}"
|
||||||
|
);
|
||||||
|
|
||||||
|
// GR-E04 key 3, Blame — and reaching it takes care, which is
|
||||||
|
// the point. A coalition's score is `sum(claimed - blame)`, so
|
||||||
|
// a Blame token lowers the score too and key 1 decides first.
|
||||||
|
// The key is only reachable when the claims COMPENSATE: 5
|
||||||
|
// claimed with one Blame ties 4 claimed with none.
|
||||||
|
s.players.get_mut(&seats[0]).expect("p").stress = 1;
|
||||||
|
s.problems.insert(1, mk(5, Some(seats[0])));
|
||||||
|
s.problems.insert(2, mk(4, Some(seats[2])));
|
||||||
|
s.players.get_mut(&seats[0]).expect("p").blame_from = vec![seats[3]];
|
||||||
|
let scored = s.score();
|
||||||
|
assert_eq!(
|
||||||
|
scored.coalitions.len(),
|
||||||
|
2,
|
||||||
|
"the fixture needs two coalitions to compare"
|
||||||
|
);
|
||||||
|
assert_eq!(
|
||||||
|
scored.coalitions[0].score, scored.coalitions[1].score,
|
||||||
|
"the Blame key is unreachable unless the scores tie: {:?}",
|
||||||
|
scored.coalitions
|
||||||
|
);
|
||||||
|
let won = scored.winners;
|
||||||
|
assert!(
|
||||||
|
won.contains(&seats[2]) && !won.contains(&seats[0]),
|
||||||
|
"GR-E04: with score and Stress level, fewer Blame must win, got {won:?}"
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
/// **`rules_delta.yaml`'s `unchanged:` list is ground-game's claim
|
/// **`rules_delta.yaml`'s `unchanged:` list is ground-game's claim
|
||||||
/// about their own experiment, and it is checkable.**
|
/// about their own experiment, and it is checkable.**
|
||||||
///
|
///
|
||||||
|
|
|
||||||
12
gates.toml
12
gates.toml
|
|
@ -137,7 +137,7 @@ retire_if = "two passes run with no finding while artifacts keep growing — tha
|
||||||
id = "CHAOS"
|
id = "CHAOS"
|
||||||
name = "the chaos roll"
|
name = "the chaos roll"
|
||||||
target = ""
|
target = ""
|
||||||
cadence = "none"
|
cadence = "manual"
|
||||||
checks = "d8 on each tier declaration, 12-declaration calibration window (window 2, opened 2026-08-03; window 1 ran at d4)"
|
checks = "d8 on each tier declaration, 12-declaration calibration window (window 2, opened 2026-08-03; window 1 ran at d4)"
|
||||||
added = "2026-07-30"
|
added = "2026-07-30"
|
||||||
review_by = "2026-11-30"
|
review_by = "2026-11-30"
|
||||||
|
|
@ -178,7 +178,7 @@ id = "CB-REV-0002/7"
|
||||||
name = "variant panels"
|
name = "variant panels"
|
||||||
target = "panels"
|
target = "panels"
|
||||||
cadence = "all"
|
cadence = "all"
|
||||||
checks = "the H1 measurement harnesses actually run, and a short cell fails rather than printing a number a reader must notice"
|
checks = "both H1 measurement harnesses run; a short cell fails, and so does a cell whose games did not reach an outcome"
|
||||||
notes = """
|
notes = """
|
||||||
Registered because the second adversarial review asked what the harness
|
Registered because the second adversarial review asked what the harness
|
||||||
would report if the work silently stopped, and the answer was "green, and
|
would report if the work silently stopped, and the answer was "green, and
|
||||||
|
|
@ -186,4 +186,12 @@ nothing else": `attack-value` and `regulation` were wired into no target,
|
||||||
so every figure in CB-EV-0030 and CB-EV-0031 came from a manual run of an
|
so every figure in CB-EV-0030 and CB-EV-0031 came from a manual run of an
|
||||||
ungated binary -- including the assertion added to catch short cells,
|
ungated binary -- including the assertion added to catch short cells,
|
||||||
which was unreachable from `make`.
|
which was unreachable from `make`.
|
||||||
|
|
||||||
|
CB-REV-0003 #1 then found this `checks` line claimed a property that held
|
||||||
|
for only one of the two harnesses: `attack-value`, which produced every
|
||||||
|
number in CB-EV-0030's table, still only warned. And #2 found that
|
||||||
|
counting games proves they STARTED: stopping the engine after one round
|
||||||
|
gave 200 games, all-zero columns and exit 0 -- the signature CB-EV-0030
|
||||||
|
says the instrumentation distinguishes from a result. Both now assert an
|
||||||
|
outcome was reached.
|
||||||
"""
|
"""
|
||||||
|
|
|
||||||
127
reviews/CB-REV-0003-h1.md
Normal file
127
reviews/CB-REV-0003-h1.md
Normal file
|
|
@ -0,0 +1,127 @@
|
||||||
|
# CB-REV-0003 — round 3
|
||||||
|
|
||||||
|
Fresh agent again. Run 2026-08-08 against the corrections made after
|
||||||
|
[`CB-REV-0002`](CB-REV-0002-h1.md).
|
||||||
|
|
||||||
|
> **Verdict: not approvable. Four FATAL, five SERIOUS, two MINOR — and
|
||||||
|
> three of the four FATAL were created or left by round 2's corrections.**
|
||||||
|
>
|
||||||
|
> **The pattern is now established over three rounds** and is the most
|
||||||
|
> useful thing this review has produced.
|
||||||
|
|
||||||
|
| round | challenges | FATAL | of which introduced by the previous round's fix |
|
||||||
|
|---|---:|---:|---:|
|
||||||
|
| 1 | 13 | 5 | — |
|
||||||
|
| 2 | 8 | 3 | 2 |
|
||||||
|
| 3 | 11 | 4 | 3 |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## The four fatal ones
|
||||||
|
|
||||||
|
### 1. The fix for round 2's #1 was applied to one of the two harnesses
|
||||||
|
|
||||||
|
`regulation.rs` got `assert_eq!(games, GAMES)`. **`attack-value.rs` — which
|
||||||
|
produced every number in CB-EV-0030's DARVO table — kept `if games != 200
|
||||||
|
{ eprintln!(..) }`.** Injected setup refusals gave exit 0 and a full table
|
||||||
|
over 195-game columns.
|
||||||
|
|
||||||
|
**And the gate registered to close the finding asserted the property for
|
||||||
|
both.** Its `checks` line was false of half of what it gates.
|
||||||
|
|
||||||
|
**Conceded.** Both harnesses assert now; the `checks` line says what is
|
||||||
|
actually checked.
|
||||||
|
|
||||||
|
### 2. Counting games proves they STARTED, not that they were PLAYED
|
||||||
|
|
||||||
|
Stopping the engine after one round (`round >= 5` → `>= 1`) gives **200
|
||||||
|
games in every cell, every value zero, exit 0, `make panels` green** —
|
||||||
|
byte for byte the signature CB-EV-0030 §3 claims the instrumentation
|
||||||
|
distinguishes from a real result.
|
||||||
|
|
||||||
|
**Conceded, and it is the sharpest finding of the three rounds.** The
|
||||||
|
counter answered *"did `play` return `Ok`"* while the claim made of it was
|
||||||
|
*"the zeros are real"*. Both harnesses now require every counted game to
|
||||||
|
have reached an outcome over five rounds; the one-round mutation fails
|
||||||
|
with the right message.
|
||||||
|
|
||||||
|
### 3. Round 2's `.csv` filter disabled the checks it was added beside
|
||||||
|
|
||||||
|
The filter was applied to **all three** loops — digest comparison, the CSV
|
||||||
|
parser check, and upstream freshness. So `catalog.yaml` and
|
||||||
|
`rules_delta.yaml`, whose missing digests were round 1's finding, were
|
||||||
|
recorded and then **never compared**, and never checked against upstream
|
||||||
|
at all. Tampering with both produced `[ok] × 12, exit 0`.
|
||||||
|
|
||||||
|
**Conceded.** Only the parser loop filters; tampering with either YAML now
|
||||||
|
fails on digest *and* freshness.
|
||||||
|
|
||||||
|
### 4. Five of six tiebreak comparators had no coverage
|
||||||
|
|
||||||
|
The strengthened oracle covered GR-E03's Stress key. **GR-E03's Bond key
|
||||||
|
and both of GR-E04's keys could be reversed with 179 tests and 26
|
||||||
|
scenarios green** — GR-E04's tiebreak never executes in any scenario,
|
||||||
|
because the only GR-E04 scenario sets `group_success: false`.
|
||||||
|
|
||||||
|
**Conceded.** All four are now covered and each was verified red under
|
||||||
|
mutation. **The Blame key took care to reach**: a coalition's score is
|
||||||
|
`sum(claimed − blame)`, so a Blame token lowers the score and key 1
|
||||||
|
decides first — the key is only reachable when claims compensate, and the
|
||||||
|
test asserts the scores tie before relying on it.
|
||||||
|
|
||||||
|
## The serious ones
|
||||||
|
|
||||||
|
| # | challenge | outcome |
|
||||||
|
|---|---|---|
|
||||||
|
| 5 | "peak Stress **held**, not assigned" — `peak_held ≡ max(start, peak_payload)` for every possible input, so the correction computed the same number | **conceded.** The fix was cosmetic and the comment was false of the new code too. The real coupling was missing, which is #6 |
|
||||||
|
| 6 | `START_STRESS` was an unchecked constant — setting it to 7 printed `peak 7` everywhere | **conceded.** Now read off the dealt state. Verified: changing setup's Stress to 3 moves the reported peak to 3 |
|
||||||
|
| 7 | `cadence = "none"` was a pure loophole — a real target that runs nowhere and passes | **conceded, removed.** It had one user, the empty-target CHAOS entry, which the loop already skips. `manual` remains self-declared and unchecked, and that is now stated rather than implied |
|
||||||
|
| 8 | sibling discovery swapped a hand-written list for three hand-written globs; `metadata.json` and `VARIANT.md` were invisible, and both are named in the package's own `changed_files` | **conceded.** Walked, not globbed. **It found both immediately**, plus nested files the globs could never reach |
|
||||||
|
| 9 | "~72,000 games" is not producible; `make panels` runs **17,600** | **conceded.** Adopted from a reviewer's message and never re-derived — in the file whose correction history is about exactly that. The conclusion (zero divergence) reproduces |
|
||||||
|
|
||||||
|
## The minor ones
|
||||||
|
|
||||||
|
- **10 — two kinds of number were presented alike.** `600/800/1200` is
|
||||||
|
exactly `seats × games` on every sample; **`363` and `29` vary** —
|
||||||
|
7.1%–11.5% inert across four windows. "92% of arms are live" is a fact
|
||||||
|
about seeds `0..200`. Now separated.
|
||||||
|
- **11 — the 29/363 citation, wrong twice, is now stated as one sample**
|
||||||
|
of a varying quantity. Both "independent confirmations" used the same
|
||||||
|
seeds, so what was confirmed was the *definition*. Whether round 1's
|
||||||
|
`StrictReactive` was the corrected policy cannot be determined: **that
|
||||||
|
reviewer's code was never kept**, which is itself worth fixing.
|
||||||
|
|
||||||
|
## What held, after being attacked
|
||||||
|
|
||||||
|
**Every one of the 24 cells in CB-EV-0030's DARVO table reproduces
|
||||||
|
exactly** — the first of three attempts at that table to produce no wrong
|
||||||
|
number. `regulation`'s `assert_eq!(c.games, GAMES)` catches both the setup
|
||||||
|
and the `play` path. The #13 fix is genuinely controlled. H1-B's mutation
|
||||||
|
coverage is real. `make all` runs `panels` and fails on it. The
|
||||||
|
1-arm-per-seat-per-game invariant holds on every sample — **and is still
|
||||||
|
unexplained.**
|
||||||
|
|
||||||
|
## What could not be checked
|
||||||
|
|
||||||
|
Upstream freshness (`../ground-game` not checked out, so that half of
|
||||||
|
`edition-check` has never run here); whether `metadata.json` and
|
||||||
|
`VARIANT.md` are load-bearing to anything; round 1's `StrictReactive`;
|
||||||
|
felt-play; cost.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## The finding that outlasts H1
|
||||||
|
|
||||||
|
Three rounds, each correcting the last, each introducing defects of the
|
||||||
|
same class. **The corrections are not getting safer.**
|
||||||
|
|
||||||
|
> A correction is written under the belief that the error is now
|
||||||
|
> understood. That belief is the condition under which this class of
|
||||||
|
> error is produced — so the correction inherits it, and the next round
|
||||||
|
> finds the same shape one level in.
|
||||||
|
|
||||||
|
**Round 4 is owed by the same argument.** The honest conclusion is not
|
||||||
|
that the work is nearly right; it is that **author-made corrections to
|
||||||
|
measurement work should be assumed defective until a fresh reader has
|
||||||
|
attacked them**, and this project should stop treating "corrected" as a
|
||||||
|
state closer to done than "found wrong".
|
||||||
|
|
@ -66,18 +66,27 @@ SIBLINGS = "../"
|
||||||
# raised FileNotFoundError from `digest` instead of failing with the
|
# raised FileNotFoundError from `digest` instead of failing with the
|
||||||
# designed message, and "the sibling packages are covered" counted lines
|
# designed message, and "the sibling packages are covered" counted lines
|
||||||
# in a Markdown file -- it passed with all three files deleted.
|
# in a Markdown file -- it passed with all three files deleted.
|
||||||
SIBLING_GLOBS = ("catalog.yaml", "experiments/*/rules_delta.yaml", "experiments/*/*.csv")
|
# Everything under `editions/` that is not the edition directory itself.
|
||||||
|
# **Walked, not globbed** (CB-REV-0003 #8): the first version listed three
|
||||||
|
# hand-written glob patterns, which is the same self-certifying shape as
|
||||||
|
# reading the list out of PROVENANCE -- one hand-written list swapped for
|
||||||
|
# another. It missed `metadata.json` and `VARIANT.md`, both named in the
|
||||||
|
# package's OWN `changed_files` manifest, and anything a directory deeper.
|
||||||
|
SIBLING_SKIP = {".DS_Store"}
|
||||||
|
|
||||||
|
|
||||||
def sibling_files():
|
def sibling_files():
|
||||||
"""Sibling packages that are on disk, as `../`-relative paths."""
|
"""Every file beside the edition, as `../`-relative paths."""
|
||||||
import glob as _glob
|
base = os.path.join(ROOT, "editions")
|
||||||
|
edition_dir = os.path.join(ROOT, EDITION)
|
||||||
root = os.path.join(ROOT, EDITION, SIBLINGS)
|
|
||||||
out = []
|
out = []
|
||||||
for pattern in SIBLING_GLOBS:
|
for dirpath, _dirs, files in os.walk(base):
|
||||||
for hit in _glob.glob(os.path.join(root, pattern)):
|
if os.path.abspath(dirpath).startswith(os.path.abspath(edition_dir)):
|
||||||
rel = os.path.relpath(hit, os.path.join(ROOT, EDITION))
|
continue
|
||||||
|
for name in files:
|
||||||
|
if name in SIBLING_SKIP:
|
||||||
|
continue
|
||||||
|
rel = os.path.relpath(os.path.join(dirpath, name), edition_dir)
|
||||||
out.append(rel.replace(os.sep, "/"))
|
out.append(rel.replace(os.sep, "/"))
|
||||||
return sorted(out)
|
return sorted(out)
|
||||||
|
|
||||||
|
|
@ -98,7 +107,7 @@ def check():
|
||||||
print(f" [FAIL] a digest is recorded for a file that is not here: {', '.join(missing)}")
|
print(f" [FAIL] a digest is recorded for a file that is not here: {', '.join(missing)}")
|
||||||
rc = 1
|
rc = 1
|
||||||
|
|
||||||
for name in [f for f in present if f.endswith(".csv")]:
|
for name in present:
|
||||||
if name not in want:
|
if name not in want:
|
||||||
continue
|
continue
|
||||||
path = os.path.join(ROOT, EDITION, name)
|
path = os.path.join(ROOT, EDITION, name)
|
||||||
|
|
@ -118,6 +127,17 @@ def check():
|
||||||
# ADR-0015 D3's falsifier, checked rather than asserted: the hand
|
# ADR-0015 D3's falsifier, checked rather than asserted: the hand
|
||||||
# reader handles commas inside quotes and NOTHING ELSE. A doubled
|
# reader handles commas inside quotes and NOTHING ELSE. A doubled
|
||||||
# quote or an embedded newline means `csv` is the answer after all.
|
# quote or an embedded newline means `csv` is the answer after all.
|
||||||
|
# ONLY this loop filters: it is ADR-0015 D3's falsifier about the
|
||||||
|
# hand-rolled CSV reader, and running it over YAML made a valid
|
||||||
|
# `catalog.yaml` line with an odd quote count fail with a message
|
||||||
|
# about a parser that never reads it (CB-REV-0002 #9).
|
||||||
|
#
|
||||||
|
# **The digest and freshness loops must NOT filter** — CB-REV-0003 #3:
|
||||||
|
# the round-2 correction applied this filter to all three, so the two
|
||||||
|
# sibling YAMLs whose missing digests were round 1's finding were
|
||||||
|
# recorded and then never compared, and never checked against
|
||||||
|
# upstream at all. `make edition-check` answered its own headline
|
||||||
|
# question with [ok] when the answer was no.
|
||||||
for name in [f for f in present if f.endswith(".csv")]:
|
for name in [f for f in present if f.endswith(".csv")]:
|
||||||
raw = open(os.path.join(ROOT, EDITION, name), encoding="utf-8-sig").read()
|
raw = open(os.path.join(ROOT, EDITION, name), encoding="utf-8-sig").read()
|
||||||
if '""' in raw:
|
if '""' in raw:
|
||||||
|
|
@ -138,7 +158,7 @@ def check():
|
||||||
print(" [----] upstream not checked out — freshness UNVERIFIED")
|
print(" [----] upstream not checked out — freshness UNVERIFIED")
|
||||||
print(f" expected {UPSTREAM_DIR}")
|
print(f" expected {UPSTREAM_DIR}")
|
||||||
return rc
|
return rc
|
||||||
for name in [f for f in present if f.endswith(".csv")]:
|
for name in present:
|
||||||
up = os.path.join(UPSTREAM_DIR, name)
|
up = os.path.join(UPSTREAM_DIR, name)
|
||||||
if not os.path.exists(up):
|
if not os.path.exists(up):
|
||||||
print(f" [FAIL] {name} is not in upstream — where did it come from?")
|
print(f" [FAIL] {name} is not in upstream — where did it come from?")
|
||||||
|
|
|
||||||
|
|
@ -316,11 +316,13 @@ def check_gate_registry(root=REPO):
|
||||||
if not target:
|
if not target:
|
||||||
continue
|
continue
|
||||||
cadence = g.get("cadence")
|
cadence = g.get("cadence")
|
||||||
if cadence not in ("all", "manual", "none"):
|
if cadence not in ("all", "manual"):
|
||||||
out.append(Finding(
|
out.append(Finding(
|
||||||
"gates", "gates.toml",
|
"gates", "gates.toml",
|
||||||
f"{target!r} declares no cadence — say whether `make all` runs "
|
f"{target!r} declares no cadence — say `all` or `manual`. "
|
||||||
f"it, or a gate can exist without ever running"))
|
f"`none` was a pure loophole: a real target that runs "
|
||||||
|
f"nowhere and passes, which is the condition this rule "
|
||||||
|
f"exists to prevent (CB-REV-0003 #7)"))
|
||||||
elif cadence == "all" and target not in deps:
|
elif cadence == "all" and target not in deps:
|
||||||
out.append(Finding(
|
out.append(Finding(
|
||||||
"gates", "gates.toml",
|
"gates", "gates.toml",
|
||||||
|
|
|
||||||
Loading…
Add table
Add a link
Reference in a new issue