CB-REV-0003: round 3, and three of four FATAL came from round 2's fixes
Some checks failed
ci / check (push) Failing after 3s

The pattern is now measured over three rounds: 5 fatal, then 3 (2 from the
previous round's corrections), then 4 (3 from them). The corrections are
not getting safer.

FATAL 1: round 2's short-cell assertion went into regulation.rs only.
attack-value.rs — which produced every number in CB-EV-0030's DARVO table
— still just warned, and the gate registered to close the finding claimed
the property for both.

FATAL 2, the sharpest of the three rounds: counting games proves they
STARTED. Stopping the engine after one round gives 200 games, all-zero
columns and exit 0 — byte for byte the signature CB-EV-0030 says the
instrumentation distinguishes from a real result. Both harnesses now
require every counted game to have reached an outcome over five rounds.

FATAL 3: round 2's `.csv` filter was applied to all three loops, so
catalog.yaml and rules_delta.yaml — whose missing digests were round 1's
finding — were recorded and then never compared, and never checked against
upstream at all. Only the parser loop filters now.

FATAL 4: five of six tiebreak comparators had no coverage. GR-E04's
tiebreak never executes in any scenario. All four are now covered and
mutation-verified; the Blame key needed compensating claims to be
reachable at all, since Blame also lowers the coalition score.

SERIOUS: "peak held" computed the same number as "peak assigned" for every
possible input — the real gap was that START_STRESS was an unchecked
constant, now read off the dealt state; cadence="none" was a pure
loophole, removed; sibling discovery swapped a hand-written list for
hand-written globs and missed metadata.json and VARIANT.md, both named in
the package's own changed_files — now walked, and it found them
immediately; and "~72,000 games" was unsourced, make panels runs 17,600.

Also separated two kinds of number that were presented alike: seats×games
is invariant, 363 and 29 vary 7.1%-11.5% across samples.

Round 4 owed. The conclusion is not that the work is nearly right — it is
that author-made corrections to measurement work should be assumed
defective until a fresh reader has attacked them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
tegwick 2026-08-08 10:26:25 +02:00
parent feb68027f9
commit c4a8a227c0
9 changed files with 386 additions and 57 deletions

View file

@ -62,6 +62,8 @@ sha256 eb21fa3237637790fef601fe6668a190a549715e9a78b7dfe47b40cd069b648e ../cat
sha256 f58e81f84ea2b0d16e39932261eb3f3d9890345cdf37ad6f0b3abc00636840be ../experiments/h1-problem-stress/rules_delta.yaml sha256 f58e81f84ea2b0d16e39932261eb3f3d9890345cdf37ad6f0b3abc00636840be ../experiments/h1-problem-stress/rules_delta.yaml
sha256 7b1cc0149122b855e827bc930576ed165bf7dd8d62707e845a9e514ce3521f8e ../experiments/h1-problem-stress/Actions.csv sha256 7b1cc0149122b855e827bc930576ed165bf7dd8d62707e845a9e514ce3521f8e ../experiments/h1-problem-stress/Actions.csv
sha256 62785f5e7e245c60171624d15de2f40187a44fec54f93c7d9705cf52584b1078 ../experiments/h1-problem-stress/Rules_Text.csv sha256 62785f5e7e245c60171624d15de2f40187a44fec54f93c7d9705cf52584b1078 ../experiments/h1-problem-stress/Rules_Text.csv
sha256 002e3adf2d038579f2803975d8c6344e2b77ee9da9eaaa1b52c8bd1666bb5201 ../experiments/h1-problem-stress/VARIANT.md
sha256 443199db94601cc889557e5e86823f374dbfdf875962fc84f95c01865605101c ../experiments/h1-problem-stress/metadata.json
``` ```
**`rules_delta.yaml` is the load-bearing one**: it is the executable **`rules_delta.yaml` is the load-bearing one**: it is the executable

View file

@ -143,12 +143,18 @@ takes the **clamp** at 5 to collapse a gap and change who wins.
**Its impact was claimed and never measured, and the measurement is **Its impact was claimed and never measured, and the measurement is
zero** (round 2, #5). This file said the pressure was invisible to *"the zero** (round 2, #5). This file said the pressure was invisible to *"the
two modes CB-EV-0030 reports on"* without checking. The reviewer ran two modes CB-EV-0030 reports on"* without checking. Reverting the fix
~72,000 games across all three modes with a divergence detector and found leaves `attack-value`'s output **byte-identical across all 24 cells and
**no divergence at all**; `attack-value`'s output is byte-identical with all three modes**, including the two whose `won` column reads through the
the fix and without it. **The defect is real in principle and its only tiebreak. **The defect is real in principle and its only witness is a
witness is a hand-constructed board.** It is worth fixing and it changed hand-constructed board.**
nothing that has been reported.
**The sample size quoted here was "~72,000 games" and was not
derivable from anything in this repo** (round 3, #9). `make panels` runs
**17,600**: `attack-value` is 2 variants × 3 modes × 3 policies × 4 seat
counts × 200 = 14,400, and `regulation` is 2 × 4 × 2 × 200 = 3,200. The
figure was adopted from a reviewer's message and never re-derived — in
the file whose whole correction history is about exactly that.
**Inert arms, and a second wrong-subject error caught on the way.** A **Inert arms, and a second wrong-subject error caught on the way.** A
DARVO arm at the End of Round 5 can never advance a stage — `GameEnded` DARVO arm at the End of Round 5 can never advance a stage — `GameEnded`
@ -162,19 +168,32 @@ follows immediately — and ground-game's criterion 1 is about DARVO
| 4p | 800 | 0 | | 4p | 800 | 0 |
| 6p | 1200 | 0 | | 6p | 1200 | 0 |
**Two kinds of number are in that table and this file did not distinguish
them** (round 3, #10). `600 / 800 / 1200` is **exactly `seats × games`**
and holds on every sample tried. `363` and `29` are **sample-specific**:
across four disjoint windows the 2p arms run 354366 and the inert share
runs **7.1%11.5%**. So "92% of arms at 2p are live" is a fact about seeds
`0..200`, not about the game — it reads 88.5% elsewhere.
The same applies to the baseline win counts quoted throughout
(`132/165/190/200`): the *identity* of greedy and reactive under baseline
holds on every sample, the **digits do not**.
**The first version of that metric was wrong and equalled `darvo` in **The first version of that metric was wrong and equalled `darvo` in
every cell**, because it tested `g.rounds >= 5` — a property of the every cell**, because it tested `g.rounds >= 5` — a property of the
*game*, not of the *event*. Every arm in every completed game was marked *game*, not of the *event*. Every arm in every completed game was marked
inert, which briefly looked like "criterion 1 is not met after all". An inert, which briefly looked like "criterion 1 is not met after all". An
arm is inert when no `RoundEnded` follows it. arm is inert when no `RoundEnded` follows it.
**29-of-363 was independently confirmed twice**, and the citation for it **The 29-of-363 citation has now been wrong twice and is stated plainly
was wrong the first time (round 2, #6): this file credited here.** Round 2 found it credited to `CB-REV-0001`, which did not contain
`CB-REV-0001`, which does not contain the figure. Round 1's reviewer *did* the figure; round 1's reviewer *had* reported it, in its message only.
report it — 29 of 363 at 2p, 8% — but only in its message, and it was Round 3 found the remaining problem (#10, #11): both "independent
never transcribed into the review file. **The record is now corrected confirmations" used **seeds 0..200**, so what was confirmed twice is the
there.** Round 2 re-derived it independently, by a different definition **definition**, not the figure — and whether round 1's `StrictReactive`
(no later DARVO event names that player), and agrees in every cell. was the corrected one-arm policy or the five-difference one **cannot be
determined from anything in this repo**, because that reviewer's code was
never kept. Treat 29/363 as one sample of a quantity that varies.
**Criterion 1 stands as met**: 92% of arms at 2p and all of them above are **Criterion 1 stands as met**: 92% of arms at 2p and all of them above are
live. live.

View file

@ -24,6 +24,10 @@ use cb_kernel::PlayerId;
use games_ground::bot::{play, Choice, GreedyPolicy, Policy}; use games_ground::bot::{play, Choice, GreedyPolicy, Policy};
use games_ground::{Action, GroundCommand, GroundState, ScoringMode}; use games_ground::{Action, GroundCommand, GroundState, ScoringMode};
/// Games per cell. Named once so the footer and the assertion cannot
/// disagree — the footer said "200 games per cell" over 195-game columns.
const GAMES: u32 = 200;
/// Greedy's ordering with ATTACK's rank as a parameter. /// Greedy's ordering with ATTACK's rank as a parameter.
/// ///
/// A copy of the ranking rather than a call into it: `GreedyPolicy::rank` /// A copy of the ranking rather than a call into it: `GreedyPolicy::rank`
@ -98,9 +102,10 @@ fn sweep(
mk: &dyn Fn(u8) -> Vec<Box<dyn Policy>>, mk: &dyn Fn(u8) -> Vec<Box<dyn Policy>>,
) -> (u32, u32, u32, u32) { ) -> (u32, u32, u32, u32) {
let (mut games, mut won, mut atk, mut darvo) = (0, 0, 0, 0); let (mut games, mut won, mut atk, mut darvo) = (0, 0, 0, 0);
let mut played = 0u32;
let mut errs: Vec<String> = Vec::new(); let mut errs: Vec<String> = Vec::new();
let mut setup_fails = 0u32; let mut setup_fails = 0u32;
for seed in 0..200u64 { for seed in 0..GAMES as u64 {
let Ok(mut st) = GroundState::setup( let Ok(mut st) = GroundState::setup(
&Setup { &Setup {
players, players,
@ -130,6 +135,9 @@ fn sweep(
} }
}; };
games += 1; games += 1;
if g.state.outcome.is_some() && g.rounds == 5 {
played += 1;
}
// In the co-op mode the table wins together. In the other two the // In the co-op mode the table wins together. In the other two the
// question is whether SEAT 0 is among the winners — because that // question is whether SEAT 0 is among the winners — because that
@ -161,20 +169,30 @@ fn sweep(
} }
} }
} }
if setup_fails > 0 { // **Asserted, not printed** (CB-REV-0003 #1). This harness produced
eprintln!(" !! {setup_fails} of 200 SETUPS failed ({players}p)"); // every number in CB-EV-0030's DARVO table, and it still only warned
} // on a short cell — the very defect fixed in `regulation.rs` and left
if games != 200 { // here, while the gate registered to close it claimed the property
eprintln!(" !! only {games} of 200 games ran ({players}p, {mode:?})"); // for both.
} assert_eq!(
if !errs.is_empty() { games,
eprintln!( GAMES,
" !! {} of 200 games did not run ({}p): first = {}", "{players}p {mode:?} {variant:?}: only {games} of {GAMES} games ran \
errs.len(), ({setup_fails} setups refused, {} play errors) the cell is short, so \
players, every number in it is over a sample nobody chose",
errs[0] errs.len()
); );
} // **And that a game RAN is not that it was PLAYED** (CB-REV-0003 #2).
// Stopping the engine after one round gave 200 games, all-zero
// columns and exit 0 — byte for byte the signature CB-EV-0030 §3
// claims to distinguish from a real result. A counted game must have
// reached an outcome.
assert_eq!(
played, GAMES,
"{players}p {mode:?} {variant:?}: {played} of {games} games reached an \
outcome the rest stopped early, and all-zero columns would read as \
a result rather than as nothing having happened"
);
(games, won, atk, darvo) (games, won, atk, darvo)
} }

View file

@ -71,9 +71,15 @@ impl Policy for Reactive {
} }
} }
/// `Scenarios.csv`: "All players start at Stress 2." Named rather than /// The starting Stress, **read off the dealt state** rather than declared.
/// inlined so the peak metric cannot silently disagree with setup. ///
const START_STRESS: u8 = 2; /// It was a `const 2` "named so the peak metric cannot silently disagree
/// with setup" — and nothing compared it to setup, so raising it to 7
/// printed `peak 7` in every cell with no warning (CB-REV-0003 #6).
/// Naming a constant is not checking it.
fn start_stress(state: &GroundState) -> u8 {
state.players.values().map(|p| p.stress).max().unwrap_or(0)
}
/// Games per cell. Named once so the assertion and the banner cannot /// Games per cell. Named once so the assertion and the banner cannot
/// disagree — the banner said "200 games per cell" over 196-game columns. /// disagree — the banner said "200 games per cell" over 196-game columns.
@ -94,6 +100,13 @@ struct Cell {
/// arm is real and can do nothing. Reported separately, because /// arm is real and can do nothing. Reported separately, because
/// ground-game's criterion 1 is asking about DARVO *mattering*. /// ground-game's criterion 1 is asking about DARVO *mattering*.
inert_arms: u32, inert_arms: u32,
/// Games that reached an outcome over the full five rounds.
///
/// **`games` counts `play` returning `Ok`, which is not the claim.**
/// Stopping the engine after one round produced 200 games, all-zero
/// columns and exit 0 — the exact signature CB-EV-0030 §3 says the
/// instrumentation distinguishes from a real result.
played: u32,
} }
fn sweep(mode: ScoringMode, variant: Variant, players: u8, reactive: bool) -> Cell { fn sweep(mode: ScoringMode, variant: Variant, players: u8, reactive: bool) -> Cell {
@ -105,6 +118,7 @@ fn sweep(mode: ScoringMode, variant: Variant, players: u8, reactive: bool) -> Ce
peak_stress: 0, peak_stress: 0,
setup_fails: 0, setup_fails: 0,
inert_arms: 0, inert_arms: 0,
played: 0,
}; };
for seed in 0..GAMES as u64 { for seed in 0..GAMES as u64 {
let Ok(mut st) = GroundState::setup( let Ok(mut st) = GroundState::setup(
@ -120,6 +134,9 @@ fn sweep(mode: ScoringMode, variant: Variant, players: u8, reactive: bool) -> Ce
}; };
st.mode = mode; st.mode = mode;
st.variant = variant; st.variant = variant;
// The dealt state, kept so the peak metric can read the real
// starting Stress rather than trust a constant.
let started = st.clone();
let mut ps: Vec<Box<dyn Policy>> = (0..players) let mut ps: Vec<Box<dyn Policy>> = (0..players)
.map(|_| { .map(|_| {
if reactive { if reactive {
@ -138,6 +155,10 @@ fn sweep(mode: ScoringMode, variant: Variant, players: u8, reactive: bool) -> Ce
} }
}; };
c.games += 1; c.games += 1;
// A game that RAN is not a game that was PLAYED (CB-REV-0003 #2).
if g.state.outcome.is_some() && g.rounds == 5 {
c.played += 1;
}
if g.state.outcome.as_ref().is_some_and(|o| o.group_success) { if g.state.outcome.as_ref().is_some_and(|o| o.group_success) {
c.won += 1; c.won += 1;
} }
@ -150,15 +171,19 @@ fn sweep(mode: ScoringMode, variant: Variant, players: u8, reactive: bool) -> Ce
c.atk += 1; c.atk += 1;
} }
} }
// Peak Stress **held**, not peak Stress *assigned*. // Peak Stress **held**.
// //
// The first version took a maximum over `StressSet` PAYLOADS. // The original took a maximum over `StressSet` PAYLOADS, and the
// Starting Stress is 2 and is written by `setup`, never by an // starting value is written by `setup` and never by an event — so
// event, so a table that sat at 2 all game reported **1**, and a // a table that sat at 2 all game reported 1.
// table with no `StressSet` at all would report 0. CB-EV-0031 §2 //
// built its headline claim on that number. Wrong subject: the // **The first correction did not change the number it computes**
// metric answered "highest value ever assigned", the prose said // (CB-REV-0003 #5): `held`'s values are exactly
// "highest Stress reached". // `{start} {payloads}`, so with the floor applied
// `peak_held ≡ max(start, peak_payload)` for every possible input,
// and the comment claiming a change of subject was false of the
// new code too. The real fix is that `start` is now read off the
// dealt state instead of asserted by a constant.
// An arm is inert when NO `RoundEnded` follows it: the game // An arm is inert when NO `RoundEnded` follows it: the game
// ended in the same End step, so the sequence never advances a // ended in the same End step, so the sequence never advances a
// stage. `g.rounds >= 5` is a property of the GAME, not of the // stage. `g.rounds >= 5` is a property of the GAME, not of the
@ -169,9 +194,10 @@ fn sweep(mode: ScoringMode, variant: Variant, players: u8, reactive: bool) -> Ce
.events .events
.iter() .iter()
.rposition(|e| matches!(e, games_ground::GroundEvent::RoundEnded { .. })); .rposition(|e| matches!(e, games_ground::GroundEvent::RoundEnded { .. }));
let start = start_stress(&started);
let mut held: std::collections::BTreeMap<PlayerId, u8> = let mut held: std::collections::BTreeMap<PlayerId, u8> =
g.state.players.keys().map(|s| (*s, START_STRESS)).collect(); g.state.players.keys().map(|s| (*s, start)).collect();
c.peak_stress = c.peak_stress.max(START_STRESS); c.peak_stress = c.peak_stress.max(start);
for (i, e) in g.events.iter().enumerate() { for (i, e) in g.events.iter().enumerate() {
if matches!(e, games_ground::GroundEvent::DarvoTriggered { .. }) { if matches!(e, games_ground::GroundEvent::DarvoTriggered { .. }) {
c.darvo += 1; c.darvo += 1;
@ -190,7 +216,7 @@ fn sweep(mode: ScoringMode, variant: Variant, players: u8, reactive: bool) -> Ce
// START_STRESS floor made the two agree. // START_STRESS floor made the two agree.
c.peak_stress = c c.peak_stress = c
.peak_stress .peak_stress
.max(held.values().copied().max().unwrap_or(START_STRESS)); .max(held.values().copied().max().unwrap_or(start));
} }
} }
// **`games`, not `games + setup_fails`** (CB-REV-0002 #1). // **`games`, not `games + setup_fails`** (CB-REV-0002 #1).
@ -207,6 +233,11 @@ fn sweep(mode: ScoringMode, variant: Variant, players: u8, reactive: bool) -> Ce
the cell is short, so every number in it is over a sample nobody chose", the cell is short, so every number in it is over a sample nobody chose",
c.games, c.setup_fails c.games, c.setup_fails
); );
assert_eq!(
c.played, GAMES,
"{players}p {variant:?}: {} of {} games reached an outcome over five rounds",
c.played, c.games
);
if c.setup_fails > 0 { if c.setup_fails > 0 {
eprintln!( eprintln!(
" !! {players}p {variant:?}: {} setups refused", " !! {players}p {variant:?}: {} setups refused",

View file

@ -2427,6 +2427,108 @@ mod tests {
); );
} }
/// **Every tiebreak comparator, both modes** (CB-REV-0003 #4).
///
/// The strengthened oracle covered **1 of 6**: GR-E03's Stress
/// key. Reversing GR-E03's *Bond* key, or **either** of GR-E04's
/// two keys, left all 179 tests and all 26 scenarios green —
/// GR-E04's tiebreak had no coverage at all, because the only
/// GR-E04 scenario sets `group_success: false`, so the winners
/// branch returns empty and the comparator never runs.
///
/// Each case is built so exactly one key decides, and the
/// expected winner is named — asserting "the set changed" is what
/// let a reversed comparator pass in the first place.
#[test]
fn every_tiebreak_key_decides_the_way_the_rules_say() {
let board = |mode: ScoringMode| {
let mut s = setup(4, Variant::Baseline, 5);
s.mode = mode;
let seats: Vec<PlayerId> = s.players.keys().copied().collect();
let ids: Vec<u32> = s.problems.keys().copied().collect();
for id in &ids {
s.problems.remove(id);
}
(s, seats)
};
let mk = |value: u8, claimed_by: Option<PlayerId>| ProblemState {
suit: Suit::Repair,
value,
face_up: true,
denied: false,
claimed_by,
protected_this_round: false,
};
// GR-E03 key 2, Stress: equal claims, lower Stress wins.
let (mut s, seats) = board(ScoringMode::CommonProblem);
s.problems.insert(1, mk(4, Some(seats[0])));
s.problems.insert(2, mk(4, Some(seats[1])));
s.players.get_mut(&seats[0]).expect("a").stress = 5;
s.players.get_mut(&seats[1]).expect("b").stress = 1;
assert_eq!(
s.score().winners,
vec![seats[1]],
"GR-E03: lower Stress must win the tiebreak"
);
// GR-E03 key 3, Bonds: equal claims AND equal Stress, more
// Bonds wins.
s.players.get_mut(&seats[0]).expect("a").stress = 1;
s.relations
.insert(Pair::new(seats[0], seats[2]), Relation::Bond);
assert_eq!(
s.score().winners,
vec![seats[0]],
"GR-E03: with claims and Stress level, more Bonds must win"
);
// GR-E04 key 2, combined Stress: two coalitions of equal
// score, the calmer one wins.
let (mut s, seats) = board(ScoringMode::BondedCoalitions);
s.problems.insert(1, mk(4, Some(seats[0])));
s.problems.insert(2, mk(4, Some(seats[2])));
s.relations
.insert(Pair::new(seats[0], seats[1]), Relation::Bond);
s.relations
.insert(Pair::new(seats[2], seats[3]), Relation::Bond);
for seat in &seats {
s.players.get_mut(seat).expect("p").stress = 1;
}
s.players.get_mut(&seats[0]).expect("p").stress = 5;
let won = s.score().winners;
assert!(
won.contains(&seats[2]) && !won.contains(&seats[0]),
"GR-E04: the coalition with lower combined Stress must win, got {won:?}"
);
// GR-E04 key 3, Blame — and reaching it takes care, which is
// the point. A coalition's score is `sum(claimed - blame)`, so
// a Blame token lowers the score too and key 1 decides first.
// The key is only reachable when the claims COMPENSATE: 5
// claimed with one Blame ties 4 claimed with none.
s.players.get_mut(&seats[0]).expect("p").stress = 1;
s.problems.insert(1, mk(5, Some(seats[0])));
s.problems.insert(2, mk(4, Some(seats[2])));
s.players.get_mut(&seats[0]).expect("p").blame_from = vec![seats[3]];
let scored = s.score();
assert_eq!(
scored.coalitions.len(),
2,
"the fixture needs two coalitions to compare"
);
assert_eq!(
scored.coalitions[0].score, scored.coalitions[1].score,
"the Blame key is unreachable unless the scores tie: {:?}",
scored.coalitions
);
let won = scored.winners;
assert!(
won.contains(&seats[2]) && !won.contains(&seats[0]),
"GR-E04: with score and Stress level, fewer Blame must win, got {won:?}"
);
}
/// **`rules_delta.yaml`'s `unchanged:` list is ground-game's claim /// **`rules_delta.yaml`'s `unchanged:` list is ground-game's claim
/// about their own experiment, and it is checkable.** /// about their own experiment, and it is checkable.**
/// ///

View file

@ -137,7 +137,7 @@ retire_if = "two passes run with no finding while artifacts keep growing — tha
id = "CHAOS" id = "CHAOS"
name = "the chaos roll" name = "the chaos roll"
target = "" target = ""
cadence = "none" cadence = "manual"
checks = "d8 on each tier declaration, 12-declaration calibration window (window 2, opened 2026-08-03; window 1 ran at d4)" checks = "d8 on each tier declaration, 12-declaration calibration window (window 2, opened 2026-08-03; window 1 ran at d4)"
added = "2026-07-30" added = "2026-07-30"
review_by = "2026-11-30" review_by = "2026-11-30"
@ -178,7 +178,7 @@ id = "CB-REV-0002/7"
name = "variant panels" name = "variant panels"
target = "panels" target = "panels"
cadence = "all" cadence = "all"
checks = "the H1 measurement harnesses actually run, and a short cell fails rather than printing a number a reader must notice" checks = "both H1 measurement harnesses run; a short cell fails, and so does a cell whose games did not reach an outcome"
notes = """ notes = """
Registered because the second adversarial review asked what the harness Registered because the second adversarial review asked what the harness
would report if the work silently stopped, and the answer was "green, and would report if the work silently stopped, and the answer was "green, and
@ -186,4 +186,12 @@ nothing else": `attack-value` and `regulation` were wired into no target,
so every figure in CB-EV-0030 and CB-EV-0031 came from a manual run of an so every figure in CB-EV-0030 and CB-EV-0031 came from a manual run of an
ungated binary -- including the assertion added to catch short cells, ungated binary -- including the assertion added to catch short cells,
which was unreachable from `make`. which was unreachable from `make`.
CB-REV-0003 #1 then found this `checks` line claimed a property that held
for only one of the two harnesses: `attack-value`, which produced every
number in CB-EV-0030's table, still only warned. And #2 found that
counting games proves they STARTED: stopping the engine after one round
gave 200 games, all-zero columns and exit 0 -- the signature CB-EV-0030
says the instrumentation distinguishes from a result. Both now assert an
outcome was reached.
""" """

127
reviews/CB-REV-0003-h1.md Normal file
View file

@ -0,0 +1,127 @@
# CB-REV-0003 — round 3
Fresh agent again. Run 2026-08-08 against the corrections made after
[`CB-REV-0002`](CB-REV-0002-h1.md).
> **Verdict: not approvable. Four FATAL, five SERIOUS, two MINOR — and
> three of the four FATAL were created or left by round 2's corrections.**
>
> **The pattern is now established over three rounds** and is the most
> useful thing this review has produced.
| round | challenges | FATAL | of which introduced by the previous round's fix |
|---|---:|---:|---:|
| 1 | 13 | 5 | — |
| 2 | 8 | 3 | 2 |
| 3 | 11 | 4 | 3 |
---
## The four fatal ones
### 1. The fix for round 2's #1 was applied to one of the two harnesses
`regulation.rs` got `assert_eq!(games, GAMES)`. **`attack-value.rs` — which
produced every number in CB-EV-0030's DARVO table — kept `if games != 200
{ eprintln!(..) }`.** Injected setup refusals gave exit 0 and a full table
over 195-game columns.
**And the gate registered to close the finding asserted the property for
both.** Its `checks` line was false of half of what it gates.
**Conceded.** Both harnesses assert now; the `checks` line says what is
actually checked.
### 2. Counting games proves they STARTED, not that they were PLAYED
Stopping the engine after one round (`round >= 5``>= 1`) gives **200
games in every cell, every value zero, exit 0, `make panels` green** —
byte for byte the signature CB-EV-0030 §3 claims the instrumentation
distinguishes from a real result.
**Conceded, and it is the sharpest finding of the three rounds.** The
counter answered *"did `play` return `Ok`"* while the claim made of it was
*"the zeros are real"*. Both harnesses now require every counted game to
have reached an outcome over five rounds; the one-round mutation fails
with the right message.
### 3. Round 2's `.csv` filter disabled the checks it was added beside
The filter was applied to **all three** loops — digest comparison, the CSV
parser check, and upstream freshness. So `catalog.yaml` and
`rules_delta.yaml`, whose missing digests were round 1's finding, were
recorded and then **never compared**, and never checked against upstream
at all. Tampering with both produced `[ok] × 12, exit 0`.
**Conceded.** Only the parser loop filters; tampering with either YAML now
fails on digest *and* freshness.
### 4. Five of six tiebreak comparators had no coverage
The strengthened oracle covered GR-E03's Stress key. **GR-E03's Bond key
and both of GR-E04's keys could be reversed with 179 tests and 26
scenarios green** — GR-E04's tiebreak never executes in any scenario,
because the only GR-E04 scenario sets `group_success: false`.
**Conceded.** All four are now covered and each was verified red under
mutation. **The Blame key took care to reach**: a coalition's score is
`sum(claimed blame)`, so a Blame token lowers the score and key 1
decides first — the key is only reachable when claims compensate, and the
test asserts the scores tie before relying on it.
## The serious ones
| # | challenge | outcome |
|---|---|---|
| 5 | "peak Stress **held**, not assigned" — `peak_held ≡ max(start, peak_payload)` for every possible input, so the correction computed the same number | **conceded.** The fix was cosmetic and the comment was false of the new code too. The real coupling was missing, which is #6 |
| 6 | `START_STRESS` was an unchecked constant — setting it to 7 printed `peak 7` everywhere | **conceded.** Now read off the dealt state. Verified: changing setup's Stress to 3 moves the reported peak to 3 |
| 7 | `cadence = "none"` was a pure loophole — a real target that runs nowhere and passes | **conceded, removed.** It had one user, the empty-target CHAOS entry, which the loop already skips. `manual` remains self-declared and unchecked, and that is now stated rather than implied |
| 8 | sibling discovery swapped a hand-written list for three hand-written globs; `metadata.json` and `VARIANT.md` were invisible, and both are named in the package's own `changed_files` | **conceded.** Walked, not globbed. **It found both immediately**, plus nested files the globs could never reach |
| 9 | "~72,000 games" is not producible; `make panels` runs **17,600** | **conceded.** Adopted from a reviewer's message and never re-derived — in the file whose correction history is about exactly that. The conclusion (zero divergence) reproduces |
## The minor ones
- **10 — two kinds of number were presented alike.** `600/800/1200` is
exactly `seats × games` on every sample; **`363` and `29` vary** —
7.1%11.5% inert across four windows. "92% of arms are live" is a fact
about seeds `0..200`. Now separated.
- **11 — the 29/363 citation, wrong twice, is now stated as one sample**
of a varying quantity. Both "independent confirmations" used the same
seeds, so what was confirmed was the *definition*. Whether round 1's
`StrictReactive` was the corrected policy cannot be determined: **that
reviewer's code was never kept**, which is itself worth fixing.
## What held, after being attacked
**Every one of the 24 cells in CB-EV-0030's DARVO table reproduces
exactly** — the first of three attempts at that table to produce no wrong
number. `regulation`'s `assert_eq!(c.games, GAMES)` catches both the setup
and the `play` path. The #13 fix is genuinely controlled. H1-B's mutation
coverage is real. `make all` runs `panels` and fails on it. The
1-arm-per-seat-per-game invariant holds on every sample — **and is still
unexplained.**
## What could not be checked
Upstream freshness (`../ground-game` not checked out, so that half of
`edition-check` has never run here); whether `metadata.json` and
`VARIANT.md` are load-bearing to anything; round 1's `StrictReactive`;
felt-play; cost.
---
## The finding that outlasts H1
Three rounds, each correcting the last, each introducing defects of the
same class. **The corrections are not getting safer.**
> A correction is written under the belief that the error is now
> understood. That belief is the condition under which this class of
> error is produced — so the correction inherits it, and the next round
> finds the same shape one level in.
**Round 4 is owed by the same argument.** The honest conclusion is not
that the work is nearly right; it is that **author-made corrections to
measurement work should be assumed defective until a fresh reader has
attacked them**, and this project should stop treating "corrected" as a
state closer to done than "found wrong".

View file

@ -66,18 +66,27 @@ SIBLINGS = "../"
# raised FileNotFoundError from `digest` instead of failing with the # raised FileNotFoundError from `digest` instead of failing with the
# designed message, and "the sibling packages are covered" counted lines # designed message, and "the sibling packages are covered" counted lines
# in a Markdown file -- it passed with all three files deleted. # in a Markdown file -- it passed with all three files deleted.
SIBLING_GLOBS = ("catalog.yaml", "experiments/*/rules_delta.yaml", "experiments/*/*.csv") # Everything under `editions/` that is not the edition directory itself.
# **Walked, not globbed** (CB-REV-0003 #8): the first version listed three
# hand-written glob patterns, which is the same self-certifying shape as
# reading the list out of PROVENANCE -- one hand-written list swapped for
# another. It missed `metadata.json` and `VARIANT.md`, both named in the
# package's OWN `changed_files` manifest, and anything a directory deeper.
SIBLING_SKIP = {".DS_Store"}
def sibling_files(): def sibling_files():
"""Sibling packages that are on disk, as `../`-relative paths.""" """Every file beside the edition, as `../`-relative paths."""
import glob as _glob base = os.path.join(ROOT, "editions")
edition_dir = os.path.join(ROOT, EDITION)
root = os.path.join(ROOT, EDITION, SIBLINGS)
out = [] out = []
for pattern in SIBLING_GLOBS: for dirpath, _dirs, files in os.walk(base):
for hit in _glob.glob(os.path.join(root, pattern)): if os.path.abspath(dirpath).startswith(os.path.abspath(edition_dir)):
rel = os.path.relpath(hit, os.path.join(ROOT, EDITION)) continue
for name in files:
if name in SIBLING_SKIP:
continue
rel = os.path.relpath(os.path.join(dirpath, name), edition_dir)
out.append(rel.replace(os.sep, "/")) out.append(rel.replace(os.sep, "/"))
return sorted(out) return sorted(out)
@ -98,7 +107,7 @@ def check():
print(f" [FAIL] a digest is recorded for a file that is not here: {', '.join(missing)}") print(f" [FAIL] a digest is recorded for a file that is not here: {', '.join(missing)}")
rc = 1 rc = 1
for name in [f for f in present if f.endswith(".csv")]: for name in present:
if name not in want: if name not in want:
continue continue
path = os.path.join(ROOT, EDITION, name) path = os.path.join(ROOT, EDITION, name)
@ -118,6 +127,17 @@ def check():
# ADR-0015 D3's falsifier, checked rather than asserted: the hand # ADR-0015 D3's falsifier, checked rather than asserted: the hand
# reader handles commas inside quotes and NOTHING ELSE. A doubled # reader handles commas inside quotes and NOTHING ELSE. A doubled
# quote or an embedded newline means `csv` is the answer after all. # quote or an embedded newline means `csv` is the answer after all.
# ONLY this loop filters: it is ADR-0015 D3's falsifier about the
# hand-rolled CSV reader, and running it over YAML made a valid
# `catalog.yaml` line with an odd quote count fail with a message
# about a parser that never reads it (CB-REV-0002 #9).
#
# **The digest and freshness loops must NOT filter** — CB-REV-0003 #3:
# the round-2 correction applied this filter to all three, so the two
# sibling YAMLs whose missing digests were round 1's finding were
# recorded and then never compared, and never checked against
# upstream at all. `make edition-check` answered its own headline
# question with [ok] when the answer was no.
for name in [f for f in present if f.endswith(".csv")]: for name in [f for f in present if f.endswith(".csv")]:
raw = open(os.path.join(ROOT, EDITION, name), encoding="utf-8-sig").read() raw = open(os.path.join(ROOT, EDITION, name), encoding="utf-8-sig").read()
if '""' in raw: if '""' in raw:
@ -138,7 +158,7 @@ def check():
print(" [----] upstream not checked out — freshness UNVERIFIED") print(" [----] upstream not checked out — freshness UNVERIFIED")
print(f" expected {UPSTREAM_DIR}") print(f" expected {UPSTREAM_DIR}")
return rc return rc
for name in [f for f in present if f.endswith(".csv")]: for name in present:
up = os.path.join(UPSTREAM_DIR, name) up = os.path.join(UPSTREAM_DIR, name)
if not os.path.exists(up): if not os.path.exists(up):
print(f" [FAIL] {name} is not in upstream — where did it come from?") print(f" [FAIL] {name} is not in upstream — where did it come from?")

View file

@ -316,11 +316,13 @@ def check_gate_registry(root=REPO):
if not target: if not target:
continue continue
cadence = g.get("cadence") cadence = g.get("cadence")
if cadence not in ("all", "manual", "none"): if cadence not in ("all", "manual"):
out.append(Finding( out.append(Finding(
"gates", "gates.toml", "gates", "gates.toml",
f"{target!r} declares no cadence — say whether `make all` runs " f"{target!r} declares no cadence — say `all` or `manual`. "
f"it, or a gate can exist without ever running")) f"`none` was a pure loophole: a real target that runs "
f"nowhere and passes, which is the condition this rule "
f"exists to prevent (CB-REV-0003 #7)"))
elif cadence == "all" and target not in deps: elif cadence == "all" and target not in deps:
out.append(Finding( out.append(Finding(
"gates", "gates.toml", "gates", "gates.toml",