CB-REV-0003: round 3, and three of four FATAL came from round 2's fixes
Some checks failed
ci / check (push) Failing after 3s

The pattern is now measured over three rounds: 5 fatal, then 3 (2 from the
previous round's corrections), then 4 (3 from them). The corrections are
not getting safer.

FATAL 1: round 2's short-cell assertion went into regulation.rs only.
attack-value.rs — which produced every number in CB-EV-0030's DARVO table
— still just warned, and the gate registered to close the finding claimed
the property for both.

FATAL 2, the sharpest of the three rounds: counting games proves they
STARTED. Stopping the engine after one round gives 200 games, all-zero
columns and exit 0 — byte for byte the signature CB-EV-0030 says the
instrumentation distinguishes from a real result. Both harnesses now
require every counted game to have reached an outcome over five rounds.

FATAL 3: round 2's `.csv` filter was applied to all three loops, so
catalog.yaml and rules_delta.yaml — whose missing digests were round 1's
finding — were recorded and then never compared, and never checked against
upstream at all. Only the parser loop filters now.

FATAL 4: five of six tiebreak comparators had no coverage. GR-E04's
tiebreak never executes in any scenario. All four are now covered and
mutation-verified; the Blame key needed compensating claims to be
reachable at all, since Blame also lowers the coalition score.

SERIOUS: "peak held" computed the same number as "peak assigned" for every
possible input — the real gap was that START_STRESS was an unchecked
constant, now read off the dealt state; cadence="none" was a pure
loophole, removed; sibling discovery swapped a hand-written list for
hand-written globs and missed metadata.json and VARIANT.md, both named in
the package's own changed_files — now walked, and it found them
immediately; and "~72,000 games" was unsourced, make panels runs 17,600.

Also separated two kinds of number that were presented alike: seats×games
is invariant, 363 and 29 vary 7.1%-11.5% across samples.

Round 4 owed. The conclusion is not that the work is nearly right — it is
that author-made corrections to measurement work should be assumed
defective until a fresh reader has attacked them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
tegwick 2026-08-08 10:26:25 +02:00
parent feb68027f9
commit c4a8a227c0
9 changed files with 386 additions and 57 deletions

View file

@ -62,6 +62,8 @@ sha256 eb21fa3237637790fef601fe6668a190a549715e9a78b7dfe47b40cd069b648e ../cat
sha256 f58e81f84ea2b0d16e39932261eb3f3d9890345cdf37ad6f0b3abc00636840be ../experiments/h1-problem-stress/rules_delta.yaml
sha256 7b1cc0149122b855e827bc930576ed165bf7dd8d62707e845a9e514ce3521f8e ../experiments/h1-problem-stress/Actions.csv
sha256 62785f5e7e245c60171624d15de2f40187a44fec54f93c7d9705cf52584b1078 ../experiments/h1-problem-stress/Rules_Text.csv
sha256 002e3adf2d038579f2803975d8c6344e2b77ee9da9eaaa1b52c8bd1666bb5201 ../experiments/h1-problem-stress/VARIANT.md
sha256 443199db94601cc889557e5e86823f374dbfdf875962fc84f95c01865605101c ../experiments/h1-problem-stress/metadata.json
```
**`rules_delta.yaml` is the load-bearing one**: it is the executable

View file

@ -143,12 +143,18 @@ takes the **clamp** at 5 to collapse a gap and change who wins.
**Its impact was claimed and never measured, and the measurement is
zero** (round 2, #5). This file said the pressure was invisible to *"the
two modes CB-EV-0030 reports on"* without checking. The reviewer ran
~72,000 games across all three modes with a divergence detector and found
**no divergence at all**; `attack-value`'s output is byte-identical with
the fix and without it. **The defect is real in principle and its only
witness is a hand-constructed board.** It is worth fixing and it changed
nothing that has been reported.
two modes CB-EV-0030 reports on"* without checking. Reverting the fix
leaves `attack-value`'s output **byte-identical across all 24 cells and
all three modes**, including the two whose `won` column reads through the
tiebreak. **The defect is real in principle and its only witness is a
hand-constructed board.**
**The sample size quoted here was "~72,000 games" and was not
derivable from anything in this repo** (round 3, #9). `make panels` runs
**17,600**: `attack-value` is 2 variants × 3 modes × 3 policies × 4 seat
counts × 200 = 14,400, and `regulation` is 2 × 4 × 2 × 200 = 3,200. The
figure was adopted from a reviewer's message and never re-derived — in
the file whose whole correction history is about exactly that.
**Inert arms, and a second wrong-subject error caught on the way.** A
DARVO arm at the End of Round 5 can never advance a stage — `GameEnded`
@ -162,19 +168,32 @@ follows immediately — and ground-game's criterion 1 is about DARVO
| 4p | 800 | 0 |
| 6p | 1200 | 0 |
**Two kinds of number are in that table and this file did not distinguish
them** (round 3, #10). `600 / 800 / 1200` is **exactly `seats × games`**
and holds on every sample tried. `363` and `29` are **sample-specific**:
across four disjoint windows the 2p arms run 354366 and the inert share
runs **7.1%11.5%**. So "92% of arms at 2p are live" is a fact about seeds
`0..200`, not about the game — it reads 88.5% elsewhere.
The same applies to the baseline win counts quoted throughout
(`132/165/190/200`): the *identity* of greedy and reactive under baseline
holds on every sample, the **digits do not**.
**The first version of that metric was wrong and equalled `darvo` in
every cell**, because it tested `g.rounds >= 5` — a property of the
*game*, not of the *event*. Every arm in every completed game was marked
inert, which briefly looked like "criterion 1 is not met after all". An
arm is inert when no `RoundEnded` follows it.
**29-of-363 was independently confirmed twice**, and the citation for it
was wrong the first time (round 2, #6): this file credited
`CB-REV-0001`, which does not contain the figure. Round 1's reviewer *did*
report it — 29 of 363 at 2p, 8% — but only in its message, and it was
never transcribed into the review file. **The record is now corrected
there.** Round 2 re-derived it independently, by a different definition
(no later DARVO event names that player), and agrees in every cell.
**The 29-of-363 citation has now been wrong twice and is stated plainly
here.** Round 2 found it credited to `CB-REV-0001`, which did not contain
the figure; round 1's reviewer *had* reported it, in its message only.
Round 3 found the remaining problem (#10, #11): both "independent
confirmations" used **seeds 0..200**, so what was confirmed twice is the
**definition**, not the figure — and whether round 1's `StrictReactive`
was the corrected one-arm policy or the five-difference one **cannot be
determined from anything in this repo**, because that reviewer's code was
never kept. Treat 29/363 as one sample of a quantity that varies.
**Criterion 1 stands as met**: 92% of arms at 2p and all of them above are
live.

View file

@ -24,6 +24,10 @@ use cb_kernel::PlayerId;
use games_ground::bot::{play, Choice, GreedyPolicy, Policy};
use games_ground::{Action, GroundCommand, GroundState, ScoringMode};
/// Games per cell. Named once so the footer and the assertion cannot
/// disagree — the footer said "200 games per cell" over 195-game columns.
const GAMES: u32 = 200;
/// Greedy's ordering with ATTACK's rank as a parameter.
///
/// A copy of the ranking rather than a call into it: `GreedyPolicy::rank`
@ -98,9 +102,10 @@ fn sweep(
mk: &dyn Fn(u8) -> Vec<Box<dyn Policy>>,
) -> (u32, u32, u32, u32) {
let (mut games, mut won, mut atk, mut darvo) = (0, 0, 0, 0);
let mut played = 0u32;
let mut errs: Vec<String> = Vec::new();
let mut setup_fails = 0u32;
for seed in 0..200u64 {
for seed in 0..GAMES as u64 {
let Ok(mut st) = GroundState::setup(
&Setup {
players,
@ -130,6 +135,9 @@ fn sweep(
}
};
games += 1;
if g.state.outcome.is_some() && g.rounds == 5 {
played += 1;
}
// In the co-op mode the table wins together. In the other two the
// question is whether SEAT 0 is among the winners — because that
@ -161,20 +169,30 @@ fn sweep(
}
}
}
if setup_fails > 0 {
eprintln!(" !! {setup_fails} of 200 SETUPS failed ({players}p)");
}
if games != 200 {
eprintln!(" !! only {games} of 200 games ran ({players}p, {mode:?})");
}
if !errs.is_empty() {
eprintln!(
" !! {} of 200 games did not run ({}p): first = {}",
errs.len(),
players,
errs[0]
);
}
// **Asserted, not printed** (CB-REV-0003 #1). This harness produced
// every number in CB-EV-0030's DARVO table, and it still only warned
// on a short cell — the very defect fixed in `regulation.rs` and left
// here, while the gate registered to close it claimed the property
// for both.
assert_eq!(
games,
GAMES,
"{players}p {mode:?} {variant:?}: only {games} of {GAMES} games ran \
({setup_fails} setups refused, {} play errors) the cell is short, so \
every number in it is over a sample nobody chose",
errs.len()
);
// **And that a game RAN is not that it was PLAYED** (CB-REV-0003 #2).
// Stopping the engine after one round gave 200 games, all-zero
// columns and exit 0 — byte for byte the signature CB-EV-0030 §3
// claims to distinguish from a real result. A counted game must have
// reached an outcome.
assert_eq!(
played, GAMES,
"{players}p {mode:?} {variant:?}: {played} of {games} games reached an \
outcome the rest stopped early, and all-zero columns would read as \
a result rather than as nothing having happened"
);
(games, won, atk, darvo)
}

View file

@ -71,9 +71,15 @@ impl Policy for Reactive {
}
}
/// `Scenarios.csv`: "All players start at Stress 2." Named rather than
/// inlined so the peak metric cannot silently disagree with setup.
const START_STRESS: u8 = 2;
/// The starting Stress, **read off the dealt state** rather than declared.
///
/// It was a `const 2` "named so the peak metric cannot silently disagree
/// with setup" — and nothing compared it to setup, so raising it to 7
/// printed `peak 7` in every cell with no warning (CB-REV-0003 #6).
/// Naming a constant is not checking it.
fn start_stress(state: &GroundState) -> u8 {
state.players.values().map(|p| p.stress).max().unwrap_or(0)
}
/// Games per cell. Named once so the assertion and the banner cannot
/// disagree — the banner said "200 games per cell" over 196-game columns.
@ -94,6 +100,13 @@ struct Cell {
/// arm is real and can do nothing. Reported separately, because
/// ground-game's criterion 1 is asking about DARVO *mattering*.
inert_arms: u32,
/// Games that reached an outcome over the full five rounds.
///
/// **`games` counts `play` returning `Ok`, which is not the claim.**
/// Stopping the engine after one round produced 200 games, all-zero
/// columns and exit 0 — the exact signature CB-EV-0030 §3 says the
/// instrumentation distinguishes from a real result.
played: u32,
}
fn sweep(mode: ScoringMode, variant: Variant, players: u8, reactive: bool) -> Cell {
@ -105,6 +118,7 @@ fn sweep(mode: ScoringMode, variant: Variant, players: u8, reactive: bool) -> Ce
peak_stress: 0,
setup_fails: 0,
inert_arms: 0,
played: 0,
};
for seed in 0..GAMES as u64 {
let Ok(mut st) = GroundState::setup(
@ -120,6 +134,9 @@ fn sweep(mode: ScoringMode, variant: Variant, players: u8, reactive: bool) -> Ce
};
st.mode = mode;
st.variant = variant;
// The dealt state, kept so the peak metric can read the real
// starting Stress rather than trust a constant.
let started = st.clone();
let mut ps: Vec<Box<dyn Policy>> = (0..players)
.map(|_| {
if reactive {
@ -138,6 +155,10 @@ fn sweep(mode: ScoringMode, variant: Variant, players: u8, reactive: bool) -> Ce
}
};
c.games += 1;
// A game that RAN is not a game that was PLAYED (CB-REV-0003 #2).
if g.state.outcome.is_some() && g.rounds == 5 {
c.played += 1;
}
if g.state.outcome.as_ref().is_some_and(|o| o.group_success) {
c.won += 1;
}
@ -150,15 +171,19 @@ fn sweep(mode: ScoringMode, variant: Variant, players: u8, reactive: bool) -> Ce
c.atk += 1;
}
}
// Peak Stress **held**, not peak Stress *assigned*.
// Peak Stress **held**.
//
// The first version took a maximum over `StressSet` PAYLOADS.
// Starting Stress is 2 and is written by `setup`, never by an
// event, so a table that sat at 2 all game reported **1**, and a
// table with no `StressSet` at all would report 0. CB-EV-0031 §2
// built its headline claim on that number. Wrong subject: the
// metric answered "highest value ever assigned", the prose said
// "highest Stress reached".
// The original took a maximum over `StressSet` PAYLOADS, and the
// starting value is written by `setup` and never by an event — so
// a table that sat at 2 all game reported 1.
//
// **The first correction did not change the number it computes**
// (CB-REV-0003 #5): `held`'s values are exactly
// `{start} {payloads}`, so with the floor applied
// `peak_held ≡ max(start, peak_payload)` for every possible input,
// and the comment claiming a change of subject was false of the
// new code too. The real fix is that `start` is now read off the
// dealt state instead of asserted by a constant.
// An arm is inert when NO `RoundEnded` follows it: the game
// ended in the same End step, so the sequence never advances a
// stage. `g.rounds >= 5` is a property of the GAME, not of the
@ -169,9 +194,10 @@ fn sweep(mode: ScoringMode, variant: Variant, players: u8, reactive: bool) -> Ce
.events
.iter()
.rposition(|e| matches!(e, games_ground::GroundEvent::RoundEnded { .. }));
let start = start_stress(&started);
let mut held: std::collections::BTreeMap<PlayerId, u8> =
g.state.players.keys().map(|s| (*s, START_STRESS)).collect();
c.peak_stress = c.peak_stress.max(START_STRESS);
g.state.players.keys().map(|s| (*s, start)).collect();
c.peak_stress = c.peak_stress.max(start);
for (i, e) in g.events.iter().enumerate() {
if matches!(e, games_ground::GroundEvent::DarvoTriggered { .. }) {
c.darvo += 1;
@ -190,7 +216,7 @@ fn sweep(mode: ScoringMode, variant: Variant, players: u8, reactive: bool) -> Ce
// START_STRESS floor made the two agree.
c.peak_stress = c
.peak_stress
.max(held.values().copied().max().unwrap_or(START_STRESS));
.max(held.values().copied().max().unwrap_or(start));
}
}
// **`games`, not `games + setup_fails`** (CB-REV-0002 #1).
@ -207,6 +233,11 @@ fn sweep(mode: ScoringMode, variant: Variant, players: u8, reactive: bool) -> Ce
the cell is short, so every number in it is over a sample nobody chose",
c.games, c.setup_fails
);
assert_eq!(
c.played, GAMES,
"{players}p {variant:?}: {} of {} games reached an outcome over five rounds",
c.played, c.games
);
if c.setup_fails > 0 {
eprintln!(
" !! {players}p {variant:?}: {} setups refused",

View file

@ -2427,6 +2427,108 @@ mod tests {
);
}
/// **Every tiebreak comparator, both modes** (CB-REV-0003 #4).
///
/// The strengthened oracle covered **1 of 6**: GR-E03's Stress
/// key. Reversing GR-E03's *Bond* key, or **either** of GR-E04's
/// two keys, left all 179 tests and all 26 scenarios green —
/// GR-E04's tiebreak had no coverage at all, because the only
/// GR-E04 scenario sets `group_success: false`, so the winners
/// branch returns empty and the comparator never runs.
///
/// Each case is built so exactly one key decides, and the
/// expected winner is named — asserting "the set changed" is what
/// let a reversed comparator pass in the first place.
#[test]
fn every_tiebreak_key_decides_the_way_the_rules_say() {
let board = |mode: ScoringMode| {
let mut s = setup(4, Variant::Baseline, 5);
s.mode = mode;
let seats: Vec<PlayerId> = s.players.keys().copied().collect();
let ids: Vec<u32> = s.problems.keys().copied().collect();
for id in &ids {
s.problems.remove(id);
}
(s, seats)
};
let mk = |value: u8, claimed_by: Option<PlayerId>| ProblemState {
suit: Suit::Repair,
value,
face_up: true,
denied: false,
claimed_by,
protected_this_round: false,
};
// GR-E03 key 2, Stress: equal claims, lower Stress wins.
let (mut s, seats) = board(ScoringMode::CommonProblem);
s.problems.insert(1, mk(4, Some(seats[0])));
s.problems.insert(2, mk(4, Some(seats[1])));
s.players.get_mut(&seats[0]).expect("a").stress = 5;
s.players.get_mut(&seats[1]).expect("b").stress = 1;
assert_eq!(
s.score().winners,
vec![seats[1]],
"GR-E03: lower Stress must win the tiebreak"
);
// GR-E03 key 3, Bonds: equal claims AND equal Stress, more
// Bonds wins.
s.players.get_mut(&seats[0]).expect("a").stress = 1;
s.relations
.insert(Pair::new(seats[0], seats[2]), Relation::Bond);
assert_eq!(
s.score().winners,
vec![seats[0]],
"GR-E03: with claims and Stress level, more Bonds must win"
);
// GR-E04 key 2, combined Stress: two coalitions of equal
// score, the calmer one wins.
let (mut s, seats) = board(ScoringMode::BondedCoalitions);
s.problems.insert(1, mk(4, Some(seats[0])));
s.problems.insert(2, mk(4, Some(seats[2])));
s.relations
.insert(Pair::new(seats[0], seats[1]), Relation::Bond);
s.relations
.insert(Pair::new(seats[2], seats[3]), Relation::Bond);
for seat in &seats {
s.players.get_mut(seat).expect("p").stress = 1;
}
s.players.get_mut(&seats[0]).expect("p").stress = 5;
let won = s.score().winners;
assert!(
won.contains(&seats[2]) && !won.contains(&seats[0]),
"GR-E04: the coalition with lower combined Stress must win, got {won:?}"
);
// GR-E04 key 3, Blame — and reaching it takes care, which is
// the point. A coalition's score is `sum(claimed - blame)`, so
// a Blame token lowers the score too and key 1 decides first.
// The key is only reachable when the claims COMPENSATE: 5
// claimed with one Blame ties 4 claimed with none.
s.players.get_mut(&seats[0]).expect("p").stress = 1;
s.problems.insert(1, mk(5, Some(seats[0])));
s.problems.insert(2, mk(4, Some(seats[2])));
s.players.get_mut(&seats[0]).expect("p").blame_from = vec![seats[3]];
let scored = s.score();
assert_eq!(
scored.coalitions.len(),
2,
"the fixture needs two coalitions to compare"
);
assert_eq!(
scored.coalitions[0].score, scored.coalitions[1].score,
"the Blame key is unreachable unless the scores tie: {:?}",
scored.coalitions
);
let won = scored.winners;
assert!(
won.contains(&seats[2]) && !won.contains(&seats[0]),
"GR-E04: with score and Stress level, fewer Blame must win, got {won:?}"
);
}
/// **`rules_delta.yaml`'s `unchanged:` list is ground-game's claim
/// about their own experiment, and it is checkable.**
///

View file

@ -137,7 +137,7 @@ retire_if = "two passes run with no finding while artifacts keep growing — tha
id = "CHAOS"
name = "the chaos roll"
target = ""
cadence = "none"
cadence = "manual"
checks = "d8 on each tier declaration, 12-declaration calibration window (window 2, opened 2026-08-03; window 1 ran at d4)"
added = "2026-07-30"
review_by = "2026-11-30"
@ -178,7 +178,7 @@ id = "CB-REV-0002/7"
name = "variant panels"
target = "panels"
cadence = "all"
checks = "the H1 measurement harnesses actually run, and a short cell fails rather than printing a number a reader must notice"
checks = "both H1 measurement harnesses run; a short cell fails, and so does a cell whose games did not reach an outcome"
notes = """
Registered because the second adversarial review asked what the harness
would report if the work silently stopped, and the answer was "green, and
@ -186,4 +186,12 @@ nothing else": `attack-value` and `regulation` were wired into no target,
so every figure in CB-EV-0030 and CB-EV-0031 came from a manual run of an
ungated binary -- including the assertion added to catch short cells,
which was unreachable from `make`.
CB-REV-0003 #1 then found this `checks` line claimed a property that held
for only one of the two harnesses: `attack-value`, which produced every
number in CB-EV-0030's table, still only warned. And #2 found that
counting games proves they STARTED: stopping the engine after one round
gave 200 games, all-zero columns and exit 0 -- the signature CB-EV-0030
says the instrumentation distinguishes from a result. Both now assert an
outcome was reached.
"""

127
reviews/CB-REV-0003-h1.md Normal file
View file

@ -0,0 +1,127 @@
# CB-REV-0003 — round 3
Fresh agent again. Run 2026-08-08 against the corrections made after
[`CB-REV-0002`](CB-REV-0002-h1.md).
> **Verdict: not approvable. Four FATAL, five SERIOUS, two MINOR — and
> three of the four FATAL were created or left by round 2's corrections.**
>
> **The pattern is now established over three rounds** and is the most
> useful thing this review has produced.
| round | challenges | FATAL | of which introduced by the previous round's fix |
|---|---:|---:|---:|
| 1 | 13 | 5 | — |
| 2 | 8 | 3 | 2 |
| 3 | 11 | 4 | 3 |
---
## The four fatal ones
### 1. The fix for round 2's #1 was applied to one of the two harnesses
`regulation.rs` got `assert_eq!(games, GAMES)`. **`attack-value.rs` — which
produced every number in CB-EV-0030's DARVO table — kept `if games != 200
{ eprintln!(..) }`.** Injected setup refusals gave exit 0 and a full table
over 195-game columns.
**And the gate registered to close the finding asserted the property for
both.** Its `checks` line was false of half of what it gates.
**Conceded.** Both harnesses assert now; the `checks` line says what is
actually checked.
### 2. Counting games proves they STARTED, not that they were PLAYED
Stopping the engine after one round (`round >= 5``>= 1`) gives **200
games in every cell, every value zero, exit 0, `make panels` green** —
byte for byte the signature CB-EV-0030 §3 claims the instrumentation
distinguishes from a real result.
**Conceded, and it is the sharpest finding of the three rounds.** The
counter answered *"did `play` return `Ok`"* while the claim made of it was
*"the zeros are real"*. Both harnesses now require every counted game to
have reached an outcome over five rounds; the one-round mutation fails
with the right message.
### 3. Round 2's `.csv` filter disabled the checks it was added beside
The filter was applied to **all three** loops — digest comparison, the CSV
parser check, and upstream freshness. So `catalog.yaml` and
`rules_delta.yaml`, whose missing digests were round 1's finding, were
recorded and then **never compared**, and never checked against upstream
at all. Tampering with both produced `[ok] × 12, exit 0`.
**Conceded.** Only the parser loop filters; tampering with either YAML now
fails on digest *and* freshness.
### 4. Five of six tiebreak comparators had no coverage
The strengthened oracle covered GR-E03's Stress key. **GR-E03's Bond key
and both of GR-E04's keys could be reversed with 179 tests and 26
scenarios green** — GR-E04's tiebreak never executes in any scenario,
because the only GR-E04 scenario sets `group_success: false`.
**Conceded.** All four are now covered and each was verified red under
mutation. **The Blame key took care to reach**: a coalition's score is
`sum(claimed blame)`, so a Blame token lowers the score and key 1
decides first — the key is only reachable when claims compensate, and the
test asserts the scores tie before relying on it.
## The serious ones
| # | challenge | outcome |
|---|---|---|
| 5 | "peak Stress **held**, not assigned" — `peak_held ≡ max(start, peak_payload)` for every possible input, so the correction computed the same number | **conceded.** The fix was cosmetic and the comment was false of the new code too. The real coupling was missing, which is #6 |
| 6 | `START_STRESS` was an unchecked constant — setting it to 7 printed `peak 7` everywhere | **conceded.** Now read off the dealt state. Verified: changing setup's Stress to 3 moves the reported peak to 3 |
| 7 | `cadence = "none"` was a pure loophole — a real target that runs nowhere and passes | **conceded, removed.** It had one user, the empty-target CHAOS entry, which the loop already skips. `manual` remains self-declared and unchecked, and that is now stated rather than implied |
| 8 | sibling discovery swapped a hand-written list for three hand-written globs; `metadata.json` and `VARIANT.md` were invisible, and both are named in the package's own `changed_files` | **conceded.** Walked, not globbed. **It found both immediately**, plus nested files the globs could never reach |
| 9 | "~72,000 games" is not producible; `make panels` runs **17,600** | **conceded.** Adopted from a reviewer's message and never re-derived — in the file whose correction history is about exactly that. The conclusion (zero divergence) reproduces |
## The minor ones
- **10 — two kinds of number were presented alike.** `600/800/1200` is
exactly `seats × games` on every sample; **`363` and `29` vary** —
7.1%11.5% inert across four windows. "92% of arms are live" is a fact
about seeds `0..200`. Now separated.
- **11 — the 29/363 citation, wrong twice, is now stated as one sample**
of a varying quantity. Both "independent confirmations" used the same
seeds, so what was confirmed was the *definition*. Whether round 1's
`StrictReactive` was the corrected policy cannot be determined: **that
reviewer's code was never kept**, which is itself worth fixing.
## What held, after being attacked
**Every one of the 24 cells in CB-EV-0030's DARVO table reproduces
exactly** — the first of three attempts at that table to produce no wrong
number. `regulation`'s `assert_eq!(c.games, GAMES)` catches both the setup
and the `play` path. The #13 fix is genuinely controlled. H1-B's mutation
coverage is real. `make all` runs `panels` and fails on it. The
1-arm-per-seat-per-game invariant holds on every sample — **and is still
unexplained.**
## What could not be checked
Upstream freshness (`../ground-game` not checked out, so that half of
`edition-check` has never run here); whether `metadata.json` and
`VARIANT.md` are load-bearing to anything; round 1's `StrictReactive`;
felt-play; cost.
---
## The finding that outlasts H1
Three rounds, each correcting the last, each introducing defects of the
same class. **The corrections are not getting safer.**
> A correction is written under the belief that the error is now
> understood. That belief is the condition under which this class of
> error is produced — so the correction inherits it, and the next round
> finds the same shape one level in.
**Round 4 is owed by the same argument.** The honest conclusion is not
that the work is nearly right; it is that **author-made corrections to
measurement work should be assumed defective until a fresh reader has
attacked them**, and this project should stop treating "corrected" as a
state closer to done than "found wrong".

View file

@ -66,18 +66,27 @@ SIBLINGS = "../"
# raised FileNotFoundError from `digest` instead of failing with the
# designed message, and "the sibling packages are covered" counted lines
# in a Markdown file -- it passed with all three files deleted.
SIBLING_GLOBS = ("catalog.yaml", "experiments/*/rules_delta.yaml", "experiments/*/*.csv")
# Everything under `editions/` that is not the edition directory itself.
# **Walked, not globbed** (CB-REV-0003 #8): the first version listed three
# hand-written glob patterns, which is the same self-certifying shape as
# reading the list out of PROVENANCE -- one hand-written list swapped for
# another. It missed `metadata.json` and `VARIANT.md`, both named in the
# package's OWN `changed_files` manifest, and anything a directory deeper.
SIBLING_SKIP = {".DS_Store"}
def sibling_files():
"""Sibling packages that are on disk, as `../`-relative paths."""
import glob as _glob
root = os.path.join(ROOT, EDITION, SIBLINGS)
"""Every file beside the edition, as `../`-relative paths."""
base = os.path.join(ROOT, "editions")
edition_dir = os.path.join(ROOT, EDITION)
out = []
for pattern in SIBLING_GLOBS:
for hit in _glob.glob(os.path.join(root, pattern)):
rel = os.path.relpath(hit, os.path.join(ROOT, EDITION))
for dirpath, _dirs, files in os.walk(base):
if os.path.abspath(dirpath).startswith(os.path.abspath(edition_dir)):
continue
for name in files:
if name in SIBLING_SKIP:
continue
rel = os.path.relpath(os.path.join(dirpath, name), edition_dir)
out.append(rel.replace(os.sep, "/"))
return sorted(out)
@ -98,7 +107,7 @@ def check():
print(f" [FAIL] a digest is recorded for a file that is not here: {', '.join(missing)}")
rc = 1
for name in [f for f in present if f.endswith(".csv")]:
for name in present:
if name not in want:
continue
path = os.path.join(ROOT, EDITION, name)
@ -118,6 +127,17 @@ def check():
# ADR-0015 D3's falsifier, checked rather than asserted: the hand
# reader handles commas inside quotes and NOTHING ELSE. A doubled
# quote or an embedded newline means `csv` is the answer after all.
# ONLY this loop filters: it is ADR-0015 D3's falsifier about the
# hand-rolled CSV reader, and running it over YAML made a valid
# `catalog.yaml` line with an odd quote count fail with a message
# about a parser that never reads it (CB-REV-0002 #9).
#
# **The digest and freshness loops must NOT filter** — CB-REV-0003 #3:
# the round-2 correction applied this filter to all three, so the two
# sibling YAMLs whose missing digests were round 1's finding were
# recorded and then never compared, and never checked against
# upstream at all. `make edition-check` answered its own headline
# question with [ok] when the answer was no.
for name in [f for f in present if f.endswith(".csv")]:
raw = open(os.path.join(ROOT, EDITION, name), encoding="utf-8-sig").read()
if '""' in raw:
@ -138,7 +158,7 @@ def check():
print(" [----] upstream not checked out — freshness UNVERIFIED")
print(f" expected {UPSTREAM_DIR}")
return rc
for name in [f for f in present if f.endswith(".csv")]:
for name in present:
up = os.path.join(UPSTREAM_DIR, name)
if not os.path.exists(up):
print(f" [FAIL] {name} is not in upstream — where did it come from?")

View file

@ -316,11 +316,13 @@ def check_gate_registry(root=REPO):
if not target:
continue
cadence = g.get("cadence")
if cadence not in ("all", "manual", "none"):
if cadence not in ("all", "manual"):
out.append(Finding(
"gates", "gates.toml",
f"{target!r} declares no cadence — say whether `make all` runs "
f"it, or a gate can exist without ever running"))
f"{target!r} declares no cadence — say `all` or `manual`. "
f"`none` was a pure loophole: a real target that runs "
f"nowhere and passes, which is the condition this rule "
f"exists to prevent (CB-REV-0003 #7)"))
elif cadence == "all" and target not in deps:
out.append(Finding(
"gates", "gates.toml",