CI: enforce every gate; close the silent-skip holes
The gates existed; CI ran half of them and tolerated the failure case. - cb-sim no longer has a "tolerable" non-zero exit. An unregistered game prefix is a failure, and a run in which nothing executed is a failure. Previously CI carried `|| test $? -eq 2`, so renaming a scenario prefix would have skipped every scenario while the pipeline stayed green. Verified with a negative control. - CI now runs make coverage (AM-1) and make dep-weight (AM-4), both added after CI was written and neither enforced until now. - dep-weight enforces its targets instead of only reporting them. - CI lints the shipped-runtime configuration separately, so the feature split cannot rot unnoticed. - Dropped the stale `make deps` target, which still measured the retired crate-count metric. The positive-control rule is now executable: CI runs `cargo bench -- --test`, which executes every benchmark once, so a workload that stalls fails the build. That step immediately found a fourth instance of the error class it was written for. The committed replay benchmark was the broken version — an earlier patch never applied, leaving a command sequence that omits Resolve, so every round produced nothing and the log-building loop spun forever. It had never run to completion; the reported AM-7 replay numbers came from a probe test instead. Fixed, given the same positive control as the round loop, and re-measured from the benchmark: 100k events fold in 2.18ms (95% CI 2.14-2.23), against a 5s budget. Evidence now reports confidence intervals rather than point estimates, so the 3% regression rule in MetricsAndScenarios is enforceable. The finding worth carrying: writing the positive-control rule into InnerLoop v1.0 did not prevent the next instance. Making it a CI step did. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
4be6e020ea
commit
72c594ee49
6 changed files with 169 additions and 70 deletions
|
|
@ -14,5 +14,27 @@ jobs:
|
||||||
- run: rustup component add rustfmt clippy
|
- run: rustup component add rustfmt clippy
|
||||||
- run: cargo fmt --all --check
|
- run: cargo fmt --all --check
|
||||||
- run: cargo clippy --workspace --all-targets -- -D warnings
|
- run: cargo clippy --workspace --all-targets -- -D warnings
|
||||||
|
|
||||||
|
# The shipped runtime is a distinct configuration (AM-4a); it must
|
||||||
|
# compile and lint on its own, or the feature split rots.
|
||||||
|
- run: cargo clippy -p games-ground --no-default-features -- -D warnings
|
||||||
|
|
||||||
- run: cargo test --workspace
|
- run: cargo test --workspace
|
||||||
- run: cargo run -q -p cb-sim -- scenarios/ground/*.yaml || test $? -eq 2
|
|
||||||
|
# No `|| test $? -eq 2`: cb-sim now fails on an unregistered game
|
||||||
|
# prefix and on a run that executed nothing. A silent skip is the
|
||||||
|
# failure this step exists to catch.
|
||||||
|
- run: cargo run -q -p cb-sim -- scenarios/ground/*.yaml
|
||||||
|
|
||||||
|
# AM-1: every numbered GR-rule is claimed by a scenario, and no
|
||||||
|
# scenario claims a rule the spec does not define.
|
||||||
|
- run: make coverage
|
||||||
|
|
||||||
|
# AM-4a/AM-4b: third-party source under audit, per configuration.
|
||||||
|
- run: make dep-weight
|
||||||
|
|
||||||
|
# InnerLoop v1.0 positive control, enforced rather than asserted in
|
||||||
|
# prose: --test runs every benchmark once, so a workload that
|
||||||
|
# stalls or produces the wrong event count fails the build instead
|
||||||
|
# of silently reporting throughput for work that never happened.
|
||||||
|
- run: cargo bench -p games-ground --bench synthetic -- --test
|
||||||
|
|
|
||||||
14
Makefile
14
Makefile
|
|
@ -3,7 +3,7 @@
|
||||||
|
|
||||||
CARGO := cargo
|
CARGO := cargo
|
||||||
|
|
||||||
.PHONY: check test sim bench deps loc all
|
.PHONY: check test sim bench bench-test coverage dep-weight loc all
|
||||||
|
|
||||||
## fmt + clippy (deny warnings) + HashMap deny-lint
|
## fmt + clippy (deny warnings) + HashMap deny-lint
|
||||||
check:
|
check:
|
||||||
|
|
@ -24,13 +24,15 @@ coverage:
|
||||||
sim:
|
sim:
|
||||||
$(CARGO) run -q -p cb-sim -- scenarios/ground/*.yaml
|
$(CARGO) run -q -p cb-sim -- scenarios/ground/*.yaml
|
||||||
|
|
||||||
## Criterion benches (AM-6/AM-7 wiring)
|
## Criterion benches (AM-6/AM-7)
|
||||||
bench:
|
bench:
|
||||||
$(CARGO) bench -p games-ground
|
$(CARGO) bench -p games-ground
|
||||||
|
|
||||||
## AM-4: transitive crate count (excludes dev/build deps)
|
## InnerLoop positive control: run every bench once, no measurement.
|
||||||
deps:
|
## Fails if a workload stalls or produces the wrong event count.
|
||||||
@$(CARGO) tree --workspace -e normal --prefix none | sed 's/ (\*)//' | sort -u | grep -vE '^(cb-|games-|cb_|$$)' | tee /dev/stderr | wc -l
|
bench-test:
|
||||||
|
$(CARGO) bench -p games-ground --bench synthetic -- --test
|
||||||
|
|
||||||
|
|
||||||
## AM-2/AM-3 input: source LOC per crate (excludes tests would need tokei)
|
## AM-2/AM-3 input: source LOC per crate (excludes tests would need tokei)
|
||||||
loc:
|
loc:
|
||||||
|
|
@ -38,4 +40,4 @@ loc:
|
||||||
printf '%-28s %s\n' $$d "$$(find $$d/src -name '*.rs' | xargs cat | grep -vcE '^\s*(//|$$)')"; \
|
printf '%-28s %s\n' $$d "$$(find $$d/src -name '*.rs' | xargs cat | grep -vcE '^\s*(//|$$)')"; \
|
||||||
done
|
done
|
||||||
|
|
||||||
all: check test sim
|
all: check test sim coverage dep-weight bench-test
|
||||||
|
|
|
||||||
|
|
@ -21,7 +21,7 @@ Machine: WSL2, Linux 6.18.33.2-microsoft-standard-WSL2, rustc 1.97.1,
|
||||||
| AM-4b dep weight, dev toolchain | ≤350,000 third-party lines | 317,021 (29 crates) | **met** |
|
| AM-4b dep weight, dev toolchain | ≤350,000 third-party lines | 317,021 (29 crates) | **met** |
|
||||||
| AM-6 throughput | ≥100,000 events/s | 1,651,400 events/s | **met, 16.5×** |
|
| AM-6 throughput | ≥100,000 events/s | 1,651,400 events/s | **met, 16.5×** |
|
||||||
| AM-7 scaling | ≥0.9× at 20× workload | 1.08× | **met** |
|
| AM-7 scaling | ≥0.9× at 20× workload | 1.08× | **met** |
|
||||||
| AM-7 replay | 100k events ≤5s | 4.13 ms | **met, 1,210×** |
|
| AM-7 replay | 100k events ≤5s | 2.18 ms (CI 2.14–2.23) | **met, 2,290×** |
|
||||||
| AM-8 determinism | zero divergence, 10 replays | 1 distinct hash / 10 runs | **met** |
|
| AM-8 determinism | zero divergence, 10 replays | 1 distinct hash / 10 runs | **met** |
|
||||||
| AM-8 lint | fmt + clippy clean | clean, `-D warnings` | **met** |
|
| AM-8 lint | fmt + clippy clean | clean, `-D warnings` | **met** |
|
||||||
| AM-10 foreign types | zero `HashMap`/`HashSet` | 0 | **met** |
|
| AM-10 foreign types | zero `HashMap`/`HashSet` | 0 | **met** |
|
||||||
|
|
@ -45,12 +45,17 @@ breaks the test rather than silently rescaling the metric.
|
||||||
| 40,000 | 1,626,000 | 1.07× |
|
| 40,000 | 1,626,000 | 1.07× |
|
||||||
| 100,000 | 1,651,400 | 1.08× |
|
| 100,000 | 1,651,400 | 1.08× |
|
||||||
|
|
||||||
Replay — folding one growing event log back into state:
|
Replay — folding one growing event log back into state (Criterion
|
||||||
|
95% CI, low/median/high):
|
||||||
|
|
||||||
| Events | Time | Rate |
|
| Events | Time (median) | 95% CI | Rate |
|
||||||
|---|---|---|
|
|---|---|---|---|
|
||||||
| 10,010 | 465 µs | 21.5M events/s |
|
| 10,010 | 184.8 µs | 180.9 – 189.3 µs | 54.2M events/s |
|
||||||
| 100,007 | 4.13 ms | 24.2M events/s |
|
| 100,072 | 2.183 ms | 2.142 – 2.226 ms | 45.8M events/s |
|
||||||
|
|
||||||
|
**Correction (2026-07-31).** These originally read 465 µs / 4.13 ms and
|
||||||
|
were taken from the `replay_probe` **test**, not from the benchmark —
|
||||||
|
because the benchmark did not work. See §2a.
|
||||||
|
|
||||||
### The comparison against boardgame.io, stated carefully
|
### The comparison against boardgame.io, stated carefully
|
||||||
|
|
||||||
|
|
@ -100,6 +105,30 @@ The benchmark now asserts the per-round event count on every round and
|
||||||
panics rather than measuring a stalled loop. The corrected figure is
|
panics rather than measuring a stalled loop. The corrected figure is
|
||||||
**5.6× lower** than the bogus one.
|
**5.6× lower** than the bogus one.
|
||||||
|
|
||||||
|
### 2a. A fourth measurement error, found by enforcing the rule
|
||||||
|
|
||||||
|
The replay benchmark committed alongside this evidence was **the broken
|
||||||
|
version**. A `python3` patch that was supposed to replace its
|
||||||
|
log-building loop never applied, leaving a sequence that omits `Resolve`
|
||||||
|
— so `EndRound` was rejected, every round produced no events, and the
|
||||||
|
`while log.len() < target` loop spun forever. It was never run to
|
||||||
|
completion; the AM-7 replay numbers were taken from a separate probe
|
||||||
|
test instead, and the dead benchmark was committed and left hanging.
|
||||||
|
|
||||||
|
Found by adding `cargo bench -- --test` to CI, which runs every
|
||||||
|
benchmark once. That is the fourth instance of one error class in this
|
||||||
|
project — a harness that appears to work while doing no work — and the
|
||||||
|
**first one caught by a gate rather than by noticing**.
|
||||||
|
|
||||||
|
The replay loop now carries the positive control the round loop already
|
||||||
|
had: it asserts each round appended events and fails rather than
|
||||||
|
spinning. Corrected figures are in the table above; both configurations
|
||||||
|
still clear the AM-7 budget by three orders of magnitude.
|
||||||
|
|
||||||
|
The lesson recorded for the loop: writing the positive-control rule into
|
||||||
|
`specs/InnerLoop.md` did **not** prevent the next instance. Making it a
|
||||||
|
CI step did. Prose rules do not enforce themselves.
|
||||||
|
|
||||||
## 3. Determinism (AM-8)
|
## 3. Determinism (AM-8)
|
||||||
|
|
||||||
- Every scenario runs twice per invocation with the same seed and fails
|
- Every scenario runs twice per invocation with the same seed and fails
|
||||||
|
|
|
||||||
|
|
@ -122,6 +122,44 @@ const FINAL_ROUND_EVENTS: usize = 12;
|
||||||
/// integers: (4 x 13 + 12) = 64.
|
/// integers: (4 x 13 + 12) = 64.
|
||||||
const EVENTS_PER_5_ROUNDS: usize = 64;
|
const EVENTS_PER_5_ROUNDS: usize = 64;
|
||||||
|
|
||||||
|
/// Play one round, appending every applied event to `log`.
|
||||||
|
fn record_round(state: &mut GroundState, log: &mut Vec<games_ground::GroundEvent>) {
|
||||||
|
let mut run = |state: &mut GroundState, actor: Actor, cmd: &GroundCommand| {
|
||||||
|
if let Ok(produced) = state.validate(actor, cmd) {
|
||||||
|
for event in &produced {
|
||||||
|
state.fold(event);
|
||||||
|
log.push(event.clone());
|
||||||
|
}
|
||||||
|
}
|
||||||
|
};
|
||||||
|
for (seat, action, target) in [
|
||||||
|
(0u8, Action::Attack, Some(PlayerId(1))),
|
||||||
|
(2, Action::Support, Some(PlayerId(1))),
|
||||||
|
(1, Action::Ground, None),
|
||||||
|
] {
|
||||||
|
run(
|
||||||
|
state,
|
||||||
|
Actor::Player(PlayerId(seat)),
|
||||||
|
&GroundCommand::SelectAction {
|
||||||
|
action,
|
||||||
|
target,
|
||||||
|
problem: None,
|
||||||
|
},
|
||||||
|
);
|
||||||
|
}
|
||||||
|
run(state, Actor::System, &GroundCommand::Reveal);
|
||||||
|
run(
|
||||||
|
state,
|
||||||
|
Actor::Player(PlayerId(1)),
|
||||||
|
&GroundCommand::ChooseGroundMode {
|
||||||
|
mode: GroundMode::Gr,
|
||||||
|
choice: None,
|
||||||
|
},
|
||||||
|
);
|
||||||
|
run(state, Actor::System, &GroundCommand::Resolve);
|
||||||
|
run(state, Actor::System, &GroundCommand::EndRound);
|
||||||
|
}
|
||||||
|
|
||||||
fn bench_synthetic(c: &mut Criterion) {
|
fn bench_synthetic(c: &mut Criterion) {
|
||||||
// Events per round is fixed by the workload, so throughput can be
|
// Events per round is fixed by the workload, so throughput can be
|
||||||
// reported in events/second — the AM-6 unit.
|
// reported in events/second — the AM-6 unit.
|
||||||
|
|
@ -146,55 +184,30 @@ fn bench_synthetic(c: &mut Criterion) {
|
||||||
// flat curve is partly by construction — this one is not, because
|
// flat curve is partly by construction — this one is not, because
|
||||||
// the log here grows without bound.
|
// the log here grows without bound.
|
||||||
let mut replay = c.benchmark_group("replay-ground-3p");
|
let mut replay = c.benchmark_group("replay-ground-3p");
|
||||||
for &events in &[10_000usize, 100_000] {
|
for &target_events in &[10_000usize, 100_000] {
|
||||||
replay.throughput(Throughput::Elements(events as u64));
|
// Build one log by playing real rounds, then measure folding it
|
||||||
replay.bench_function(format!("fold-{events}-events"), |b| {
|
// back. Must use the full command sequence: a shortened one
|
||||||
// Build one log of `events` events, then measure folding it.
|
// stalls, because Reveal needs every seat's selection and
|
||||||
let mut source = setup(42);
|
// EndRound is gated on Resolve.
|
||||||
let mut log = Vec::with_capacity(events);
|
let mut log = Vec::with_capacity(target_events);
|
||||||
while log.len() < events {
|
let mut source = setup(42);
|
||||||
if source.outcome.is_some() {
|
let mut games = 0u64;
|
||||||
source = setup(43);
|
while log.len() < target_events {
|
||||||
}
|
if source.outcome.is_some() {
|
||||||
let picks = [
|
games += 1;
|
||||||
(
|
source = setup(42 + games);
|
||||||
PlayerId(0),
|
|
||||||
GroundCommand::SelectAction {
|
|
||||||
action: Action::Attack,
|
|
||||||
target: Some(PlayerId(1)),
|
|
||||||
problem: None,
|
|
||||||
},
|
|
||||||
),
|
|
||||||
(
|
|
||||||
PlayerId(1),
|
|
||||||
GroundCommand::SelectAction {
|
|
||||||
action: Action::Support,
|
|
||||||
target: Some(PlayerId(2)),
|
|
||||||
problem: None,
|
|
||||||
},
|
|
||||||
),
|
|
||||||
];
|
|
||||||
for (seat, cmd) in picks {
|
|
||||||
if let Ok(produced) = source.validate(Actor::Player(seat), &cmd) {
|
|
||||||
for e in &produced {
|
|
||||||
source.fold(e);
|
|
||||||
log.push(e.clone());
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
if let Ok(produced) = source.validate(Actor::System, &GroundCommand::Reveal) {
|
|
||||||
for e in &produced {
|
|
||||||
source.fold(e);
|
|
||||||
log.push(e.clone());
|
|
||||||
}
|
|
||||||
}
|
|
||||||
if let Ok(produced) = source.validate(Actor::System, &GroundCommand::EndRound) {
|
|
||||||
for e in &produced {
|
|
||||||
source.fold(e);
|
|
||||||
log.push(e.clone());
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
}
|
||||||
|
let before = log.len();
|
||||||
|
record_round(&mut source, &mut log);
|
||||||
|
// Positive control: a round that yields nothing means the
|
||||||
|
// workload stalled, and the loop above would spin forever.
|
||||||
|
assert!(
|
||||||
|
log.len() > before,
|
||||||
|
"replay workload stalled: a round produced no events"
|
||||||
|
);
|
||||||
|
}
|
||||||
|
replay.throughput(Throughput::Elements(log.len() as u64));
|
||||||
|
replay.bench_function(format!("fold-{target_events}-events"), |b| {
|
||||||
b.iter(|| {
|
b.iter(|| {
|
||||||
let mut state = setup(42);
|
let mut state = setup(42);
|
||||||
for event in &log {
|
for event in &log {
|
||||||
|
|
|
||||||
|
|
@ -1,7 +1,10 @@
|
||||||
//! cb-sim — scenario runner binary (GameKernel K17; precursor of `cb sim`).
|
//! cb-sim — scenario runner binary (GameKernel K17; precursor of `cb sim`).
|
||||||
//! Parses scenario files, dispatches each to its game by the `<game>/`
|
//! Parses scenario files, dispatches each to its game by the `<game>/`
|
||||||
//! prefix of its `scenario` field, and executes it. Exit codes: 0 all
|
//! prefix of its `scenario` field, and executes it. Exit codes: 0 all
|
||||||
//! passed, 1 failures, 2 unknown game, 64 usage error.
|
//! passed, 1 any failure — including a scenario whose game prefix is not
|
||||||
|
//! registered, and a run in which nothing executed. There is deliberately
|
||||||
|
//! no "tolerable" non-zero exit: a silent skip is the failure mode this
|
||||||
|
//! binary exists to catch.
|
||||||
|
|
||||||
use cb_game_runtime::{scenario, RunOutcome, ScenarioFile};
|
use cb_game_runtime::{scenario, RunOutcome, ScenarioFile};
|
||||||
use games_ground::GroundState;
|
use games_ground::GroundState;
|
||||||
|
|
@ -13,7 +16,6 @@ fn main() {
|
||||||
std::process::exit(64);
|
std::process::exit(64);
|
||||||
}
|
}
|
||||||
|
|
||||||
let mut unknown_game = false;
|
|
||||||
let mut failed = false;
|
let mut failed = false;
|
||||||
let mut passed = 0usize;
|
let mut passed = 0usize;
|
||||||
let mut covered: Vec<String> = Vec::new();
|
let mut covered: Vec<String> = Vec::new();
|
||||||
|
|
@ -39,8 +41,14 @@ fn main() {
|
||||||
let outcome = match sc.scenario.split('/').next() {
|
let outcome = match sc.scenario.split('/').next() {
|
||||||
Some("ground") => scenario::run::<GroundState>(&sc),
|
Some("ground") => scenario::run::<GroundState>(&sc),
|
||||||
_ => {
|
_ => {
|
||||||
println!("SKIP {} — no game registered for this prefix", sc.scenario);
|
// A renamed or typo'd prefix would otherwise skip every
|
||||||
unknown_game = true;
|
// scenario while the run still looked clean.
|
||||||
|
eprintln!(
|
||||||
|
"FAIL {} — no game registered for prefix {:?}",
|
||||||
|
sc.scenario,
|
||||||
|
sc.scenario.split('/').next().unwrap_or("")
|
||||||
|
);
|
||||||
|
failed = true;
|
||||||
continue;
|
continue;
|
||||||
}
|
}
|
||||||
};
|
};
|
||||||
|
|
@ -67,10 +75,14 @@ fn main() {
|
||||||
covered.dedup();
|
covered.dedup();
|
||||||
println!("{passed} passed, {} rules covered", covered.len());
|
println!("{passed} passed, {} rules covered", covered.len());
|
||||||
|
|
||||||
|
// Positive control: a run that executed nothing must not pass. This
|
||||||
|
// is the same class of error as a benchmark timing rejected work.
|
||||||
|
if passed == 0 {
|
||||||
|
eprintln!("FAIL — no scenario executed; refusing to report success");
|
||||||
|
failed = true;
|
||||||
|
}
|
||||||
|
|
||||||
if failed {
|
if failed {
|
||||||
std::process::exit(1);
|
std::process::exit(1);
|
||||||
}
|
}
|
||||||
if unknown_game {
|
|
||||||
std::process::exit(2);
|
|
||||||
}
|
|
||||||
}
|
}
|
||||||
|
|
|
||||||
|
|
@ -32,6 +32,13 @@ CONFIGS = {
|
||||||
"dev-toolchain": [],
|
"dev-toolchain": [],
|
||||||
}
|
}
|
||||||
|
|
||||||
|
# AM-4a / AM-4b targets from specs/GameKernel.md §4. Breaching one fails
|
||||||
|
# the build: a gate that only reports is a suggestion.
|
||||||
|
TARGETS = {
|
||||||
|
"shipped-runtime": 250_000,
|
||||||
|
"dev-toolchain": 350_000,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
def crates(extra_args):
|
def crates(extra_args):
|
||||||
"""Third-party crates in the normal (non-dev) dependency graph."""
|
"""Third-party crates in the normal (non-dev) dependency graph."""
|
||||||
|
|
@ -116,9 +123,12 @@ def main():
|
||||||
print(f" own source {own:>9,} lines")
|
print(f" own source {own:>9,} lines")
|
||||||
for label in CONFIGS:
|
for label in CONFIGS:
|
||||||
r = report[label]
|
r = report[label]
|
||||||
|
limit = TARGETS[label]
|
||||||
|
mark = "ok " if r["third_party_loc"] <= limit else "FAIL"
|
||||||
print(
|
print(
|
||||||
f" {label:<18}{r['crates']:>3} crates "
|
f" {label:<18}{r['crates']:>3} crates "
|
||||||
f"{r['third_party_loc']:>9,} lines third-party"
|
f"{r['third_party_loc']:>9,} lines third-party "
|
||||||
|
f"[{mark} target {limit:,}]"
|
||||||
)
|
)
|
||||||
delta = (
|
delta = (
|
||||||
report["dev-toolchain"]["third_party_loc"]
|
report["dev-toolchain"]["third_party_loc"]
|
||||||
|
|
@ -131,7 +141,18 @@ def main():
|
||||||
# shrink the total, so refuse to report rather than under-report.
|
# shrink the total, so refuse to report rather than under-report.
|
||||||
print("\nERROR — source not found for:", ", ".join(missing), file=sys.stderr)
|
print("\nERROR — source not found for:", ", ".join(missing), file=sys.stderr)
|
||||||
return 1
|
return 1
|
||||||
return 0
|
|
||||||
|
breached = [
|
||||||
|
(label, report[label]["third_party_loc"], limit)
|
||||||
|
for label, limit in TARGETS.items()
|
||||||
|
if report[label]["third_party_loc"] > limit
|
||||||
|
]
|
||||||
|
for label, actual, limit in breached:
|
||||||
|
print(
|
||||||
|
f"\nFAIL AM-4 — {label}: {actual:,} lines exceeds target {limit:,}",
|
||||||
|
file=sys.stderr,
|
||||||
|
)
|
||||||
|
return 1 if breached else 0
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
if __name__ == "__main__":
|
||||||
|
|
|
||||||
Loading…
Add table
Add a link
Reference in a new issue