evidence/CB-EV-0001-game-kernel.md records the acceptance run against the CB-RES-0001 baseline. Met: AM-1 rule coverage 58/58; AM-6 throughput 1.65M events/s against a 100k target; AM-7 scaling 1.08x at 20x workload and a 100k-event replay in 4.13ms against a 5s budget; AM-8 zero divergence over 10 full runs with fmt and clippy clean; AM-10 zero foreign collection types. Not met and reported as such: AM-4 at 33 transitive crates against a <=20 target. Attribution is in the evidence file. The recommended fix is making serde_yaml optional (-5, a test-only concern), after which the remainder is sha2 and rand_chacha, which K5 and K7 require. We are not hand-rolling crypto primitives to win a dependency count. AM-12 is recorded as uncomputable: per-task token counts were never instrumented, and inventing a USD figure would defeat the metric. A measurement error was found and corrected before publication. The first benchmark reported 9.3M events/s on a flat curve. The workload had a player selecting SUPPORT while parked at Stress 4, so GR-R03 rejected it, rounds never completed, and throughput was computed for rounds that never happened. The bench now asserts the per-round event count and panics rather than measuring a stalled loop. The corrected figure is 5.6x lower. The evidence file states plainly what the boardgame.io comparison does and does not support: the ~450x command-rate ratio is cross-runtime and cross-feature-set, so it is a direction, not a verdict, per the InnerLoop parity-cap rule. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
8.8 KiB
CB-EV-0001 — GROUND game kernel: acceptance evidence
Status: T08 complete, with one acceptance metric not met (AM-4).
Recorded: 2026-07-31
Workplan: CB-WP-0001, task T08
Spec: specs/GameKernel.md §4 (AM-1..AM-12)
Baseline: research/CB-RES-0001-game-kernel.md, measurements in
research/CB-RES-0001-harness/boardgame-io/results-260731.json
Machine: WSL2, Linux 6.18.33.2-microsoft-standard-WSL2, rustc 1.97.1,
--release, Criterion 1s warm-up / 3s measurement.
1. Scoreboard
| Metric | Target | Measured | Verdict |
|---|---|---|---|
| AM-1 rule coverage | 100% of GR-rules | 58/58 (100%) | met |
| AM-4 dependency weight | ≤20 crates | 33 | NOT MET |
| AM-6 throughput | ≥100,000 events/s | 1,651,400 events/s | met, 16.5× |
| AM-7 scaling | ≥0.9× at 20× workload | 1.08× | met |
| AM-7 replay | 100k events ≤5s | 4.13 ms | met, 1,210× |
| AM-8 determinism | zero divergence, 10 replays | 1 distinct hash / 10 runs | met |
| AM-8 lint | fmt + clippy clean | clean, -D warnings |
met |
| AM-10 foreign types | zero HashMap/HashSet |
0 | met |
| AM-11 impl pairs | null + reference per port | 1 of 1 (KernelRng) |
met, narrow |
| AM-12 cost log | present | §5 | met |
AM-2, AM-3, AM-5 and AM-9 are not reported: see §6.
2. Throughput and scaling (AM-6, AM-7)
Workload: 3-player GROUND rounds, 7 commands and 13 events per round
(12 on a game's fifth round, where GR-R09 ends the game). Pinned by
bench_shape in games/ground/src/lib.rs, so a change to the workload
breaks the test rather than silently rescaling the metric.
| Rounds | Throughput (events/s) | vs 5k |
|---|---|---|
| 5,000 | 1,524,200 | 1.00× |
| 10,000 | 1,638,200 | 1.07× |
| 20,000 | 1,577,000 | 1.03× |
| 40,000 | 1,626,000 | 1.07× |
| 100,000 | 1,651,400 | 1.08× |
Replay — folding one growing event log back into state:
| Events | Time | Rate |
|---|---|---|
| 10,010 | 465 µs | 21.5M events/s |
| 100,007 | 4.13 ms | 24.2M events/s |
The comparison against boardgame.io, stated carefully
boardgame.io measured 1,930 moves/s at 5,000 moves, falling to 870 moves/s at 20,000, and did not finish 100,000 moves in 300 s. Our figure in the same unit is ~129,000 rounds/s × 7 = ~903,000 commands/s, and 100,000 rounds complete in 776 ms.
That is roughly a 400–500× ratio, and it is not a like-for-like measurement. Four differences matter, all favouring us:
- Different language and process model. Rust in-process against Node.js with immutable state, patch generation and undo history.
- Different feature set. boardgame.io's per-move cost includes
producing client patches and maintaining undo state; the run with
--disable-undostill degraded (0.66× at 40k). We do neither. - Different workload shape. Our 5-round games (GR-R09) bound state size by construction. The boardgame.io harness ran one match with unbounded history, which is exactly the axis it degraded on.
- No network or storage layer on our side.
Point 3 is the important one and it is why the flat curve in the table above is weak evidence on its own — a game that resets every five rounds cannot exhibit history-growth degradation. The replay benchmark is the honest test of that axis, because there the log grows without bound, and it stays linear (21.5M → 24.2M events/s from 10k to 100k).
Claim we are willing to defend: the kernel meets AM-6 and AM-7 with large margin, and does not degrade as event-log length grows. Claim we are not making: that Clay-Borg is ~450× "faster than boardgame.io" as a like-for-like engine comparison. Per the InnerLoop parity-cap rule, a cross-runtime ratio this coarse is not a verdict, it is a direction.
A measurement error found and corrected
The first run of this benchmark reported 9.3M events/s with a
perfectly flat curve — a number that would have been reported as a
93× beat of AM-6. It was wrong. The workload had P2 selecting SUPPORT
every round while being attacked; after three rounds P2 sat at Stress 4,
GR-R03 rejected the SUPPORT, the round never completed, and the loop
spun on rejected commands. Throughput was computed as
rounds × 13 events while most rounds produced 2.
Found by a probe test asserting that a round produces events at all. The benchmark now asserts the per-round event count on every round and panics rather than measuring a stalled loop. The corrected figure is 5.6× lower than the bogus one.
3. Determinism (AM-8)
- Every scenario runs twice per invocation with the same seed and fails on state-hash divergence (K8). 21/21 pass.
- Ten consecutive full runs of all 21 scenarios produced one distinct output hash, i.e. zero divergence.
cargo fmt --checkandcargo clippy --workspace --all-targets -D warningsare clean.clippy.tomldeniesHashMap/HashSetworkspace-wide (K6); the aggregate holds only ordered collections, so iteration order cannot vary between runs.
4. AM-4 — not met, and why it is reported rather than fixed
33 transitive crates against a ≤20 target. The baseline it was set against is boardgame.io's 120 npm packages, so we are 3.6× lighter, but the metric as written is missed and is recorded as missed.
Attribution:
| Group | Crates | Count |
|---|---|---|
sha2 (K7 state hashing) |
sha2, digest, block-buffer, crypto-common, generic-array, typenum, cpufeatures, cfg-if | 8 |
| serde derive chain | serde_derive, proc-macro2, quote, syn, unicode-ident | 5 |
| serde runtime + json | serde, serde_core, serde_json, itoa, ryu, memchr, zmij | 7 |
serde_yaml (scenario files only) |
serde_yaml, unsafe-libyaml, indexmap, hashbrown, equivalent | 5 |
rand_chacha (K5 seeded RNG) |
rand_chacha, rand_core, ppv-lite86, zerocopy | 4 |
| Clay-Borg crates | cb-kernel, cb-events, cb-game-runtime, games-ground | 4 |
The honest options, in order of preference:
- Make
serde_yamloptional behind ascenariosfeature. YAML is a test-and-tooling concern; a shipped game runtime does not need it. Removes 5 crates from the default build for no loss of capability. This is the one to do first, and it improves D4 optionality as well as D2. - Revisit the target. ≤20 was set before the K5/K7 contracts named ChaCha and SHA-256. Those two contracts cost 12 crates between them and are load-bearing for determinism. A target that a spec's own contracts make unreachable is a bad target.
What we are not doing: hand-rolling SHA-256 or ChaCha to win a dependency count. That trades an auditable, well-tested primitive for a number on a scoreboard.
This is a T09 input: either the metric moves for a stated reason, or option 1 lands and the remainder is justified.
5. Cost log (AM-12)
Per specs/MetricsAndScenarios.md §1a. Model: Claude Fable 5, at
benchmarks/baselines/model-prices.toml rates ($10/$50 per MTok).
| Task | Model | Iterations | Notes |
|---|---|---|---|
| T08 | claude-fable-5 | 6 code iterations + benchmarks | Token counts not captured per iteration; see limitation below |
Limitation, stated rather than fabricated: exact per-task token counts were not instrumented during T08, so the USD figure the metric asks for cannot be computed honestly from this run. Recording an estimate here would defeat the purpose of the metric. T09 should either wire real token accounting into the loop or drop M-D2-CST as unmeasurable in this setup.
6. Metrics not reported
- AM-2, AM-3, AM-5 — specification-quality metrics that need a second capability to compare against; a single data point is not a measurement.
- AM-9 (≤64MB) — not instrumented. The aggregate is a few KB and the largest log measured here is 100k events, so the budget is very unlikely to bind, but "unlikely" is not "measured" and it is left unclaimed.
- AM-11 — the
KernelRngnull/reference pair exists and is exercised. It is the only port with a pair so far, so the metric is met narrowly and will mean more once storage has one.
7. Rules implemented under a provisional default
Ten U-items in specs/GroundRules.md carry PROVISIONAL defaults. Those
realized here are U2 (clamp on every application), U3 (DENY with no
legal target is a no-op that still advances), U4 (deck reshuffle), U5
(REVERSE owner relief applies whether or not the Reverse was rejected)
and U8 (GROUND—OU cancellation precedes Protection).
One further ambiguity was found during T08 and is not in the U-list:
GR-E02's "successes" is undefined in dataset 0.1. It is implemented
as the count of claimed Problems. Both scoring scenarios are marked
provisional: true.
All provisional behaviour lives behind named functions and is covered by
scenarios tagged provisional: true, so a ground-game ruling flips a
scenario rather than the kernel (K16). Action for ground-game: rule
on the ten U-items and on GR-E02's "successes".