# Game Kernel — specification and acceptance metrics Status: **v0.1 draft** — implements ADR-0002 (reimplement in Rust, assimilate patterns). Governs crates `cb-kernel`, `cb-events`, `cb-game-runtime`, and the first game package `games/ground`. Baselines referenced: research/CB-RES-0001-game-kernel.md. Instruments: specs/MetricsAndScenarios.md. Rules content: specs/GroundRules.md. The kernel is the **authoritative semantic layer** of the Clay-Borg architecture (Blueprint §1): headless, deterministic, replayable. No rendering, no physics, no networking — those attach later through ports and must never leak types into anything specified here (M-D4-LEAK = 0). --- ## 1. Crate boundaries | Crate | Owns | Must not know about | |---|---|---| | `cb-kernel` | canonical IDs, RNG service, command/event/fold traits, clock-free tick model | any concrete game, serialization format details | | `cb-events` | event envelope, append-only log, snapshots, canonical serialization, state hashing | game semantics | | `cb-game-runtime` | round/phase machinery, simultaneous commit windows, per-player projections, scenario-runner API | GROUND specifics | | `games/ground` | GROUND rules: state aggregate, commands, events, reducers per specs/GroundRules.md | anything below `cb-game-runtime`'s public API | Dependency rule: `games/ground → cb-game-runtime → cb-events → cb-kernel`. No cycle, no skip that bypasses a public API. External crates allowed in the headless kernel workspace: serde (+format crate), a seedable RNG (e.g. chacha), a hash (sha2), thiserror-class error derive. Weight is budgeted as **third-party source under audit**, split by build configuration — see AM-4. **Scenario parsing is dev-only.** `cb-game-runtime`'s `scenarios` feature carries the YAML dependency; a shipped game runtime builds with `--no-default-features` and parses no YAML. Workspace dependencies on `cb-game-runtime` and `games-ground` therefore set `default-features = false`, and consumers that need scenarios (currently `cb-sim`) opt in explicitly. ## 2. Canonical model ### 2.1 Identifiers (`cb-kernel`) Newtyped, copyable, ordered: `GameId`, `PlayerId`, `EntityId` (problems, tokens, relations), `CommandId`, `EventId` (monotonic `u64` sequence per game), `Seed` (`u64`). No UUIDs inside the hot path; string forms only at the boundary. ### 2.2 The mutation pipeline (assimilated: cqrs-es pattern) ```text Command (actor-tagged intent) → validate(&State, &Command) → Result, Rejection> → fold(&mut State, &Event) # infallible, total → EventLog::append(events) # with per-event hash chaining ``` - **K1** All state change flows through events; there is no other mutator. `fold` is infallible: every validation that could fail happens in `validate`, so a logged event always applies. - **K2** Commands carry: `CommandId`, issuing `PlayerId` (or `System`), round number, and payload. Duplicate `CommandId` within a game is rejected (idempotency). - **K3** Rejections are typed and testable (`Rejection::StressGate`, `::NotYourTarget`, `::SlotOccupied`, …); scenarios assert on them (`expect.rejects`). - **K4** Events are closed enums per game package, serialized through a versioned envelope: `{seq, game_id, round, kind, payload, schema_ver}`. ### 2.3 Determinism (assimilated: Rune's enforced discipline) - **K5** The only randomness source is the kernel RNG service, seeded from `Seed`; RNG draws are themselves events (or deterministically derived from the event sequence) so replay never re-rolls. - **K6** No ambient time, no OS entropy, no pointer/hash iteration order: all iterated collections in semantic state are ordered (`BTreeMap`/ `Vec`); `HashMap` is forbidden in `games/*` state. Enforced by lint/deny in CI, not convention (AM-8). - **K7** State hash = SHA-256 over the canonical serialization of the aggregate; computed on demand and recorded at round End events. - **K8** Double-run invariant: the scenario runner executes every scenario twice with the same seed and fails on any hash divergence (MetricsAndScenarios §2). ### 2.4 Snapshots and replay (`cb-events`) - **K9** A snapshot is the canonical serialization of the full aggregate + the `EventId` it includes. `snapshot + remaining events → state` must be hash-identical to a from-genesis fold (AM-7). - **K10** A replay bundle (`.cbreplay`, MetricsAndScenarios §4) contains manifest, command log, initial snapshot, and failed expectations; the runner's `--replay` re-executes it bit-identically. - **K11** Log format is append-only, length-prefixed, versioned; a truncated tail is detected, not silently accepted. ### 2.5 Simultaneity and hidden information (`cb-game-runtime`) - **K12** Commit window primitive (assimilated: GROUND's Select/Reveal, boardgame.io `activePlayers` shape): the runtime opens a window naming the players who must submit; submissions are recorded as **commitment events** whose payload is hidden in projections until the window's reveal event; late/duplicate submissions are Rejections. - **K13** Projection (assimilated: OpenSpiel information states): for each `PlayerId` (and `Spectator`), a total function `project(&State) → PlayerView` that structurally cannot include: other players' unrevealed commitments, face-down problem identities, other players' hands. Projections are derived views — never inputs to `validate`/`fold`. - **K14** *(amended 2026-08-01, CB-WP-0006 T07 — see §2.5a)* The GROUND round (GR-R01..R09) **implements the commit-window contract**: one window per Select, duplicate submission rejected, completeness required before Reveal, then fixed-order resolution steps with Lead-order iteration (GR-R06/R07) driven by system commands. It implements that contract **in its own aggregate**; `CommitWindow` in `cb-game-runtime` is the extracted form, and is **provisional until a second game uses it**. ### 2.6 GROUND aggregate (`games/ground`) #### 2.5a Why K14 was amended rather than implemented The original K14 said the round *is expressed through* runtime primitives. It was not: `CommitWindow` had **zero non-test users** and `games/ground` did not import it. GROUND collects selections in its own `BTreeMap` and enforces the same contract inline — duplicate submission (`second_selection_is_a_duplicate`), completeness before Reveal (`GR-R04`), and ordered reveal. Two options were available and the choice is recorded because it could reasonably have gone the other way. **Wiring GROUND through `CommitWindow` was rejected.** It would change the serialized shape of `selections`, which four scenario files assert by dot-path (`selections.0.action`) and which every state hash depends on — a large, risky refactor whose only benefit is making a sentence literally true. More importantly it inverts INTENT: *"abstractions are extracted from working games... rather than invented in isolation"*, and *"No concept becomes canonical merely because it looks general. It becomes canonical after surviving a second concrete use."* `CommitWindow` was invented before any game needed it and has survived **zero** uses. Imposing it on GROUND now would manufacture the first use rather than discover it. **So the rule was amended to state what is actually guaranteed**, and the primitive is kept and marked provisional. It costs ~60 lines, it documents the seam stage 3 (networked sessions) and stage 4 (packaging) will need, and re-extracting it from GROUND when a second game exists will be better-informed than keeping it aligned by hand now. **Open, with a date:** if no second game uses `CommitWindow` by **2026-12-31**, it should be deleted rather than carried — a primitive with one hypothetical user and a test that exercises only itself is the AM-11 shape (a claim resting on a pair with no consumer), and this project has now paid for that shape twice. - **K15** State implements specs/GroundRules.md §1 exactly; every GR-rule is realized in `validate`/`fold` and cross-referenced by rule ID in doc comments, giving a greppable rule→code→scenario chain. - **K16** U-item defaults (GroundRules §Underdetermined) are implemented behind clearly named functions so a ground-game ruling is a localized change; their scenarios carry `provisional: true`. ## 3. Scenario runner and CLI surface - **K17** `cb-game-runtime` exposes the scenario runner as a library; a thin binary (`cb-sim`, precursor of `cb sim`) runs `scenarios/ground/*.yaml` per MetricsAndScenarios §2: setup presets, ordered actor-tagged commands, partial end-state assertions, `covers`, `rejects`, optional `state_hash`. - **K18** Benchmarks are Criterion benches driving the same scenario format at scale (MetricsAndScenarios §3); the synthetic workload mirrors the CB-RES-0001 harness shape (3 players, commit/reveal rounds) for same-shape comparison. --- ## 4. Acceptance metrics **On AM-4's retarget (2026-07-31).** *Ratified by [ADR-0004](../decisions/ADR-0004-am4-ratification.md) on 2026-07-31; the headroom argument for why these ceilings bind on future work lives there.* AM-4 originally read "≤20 transitive crates", set against boardgame.io's 120 npm packages. That target was retired for two measured reasons. First, it was unreachable without undoing this spec's own contracts: K5 (seeded ChaCha) and K7 (SHA-256) cost 12 crates between them, and the measured ladder showed nothing reached 20 except reimplementing one of those primitives — which trades an audited implementation for a scoreboard number. Second, crate count does not compare across ecosystems: Rust splits crates far more finely than npm, so the original 33-vs-120 comparison flattered us while the ≤20 target punished us, both for the same reason. Third-party source under audit is what the count was a proxy for, it is comparable across ecosystems, and it cannot be gamed by crate granularity. Splitting it by build configuration also makes the dev/shipped distinction visible, which the single number hid. (the code loop's exit condition) Per InnerLoop step 4/5: T08 iterates until every row meets its target; evidence lands in `evidence/CB-EV-0001-game-kernel.md` with no `unmeasured`. Baselines from CB-RES-0001; cited-only rows cap at parity. | ID | Metric | Baseline (CB-RES-0001) | Target | Verdict basis | |---|---|---|---|---| | AM-1 | M-D1-COV: GR-rules covered by ≥1 passing scenario | no candidate has any (observation) | **100%** of GR + U rules | measured by runner report | | AM-2 | M-D1-SPL: spec lines per rule in `games/ground` rules code (impl LOC ÷ rule count) | boardgame.io ~36 LOC for the 2-move synthetic game | ≤ 40 LOC/rule, paired with AM-1 (anti-gaming pair) | measured (tokei + rule count) | | AM-3 | Synthetic-workload definition size: LOC to express the CB-RES-0001 synthetic game on our kernel | ~36 LOC (boardgame.io, measured) | ≤ 50 LOC | measured | | AM-4a | M-D2-DEP: third-party LOC, **shipped runtime** (`--no-default-features --edges normal,no-proc-macro`) | boardgame.io: 120 npm packages / 3.9M LOC | **≤ 161,000 lines** (ADR-0008 D3, was 250,000) | measured (`make dep-weight`) | | AM-4b | M-D2-DEP: third-party LOC, **what a contributor acquires** (`--workspace --edges normal,dev`) | as above | **≤ 745,000 lines** (CB-WP-0019; was 350,000 against a graph that measured 317,021 of the real 725,258) | measured (`make dep-weight`) — see §5c | | ~~AM-4c~~ | M-D2-DEP: own source per third-party 100k lines | — | **WITHDRAWN from the acceptance table 2026-08-01 (CB-WP-0006 T04)** — retained as a reported diagnostic in `make dep-weight`; see §5a | diagnostic | | AM-5 | M-D2-BLD: clean release build of headless workspace | n/a (npm install ~seconds; not comparable) | ≤ 60 s on bnt-lap001, recorded not gated | measured | | AM-6 | M-D3-THR: applied events/s, synthetic workload, same machine | boardgame.io ~1,100–1,900 moves/s (best config, degrading) | **≥ 100,000/s** (stipulated target, ADR-0002) | measured | | AM-7 | M-D3 scaling: throughput @100k events vs @5k; and snapshot+replay of 100k events | boardgame.io 0.45–0.66× @20–40k, DNF @100k | **≥ 0.9×** (flat), replay of 100k events ≤ 5 s, hash-identical | measured (`make am7`, `make test`) | | AM-8 | Determinism invariant: N=10 same-seed replays, bit-identical hashes; HashMap-in-state deny lint clean | Rune: enforced by tooling (cited) | zero divergence, lint clean in CI | measured (`make am8`, `make check`) — see §5b | | AM-9 | M-D3-MEM: peak RSS, 100k-event synthetic run | boardgame.io ~100→232 MB @5k→40k (indicative) | ≤ 64 MB, flat with history given snapshot interval | measured (indicative label, same method) | | AM-10 | M-D4-LEAK **(withdrawn 2026-07-31, ADR-0005 §4 — no `cb-*-api` crate exists, so the population is empty; the clippy `HashMap`/`HashSet` deny that stood in for it cites K6 determinism and is now reported as AM-10′)**: foreign types in canonical-interface signatures | boardgame.io: JS-ecosystem-locked | **0** | measured (grep/deny rule) | | AM-11 | M-D4-SWAP **(met 2026-08-01 — `KernelRng` and `LogStore` each drive one shared `conformance()`; CB-WP-0006 T05)**: null + reference impls passing one conformance suite | no candidate has the pattern | RNG and log storage each have ≥2 impls (real + test/null) under one suite | measured (bool) | | AM-12 | M-D2-TOK / M-D2-CST: tokens and USD per completed task | n/a — first pass sets our own baseline | recorded per task in the evidence cost log (price sheet 2026-07-31) | recorded, not gated | ### 5a. Why AM-4c was withdrawn from the acceptance table *(CB-WP-0006 T04, 2026-08-01. Tier S — amends one row, creates no capability; chaos d4=2, no override. Follows the precedent ADR-0005 §4 set for AM-10.)* AM-4c was `reported, not targeted`, so nothing could fail and it counted against M-D1-MUT. The task was to give it a threshold or drop it. It is dropped, for a reason that a threshold cannot fix: **The ratio has no monotone better direction.** INTENT's rule is *own the semantics; assimilate the implementation*. A **rising** ratio can mean we are properly owning semantics, or that we are reimplementing things we should have assimilated. A **falling** ratio can mean good leverage, or dependency bloat and implementation leaking into our semantics. Both directions are ambiguous, and a target requires knowing which way is better. **It is also redundant.** AM-4a and AM-4b already bound the denominator (third-party LOC ceilings, ratified in ADR-0004) and AM-2 bounds own-source density per rule. AM-4c is the ratio of two quantities that are each already targeted; any threshold on it would be implied by those two or would contradict them. Measured at withdrawal: **1,426** own lines per 100k third-party (shipped runtime), **1,107** (dev toolchain). **Decided before Phase B, deliberately.** ADR-0005 predicts own-source growth from the kernel work, which will move this ratio. Setting a threshold after seeing that movement would be the retarget InnerLoop §Step 4 forbids — so the decision was taken while the number was still unaffected by the work that will change it. **M-D1-MUT keeps AM-4c in its denominator.** Withdrawing a row would otherwise improve the metric from 7/14 to 7/13 without enforcing anything — a score improved by deleting the question. Comparisons against the event-sourcing 10⁵–10⁶/s estimate stay **parity** until a local Rust comparator is measured (open follow-up from the adversarial review). ### 5c. What each AM-4 budget asks, and why they differ *(CB-WP-0019 T01/T02, 2026-08-03. Tier M.)* The two budgets had the **same scope** — one package, no dev edges — while claiming to bound different things. That left AM-4b blind to **28 crates and 408,237 lines**, more source than its own target, and it is how `quick-js` entered in CB-WP-0014 without moving the number that governs dependencies (ADR-0009 withdrew its own cost argument over it). | | the question it asks | scope | proc-macros | |---|---|---|---| | **AM-4a** | what does a game **ship**? | `-p games-ground --no-default-features` | **excluded** | | **AM-4b** | what does a contributor **acquire**? | `--workspace --edges normal,dev` | **counted** | **The proc-macro treatments are opposite on purpose.** AM-4a excludes them because they run in the compiler and never reach a shipped binary — counting them in *"what a game ships"* was simply false. AM-4b counts them, because ADR-0007 D3's acquisition rule counts what the build causes to be **fetched**, and a proc-macro is fetched, compiled and unaudited on a contributor's machine like anything else. *"It does not ship"* is no answer to *"we downloaded it"*. **When the two rules disagree, the question each budget asks decides.** That is the rule ADR-0008 D2 left open, and it is why that decision refused to reuse AM-4a's measured 36.2% share for AM-4b: the real share is **15.1%** (109,585 lines), so borrowing would have been wrong by more than a factor of two. **The target moved to fit the measurement, never the reverse.** 745,000 keeps ~2.7% of room on ADR-0008 D3's reasoning that ~1.5% fails on a dependency's patch release — the same margin AM-4a received, applied to a number that grew because the instrument was repaired rather than because anything was added. ### 5b. Where AM-8's ten runs live, and why not everywhere *(CB-WP-0015 T02, 2026-08-02. Tier S. The spec value N=10 is **not** amended — this records where it is enforced.)* For eight passes the runner executed each scenario **twice** (K8) while this row said ten, and `mutation-check` reported the count inert every run. Closing it needed an argument about what the extra runs buy, because "the spec says ten" is not one. **A deterministic divergence does not need ten runs.** A seed threaded wrong or a fold that depends on insertion order diverges on run 2 exactly as reliably as on run 10. For that class K8's double-run is sufficient and the other eight are repetitions of an answered question — 47 s per build across 25 scenarios. **A late-onset or probabilistic divergence does.** Measured, on `gr-r06-round-resolve` with the RNG perturbed only from its fourth construction onward: `--runs 2` **passes**; `--runs 10` fails with *"run 1 hash … != run 4 hash … (of 10)"*. That is a real class the double-run structurally cannot see, and it is deterministic rather than flaky, so it can be a control rather than a coin flip. So both stay, at their own costs: **K8's two runs on every scenario** (broad, cheap, `make sim`) and **ten runs on one scenario** (deep, ~2 s, `make am8`). The clause is now measured by that mutation rather than declared by its author. The primary defence against the probabilistic class remains the `HashMap`/`HashSet` deny lint under K6 — this row's other clause, already live. The ten runs are defence in depth against that exclusion failing, which is why one workload's worth is proportionate. ## 5. Out of scope for this pass Networking/session protocol, WIT/Wasm game boundary, ECS world layer, rendering, persistence beyond file snapshots, host migration, and any second game. Each arrives through its own loop pass with its own survey.