clay-borg/evidence/CB-EV-0001-game-kernel.md
tegwick 8e11fc412e
Some checks failed
ci / check (push) Failing after 4s
Amend CB-EV-0001; add CB-WP-0002 for cost accounting
AM-4: measured what each remediation option actually buys, rather than
leaving one recommendation unquantified. serde_yaml optional removes 6
crates, not 5 — ryu belongs to that group, since serde_json now uses
zmij for floats. Full ladder: -6 to 27, serde_json -4 more to 23,
inlining SHA-256 -8 to 19, inlining ChaCha12 -4 to 15. Only
reimplementing a primitive gets under 20, so the target is unreachable
without undoing K5/K7.

Also records that crate count compares badly across ecosystems, and
offers the alternative the count is a proxy for: 307,317 lines of
third-party source under audit against 3,398 of our own.

AM-12: corrected from "uncomputable" to measured. The refusal to
estimate was right; the claim that no instrument existed was wrong.
Session transcripts carry exact per-message usage including the cache
breakdown. This session cost $248.46 at Fable 5 rates, of which 53% is
cache reads — cost is driven by context size times turn count, not by
output volume. What is still missing is per-task attribution, since
nothing marks task boundaries in a transcript.

CB-WP-0002 makes cost measurable and attributable: survey the
instruments, decide the attribution model by ADR, spec metrics that
include cost composition rather than a bare total, build a collector
whose positive control refuses to emit unreconciled numbers, and prove
it by answering a question that could not be answered before.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 03:26:03 +02:00

10 KiB
Raw Blame History

CB-EV-0001 — GROUND game kernel: acceptance evidence

Status: T08 complete, with one acceptance metric not met (AM-4). Recorded: 2026-07-31. Amended 2026-07-31 — §4 gains measured savings per remediation option, and §5 corrects AM-12 from "uncomputable" to measured-at-session-level; see CB-WP-0002. Workplan: CB-WP-0001, task T08 Spec: specs/GameKernel.md §4 (AM-1..AM-12) Baseline: research/CB-RES-0001-game-kernel.md, measurements in research/CB-RES-0001-harness/boardgame-io/results-260731.json

Machine: WSL2, Linux 6.18.33.2-microsoft-standard-WSL2, rustc 1.97.1, --release, Criterion 1s warm-up / 3s measurement.

1. Scoreboard

Metric Target Measured Verdict
AM-1 rule coverage 100% of GR-rules 58/58 (100%) met
AM-4 dependency weight ≤20 crates 33 NOT MET
AM-6 throughput ≥100,000 events/s 1,651,400 events/s met, 16.5×
AM-7 scaling ≥0.9× at 20× workload 1.08× met
AM-7 replay 100k events ≤5s 4.13 ms met, 1,210×
AM-8 determinism zero divergence, 10 replays 1 distinct hash / 10 runs met
AM-8 lint fmt + clippy clean clean, -D warnings met
AM-10 foreign types zero HashMap/HashSet 0 met
AM-11 impl pairs null + reference per port 1 of 1 (KernelRng) met, narrow
AM-12 cost per-task USD $248.46 session; per-task pending partial

AM-2, AM-3, AM-5 and AM-9 are not reported: see §6.

2. Throughput and scaling (AM-6, AM-7)

Workload: 3-player GROUND rounds, 7 commands and 13 events per round (12 on a game's fifth round, where GR-R09 ends the game). Pinned by bench_shape in games/ground/src/lib.rs, so a change to the workload breaks the test rather than silently rescaling the metric.

Rounds Throughput (events/s) vs 5k
5,000 1,524,200 1.00×
10,000 1,638,200 1.07×
20,000 1,577,000 1.03×
40,000 1,626,000 1.07×
100,000 1,651,400 1.08×

Replay — folding one growing event log back into state:

Events Time Rate
10,010 465 µs 21.5M events/s
100,007 4.13 ms 24.2M events/s

The comparison against boardgame.io, stated carefully

boardgame.io measured 1,930 moves/s at 5,000 moves, falling to 870 moves/s at 20,000, and did not finish 100,000 moves in 300 s. Our figure in the same unit is ~129,000 rounds/s × 7 = ~903,000 commands/s, and 100,000 rounds complete in 776 ms.

That is roughly a 400500× ratio, and it is not a like-for-like measurement. Four differences matter, all favouring us:

  1. Different language and process model. Rust in-process against Node.js with immutable state, patch generation and undo history.
  2. Different feature set. boardgame.io's per-move cost includes producing client patches and maintaining undo state; the run with --disable-undo still degraded (0.66× at 40k). We do neither.
  3. Different workload shape. Our 5-round games (GR-R09) bound state size by construction. The boardgame.io harness ran one match with unbounded history, which is exactly the axis it degraded on.
  4. No network or storage layer on our side.

Point 3 is the important one and it is why the flat curve in the table above is weak evidence on its own — a game that resets every five rounds cannot exhibit history-growth degradation. The replay benchmark is the honest test of that axis, because there the log grows without bound, and it stays linear (21.5M → 24.2M events/s from 10k to 100k).

Claim we are willing to defend: the kernel meets AM-6 and AM-7 with large margin, and does not degrade as event-log length grows. Claim we are not making: that Clay-Borg is ~450× "faster than boardgame.io" as a like-for-like engine comparison. Per the InnerLoop parity-cap rule, a cross-runtime ratio this coarse is not a verdict, it is a direction.

A measurement error found and corrected

The first run of this benchmark reported 9.3M events/s with a perfectly flat curve — a number that would have been reported as a 93× beat of AM-6. It was wrong. The workload had P2 selecting SUPPORT every round while being attacked; after three rounds P2 sat at Stress 4, GR-R03 rejected the SUPPORT, the round never completed, and the loop spun on rejected commands. Throughput was computed as rounds × 13 events while most rounds produced 2.

Found by a probe test asserting that a round produces events at all. The benchmark now asserts the per-round event count on every round and panics rather than measuring a stalled loop. The corrected figure is 5.6× lower than the bogus one.

3. Determinism (AM-8)

  • Every scenario runs twice per invocation with the same seed and fails on state-hash divergence (K8). 21/21 pass.
  • Ten consecutive full runs of all 21 scenarios produced one distinct output hash, i.e. zero divergence.
  • cargo fmt --check and cargo clippy --workspace --all-targets -D warnings are clean.
  • clippy.toml denies HashMap/HashSet workspace-wide (K6); the aggregate holds only ordered collections, so iteration order cannot vary between runs.

4. AM-4 — not met, and why it is reported rather than fixed

33 transitive crates against a ≤20 target. The baseline it was set against is boardgame.io's 120 npm packages, so we are 3.6× lighter, but the metric as written is missed and is recorded as missed.

Attribution:

Group Crates Count
sha2 (K7 state hashing) sha2, digest, block-buffer, crypto-common, generic-array, typenum, cpufeatures, cfg-if 8
serde derive chain serde_derive, proc-macro2, quote, syn, unicode-ident 5
serde runtime + json serde, serde_core, serde_json, itoa, memchr, zmij 6
serde_yaml (scenario files only) serde_yaml, unsafe-libyaml, indexmap, hashbrown, equivalent, ryu 6
rand_chacha (K5 seeded RNG) rand_chacha, rand_core, ppv-lite86, zerocopy 4
Clay-Borg crates cb-kernel, cb-events, cb-game-runtime, games-ground 4

The honest options, in order of preference:

  1. Make serde_yaml optional behind a scenarios feature. YAML is a test-and-tooling concern; a shipped game runtime does not need it. Removes 6 crates from the default build for no loss of capability (ryu belongs to this group, not to serde_json, which uses zmij for floats — corrected after measuring the reverse-dependency graph). This is the one to do first, and it improves D4 optionality as well as D2.

  2. Revisit the target, and what it measures. Measured savings per option: serde_yaml optional 6 (→27); replacing serde_json 4 more (→23); inlining SHA-256 8 (→19); inlining ChaCha12 4 (→15). Only reimplementing SHA-256 or ChaCha gets under 20, so the target is unreachable without undoing K5/K7.

    Crate count also compares badly across ecosystems: Rust splits crates far more finely than npm, so "33 vs 120 npm packages" flatters us. The measurable thing crate count proxies for is third-party source under audit: 307,317 lines across all five groups, against 3,398 of our own. Retargeting AM-4 on audited third-party LOC, split into shipped-runtime and dev-toolchain, measures the real concern and cannot be gamed by crate granularity.

What we are not doing: hand-rolling SHA-256 or ChaCha to win a dependency count. That trades an auditable, well-tested primitive for a number on a scoreboard.

Carried forward as an open decision (see the note at the head of this file): AM-4 is re-measured once the option is chosen.

5. Cost log (AM-12)

Per specs/MetricsAndScenarios.md §1a. Model: Claude Fable 5, at benchmarks/baselines/model-prices.toml rates ($10/$50 per MTok).

Task Model Iterations Notes
CB-WP-0001 (whole session, T01T09) claude-fable-5 6 T08 code iterations + benchmarks $248.46 measured; per-task split pending CB-WP-0002

Correction (2026-07-31). This section originally recorded AM-12 as uncomputable. That was wrong. The declining to estimate was right; the conclusion that no instrument existed was not. Every session transcript (~/.claude/projects/<slug>/<session>.jsonl) carries exact per-message usage including the cache breakdown. Read for this session:

Component Tokens Cost (Fable 5)
Output 585,528 $29.28
Cache read 131,863,164 $131.86
Cache write (1h) 4,365,668 $87.31
Input 1,090 $0.01
Session total $248.46 (~$124 on Opus 5)

53% of the cost is cache reads, not output. Cost in an agentic loop is driven by context size × turn count, which no "tokens per task" metric would have surfaced.

Still missing is attribution: this is a whole-session figure, not a per-task one, because nothing marks task boundaries in the transcript. That is what CB-WP-0002 is for. The AM-12 row above should be read as "session-level cost measured; per-task attribution pending CB-WP-0002".

6. Metrics not reported

  • AM-2, AM-3, AM-5 — specification-quality metrics that need a second capability to compare against; a single data point is not a measurement.
  • AM-9 (≤64MB) — not instrumented. The aggregate is a few KB and the largest log measured here is 100k events, so the budget is very unlikely to bind, but "unlikely" is not "measured" and it is left unclaimed.
  • AM-11 — the KernelRng null/reference pair exists and is exercised. It is the only port with a pair so far, so the metric is met narrowly and will mean more once storage has one.

7. Rules implemented under a provisional default

Ten U-items in specs/GroundRules.md carry PROVISIONAL defaults. Those realized here are U2 (clamp on every application), U3 (DENY with no legal target is a no-op that still advances), U4 (deck reshuffle), U5 (REVERSE owner relief applies whether or not the Reverse was rejected) and U8 (GROUND—OU cancellation precedes Protection).

One further ambiguity was found during T08 and is not in the U-list: GR-E02's "successes" is undefined in dataset 0.1. It is implemented as the count of claimed Problems. Both scoring scenarios are marked provisional: true.

All provisional behaviour lives behind named functions and is covered by scenarios tagged provisional: true, so a ground-game ruling flips a scenario rather than the kernel (K16). Action for ground-game: rule on the ten U-items and on GR-E02's "successes".