clay-borg/evidence/CB-EV-0001-game-kernel.md
tegwick 0b6f7c5bc8
Some checks failed
ci / check (push) Failing after 4s
CB-WP-0019 T01/T02: AM-4b asks what a contributor acquires
The two AM-4 budgets had the SAME scope -- one package, no dev edges --
while claiming to bound different things. AM-4b now measures the
workspace with dev edges: 57 crates / 725,258 lines where it read 29 /
317,021, having been blind to 28 crates and 408,237 lines, more source
than its own target.

Target 745,000, ~2.7% of room -- the same margin ADR-0008 D3 gave AM-4a,
applied to a number that grew because the instrument was repaired, not
because anything was added. The target moved to fit the measurement.

T02: proc-macros are COUNTED here and excluded from AM-4a, on purpose.
AM-4a asks what ships and a proc-macro never ships. AM-4b asks what is
acquired, and ADR-0007 D3's acquisition rule counts what the build
fetches -- 'it does not ship' is no answer to 'we downloaded it'. When
the rules disagree, the question each budget asks decides. Measured
share 109,585 lines / 15.1% against AM-4a's 36.2%, so ADR-0008 D2's
refusal to borrow the ratio was right by more than a factor of two.

Caught by this project's own earlier work twice: the mutation
find-string went stale and --self-test reported it BUILD-FREE (the check
CB-WP-0015 added after AM-4a's rotted for two passes), then the DFD gate
caught facts.toml carrying the old numbers.

CB-EV-0001 and ADR-0004 carried live fact: tags on historical readings.
A dated record asserting a CURRENT value is a category error, so those
occurrences are marked as-measured instead of retro-edited, and ADR-0004
gains a supersession note.

make all exits 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 19:04:54 +02:00

19 KiB
Raw Blame History

CB-EV-0001 — GROUND game kernel: acceptance evidence

Superseded measurement (ADR-0008 D2/D3, 2026-08-02). The AM-4a figures below were taken with an instrument that counted proc-macro crates — 89,048 lines, 36.2% — which run in the compiler and never reach a binary. Corrected, the same tree measures 157,202 against a target moved to 161,000. The numbers below are left as the record of what was measured then, and are no longer live facts.

Status: T08 complete. AM-4 remediated and re-measured 2026-07-31. Recorded: 2026-07-31. Amended 2026-07-31 — §4 gains measured savings per remediation option, and §5 corrects AM-12 from "uncomputable" to measured-at-session-level; see CB-WP-0002. Workplan: CB-WP-0001, task T08 Spec: specs/GameKernel.md §4 (AM-1..AM-12) Baseline: research/CB-RES-0001-game-kernel.md, measurements in research/CB-RES-0001-harness/boardgame-io/results-260731.json

Machine: WSL2, Linux 6.18.33.2-microsoft-standard-WSL2, rustc 1.97.1, --release, Criterion 1s warm-up / 3s measurement.

1. Scoreboard

Corrected 2026-07-31 (CB-WP-0005 T03, per ADR-0005 §4). Four rows below were wrong or misleading as originally committed. They are corrected in place with this note, not silently edited — the correction trail is the artifact. What changed, and why, is in §1a.

The Enforced column is new and is the point of the correction. It carries M-D1-MUT (make mutation-check): does anything actually fail when the property is false? A row can be measured and still enforce nothing.

Metric Target Measured Verdict Enforced
AM-1 rule coverage (GR only) 100% of GR-rules 58/58 (100%) met yes
AM-1b link, ground 58 claimed rules named in the aggregate 49/58 unmet reported
AM-1b link, kernel 18 K-rules named in source 15/18 unmet reported until 2026-08-31
AM-4a dep weight, shipped runtime ≤250,000 third-party lines 246,250 (23 crates) met yes
AM-4b dep weight, dev toolchain ≤350,000 third-party lines 317,021 (29 crates) met yes
AM-6 throughput ≥100,000 events/s 2,017,009 events/s (make am6) met, 20.2× yes — CB-WP-0006 T01
AM-7 scaling ≥0.9× at 20× workload 1.08× met no — no code computes the ratio
AM-7 replay, timing 100k events ≤5s 2.18 ms (CI 2.142.23) met, 2,290× yes
AM-7 replay, hash-identical bit-identical fold per-segment replay reproduces each recorded hash re-earned 2026-08-01 yes — CB-WP-0006 T06
AM-8 determinism zero divergence, 10 replays 1 distinct hash / 10 runs (one-off probe) met, narrow partial — the runner double-runs; nothing enforces N=10
AM-8 lint fmt + clippy clean clean, -D warnings met yes
AM-10 foreign types 0 in cb-*-api signatures WITHDRAWN no — no such crate; population empty
AM-10 determinism lint (K6) zero HashMap/HashSet in game state 0 met yes
AM-11 impl pairs ≥2 impls under one conformance suite 2 ports, 2 impls each, one shared suite per port met 2026-08-01 yes — break one impl and the shared suite fails
AM-2 LOC per rule ≤40 27.2 (1,575 impl lines / 58 rules) met yes — CB-WP-0006 T02
AM-3 synthetic workload LOC ≤50 blocked — the artifact has never been built no
AM-5 clean release build ≤60 s on bnt-lap001 37.3 s dev / 41.2 s shipped, best of 3, quiet met, 1.6× no — spec declares it ungated
AM-9 peak RSS ≤64 MB 13.4 MB met, 4.8× yes — CB-WP-0006 T03
AM-12 cost per-task USD $93.15 pinned, per task via make cost met yes

M-D1-MUT over the whole acceptance table: 8 of 14 rows enforced (make mutation-check), up from 4 when the instrument was first run. AM-4c is retained in that denominator after its withdrawal, deliberately: a score improved by deleting the question is not an improvement.

Of the six not enforced: AM-3 is blocked on an artifact that was never built; AM-4c and AM-5 cannot fail because the spec declares them untargeted/ungated; AM-10 was withdrawn; AM-7 and AM-8 are partial — some clauses live, some inert. §6's note that AM-2/AM-5/AM-9 are "not reported" was true until CB-WP-0006 and is superseded by the rows above.

1a. What was corrected, and why

row as committed corrected to found by
AM-7 replay met, 2,290× split: timing met, hash-identical withdrawnre-earned 2026-08-01, CB-WP-0006 T06 mutation — folding the log from fresh(999) instead of fresh(42), an unrelated genesis state, leaves the test green. The hash reaches only a println!. It could not be asserted as written anyway: the log spans games seeded 42, 43, 44… so it is not a replay of anything
AM-10 met withdrawn as written; restated as AM-10 adversarial review — there is no cb-*-api crate, so 0 foreign types was true over an empty set. What was actually measured is a clippy.toml deny of HashMap/HashSet whose stated reason cites K6 (determinism), reported under a D4 leak row
AM-11 met, narrow unmet adversarial review — M-D4-SWAP is a bool over "the same conformance suite"; grep -rn conformance over every .rs returns one doc comment describing future work, and the RNG pair is exercised by two separate, non-shared tests
AM-12 $248.46 session $93.15 CB-WP-0002 re-derivation — the original figure double-counted per-line and priced a three-model session at one model's rate. It had been stale in this file since, untagged and therefore invisible to facts-check
AM-1b absent from the scoreboard added, both denominators adversarial review — make coverage printed 49/58 two lines below the 100% this table carried, and the table kept the flattering half

Nothing here was found by re-running the command that printed the number. Three of the five were found by an adversarial reviewer reading the assertion behind the number, and two of those required a mutation to settle. That is the loop change recorded in CB-WP-0005 T08.

2. Throughput and scaling (AM-6, AM-7)

Workload: 3-player GROUND rounds, 7 commands and 13 events per round (12 on a game's fifth round, where GR-R09 ends the game). Pinned by bench_shape in games/ground/src/lib.rs, so a change to the workload breaks the test rather than silently rescaling the metric.

Rounds Throughput (events/s) vs 5k
5,000 1,524,200 1.00×
10,000 1,638,200 1.07×
20,000 1,577,000 1.03×
40,000 1,626,000 1.07×
100,000 1,651,400 1.08×

Replay — folding one growing event log back into state (Criterion 95% CI, low/median/high):

Events Time (median) 95% CI Rate
10,010 184.8 µs 180.9 189.3 µs 54.2M events/s
100,072 2.183 ms 2.142 2.226 ms 45.8M events/s

Correction (2026-07-31). These originally read 465 µs / 4.13 ms and were taken from the replay_probe test, not from the benchmark — because the benchmark did not work. See §2a.

The comparison against boardgame.io, stated carefully

boardgame.io measured 1,930 moves/s at 5,000 moves, falling to 870 moves/s at 20,000, and did not finish 100,000 moves in 300 s. Our figure in the same unit is ~129,000 rounds/s × 7 = ~903,000 commands/s, and 100,000 rounds complete in 776 ms.

That is roughly a 400500× ratio, and it is not a like-for-like measurement. Four differences matter, all favouring us:

  1. Different language and process model. Rust in-process against Node.js with immutable state, patch generation and undo history.
  2. Different feature set. boardgame.io's per-move cost includes producing client patches and maintaining undo state; the run with --disable-undo still degraded (0.66× at 40k). We do neither.
  3. Different workload shape. Our 5-round games (GR-R09) bound state size by construction. The boardgame.io harness ran one match with unbounded history, which is exactly the axis it degraded on.
  4. No network or storage layer on our side.

Point 3 is the important one and it is why the flat curve in the table above is weak evidence on its own — a game that resets every five rounds cannot exhibit history-growth degradation. The replay benchmark is the honest test of that axis, because there the log grows without bound, and it stays linear (21.5M → 24.2M events/s from 10k to 100k).

Claim we are willing to defend: the kernel meets AM-6 and AM-7 with large margin, and does not degrade as event-log length grows. Claim we are not making: that Clay-Borg is ~450× "faster than boardgame.io" as a like-for-like engine comparison. Per the InnerLoop parity-cap rule, a cross-runtime ratio this coarse is not a verdict, it is a direction.

A measurement error found and corrected

The first run of this benchmark reported 9.3M events/s with a perfectly flat curve — a number that would have been reported as a 93× beat of AM-6. It was wrong. The workload had P2 selecting SUPPORT every round while being attacked; after three rounds P2 sat at Stress 4, GR-R03 rejected the SUPPORT, the round never completed, and the loop spun on rejected commands. Throughput was computed as rounds × 13 events while most rounds produced 2.

Found by a probe test asserting that a round produces events at all. The benchmark now asserts the per-round event count on every round and panics rather than measuring a stalled loop. The corrected figure is 5.6× lower than the bogus one.

2a. A fourth measurement error, found by enforcing the rule

The replay benchmark committed alongside this evidence was the broken version. A python3 patch that was supposed to replace its log-building loop never applied, leaving a sequence that omits Resolve — so EndRound was rejected, every round produced no events, and the while log.len() < target loop spun forever. It was never run to completion; the AM-7 replay numbers were taken from a separate probe test instead, and the dead benchmark was committed and left hanging.

Found by adding cargo bench -- --test to CI, which runs every benchmark once. That is the fourth instance of one error class in this project — a harness that appears to work while doing no work — and the first one caught by a gate rather than by noticing.

The replay loop now carries the positive control the round loop already had: it asserts each round appended events and fails rather than spinning. Corrected figures are in the table above; both configurations still clear the AM-7 budget by three orders of magnitude.

The lesson recorded for the loop: writing the positive-control rule into specs/InnerLoop.md did not prevent the next instance. Making it a CI step did. Prose rules do not enforce themselves.

3. Determinism (AM-8)

  • Every scenario runs twice per invocation with the same seed and fails on state-hash divergence (K8). 21/21 pass.
  • Ten consecutive full runs of all 21 scenarios produced one distinct output hash, i.e. zero divergence.
  • cargo fmt --check and cargo clippy --workspace --all-targets -D warnings are clean.
  • clippy.toml denies HashMap/HashSet workspace-wide (K6); the aggregate holds only ordered collections, so iteration order cannot vary between runs.

4. AM-4 — remediated and re-measured

Original result: NOT MET, 33 transitive crates against a ≤20 target. Resolved by adopting both remediations (maintainer decision, 2026-07-31): serde_yaml was made optional, and the metric was retargeted onto third-party source under audit.

Re-measurement (make dep-weight)

Configuration Crates Third-party LOC Target Verdict
Shipped runtime (--no-default-features) 23 246,250 ≤250,000 met
Dev toolchain (default features) 29 317,021 ≤350,000 met
Our own source 3,408

Scenario tooling costs 70,771 lines that a shipped game never compiles. That split is the substantive result: the single number previously reported conflated a runtime concern with a test concern.

What actually changed in the build. cb-game-runtime gained a scenarios feature carrying serde_yaml; the scenario module, the ScenarioGame impl and the string parsers behind it are #[cfg]-gated. Both configurations compile and lint clean under -D warnings.

One trap worth recording: setting default-features = false on a member dependency is silently ignored when the workspace dependency does not specify it, so the first attempt gated nothing while appearing to work — cargo tree still showed all six YAML crates. The fix was setting default-features = false on the workspace dependency itself, with cb-sim opting into scenarios explicitly. This is exactly the class of error InnerLoop v1.0's positive-control rule targets: the build succeeded and the feature flag looked applied. It was caught by checking the dependency graph rather than trusting that the edit had worked.

Why the target moved, and why that is not moving the goalposts

The ≤20 crate target was retired for two measured reasons, both recorded before the decision was taken:

Attribution:

Group Crates Count
sha2 (K7 state hashing) sha2, digest, block-buffer, crypto-common, generic-array, typenum, cpufeatures, cfg-if 8
serde derive chain serde_derive, proc-macro2, quote, syn, unicode-ident 5
serde runtime + json serde, serde_core, serde_json, itoa, memchr, zmij 6
serde_yaml (scenario files only) serde_yaml, unsafe-libyaml, indexmap, hashbrown, equivalent, ryu 6
rand_chacha (K5 seeded RNG) rand_chacha, rand_core, ppv-lite86, zerocopy 4
Clay-Borg crates cb-kernel, cb-events, cb-game-runtime, games-ground 4
  1. It was unreachable without undoing the spec's own contracts. Measured ladder: serde_yaml optional 6 (→27), dropping serde_json 4 (→23), inlining SHA-256 8 (→19), inlining ChaCha12 4 (→15). Nothing reaches 20 except reimplementing a primitive that K5 or K7 requires — trading an audited implementation for a scoreboard number.
  2. Crate count does not compare across ecosystems. Rust splits crates far more finely than npm. The same granularity difference made "33 vs 120 npm packages" flatter us and made ≤20 punish us.

Third-party source under audit is what the count was proxying for, is comparable across ecosystems, and cannot be gamed by granularity. The new targets are set at roughly the current measurement plus headroom, so they bind on future growth rather than retroactively passing something that failed: adding another serde_yaml-sized dependency to the shipped runtime would breach AM-4a.

What we did not do: hand-roll SHA-256 or ChaCha to win a count.

5. Cost log (AM-12)

Per specs/MetricsAndScenarios.md §1a. Model: Claude Fable 5, at benchmarks/baselines/model-prices.toml rates ($10/$50 per MTok).

Task Model Iterations Notes
CB-WP-0001 (whole session, T01T09) claude-fable-5 6 T08 code iterations + benchmarks $248.46 measured; per-task split pending CB-WP-0002

Correction (2026-07-31). This section originally recorded AM-12 as uncomputable. That was wrong. The declining to estimate was right; the conclusion that no instrument existed was not. Every session transcript (~/.claude/projects/<slug>/<session>.jsonl) carries exact per-message usage including the cache breakdown. Read for this session:

Component Tokens Cost (Fable 5)
Output 585,528 $29.28
Cache read 131,863,164 $131.86
Cache write (1h) 4,365,668 $87.31
Input 1,090 $0.01
Session total $248.46 (~$124 on Opus 5)

53% of the cost is cache reads, not output. Cost in an agentic loop is driven by context size × turn count, which no "tokens per task" metric would have surfaced.

Still missing is attribution: this is a whole-session figure, not a per-task one, because nothing marks task boundaries in the transcript. That is what CB-WP-0002 is for. The AM-12 row above should be read as "session-level cost measured; per-task attribution pending CB-WP-0002".

6. Metrics not reported

  • AM-2, AM-3, AM-5 — specification-quality metrics that need a second capability to compare against; a single data point is not a measurement.
  • AM-9 (≤64MB) — not instrumented. The aggregate is a few KB and the largest log measured here is 100k events, so the budget is very unlikely to bind, but "unlikely" is not "measured" and it is left unclaimed.
  • AM-11the KernelRng null/reference pair exists and is exercised. It is the only port with a pair so far, so the metric is met narrowly and will mean more once storage has one. Corrected 2026-07-31: the pair exists; the conformance suite does not. Resolved 2026-08-01 (CB-WP-0006 T05): cb_kernel::rng::conformance drives ChaChaRng and NullRng; cb_events::store::conformance drives MemLogStore and FileLogStore. One suite per port, both impls through it. AM-11 is met and mutation-verified — breaking NullRng::draw to return its bound fails the shared suite. The original text follows. M-D4-SWAP is a bool over impls "passing the same conformance suite", and the two impls are exercised by two separate, non-shared tests — there is no fn conformance<R: KernelRng>(…) that both are driven through. grep -rn conformance over every .rs returns a single doc comment describing future work. AM-11 is unmet, and was never earned. Building the suite is deferred to the re-planned pass, not to this one.

7. Rules implemented under a provisional default

Ten U-items in specs/GroundRules.md carry PROVISIONAL defaults. Those realized here are U2 (clamp on every application), U3 (DENY with no legal target is a no-op that still advances), U4 (deck reshuffle), U5 (REVERSE owner relief applies whether or not the Reverse was rejected) and U8 (GROUND—OU cancellation precedes Protection).

One further ambiguity was found during T08 and is not in the U-list: GR-E02's "successes" is undefined in dataset 0.1. It is implemented as the count of claimed Problems. Both scoring scenarios are marked provisional: true.

All provisional behaviour lives behind named functions and is covered by scenarios tagged provisional: true, so a ground-game ruling flips a scenario rather than the kernel (K16). Action for ground-game: rule on the ten U-items and on GR-E02's "successes".