INTENT design decision 8 of 10, unimplemented for six passes. cb-sim had no flag parsing at all, so --replay had nowhere to go. The bundle is manifest + commands.log + initial.snapshot + expected.yaml, dev-only behind the scenarios feature and charged to AM-4b. The command stream goes through the K11 framing built in T05, so a truncated bundle is detected rather than replayed short — the two tasks compose rather than duplicating. The reviewer's D2 correction was real: this was not "a directory of four files". Pass carried only the end state, RunOutcome::Failed was a formatted String, and scenario.rs created an EventLog, appended to it and never read it. All three had to change. The first round trip failed to reproduce, and the cause is worth keeping: state_hash_hex over a serde_json::Value is a different canonical form than over the typed aggregate — Value's map is key-sorted, a struct serializes in declaration order. The bundle was written with one basis and verified with the other. A round trip written to recompute its own comparison value would have PASSED this bug; it failed because the recorded hash came from the producing process, which is control 2's entire purpose. make replay-test implements ADR-0005 §6's four controls, 14/14: a committed deliberately-failing fixture outside the corpus with covers: [] so it neither fails `make sim` nor inflates AM-1; a tampered recorded hash must fail; a log short by one byte and a corrupted length prefix must be rejected; and a mutated manifest seed must fail — which bites only because replay re-derives the initial state from seed+setup and checks it against the recorded snapshot, since restoring from the snapshot alone would leave the seed inert. Plus a control on the controls: the bundle must still replay after every mutation is reverted. AM-7's hash-identical clause is re-earned. The probe records a hash per per-game segment and replays each from its own genesis; folding from the wrong seed now fails. That is the clause ADR-0005 §4 withdrew as mutation-proven inert. The scaling >= 0.9x clause is still unenforced, so AM-7 stays PARTIAL — reported, not rounded up. Kernel coverage 15/18 -> 16/18. facts-check immediately caught the spec's copy of that number going stale, on a number that moved the same hour. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
329 lines
18 KiB
Markdown
329 lines
18 KiB
Markdown
# CB-EV-0001 — GROUND game kernel: acceptance evidence
|
||
|
||
Status: **T08 complete. AM-4 remediated and re-measured 2026-07-31.**
|
||
Recorded: 2026-07-31. Amended 2026-07-31 — §4 gains measured savings per
|
||
remediation option, and §5 corrects AM-12 from "uncomputable" to
|
||
measured-at-session-level; see CB-WP-0002.
|
||
Workplan: CB-WP-0001, task T08
|
||
Spec: `specs/GameKernel.md` §4 (AM-1..AM-12)
|
||
Baseline: `research/CB-RES-0001-game-kernel.md`, measurements in
|
||
`research/CB-RES-0001-harness/boardgame-io/results-260731.json`
|
||
|
||
Machine: WSL2, Linux 6.18.33.2-microsoft-standard-WSL2, rustc 1.97.1,
|
||
`--release`, Criterion 1s warm-up / 3s measurement.
|
||
|
||
## 1. Scoreboard
|
||
|
||
> **Corrected 2026-07-31 (CB-WP-0005 T03, per ADR-0005 §4).** Four rows
|
||
> below were wrong or misleading as originally committed. They are
|
||
> corrected **in place with this note**, not silently edited — the
|
||
> correction trail is the artifact. What changed, and why, is in §1a.
|
||
|
||
The **Enforced** column is new and is the point of the correction. It
|
||
carries M-D1-MUT (`make mutation-check`): does anything actually fail when
|
||
the property is false? A row can be *measured* and still enforce nothing.
|
||
|
||
| Metric | Target | Measured | Verdict | Enforced |
|
||
|---|---|---|---|---|
|
||
| AM-1 rule coverage (GR only) | 100% of GR-rules | 58/58 (100%) | **met** | **yes** |
|
||
| AM-1b link, ground | 58 claimed rules named in the aggregate | 49/58 | **unmet** | reported |
|
||
| AM-1b link, kernel | 18 K-rules named in source | 15/18 | **unmet** | reported until 2026-08-31 |
|
||
| AM-4a dep weight, shipped runtime | ≤250,000 third-party lines | 246,250 (23 crates) | **met** | **yes** | <!-- fact:am4a_loc -->
|
||
| AM-4b dep weight, dev toolchain | ≤350,000 third-party lines | 317,021 (29 crates) | **met** | **yes** | <!-- fact:am4b_loc -->
|
||
| AM-6 throughput | ≥100,000 events/s | 2,017,009 events/s (`make am6`) | **met, 20.2×** | **yes** — CB-WP-0006 T01 |
|
||
| AM-7 scaling | ≥0.9× at 20× workload | 1.08× | **met** | **no** — no code computes the ratio |
|
||
| AM-7 replay, timing | 100k events ≤5s | 2.18 ms (CI 2.14–2.23) | **met, 2,290×** | **yes** |
|
||
| AM-7 replay, hash-identical | bit-identical fold | per-segment replay reproduces each recorded hash | **re-earned 2026-08-01** | **yes** — CB-WP-0006 T06 |
|
||
| AM-8 determinism | zero divergence, 10 replays | 1 distinct hash / 10 runs (one-off probe) | **met, narrow** | **partial** — the runner double-runs; nothing enforces N=10 |
|
||
| AM-8 lint | fmt + clippy clean | clean, `-D warnings` | **met** | **yes** |
|
||
| AM-10 foreign types | 0 in `cb-*-api` signatures | — | **WITHDRAWN** | **no** — no such crate; population empty |
|
||
| AM-10′ determinism lint (K6) | zero `HashMap`/`HashSet` in game state | 0 | **met** | **yes** |
|
||
| AM-11 impl pairs | ≥2 impls under **one conformance suite** | 2 ports, 2 impls each, one shared suite per port | **met 2026-08-01** | **yes** — break one impl and the shared suite fails |
|
||
| AM-2 LOC per rule | ≤40 | 27.2 (1,575 impl lines / 58 rules) | **met** | **yes** — CB-WP-0006 T02 |
|
||
| AM-3 synthetic workload LOC | ≤50 | — | **blocked** — the artifact has never been built | **no** |
|
||
| AM-5 clean release build | ≤60 s on bnt-lap001 | 37.3 s dev / 41.2 s shipped, best of 3, quiet | **met, 1.6×** | **no** — spec declares it ungated |
|
||
| AM-9 peak RSS | ≤64 MB | 13.4 MB | **met, 4.8×** | **yes** — CB-WP-0006 T03 |
|
||
| AM-12 cost | per-task USD | **$93.15** pinned, per task via `make cost` | **met** | **yes** | <!-- fact:pinned_total -->
|
||
|
||
**M-D1-MUT over the whole acceptance table: 8 of 14 rows enforced**
|
||
(`make mutation-check`), up from 4 when the instrument was first run.
|
||
AM-4c is retained in that denominator after its withdrawal, deliberately:
|
||
a score improved by deleting the question is not an improvement.
|
||
|
||
Of the six not enforced: **AM-3** is blocked on an artifact that was never
|
||
built; **AM-4c** and **AM-5** cannot fail because the spec declares them
|
||
untargeted/ungated; **AM-10** was withdrawn; **AM-7** and **AM-8** are
|
||
partial — some clauses live, some inert. §6's note that AM-2/AM-5/AM-9 are
|
||
"not reported" was true until CB-WP-0006 and is superseded by the rows
|
||
above.
|
||
|
||
### 1a. What was corrected, and why
|
||
|
||
| row | as committed | corrected to | found by |
|
||
|---|---|---|---|
|
||
| **AM-7 replay** | `met, 2,290×` | split: timing **met**, `hash-identical` **withdrawn** — *re-earned 2026-08-01, CB-WP-0006 T06* | mutation — folding the log from `fresh(999)` instead of `fresh(42)`, an unrelated genesis state, leaves the test green. The hash reaches only a `println!`. It could not be asserted as written anyway: the log spans games seeded 42, 43, 44… so it is not a replay of anything |
|
||
| **AM-10** | `met` | **withdrawn as written**; restated as AM-10′ | adversarial review — there is no `cb-*-api` crate, so `0 foreign types` was true over an empty set. What was actually measured is a `clippy.toml` deny of `HashMap`/`HashSet` whose stated reason cites **K6 (determinism)**, reported under a **D4 leak** row |
|
||
| **AM-11** | `met, narrow` | **unmet** | adversarial review — M-D4-SWAP is a bool over "the same conformance suite"; `grep -rn conformance` over every `.rs` returns one doc comment describing future work, and the RNG pair is exercised by two separate, non-shared tests |
|
||
| **AM-12** | `$248.46 session` | **$93.15** | CB-WP-0002 re-derivation — the original figure double-counted per-line and priced a three-model session at one model's rate. It had been stale in this file since, untagged and therefore invisible to `facts-check` |
|
||
| **AM-1b** | *absent from the scoreboard* | **added, both denominators** | adversarial review — `make coverage` printed `49/58` two lines below the `100%` this table carried, and the table kept the flattering half |
|
||
|
||
Nothing here was found by re-running the command that printed the number.
|
||
Three of the five were found by an adversarial reviewer reading the
|
||
assertion behind the number, and two of those required a mutation to
|
||
settle. That is the loop change recorded in CB-WP-0005 T08.
|
||
|
||
## 2. Throughput and scaling (AM-6, AM-7)
|
||
|
||
Workload: 3-player GROUND rounds, 7 commands and 13 events per round
|
||
(12 on a game's fifth round, where GR-R09 ends the game). Pinned by
|
||
`bench_shape` in `games/ground/src/lib.rs`, so a change to the workload
|
||
breaks the test rather than silently rescaling the metric.
|
||
|
||
| Rounds | Throughput (events/s) | vs 5k |
|
||
|---|---|---|
|
||
| 5,000 | 1,524,200 | 1.00× |
|
||
| 10,000 | 1,638,200 | 1.07× |
|
||
| 20,000 | 1,577,000 | 1.03× |
|
||
| 40,000 | 1,626,000 | 1.07× |
|
||
| 100,000 | 1,651,400 | 1.08× |
|
||
|
||
Replay — folding one growing event log back into state (Criterion
|
||
95% CI, low/median/high):
|
||
|
||
| Events | Time (median) | 95% CI | Rate |
|
||
|---|---|---|---|
|
||
| 10,010 | 184.8 µs | 180.9 – 189.3 µs | 54.2M events/s |
|
||
| 100,072 | 2.183 ms | 2.142 – 2.226 ms | 45.8M events/s |
|
||
|
||
**Correction (2026-07-31).** These originally read 465 µs / 4.13 ms and
|
||
were taken from the `replay_probe` **test**, not from the benchmark —
|
||
because the benchmark did not work. See §2a.
|
||
|
||
### The comparison against boardgame.io, stated carefully
|
||
|
||
boardgame.io measured **1,930 moves/s at 5,000 moves**, falling to
|
||
**870 moves/s at 20,000**, and **did not finish 100,000 moves in 300 s**.
|
||
Our figure in the same unit is ~129,000 rounds/s × 7 = **~903,000
|
||
commands/s**, and 100,000 rounds complete in 776 ms.
|
||
|
||
That is roughly a 400–500× ratio, and it is **not a like-for-like
|
||
measurement**. Four differences matter, all favouring us:
|
||
|
||
1. **Different language and process model.** Rust in-process against
|
||
Node.js with immutable state, patch generation and undo history.
|
||
2. **Different feature set.** boardgame.io's per-move cost includes
|
||
producing client patches and maintaining undo state; the run with
|
||
`--disable-undo` still degraded (0.66× at 40k). We do neither.
|
||
3. **Different workload shape.** Our 5-round games (GR-R09) bound state
|
||
size by construction. The boardgame.io harness ran one match with
|
||
unbounded history, which is exactly the axis it degraded on.
|
||
4. **No network or storage layer** on our side.
|
||
|
||
Point 3 is the important one and it is why the flat curve in the table
|
||
above is *weak evidence on its own* — a game that resets every five
|
||
rounds cannot exhibit history-growth degradation. The replay benchmark
|
||
is the honest test of that axis, because there the log grows without
|
||
bound, and it stays linear (21.5M → 24.2M events/s from 10k to 100k).
|
||
|
||
**Claim we are willing to defend:** the kernel meets AM-6 and AM-7 with
|
||
large margin, and does not degrade as event-log length grows.
|
||
**Claim we are not making:** that Clay-Borg is ~450× "faster than
|
||
boardgame.io" as a like-for-like engine comparison. Per the
|
||
InnerLoop parity-cap rule, a cross-runtime ratio this coarse is not a
|
||
verdict, it is a direction.
|
||
|
||
### A measurement error found and corrected
|
||
|
||
The first run of this benchmark reported **9.3M events/s with a
|
||
perfectly flat curve** — a number that would have been reported as a
|
||
93× beat of AM-6. It was wrong. The workload had P2 selecting SUPPORT
|
||
every round while being attacked; after three rounds P2 sat at Stress 4,
|
||
GR-R03 rejected the SUPPORT, the round never completed, and the loop
|
||
spun on rejected commands. Throughput was computed as
|
||
`rounds × 13 events` while most rounds produced 2.
|
||
|
||
Found by a probe test asserting that a round produces events at all.
|
||
The benchmark now asserts the per-round event count on every round and
|
||
panics rather than measuring a stalled loop. The corrected figure is
|
||
**5.6× lower** than the bogus one.
|
||
|
||
### 2a. A fourth measurement error, found by enforcing the rule
|
||
|
||
The replay benchmark committed alongside this evidence was **the broken
|
||
version**. A `python3` patch that was supposed to replace its
|
||
log-building loop never applied, leaving a sequence that omits `Resolve`
|
||
— so `EndRound` was rejected, every round produced no events, and the
|
||
`while log.len() < target` loop spun forever. It was never run to
|
||
completion; the AM-7 replay numbers were taken from a separate probe
|
||
test instead, and the dead benchmark was committed and left hanging.
|
||
|
||
Found by adding `cargo bench -- --test` to CI, which runs every
|
||
benchmark once. That is the fourth instance of one error class in this
|
||
project — a harness that appears to work while doing no work — and the
|
||
**first one caught by a gate rather than by noticing**.
|
||
|
||
The replay loop now carries the positive control the round loop already
|
||
had: it asserts each round appended events and fails rather than
|
||
spinning. Corrected figures are in the table above; both configurations
|
||
still clear the AM-7 budget by three orders of magnitude.
|
||
|
||
The lesson recorded for the loop: writing the positive-control rule into
|
||
`specs/InnerLoop.md` did **not** prevent the next instance. Making it a
|
||
CI step did. Prose rules do not enforce themselves.
|
||
|
||
## 3. Determinism (AM-8)
|
||
|
||
- Every scenario runs twice per invocation with the same seed and fails
|
||
on state-hash divergence (K8). 21/21 pass.
|
||
- Ten consecutive full runs of all 21 scenarios produced **one distinct
|
||
output hash**, i.e. zero divergence.
|
||
- `cargo fmt --check` and `cargo clippy --workspace --all-targets
|
||
-D warnings` are clean.
|
||
- `clippy.toml` denies `HashMap`/`HashSet` workspace-wide (K6); the
|
||
aggregate holds only ordered collections, so iteration order cannot
|
||
vary between runs.
|
||
|
||
## 4. AM-4 — remediated and re-measured
|
||
|
||
**Original result: NOT MET, 33 transitive crates against a ≤20 target.**
|
||
Resolved by adopting both remediations (maintainer decision,
|
||
2026-07-31): `serde_yaml` was made optional, and the metric was
|
||
retargeted onto third-party source under audit.
|
||
|
||
### Re-measurement (`make dep-weight`)
|
||
|
||
| Configuration | Crates | Third-party LOC | Target | Verdict |
|
||
|---|---|---|---|---|
|
||
| Shipped runtime (`--no-default-features`) | 23 | 246,250 | ≤250,000 | **met** | <!-- fact:am4a_loc --><!-- fact:am4a_target -->
|
||
| Dev toolchain (default features) | 29 | 317,021 | ≤350,000 | **met** | <!-- fact:am4b_loc --><!-- fact:am4b_target -->
|
||
| Our own source | — | 3,408 | — | — |
|
||
|
||
Scenario tooling costs **70,771 lines that a shipped game never
|
||
compiles**. That split is the substantive result: the single number
|
||
previously reported conflated a runtime concern with a test concern.
|
||
|
||
**What actually changed in the build.** `cb-game-runtime` gained a
|
||
`scenarios` feature carrying `serde_yaml`; the scenario module, the
|
||
`ScenarioGame` impl and the string parsers behind it are `#[cfg]`-gated.
|
||
Both configurations compile and lint clean under `-D warnings`.
|
||
|
||
One trap worth recording: setting `default-features = false` on a
|
||
*member* dependency is silently ignored when the workspace dependency
|
||
does not specify it, so the first attempt gated nothing while appearing
|
||
to work — `cargo tree` still showed all six YAML crates. The fix was
|
||
setting `default-features = false` on the workspace dependency itself,
|
||
with `cb-sim` opting into `scenarios` explicitly. This is exactly the
|
||
class of error InnerLoop v1.0's positive-control rule targets: the build
|
||
succeeded and the feature flag looked applied. It was caught by checking
|
||
the dependency graph rather than trusting that the edit had worked.
|
||
|
||
### Why the target moved, and why that is not moving the goalposts
|
||
|
||
The ≤20 crate target was retired for two measured reasons, both
|
||
recorded before the decision was taken:
|
||
|
||
Attribution:
|
||
|
||
| Group | Crates | Count |
|
||
|---|---|---|
|
||
| `sha2` (K7 state hashing) | sha2, digest, block-buffer, crypto-common, generic-array, typenum, cpufeatures, cfg-if | 8 |
|
||
| serde derive chain | serde_derive, proc-macro2, quote, syn, unicode-ident | 5 |
|
||
| serde runtime + json | serde, serde_core, serde_json, itoa, memchr, zmij | 6 |
|
||
| `serde_yaml` (scenario files only) | serde_yaml, unsafe-libyaml, indexmap, hashbrown, equivalent, ryu | 6 |
|
||
| `rand_chacha` (K5 seeded RNG) | rand_chacha, rand_core, ppv-lite86, zerocopy | 4 |
|
||
| Clay-Borg crates | cb-kernel, cb-events, cb-game-runtime, games-ground | 4 |
|
||
|
||
1. **It was unreachable without undoing the spec's own contracts.**
|
||
Measured ladder: `serde_yaml` optional −6 (→27), dropping
|
||
`serde_json` −4 (→23), inlining SHA-256 −8 (→19), inlining ChaCha12
|
||
−4 (→15). Nothing reaches 20 except reimplementing a primitive that
|
||
K5 or K7 requires — trading an audited implementation for a
|
||
scoreboard number.
|
||
2. **Crate count does not compare across ecosystems.** Rust splits
|
||
crates far more finely than npm. The same granularity difference made
|
||
"33 vs 120 npm packages" flatter us *and* made ≤20 punish us.
|
||
|
||
Third-party source under audit is what the count was proxying for, is
|
||
comparable across ecosystems, and cannot be gamed by granularity. The
|
||
new targets are set at roughly the current measurement plus headroom,
|
||
so they bind on future growth rather than retroactively passing
|
||
something that failed: adding another `serde_yaml`-sized dependency to
|
||
the shipped runtime would breach AM-4a.
|
||
|
||
What we did **not** do: hand-roll SHA-256 or ChaCha to win a count.
|
||
|
||
## 5. Cost log (AM-12)
|
||
|
||
Per `specs/MetricsAndScenarios.md` §1a. Model: Claude Fable 5, at
|
||
`benchmarks/baselines/model-prices.toml` rates ($10/$50 per MTok).
|
||
|
||
| Task | Model | Iterations | Notes |
|
||
|---|---|---|---|
|
||
| CB-WP-0001 (whole session, T01–T09) | claude-fable-5 | 6 T08 code iterations + benchmarks | $248.46 measured; per-task split pending CB-WP-0002 |
|
||
|
||
**Correction (2026-07-31).** This section originally recorded AM-12 as
|
||
*uncomputable*. That was wrong. The declining to estimate was right; the
|
||
conclusion that no instrument existed was not. Every session transcript
|
||
(`~/.claude/projects/<slug>/<session>.jsonl`) carries exact per-message
|
||
`usage` including the cache breakdown. Read for this session:
|
||
|
||
| Component | Tokens | Cost (Fable 5) |
|
||
|---|---|---|
|
||
| Output | 585,528 | $29.28 |
|
||
| Cache read | 131,863,164 | $131.86 |
|
||
| Cache write (1h) | 4,365,668 | $87.31 |
|
||
| Input | 1,090 | $0.01 |
|
||
| **Session total** | | **$248.46** (~$124 on Opus 5) |
|
||
|
||
**53% of the cost is cache reads**, not output. Cost in an agentic loop
|
||
is driven by context size × turn count, which no "tokens per task"
|
||
metric would have surfaced.
|
||
|
||
Still missing is *attribution*: this is a whole-session figure, not a
|
||
per-task one, because nothing marks task boundaries in the transcript.
|
||
That is what CB-WP-0002 is for. The AM-12 row above should be read as
|
||
"session-level cost measured; per-task attribution pending CB-WP-0002".
|
||
|
||
## 6. Metrics not reported
|
||
|
||
- **AM-2, AM-3, AM-5** — specification-quality metrics that need a
|
||
second capability to compare against; a single data point is not a
|
||
measurement.
|
||
- **AM-9 (≤64MB)** — not instrumented. The aggregate is a few KB and
|
||
the largest log measured here is 100k events, so the budget is very
|
||
unlikely to bind, but "unlikely" is not "measured" and it is left
|
||
unclaimed.
|
||
- **AM-11** — ~~the `KernelRng` null/reference pair exists and is
|
||
exercised. It is the only port with a pair so far, so the metric is
|
||
met narrowly and will mean more once storage has one.~~
|
||
**Corrected 2026-07-31:** the pair exists; the **conformance suite does
|
||
not**. **Resolved 2026-08-01 (CB-WP-0006 T05):** `cb_kernel::rng::conformance`
|
||
drives `ChaChaRng` and `NullRng`; `cb_events::store::conformance` drives
|
||
`MemLogStore` and `FileLogStore`. One suite per port, both impls through
|
||
it. AM-11 is **met** and mutation-verified — breaking `NullRng::draw` to
|
||
return its bound fails the shared suite. The original text follows.
|
||
<!-- historical --> M-D4-SWAP is a bool over impls "passing the same conformance
|
||
suite", and the two impls are exercised by two separate, non-shared
|
||
tests — there is no `fn conformance<R: KernelRng>(…)` that both are
|
||
driven through. `grep -rn conformance` over every `.rs` returns a single
|
||
doc comment describing future work. **AM-11 is unmet**, and was never
|
||
earned. Building the suite is deferred to the re-planned pass, not to
|
||
this one.
|
||
|
||
## 7. Rules implemented under a provisional default
|
||
|
||
Ten U-items in `specs/GroundRules.md` carry PROVISIONAL defaults. Those
|
||
realized here are U2 (clamp on every application), U3 (DENY with no
|
||
legal target is a no-op that still advances), U4 (deck reshuffle), U5
|
||
(REVERSE owner relief applies whether or not the Reverse was rejected)
|
||
and U8 (GROUND—OU cancellation precedes Protection).
|
||
|
||
One further ambiguity was found during T08 and is **not** in the U-list:
|
||
**GR-E02's "successes"** is undefined in dataset 0.1. It is implemented
|
||
as the count of claimed Problems. Both scoring scenarios are marked
|
||
`provisional: true`.
|
||
|
||
All provisional behaviour lives behind named functions and is covered by
|
||
scenarios tagged `provisional: true`, so a ground-game ruling flips a
|
||
scenario rather than the kernel (K16). **Action for ground-game:** rule
|
||
on the ten U-items and on GR-E02's "successes".
|