CB-WP-0005 T03: correct the record, and defer Phase C

Four verdicts in evidence/CB-EV-0001 corrected in place with a dated
note, per ADR-0005 §4: AM-7 replay split (timing met, hash-identical
withdrawn), AM-10 withdrawn as written and restated as AM-10' (the K6
determinism lint it actually measured), AM-11 downgraded to unmet, and
AM-1b added to the scoreboard it was missing from.

The scoreboard gains an Enforced column carrying M-D1-MUT, because a row
can be measured and still enforce nothing and the table had no way to say
so. AM-6 now reads "met, 16.5x" alongside "not enforced — nothing
compares any number to 100,000".

A fifth correction surfaced that ADR-0005 did not list: AM-12 still read
$248.46, the figure CB-WP-0002 disproved and corrected to $93.15 four
workplans ago. It was stale in the evidence file ever since — untagged,
and therefore invisible to facts-check. Now tagged. A duplicated-fact
instance that survived the gate built to catch duplicated facts, because
that gate only checks copies that opted in. Recorded for T07.

GameKernel §5 carries the AM-10 withdrawal and AM-11 downgrade inline so
a reader of the spec cannot reach the old claim.

Phase C is deferred before starting, per the stop condition T02 wrote and
the maintainer's decision. It is scoped to five rules; the measurement
says eight acceptance rows have no instrument at all. Building it as
written would proceed on a diagnosis the instrument had just
contradicted. T04-T06 stay in the file with their analysis intact and
move to CB-WP-0006.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
tegwick 2026-07-31 17:54:29 +02:00
parent 4e8d89f482
commit 145611e3b6
3 changed files with 86 additions and 19 deletions

View file

@ -14,21 +14,52 @@ Machine: WSL2, Linux 6.18.33.2-microsoft-standard-WSL2, rustc 1.97.1,
## 1. Scoreboard
| Metric | Target | Measured | Verdict |
|---|---|---|---|
| AM-1 rule coverage | 100% of GR-rules | 58/58 (100%) | **met** |
| AM-4a dep weight, shipped runtime | ≤250,000 third-party lines | 246,250 (23 crates) | **met** | <!-- fact:am4a_loc -->
| AM-4b dep weight, dev toolchain | ≤350,000 third-party lines | 317,021 (29 crates) | **met** | <!-- fact:am4b_loc -->
| AM-6 throughput | ≥100,000 events/s | 1,651,400 events/s | **met, 16.5×** |
| AM-7 scaling | ≥0.9× at 20× workload | 1.08× | **met** |
| AM-7 replay | 100k events ≤5s | 2.18 ms (CI 2.142.23) | **met, 2,290×** |
| AM-8 determinism | zero divergence, 10 replays | 1 distinct hash / 10 runs | **met** |
| AM-8 lint | fmt + clippy clean | clean, `-D warnings` | **met** |
| AM-10 foreign types | zero `HashMap`/`HashSet` | 0 | **met** |
| AM-11 impl pairs | null + reference per port | 1 of 1 (`KernelRng`) | **met, narrow** |
| AM-12 cost | per-task USD | $248.46 session; per-task pending | **partial** |
> **Corrected 2026-07-31 (CB-WP-0005 T03, per ADR-0005 §4).** Four rows
> below were wrong or misleading as originally committed. They are
> corrected **in place with this note**, not silently edited — the
> correction trail is the artifact. What changed, and why, is in §1a.
AM-2, AM-3, AM-5 and AM-9 are not reported: see §6.
The **Enforced** column is new and is the point of the correction. It
carries M-D1-MUT (`make mutation-check`): does anything actually fail when
the property is false? A row can be *measured* and still enforce nothing.
| Metric | Target | Measured | Verdict | Enforced |
|---|---|---|---|---|
| AM-1 rule coverage (GR only) | 100% of GR-rules | 58/58 (100%) | **met** | **yes** |
| AM-1b link, ground | 58 claimed rules named in the aggregate | 49/58 | **unmet** | reported |
| AM-1b link, kernel | 18 K-rules named in source | 15/18 | **unmet** | reported until 2026-08-31 |
| AM-4a dep weight, shipped runtime | ≤250,000 third-party lines | 246,250 (23 crates) | **met** | **yes** | <!-- fact:am4a_loc -->
| AM-4b dep weight, dev toolchain | ≤350,000 third-party lines | 317,021 (29 crates) | **met** | **yes** | <!-- fact:am4b_loc -->
| AM-6 throughput | ≥100,000 events/s | 1,651,400 events/s | **met, 16.5×** | **no** — nothing compares any number to 100,000 |
| AM-7 scaling | ≥0.9× at 20× workload | 1.08× | **met** | **no** — no code computes the ratio |
| AM-7 replay, timing | 100k events ≤5s | 2.18 ms (CI 2.142.23) | **met, 2,290×** | **yes** |
| AM-7 replay, hash-identical | bit-identical fold | — | **WITHDRAWN** | **no** — mutation-proven inert |
| AM-8 determinism | zero divergence, 10 replays | 1 distinct hash / 10 runs (one-off probe) | **met, narrow** | **partial** — the runner double-runs; nothing enforces N=10 |
| AM-8 lint | fmt + clippy clean | clean, `-D warnings` | **met** | **yes** |
| AM-10 foreign types | 0 in `cb-*-api` signatures | — | **WITHDRAWN** | **no** — no such crate; population empty |
| AM-10 determinism lint (K6) | zero `HashMap`/`HashSet` in game state | 0 | **met** | **yes** |
| AM-11 impl pairs | ≥2 impls under **one conformance suite** | pair yes, suite **none** | **UNMET** | **no** — the suite does not exist |
| AM-12 cost | per-task USD | **$93.15** pinned, per task via `make cost` | **met** | **yes** | <!-- fact:pinned_total -->
AM-2, AM-3, AM-5 and AM-9 are not reported: see §6. All four are
`unmutatable` — no instrument exists to invert.
**M-D1-MUT over the whole acceptance table: 4 of 14 rows enforced.**
### 1a. What was corrected, and why
| row | as committed | corrected to | found by |
|---|---|---|---|
| **AM-7 replay** | `met, 2,290×` | split: timing **met**, `hash-identical` **withdrawn** | mutation — folding the log from `fresh(999)` instead of `fresh(42)`, an unrelated genesis state, leaves the test green. The hash reaches only a `println!`. It could not be asserted as written anyway: the log spans games seeded 42, 43, 44… so it is not a replay of anything |
| **AM-10** | `met` | **withdrawn as written**; restated as AM-10 | adversarial review — there is no `cb-*-api` crate, so `0 foreign types` was true over an empty set. What was actually measured is a `clippy.toml` deny of `HashMap`/`HashSet` whose stated reason cites **K6 (determinism)**, reported under a **D4 leak** row |
| **AM-11** | `met, narrow` | **unmet** | adversarial review — M-D4-SWAP is a bool over "the same conformance suite"; `grep -rn conformance` over every `.rs` returns one doc comment describing future work, and the RNG pair is exercised by two separate, non-shared tests |
| **AM-12** | `$248.46 session` | **$93.15** | CB-WP-0002 re-derivation — the original figure double-counted per-line and priced a three-model session at one model's rate. It had been stale in this file since, untagged and therefore invisible to `facts-check` |
| **AM-1b** | *absent from the scoreboard* | **added, both denominators** | adversarial review — `make coverage` printed `49/58` two lines below the `100%` this table carried, and the table kept the flattering half |
Nothing here was found by re-running the command that printed the number.
Three of the five were found by an adversarial reviewer reading the
assertion behind the number, and two of those required a mutation to
settle. That is the loop change recorded in CB-WP-0005 T08.
## 2. Throughput and scaling (AM-6, AM-7)
@ -251,9 +282,17 @@ That is what CB-WP-0002 is for. The AM-12 row above should be read as
the largest log measured here is 100k events, so the budget is very
unlikely to bind, but "unlikely" is not "measured" and it is left
unclaimed.
- **AM-11** — the `KernelRng` null/reference pair exists and is
- **AM-11**~~the `KernelRng` null/reference pair exists and is
exercised. It is the only port with a pair so far, so the metric is
met narrowly and will mean more once storage has one.
met narrowly and will mean more once storage has one.~~
**Corrected 2026-07-31:** the pair exists; the **conformance suite does
not**. M-D4-SWAP is a bool over impls "passing the same conformance
suite", and the two impls are exercised by two separate, non-shared
tests — there is no `fn conformance<R: KernelRng>(…)` that both are
driven through. `grep -rn conformance` over every `.rs` returns a single
doc comment describing future work. **AM-11 is unmet**, and was never
earned. Building the suite is deferred to the re-planned pass, not to
this one.
## 7. Rules implemented under a provisional default

View file

@ -171,8 +171,8 @@ evidence lands in `evidence/CB-EV-0001-game-kernel.md` with no
| AM-7 | M-D3 scaling: throughput @100k events vs @5k; and snapshot+replay of 100k events | boardgame.io 0.450.66× @2040k, DNF @100k | **≥ 0.9×** (flat), replay of 100k events ≤ 5 s, hash-identical | measured |
| AM-8 | Determinism invariant: N=10 same-seed replays, bit-identical hashes; HashMap-in-state deny lint clean | Rune: enforced by tooling (cited) | zero divergence, lint clean in CI | measured (invariant, not a verdict row) |
| AM-9 | M-D3-MEM: peak RSS, 100k-event synthetic run | boardgame.io ~100→232 MB @5k→40k (indicative) | ≤ 64 MB, flat with history given snapshot interval | measured (indicative label, same method) |
| AM-10 | M-D4-LEAK: foreign types in `cb-*-api`-visible signatures | boardgame.io: JS-ecosystem-locked | **0** | measured (grep/deny rule) |
| AM-11 | M-D4-SWAP: null + reference impls passing one conformance suite | no candidate has the pattern | RNG and log storage each have ≥2 impls (real + test/null) under one suite | measured (bool) |
| AM-10 | M-D4-LEAK **(withdrawn 2026-07-31, ADR-0005 §4 — no `cb-*-api` crate exists, so the population is empty; the clippy `HashMap`/`HashSet` deny that stood in for it cites K6 determinism and is now reported as AM-10)**: foreign types in canonical-interface signatures | boardgame.io: JS-ecosystem-locked | **0** | measured (grep/deny rule) |
| AM-11 | M-D4-SWAP **(unmet 2026-07-31 — the pair exists, the suite does not)**: null + reference impls passing one conformance suite | no candidate has the pattern | RNG and log storage each have ≥2 impls (real + test/null) under one suite | measured (bool) |
| AM-12 | M-D2-TOK / M-D2-CST: tokens and USD per completed task | n/a — first pass sets our own baseline | recorded per task in the evidence cost log (price sheet 2026-07-31) | recorded, not gated |
Comparisons against the event-sourcing 10⁵10⁶/s estimate stay **parity**

View file

@ -211,7 +211,35 @@ flattering half. Update `specs/GameKernel.md` §5 and
`specs/MetricsAndScenarios.md` §1 to match, and tag the new figures under
`make facts-check` so they cannot drift.
## Phase C — implement what the instruments expose
**Delivered.** `evidence/CB-EV-0001` §1 now carries a dated correction
note and a new **Enforced** column carrying M-D1-MUT — because a row can
be *measured* and still enforce nothing, and the scoreboard had no way to
say so. §1a records what changed and how each was found.
**A fifth correction surfaced that ADR-0005 did not list:** AM-12 still
read **$248.46**, the figure CB-WP-0002 disproved and corrected to
$93.15. It had been stale in the evidence file ever since — **untagged,
and therefore invisible to `facts-check`**. It is now tagged
`fact:pinned_total`. A DFD instance that survived the gate built to catch
DFD, because that gate only checks copies that opted in.
`specs/GameKernel.md` §5 carries the AM-10 withdrawal and the AM-11
downgrade inline, so a reader of the spec cannot reach the old claim.
## Phase C — DEFERRED before starting (see T02)
> **Deferred 2026-07-31 by the stop condition in T02, maintainer decision.**
> M-D1-MUT measured **4 of 14** (28.6%) against a 25% floor and a 71%
> prediction. Phase C is scoped to K9/K10/K11/K14/K18 — five rules — but
> the measurement says **8 of 14 acceptance rows have no instrument at
> all**, which Phase C does not address. Building it as written would mean
> proceeding on a diagnosis the instrument had just contradicted.
>
> T04, T05 and T06 stay in this file, unstarted, with their analysis
> intact; they move to **CB-WP-0006**, which is scoped to the finding that
> was actually measured rather than the one that was predicted.
## Phase C — implement what the instruments expose (deferred)
## Task: K9 and K11 — durable log, LogStore port, real conformance suite