clay-borg/workplans/CB-WP-0006-instrument-the-table.md
tegwick c43754f0fe CB-WP-0006 T01: assert the AM-6 throughput target
Nothing in the workspace compared any number to 100,000 events/s while
the evidence file reported "AM-6 | met, 16.5x". Now a test does — a test,
not a bench, because Criterion reports throughput and asserts nothing,
which is why this row measured nothing for six passes.

Measured on bnt-lap001: 341,280 ev/s in debug (3.4x the target), ~2.4-3.1M
in release. The spec target holds even in an unoptimized build, so the
gate needs no cfg split and runs in the ordinary `make test`.

The trap this task named — loosening a flaky timing assertion until it
never fires — is avoided by construction. The threshold is the spec value,
untouched; the constant says lowering it requires an ADR; and the failure
message repeats that, states measured headroom, and names reference
figures, so an agent hitting a red AM-6 is told not to tune it in the
place they are actually reading. Robustness comes from best-of-N, not from
a lower bar: a throughput floor asks whether the machine is capable, so
transient load should not fail the build.

Two positive controls in the test: a run that applied fewer than 50,000
events, or measured zero elapsed time, fails rather than scoring as
infinite throughput.

Verified by a PROPERTY mutation — 4,000 black_box iterations injected into
GroundState::fold, the hot path — not a threshold tweak, which would only
prove the comparison runs.

And the FA class found last pass is now gated. mutation-check rows gained
an `expect` field: the mutant's output must contain the row's stated
failure string or the verdict is WRONG-REASON, not red. Without it a
mutation that merely failed to compile would credit its row with an
assertion it does not have. Verified by pointing expect at a string the
verifier never prints and watching the verdict flip. This is remedy (2)
from the CB-WP-0005 retrospective, built a task earlier than planned
because the class it guards is the newest and most dangerous.

M-D1-MUT: 4 -> 5 of 14.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 18:32:16 +02:00

12 KiB
Raw Blame History

id title status state_hub_workstream_id
CB-WP-0006 Instrument the acceptance table, then implement what it exposes in_progress 8a6327cc-fd5c-4e2c-a29b-b437c27d1e71

Purpose

CB-WP-0005 T02 measured M-D1-MUT for the first time:

  M-D1-MUT: 4/14 rows enforced
    PARTIAL       2      (AM-7, AM-8 — some clauses live, some inert)
    unmutatable   8      (no property to invert)

4 of 14, against a predicted 10 of 14. The prediction was wrong by 2.5×, and wrong about what as well as how much: CB-RES-0004 diagnosed three absent kernel rules, and the measurement found that more than half the acceptance table has no instrument behind it at all.

CB-WP-0005 Phase C was deferred before starting on that basis (T02's stop condition, maintainer decision 2026-07-31). It was scoped to five rules — K9, K10, K11, K14, K18 — which does not address the eight.

This workplan is scoped to what was measured rather than what was predicted. Phase A instruments the unmeasured rows; Phase B implements the kernel work inherited from CB-WP-0005 T04T06. That order is the whole point: the previous pass built a fix before its instrument existed, and the instrument then contradicted the fix's premise.

Per InnerLoop §Step 4, no target moves in the commit that measures it.

Phase A — instrument the rows that measure nothing

Task: AM-6 — assert the throughput target

id: CB-WP-0006-T01
status: progress
priority: high
state_hub_task_id: "a4513794-2f5f-480c-be30-56658711796f"

The sharpest single finding of CB-WP-0005. Nothing in the workspace compares any number to 100,000 events/s — the project's headline throughput claim, reported as met, 16.5×. The Criterion bench emits throughput; the only asserts in synthetic.rs are the stress-gate shape and the events-per-round pin.

Deliver a standing assertion — a test, not a bench — that measures applied events/s on the synthetic workload and fails below the AM-6 target.

The trap, stated up front: a timing assertion in CI is flaky by nature, and the usual response is to loosen it until it never fires, which would reproduce exactly the defect being fixed. So the threshold is pinned well below the measured 1.65M/s (AM-6 asks for 100k), the assertion states its own headroom in the failure message, and a loosening requires an ADR. A gate tuned until it cannot fire is a gate that counts names.

Verified by: make mutation-check --row AM-6 goes from unmutatable to red.

Delivered. am6_throughput_clears_the_spec_target in games/ground/src/lib.rs — a test, not a bench. Best of 3 samples of 50,000 applied events each.

Measured on bnt-lap001 2026-07-31: 341,280 ev/s in debug (3.4× the target), ~2.43.1M in release (~2430×). The spec target holds even in an unoptimized build, so the gate needed no cfg split and runs in the ordinary make test.

The trap was avoided by construction, not by intention. The threshold is the spec value 100_000, untouched; the constant carries a comment saying lowering it requires an ADR; and the failure message repeats that, states the measured headroom, and names the reference figures — so a future agent hitting a red AM-6 is told not to tune it, in the place they will actually be reading. Robustness comes from best-of-N, not from a lower bar: a throughput floor asks "is this machine capable", so transient load should not fail the build.

Two positive controls in the test itself: a run that applied fewer than 50,000 events, or measured zero elapsed time, fails rather than scoring as infinite throughput.

Verified: make mutation-check --row AM-6red, via a property mutation (4,000 black_box iterations injected into GroundState::fold, the hot path) rather than a threshold tweak — raising the target would only prove the comparison runs.

And the FA class is now gated. mutation-check rows gained an expect field: the mutant's output must contain the row's stated failure string, or the verdict is WRONG-REASON, not red. Without it, a mutation that merely failed to compile would credit its row with an assertion it does not have. Verified by pointing expect at a string the verifier never prints and confirming the verdict flips. This is remedy (2) from the CB-WP-0005 retrospective, built one task earlier than T08 planned because the class it guards is the newest and the most dangerous.

M-D1-MUT: 4 → 5 of 14.

Task: AM-2, AM-3 — instrument the size metrics

id: CB-WP-0006-T02
status: todo
priority: medium
state_hub_task_id: "d1572a4f-46e5-49dd-b1ec-7c371a780bb7"

AM-2 (≤40 spec lines per rule) and AM-3 (≤50 LOC for the synthetic workload) are both unmutatable for the same reason: make loc prints line counts and nothing divides, compares, or gates.

AM-2 is the anti-gaming pair for AM-1 — 100% rule coverage means nothing if the rules are trivially small — so an uninstrumented AM-2 leaves AM-1 gameable, which is the specific risk MetricsAndScenarios §1 names when it introduces M-D1-SPL.

AM-3 depends on K18 (benches driven from scenario files); if K18 resolves toward amending the spec rather than implementing it, AM-3 must be restated or withdrawn with an argument rather than left unmeasured.

Task: AM-5, AM-9 — measure or withdraw, but stop leaving them blank

id: CB-WP-0006-T03
status: todo
priority: medium
state_hub_task_id: "6c84a220-cab3-40f1-b35f-b8df79106f6e"

AM-5 (clean build ≤60 s) is recorded not gated — and not recorded either. AM-9 (peak RSS ≤64 MB) is declared "very unlikely to bind", which CB-EV-0001 itself calls out as an unmeasured judgment call.

Both are cheap to measure: one time around a clean build, one /usr/bin/time -v around the 100k-event run. Or withdraw them with an argument, which is a legitimate outcome and possibly the right one for AM-5 on a machine-dependent number. What is not legitimate is a third pass leaving them blank while the scoreboard reports elsewhere.

Task: AM-4c — target it or drop it

id: CB-WP-0006-T04
status: todo
priority: low
state_hub_task_id: "a2571314-4c02-41cf-9473-96d3ae2d7b2b"

reported, not targeted, so there is no property to invert and it counts against M-D1-MUT by ADR-0005 §1.

Two honest options, and the task must pick one with an argument: give it a threshold (making it enforceable), or remove it from the acceptance table and keep it as a reported diagnostic. A row in an acceptance table that cannot fail is decoration, and this pass has learned to say so.

Note the incentive: ADR-0005 already predicts own-source growth from the kernel work in Phase B, so AM-4c will move against us. Setting its threshold after seeing that movement would be the retarget InnerLoop §Step 4 forbids — so it is set here, before Phase B, or not at all.

Phase B — the kernel work inherited from CB-WP-0005

These three were written in CB-WP-0005 Phase C and cancelled unstarted. Their analysis lives here, next to the work — CB-WP-0005 keeps only a forward pointer, so there is one copy. The normative controls are in ADR-0005 §6.

Task: K9 and K11 — durable log, LogStore port, real conformance suite

id: CB-WP-0006-T05
status: todo
priority: high
state_hub_task_id: "2a357dbf-3861-499d-a12e-f7682b7332f6"

K9 — the acceptance property is snapshot at seq N + events N+1..M ≡ genesis fold, hash-compared, on GroundState. What exists instead is one test round-tripping a BTreeMap<String, u8> with EventSeq(17) as a literal: no game aggregate, no events applied, no from-genesis comparison. Mutation-proven — making Snapshot::take discard its EventSeq leaves the test green, so the half of K9 that says "+ the EventId it includes" is unverified. Snapshot has no callers outside its own test.

K11 — append-only, length-prefixed, versioned framing, with a truncated tail detected; the operative clause is detection, so truncate-by-one- byte and corrupt-length-prefix rejection are required, not optional. Reimplemented rather than assimilated (ADR-0005 §2): ~100 lines against a format Kafka and EventStore converged on independently, adding no dependency, charged to AM-4a (shipped runtime, 1.5% headroom). The LogStore port with a shared fn conformance<S: LogStore>(…), retro-fitted to KernelRng, is what finally earns AM-11 — currently downgraded to unmet because the pair exists and the suite does not.

Verified by: make mutation-check --row AM-11 goes from unmutatable to red.

Task: K10 — the replay bundle and --replay

id: CB-WP-0006-T06
status: todo
priority: high
state_hub_task_id: "d459444d-41be-40c5-9101-7a8420b42fb6"

INTENT design decision 8 of 10, unimplemented. cb-sim has no flag parsing at all, so --replay has nowhere to go yet.

This is not "a directory of four files" — the D2 correction the adversarial review forced. scenario.rs creates an EventLog, appends to it and never reads it: the one production instantiation of the K11 log is a write-only sink. Pass carries the end state, not an initial snapshot, and failures are a formatted String, not structured expected-vs-actual. Plumb the log out of execute, capture an initial snapshot, restructure RunOutcome::Failed. The writer is dev-only behind the scenarios feature and is charged to AM-4b (9.4% headroom), not AM-4a.

All four controls from ADR-0005 §6 are required verbatim, including the two the reviewer supplied: truncate-by-one-byte rejection, and a mutated-seed control so the round-trip is capable of failing.

Fixing this is also what lets AM-7's withdrawn hash-identical clause be re-earned, since a real replay is what the clause was always about.

Task: K14 and K18 — implement, or amend the spec and say why

id: CB-WP-0006-T07
status: todo
priority: medium
state_hub_task_id: "d76ad893-520b-44bd-9b77-07dbfb651ce3"

Inherited from CB-WP-0005 T06. CommitWindow has zero non-test users; the bench never touches ScenarioFile. Either wire them up or amend the rules with a recorded argument. Deleting a rule to make a gate green is forbidden; deleting one with an argument is legitimate and probably right for K18.

Phase C — measure, and answer the question this pass inherited

Task: control loop — M-D1-MUT before and after

id: CB-WP-0006-T08
status: todo
priority: high
state_hub_task_id: "1d4455c3-c642-47d8-916b-ab35bc512207"

Commit evidence/CB-EV-0004. The baseline is 4 of 14, committed and reproducible via make mutation-check.

Three tests, all reported:

  1. Did the enforced count rise, per row? Against the per-task predictions above, unmet reported as unmet.
  2. Did any row regress from red to SURVIVED? The four currently enforced rows are the ones with something to lose.
  3. Were the new mutations strong? CB-WP-0005 T02 wrote two no-op mutations on its first attempt, either of which would have published a false accusation against working code. Every mutation added by this pass must be shown to fail for the stated reason — record how that was checked.

Task: retrospective — is a weak mutation the new grep?

id: CB-WP-0006-T09
status: todo
priority: medium
state_hub_task_id: "4f291d5a-90f0-4f1f-bcaf-7ec28856bfcf"

The question CB-WP-0005 T08 raised and this pass is positioned to answer.

CB-WP-0004 concluded that tooling recovers capacity only where it removes the manual path. M-D1-MUT does not remove a manual path — writing a weak mutation is exactly as easy as writing a strong one, and the harness cannot tell the difference. CB-WP-0005 produced two no-op mutations on its first attempt, so the failure mode is demonstrated, not hypothetical.

And it fails in the more dangerous direction: a weak mutation does not merely miss a defect, it indicts working code. A false SURVIVED sends the next pass to fix something that is not broken.

So: is M-D1-MUT a real instrument, or a name-counter with extra steps? If the latter, say so and propose what would actually bind — the candidate being that the mutation must be written by someone other than the author of the assertion, which is the adversarial-review principle applied one level down.