clay-borg/workplans/CB-WP-0015-the-inert-clauses.md
tegwick aead0f2ea6
Some checks failed
ci / check (push) Failing after 3s
Sync hub IDs and work-record index for CB-WP-0015
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 14:08:30 +02:00

259 lines
12 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
id: CB-WP-0015
kind: product
title: "The two inert clauses: AM-7 scaling and AM-8 N=10"
status: done
state_hub_workstream_id: "84bc173e-bb96-4b66-b494-6630622e2a1e"
---
# Purpose
```
structural tier S (acceptance rows measure the product, not how we
work — `bench-test` is in gates.toml's
not_control_gates list, so the M trigger about the
loop's own constraints does not fire)
chaos d4 = 2 → no override
declared tier S
```
Declaration 10 of 12. Tier S: one provenance paragraph in the commit
message, no survey and no ADR. Steps 45 are not skipped.
## The two clauses nobody has touched in seven passes
`tools/mutation-check.py` has carried these two verdicts since CB-WP-0005,
in its own words:
> **AM-7 `scaling >= 0.9x`** — *"no code computes the ratio of throughput
> @100k to @5k or compares it to 0.9; Criterion reports both and nothing
> relates them."*
>
> **AM-8 `N=10 same-seed replays`** — *"the runner executes each scenario
> TWICE (K8), not ten times; the spec says N=10. The double-run is live
> and the count is not."*
Both have been reported honestly every run and neither has been worked.
AM-7 moved 1/3 → 2/3 in CB-WP-0006; AM-8 has been 1/2 since it was
instrumented. They are the last two PARTIAL rows in the acceptance table.
**And the test is named for the property it does not check.**
`replay_100k_events_is_linear_and_fast` folds a log at 10,000 and at
100,000 events, prints a throughput for each, and asserts each is under
five seconds *independently*. The two numbers are computed, printed, and
never compared. At 10,000 the five-second assertion is close to vacuous on
its own.
This is the same family as the last four passes' central findings, one
step further along: not a claim nobody checked, but **a number measured
and then not used**. The name asserts linearity; the body does not.
## The risk this pass has to face rather than route around
AM-7's clause was written against a measured boardgame.io defect —
throughput halving as history doubled. Our aggregate may not be able to
exhibit that defect at all: if GROUND state is bounded and `fold` does not
scan history, replay is O(n) by construction and the ratio is ~1.0 no
matter what.
**If so, the honest result is not a green row.** It is that AM-7's scaling
clause measures something our design makes structurally impossible to
violate, and a clause that cannot fail is the `unmutatable` verdict this
project already counts *against* itself (ADR-0005 §1). Making it green by
asserting a tautology would be the worse outcome of the two.
So this task may end in an amendment rather than an assertion. Decide it on
what the mutation says, not on which outcome closes more rows.
## Task: make AM-7's scaling clause live, or show it cannot be
```task
id: CB-WP-0015-T01
status: done
priority: high
state_hub_task_id: "9f6b6bfd-a7c0-4f0e-ab91-4e2e5f2e8eed"
```
Compute the ratio the spec names — throughput folding a growing log at
100k events against the same at 5k — and assert it against 0.9.
**The measurement has a known trap, and it is documented in this repo
already.** AM-6's assertion first ran inside `make all` and read 38,753
ev/s against 341,280 in isolation, a 9× error, because `cargo test` runs
binaries and threads concurrently: *"a throughput assertion inside a
parallel harness measures contention, not throughput."* A **ratio** of two
timings taken under varying contention is worse than a single reading, not
better — the noise does not cancel, it multiplies. Whatever this lands in
must run where AM-6 runs (release, `--test-threads=1`, its own target) or
not be gated at all.
Required, and the pass fails without them:
- **The mutation must go red for the stated reason.** Slowing `fold` in
proportion to log length — not by a constant, which AM-6 already
catches — must break the ratio and nothing else.
- **A positive control on the ratio itself.** A harness that measured the
same size twice would report 1.0 and look excellent. Prove the two legs
are different sizes and that the denominator is the small one.
- **State the variance.** One reading of a ratio is not a measurement.
Report the spread across samples, and if the spread crosses 0.9 then
the threshold is not gateable at this sample count — say that instead of
taking the best reading.
If the ratio is ~1.0 because nothing in the design can make it otherwise,
write that up and amend the row, following the precedent AM-4c set in
`specs/GameKernel.md` §5a — the argument goes in the spec, at tier S,
with the row retained in the denominator.
**Done 2026-08-02.** `make am7`; AM-7 is now **red, 3/3 clauses**.
The clause is live and no amendment was needed. Measured ratio **0.987 /
0.991 / 0.989** across three runs against the 0.9 floor, and **0.989 under
eight-way CPU contention** — the design is flat, as the bounded-state
analysis predicted, with a few percent of cache cost on the longer log.
**Two estimators were wrong before the third, and both looked
reasonable.** Best-of-N per leg then divide — AM-6's estimator, correct for
a floor on one number — gave **0.581 to 1.085** on an unchanged binary.
Running the legs back-to-back inside one sample gave medians 1.004 /
0.931 / 0.956: closer, but a 50 ms leg samples a *point* on this machine's
drift rather than averaging it. Interleaving at fold granularity with 10 M
events a leg makes the drift common-mode, and it divides out. Absolute
throughput still swings **11 M 55 M ev/s**; the ratio does not.
**The INDETERMINATE guard was wrong for the same reason the estimator
was.** It demanded unanimity. Under contention the median read 0.971 and
one sample read 0.899 — a thousandth under — and it failed the build; it
also fired intermittently inside `mutation-check`, which runs this test
straight after a 50-second rebuild. It now requires a two-thirds majority
agreeing with the median. A gate that fails when the machine is busy is a
flake, and a flake gets suppressed rather than fixed.
**The mutation, and the control that matters more.** Making fold cost grow
with history drives the ratio to **0.751** — red, tight, no straddle.
AM-6's *constant*-cost mutation halves throughput (28M → 15M ev/s) and
leaves this ratio at **0.999× — green**. Without that control, "AM-7 goes
red" would have been consistent with AM-7 being a redundant copy of AM-6.
## Task: settle AM-8's N, at 10 or at 2, with the argument
```task
id: CB-WP-0015-T02
status: done
priority: high
state_hub_task_id: "b6640044-5d17-489e-b97e-b6c82fef3554"
```
The spec says ten same-seed replays; the runner does two. Close the gap in
whichever direction survives the argument — and the argument is the
deliverable here, not the number.
**Ten is doing real work only if the failure it hunts is probabilistic.**
A deterministic divergence — a seeded RNG threaded wrong, a fold that
depends on insertion order — fails on run two exactly as reliably as on
run ten. A divergence from ambient nondeterminism — hash iteration order,
pointer-dependent hashing, thread scheduling — appears with some
probability per run, and there ten runs are eight more chances than two.
So: which does AM-8 exist to catch, and what does the repo already do
about the other? `clippy.toml` denies `HashMap`/`HashSet` and `make check`
runs `-D warnings`, which is AM-8's other clause and already live. If the
lint removes the probabilistic class structurally, ten runs are buying
repetitions of a check that cannot fail — and that costs time on every
`make all`.
Whichever way it goes, the mutation must still go red: the existing
seed-perturbation mutation is the control, and if raising or lowering N
does not change what that mutation catches, that is itself the finding.
Do not amend the spec silently. If N becomes 2, GameKernel §5 says so and
says why, in the same shape as §5a.
**Done 2026-08-02.** `make am8`; AM-8 is now **red, 2/2 clauses**.
`GameKernel.md` §5b records where the ten runs live.
**N stays at 10, and the measurement is why.** The argument above is
correct as far as it goes — a deterministic divergence shows on run 2
exactly as reliably as on run 10 — but it is not the whole population.
Perturbing the RNG only from its **fourth** construction onward, on
`gr-r06-round-resolve`:
| | result |
|---|---|
| `cb-sim --runs 2` | **PASS** |
| `cb-sim --runs 10` | FAIL — *"run 1 hash … != run 4 hash … (of 10)"* |
A **late-onset** divergence is a real class the double-run structurally
cannot see, and — the part that makes it usable — it is *deterministic,
not probabilistic*. So it can be a control rather than a coin flip, and
the `N=10` clause is now measured by it rather than declared.
The cost objection stands and is answered by placement, not by amendment:
ten runs across all 25 scenarios would cost **47 s per build** to
re-answer an answered question, so `make sim` keeps K8's double-run over
everything and `make am8` runs ten on one scenario in ~2 s. The primary
defence against the *probabilistic* class is still the `HashMap`/`HashSet`
deny lint — this row's other clause, already live — and the ten runs are
defence in depth against that exclusion failing.
Also here: `scenario::run_n`, `cb-sim --runs <n>` (which refuses `n < 2`
rather than falling back to 2 — a flag that silently degraded would report
an N=10 result after an N=2 check), and later runs compared to the **first**
rather than to their predecessor, so a divergence appearing at run 5 and
persisting cannot hide after run 6.
## Task: evidence, and what the acceptance table now claims
```task
id: CB-WP-0015-T03
status: done
priority: high
state_hub_task_id: "a5ceddbe-acde-47d3-81e0-e29bb73a4281"
```
`evidence/CB-EV-0013-*.md`.
Report the mutation table before and after, and be exact about the
denominator — CB-EV-0005 established that a score improved by deleting the
question is not an improvement, and this pass can improve its score two
ways, one of which is deleting a question.
Also due here:
- **Quote CB-WP-0014's final cost, not this pass's.** The rule adopted in
CB-EV-0012. CB-WP-0014 closed at **$7.47 / 34 responses** = 0.220.
- **CB-EV-0009's standing prediction is now testable for the first
time.** It predicted a pass opening above the SH-1 hard line would cost
more than 0.123 per response. CB-WP-0014 opened above the line and cost
0.220. Report it as one data point, and say plainly that one point past
a threshold is not a confirmation — the prediction has a direction and
no control.
- **Whether the acceptance table still has a PARTIAL row**, and if the
answer is no, whether that is because both clauses became live or
because one became a spec amendment.
**Done 2026-08-02.**
[CB-EV-0013](../evidence/CB-EV-0013-the-inert-clauses.md). `make all`
exits 0.
- **No PARTIAL rows remain, and neither clause was amended away.** AM-7
red 3/3, AM-8 red 2/2. **M-D1-MUT 10/14**, measured, and ADR-0005's
standing prediction (≥10 of 14) is **MET for the first time**.
- **The denominator did not move**, which was the live risk. This pass
could have improved its score by deleting a question — amending AM-8's
N down to 2, or splitting AM-7's scaling into a fifteenth row — and both
were considered and refused. The four unenforced rows are the four that
were already unenforceable.
- **The full run found a third defect: AM-4a's mutation was stale**, still
pointing at the 250,000 target ADR-0008 D3 replaced with 161,000. It
reported HARNESS-BROKEN and refused to publish a score. The build-free
half of that check is now a `--self-test` assertion, so `make all`
catches the next one instead of only a full mutation run.
- **`mutation-check` now measures a claim it used to assert.** A clause
may carry its own verify and mutation; then its `enforced` flag is
measured and a declaration disagreeing with its measurement is refused.
Two clauses carry one, and print as `red*`.
- **CB-EV-0009's prediction came due**: CB-WP-0014 opened above the SH-1
hard line and cost 0.220 $/response against the predicted floor of
0.123. Unfalsified, **not confirmed** — there is no control, because no
pass has opened below the line since it was made.