CB-WP-0015: the two inert clauses, AM-7 scaling and AM-8 N=10
Some checks failed
ci / check (push) Failing after 3s
Some checks failed
ci / check (push) Failing after 3s
Provenance (tier S, one paragraph in lieu of survey and ADR): the two clauses mutation-check has reported inert since CB-WP-0005. AM-7's scaling ratio was held up by a test literally named replay_100k_events_is_linear_and_fast that computed both throughputs, printed both, and never divided one by the other. AM-8's N=10 was held up by a runner that does two. Both are now red. AM-7 3/3, AM-8 2/2, M-D1-MUT 10/14, and ADR-0005's >=10-of-14 prediction MET for the first time. Neither was closed by amending the question away, which was the live risk: the denominator is unchanged and the four unenforced rows are the four already unenforceable. AM-7 needed three estimators. Best-of-N per leg then divide (AM-6's, correct for a floor on one number) gave 0.581-1.085 on an unchanged binary; legs back-to-back gave medians 0.931-1.004; legs interleaved at fold granularity give 0.987/0.991/0.989, and 0.989 under 8-way CPU contention while absolute throughput fell 4x. The INDETERMINATE guard demanded unanimity and failed a good measurement over one sample 0.001 under the floor; it now requires a two-thirds majority. The control that matters: AM-6's constant-cost mutation halves throughput and leaves this ratio at 0.999x green, so AM-7 is not a second AM-6. AM-8 kept N=10 because the measurement said so. Perturbing the RNG only from its fourth construction on: --runs 2 PASSES, --runs 10 fails. A late-onset divergence is deterministic, not flaky, so it is a control rather than a coin flip. Ten runs live on one scenario (make am8, ~2s) rather than all 25 (47s a build). GameKernel 5b records it. The full run also found AM-4a's own mutation stale since ADR-0008 D3 moved the target 250,000 -> 161,000 in CB-WP-0013 -- reported HARNESS-BROKEN, no score published. The build-free half of that check is now a --self-test assertion, so make all catches the next one. mutation-check clauses may now carry their own verify and mutation, and then the enforced flag is measured rather than declared; a declaration disagreeing with its measurement is refused. make all exits 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
7e9ab221a7
commit
ee37b82675
9 changed files with 802 additions and 51 deletions
201
evidence/CB-EV-0013-the-inert-clauses.md
Normal file
201
evidence/CB-EV-0013-the-inert-clauses.md
Normal file
|
|
@ -0,0 +1,201 @@
|
|||
# CB-EV-0013 — two inert clauses, and an estimator that was wrong twice
|
||||
|
||||
CB-WP-0015 T03. Measured 2026-08-02 at `7e9ab22`+. Pass kind `product`,
|
||||
tier **S** (chaos d4=2, no override). Declaration 10 of 12.
|
||||
|
||||
Per the rule adopted in CB-EV-0012, the cost table quotes **CB-WP-0014's**
|
||||
final figure and declines to quote this pass's own. See §5.
|
||||
|
||||
---
|
||||
|
||||
## 1. The machine wanders by 2.5×, and that is the whole story
|
||||
|
||||
Before any of this could be measured, the instrument had to survive the
|
||||
machine. Folding the same log on an unchanged binary, seconds apart:
|
||||
|
||||
| leg | observed range, one afternoon |
|
||||
|---|---|
|
||||
| 5,000-event fold | 11.2 M – 55.7 M ev/s |
|
||||
| 100,000-event fold | 9.7 M – 45.9 M ev/s |
|
||||
|
||||
**A 5× swing in absolute throughput.** Any single reading of either number
|
||||
is worthless, and AM-6 already knew this — it takes best-of-3 and says so.
|
||||
What AM-6's estimator cannot do is survive being turned into a ratio.
|
||||
|
||||
## 2. AM-7's scaling clause: three estimators, two of them wrong
|
||||
|
||||
`mutation-check` had said the same thing every run since CB-WP-0005: *"no
|
||||
code computes the ratio of throughput @100k to @5k or compares it to 0.9;
|
||||
Criterion reports both and nothing relates them."* Worse, the test next to
|
||||
it is called `replay_100k_events_is_linear_and_fast` and checks the two
|
||||
sizes **independently** — it computes both numbers, prints both, and never
|
||||
divides one by the other. **The name asserts the property the body does
|
||||
not test.**
|
||||
|
||||
| estimator | result on an unchanged binary | why it failed |
|
||||
|---|---|---|
|
||||
| best-of-5 per leg, then divide (AM-6's) | **0.581 – 1.085** | the legs are measured minutes apart; noise multiplies rather than cancels |
|
||||
| legs back-to-back inside one sample | medians 1.004 / 0.931 / 0.956 | closer, but a 50 ms leg samples a *point* on the drift rather than averaging it |
|
||||
| legs interleaved per fold, 10 M events each, median of 9 | **0.987 / 0.991 / 0.989** | drift becomes common-mode and divides out |
|
||||
|
||||
The third one holds up under abuse: with eight busy-loops pinned against
|
||||
eight cores, absolute throughput fell 4× and the **median ratio stayed at
|
||||
0.989**. Inside a real `make all`, after the build has just hammered every
|
||||
core, it read **0.965** (worst sample 0.944) — the condition the first
|
||||
guard failed on, now passing with margin.
|
||||
|
||||
**The measured answer is ~0.97–0.99 against a 0.9 floor.** Nothing in
|
||||
`GroundState` grows with log length, so the fold is O(1) per event by
|
||||
construction; the residual few percent is cache residency on streaming a
|
||||
20×-longer `Vec`.
|
||||
|
||||
### The guard that was wrong for the same reason the estimator was
|
||||
|
||||
The first version declared INDETERMINATE if *any* sample fell on the other
|
||||
side of the floor — unanimity. Under contention the median read 0.971,
|
||||
which is a good measurement, and one sample read **0.899** — a thousandth
|
||||
under — and the guard failed the build. It also fired intermittently
|
||||
inside `mutation-check`, where this test runs straight after a 50-second
|
||||
rebuild.
|
||||
|
||||
A gate that fails when the machine is busy is a flake, and a flake gets
|
||||
suppressed rather than fixed. It now requires a **two-thirds majority**
|
||||
agreeing with the median: the guard keeps its purpose — refusing to read a
|
||||
coin-flip as a verdict — without treating one outlier as one.
|
||||
|
||||
### The mutation, and the control that says AM-7 is not a second AM-6
|
||||
|
||||
To make the property false you must make fold cost grow *with history*,
|
||||
which is precisely boardgame.io's measured defect. Growing the deck by one
|
||||
card per fold and scanning it drives the ratio to **0.751** — red, tight,
|
||||
no straddle.
|
||||
|
||||
**And the control matters more than the mutation.** AM-6's mutation adds a
|
||||
*constant* per-event cost. It halves throughput, 28 M → 15 M ev/s, and
|
||||
leaves this ratio at **0.999× — green**. The two rows catch different
|
||||
things: a constant slowdown is AM-6's, a history-proportional one is
|
||||
AM-7's. Without that check, "AM-7 goes red" would have been consistent
|
||||
with AM-7 being a redundant copy of a row we already had.
|
||||
|
||||
## 3. AM-8's N: the spec was right and the runner was wrong
|
||||
|
||||
Eight passes of the runner executing each scenario **twice** while the row
|
||||
said ten. Closing it needed an argument, because *"the spec says ten"* is
|
||||
not one.
|
||||
|
||||
A **deterministic** divergence — a seed threaded wrong, an order-dependent
|
||||
fold — shows on run 2 exactly as reliably as on run 10. For that class the
|
||||
double-run is sufficient and eight more runs across 25 scenarios cost 47 s
|
||||
a build to re-answer an answered question.
|
||||
|
||||
**The measurement settled it.** Perturbing the RNG only from its fourth
|
||||
construction onward, on `gr-r06-round-resolve`:
|
||||
|
||||
| | result |
|
||||
|---|---|
|
||||
| `cb-sim --runs 2` | **PASS** |
|
||||
| `cb-sim --runs 10` | FAIL — *"run 1 hash … != run 4 hash … (of 10)"* |
|
||||
|
||||
A late-onset divergence is a real class the double-run structurally cannot
|
||||
see, and it is **deterministic, not probabilistic** — so it can be a
|
||||
control rather than a coin flip. The spec value stands unamended; what
|
||||
changed is that it is now enforced, on one scenario in `make am8` (~2 s)
|
||||
rather than on all 25.
|
||||
|
||||
The primary defence against the *probabilistic* class remains the
|
||||
`HashMap`/`HashSet` deny lint — this row's other clause, already live.
|
||||
Ten runs are defence in depth against that exclusion failing, which is why
|
||||
one workload's worth is proportionate rather than 25.
|
||||
|
||||
## 4. What the acceptance table now claims
|
||||
|
||||
| row | before | after |
|
||||
|---|---|---|
|
||||
| AM-7 | PARTIAL 2/3 | **red, 3/3** |
|
||||
| AM-8 | PARTIAL 1/2 | **red, 2/2** |
|
||||
|
||||
**M-D1-MUT: 10 of 14 rows enforced**, from 8. The prediction of record
|
||||
(≥10 of 14, from ADR-0005's 9-of-12) is **MET** for the first time.
|
||||
|
||||
**The denominator did not move, and that was the live risk.** This pass
|
||||
could have improved its score two ways, and one of them was deleting a
|
||||
question — amending AM-8's N down to 2, or splitting AM-7's scaling clause
|
||||
into a fifteenth row. Both were considered and both were refused. The four
|
||||
rows that remain unenforced are the four that were already unenforceable:
|
||||
AM-3 blocked on an unbuilt artifact, AM-4c withdrawn and retained on
|
||||
purpose, AM-5 declared ungated by the spec, AM-10 withdrawn.
|
||||
|
||||
### And the run found a third thing, in the instrument itself
|
||||
|
||||
The first full `mutation-check` of this pass reported **AM-4a
|
||||
HARNESS-BROKEN**, and correctly refused to publish any score at all. Its
|
||||
mutation still pointed at `"shipped-runtime": 250_000` — the target
|
||||
**ADR-0008 D3 moved to 161,000 in CB-WP-0013**. The find-string had been
|
||||
stale ever since, because no full mutation-check had been run in between.
|
||||
|
||||
Positive control 1 doing exactly its job. But it only *can* do that job on
|
||||
a full run, and a full run is deliberately not in `make all` — it rebuilds
|
||||
the workspace once per mutated row. So a mutation can rot for passes at a
|
||||
time while every build stays green.
|
||||
|
||||
The cheap half of that check needs no build at all: does each find-string
|
||||
still occur in the file it names? That is now a `--self-test` assertion,
|
||||
and `self-tests` **is** in `make all`. Verified both ways — perturbing the
|
||||
target by one digit turns it red naming `AM-4a`, restoring it turns it
|
||||
green.
|
||||
|
||||
Two rows were repaired by this pass and a third by the run that measured
|
||||
it. The AM-4 family has now produced four defects, none of them found by
|
||||
looking for them.
|
||||
|
||||
### The tool now measures a claim it used to assert
|
||||
|
||||
`mutation-check`'s clause flags were hand-maintained booleans saying
|
||||
whether an assertion existed — which is the same shape of claim the tool
|
||||
was built to stop trusting. A clause may now carry its own verify command
|
||||
and mutation, and then its `enforced` flag is **measured**; a declaration
|
||||
that disagrees with its own measurement is refused as HARNESS-BROKEN
|
||||
rather than reported as either verdict. Two clauses carry one so far, and
|
||||
they print as `red*`. The rest still print as the author's word, which is
|
||||
what they always were.
|
||||
|
||||
## 5. Cost, and the prediction that came due
|
||||
|
||||
| pass | kind | responses | cost | $/response |
|
||||
|---|---|---|---|---|
|
||||
| CB-WP-0013 | meta | 47 | $8.26 | 0.176 |
|
||||
| **CB-WP-0014** (final) | product | 34 | **$7.47** | **0.220** |
|
||||
| CB-WP-0015 | product | *provisional — not quoted, per CB-EV-0012 §5* | | |
|
||||
|
||||
**CB-EV-0009's standing prediction came due and is consistent — once.** It
|
||||
predicted that a pass opening above the SH-1 hard line would cost more than
|
||||
0.123 $/response. CB-WP-0014 opened above the line and cost **0.220**.
|
||||
|
||||
That is one point past a threshold, and it is worth being explicit about
|
||||
what it is not: there is **no control**. No pass has opened *below* the
|
||||
line since the prediction was made, the two passes either side differ in
|
||||
tier, kind and subject, and 0.220 is inside the range product passes have
|
||||
shown anyway. The prediction is unfalsified, not confirmed. It needs a
|
||||
pass that opens below the line to mean anything.
|
||||
|
||||
## 6. Open
|
||||
|
||||
- **AM-4b's scope defect (408,237 uncounted lines)** and its unmeasured
|
||||
proc-macro share. Unchanged from CB-EV-0012 §3.
|
||||
- **INTENT stage 1: one human verification** — `cb-play --serve 0`, open
|
||||
the URL, confirm the table reads and a drag works. Unchanged.
|
||||
- **`python3` as a toolchain dependency was never argued.**
|
||||
- **AM-4a still cannot survive stage 2** — 1,741,979 against 161,000.
|
||||
- **ADR-0007 D3's acquisition rule** remains unratified, having now
|
||||
decided two dependency questions.
|
||||
- **AM-3 stays blocked** on an artifact nobody has built, and it is now
|
||||
one of only four unenforced rows rather than one of six.
|
||||
- **`cb-cost --self-test` failed once and passed on re-run**, on
|
||||
*"pass_costs windows sum to the unwindowed total"*, inside a `make all`
|
||||
during this session. The obvious hypothesis is that it reads the session
|
||||
transcript while this session is appending to it — a self-test racing
|
||||
the file it measures. **That is a hypothesis, not a finding: it has not
|
||||
been reproduced or instrumented.** Recorded because an intermittent
|
||||
positive control is worth more attention than a failing one, and because
|
||||
a flake in the tool that prices every pass would be quietly corrosive.
|
||||
- **Chaos: 10 of 12 declarations, 1 override.**
|
||||
Loading…
Add table
Add a link
Reference in a new issue