CB-WP-0006 T04: withdraw AM-4c; and fix where AM-6 is measured

AM-4c is withdrawn from the acceptance table and retained as a reported
diagnostic. GameKernel §5a carries the argument.

The ratio has no monotone better direction. INTENT's rule is "own the
semantics, assimilate the implementation": rising can mean owning
semantics properly or reimplementing what should have been assimilated;
falling can mean leverage or dependency bloat. A target requires knowing
which way is better. It is also redundant — AM-4a/AM-4b bound the
denominator and AM-2 bounds own-source density, so AM-4c is a ratio of two
already-targeted quantities.

Measured at withdrawal: 1,426 own lines per 100k third-party (shipped),
1,107 (dev). make dep-weight now prints both, labelled diagnostic — the
row was never actually reported before.

M-D1-MUT keeps AM-4c in its denominator on purpose and says so in the
output. Dropping it would move the score 7/14 -> 7/13 without enforcing
anything: a score improved by deleting the question.

Decided before Phase B deliberately, since ADR-0005 predicts own-source
growth that will move this ratio; deciding after would be the retarget
§Step 4 forbids.

A T01 correction found here. The AM-6 gate failed inside `make all` at
38,753 ev/s against 341,280 in isolation — a 9x drop, because cargo test
runs binaries and threads concurrently. A throughput assertion inside a
parallel harness measures contention, not throughput. T01's measurement
was valid; its gate placement was not.

Fixed by running it only where valid — #[ignore] plus `make am6` in
release with --test-threads=1, now 2.0M ev/s at 20.2x headroom — and not
by lowering the target, which T01 forbade. My first attempt did drift that
way, adding a debug "sanity floor" of 50,000, and was backed out: a second
threshold is still a second chance to tune.

The mutation then went SURVIVED on the first run after the move. 4,000
black_box iterations were calibrated against debug's 3.4x headroom and are
invisible against release's 20x. Raised to 100,000; back to red. A weak
mutation is not a fixed property of a row — it can become weak when the
row's measurement conditions change.

Tier S (amends one row, creates no capability), chaos d4=2, no override.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
tegwick 2026-08-01 10:06:00 +02:00
parent 83db14f2e2
commit db6445ae37
10 changed files with 145 additions and 17 deletions

View file

@ -165,7 +165,7 @@ evidence lands in `evidence/CB-EV-0001-game-kernel.md` with no
| AM-3 | Synthetic-workload definition size: LOC to express the CB-RES-0001 synthetic game on our kernel | ~36 LOC (boardgame.io, measured) | ≤ 50 LOC | measured |
| AM-4a | M-D2-DEP: third-party LOC, **shipped runtime** (`--no-default-features`) | boardgame.io: 120 npm packages / 3.9M LOC | **≤ 250,000 lines** | measured (`make dep-weight`) |
| AM-4b | M-D2-DEP: third-party LOC, **dev toolchain** (default features) | as above | **≤ 350,000 lines** | measured (`make dep-weight`) |
| AM-4c | M-D2-DEP: own source per third-party 100k lines | — | reported, not targeted | measured (`make dep-weight`) |
| ~~AM-4c~~ | M-D2-DEP: own source per third-party 100k lines | — | **WITHDRAWN from the acceptance table 2026-08-01 (CB-WP-0006 T04)** — retained as a reported diagnostic in `make dep-weight`; see §5a | diagnostic |
| AM-5 | M-D2-BLD: clean release build of headless workspace | n/a (npm install ~seconds; not comparable) | ≤ 60 s on bnt-lap001, recorded not gated | measured |
| AM-6 | M-D3-THR: applied events/s, synthetic workload, same machine | boardgame.io ~1,1001,900 moves/s (best config, degrading) | **≥ 100,000/s** (stipulated target, ADR-0002) | measured |
| AM-7 | M-D3 scaling: throughput @100k events vs @5k; and snapshot+replay of 100k events | boardgame.io 0.450.66× @2040k, DNF @100k | **≥ 0.9×** (flat), replay of 100k events ≤ 5 s, hash-identical | measured |
@ -175,6 +175,43 @@ evidence lands in `evidence/CB-EV-0001-game-kernel.md` with no
| AM-11 | M-D4-SWAP **(unmet 2026-07-31 — the pair exists, the suite does not)**: null + reference impls passing one conformance suite | no candidate has the pattern | RNG and log storage each have ≥2 impls (real + test/null) under one suite | measured (bool) |
| AM-12 | M-D2-TOK / M-D2-CST: tokens and USD per completed task | n/a — first pass sets our own baseline | recorded per task in the evidence cost log (price sheet 2026-07-31) | recorded, not gated |
### 5a. Why AM-4c was withdrawn from the acceptance table
*(CB-WP-0006 T04, 2026-08-01. Tier S — amends one row, creates no
capability; chaos d4=2, no override. Follows the precedent ADR-0005 §4 set
for AM-10.)*
AM-4c was `reported, not targeted`, so nothing could fail and it counted
against M-D1-MUT. The task was to give it a threshold or drop it. It is
dropped, for a reason that a threshold cannot fix:
**The ratio has no monotone better direction.** INTENT's rule is *own the
semantics; assimilate the implementation*. A **rising** ratio can mean we
are properly owning semantics, or that we are reimplementing things we
should have assimilated. A **falling** ratio can mean good leverage, or
dependency bloat and implementation leaking into our semantics. Both
directions are ambiguous, and a target requires knowing which way is
better.
**It is also redundant.** AM-4a and AM-4b already bound the denominator
(third-party LOC ceilings, ratified in ADR-0004) and AM-2 bounds own-source
density per rule. AM-4c is the ratio of two quantities that are each
already targeted; any threshold on it would be implied by those two or
would contradict them.
Measured at withdrawal: **1,426** own lines per 100k third-party (shipped
runtime), **1,107** (dev toolchain).
**Decided before Phase B, deliberately.** ADR-0005 predicts own-source
growth from the kernel work, which will move this ratio. Setting a
threshold after seeing that movement would be the retarget InnerLoop
§Step 4 forbids — so the decision was taken while the number was still
unaffected by the work that will change it.
**M-D1-MUT keeps AM-4c in its denominator.** Withdrawing a row would
otherwise improve the metric from 7/14 to 7/13 without enforcing anything —
a score improved by deleting the question.
Comparisons against the event-sourcing 10⁵10⁶/s estimate stay **parity**
until a local Rust comparator is measured (open follow-up from the
adversarial review).