clay-borg/evidence/CB-EV-0019-import-the-edition.md
tegwick b1f4a85a4a CB-WP-0021 T04: evidence -- the endings are tight, and the cost chain snapped
Closes CB-WP-0021, four days after its code landed.

The budget was the wrong one and ADR-0011 corrected it: csv costs 17,651
against AM-4b's 19,742, so it FITS, and was refused on proportion anyway
(89% of the remainder to read 20 rows). Third premise this pass corrected,
second against a mis-chosen denominator. That pattern now has three
instances and no control -- facts-check catches copies that disagree,
nothing catches a number computed correctly against the wrong base.

The hash blast radius was PREDICTED, not discovered. T01's brief required
it be established before writing code and ADR-0011 D3 measured it: 0
scenario files pin a state hash. Verified today, replay-test 14/14 and
AM-7 at 0.985/0.998/1.018 against a 0.9 floor. What actually broke was six
scenarios' data, not any logic.

The endings mean something and they are tight. Played 2p all-bots: total 4
/ threshold 5, group failure, 6 available. T02's control predicted a full
clear is required at 2p and play confirms it -- leaving any single Problem
unclaimed loses at two seats.

T04's last bullet cannot be satisfied. cb-cost.py --slug CB-WP-0020 ABORTS
-- no transcripts retained -- and the instrument refusing to estimate is
correct behaviour, so the number is absent rather than wrong. 41.6% of
spend is UNATTRIBUTED. The quote-the-previous-pass rule breaks after about
four passes and needs a bound; this is the first pass where the chain
actually snapped.

loop-lint's own-cost rule caught the first draft quoting this pass's own
cost as final. Marked provisional -- the rule working.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:56:31 +02:00

176 lines
7.2 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# CB-EV-0019 — import the edition
CB-WP-0021 T04. Tier M (structural M — adds or refuses an external
dependency and changes how a game is set up; chaos d8=7 → no override).
Closed 2026-08-05, four days after the code landed.
**Delivered:** [ADR-0011](../decisions/ADR-0011-vendor-the-edition.md),
`editions/ground-darvo-r0/Problems.csv` with a digest, a ~150-line reader
in `games/ground/src/edition.rs`, `tools/edition-check.py`, and the ruled
deal — Surface hidden `1..=k`.
---
## 1. What the import cost, and the budget it was measured against
**The declaration's constraint was measured against the wrong
denominator, and the ADR corrected it.** CB-WP-0021 said a CSV crate costs
21,613 lines against AM-4a's **3,798** of headroom — *"5.7× over, settled
by measurement rather than preference."*
`setup` and `problem_priorities` are `#[cfg(feature = "scenarios")]`.
**They are not in the shipped runtime**, so AM-4a never sees them. Against
AM-4b's graph:
| | lines |
|---|---:|
| AM-4b headroom | 19,742 |
| `csv` marginal (`csv` + `csv-core`; `ryu`/`itoa`/`memchr` already present) | **17,651** |
**It fits, with 2,091 to spare — and was refused anyway**, on proportion:
89% of everything the budget had left, to read 20 rows. What shipped is a
hand-rolled reader we own.
**That is the third premise this pass had to correct**, and the second
where a real number was computed against a mis-chosen denominator. The
pattern is now well enough attested to name: *this project's arithmetic
errors are not in the arithmetic.* CB-WP-0022 T02 produced two more of the
same shape four days later.
Current dependency position, `make dep-weight`:
```
AM-4c shipped-runtime 7,976 own lines per 100k third-party (diagnostic)
AM-4c dev-toolchain 1,729 own lines per 100k third-party (diagnostic)
```
## 2. What broke — and the hash blast radius was predicted, not discovered
The task asked which it was. **Predicted, and then measured to be almost
nothing.**
T01's brief said Problem values and suits enter `GroundState`, which is
hashed (K7), so *"every recorded state hash changes"*, and required the
blast radius be established **before writing code**. ADR-0011 §Decision 3
did exactly that:
| | |
|---|---:|
| scenario files pinning a state hash | **0** |
So the feared radius was ~0. Scenarios assert on state *blocks*, not
hashes; `replay-test` re-executes bundles and restores them; AM-7's probe
asserts per-segment hashes within a run rather than against recorded
literals.
**No hash was grandfathered and none recorded as "was"** — ADR-0011's
words: *a recorded hash that outliving the content it describes is a lie
with a timestamp.*
Verified today:
- `make replay-test`**14/14 controls passed**, including *"the bundle
still replays after every control restored it"*;
- `make am7` — scaling **0.985 / 0.998 / 1.018** (worst/median/best of 9)
against a 0.9 floor. Green, after T06 fixed the *measurement* rather
than lowering the floor.
**What did break** was the scenario corpus's data assumptions: six
scenario files needed value/suit updates (`gr-a02-solve`,
`gr-e03-common-problem`, `gr-e04-coalitions`, `gr-f02-no-gate`,
`gr-p05-solve-legality`) because they encoded stand-in numbers. That is
the expected cost of replacing invented data with real data, and every one
was a data edit rather than a logic change.
## 3. Do the endings mean anything now? Yes — and the margin is one point
This is the question the pass existed for. Played today, 2 players,
all-bots:
```
OUTCOME total 4 / threshold 5 group failure
personal: P1 2 P2 2
mastery 2
problems: [1] Repair 2 claimed by P2 [2] Clarify 2 [3] Boundary 2 claimed by P1
```
**Three Problems worth 2 each — 6 available against a threshold of 5.**
The bots claimed two of three, scored 4, and lost by one point.
**This confirms T02's control by play.** That control said: *"at 2p that
means a **full clear**, since 2+2+2 against 5 leaves no slack."* It is
exactly right — at two seats, leaving any single Problem unclaimed loses
the game. The endings are not merely reachable; they are **tight**, which
is a far better answer than "reachable" and was not guaranteed by the
ruling.
**And it retires the finding this pass was built around.** GR-E01 was
raised as *"unreachable below 5 seats"* on 4/6/9 against 5/7/9. The real
deal is 6/9/12 against 5/7/9. `gr-e01-threshold-unreachable-2p` is now
`gr-e01-threshold-reachable-2p`, non-provisional, and the finding was
**formally withdrawn to ground-game on 2026-08-05** (CB-WP-0022 T06) —
because it took an adversarial review, four days later, for anyone to
notice that this pass had already falsified it.
**That is the sharpest thing in this evidence file.** The work that
disproved GR-E01 is *in this workplan*. The workplan that quoted GR-E01 as
its founding example was written the next day. Nothing connected them
until a separate agent ran `git log`.
## 4. Quoting CB-WP-0020's cost — the instrument refuses
T04 asked for CB-WP-0020's cost by re-running the instrument. It cannot be
supplied:
```
$ python3 tools/cb-cost.py --slug CB-WP-0020 --by-task
ABORT — no transcripts found for CB-WP-0020
```
**The instrument aborts rather than estimating, which is correct
behaviour** and is why the number is absent instead of wrong. CB-WP-0020's
transcripts are outside the retained window.
What the instrument does report, run today:
```
main $ 548.03 output 2,125,069 tok $ 56.16 10.1%
subagent tree $ 5.50 cache_read 750,152,134 tok $394.44 71.3%
TOTAL $ 553.53 write_1h 9,233,711 tok $ 97.48 17.6%
UNATTRIBUTED $230.34 41.6%
CB-WP-0021-T01 $ 17.15 3.1% PROVISIONAL — this pass's own
CB-WP-0021-T06 $ 14.31 2.6% PROVISIONAL — this pass's own
```
**The two CB-WP-0021 rows are provisional and must not be quoted as
final.** They are this pass's own spend, measured while writing this file,
so the number is still moving — `loop-lint`'s `own-cost` rule caught the
first draft asserting them flat, which is the rule working.
**41.6% unattributed is the number worth carrying forward**, not any
per-task figure. A cost instrument that cannot say where two-fifths of the
spend went is a partial instrument, and the retention window is why the
chain of *"quote the previous pass's cost"* — a rule designed to keep
costs comparable across passes — **breaks after roughly four passes.**
That rule should either be bounded to what the window holds, or the window
extended. It is not this pass's to decide; it is recorded here because
this is the first pass where the chain actually snapped.
## 5. Chaos window 2
**Declaration 4 of 12.** Structural M, d8 = 7, no override. The roll is
recorded per §Loop tiers even though it changed nothing.
Window 2 standing: no override has yet changed an outcome. The retirement
condition — *retire if an override changes nothing twice running* — has
not been triggered, because there have been no overrides to evaluate.
## Open after this pass
- **The cost chain is broken** (§4) and the rule that depends on it needs a
bound.
- The wrong-denominator pattern (§1) now has three instances and no
control. `facts-check` catches copies that disagree; nothing catches a
number computed correctly against the wrong base.