CB-WP-0021 T04: evidence -- the endings are tight, and the cost chain snapped
Closes CB-WP-0021, four days after its code landed. The budget was the wrong one and ADR-0011 corrected it: csv costs 17,651 against AM-4b's 19,742, so it FITS, and was refused on proportion anyway (89% of the remainder to read 20 rows). Third premise this pass corrected, second against a mis-chosen denominator. That pattern now has three instances and no control -- facts-check catches copies that disagree, nothing catches a number computed correctly against the wrong base. The hash blast radius was PREDICTED, not discovered. T01's brief required it be established before writing code and ADR-0011 D3 measured it: 0 scenario files pin a state hash. Verified today, replay-test 14/14 and AM-7 at 0.985/0.998/1.018 against a 0.9 floor. What actually broke was six scenarios' data, not any logic. The endings mean something and they are tight. Played 2p all-bots: total 4 / threshold 5, group failure, 6 available. T02's control predicted a full clear is required at 2p and play confirms it -- leaving any single Problem unclaimed loses at two seats. T04's last bullet cannot be satisfied. cb-cost.py --slug CB-WP-0020 ABORTS -- no transcripts retained -- and the instrument refusing to estimate is correct behaviour, so the number is absent rather than wrong. 41.6% of spend is UNATTRIBUTED. The quote-the-previous-pass rule breaks after about four passes and needs a bound; this is the first pass where the chain actually snapped. loop-lint's own-cost rule caught the first draft quoting this pass's own cost as final. Marked provisional -- the rule working. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
5e5894f47d
commit
b1f4a85a4a
2 changed files with 201 additions and 2 deletions
176
evidence/CB-EV-0019-import-the-edition.md
Normal file
176
evidence/CB-EV-0019-import-the-edition.md
Normal file
|
|
@ -0,0 +1,176 @@
|
|||
# CB-EV-0019 — import the edition
|
||||
|
||||
CB-WP-0021 T04. Tier M (structural M — adds or refuses an external
|
||||
dependency and changes how a game is set up; chaos d8=7 → no override).
|
||||
Closed 2026-08-05, four days after the code landed.
|
||||
|
||||
**Delivered:** [ADR-0011](../decisions/ADR-0011-vendor-the-edition.md),
|
||||
`editions/ground-darvo-r0/Problems.csv` with a digest, a ~150-line reader
|
||||
in `games/ground/src/edition.rs`, `tools/edition-check.py`, and the ruled
|
||||
deal — Surface ∪ hidden `1..=k`.
|
||||
|
||||
---
|
||||
|
||||
## 1. What the import cost, and the budget it was measured against
|
||||
|
||||
**The declaration's constraint was measured against the wrong
|
||||
denominator, and the ADR corrected it.** CB-WP-0021 said a CSV crate costs
|
||||
21,613 lines against AM-4a's **3,798** of headroom — *"5.7× over, settled
|
||||
by measurement rather than preference."*
|
||||
|
||||
`setup` and `problem_priorities` are `#[cfg(feature = "scenarios")]`.
|
||||
**They are not in the shipped runtime**, so AM-4a never sees them. Against
|
||||
AM-4b's graph:
|
||||
|
||||
| | lines |
|
||||
|---|---:|
|
||||
| AM-4b headroom | 19,742 |
|
||||
| `csv` marginal (`csv` + `csv-core`; `ryu`/`itoa`/`memchr` already present) | **17,651** |
|
||||
|
||||
**It fits, with 2,091 to spare — and was refused anyway**, on proportion:
|
||||
89% of everything the budget had left, to read 20 rows. What shipped is a
|
||||
hand-rolled reader we own.
|
||||
|
||||
**That is the third premise this pass had to correct**, and the second
|
||||
where a real number was computed against a mis-chosen denominator. The
|
||||
pattern is now well enough attested to name: *this project's arithmetic
|
||||
errors are not in the arithmetic.* CB-WP-0022 T02 produced two more of the
|
||||
same shape four days later.
|
||||
|
||||
Current dependency position, `make dep-weight`:
|
||||
|
||||
```
|
||||
AM-4c shipped-runtime 7,976 own lines per 100k third-party (diagnostic)
|
||||
AM-4c dev-toolchain 1,729 own lines per 100k third-party (diagnostic)
|
||||
```
|
||||
|
||||
## 2. What broke — and the hash blast radius was predicted, not discovered
|
||||
|
||||
The task asked which it was. **Predicted, and then measured to be almost
|
||||
nothing.**
|
||||
|
||||
T01's brief said Problem values and suits enter `GroundState`, which is
|
||||
hashed (K7), so *"every recorded state hash changes"*, and required the
|
||||
blast radius be established **before writing code**. ADR-0011 §Decision 3
|
||||
did exactly that:
|
||||
|
||||
| | |
|
||||
|---|---:|
|
||||
| scenario files pinning a state hash | **0** |
|
||||
|
||||
So the feared radius was ~0. Scenarios assert on state *blocks*, not
|
||||
hashes; `replay-test` re-executes bundles and restores them; AM-7's probe
|
||||
asserts per-segment hashes within a run rather than against recorded
|
||||
literals.
|
||||
|
||||
**No hash was grandfathered and none recorded as "was"** — ADR-0011's
|
||||
words: *a recorded hash that outliving the content it describes is a lie
|
||||
with a timestamp.*
|
||||
|
||||
Verified today:
|
||||
|
||||
- `make replay-test` — **14/14 controls passed**, including *"the bundle
|
||||
still replays after every control restored it"*;
|
||||
- `make am7` — scaling **0.985 / 0.998 / 1.018** (worst/median/best of 9)
|
||||
against a 0.9 floor. Green, after T06 fixed the *measurement* rather
|
||||
than lowering the floor.
|
||||
|
||||
**What did break** was the scenario corpus's data assumptions: six
|
||||
scenario files needed value/suit updates (`gr-a02-solve`,
|
||||
`gr-e03-common-problem`, `gr-e04-coalitions`, `gr-f02-no-gate`,
|
||||
`gr-p05-solve-legality`) because they encoded stand-in numbers. That is
|
||||
the expected cost of replacing invented data with real data, and every one
|
||||
was a data edit rather than a logic change.
|
||||
|
||||
## 3. Do the endings mean anything now? Yes — and the margin is one point
|
||||
|
||||
This is the question the pass existed for. Played today, 2 players,
|
||||
all-bots:
|
||||
|
||||
```
|
||||
OUTCOME total 4 / threshold 5 group failure
|
||||
personal: P1 2 P2 2
|
||||
mastery 2
|
||||
problems: [1] Repair 2 claimed by P2 [2] Clarify 2 [3] Boundary 2 claimed by P1
|
||||
```
|
||||
|
||||
**Three Problems worth 2 each — 6 available against a threshold of 5.**
|
||||
The bots claimed two of three, scored 4, and lost by one point.
|
||||
|
||||
**This confirms T02's control by play.** That control said: *"at 2p that
|
||||
means a **full clear**, since 2+2+2 against 5 leaves no slack."* It is
|
||||
exactly right — at two seats, leaving any single Problem unclaimed loses
|
||||
the game. The endings are not merely reachable; they are **tight**, which
|
||||
is a far better answer than "reachable" and was not guaranteed by the
|
||||
ruling.
|
||||
|
||||
**And it retires the finding this pass was built around.** GR-E01 was
|
||||
raised as *"unreachable below 5 seats"* on 4/6/9 against 5/7/9. The real
|
||||
deal is 6/9/12 against 5/7/9. `gr-e01-threshold-unreachable-2p` is now
|
||||
`gr-e01-threshold-reachable-2p`, non-provisional, and the finding was
|
||||
**formally withdrawn to ground-game on 2026-08-05** (CB-WP-0022 T06) —
|
||||
because it took an adversarial review, four days later, for anyone to
|
||||
notice that this pass had already falsified it.
|
||||
|
||||
**That is the sharpest thing in this evidence file.** The work that
|
||||
disproved GR-E01 is *in this workplan*. The workplan that quoted GR-E01 as
|
||||
its founding example was written the next day. Nothing connected them
|
||||
until a separate agent ran `git log`.
|
||||
|
||||
## 4. Quoting CB-WP-0020's cost — the instrument refuses
|
||||
|
||||
T04 asked for CB-WP-0020's cost by re-running the instrument. It cannot be
|
||||
supplied:
|
||||
|
||||
```
|
||||
$ python3 tools/cb-cost.py --slug CB-WP-0020 --by-task
|
||||
ABORT — no transcripts found for CB-WP-0020
|
||||
```
|
||||
|
||||
**The instrument aborts rather than estimating, which is correct
|
||||
behaviour** and is why the number is absent instead of wrong. CB-WP-0020's
|
||||
transcripts are outside the retained window.
|
||||
|
||||
What the instrument does report, run today:
|
||||
|
||||
```
|
||||
main $ 548.03 output 2,125,069 tok $ 56.16 10.1%
|
||||
subagent tree $ 5.50 cache_read 750,152,134 tok $394.44 71.3%
|
||||
TOTAL $ 553.53 write_1h 9,233,711 tok $ 97.48 17.6%
|
||||
|
||||
UNATTRIBUTED $230.34 41.6%
|
||||
CB-WP-0021-T01 $ 17.15 3.1% PROVISIONAL — this pass's own
|
||||
CB-WP-0021-T06 $ 14.31 2.6% PROVISIONAL — this pass's own
|
||||
```
|
||||
|
||||
**The two CB-WP-0021 rows are provisional and must not be quoted as
|
||||
final.** They are this pass's own spend, measured while writing this file,
|
||||
so the number is still moving — `loop-lint`'s `own-cost` rule caught the
|
||||
first draft asserting them flat, which is the rule working.
|
||||
|
||||
**41.6% unattributed is the number worth carrying forward**, not any
|
||||
per-task figure. A cost instrument that cannot say where two-fifths of the
|
||||
spend went is a partial instrument, and the retention window is why the
|
||||
chain of *"quote the previous pass's cost"* — a rule designed to keep
|
||||
costs comparable across passes — **breaks after roughly four passes.**
|
||||
|
||||
That rule should either be bounded to what the window holds, or the window
|
||||
extended. It is not this pass's to decide; it is recorded here because
|
||||
this is the first pass where the chain actually snapped.
|
||||
|
||||
## 5. Chaos window 2
|
||||
|
||||
**Declaration 4 of 12.** Structural M, d8 = 7, no override. The roll is
|
||||
recorded per §Loop tiers even though it changed nothing.
|
||||
|
||||
Window 2 standing: no override has yet changed an outcome. The retirement
|
||||
condition — *retire if an override changes nothing twice running* — has
|
||||
not been triggered, because there have been no overrides to evaluate.
|
||||
|
||||
## Open after this pass
|
||||
|
||||
- **The cost chain is broken** (§4) and the rule that depends on it needs a
|
||||
bound.
|
||||
- The wrong-denominator pattern (§1) now has three instances and no
|
||||
control. `facts-check` catches copies that disagree; nothing catches a
|
||||
number computed correctly against the wrong base.
|
||||
|
|
@ -2,7 +2,7 @@
|
|||
id: CB-WP-0021
|
||||
kind: product
|
||||
title: "Import the edition: the game plays its own data"
|
||||
status: active
|
||||
status: done
|
||||
state_hub_workstream_id: "782b1c37-f3a7-469b-87a3-fa73ebe758d2"
|
||||
---
|
||||
|
||||
|
|
@ -278,7 +278,7 @@ collections.
|
|||
|
||||
```task
|
||||
id: CB-WP-0021-T04
|
||||
status: todo
|
||||
status: done
|
||||
priority: high
|
||||
state_hub_task_id: "ba76138d-a225-470d-bb2a-3a6881f4ca82"
|
||||
```
|
||||
|
|
@ -293,3 +293,26 @@ state_hub_task_id: "ba76138d-a225-470d-bb2a-3a6881f4ca82"
|
|||
score against the threshold.
|
||||
- **Quote CB-WP-0020's cost by re-running the instrument.**
|
||||
- **Chaos: 4 of 12 in window 2.**
|
||||
|
||||
**Done 2026-08-05.**
|
||||
[CB-EV-0019](../evidence/CB-EV-0019-import-the-edition.md).
|
||||
|
||||
- **The budget was the wrong one, and the ADR corrected it**: `csv` costs
|
||||
17,651 against AM-4b's 19,742, so it *fits* — and is refused on
|
||||
proportion (89% of the remainder to read 20 rows). Third premise this
|
||||
pass corrected, second against a mis-chosen denominator.
|
||||
- **The hash blast radius was predicted, then measured to be ~0** — ADR-0011
|
||||
§D3 found **0 scenario files pin a state hash**. Verified today:
|
||||
`replay-test` 14/14, AM-7 0.985/0.998/1.018 against a 0.9 floor. What
|
||||
actually broke was six scenarios' *data*, not any logic.
|
||||
- **The endings mean something, and they are tight.** Played 2p all-bots:
|
||||
**total 4 / threshold 5, group failure**, 6 available. T02's control
|
||||
predicted a full clear is required at 2p — confirmed by play.
|
||||
- **The cost chain snapped.** `cb-cost.py --slug CB-WP-0020` **aborts** —
|
||||
no transcripts retained — so T04's last bullet cannot be satisfied, and
|
||||
the instrument refusing to estimate is the correct behaviour. 41.6% of
|
||||
spend is UNATTRIBUTED. The quote-the-previous-pass rule needs a bound.
|
||||
- **The sharpest thing in it**: the work that disproved GR-E01 is *in this
|
||||
workplan*, and CB-WP-0022 was written the next day quoting GR-E01 as its
|
||||
founding example. Nothing connected them until a separate agent ran
|
||||
`git log` four days later.
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue