clay-borg/workplans/CB-WP-0021-import-the-edition.md
tegwick b1f4a85a4a CB-WP-0021 T04: evidence -- the endings are tight, and the cost chain snapped
Closes CB-WP-0021, four days after its code landed.

The budget was the wrong one and ADR-0011 corrected it: csv costs 17,651
against AM-4b's 19,742, so it FITS, and was refused on proportion anyway
(89% of the remainder to read 20 rows). Third premise this pass corrected,
second against a mis-chosen denominator. That pattern now has three
instances and no control -- facts-check catches copies that disagree,
nothing catches a number computed correctly against the wrong base.

The hash blast radius was PREDICTED, not discovered. T01's brief required
it be established before writing code and ADR-0011 D3 measured it: 0
scenario files pin a state hash. Verified today, replay-test 14/14 and
AM-7 at 0.985/0.998/1.018 against a 0.9 floor. What actually broke was six
scenarios' data, not any logic.

The endings mean something and they are tight. Played 2p all-bots: total 4
/ threshold 5, group failure, 6 available. T02's control predicted a full
clear is required at 2p and play confirms it -- leaving any single Problem
unclaimed loses at two seats.

T04's last bullet cannot be satisfied. cb-cost.py --slug CB-WP-0020 ABORTS
-- no transcripts retained -- and the instrument refusing to estimate is
correct behaviour, so the number is absent rather than wrong. 41.6% of
spend is UNATTRIBUTED. The quote-the-previous-pass rule breaks after about
four passes and needs a bound; this is the first pass where the chain
actually snapped.

loop-lint's own-cost rule caught the first draft quoting this pass's own
cost as final. Marked provisional -- the rule working.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:56:31 +02:00

318 lines
13 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
id: CB-WP-0021
kind: product
title: "Import the edition: the game plays its own data"
status: done
state_hub_workstream_id: "782b1c37-f3a7-469b-87a3-fa73ebe758d2"
---
# Purpose
```
structural tier M (adds or refuses an external dependency, and changes
how a game is set up — a canonical interface)
chaos d8 = 7 → no override
declared tier M
```
Declaration 4 of chaos window 2.
## The engine has been playing a stand-in
`games/ground/src/lib.rs` builds Problems with `value: priority` and suits
cycled by index. `editions/ground-darvo-r0/Problems.csv` has carried the
real thing since 2026-07-31, and **GROUND-WP-0002 T01 ruled it
authoritative on 2026-08-03**:
| | problems/scenario | values | total | suits |
|---|---:|---|---:|---|
| `Problems.csv` | **5** | 2, 2, 2, 3, 3 | **12** | `required_solution` per problem |
| the stand-in | 3 | 1, 2, 3 | **6** | cycled by index |
## CORRECTION, before any code: the import does not fix GR-E01
**This declaration opened by claiming it would, and the data says
otherwise.** The claim was *"ordinary against 12, unreachable against 6 —
it was never a rules gap."* Measured across all four scenarios:
**GR-S01 deals 2 / 3 / 4 problems by player count, not all five.** So the
points actually in play are never 12:
| players | dealt | available | GR-E01 threshold | |
|---:|---:|---:|---:|---|
| 2 | 2 | **4** | 5 | unreachable |
| 34 | 3 | **6** | 7 | unreachable |
| 56 | 4 | **9** | 9 | reachable, exactly |
Identical in shape to the stand-in, which gave 3 / 6 / 10 against the same
5 / 7 / 9. **So `GR-E01 unreachable below 5 seats` is a real property of
the game, not an artifact of the stand-in**, and
`gr-e01-threshold-unreachable-2p` is asserting something true.
The error was mine and it was the cheap kind to make: 12 points exist in
the file, so I assumed 12 points are in play. **One command over the CSV
settled it, and it was not run until after the declaration was
committed** — the characteristic error of this project, in the pass that
followed a ruling obtained *because* of it.
**CB-EV-0018 is corrected too.** It said the `0` scores were *"confirmed
as the stand-in's doing, not a scoring bug."* Overstated: the zero came
from no Problem being claimed at all, and the threshold gap is
independent of which dataset is loaded.
**What this changes.** The import is still right — the suits, values and
visibility are authoritative and the engine should stop inventing them.
But it resolves nothing about GR-E01, and T03 must send ground-game a
*sharper* question rather than a retirement: **either GR-S01's deal count
is wrong or GR-E01's thresholds are, and no dataset can reconcile them.**
## SUPERSEDED 2026-08-04 — ground-game ruled, and it was the deal
The question above was sent and **answered** (GROUND-WP-0004, ruled
2026-08-04). It was GR-S01:
> **Surface always hidden priorities `1..k`**, with `k` = 2 (2p), 3
> (34p), 4 (56p). *"Surface is never one of the hidden priority slots."*
So the deal is **3 / 4 / 5** Problems, not 2 / 3 / 4. Available points
become **6 / 9 / 12**, thresholds **5 / 7 / 9 stand**, and SHARED GROUND
is a **26 player** mode as printed — not 56p only.
**The import is now load-bearing rather than merely correct.** The ruled
6 / 9 / 12 holds *only* with `Problems.csv` values (2,2,2,3,3). With the
stand-in's `value = priority` the same deal gives **6 / 10 / 15** — a
different game that happens to also be winnable. **Fixing the deal without
importing the data would produce numbers ground-game did not rule on.**
The two are therefore one pass, and this workplan is re-scoped to do both.
**A subtlety the ruling keeps, worth not smoothing away:** at 2p, values
2+2+2 against a threshold of 5 still require a **full clear** — any two
cards sum to 4. *"Reachable ≠ forgiving."*
## The constraint, measured before declaring
**AM-4a has 3,798 lines of headroom.** A CSV crate costs, marginally
against the shipped graph:
| crate | lines |
|---|---:|
| `csv` | 14,291 |
| `csv-core` | 3,360 |
| `ryu` | 3,962 |
| **marginal total** | **21,613** |
**5.7× the available headroom.** `itoa`, `memchr`, `serde` and
`serde_core` are already present and cost nothing; the parser itself is
what does not fit.
So the shipped runtime cannot gain a CSV parser, and that is settled by
measurement rather than by preference. What remains is *which* of the
alternatives to take, and that is the ADR.
## Task: decide how the data reaches the game
```task
id: CB-WP-0021-T01
status: done
priority: high
state_hub_task_id: "68e4fe63-eec6-4fb8-a84f-32c7edee19af"
```
Write `decisions/ADR-0011-*.md` (tier M: survey and decision in one).
**Three questions, and the third is the one that bites.**
**(a) Where does the data live?** `ground-game` is a separate repository.
Depending on a sibling checkout makes the build depend on a path that may
not exist; vendoring a copy makes clay-borg carry content it does not own.
Whichever is chosen must say how a copy is known to be current, because a
silently stale copy is worse than no copy.
**(b) How is it parsed, given a CSV parser does not fit?** Candidates, and
each needs its marginal cost stated rather than assumed:
- hand-rolled reader in `games/ground` — small, but it is a parser we
then own, and `problem_text` contains commas;
- **bake it**: a dev-time step converts the CSV into a generated Rust
module or compact literal, so the shipped runtime parses **nothing**.
Then the generated artifact must be proven to match its source, which
is the DFD class this repo already has machinery for;
- load nothing at runtime and treat the data as scenario input.
**(c) What does this do to determinism?** Problem values and required
suits become part of `GroundState`, which is hashed (K7). **Every recorded
state hash changes.** Before writing code, establish what actually pins a
hash today — `make sim` runs 25 scenarios, `replay-test` re-executes
bundles, and AM-7's probe asserts per-segment hashes. Say which of those
break, and whether any of them are *supposed* to be stable across a
content change.
**This is the reason the ADR exists.** A content import that quietly
invalidates every recorded hash, in a project whose central invariant is
replay determinism, is not a data-loading change.
## Task: deal what GR-S01 now says
```task
id: CB-WP-0021-T05
status: done
priority: high
state_hub_task_id: "c77c0b39-0841-40bf-8078-135486d7ed55"
```
`problem_priorities` returns 2 / 3 / 4 and `setup` deals priorities
`1..=N` with priority 1 as the face-up Surface. That is **Surface +
(N1) hidden**. The ruling is **Surface + k hidden**, k = 2 / 3 / 4.
Do this **with** T02, not before it: the ruled arithmetic is a property of
the deal *and* the data together, and landing one without the other gives
numbers nobody ruled on.
**Controls:**
- for every seat band, `sum(point_value of dealt) ≥ threshold` — this is
the check ground-game asked for by name, and it replaces
`gr-e01-threshold-unreachable-2p`;
- the Surface Problem is dealt at every seat count and is **never** one of
the hidden slots;
- **`gd0001` must be inverted**, not deleted. It asserts unreachability at
2/3/4 seats and is the record of why this changed; inverting it is how a
reader learns the game became winnable rather than that a test vanished.
## Task: import it, and let the thresholds mean something
```task
id: CB-WP-0021-T02
status: done
priority: high
state_hub_task_id: "28c3ff2c-16ae-47b5-9474-10e754936c60"
```
Replace the stand-in. `setup` must build Problems from the edition data:
`point_value`, `required_solution`, `visibility` (Surface → face up,
Hidden → face down) and `hidden_priority`.
**Controls:**
- **The stand-in must become unreachable.** A test that the fixture
constants no longer appear — a `value: priority` that survives beside
real data is a fallback nobody will notice until the numbers look odd
again.
- **A scenario must be able to reach GR-E01's threshold at 2 players**,
which is the specific thing that was impossible. Assert the arithmetic,
not just that a game runs — and at 2p that means a **full clear**, since
2+2+2 against 5 leaves no slack.
- The deal shape must survive: code that assumed 2/3/4 Problems will not
announce itself when it becomes 3/4/5.
## Task: retire the rules gap that was never one
```task
id: CB-WP-0021-T03
status: done
priority: medium
state_hub_task_id: "ff3bd923-9066-49ce-aadd-a3552e4964ff"
```
`gr-e01-threshold-unreachable-2p` is tagged `provisional: true` with
`provisional_owner: ground-game`, and its description says group success
*"is unreachable at 2, 3 and 4 players"*. If T02 lands, that scenario is
asserting a property of the **stand-in**, not of the game.
**Ruled 2026-08-04:** *"Not a design gap. Rewrite as a non-provisional
import/fixture check: for every seat band, `sum(point_value of dealt
Problems) ≥ threshold`."*
So it is **rewritten, not retired**, and it stops being provisional —
`provisional_owner: ground-game` comes off because ground-game has now
ruled. Message `ground-game` with the outcome.
Do **not** quietly delete a failing-in-fact scenario. CB-EV-0005: *a score
improved by deleting the question is not an improvement.*
## Task: fix AM-7's measurement, not its floor
```task
id: CB-WP-0021-T06
status: done
priority: high
state_hub_task_id: "be065f59-ac0e-46b1-840d-65c0754f7b8f"
```
T05 turned AM-7 red: median **0.845** against a 0.9 floor, all nine
samples below. The maintainer chose to **fix the measurement** rather than
ADR the floor or optimise the fold.
**The row was measuring the wrong thing.** It folded a 5,000-event log and
a 100,000-event log and compared throughputs, which confounds:
1. does cost per event grow with how many events have been folded? — the
property AM-7 claims; and
2. does streaming a 20× longer `Vec` cost more per element? — a
memory-hierarchy fact true of any program.
It measured (2) and reported it as (1). Importing the edition enlarged the
aggregate and the ratio fell, **with the state bounded**.
**Corrected:** time a 5,000-event window on a state at depth 0, and *the
same events* on a state at depth 100,000. Equal windows, equal event mix;
the only difference left is history depth.
| | clean | history-proportional mutation |
|---|---|---|
| corrected | **1.004** | **0.589 — red** |
| old | 0.845 (red on healthy code) | 0.751 |
Renamed `am7_cost_per_event_does_not_grow_with_history`, because the old
name described the confounded measurement.
**Two of my own measurements in this task were wrong, and both were caught
by measuring again.** A 2-minute timeout killed the shell line before its
restoring `cp` ran, so the next three readings were taken on **mutated
code** — I diagnosed a "second confound" from event-mix that did not
exist, and "fixed" it by folding identical events on both legs. That
change is kept, on its own merits: identical events remove a real
potential confound. But the justification I gave for it was fiction, and
the probe that proved state was bounded had only checked four of eleven
collections.
## Task: evidence
```task
id: CB-WP-0021-T04
status: done
priority: high
state_hub_task_id: "ba76138d-a225-470d-bb2a-3a6881f4ca82"
```
`evidence/CB-EV-0019-*.md`.
- **What the import cost**, against the 3,798 lines that were available.
- **What broke**, especially hashes, and whether the blast radius was
predicted in T01 or discovered in T02. If it was discovered, say so —
that is the ADR having missed something.
- **Whether the endings now mean anything**: play one and report the
score against the threshold.
- **Quote CB-WP-0020's cost by re-running the instrument.**
- **Chaos: 4 of 12 in window 2.**
**Done 2026-08-05.**
[CB-EV-0019](../evidence/CB-EV-0019-import-the-edition.md).
- **The budget was the wrong one, and the ADR corrected it**: `csv` costs
17,651 against AM-4b's 19,742, so it *fits* — and is refused on
proportion (89% of the remainder to read 20 rows). Third premise this
pass corrected, second against a mis-chosen denominator.
- **The hash blast radius was predicted, then measured to be ~0** — ADR-0011
§D3 found **0 scenario files pin a state hash**. Verified today:
`replay-test` 14/14, AM-7 0.985/0.998/1.018 against a 0.9 floor. What
actually broke was six scenarios' *data*, not any logic.
- **The endings mean something, and they are tight.** Played 2p all-bots:
**total 4 / threshold 5, group failure**, 6 available. T02's control
predicted a full clear is required at 2p — confirmed by play.
- **The cost chain snapped.** `cb-cost.py --slug CB-WP-0020` **aborts**
no transcripts retained — so T04's last bullet cannot be satisfied, and
the instrument refusing to estimate is the correct behaviour. 41.6% of
spend is UNATTRIBUTED. The quote-the-previous-pass rule needs a bound.
- **The sharpest thing in it**: the work that disproved GR-E01 is *in this
workplan*, and CB-WP-0022 was written the next day quoting GR-E01 as its
founding example. Nothing connected them until a separate agent ran
`git log` four days later.