--- id: CB-WP-0022 kind: product title: "The design instrument: findings about the game, with their reproductions" status: done state_hub_workstream_id: "fda16340-0049-4acf-884b-a5cfdbde47c0" --- # Purpose ``` structural tier L (named a high-leverage pass by the maintainer, and it amends INTENT — clay-borg gains a stated aspect) chaos d8 = 6 → no override declared tier L ``` Declaration 5 of chaos window 2. Tier L: separate survey, **adversarial review**, ADR, then spec, then code. ## The insight, in the maintainer's words > *"We should consider ourselves testing the game and document > inconsistencies to report them back to the ground-game repo, so that the > game designer can improve the rules accordingly… We should have the > rigorous game simulation engine and a meta scope to capture notes about > game design flaws, questions, results and protocols about trial games… > This will provide clay-borg an additional aspect as a valuable game > design tool."* **This is already happening and has no home.** In five passes the engine has produced, as a by-product of being rigorous: | finding | how it surfaced | |---|---| | ten underdetermined rules points (U1–U10) | formalizing the dataset into testable rules | | SOLVE offered on a face-down Problem, always inert | a human dragging it three rounds running | | GR-A13 "wasted SOLVE" on a claimed Problem | a scenario that had to pick a default | | ~~GR-E01 unreachable below 5 seats~~ | arithmetic over the deal count — **withdrawn 2026-08-05, it was wrong (C1)** | | ~~six provisional scenario defaults~~ | **five, and double-counted with GR-E01 (C3)** — they are reproductions, not a finding | | GR-E03 / GR-E04 never played to the end | nobody noticed for nineteen passes | Every one was found by *building the simulator*, not by playing. That is the thing worth naming: **a simulator rigorous enough to refuse ambiguity is a design instrument, because it cannot proceed past a rule that does not decide.** And every one of them has been carried in prose, in six different places, and one sat unread in an inbox for four days. ## The load-bearing rule this must have The project's standing failure is *unexecuted verification*. A design register that collects opinions would reproduce it in a new medium. > **A design finding is not admissible without its reproduction.** Concretely: a scenario that fails, an arithmetic check that prints the contradiction, a recorded game the reader can replay, or a named test. *"This feels unbalanced"* is a note, not a finding. **The SOLVE inertness is admissible because a recorded session shows three no-ops.** > **The example that stood here was GR-E01, and the review killed it > (C1).** *"4/6/9 against 5/7/9 is a computation anyone can rerun"* had > already been rerun: `2da19a4` measured **6/9/12**, and the scenario was > renamed `-unreachable-` → `-reachable-`. The conclusion inverted — and > it was one of the two findings that **passed** this rule. So existence > is not what was missing. See [ADR-0012](../decisions/ADR-0012-the-design-instrument.md) > D3: the rule gains **shape**, and **a reproduction must be able to > fail.** Ours went green and stayed admissible. This is what would make clay-borg a design tool rather than a suggestion box, and it is the one part of this proposal that must not be traded away for convenience. ## The judgment I want reviewed, not assumed The maintainer asked whether this should extend to *"a meta about the clay-borg engine evolution itself."* **My answer is no, and it should be argued rather than accepted.** That register already exists and is load-bearing: `evidence/CB-EV-*`, `decisions/ADR-*`, `gates.toml` and the workplans capture nineteen passes with dates, costs and falsifiers. A second register for the same subject would be ceremony. The asymmetry is the point: engine evolution has a home and game design does not. > **Settled by [ADR-0012](../decisions/ADR-0012-the-design-instrument.md) > D7: no register — but the argument above did not survive.** C5 found the > "third thing" the maintainer meant is visible in > `specs/InnerLoopReference.md` and `history/`'s retrospectives, **neither > of which this inventory names.** Conclusion narrowed, not settled: if > InnerLoopReference keeps absorbing material that is neither a decision > nor a finding, revisit. ## Task: survey how this is done elsewhere, and what we already have ```task id: CB-WP-0022-T01 status: done priority: high state_hub_task_id: "6b8663b8-4143-4361-ad9e-7df0c4f4a19d" ``` `research/CB-RES-0007-*.md`. **Do not survey issue trackers.** The question is narrower and more interesting: how do rigorous rule systems record *the ambiguity they found*? Candidates worth a benchmark-to-beat: - **Errata and rulings practice** in published games (Magic's comprehensive-rules + rulings split, Netrunner's NAPD card rulings) — what makes a ruling *findable* years later. - **Formal-methods counterexample traces** — a model checker's output is precisely a reproduction attached to a claim, which is the shape wanted here. - **Conformance-suite provisional behaviour** — how W3C/WHATWG mark "implementation-defined" and how a spec later absorbs it. - **What this repo already has**: `provisional: true` scenarios, `§Underdetermined`, `gates.toml`'s `caught`/`retire_if` shape, and the hub message that went unread. **The register must reuse the provisional machinery rather than compete with it.** Name, per dimension, the property to beat — findability, reproducibility, and whether a ruling can *close* a finding mechanically. **Done 2026-08-03.** [CB-RES-0007](../research/CB-RES-0007-design-instrument.md), with a runnable baseline (`tools/design-baseline.py`). **Its numbers were withdrawn by T02 and must not be quoted from here.** The survey reported *6 findings, 2/6 = 33% reproduced, 11 files, 4 days*. C2 showed the instrument counted itself and its reproduction check never stat'd the file; C3 showed *"six provisional defaults"* is five and GR-E01 was double-counted; T05's backfill contradicted *"six of the ten have provisional scenarios"* — **one** does. What survives is direction: many files, no index, 0 of 10 ruled. The first honest figures are T05's. **Magic corrected an assumption this pass was about to build on.** Rulings are *"reminder information with no actual weight or rules meaning"*; the authoritative fix folds into the **Oracle** card text. **A finding closes when the source changes, not when an annotation is added** — the register is a queue that empties. Model checkers supplied the reproduction rule independently, and W3C's *implementation-defined* mark is machinery we already have and must reuse rather than duplicate. ## Task: adversarial review ```task id: CB-WP-0022-T02 status: done priority: high state_hub_task_id: "d2597895-fe2e-4f1e-a3db-a5fb8833434c" ``` Tier L requires it. Give the reviewer the survey **and** the §judgment above, and require an attempt at: - **that the engine-evolution register is redundant** — the strongest counter is that ADRs record *decisions* and evidence records *findings*, but nothing records *what we learned about building engines*, which is a third thing; - **that "carries its reproduction" is affordable** — if half the real findings cannot be reproduced cheaply, the rule will be quietly dropped and the register becomes a suggestion box anyway. *(Hardened before the review ran: two findings had already reached ground-game on wrong premises, so the reviewer was told to press whether the rule is **sufficient**, not whether it is affordable.)* - **that a register is needed at all**, rather than one more section in `GroundRules.md §Underdetermined`, which already exists and already works. Record the trail in `history/`, unpolished. **Done 2026-08-05.** Trail: [challenge](../history/260805-design-instrument-challenge.md), [response](../history/260805-design-instrument-response.md). **Run by a separate agent** — the first in this repo that was. CB-RES-0006's review opened by conceding it could not be, and called its own findings *"a lower bound on what a genuinely separate reviewer would find."* That was measurable, and this is the measurement: the separate reviewer ran `git log` against the survey's central example and found our own commit had falsified it four days earlier, while the author — who wrote that commit — quoted the dead number twice. **Seven challenges: four conceded, two conceded in part, one answered.** **C1 changed the design** — the rule's showcase finding was false and had *passed* the rule, so existence is not what was missing — and **caught a defect in flight**, T06's payload still naming the dead number. C2 withdrew the baseline's precision, C3 its arithmetic, C4 flipped T03's burden toward extending `§Underdetermined`, C5 corrected the redundancy inventory. Survived: affordability, and reuse of the provisional machinery. Full account: [CB-EV-0021](../evidence/CB-EV-0021-the-design-instrument.md) §1–§3. ## Task: decide ```task id: CB-WP-0022-T03 status: done priority: high state_hub_task_id: "a2e85810-e949-4ea0-81fc-36b912af326c" ``` `decisions/ADR-0012-*.md`. At minimum: - whether clay-borg's **INTENT gains a stated aspect** as a design instrument, and in what words — this is the change with the longest half-life in the pass; - the finding **taxonomy**, and it should be grounded in the six findings above rather than invented: *underdetermined* (rules do not say), *inconsistent* (rules disagree with each other or the data), *inert* (a rule that cannot fire), *degenerate* (fires, but collapses play), *unplayed* (implemented, never played); - the **lifecycle** and who owns each state: raised → reported → ruled → applied, or withdrawn; - whether a finding without a reproduction is **rejected** or **admitted as a note** — and if admitted, how it is prevented from aging into an apparent finding. **Done 2026-08-05.** [ADR-0012](../decisions/ADR-0012-the-design-instrument.md), nine decisions. The two not on this list are the two the review forced: - **D2 — `§Underdetermined` *is* the register; nothing parallel is built.** Against the survey's own five benchmarks the incumbent already delivers four, including the Oracle property the survey went to Magic to find and we had written ourselves eight days earlier (`GroundRules.md:231-233`). What it lacks is reproductions. So this pass **extends** a section — no new file, no new schema. - **D3 — admissibility is three clauses.** Exists, has the ruled shape (row-level table, never a sum), **and can fail.** GR-E01's artifact went green and the finding stayed admissible and stayed queued, because nothing said a passing artifact was a signal. **A green reproduction is an alarm.** The rest, in one line each: **D1** INTENT gains property 4, *Instrument*, applied with its falsifier. **D4** five kinds, each forced by an existing finding. **D5** `applied` means the source changed; withdrawals are reported, not deleted. **D6** notes admitted but never reportable, 30-day expiry. **D7** no engine-evolution register, on an inventory C5 corrected. **D8** `design-baseline.py` retired. **D9** the artifact stays here, ground-game gets a generated file under its own workplan. ## Task: specify ```task id: CB-WP-0022-T04 status: done priority: high state_hub_task_id: "60ddfeab-81b1-45e4-9ec9-91cc8ea7fb72" ``` `specs/GameDesign.md`, with metrics, because a spec without them is prose. Candidate measures, to be argued not adopted: - **findings with a runnable reproduction** — target 100%, and the denominator includes withdrawn ones; - **time from raised to reported** — the U-items took four days to be *read*; that is the number this exists to fix; - **findings closed by a ruling** vs **findings still open**, with age. **ground-game has ruled on what a finding must carry** (GROUND-WP-0004 T02, 2026-08-04), and it is stricter than this pass proposed. Adopt it: > Arithmetic findings ship a **runnable reproduction** *and* a > **row-level deal table** — never only "sum of file" or "deal depth N"; > and ground-game's arithmetic rulings cite that reproduction by path. The second half is theirs to keep. **So the reproduction rule gains a shape requirement, not just an existence one** — a finding that ships a passing test but describes the wrong quantity is still a bad finding, which is exactly what happened twice. Also specify the **trial protocol**: a trial game is a `--record`ed session plus an observation log, so *"we played it and X happened"* is replayable rather than remembered. It must cost almost nothing or it will not be done. **Done 2026-08-05.** [specs/GameDesign.md](../specs/GameDesign.md) v1.0 — not a register (ADR-0012 D2 put that in `§Underdetermined`). **§1.2 is written against evidence rather than principle**: a finding must print the rows behind any number it claims. *"12" was arithmetically defensible and still wrong about the game.* **§1.3's target is `0` reproductions gone green while open** — what GR-E01 would have tripped four days before a human caught it. **No baseline rate is quoted.** **The trial protocol costs one flag**: `cb-play --record` plus a sibling `.md` in the player's own words. An observation is a **note** until it has a reproduction — *"I felt it was too easy but then we lost"* is the case it is shaped around, and a schema at the moment of observation would lose it. ## Task: build it, and backfill what is already known ```task id: CB-WP-0022-T05 status: done priority: high state_hub_task_id: "469413ba-ed47-4840-a60e-ee7f95dc06ff" ``` The register, the tool, and then **the six findings above entered into it** — backfilling is the test. A register that cannot express findings the project already has is the wrong register, and discovering that after designing it is the point of doing it in this order. `make design` (or equivalent) must report: open findings by kind, those without a reproduction, and those never reported to their owner. **Done 2026-08-05.** `tools/design.py`, `make design`, and the register in [`GroundRules.md`](../specs/GroundRules.md) — **14 rows, no new file.** Backfill was the test. The taxonomy held (five kinds, no sixth), and it **produced a `role` column ADR-0012 does not have**: the first report alarmed on U2, wrongly — a green *default* is expected, a green *counterexample* is the alarm. Folded into GameDesign §1.3. It also contradicted the survey: **one** U-item names itself in a scenario, not six. Detail and figures: [CB-EV-0021](../evidence/CB-EV-0021-the-design-instrument.md) §4, §6. ## Task: report to ground-game, mechanically ```task id: CB-WP-0022-T06 status: done priority: high state_hub_task_id: "c6a659bd-c495-4754-9332-baadad51012a" ``` Generate the report and send it. **The message that sat unread for four days is the baseline to beat** — the failure was not the message, it was that nothing pointed at it. So the report lands as a file in `ground-game` under its own workplan, extending GROUND-WP-0002 rather than duplicating it. Include the findings this pass has sharpened: - **SOLVE's legality** against a face-down Problem or an unmatchable suit — and note that the case we *reported* was not the case that fired (CB-WP-0023 T01). - ~~**GR-E01 vs GR-S01** — 4/6/9 against 5/7/9~~ — **withdrawn 2026-08-05, before sending.** C1 found `2da19a4` had measured **6/9/12**: the dataset reconciles them. It would have been the **fourth** wrong premise to reach `ground-game` and is the only one caught before transmission. **Report the withdrawal** — a claim retracted silently is how the first three survived. **Done 2026-08-05.** [`ground-game/workplans/GROUND-WP-0002-clay-borg-report-260805.md`](../../ground-game/workplans/GROUND-WP-0002-clay-borg-report-260805.md), committed there, with a hub message that only *points at* the file. **The report asks for no ruling.** It carries GR-E01's withdrawal, our own reproduction debt, and two notes that are explicitly not findings. **And it acknowledged something the pass did not expect.** GROUND-WP-0002 is `finished`: **all ten U-items were ruled 2026-08-03**, every one confirmed as the default we simulate. CB-RES-0007 said *"0 of 10 ruled"* two days later. **The unread-inbox failure running in the opposite direction** — they answered and we did not collect it. The instrument's first run surfaced it. ## Task: evidence ```task id: CB-WP-0022-T07 status: done priority: high state_hub_task_id: "4fb77911-486d-4d9e-a53d-c4e486c74ce2" ``` `evidence/CB-EV-0021-*.md`. *(Was CB-EV-0020 when written; CB-WP-0023 shipped that number first — `evidence/CB-EV-0020-solve-legality.md` — so this one moves rather than collides.)* - **Whether backfilling changed the design** — if all six findings fit the first taxonomy, say so and be suspicious of it. - **What tier L cost against what it caught**, since this is the second full-weight L pass and CB-WP-0012's deleted its own structural trigger. - **The engine-evolution question**, as the review left it. - **Quote CB-WP-0021's cost by re-running the instrument.** **Done 2026-08-05.** [CB-EV-0021](../evidence/CB-EV-0021-the-design-instrument.md). **Backfill did change the design** — and the honest answer to *"be suspicious if all six fit"* is that only **five** were entered (one was a double-count), so fitting them is close to circular. The taxonomy's real test is the seventh finding. **Tier L's cost against what it caught**: four of six catches came only from the separate reviewer, and **two came from execution rather than process** — the `role` distinction from building it, the ten uncollected rulings from running it. That is InnerLoop §Design goal's prediction holding, and an argument against front-loading more review rather than less.