Some checks failed
ci / check (push) Has been cancelled
T05. tools/design.py, make design, and the register in GroundRules.md --
14 rows, no new file, because ADR-0012 D2 made §Underdetermined the
register rather than building one beside it.
Backfill was the test and it caught two things the ADR did not have.
First, a `role` column. The first report alarmed on U2 and was wrong to:
U2's scenario is green BECAUSE the provisional default it documents is
implemented, which says nothing about whether ground-game agrees. GR-E01's
was a counterexample that went green. Same colour, opposite meaning -- a
register that cannot tell them apart either alarms constantly or never.
Only a green counterexample alarms. Folded back into GameDesign §1.3.
Second, it contradicted the survey. CB-RES-0007 said "six of the ten
already have provisional scenarios." Measured -- grep -lE "\bU<n>\b" over
scenarios/ground -- exactly ONE U-item names itself. Five provisional
scenarios exist and four probably encode U-item defaults, but the mapping
is not written down, so it is not checkable. Same defect class as the
wrong premises, found inside the survey that proposed the fix. Now a
reported debt: open, lacking a reproduction: 9, target 0.
design.py carries the control design-baseline.py never had, asserted
directly: a row citing a nonexistent file must not count as reproduced,
using the exact path 2da19a4 deleted -- which the old tool called green.
design-baseline.py is marked superseded rather than deleted; it is the
evidence for how a wrong number got into a survey.
T06. The report is a FILE in ground-game under GROUND-WP-0002, committed
there, with a hub message that only points at it. It asks for no ruling:
it carries GR-E01's withdrawal, our reproduction debt, and two notes that
are explicitly not findings.
And it had to acknowledge something nobody anticipated. GROUND-WP-0002 is
finished -- all ten U-items were RULED 2026-08-03, every one confirmed as
the default we simulate, plus five of six provisional scenarios. The
survey said "0 of 10 ruled" two days later and this register was built
saying `reported`. That is the unread-inbox failure running in the
opposite direction: they answered and we did not collect it. The
instrument's first run surfaced it. They are `ruled`, not `applied` --
lifting the now-settled provisional flags is owed and is not done, and
make design shows them open until it is.
T07. evidence/CB-EV-0021. Two of six catches in this pass came from
execution rather than process (the role distinction from building it, the
ten uncollected rulings from running it), which is InnerLoop §Design
goal's prediction holding.
make self-tests, facts-check, loop-lint: clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
398 lines
18 KiB
Markdown
398 lines
18 KiB
Markdown
---
|
||
id: CB-WP-0022
|
||
kind: product
|
||
title: "The design instrument: findings about the game, with their reproductions"
|
||
status: done
|
||
state_hub_workstream_id: "fda16340-0049-4acf-884b-a5cfdbde47c0"
|
||
---
|
||
|
||
# Purpose
|
||
|
||
```
|
||
structural tier L (named a high-leverage pass by the maintainer, and it
|
||
amends INTENT — clay-borg gains a stated aspect)
|
||
chaos d8 = 6 → no override
|
||
declared tier L
|
||
```
|
||
|
||
Declaration 5 of chaos window 2. Tier L: separate survey, **adversarial
|
||
review**, ADR, then spec, then code.
|
||
|
||
## The insight, in the maintainer's words
|
||
|
||
> *"We should consider ourselves testing the game and document
|
||
> inconsistencies to report them back to the ground-game repo, so that the
|
||
> game designer can improve the rules accordingly… We should have the
|
||
> rigorous game simulation engine and a meta scope to capture notes about
|
||
> game design flaws, questions, results and protocols about trial games…
|
||
> This will provide clay-borg an additional aspect as a valuable game
|
||
> design tool."*
|
||
|
||
**This is already happening and has no home.** In five passes the engine
|
||
has produced, as a by-product of being rigorous:
|
||
|
||
| finding | how it surfaced |
|
||
|---|---|
|
||
| ten underdetermined rules points (U1–U10) | formalizing the dataset into testable rules |
|
||
| SOLVE offered on a face-down Problem, always inert | a human dragging it three rounds running |
|
||
| GR-A13 "wasted SOLVE" on a claimed Problem | a scenario that had to pick a default |
|
||
| ~~GR-E01 unreachable below 5 seats~~ | arithmetic over the deal count — **withdrawn 2026-08-05, it was wrong (C1)** |
|
||
| ~~six provisional scenario defaults~~ | **five, and double-counted with GR-E01 (C3)** — they are reproductions, not a finding |
|
||
| GR-E03 / GR-E04 never played to the end | nobody noticed for nineteen passes |
|
||
|
||
Every one was found by *building the simulator*, not by playing. That is
|
||
the thing worth naming: **a simulator rigorous enough to refuse ambiguity
|
||
is a design instrument, because it cannot proceed past a rule that does
|
||
not decide.**
|
||
|
||
And every one of them has been carried in prose, in six different places,
|
||
and one sat unread in an inbox for four days.
|
||
|
||
## The load-bearing rule this must have
|
||
|
||
The project's standing failure is *unexecuted verification*. A design
|
||
register that collects opinions would reproduce it in a new medium.
|
||
|
||
> **A design finding is not admissible without its reproduction.**
|
||
|
||
Concretely: a scenario that fails, an arithmetic check that prints the
|
||
contradiction, a recorded game the reader can replay, or a named test.
|
||
*"This feels unbalanced"* is a note, not a finding. **The SOLVE inertness
|
||
is admissible because a recorded session shows three no-ops.**
|
||
|
||
> **The example that stood here was GR-E01, and the review killed it
|
||
> (C1).** *"4/6/9 against 5/7/9 is a computation anyone can rerun"* had
|
||
> already been rerun: `2da19a4` measured **6/9/12**, and the scenario was
|
||
> renamed `-unreachable-` → `-reachable-`. The conclusion inverted — and
|
||
> it was one of the two findings that **passed** this rule. So existence
|
||
> is not what was missing. See [ADR-0012](../decisions/ADR-0012-the-design-instrument.md)
|
||
> D3: the rule gains **shape**, and **a reproduction must be able to
|
||
> fail.** Ours went green and stayed admissible.
|
||
|
||
This is what would make clay-borg a design tool rather than a suggestion
|
||
box, and it is the one part of this proposal that must not be traded away
|
||
for convenience.
|
||
|
||
## The judgment I want reviewed, not assumed
|
||
|
||
The maintainer asked whether this should extend to *"a meta about the
|
||
clay-borg engine evolution itself."*
|
||
|
||
**My answer is no, and it should be argued rather than accepted.** That
|
||
register already exists and is load-bearing: `evidence/CB-EV-*`,
|
||
`decisions/ADR-*`, `gates.toml` and the workplans capture nineteen passes
|
||
with dates, costs and falsifiers. A second register for the same subject
|
||
would be ceremony. The asymmetry is the point: engine evolution has a home
|
||
and game design does not.
|
||
|
||
> **Settled by [ADR-0012](../decisions/ADR-0012-the-design-instrument.md)
|
||
> D7: no register — but the argument above did not survive.** C5 found the
|
||
> "third thing" the maintainer meant is visible in
|
||
> `specs/InnerLoopReference.md` and `history/`'s retrospectives, **neither
|
||
> of which this inventory names.** Conclusion narrowed, not settled: if
|
||
> InnerLoopReference keeps absorbing material that is neither a decision
|
||
> nor a finding, revisit.
|
||
|
||
## Task: survey how this is done elsewhere, and what we already have
|
||
|
||
```task
|
||
id: CB-WP-0022-T01
|
||
status: done
|
||
priority: high
|
||
state_hub_task_id: "6b8663b8-4143-4361-ad9e-7df0c4f4a19d"
|
||
```
|
||
|
||
`research/CB-RES-0007-*.md`.
|
||
|
||
**Do not survey issue trackers.** The question is narrower and more
|
||
interesting: how do rigorous rule systems record *the ambiguity they
|
||
found*? Candidates worth a benchmark-to-beat:
|
||
|
||
- **Errata and rulings practice** in published games (Magic's
|
||
comprehensive-rules + rulings split, Netrunner's NAPD card rulings) —
|
||
what makes a ruling *findable* years later.
|
||
- **Formal-methods counterexample traces** — a model checker's output is
|
||
precisely a reproduction attached to a claim, which is the shape wanted
|
||
here.
|
||
- **Conformance-suite provisional behaviour** — how W3C/WHATWG mark
|
||
"implementation-defined" and how a spec later absorbs it.
|
||
- **What this repo already has**: `provisional: true` scenarios,
|
||
`§Underdetermined`, `gates.toml`'s `caught`/`retire_if` shape, and the
|
||
hub message that went unread. **The register must reuse the provisional
|
||
machinery rather than compete with it.**
|
||
|
||
Name, per dimension, the property to beat — findability, reproducibility,
|
||
and whether a ruling can *close* a finding mechanically.
|
||
|
||
**Done 2026-08-03.**
|
||
[CB-RES-0007](../research/CB-RES-0007-design-instrument.md), with a
|
||
runnable baseline (`tools/design-baseline.py`).
|
||
|
||
**Its numbers were withdrawn by T02 and must not be quoted from here.**
|
||
The survey reported *6 findings, 2/6 = 33% reproduced, 11 files, 4 days*.
|
||
C2 showed the instrument counted itself and its reproduction check never
|
||
stat'd the file; C3 showed *"six provisional defaults"* is five and GR-E01
|
||
was double-counted; T05's backfill contradicted *"six of the ten have
|
||
provisional scenarios"* — **one** does. What survives is direction: many
|
||
files, no index, 0 of 10 ruled. The first honest figures are T05's.
|
||
|
||
**Magic corrected an assumption this pass was about to build on.** Rulings
|
||
are *"reminder information with no actual weight or rules meaning"*; the
|
||
authoritative fix folds into the **Oracle** card text. **A finding closes
|
||
when the source changes, not when an annotation is added** — the register
|
||
is a queue that empties. Model checkers supplied the reproduction rule
|
||
independently, and W3C's *implementation-defined* mark is machinery we
|
||
already have and must reuse rather than duplicate.
|
||
|
||
## Task: adversarial review
|
||
|
||
```task
|
||
id: CB-WP-0022-T02
|
||
status: done
|
||
priority: high
|
||
state_hub_task_id: "d2597895-fe2e-4f1e-a3db-a5fb8833434c"
|
||
```
|
||
|
||
Tier L requires it. Give the reviewer the survey **and** the §judgment
|
||
above, and require an attempt at:
|
||
|
||
- **that the engine-evolution register is redundant** — the strongest
|
||
counter is that ADRs record *decisions* and evidence records *findings*,
|
||
but nothing records *what we learned about building engines*, which is a
|
||
third thing;
|
||
- **that "carries its reproduction" is affordable** — if half the real
|
||
findings cannot be reproduced cheaply, the rule will be quietly dropped
|
||
and the register becomes a suggestion box anyway. *(Hardened before the review ran: two findings had already reached
|
||
ground-game on wrong premises, so the reviewer was told to press whether
|
||
the rule is **sufficient**, not whether it is affordable.)*
|
||
- **that a register is needed at all**, rather than one more section in
|
||
`GroundRules.md §Underdetermined`, which already exists and already
|
||
works.
|
||
|
||
Record the trail in `history/`, unpolished.
|
||
|
||
**Done 2026-08-05.** Trail:
|
||
[challenge](../history/260805-design-instrument-challenge.md),
|
||
[response](../history/260805-design-instrument-response.md).
|
||
|
||
**Run by a separate agent** — the first in this repo that was. CB-RES-0006's
|
||
review opened by conceding it could not be, and called its own findings
|
||
*"a lower bound on what a genuinely separate reviewer would find."* That
|
||
was measurable, and this is the measurement: the separate reviewer ran
|
||
`git log` against the survey's central example and found our own commit
|
||
had falsified it four days earlier, while the author — who wrote that
|
||
commit — quoted the dead number twice.
|
||
|
||
**Seven challenges: four conceded, two conceded in part, one answered.**
|
||
**C1 changed the design** — the rule's showcase finding was false and had
|
||
*passed* the rule, so existence is not what was missing — and **caught a
|
||
defect in flight**, T06's payload still naming the dead number. C2
|
||
withdrew the baseline's precision, C3 its arithmetic, C4 flipped T03's
|
||
burden toward extending `§Underdetermined`, C5 corrected the redundancy
|
||
inventory. Survived: affordability, and reuse of the provisional
|
||
machinery. Full account:
|
||
[CB-EV-0021](../evidence/CB-EV-0021-the-design-instrument.md) §1–§3.
|
||
|
||
## Task: decide
|
||
|
||
```task
|
||
id: CB-WP-0022-T03
|
||
status: done
|
||
priority: high
|
||
state_hub_task_id: "a2e85810-e949-4ea0-81fc-36b912af326c"
|
||
```
|
||
|
||
`decisions/ADR-0012-*.md`. At minimum:
|
||
|
||
- whether clay-borg's **INTENT gains a stated aspect** as a design
|
||
instrument, and in what words — this is the change with the longest
|
||
half-life in the pass;
|
||
- the finding **taxonomy**, and it should be grounded in the six findings
|
||
above rather than invented: *underdetermined* (rules do not say),
|
||
*inconsistent* (rules disagree with each other or the data), *inert* (a
|
||
rule that cannot fire), *degenerate* (fires, but collapses play),
|
||
*unplayed* (implemented, never played);
|
||
- the **lifecycle** and who owns each state: raised → reported → ruled →
|
||
applied, or withdrawn;
|
||
- whether a finding without a reproduction is **rejected** or **admitted
|
||
as a note** — and if admitted, how it is prevented from aging into an
|
||
apparent finding.
|
||
|
||
**Done 2026-08-05.**
|
||
[ADR-0012](../decisions/ADR-0012-the-design-instrument.md), nine
|
||
decisions. The two not on this list are the two the review forced:
|
||
|
||
- **D2 — `§Underdetermined` *is* the register; nothing parallel is built.**
|
||
Against the survey's own five benchmarks the incumbent already delivers
|
||
four, including the Oracle property the survey went to Magic to find and
|
||
we had written ourselves eight days earlier (`GroundRules.md:231-233`).
|
||
What it lacks is reproductions. So this pass **extends** a section — no
|
||
new file, no new schema.
|
||
- **D3 — admissibility is three clauses.** Exists, has the ruled shape
|
||
(row-level table, never a sum), **and can fail.** GR-E01's artifact went
|
||
green and the finding stayed admissible and stayed queued, because
|
||
nothing said a passing artifact was a signal. **A green reproduction is
|
||
an alarm.**
|
||
|
||
The rest, in one line each: **D1** INTENT gains property 4, *Instrument*,
|
||
applied with its falsifier. **D4** five kinds, each forced by an existing
|
||
finding. **D5** `applied` means the source changed; withdrawals are
|
||
reported, not deleted. **D6** notes admitted but never reportable, 30-day
|
||
expiry. **D7** no engine-evolution register, on an inventory C5 corrected.
|
||
**D8** `design-baseline.py` retired. **D9** the artifact stays here,
|
||
ground-game gets a generated file under its own workplan.
|
||
|
||
## Task: specify
|
||
|
||
```task
|
||
id: CB-WP-0022-T04
|
||
status: done
|
||
priority: high
|
||
state_hub_task_id: "60ddfeab-81b1-45e4-9ec9-91cc8ea7fb72"
|
||
```
|
||
|
||
`specs/GameDesign.md`, with metrics, because a spec without them is prose.
|
||
|
||
Candidate measures, to be argued not adopted:
|
||
|
||
- **findings with a runnable reproduction** — target 100%, and the
|
||
denominator includes withdrawn ones;
|
||
- **time from raised to reported** — the U-items took four days to be
|
||
*read*; that is the number this exists to fix;
|
||
- **findings closed by a ruling** vs **findings still open**, with age.
|
||
|
||
**ground-game has ruled on what a finding must carry** (GROUND-WP-0004
|
||
T02, 2026-08-04), and it is stricter than this pass proposed. Adopt it:
|
||
|
||
> Arithmetic findings ship a **runnable reproduction** *and* a
|
||
> **row-level deal table** — never only "sum of file" or "deal depth N";
|
||
> and ground-game's arithmetic rulings cite that reproduction by path.
|
||
|
||
The second half is theirs to keep. **So the reproduction rule gains a
|
||
shape requirement, not just an existence one** — a finding that ships a
|
||
passing test but describes the wrong quantity is still a bad finding,
|
||
which is exactly what happened twice.
|
||
|
||
Also specify the **trial protocol**: a trial game is a `--record`ed
|
||
session plus an observation log, so *"we played it and X happened"* is
|
||
replayable rather than remembered. It must cost almost nothing or it will
|
||
not be done.
|
||
|
||
**Done 2026-08-05.** [specs/GameDesign.md](../specs/GameDesign.md) v1.0 —
|
||
not a register (ADR-0012 D2 put that in `§Underdetermined`).
|
||
|
||
**§1.2 is written against evidence rather than principle**: a finding must
|
||
print the rows behind any number it claims. *"12" was arithmetically
|
||
defensible and still wrong about the game.* **§1.3's target is `0`
|
||
reproductions gone green while open** — what GR-E01 would have tripped
|
||
four days before a human caught it. **No baseline rate is quoted.**
|
||
|
||
**The trial protocol costs one flag**: `cb-play --record` plus a sibling
|
||
`.md` in the player's own words. An observation is a **note** until it has
|
||
a reproduction — *"I felt it was too easy but then we lost"* is the case
|
||
it is shaped around, and a schema at the moment of observation would lose
|
||
it.
|
||
|
||
## Task: build it, and backfill what is already known
|
||
|
||
```task
|
||
id: CB-WP-0022-T05
|
||
status: done
|
||
priority: high
|
||
state_hub_task_id: "469413ba-ed47-4840-a60e-ee7f95dc06ff"
|
||
```
|
||
|
||
The register, the tool, and then **the six findings above entered into
|
||
it** — backfilling is the test. A register that cannot express findings
|
||
the project already has is the wrong register, and discovering that after
|
||
designing it is the point of doing it in this order.
|
||
|
||
`make design` (or equivalent) must report: open findings by kind, those
|
||
without a reproduction, and those never reported to their owner.
|
||
|
||
**Done 2026-08-05.** `tools/design.py`, `make design`, and the register in
|
||
[`GroundRules.md`](../specs/GroundRules.md) — **14 rows, no new file.**
|
||
|
||
Backfill was the test. The taxonomy held (five kinds, no sixth), and it
|
||
**produced a `role` column ADR-0012 does not have**: the first report
|
||
alarmed on U2, wrongly — a green *default* is expected, a green
|
||
*counterexample* is the alarm. Folded into GameDesign §1.3. It also
|
||
contradicted the survey: **one** U-item names itself in a scenario, not
|
||
six. Detail and figures:
|
||
[CB-EV-0021](../evidence/CB-EV-0021-the-design-instrument.md) §4, §6.
|
||
|
||
## Task: report to ground-game, mechanically
|
||
|
||
```task
|
||
id: CB-WP-0022-T06
|
||
status: done
|
||
priority: high
|
||
state_hub_task_id: "c6a659bd-c495-4754-9332-baadad51012a"
|
||
```
|
||
|
||
Generate the report and send it. **The message that sat unread for four
|
||
days is the baseline to beat** — the failure was not the message, it was
|
||
that nothing pointed at it. So the report lands as a file in `ground-game`
|
||
under its own workplan, extending GROUND-WP-0002 rather than duplicating
|
||
it.
|
||
|
||
Include the findings this pass has sharpened:
|
||
|
||
- **SOLVE's legality** against a face-down Problem or an unmatchable suit —
|
||
and note that the case we *reported* was not the case that fired
|
||
(CB-WP-0023 T01).
|
||
- ~~**GR-E01 vs GR-S01** — 4/6/9 against 5/7/9~~ — **withdrawn
|
||
2026-08-05, before sending.** C1 found `2da19a4` had measured **6/9/12**:
|
||
the dataset reconciles them. It would have been the **fourth** wrong
|
||
premise to reach `ground-game` and is the only one caught before
|
||
transmission. **Report the withdrawal** — a claim retracted silently is
|
||
how the first three survived.
|
||
|
||
**Done 2026-08-05.**
|
||
[`ground-game/workplans/GROUND-WP-0002-clay-borg-report-260805.md`](../../ground-game/workplans/GROUND-WP-0002-clay-borg-report-260805.md),
|
||
committed there, with a hub message that only *points at* the file.
|
||
|
||
**The report asks for no ruling.** It carries GR-E01's withdrawal, our own
|
||
reproduction debt, and two notes that are explicitly not findings.
|
||
|
||
**And it acknowledged something the pass did not expect.**
|
||
GROUND-WP-0002 is `finished`: **all ten U-items were ruled 2026-08-03**,
|
||
every one confirmed as the default we simulate. CB-RES-0007 said *"0 of 10
|
||
ruled"* two days later. **The unread-inbox failure running in the opposite
|
||
direction** — they answered and we did not collect it. The instrument's
|
||
first run surfaced it.
|
||
|
||
## Task: evidence
|
||
|
||
```task
|
||
id: CB-WP-0022-T07
|
||
status: done
|
||
priority: high
|
||
state_hub_task_id: "4fb77911-486d-4d9e-a53d-c4e486c74ce2"
|
||
```
|
||
|
||
`evidence/CB-EV-0021-*.md`. *(Was CB-EV-0020 when written; CB-WP-0023
|
||
shipped that number first — `evidence/CB-EV-0020-solve-legality.md` — so
|
||
this one moves rather than collides.)*
|
||
|
||
- **Whether backfilling changed the design** — if all six findings fit the
|
||
first taxonomy, say so and be suspicious of it.
|
||
- **What tier L cost against what it caught**, since this is the second
|
||
full-weight L pass and CB-WP-0012's deleted its own structural trigger.
|
||
- **The engine-evolution question**, as the review left it.
|
||
- **Quote CB-WP-0021's cost by re-running the instrument.**
|
||
|
||
**Done 2026-08-05.**
|
||
[CB-EV-0021](../evidence/CB-EV-0021-the-design-instrument.md).
|
||
|
||
**Backfill did change the design** — and the honest answer to *"be
|
||
suspicious if all six fit"* is that only **five** were entered (one was a
|
||
double-count), so fitting them is close to circular. The taxonomy's real
|
||
test is the seventh finding.
|
||
|
||
**Tier L's cost against what it caught**: four of six catches came only
|
||
from the separate reviewer, and **two came from execution rather than
|
||
process** — the `role` distinction from building it, the ten uncollected
|
||
rulings from running it. That is InnerLoop §Design goal's prediction
|
||
holding, and an argument against front-loading more review rather than
|
||
less.
|