clay-borg/workplans/CB-WP-0022-the-design-instrument.md
tegwick abd62c5567 CB-EV-0020: a gate moved the rule, and the report was wrong
The AM-1 coverage gate failed the build on GR-P05 being uncovered, which
is what showed the rule was in the offer layer rather than in validate.
A rule enforced only by the offer is enforced only for clients that ask
what is legal. The gate did not catch a bug, it caught a design error.

And the reported case was not the one reported. CB-WP-0018, CB-EV-0016
and the message to ground-game all described SOLVE offered on a
face-down Problem; validate already rejected face-down, so it never was.
Problem 1 is the Surface Problem, face-up from the deal, so the three
inert SOLVEs were the HAND case. The ruling covers both so nothing is
invalidated, but a ruling was requested on a wrong description -- the
second time in three passes that a premise reached ground-game
unchecked, after the '12 points available' that voided GR-E01.

Two of two. The pattern is not careless analysis; it is that a claim gets
SENT the moment it is interesting and checked afterwards. Unexecuted
verification, one step further out: not a belief acted on, but a belief
published. CB-WP-0022's reproduction rule would have caught both.

An earlier mutation run reported three survivors and was wrong -- the
replacement strings did not match, so nothing was mutated. It proved
nothing and looked like a result.

Also renames CB-WP-0022-T06B to T07; the hub flagged it as an
unregistered species.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 00:21:21 +02:00

279 lines
11 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
id: CB-WP-0022
kind: product
title: "The design instrument: findings about the game, with their reproductions"
status: active
state_hub_workstream_id: "fda16340-0049-4acf-884b-a5cfdbde47c0"
---
# Purpose
```
structural tier L (named a high-leverage pass by the maintainer, and it
amends INTENT — clay-borg gains a stated aspect)
chaos d8 = 6 → no override
declared tier L
```
Declaration 5 of chaos window 2. Tier L: separate survey, **adversarial
review**, ADR, then spec, then code.
## The insight, in the maintainer's words
> *"We should consider ourselves testing the game and document
> inconsistencies to report them back to the ground-game repo, so that the
> game designer can improve the rules accordingly… We should have the
> rigorous game simulation engine and a meta scope to capture notes about
> game design flaws, questions, results and protocols about trial games…
> This will provide clay-borg an additional aspect as a valuable game
> design tool."*
**This is already happening and has no home.** In five passes the engine
has produced, as a by-product of being rigorous:
| finding | how it surfaced |
|---|---|
| ten underdetermined rules points (U1U10) | formalizing the dataset into testable rules |
| SOLVE offered on a face-down Problem, always inert | a human dragging it three rounds running |
| GR-A13 "wasted SOLVE" on a claimed Problem | a scenario that had to pick a default |
| GR-E01 unreachable below 5 seats | arithmetic over the deal count |
| six provisional scenario defaults | scenarios that could not be written without deciding something |
| GR-E03 / GR-E04 never played to the end | nobody noticed for nineteen passes |
Every one was found by *building the simulator*, not by playing. That is
the thing worth naming: **a simulator rigorous enough to refuse ambiguity
is a design instrument, because it cannot proceed past a rule that does
not decide.**
And every one of them has been carried in prose, in six different places,
and one sat unread in an inbox for four days.
## The load-bearing rule this must have
The project's standing failure is *unexecuted verification*. A design
register that collects opinions would reproduce it in a new medium.
> **A design finding is not admissible without its reproduction.**
Concretely: a scenario that fails, an arithmetic check that prints the
contradiction, a recorded game the reader can replay, or a named test.
*"This feels unbalanced"* is a note, not a finding. **GR-E01 is admissible
because 4/6/9 against 5/7/9 is a computation anyone can rerun; the SOLVE
inertness is admissible because a recorded session shows three no-ops.**
This is what would make clay-borg a design tool rather than a suggestion
box, and it is the one part of this proposal that must not be traded away
for convenience.
## The judgment I want reviewed, not assumed
The maintainer asked whether this should extend to *"a meta about the
clay-borg engine evolution itself."*
**My answer is no, and it should be argued rather than accepted.** That
register already exists and is load-bearing: `evidence/CB-EV-*` records
what each pass found, `decisions/ADR-*` records what was decided and what
was rejected, `gates.toml` records what every control has caught and what
would retire it, and workplans record what was attempted. Nineteen passes
of engine evolution are already captured, with dates, costs and
falsifiers.
**Building a second register for the same subject would be ceremony**, and
this project has a standing rule that a gate must cash out. The asymmetry
is the point: engine evolution has a home and game design does not.
If the adversarial review disagrees, that is exactly the kind of thing
tier L exists to surface.
## Task: survey how this is done elsewhere, and what we already have
```task
id: CB-WP-0022-T01
status: done
priority: high
state_hub_task_id: "6b8663b8-4143-4361-ad9e-7df0c4f4a19d"
```
`research/CB-RES-0007-*.md`.
**Do not survey issue trackers.** The question is narrower and more
interesting: how do rigorous rule systems record *the ambiguity they
found*? Candidates worth a benchmark-to-beat:
- **Errata and rulings practice** in published games (Magic's
comprehensive-rules + rulings split, Netrunner's NAPD card rulings) —
what makes a ruling *findable* years later.
- **Formal-methods counterexample traces** — a model checker's output is
precisely a reproduction attached to a claim, which is the shape wanted
here.
- **Conformance-suite provisional behaviour** — how W3C/WHATWG mark
"implementation-defined" and how a spec later absorbs it.
- **What this repo already has**: `provisional: true` scenarios,
`§Underdetermined`, `gates.toml`'s `caught`/`retire_if` shape, and the
hub message that went unread. **The register must reuse the provisional
machinery rather than compete with it.**
Name, per dimension, the property to beat — findability, reproducibility,
and whether a ruling can *close* a finding mechanically.
**Done 2026-08-03.**
[CB-RES-0007](../research/CB-RES-0007-design-instrument.md), with a
runnable baseline (`tools/design-baseline.py`).
**The baseline is us, and it is measured**: 6 findings across **11 files**
with no index, **2 of 6 (33%) with a runnable reproduction**, and U1U10
raised 2026-07-30, first *read* 2026-08-03 — **4 days**, 0 of 10 ruled.
**The uncomfortable number is stated up front rather than left for the
review to find: the reproduction rule would reject four of our six
existing findings.** The survey answers it — none of the four is
expensive to reproduce, so the 33% is evidence that nobody was ever asked
for one, not that the rule is unaffordable.
**Magic corrected an assumption this pass was about to build on.** I
expected a ruling to be the authoritative resolution. It is not: rulings
are *"reminder information with no actual weight or rules meaning"*, and
the authoritative fix folds into the **Oracle** card text. **A finding
closes when the source changes, not when an annotation is added** — so
the register must be a queue that empties, not an archive that grows.
That is now a constraint on T03's lifecycle.
Model checkers supplied the reproduction rule independently (a
counterexample trace *is* the finding), and W3C's *implementation-defined*
mark is the one piece of machinery we already have and must reuse rather
than duplicate.
## Task: adversarial review
```task
id: CB-WP-0022-T02
status: todo
priority: high
state_hub_task_id: "d2597895-fe2e-4f1e-a3db-a5fb8833434c"
```
Tier L requires it. Give the reviewer the survey **and** the §judgment
above, and require an attempt at:
- **that the engine-evolution register is redundant** — the strongest
counter is that ADRs record *decisions* and evidence records *findings*,
but nothing records *what we learned about building engines*, which is a
third thing;
- **that "carries its reproduction" is affordable** — if half the real
findings cannot be reproduced cheaply, the rule will be quietly dropped
and the register becomes a suggestion box anyway;
- **that a register is needed at all**, rather than one more section in
`GroundRules.md §Underdetermined`, which already exists and already
works.
Record the trail in `history/`, unpolished.
## Task: decide
```task
id: CB-WP-0022-T03
status: todo
priority: high
state_hub_task_id: "a2e85810-e949-4ea0-81fc-36b912af326c"
```
`decisions/ADR-0012-*.md`. At minimum:
- whether clay-borg's **INTENT gains a stated aspect** as a design
instrument, and in what words — this is the change with the longest
half-life in the pass;
- the finding **taxonomy**, and it should be grounded in the six findings
above rather than invented: *underdetermined* (rules do not say),
*inconsistent* (rules disagree with each other or the data), *inert* (a
rule that cannot fire), *degenerate* (fires, but collapses play),
*unplayed* (implemented, never played);
- the **lifecycle** and who owns each state: raised → reported → ruled →
applied, or withdrawn;
- whether a finding without a reproduction is **rejected** or **admitted
as a note** — and if admitted, how it is prevented from aging into an
apparent finding.
## Task: specify
```task
id: CB-WP-0022-T04
status: todo
priority: high
state_hub_task_id: "60ddfeab-81b1-45e4-9ec9-91cc8ea7fb72"
```
`specs/GameDesign.md`, with metrics, because a spec without them is prose.
Candidate measures, to be argued not adopted:
- **findings with a runnable reproduction** — target 100%, and the
denominator includes withdrawn ones;
- **time from raised to reported** — the U-items took four days to be
*read*; that is the number this exists to fix;
- **findings closed by a ruling** vs **findings still open**, with age.
Also specify the **trial protocol**, which is the part with no precedent
here: a trial game is a `--record`ed session plus an observation log, so
*"we played it and X happened"* is replayable rather than remembered. The
engine already records sessions as scenarios; a trial is that plus notes,
and it must cost almost nothing or it will not be done.
## Task: build it, and backfill what is already known
```task
id: CB-WP-0022-T05
status: todo
priority: high
state_hub_task_id: "469413ba-ed47-4840-a60e-ee7f95dc06ff"
```
The register, the tool, and then **the six findings above entered into
it** — backfilling is the test. A register that cannot express findings
the project already has is the wrong register, and discovering that after
designing it is the point of doing it in this order.
`make design` (or equivalent) must report: open findings by kind, those
without a reproduction, and those never reported to their owner.
## Task: report to ground-game, mechanically
```task
id: CB-WP-0022-T06
status: todo
priority: high
state_hub_task_id: "c6a659bd-c495-4754-9332-baadad51012a"
```
Generate the report and send it. **The message that sat unread for four
days is the baseline to beat**, and the failure was not the message — it
was that nothing pointed at it and nothing tracked whether it was
answered.
So the report must land somewhere that persists: a file in `ground-game`
under its own workplan, not only an inbox entry. GROUND-WP-0002 already
holds the ten U-items; this should extend it rather than duplicate it.
Include the two sharpened findings this pass has already produced:
- **GR-E01 vs GR-S01** — the deal count puts 4/6/9 points in play against
thresholds of 5/7/9, so either the count or the thresholds are wrong and
no dataset reconciles them;
- **SOLVE's legality** against a face-down Problem or an unmatchable suit.
## Task: evidence
```task
id: CB-WP-0022-T07
status: todo
priority: high
state_hub_task_id: "4fb77911-486d-4d9e-a53d-c4e486c74ce2"
```
`evidence/CB-EV-0020-*.md`.
- **Whether backfilling changed the design** — if all six findings fit the
first taxonomy, say so and be suspicious of it.
- **What tier L cost against what it caught**, since this is the second
full-weight L pass and CB-WP-0012's deleted its own structural trigger.
- **The engine-evolution question**, as the review left it.
- **Quote CB-WP-0021's cost by re-running the instrument.**