First adversarial review in this repo run by a genuinely separate agent.
CB-RES-0006's reviewer opened by conceding it could not be, and called its
own findings "a lower bound on what a genuinely separate reviewer would
find." This is the measurement: the separate reviewer ran git log against
the survey's central example and found 2da19a4 had falsified it four days
earlier, while the author -- who wrote that commit -- quoted the dead
number twice.
Seven challenges: four conceded, two conceded in part, one answered.
C1 changes the design. "GR-E01 is admissible because 4/6/9 against 5/7/9
is a computation anyone can rerun" was a computation already rerun: the
edition import measured 6/9/12, the conclusion inverted, and the scenario
was renamed -unreachable- to -reachable-. That finding was one of the TWO
that passed the reproduction rule. So three wrong premises have now
reached ground-game and the third satisfied an existence test -- existence
is not the property that was missing. The rule gains shape (ground-game's
row-level deal table, promoted from a T04 addendum) and a clause the
survey never contemplated: a reproduction must be able to fail. Ours went
green and stayed admissible.
C1 also caught a defect in flight. T06's payload, status todo, still named
4/6/9 and was queued to send it to ground-game as "no dataset reconciles
them." Withdrawn before sending -- the fourth wrong premise, and the only
one stopped.
C2 withdraws the baseline's precision: design-baseline.py is a
hand-maintained dict counting itself, has_reproduction never checks the
file exists (its YES-control is green against a deleted path), and
Makefile:127 runs only --self-test so the reporting path has no CI. The
direction survives; 33% is not a measured rate and T05 must not build on
it. C3: "six provisional defaults" is five, GR-E01 double-counted. C4:
GroundRules §Underdetermined was never evaluated as a candidate and
already delivers four of five benchmarks -- T03's burden flips to arguing
extension over replacement.
Survived: the rule's affordability, and reuse of the provisional
machinery.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
366 lines
16 KiB
Markdown
366 lines
16 KiB
Markdown
---
|
||
id: CB-WP-0022
|
||
kind: product
|
||
title: "The design instrument: findings about the game, with their reproductions"
|
||
status: active
|
||
state_hub_workstream_id: "fda16340-0049-4acf-884b-a5cfdbde47c0"
|
||
---
|
||
|
||
# Purpose
|
||
|
||
```
|
||
structural tier L (named a high-leverage pass by the maintainer, and it
|
||
amends INTENT — clay-borg gains a stated aspect)
|
||
chaos d8 = 6 → no override
|
||
declared tier L
|
||
```
|
||
|
||
Declaration 5 of chaos window 2. Tier L: separate survey, **adversarial
|
||
review**, ADR, then spec, then code.
|
||
|
||
## The insight, in the maintainer's words
|
||
|
||
> *"We should consider ourselves testing the game and document
|
||
> inconsistencies to report them back to the ground-game repo, so that the
|
||
> game designer can improve the rules accordingly… We should have the
|
||
> rigorous game simulation engine and a meta scope to capture notes about
|
||
> game design flaws, questions, results and protocols about trial games…
|
||
> This will provide clay-borg an additional aspect as a valuable game
|
||
> design tool."*
|
||
|
||
**This is already happening and has no home.** In five passes the engine
|
||
has produced, as a by-product of being rigorous:
|
||
|
||
| finding | how it surfaced |
|
||
|---|---|
|
||
| ten underdetermined rules points (U1–U10) | formalizing the dataset into testable rules |
|
||
| SOLVE offered on a face-down Problem, always inert | a human dragging it three rounds running |
|
||
| GR-A13 "wasted SOLVE" on a claimed Problem | a scenario that had to pick a default |
|
||
| GR-E01 unreachable below 5 seats | arithmetic over the deal count |
|
||
| six provisional scenario defaults | scenarios that could not be written without deciding something |
|
||
| GR-E03 / GR-E04 never played to the end | nobody noticed for nineteen passes |
|
||
|
||
Every one was found by *building the simulator*, not by playing. That is
|
||
the thing worth naming: **a simulator rigorous enough to refuse ambiguity
|
||
is a design instrument, because it cannot proceed past a rule that does
|
||
not decide.**
|
||
|
||
And every one of them has been carried in prose, in six different places,
|
||
and one sat unread in an inbox for four days.
|
||
|
||
## The load-bearing rule this must have
|
||
|
||
The project's standing failure is *unexecuted verification*. A design
|
||
register that collects opinions would reproduce it in a new medium.
|
||
|
||
> **A design finding is not admissible without its reproduction.**
|
||
|
||
Concretely: a scenario that fails, an arithmetic check that prints the
|
||
contradiction, a recorded game the reader can replay, or a named test.
|
||
*"This feels unbalanced"* is a note, not a finding. **The SOLVE inertness
|
||
is admissible because a recorded session shows three no-ops.**
|
||
|
||
> **The example that stood here was GR-E01, and the adversarial review
|
||
> killed it (C1, 2026-08-05).** *"4/6/9 against 5/7/9 is a computation
|
||
> anyone can rerun"* was a computation that had already been rerun:
|
||
> `2da19a4` measured **6/9/12 against 5/7/9** and renamed the scenario
|
||
> `-unreachable-` → `-reachable-`. The finding's conclusion inverted, and
|
||
> it was one of the two findings that **passed** this rule.
|
||
>
|
||
> So existence is not the property that was missing — three wrong premises
|
||
> have now reached `ground-game`, and the third satisfied an existence
|
||
> test. T03 must adopt the shape requirement as part of the rule, plus a
|
||
> clause the survey never contemplated: **a reproduction must be able to
|
||
> fail.** Ours went green and stayed admissible.
|
||
|
||
This is what would make clay-borg a design tool rather than a suggestion
|
||
box, and it is the one part of this proposal that must not be traded away
|
||
for convenience.
|
||
|
||
## The judgment I want reviewed, not assumed
|
||
|
||
The maintainer asked whether this should extend to *"a meta about the
|
||
clay-borg engine evolution itself."*
|
||
|
||
**My answer is no, and it should be argued rather than accepted.** That
|
||
register already exists and is load-bearing: `evidence/CB-EV-*` records
|
||
what each pass found, `decisions/ADR-*` records what was decided and what
|
||
was rejected, `gates.toml` records what every control has caught and what
|
||
would retire it, and workplans record what was attempted. Nineteen passes
|
||
of engine evolution are already captured, with dates, costs and
|
||
falsifiers.
|
||
|
||
**Building a second register for the same subject would be ceremony**, and
|
||
this project has a standing rule that a gate must cash out. The asymmetry
|
||
is the point: engine evolution has a home and game design does not.
|
||
|
||
If the adversarial review disagrees, that is exactly the kind of thing
|
||
tier L exists to surface.
|
||
|
||
## Task: survey how this is done elsewhere, and what we already have
|
||
|
||
```task
|
||
id: CB-WP-0022-T01
|
||
status: done
|
||
priority: high
|
||
state_hub_task_id: "6b8663b8-4143-4361-ad9e-7df0c4f4a19d"
|
||
```
|
||
|
||
`research/CB-RES-0007-*.md`.
|
||
|
||
**Do not survey issue trackers.** The question is narrower and more
|
||
interesting: how do rigorous rule systems record *the ambiguity they
|
||
found*? Candidates worth a benchmark-to-beat:
|
||
|
||
- **Errata and rulings practice** in published games (Magic's
|
||
comprehensive-rules + rulings split, Netrunner's NAPD card rulings) —
|
||
what makes a ruling *findable* years later.
|
||
- **Formal-methods counterexample traces** — a model checker's output is
|
||
precisely a reproduction attached to a claim, which is the shape wanted
|
||
here.
|
||
- **Conformance-suite provisional behaviour** — how W3C/WHATWG mark
|
||
"implementation-defined" and how a spec later absorbs it.
|
||
- **What this repo already has**: `provisional: true` scenarios,
|
||
`§Underdetermined`, `gates.toml`'s `caught`/`retire_if` shape, and the
|
||
hub message that went unread. **The register must reuse the provisional
|
||
machinery rather than compete with it.**
|
||
|
||
Name, per dimension, the property to beat — findability, reproducibility,
|
||
and whether a ruling can *close* a finding mechanically.
|
||
|
||
**Done 2026-08-03.**
|
||
[CB-RES-0007](../research/CB-RES-0007-design-instrument.md), with a
|
||
runnable baseline (`tools/design-baseline.py`).
|
||
|
||
**The baseline is us, and it is measured**: 6 findings across **11 files**
|
||
with no index, **2 of 6 (33%) with a runnable reproduction**, and U1–U10
|
||
raised 2026-07-30, first *read* 2026-08-03 — **4 days**, 0 of 10 ruled.
|
||
|
||
**The uncomfortable number is stated up front rather than left for the
|
||
review to find: the reproduction rule would reject four of our six
|
||
existing findings.** The survey answers it — none of the four is
|
||
expensive to reproduce, so the 33% is evidence that nobody was ever asked
|
||
for one, not that the rule is unaffordable.
|
||
|
||
**Magic corrected an assumption this pass was about to build on.** I
|
||
expected a ruling to be the authoritative resolution. It is not: rulings
|
||
are *"reminder information with no actual weight or rules meaning"*, and
|
||
the authoritative fix folds into the **Oracle** card text. **A finding
|
||
closes when the source changes, not when an annotation is added** — so
|
||
the register must be a queue that empties, not an archive that grows.
|
||
That is now a constraint on T03's lifecycle.
|
||
|
||
Model checkers supplied the reproduction rule independently (a
|
||
counterexample trace *is* the finding), and W3C's *implementation-defined*
|
||
mark is the one piece of machinery we already have and must reuse rather
|
||
than duplicate.
|
||
|
||
## Task: adversarial review
|
||
|
||
```task
|
||
id: CB-WP-0022-T02
|
||
status: done
|
||
priority: high
|
||
state_hub_task_id: "d2597895-fe2e-4f1e-a3db-a5fb8833434c"
|
||
```
|
||
|
||
Tier L requires it. Give the reviewer the survey **and** the §judgment
|
||
above, and require an attempt at:
|
||
|
||
- **that the engine-evolution register is redundant** — the strongest
|
||
counter is that ADRs record *decisions* and evidence records *findings*,
|
||
but nothing records *what we learned about building engines*, which is a
|
||
third thing;
|
||
- **that "carries its reproduction" is affordable** — if half the real
|
||
findings cannot be reproduced cheaply, the rule will be quietly dropped
|
||
and the register becomes a suggestion box anyway. *(Since this task was
|
||
written the answer has hardened: two findings reached ground-game on
|
||
wrong premises, and ground-game has itself ruled that reproductions are
|
||
required. The reviewer should now press the opposite way — whether the
|
||
rule is **sufficient**, given both bad premises would have shipped a
|
||
passing test alongside a wrong description.)*
|
||
- **that a register is needed at all**, rather than one more section in
|
||
`GroundRules.md §Underdetermined`, which already exists and already
|
||
works.
|
||
|
||
Record the trail in `history/`, unpolished.
|
||
|
||
**Done 2026-08-05.** Trail:
|
||
[challenge](../history/260805-design-instrument-challenge.md),
|
||
[response](../history/260805-design-instrument-response.md).
|
||
|
||
**Run by a separate agent** — the first in this repo that was. CB-RES-0006's
|
||
review opened by conceding it could not be, and called its own findings
|
||
*"a lower bound on what a genuinely separate reviewer would find."* That
|
||
was measurable, and this is the measurement: the separate reviewer ran
|
||
`git log` against the survey's central example and found our own commit
|
||
had falsified it four days earlier, while the author — who wrote that
|
||
commit — quoted the dead number twice.
|
||
|
||
**Seven challenges: four conceded, two conceded in part, one answered.**
|
||
|
||
- **C1 lands hardest and changed the design.** The rule's showcase finding
|
||
was false and had *passed* the rule. Existence is not the missing
|
||
property; **shape** and **falsifiability** are. Folded into §The
|
||
load-bearing rule above, and it is T03's to settle.
|
||
- **C1 also caught a defect in flight** — T06's payload, `todo`, still
|
||
named the dead number. Withdrawn above before sending.
|
||
- **C2 withdrew the baseline's precision.** `tools/design-baseline.py` is a
|
||
hand-maintained dict counting itself (`:16-36`, `:89`); `has_reproduction`
|
||
(`:38-43`) never checks the file exists, so the self-test's YES-control
|
||
(`:63`) is green against a path `2da19a4` deleted. `Makefile:127` runs
|
||
only `--self-test`, so the reporting path has no CI. The direction
|
||
stands — 11 files, no index, 0 of 10 ruled are all checkable without the
|
||
tool — but **33% is not a measured rate and T05 must not build on it.**
|
||
- **C3**: "six provisional defaults" is five, and GR-E01 is double-counted
|
||
in the `2/6`. No corrected rate is quoted here; the instrument that would
|
||
produce it is the one C2 withdrew.
|
||
- **C4**: `§Underdetermined` was never evaluated as a candidate, and it
|
||
already delivers four of five benchmarks including the Oracle property
|
||
the survey went to Magic to find. **T03's burden flips: argue why it is
|
||
extended, not replaced.**
|
||
- **C5**: the engine-evolution "third thing" is visible in
|
||
`specs/InnerLoopReference.md` and `history/`'s retrospectives, neither of
|
||
which my redundancy inventory named. Conclusion narrowed, not settled.
|
||
- **Survived**: the reproduction rule's *affordability*, and §4's reuse of
|
||
the provisional machinery. Both with stated falsifiers.
|
||
|
||
## Task: decide
|
||
|
||
```task
|
||
id: CB-WP-0022-T03
|
||
status: todo
|
||
priority: high
|
||
state_hub_task_id: "a2e85810-e949-4ea0-81fc-36b912af326c"
|
||
```
|
||
|
||
`decisions/ADR-0012-*.md`. At minimum:
|
||
|
||
- whether clay-borg's **INTENT gains a stated aspect** as a design
|
||
instrument, and in what words — this is the change with the longest
|
||
half-life in the pass;
|
||
- the finding **taxonomy**, and it should be grounded in the six findings
|
||
above rather than invented: *underdetermined* (rules do not say),
|
||
*inconsistent* (rules disagree with each other or the data), *inert* (a
|
||
rule that cannot fire), *degenerate* (fires, but collapses play),
|
||
*unplayed* (implemented, never played);
|
||
- the **lifecycle** and who owns each state: raised → reported → ruled →
|
||
applied, or withdrawn;
|
||
- whether a finding without a reproduction is **rejected** or **admitted
|
||
as a note** — and if admitted, how it is prevented from aging into an
|
||
apparent finding.
|
||
|
||
## Task: specify
|
||
|
||
```task
|
||
id: CB-WP-0022-T04
|
||
status: todo
|
||
priority: high
|
||
state_hub_task_id: "60ddfeab-81b1-45e4-9ec9-91cc8ea7fb72"
|
||
```
|
||
|
||
`specs/GameDesign.md`, with metrics, because a spec without them is prose.
|
||
|
||
Candidate measures, to be argued not adopted:
|
||
|
||
- **findings with a runnable reproduction** — target 100%, and the
|
||
denominator includes withdrawn ones;
|
||
- **time from raised to reported** — the U-items took four days to be
|
||
*read*; that is the number this exists to fix;
|
||
- **findings closed by a ruling** vs **findings still open**, with age.
|
||
|
||
**ground-game has ruled on what a finding must carry** (GROUND-WP-0004
|
||
T02, 2026-08-04), and it is stricter than this pass proposed. Adopt it:
|
||
|
||
> 1. Arithmetic findings must ship a **runnable reproduction** *and* a
|
||
> **row-level deal table** listing Surface and each hidden priority
|
||
> separately — **never only "sum of file" or "deal depth N"**.
|
||
> 2. ground-game rulings that depend on arithmetic **should cite** that
|
||
> reproduction id or path, so a later reader can re-run it.
|
||
|
||
The second half is theirs to keep. The first is ours, and it is a direct
|
||
consequence of both premises we got wrong: *"12 in the file"* was a sum
|
||
without a deal table, and *"offered on a face-down Problem"* was a
|
||
condition named without checking which one actually fired.
|
||
|
||
**So the reproduction rule gains a shape requirement, not just an
|
||
existence one.** A finding that ships a passing test but describes the
|
||
wrong quantity is still a bad finding — which is exactly what happened
|
||
twice. The spec must say what a reproduction has to *show*, not only that
|
||
one exists.
|
||
|
||
Also specify the **trial protocol**, which is the part with no precedent
|
||
here: a trial game is a `--record`ed session plus an observation log, so
|
||
*"we played it and X happened"* is replayable rather than remembered. The
|
||
engine already records sessions as scenarios; a trial is that plus notes,
|
||
and it must cost almost nothing or it will not be done.
|
||
|
||
## Task: build it, and backfill what is already known
|
||
|
||
```task
|
||
id: CB-WP-0022-T05
|
||
status: todo
|
||
priority: high
|
||
state_hub_task_id: "469413ba-ed47-4840-a60e-ee7f95dc06ff"
|
||
```
|
||
|
||
The register, the tool, and then **the six findings above entered into
|
||
it** — backfilling is the test. A register that cannot express findings
|
||
the project already has is the wrong register, and discovering that after
|
||
designing it is the point of doing it in this order.
|
||
|
||
`make design` (or equivalent) must report: open findings by kind, those
|
||
without a reproduction, and those never reported to their owner.
|
||
|
||
## Task: report to ground-game, mechanically
|
||
|
||
```task
|
||
id: CB-WP-0022-T06
|
||
status: todo
|
||
priority: high
|
||
state_hub_task_id: "c6a659bd-c495-4754-9332-baadad51012a"
|
||
```
|
||
|
||
Generate the report and send it. **The message that sat unread for four
|
||
days is the baseline to beat**, and the failure was not the message — it
|
||
was that nothing pointed at it and nothing tracked whether it was
|
||
answered.
|
||
|
||
So the report must land somewhere that persists: a file in `ground-game`
|
||
under its own workplan, not only an inbox entry. GROUND-WP-0002 already
|
||
holds the ten U-items; this should extend it rather than duplicate it.
|
||
|
||
Include the findings this pass has sharpened:
|
||
|
||
- **SOLVE's legality** against a face-down Problem or an unmatchable suit —
|
||
and note that the case we *reported* was not the case that fired
|
||
(CB-WP-0023 T01).
|
||
- ~~**GR-E01 vs GR-S01** — 4/6/9 against 5/7/9, no dataset reconciles
|
||
them~~ — **withdrawn 2026-08-05, before sending.** The adversarial
|
||
review (C1) found `2da19a4` had already measured **6/9/12 against
|
||
5/7/9**: the dataset reconciles them and the scenario is now
|
||
`-reachable-`. Sending this would have been the **fourth** wrong premise
|
||
to reach `ground-game`, and the only one caught before transmission.
|
||
**Report the withdrawal, not the finding** — GROUND-WP-0002 holds the
|
||
original, and a claim retracted silently is how the first three
|
||
survived.
|
||
|
||
## Task: evidence
|
||
|
||
```task
|
||
id: CB-WP-0022-T07
|
||
status: todo
|
||
priority: high
|
||
state_hub_task_id: "4fb77911-486d-4d9e-a53d-c4e486c74ce2"
|
||
```
|
||
|
||
`evidence/CB-EV-0021-*.md`. *(Was CB-EV-0020 when written; CB-WP-0023
|
||
shipped that number first — `evidence/CB-EV-0020-solve-legality.md` — so
|
||
this one moves rather than collides.)*
|
||
|
||
- **Whether backfilling changed the design** — if all six findings fit the
|
||
first taxonomy, say so and be suspicious of it.
|
||
- **What tier L cost against what it caught**, since this is the second
|
||
full-weight L pass and CB-WP-0012's deleted its own structural trigger.
|
||
- **The engine-evolution question**, as the review left it.
|
||
- **Quote CB-WP-0021's cost by re-running the instrument.**
|