clay-borg/workplans/CB-WP-0022-the-design-instrument.md
tegwick 04c3a4977f CB-WP-0022 T02: the separate reviewer found the showcase finding was false
First adversarial review in this repo run by a genuinely separate agent.
CB-RES-0006's reviewer opened by conceding it could not be, and called its
own findings "a lower bound on what a genuinely separate reviewer would
find." This is the measurement: the separate reviewer ran git log against
the survey's central example and found 2da19a4 had falsified it four days
earlier, while the author -- who wrote that commit -- quoted the dead
number twice.

Seven challenges: four conceded, two conceded in part, one answered.

C1 changes the design. "GR-E01 is admissible because 4/6/9 against 5/7/9
is a computation anyone can rerun" was a computation already rerun: the
edition import measured 6/9/12, the conclusion inverted, and the scenario
was renamed -unreachable- to -reachable-. That finding was one of the TWO
that passed the reproduction rule. So three wrong premises have now
reached ground-game and the third satisfied an existence test -- existence
is not the property that was missing. The rule gains shape (ground-game's
row-level deal table, promoted from a T04 addendum) and a clause the
survey never contemplated: a reproduction must be able to fail. Ours went
green and stayed admissible.

C1 also caught a defect in flight. T06's payload, status todo, still named
4/6/9 and was queued to send it to ground-game as "no dataset reconciles
them." Withdrawn before sending -- the fourth wrong premise, and the only
one stopped.

C2 withdraws the baseline's precision: design-baseline.py is a
hand-maintained dict counting itself, has_reproduction never checks the
file exists (its YES-control is green against a deleted path), and
Makefile:127 runs only --self-test so the reporting path has no CI. The
direction survives; 33% is not a measured rate and T05 must not build on
it. C3: "six provisional defaults" is five, GR-E01 double-counted. C4:
GroundRules §Underdetermined was never evaluated as a candidate and
already delivers four of five benchmarks -- T03's burden flips to arguing
extension over replacement.

Survived: the rule's affordability, and reuse of the provisional
machinery.

loop-lint: no findings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:03:12 +02:00

16 KiB
Raw Blame History

id kind title status state_hub_workstream_id
CB-WP-0022 product The design instrument: findings about the game, with their reproductions active fda16340-0049-4acf-884b-a5cfdbde47c0

Purpose

structural tier  L   (named a high-leverage pass by the maintainer, and it
                      amends INTENT — clay-borg gains a stated aspect)
chaos            d8 = 6  → no override
declared tier    L

Declaration 5 of chaos window 2. Tier L: separate survey, adversarial review, ADR, then spec, then code.

The insight, in the maintainer's words

"We should consider ourselves testing the game and document inconsistencies to report them back to the ground-game repo, so that the game designer can improve the rules accordingly… We should have the rigorous game simulation engine and a meta scope to capture notes about game design flaws, questions, results and protocols about trial games… This will provide clay-borg an additional aspect as a valuable game design tool."

This is already happening and has no home. In five passes the engine has produced, as a by-product of being rigorous:

finding how it surfaced
ten underdetermined rules points (U1U10) formalizing the dataset into testable rules
SOLVE offered on a face-down Problem, always inert a human dragging it three rounds running
GR-A13 "wasted SOLVE" on a claimed Problem a scenario that had to pick a default
GR-E01 unreachable below 5 seats arithmetic over the deal count
six provisional scenario defaults scenarios that could not be written without deciding something
GR-E03 / GR-E04 never played to the end nobody noticed for nineteen passes

Every one was found by building the simulator, not by playing. That is the thing worth naming: a simulator rigorous enough to refuse ambiguity is a design instrument, because it cannot proceed past a rule that does not decide.

And every one of them has been carried in prose, in six different places, and one sat unread in an inbox for four days.

The load-bearing rule this must have

The project's standing failure is unexecuted verification. A design register that collects opinions would reproduce it in a new medium.

A design finding is not admissible without its reproduction.

Concretely: a scenario that fails, an arithmetic check that prints the contradiction, a recorded game the reader can replay, or a named test. "This feels unbalanced" is a note, not a finding. The SOLVE inertness is admissible because a recorded session shows three no-ops.

The example that stood here was GR-E01, and the adversarial review killed it (C1, 2026-08-05). "4/6/9 against 5/7/9 is a computation anyone can rerun" was a computation that had already been rerun: 2da19a4 measured 6/9/12 against 5/7/9 and renamed the scenario -unreachable--reachable-. The finding's conclusion inverted, and it was one of the two findings that passed this rule.

So existence is not the property that was missing — three wrong premises have now reached ground-game, and the third satisfied an existence test. T03 must adopt the shape requirement as part of the rule, plus a clause the survey never contemplated: a reproduction must be able to fail. Ours went green and stayed admissible.

This is what would make clay-borg a design tool rather than a suggestion box, and it is the one part of this proposal that must not be traded away for convenience.

The judgment I want reviewed, not assumed

The maintainer asked whether this should extend to "a meta about the clay-borg engine evolution itself."

My answer is no, and it should be argued rather than accepted. That register already exists and is load-bearing: evidence/CB-EV-* records what each pass found, decisions/ADR-* records what was decided and what was rejected, gates.toml records what every control has caught and what would retire it, and workplans record what was attempted. Nineteen passes of engine evolution are already captured, with dates, costs and falsifiers.

Building a second register for the same subject would be ceremony, and this project has a standing rule that a gate must cash out. The asymmetry is the point: engine evolution has a home and game design does not.

If the adversarial review disagrees, that is exactly the kind of thing tier L exists to surface.

Task: survey how this is done elsewhere, and what we already have

id: CB-WP-0022-T01
status: done
priority: high
state_hub_task_id: "6b8663b8-4143-4361-ad9e-7df0c4f4a19d"

research/CB-RES-0007-*.md.

Do not survey issue trackers. The question is narrower and more interesting: how do rigorous rule systems record the ambiguity they found? Candidates worth a benchmark-to-beat:

  • Errata and rulings practice in published games (Magic's comprehensive-rules + rulings split, Netrunner's NAPD card rulings) — what makes a ruling findable years later.
  • Formal-methods counterexample traces — a model checker's output is precisely a reproduction attached to a claim, which is the shape wanted here.
  • Conformance-suite provisional behaviour — how W3C/WHATWG mark "implementation-defined" and how a spec later absorbs it.
  • What this repo already has: provisional: true scenarios, §Underdetermined, gates.toml's caught/retire_if shape, and the hub message that went unread. The register must reuse the provisional machinery rather than compete with it.

Name, per dimension, the property to beat — findability, reproducibility, and whether a ruling can close a finding mechanically.

Done 2026-08-03. CB-RES-0007, with a runnable baseline (tools/design-baseline.py).

The baseline is us, and it is measured: 6 findings across 11 files with no index, 2 of 6 (33%) with a runnable reproduction, and U1U10 raised 2026-07-30, first read 2026-08-03 — 4 days, 0 of 10 ruled.

The uncomfortable number is stated up front rather than left for the review to find: the reproduction rule would reject four of our six existing findings. The survey answers it — none of the four is expensive to reproduce, so the 33% is evidence that nobody was ever asked for one, not that the rule is unaffordable.

Magic corrected an assumption this pass was about to build on. I expected a ruling to be the authoritative resolution. It is not: rulings are "reminder information with no actual weight or rules meaning", and the authoritative fix folds into the Oracle card text. A finding closes when the source changes, not when an annotation is added — so the register must be a queue that empties, not an archive that grows. That is now a constraint on T03's lifecycle.

Model checkers supplied the reproduction rule independently (a counterexample trace is the finding), and W3C's implementation-defined mark is the one piece of machinery we already have and must reuse rather than duplicate.

Task: adversarial review

id: CB-WP-0022-T02
status: done
priority: high
state_hub_task_id: "d2597895-fe2e-4f1e-a3db-a5fb8833434c"

Tier L requires it. Give the reviewer the survey and the §judgment above, and require an attempt at:

  • that the engine-evolution register is redundant — the strongest counter is that ADRs record decisions and evidence records findings, but nothing records what we learned about building engines, which is a third thing;
  • that "carries its reproduction" is affordable — if half the real findings cannot be reproduced cheaply, the rule will be quietly dropped and the register becomes a suggestion box anyway. (Since this task was written the answer has hardened: two findings reached ground-game on wrong premises, and ground-game has itself ruled that reproductions are required. The reviewer should now press the opposite way — whether the rule is sufficient, given both bad premises would have shipped a passing test alongside a wrong description.)
  • that a register is needed at all, rather than one more section in GroundRules.md §Underdetermined, which already exists and already works.

Record the trail in history/, unpolished.

Done 2026-08-05. Trail: challenge, response.

Run by a separate agent — the first in this repo that was. CB-RES-0006's review opened by conceding it could not be, and called its own findings "a lower bound on what a genuinely separate reviewer would find." That was measurable, and this is the measurement: the separate reviewer ran git log against the survey's central example and found our own commit had falsified it four days earlier, while the author — who wrote that commit — quoted the dead number twice.

Seven challenges: four conceded, two conceded in part, one answered.

  • C1 lands hardest and changed the design. The rule's showcase finding was false and had passed the rule. Existence is not the missing property; shape and falsifiability are. Folded into §The load-bearing rule above, and it is T03's to settle.
  • C1 also caught a defect in flight — T06's payload, todo, still named the dead number. Withdrawn above before sending.
  • C2 withdrew the baseline's precision. tools/design-baseline.py is a hand-maintained dict counting itself (:16-36, :89); has_reproduction (:38-43) never checks the file exists, so the self-test's YES-control (:63) is green against a path 2da19a4 deleted. Makefile:127 runs only --self-test, so the reporting path has no CI. The direction stands — 11 files, no index, 0 of 10 ruled are all checkable without the tool — but 33% is not a measured rate and T05 must not build on it.
  • C3: "six provisional defaults" is five, and GR-E01 is double-counted in the 2/6. No corrected rate is quoted here; the instrument that would produce it is the one C2 withdrew.
  • C4: §Underdetermined was never evaluated as a candidate, and it already delivers four of five benchmarks including the Oracle property the survey went to Magic to find. T03's burden flips: argue why it is extended, not replaced.
  • C5: the engine-evolution "third thing" is visible in specs/InnerLoopReference.md and history/'s retrospectives, neither of which my redundancy inventory named. Conclusion narrowed, not settled.
  • Survived: the reproduction rule's affordability, and §4's reuse of the provisional machinery. Both with stated falsifiers.

Task: decide

id: CB-WP-0022-T03
status: todo
priority: high
state_hub_task_id: "a2e85810-e949-4ea0-81fc-36b912af326c"

decisions/ADR-0012-*.md. At minimum:

  • whether clay-borg's INTENT gains a stated aspect as a design instrument, and in what words — this is the change with the longest half-life in the pass;
  • the finding taxonomy, and it should be grounded in the six findings above rather than invented: underdetermined (rules do not say), inconsistent (rules disagree with each other or the data), inert (a rule that cannot fire), degenerate (fires, but collapses play), unplayed (implemented, never played);
  • the lifecycle and who owns each state: raised → reported → ruled → applied, or withdrawn;
  • whether a finding without a reproduction is rejected or admitted as a note — and if admitted, how it is prevented from aging into an apparent finding.

Task: specify

id: CB-WP-0022-T04
status: todo
priority: high
state_hub_task_id: "60ddfeab-81b1-45e4-9ec9-91cc8ea7fb72"

specs/GameDesign.md, with metrics, because a spec without them is prose.

Candidate measures, to be argued not adopted:

  • findings with a runnable reproduction — target 100%, and the denominator includes withdrawn ones;
  • time from raised to reported — the U-items took four days to be read; that is the number this exists to fix;
  • findings closed by a ruling vs findings still open, with age.

ground-game has ruled on what a finding must carry (GROUND-WP-0004 T02, 2026-08-04), and it is stricter than this pass proposed. Adopt it:

  1. Arithmetic findings must ship a runnable reproduction and a row-level deal table listing Surface and each hidden priority separately — never only "sum of file" or "deal depth N".
  2. ground-game rulings that depend on arithmetic should cite that reproduction id or path, so a later reader can re-run it.

The second half is theirs to keep. The first is ours, and it is a direct consequence of both premises we got wrong: "12 in the file" was a sum without a deal table, and "offered on a face-down Problem" was a condition named without checking which one actually fired.

So the reproduction rule gains a shape requirement, not just an existence one. A finding that ships a passing test but describes the wrong quantity is still a bad finding — which is exactly what happened twice. The spec must say what a reproduction has to show, not only that one exists.

Also specify the trial protocol, which is the part with no precedent here: a trial game is a --recorded session plus an observation log, so "we played it and X happened" is replayable rather than remembered. The engine already records sessions as scenarios; a trial is that plus notes, and it must cost almost nothing or it will not be done.

Task: build it, and backfill what is already known

id: CB-WP-0022-T05
status: todo
priority: high
state_hub_task_id: "469413ba-ed47-4840-a60e-ee7f95dc06ff"

The register, the tool, and then the six findings above entered into it — backfilling is the test. A register that cannot express findings the project already has is the wrong register, and discovering that after designing it is the point of doing it in this order.

make design (or equivalent) must report: open findings by kind, those without a reproduction, and those never reported to their owner.

Task: report to ground-game, mechanically

id: CB-WP-0022-T06
status: todo
priority: high
state_hub_task_id: "c6a659bd-c495-4754-9332-baadad51012a"

Generate the report and send it. The message that sat unread for four days is the baseline to beat, and the failure was not the message — it was that nothing pointed at it and nothing tracked whether it was answered.

So the report must land somewhere that persists: a file in ground-game under its own workplan, not only an inbox entry. GROUND-WP-0002 already holds the ten U-items; this should extend it rather than duplicate it.

Include the findings this pass has sharpened:

  • SOLVE's legality against a face-down Problem or an unmatchable suit — and note that the case we reported was not the case that fired (CB-WP-0023 T01).
  • GR-E01 vs GR-S01 — 4/6/9 against 5/7/9, no dataset reconciles themwithdrawn 2026-08-05, before sending. The adversarial review (C1) found 2da19a4 had already measured 6/9/12 against 5/7/9: the dataset reconciles them and the scenario is now -reachable-. Sending this would have been the fourth wrong premise to reach ground-game, and the only one caught before transmission. Report the withdrawal, not the finding — GROUND-WP-0002 holds the original, and a claim retracted silently is how the first three survived.

Task: evidence

id: CB-WP-0022-T07
status: todo
priority: high
state_hub_task_id: "4fb77911-486d-4d9e-a53d-c4e486c74ce2"

evidence/CB-EV-0021-*.md. (Was CB-EV-0020 when written; CB-WP-0023 shipped that number first — evidence/CB-EV-0020-solve-legality.md — so this one moves rather than collides.)

  • Whether backfilling changed the design — if all six findings fit the first taxonomy, say so and be suspicious of it.
  • What tier L cost against what it caught, since this is the second full-weight L pass and CB-WP-0012's deleted its own structural trigger.
  • The engine-evolution question, as the review left it.
  • Quote CB-WP-0021's cost by re-running the instrument.