clay-borg/workplans/CB-WP-0022-the-design-instrument.md

399 lines
18 KiB
Markdown
Raw Normal View History

Declare CB-WP-0022: the design instrument, tier L The maintainer named a new aspect: clay-borg as a game design tool, with a register for design flaws, questions, results and trial protocols. Structural L on the maintainer-named-high-leverage trigger, and it amends INTENT. Chaos d8=6, no override. The insight is that this is already happening with no home. Five passes have produced ten underdetermined rules points, SOLVE offered on a face-down problem and always inert, GR-A13's wasted SOLVE, GR-E01 unreachable below 5 seats, six provisional scenario defaults, and two scoring modes never played to the end -- every one found by BUILDING the simulator rather than by playing it. A simulator rigorous enough to refuse ambiguity is a design instrument, because it cannot proceed past a rule that does not decide. All of it has been carried in prose in six places and one sat unread in an inbox for four days. The load-bearing rule: a design finding is not admissible without its reproduction. A register that collects opinions would reproduce this project's standing failure -- unexecuted verification -- in a new medium. Recorded as a judgment for the adversarial review rather than assumed: the engine-evolution meta the maintainer also asked about should NOT be built, because evidence/, decisions/, gates.toml and workplans already carry nineteen passes of it with dates, costs and falsifiers. A second register for the same subject is ceremony. The asymmetry is the point -- engine evolution has a home and game design does not. T05 backfills the six known findings as the TEST of the register: one that cannot express findings the project already has is the wrong register, and discovering that after designing it is why the order is survey, review, decide, specify, build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 20:48:08 +02:00
---
id: CB-WP-0022
kind: product
title: "The design instrument: findings about the game, with their reproductions"
CB-WP-0022 T05/T06/T07: the register, and what its first run found T05. tools/design.py, make design, and the register in GroundRules.md -- 14 rows, no new file, because ADR-0012 D2 made §Underdetermined the register rather than building one beside it. Backfill was the test and it caught two things the ADR did not have. First, a `role` column. The first report alarmed on U2 and was wrong to: U2's scenario is green BECAUSE the provisional default it documents is implemented, which says nothing about whether ground-game agrees. GR-E01's was a counterexample that went green. Same colour, opposite meaning -- a register that cannot tell them apart either alarms constantly or never. Only a green counterexample alarms. Folded back into GameDesign §1.3. Second, it contradicted the survey. CB-RES-0007 said "six of the ten already have provisional scenarios." Measured -- grep -lE "\bU<n>\b" over scenarios/ground -- exactly ONE U-item names itself. Five provisional scenarios exist and four probably encode U-item defaults, but the mapping is not written down, so it is not checkable. Same defect class as the wrong premises, found inside the survey that proposed the fix. Now a reported debt: open, lacking a reproduction: 9, target 0. design.py carries the control design-baseline.py never had, asserted directly: a row citing a nonexistent file must not count as reproduced, using the exact path 2da19a4 deleted -- which the old tool called green. design-baseline.py is marked superseded rather than deleted; it is the evidence for how a wrong number got into a survey. T06. The report is a FILE in ground-game under GROUND-WP-0002, committed there, with a hub message that only points at it. It asks for no ruling: it carries GR-E01's withdrawal, our reproduction debt, and two notes that are explicitly not findings. And it had to acknowledge something nobody anticipated. GROUND-WP-0002 is finished -- all ten U-items were RULED 2026-08-03, every one confirmed as the default we simulate, plus five of six provisional scenarios. The survey said "0 of 10 ruled" two days later and this register was built saying `reported`. That is the unread-inbox failure running in the opposite direction: they answered and we did not collect it. The instrument's first run surfaced it. They are `ruled`, not `applied` -- lifting the now-settled provisional flags is owed and is not done, and make design shows them open until it is. T07. evidence/CB-EV-0021. Two of six catches in this pass came from execution rather than process (the role distinction from building it, the ten uncollected rulings from running it), which is InnerLoop §Design goal's prediction holding. make self-tests, facts-check, loop-lint: clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:22:33 +02:00
status: done
2026-08-04 00:21:21 +02:00
state_hub_workstream_id: "fda16340-0049-4acf-884b-a5cfdbde47c0"
Declare CB-WP-0022: the design instrument, tier L The maintainer named a new aspect: clay-borg as a game design tool, with a register for design flaws, questions, results and trial protocols. Structural L on the maintainer-named-high-leverage trigger, and it amends INTENT. Chaos d8=6, no override. The insight is that this is already happening with no home. Five passes have produced ten underdetermined rules points, SOLVE offered on a face-down problem and always inert, GR-A13's wasted SOLVE, GR-E01 unreachable below 5 seats, six provisional scenario defaults, and two scoring modes never played to the end -- every one found by BUILDING the simulator rather than by playing it. A simulator rigorous enough to refuse ambiguity is a design instrument, because it cannot proceed past a rule that does not decide. All of it has been carried in prose in six places and one sat unread in an inbox for four days. The load-bearing rule: a design finding is not admissible without its reproduction. A register that collects opinions would reproduce this project's standing failure -- unexecuted verification -- in a new medium. Recorded as a judgment for the adversarial review rather than assumed: the engine-evolution meta the maintainer also asked about should NOT be built, because evidence/, decisions/, gates.toml and workplans already carry nineteen passes of it with dates, costs and falsifiers. A second register for the same subject is ceremony. The asymmetry is the point -- engine evolution has a home and game design does not. T05 backfills the six known findings as the TEST of the register: one that cannot express findings the project already has is the wrong register, and discovering that after designing it is why the order is survey, review, decide, specify, build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 20:48:08 +02:00
---
# Purpose
```
structural tier L (named a high-leverage pass by the maintainer, and it
amends INTENT — clay-borg gains a stated aspect)
chaos d8 = 6 → no override
declared tier L
```
Declaration 5 of chaos window 2. Tier L: separate survey, **adversarial
review**, ADR, then spec, then code.
## The insight, in the maintainer's words
> *"We should consider ourselves testing the game and document
> inconsistencies to report them back to the ground-game repo, so that the
> game designer can improve the rules accordingly… We should have the
> rigorous game simulation engine and a meta scope to capture notes about
> game design flaws, questions, results and protocols about trial games…
> This will provide clay-borg an additional aspect as a valuable game
> design tool."*
**This is already happening and has no home.** In five passes the engine
has produced, as a by-product of being rigorous:
| finding | how it surfaced |
|---|---|
| ten underdetermined rules points (U1U10) | formalizing the dataset into testable rules |
| SOLVE offered on a face-down Problem, always inert | a human dragging it three rounds running |
| GR-A13 "wasted SOLVE" on a claimed Problem | a scenario that had to pick a default |
CB-WP-0022 T05/T06/T07: the register, and what its first run found T05. tools/design.py, make design, and the register in GroundRules.md -- 14 rows, no new file, because ADR-0012 D2 made §Underdetermined the register rather than building one beside it. Backfill was the test and it caught two things the ADR did not have. First, a `role` column. The first report alarmed on U2 and was wrong to: U2's scenario is green BECAUSE the provisional default it documents is implemented, which says nothing about whether ground-game agrees. GR-E01's was a counterexample that went green. Same colour, opposite meaning -- a register that cannot tell them apart either alarms constantly or never. Only a green counterexample alarms. Folded back into GameDesign §1.3. Second, it contradicted the survey. CB-RES-0007 said "six of the ten already have provisional scenarios." Measured -- grep -lE "\bU<n>\b" over scenarios/ground -- exactly ONE U-item names itself. Five provisional scenarios exist and four probably encode U-item defaults, but the mapping is not written down, so it is not checkable. Same defect class as the wrong premises, found inside the survey that proposed the fix. Now a reported debt: open, lacking a reproduction: 9, target 0. design.py carries the control design-baseline.py never had, asserted directly: a row citing a nonexistent file must not count as reproduced, using the exact path 2da19a4 deleted -- which the old tool called green. design-baseline.py is marked superseded rather than deleted; it is the evidence for how a wrong number got into a survey. T06. The report is a FILE in ground-game under GROUND-WP-0002, committed there, with a hub message that only points at it. It asks for no ruling: it carries GR-E01's withdrawal, our reproduction debt, and two notes that are explicitly not findings. And it had to acknowledge something nobody anticipated. GROUND-WP-0002 is finished -- all ten U-items were RULED 2026-08-03, every one confirmed as the default we simulate, plus five of six provisional scenarios. The survey said "0 of 10 ruled" two days later and this register was built saying `reported`. That is the unread-inbox failure running in the opposite direction: they answered and we did not collect it. The instrument's first run surfaced it. They are `ruled`, not `applied` -- lifting the now-settled provisional flags is owed and is not done, and make design shows them open until it is. T07. evidence/CB-EV-0021. Two of six catches in this pass came from execution rather than process (the role distinction from building it, the ten uncollected rulings from running it), which is InnerLoop §Design goal's prediction holding. make self-tests, facts-check, loop-lint: clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:22:33 +02:00
| ~~GR-E01 unreachable below 5 seats~~ | arithmetic over the deal count — **withdrawn 2026-08-05, it was wrong (C1)** |
| ~~six provisional scenario defaults~~ | **five, and double-counted with GR-E01 (C3)** — they are reproductions, not a finding |
Declare CB-WP-0022: the design instrument, tier L The maintainer named a new aspect: clay-borg as a game design tool, with a register for design flaws, questions, results and trial protocols. Structural L on the maintainer-named-high-leverage trigger, and it amends INTENT. Chaos d8=6, no override. The insight is that this is already happening with no home. Five passes have produced ten underdetermined rules points, SOLVE offered on a face-down problem and always inert, GR-A13's wasted SOLVE, GR-E01 unreachable below 5 seats, six provisional scenario defaults, and two scoring modes never played to the end -- every one found by BUILDING the simulator rather than by playing it. A simulator rigorous enough to refuse ambiguity is a design instrument, because it cannot proceed past a rule that does not decide. All of it has been carried in prose in six places and one sat unread in an inbox for four days. The load-bearing rule: a design finding is not admissible without its reproduction. A register that collects opinions would reproduce this project's standing failure -- unexecuted verification -- in a new medium. Recorded as a judgment for the adversarial review rather than assumed: the engine-evolution meta the maintainer also asked about should NOT be built, because evidence/, decisions/, gates.toml and workplans already carry nineteen passes of it with dates, costs and falsifiers. A second register for the same subject is ceremony. The asymmetry is the point -- engine evolution has a home and game design does not. T05 backfills the six known findings as the TEST of the register: one that cannot express findings the project already has is the wrong register, and discovering that after designing it is why the order is survey, review, decide, specify, build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 20:48:08 +02:00
| GR-E03 / GR-E04 never played to the end | nobody noticed for nineteen passes |
Every one was found by *building the simulator*, not by playing. That is
the thing worth naming: **a simulator rigorous enough to refuse ambiguity
is a design instrument, because it cannot proceed past a rule that does
not decide.**
And every one of them has been carried in prose, in six different places,
and one sat unread in an inbox for four days.
## The load-bearing rule this must have
The project's standing failure is *unexecuted verification*. A design
register that collects opinions would reproduce it in a new medium.
> **A design finding is not admissible without its reproduction.**
Concretely: a scenario that fails, an arithmetic check that prints the
contradiction, a recorded game the reader can replay, or a named test.
CB-WP-0022 T02: the separate reviewer found the showcase finding was false First adversarial review in this repo run by a genuinely separate agent. CB-RES-0006's reviewer opened by conceding it could not be, and called its own findings "a lower bound on what a genuinely separate reviewer would find." This is the measurement: the separate reviewer ran git log against the survey's central example and found 2da19a4 had falsified it four days earlier, while the author -- who wrote that commit -- quoted the dead number twice. Seven challenges: four conceded, two conceded in part, one answered. C1 changes the design. "GR-E01 is admissible because 4/6/9 against 5/7/9 is a computation anyone can rerun" was a computation already rerun: the edition import measured 6/9/12, the conclusion inverted, and the scenario was renamed -unreachable- to -reachable-. That finding was one of the TWO that passed the reproduction rule. So three wrong premises have now reached ground-game and the third satisfied an existence test -- existence is not the property that was missing. The rule gains shape (ground-game's row-level deal table, promoted from a T04 addendum) and a clause the survey never contemplated: a reproduction must be able to fail. Ours went green and stayed admissible. C1 also caught a defect in flight. T06's payload, status todo, still named 4/6/9 and was queued to send it to ground-game as "no dataset reconciles them." Withdrawn before sending -- the fourth wrong premise, and the only one stopped. C2 withdraws the baseline's precision: design-baseline.py is a hand-maintained dict counting itself, has_reproduction never checks the file exists (its YES-control is green against a deleted path), and Makefile:127 runs only --self-test so the reporting path has no CI. The direction survives; 33% is not a measured rate and T05 must not build on it. C3: "six provisional defaults" is five, GR-E01 double-counted. C4: GroundRules §Underdetermined was never evaluated as a candidate and already delivers four of five benchmarks -- T03's burden flips to arguing extension over replacement. Survived: the rule's affordability, and reuse of the provisional machinery. loop-lint: no findings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:03:12 +02:00
*"This feels unbalanced"* is a note, not a finding. **The SOLVE inertness
is admissible because a recorded session shows three no-ops.**
> **The example that stood here was GR-E01, and the review killed it
> (C1).** *"4/6/9 against 5/7/9 is a computation anyone can rerun"* had
> already been rerun: `2da19a4` measured **6/9/12**, and the scenario was
> renamed `-unreachable-` → `-reachable-`. The conclusion inverted — and
> it was one of the two findings that **passed** this rule. So existence
> is not what was missing. See [ADR-0012](../decisions/ADR-0012-the-design-instrument.md)
> D3: the rule gains **shape**, and **a reproduction must be able to
CB-WP-0022 T02: the separate reviewer found the showcase finding was false First adversarial review in this repo run by a genuinely separate agent. CB-RES-0006's reviewer opened by conceding it could not be, and called its own findings "a lower bound on what a genuinely separate reviewer would find." This is the measurement: the separate reviewer ran git log against the survey's central example and found 2da19a4 had falsified it four days earlier, while the author -- who wrote that commit -- quoted the dead number twice. Seven challenges: four conceded, two conceded in part, one answered. C1 changes the design. "GR-E01 is admissible because 4/6/9 against 5/7/9 is a computation anyone can rerun" was a computation already rerun: the edition import measured 6/9/12, the conclusion inverted, and the scenario was renamed -unreachable- to -reachable-. That finding was one of the TWO that passed the reproduction rule. So three wrong premises have now reached ground-game and the third satisfied an existence test -- existence is not the property that was missing. The rule gains shape (ground-game's row-level deal table, promoted from a T04 addendum) and a clause the survey never contemplated: a reproduction must be able to fail. Ours went green and stayed admissible. C1 also caught a defect in flight. T06's payload, status todo, still named 4/6/9 and was queued to send it to ground-game as "no dataset reconciles them." Withdrawn before sending -- the fourth wrong premise, and the only one stopped. C2 withdraws the baseline's precision: design-baseline.py is a hand-maintained dict counting itself, has_reproduction never checks the file exists (its YES-control is green against a deleted path), and Makefile:127 runs only --self-test so the reporting path has no CI. The direction survives; 33% is not a measured rate and T05 must not build on it. C3: "six provisional defaults" is five, GR-E01 double-counted. C4: GroundRules §Underdetermined was never evaluated as a candidate and already delivers four of five benchmarks -- T03's burden flips to arguing extension over replacement. Survived: the rule's affordability, and reuse of the provisional machinery. loop-lint: no findings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:03:12 +02:00
> fail.** Ours went green and stayed admissible.
Declare CB-WP-0022: the design instrument, tier L The maintainer named a new aspect: clay-borg as a game design tool, with a register for design flaws, questions, results and trial protocols. Structural L on the maintainer-named-high-leverage trigger, and it amends INTENT. Chaos d8=6, no override. The insight is that this is already happening with no home. Five passes have produced ten underdetermined rules points, SOLVE offered on a face-down problem and always inert, GR-A13's wasted SOLVE, GR-E01 unreachable below 5 seats, six provisional scenario defaults, and two scoring modes never played to the end -- every one found by BUILDING the simulator rather than by playing it. A simulator rigorous enough to refuse ambiguity is a design instrument, because it cannot proceed past a rule that does not decide. All of it has been carried in prose in six places and one sat unread in an inbox for four days. The load-bearing rule: a design finding is not admissible without its reproduction. A register that collects opinions would reproduce this project's standing failure -- unexecuted verification -- in a new medium. Recorded as a judgment for the adversarial review rather than assumed: the engine-evolution meta the maintainer also asked about should NOT be built, because evidence/, decisions/, gates.toml and workplans already carry nineteen passes of it with dates, costs and falsifiers. A second register for the same subject is ceremony. The asymmetry is the point -- engine evolution has a home and game design does not. T05 backfills the six known findings as the TEST of the register: one that cannot express findings the project already has is the wrong register, and discovering that after designing it is why the order is survey, review, decide, specify, build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 20:48:08 +02:00
This is what would make clay-borg a design tool rather than a suggestion
box, and it is the one part of this proposal that must not be traded away
for convenience.
## The judgment I want reviewed, not assumed
The maintainer asked whether this should extend to *"a meta about the
clay-borg engine evolution itself."*
**My answer is no, and it should be argued rather than accepted.** That
CB-WP-0022 T05/T06/T07: the register, and what its first run found T05. tools/design.py, make design, and the register in GroundRules.md -- 14 rows, no new file, because ADR-0012 D2 made §Underdetermined the register rather than building one beside it. Backfill was the test and it caught two things the ADR did not have. First, a `role` column. The first report alarmed on U2 and was wrong to: U2's scenario is green BECAUSE the provisional default it documents is implemented, which says nothing about whether ground-game agrees. GR-E01's was a counterexample that went green. Same colour, opposite meaning -- a register that cannot tell them apart either alarms constantly or never. Only a green counterexample alarms. Folded back into GameDesign §1.3. Second, it contradicted the survey. CB-RES-0007 said "six of the ten already have provisional scenarios." Measured -- grep -lE "\bU<n>\b" over scenarios/ground -- exactly ONE U-item names itself. Five provisional scenarios exist and four probably encode U-item defaults, but the mapping is not written down, so it is not checkable. Same defect class as the wrong premises, found inside the survey that proposed the fix. Now a reported debt: open, lacking a reproduction: 9, target 0. design.py carries the control design-baseline.py never had, asserted directly: a row citing a nonexistent file must not count as reproduced, using the exact path 2da19a4 deleted -- which the old tool called green. design-baseline.py is marked superseded rather than deleted; it is the evidence for how a wrong number got into a survey. T06. The report is a FILE in ground-game under GROUND-WP-0002, committed there, with a hub message that only points at it. It asks for no ruling: it carries GR-E01's withdrawal, our reproduction debt, and two notes that are explicitly not findings. And it had to acknowledge something nobody anticipated. GROUND-WP-0002 is finished -- all ten U-items were RULED 2026-08-03, every one confirmed as the default we simulate, plus five of six provisional scenarios. The survey said "0 of 10 ruled" two days later and this register was built saying `reported`. That is the unread-inbox failure running in the opposite direction: they answered and we did not collect it. The instrument's first run surfaced it. They are `ruled`, not `applied` -- lifting the now-settled provisional flags is owed and is not done, and make design shows them open until it is. T07. evidence/CB-EV-0021. Two of six catches in this pass came from execution rather than process (the role distinction from building it, the ten uncollected rulings from running it), which is InnerLoop §Design goal's prediction holding. make self-tests, facts-check, loop-lint: clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:22:33 +02:00
register already exists and is load-bearing: `evidence/CB-EV-*`,
`decisions/ADR-*`, `gates.toml` and the workplans capture nineteen passes
with dates, costs and falsifiers. A second register for the same subject
would be ceremony. The asymmetry is the point: engine evolution has a home
and game design does not.
> **Settled by [ADR-0012](../decisions/ADR-0012-the-design-instrument.md)
> D7: no register — but the argument above did not survive.** C5 found the
> "third thing" the maintainer meant is visible in
> `specs/InnerLoopReference.md` and `history/`'s retrospectives, **neither
> of which this inventory names.** Conclusion narrowed, not settled: if
> InnerLoopReference keeps absorbing material that is neither a decision
> nor a finding, revisit.
Declare CB-WP-0022: the design instrument, tier L The maintainer named a new aspect: clay-borg as a game design tool, with a register for design flaws, questions, results and trial protocols. Structural L on the maintainer-named-high-leverage trigger, and it amends INTENT. Chaos d8=6, no override. The insight is that this is already happening with no home. Five passes have produced ten underdetermined rules points, SOLVE offered on a face-down problem and always inert, GR-A13's wasted SOLVE, GR-E01 unreachable below 5 seats, six provisional scenario defaults, and two scoring modes never played to the end -- every one found by BUILDING the simulator rather than by playing it. A simulator rigorous enough to refuse ambiguity is a design instrument, because it cannot proceed past a rule that does not decide. All of it has been carried in prose in six places and one sat unread in an inbox for four days. The load-bearing rule: a design finding is not admissible without its reproduction. A register that collects opinions would reproduce this project's standing failure -- unexecuted verification -- in a new medium. Recorded as a judgment for the adversarial review rather than assumed: the engine-evolution meta the maintainer also asked about should NOT be built, because evidence/, decisions/, gates.toml and workplans already carry nineteen passes of it with dates, costs and falsifiers. A second register for the same subject is ceremony. The asymmetry is the point -- engine evolution has a home and game design does not. T05 backfills the six known findings as the TEST of the register: one that cannot express findings the project already has is the wrong register, and discovering that after designing it is why the order is survey, review, decide, specify, build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 20:48:08 +02:00
## Task: survey how this is done elsewhere, and what we already have
```task
id: CB-WP-0022-T01
CB-WP-0022-T01: survey — how rule systems record the ambiguity they find CB-RES-0007 plus a runnable baseline harness. Tier L invokes the runnable-baseline option; the external candidates are practices rather than software, so their rows are directional and cap at parity, and the row that CAN be run is our own. Measured: 6 findings across 11 files with no index, 2 of 6 (33%) with a runnable reproduction, U1-U10 raised 2026-07-30 and first READ 2026-08-03 -- 4 days, 0 of 10 ruled. The uncomfortable number is stated before the review can find it: the proposed 'no finding without its reproduction' rule would reject four of our six existing findings. The survey answers rather than routes around it -- none of the four is expensive to reproduce, so 33% is evidence nobody was ever asked for one. Magic corrected an assumption this pass was about to build on. Rulings are NOT authoritative -- they are 'reminder information with no actual weight or rules meaning' -- and the authoritative fix folds into the Oracle card text. So a finding closes when the SOURCE changes, not when an annotation is added, and the register must be a queue that empties rather than an archive that grows. That is now a constraint on the ADR's lifecycle. Model checkers supply the reproduction rule independently: a counterexample trace IS the finding. W3C's implementation-defined mark is the machinery we already have in provisional: scenarios and must reuse. The loop-lint gate caught the new tool with no --self-test; it has one, pinning the 2-of-6 baseline so a later edit cannot move it silently. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 21:19:12 +02:00
status: done
Declare CB-WP-0022: the design instrument, tier L The maintainer named a new aspect: clay-borg as a game design tool, with a register for design flaws, questions, results and trial protocols. Structural L on the maintainer-named-high-leverage trigger, and it amends INTENT. Chaos d8=6, no override. The insight is that this is already happening with no home. Five passes have produced ten underdetermined rules points, SOLVE offered on a face-down problem and always inert, GR-A13's wasted SOLVE, GR-E01 unreachable below 5 seats, six provisional scenario defaults, and two scoring modes never played to the end -- every one found by BUILDING the simulator rather than by playing it. A simulator rigorous enough to refuse ambiguity is a design instrument, because it cannot proceed past a rule that does not decide. All of it has been carried in prose in six places and one sat unread in an inbox for four days. The load-bearing rule: a design finding is not admissible without its reproduction. A register that collects opinions would reproduce this project's standing failure -- unexecuted verification -- in a new medium. Recorded as a judgment for the adversarial review rather than assumed: the engine-evolution meta the maintainer also asked about should NOT be built, because evidence/, decisions/, gates.toml and workplans already carry nineteen passes of it with dates, costs and falsifiers. A second register for the same subject is ceremony. The asymmetry is the point -- engine evolution has a home and game design does not. T05 backfills the six known findings as the TEST of the register: one that cannot express findings the project already has is the wrong register, and discovering that after designing it is why the order is survey, review, decide, specify, build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 20:48:08 +02:00
priority: high
2026-08-04 00:21:21 +02:00
state_hub_task_id: "6b8663b8-4143-4361-ad9e-7df0c4f4a19d"
Declare CB-WP-0022: the design instrument, tier L The maintainer named a new aspect: clay-borg as a game design tool, with a register for design flaws, questions, results and trial protocols. Structural L on the maintainer-named-high-leverage trigger, and it amends INTENT. Chaos d8=6, no override. The insight is that this is already happening with no home. Five passes have produced ten underdetermined rules points, SOLVE offered on a face-down problem and always inert, GR-A13's wasted SOLVE, GR-E01 unreachable below 5 seats, six provisional scenario defaults, and two scoring modes never played to the end -- every one found by BUILDING the simulator rather than by playing it. A simulator rigorous enough to refuse ambiguity is a design instrument, because it cannot proceed past a rule that does not decide. All of it has been carried in prose in six places and one sat unread in an inbox for four days. The load-bearing rule: a design finding is not admissible without its reproduction. A register that collects opinions would reproduce this project's standing failure -- unexecuted verification -- in a new medium. Recorded as a judgment for the adversarial review rather than assumed: the engine-evolution meta the maintainer also asked about should NOT be built, because evidence/, decisions/, gates.toml and workplans already carry nineteen passes of it with dates, costs and falsifiers. A second register for the same subject is ceremony. The asymmetry is the point -- engine evolution has a home and game design does not. T05 backfills the six known findings as the TEST of the register: one that cannot express findings the project already has is the wrong register, and discovering that after designing it is why the order is survey, review, decide, specify, build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 20:48:08 +02:00
```
`research/CB-RES-0007-*.md`.
**Do not survey issue trackers.** The question is narrower and more
interesting: how do rigorous rule systems record *the ambiguity they
found*? Candidates worth a benchmark-to-beat:
- **Errata and rulings practice** in published games (Magic's
comprehensive-rules + rulings split, Netrunner's NAPD card rulings) —
what makes a ruling *findable* years later.
- **Formal-methods counterexample traces** — a model checker's output is
precisely a reproduction attached to a claim, which is the shape wanted
here.
- **Conformance-suite provisional behaviour** — how W3C/WHATWG mark
"implementation-defined" and how a spec later absorbs it.
- **What this repo already has**: `provisional: true` scenarios,
`§Underdetermined`, `gates.toml`'s `caught`/`retire_if` shape, and the
hub message that went unread. **The register must reuse the provisional
machinery rather than compete with it.**
Name, per dimension, the property to beat — findability, reproducibility,
and whether a ruling can *close* a finding mechanically.
CB-WP-0022-T01: survey — how rule systems record the ambiguity they find CB-RES-0007 plus a runnable baseline harness. Tier L invokes the runnable-baseline option; the external candidates are practices rather than software, so their rows are directional and cap at parity, and the row that CAN be run is our own. Measured: 6 findings across 11 files with no index, 2 of 6 (33%) with a runnable reproduction, U1-U10 raised 2026-07-30 and first READ 2026-08-03 -- 4 days, 0 of 10 ruled. The uncomfortable number is stated before the review can find it: the proposed 'no finding without its reproduction' rule would reject four of our six existing findings. The survey answers rather than routes around it -- none of the four is expensive to reproduce, so 33% is evidence nobody was ever asked for one. Magic corrected an assumption this pass was about to build on. Rulings are NOT authoritative -- they are 'reminder information with no actual weight or rules meaning' -- and the authoritative fix folds into the Oracle card text. So a finding closes when the SOURCE changes, not when an annotation is added, and the register must be a queue that empties rather than an archive that grows. That is now a constraint on the ADR's lifecycle. Model checkers supply the reproduction rule independently: a counterexample trace IS the finding. W3C's implementation-defined mark is the machinery we already have in provisional: scenarios and must reuse. The loop-lint gate caught the new tool with no --self-test; it has one, pinning the 2-of-6 baseline so a later edit cannot move it silently. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 21:19:12 +02:00
**Done 2026-08-03.**
[CB-RES-0007](../research/CB-RES-0007-design-instrument.md), with a
runnable baseline (`tools/design-baseline.py`).
CB-WP-0022 T05/T06/T07: the register, and what its first run found T05. tools/design.py, make design, and the register in GroundRules.md -- 14 rows, no new file, because ADR-0012 D2 made §Underdetermined the register rather than building one beside it. Backfill was the test and it caught two things the ADR did not have. First, a `role` column. The first report alarmed on U2 and was wrong to: U2's scenario is green BECAUSE the provisional default it documents is implemented, which says nothing about whether ground-game agrees. GR-E01's was a counterexample that went green. Same colour, opposite meaning -- a register that cannot tell them apart either alarms constantly or never. Only a green counterexample alarms. Folded back into GameDesign §1.3. Second, it contradicted the survey. CB-RES-0007 said "six of the ten already have provisional scenarios." Measured -- grep -lE "\bU<n>\b" over scenarios/ground -- exactly ONE U-item names itself. Five provisional scenarios exist and four probably encode U-item defaults, but the mapping is not written down, so it is not checkable. Same defect class as the wrong premises, found inside the survey that proposed the fix. Now a reported debt: open, lacking a reproduction: 9, target 0. design.py carries the control design-baseline.py never had, asserted directly: a row citing a nonexistent file must not count as reproduced, using the exact path 2da19a4 deleted -- which the old tool called green. design-baseline.py is marked superseded rather than deleted; it is the evidence for how a wrong number got into a survey. T06. The report is a FILE in ground-game under GROUND-WP-0002, committed there, with a hub message that only points at it. It asks for no ruling: it carries GR-E01's withdrawal, our reproduction debt, and two notes that are explicitly not findings. And it had to acknowledge something nobody anticipated. GROUND-WP-0002 is finished -- all ten U-items were RULED 2026-08-03, every one confirmed as the default we simulate, plus five of six provisional scenarios. The survey said "0 of 10 ruled" two days later and this register was built saying `reported`. That is the unread-inbox failure running in the opposite direction: they answered and we did not collect it. The instrument's first run surfaced it. They are `ruled`, not `applied` -- lifting the now-settled provisional flags is owed and is not done, and make design shows them open until it is. T07. evidence/CB-EV-0021. Two of six catches in this pass came from execution rather than process (the role distinction from building it, the ten uncollected rulings from running it), which is InnerLoop §Design goal's prediction holding. make self-tests, facts-check, loop-lint: clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:22:33 +02:00
**Its numbers were withdrawn by T02 and must not be quoted from here.**
The survey reported *6 findings, 2/6 = 33% reproduced, 11 files, 4 days*.
C2 showed the instrument counted itself and its reproduction check never
stat'd the file; C3 showed *"six provisional defaults"* is five and GR-E01
was double-counted; T05's backfill contradicted *"six of the ten have
provisional scenarios"* — **one** does. What survives is direction: many
files, no index, 0 of 10 ruled. The first honest figures are T05's.
**Magic corrected an assumption this pass was about to build on.** Rulings
are *"reminder information with no actual weight or rules meaning"*; the
authoritative fix folds into the **Oracle** card text. **A finding closes
when the source changes, not when an annotation is added** — the register
is a queue that empties. Model checkers supplied the reproduction rule
independently, and W3C's *implementation-defined* mark is machinery we
already have and must reuse rather than duplicate.
CB-WP-0022-T01: survey — how rule systems record the ambiguity they find CB-RES-0007 plus a runnable baseline harness. Tier L invokes the runnable-baseline option; the external candidates are practices rather than software, so their rows are directional and cap at parity, and the row that CAN be run is our own. Measured: 6 findings across 11 files with no index, 2 of 6 (33%) with a runnable reproduction, U1-U10 raised 2026-07-30 and first READ 2026-08-03 -- 4 days, 0 of 10 ruled. The uncomfortable number is stated before the review can find it: the proposed 'no finding without its reproduction' rule would reject four of our six existing findings. The survey answers rather than routes around it -- none of the four is expensive to reproduce, so 33% is evidence nobody was ever asked for one. Magic corrected an assumption this pass was about to build on. Rulings are NOT authoritative -- they are 'reminder information with no actual weight or rules meaning' -- and the authoritative fix folds into the Oracle card text. So a finding closes when the SOURCE changes, not when an annotation is added, and the register must be a queue that empties rather than an archive that grows. That is now a constraint on the ADR's lifecycle. Model checkers supply the reproduction rule independently: a counterexample trace IS the finding. W3C's implementation-defined mark is the machinery we already have in provisional: scenarios and must reuse. The loop-lint gate caught the new tool with no --self-test; it has one, pinning the 2-of-6 baseline so a later edit cannot move it silently. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 21:19:12 +02:00
Declare CB-WP-0022: the design instrument, tier L The maintainer named a new aspect: clay-borg as a game design tool, with a register for design flaws, questions, results and trial protocols. Structural L on the maintainer-named-high-leverage trigger, and it amends INTENT. Chaos d8=6, no override. The insight is that this is already happening with no home. Five passes have produced ten underdetermined rules points, SOLVE offered on a face-down problem and always inert, GR-A13's wasted SOLVE, GR-E01 unreachable below 5 seats, six provisional scenario defaults, and two scoring modes never played to the end -- every one found by BUILDING the simulator rather than by playing it. A simulator rigorous enough to refuse ambiguity is a design instrument, because it cannot proceed past a rule that does not decide. All of it has been carried in prose in six places and one sat unread in an inbox for four days. The load-bearing rule: a design finding is not admissible without its reproduction. A register that collects opinions would reproduce this project's standing failure -- unexecuted verification -- in a new medium. Recorded as a judgment for the adversarial review rather than assumed: the engine-evolution meta the maintainer also asked about should NOT be built, because evidence/, decisions/, gates.toml and workplans already carry nineteen passes of it with dates, costs and falsifiers. A second register for the same subject is ceremony. The asymmetry is the point -- engine evolution has a home and game design does not. T05 backfills the six known findings as the TEST of the register: one that cannot express findings the project already has is the wrong register, and discovering that after designing it is why the order is survey, review, decide, specify, build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 20:48:08 +02:00
## Task: adversarial review
```task
id: CB-WP-0022-T02
CB-WP-0022 T02: the separate reviewer found the showcase finding was false First adversarial review in this repo run by a genuinely separate agent. CB-RES-0006's reviewer opened by conceding it could not be, and called its own findings "a lower bound on what a genuinely separate reviewer would find." This is the measurement: the separate reviewer ran git log against the survey's central example and found 2da19a4 had falsified it four days earlier, while the author -- who wrote that commit -- quoted the dead number twice. Seven challenges: four conceded, two conceded in part, one answered. C1 changes the design. "GR-E01 is admissible because 4/6/9 against 5/7/9 is a computation anyone can rerun" was a computation already rerun: the edition import measured 6/9/12, the conclusion inverted, and the scenario was renamed -unreachable- to -reachable-. That finding was one of the TWO that passed the reproduction rule. So three wrong premises have now reached ground-game and the third satisfied an existence test -- existence is not the property that was missing. The rule gains shape (ground-game's row-level deal table, promoted from a T04 addendum) and a clause the survey never contemplated: a reproduction must be able to fail. Ours went green and stayed admissible. C1 also caught a defect in flight. T06's payload, status todo, still named 4/6/9 and was queued to send it to ground-game as "no dataset reconciles them." Withdrawn before sending -- the fourth wrong premise, and the only one stopped. C2 withdraws the baseline's precision: design-baseline.py is a hand-maintained dict counting itself, has_reproduction never checks the file exists (its YES-control is green against a deleted path), and Makefile:127 runs only --self-test so the reporting path has no CI. The direction survives; 33% is not a measured rate and T05 must not build on it. C3: "six provisional defaults" is five, GR-E01 double-counted. C4: GroundRules §Underdetermined was never evaluated as a candidate and already delivers four of five benchmarks -- T03's burden flips to arguing extension over replacement. Survived: the rule's affordability, and reuse of the provisional machinery. loop-lint: no findings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:03:12 +02:00
status: done
Declare CB-WP-0022: the design instrument, tier L The maintainer named a new aspect: clay-borg as a game design tool, with a register for design flaws, questions, results and trial protocols. Structural L on the maintainer-named-high-leverage trigger, and it amends INTENT. Chaos d8=6, no override. The insight is that this is already happening with no home. Five passes have produced ten underdetermined rules points, SOLVE offered on a face-down problem and always inert, GR-A13's wasted SOLVE, GR-E01 unreachable below 5 seats, six provisional scenario defaults, and two scoring modes never played to the end -- every one found by BUILDING the simulator rather than by playing it. A simulator rigorous enough to refuse ambiguity is a design instrument, because it cannot proceed past a rule that does not decide. All of it has been carried in prose in six places and one sat unread in an inbox for four days. The load-bearing rule: a design finding is not admissible without its reproduction. A register that collects opinions would reproduce this project's standing failure -- unexecuted verification -- in a new medium. Recorded as a judgment for the adversarial review rather than assumed: the engine-evolution meta the maintainer also asked about should NOT be built, because evidence/, decisions/, gates.toml and workplans already carry nineteen passes of it with dates, costs and falsifiers. A second register for the same subject is ceremony. The asymmetry is the point -- engine evolution has a home and game design does not. T05 backfills the six known findings as the TEST of the register: one that cannot express findings the project already has is the wrong register, and discovering that after designing it is why the order is survey, review, decide, specify, build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 20:48:08 +02:00
priority: high
2026-08-04 00:21:21 +02:00
state_hub_task_id: "d2597895-fe2e-4f1e-a3db-a5fb8833434c"
Declare CB-WP-0022: the design instrument, tier L The maintainer named a new aspect: clay-borg as a game design tool, with a register for design flaws, questions, results and trial protocols. Structural L on the maintainer-named-high-leverage trigger, and it amends INTENT. Chaos d8=6, no override. The insight is that this is already happening with no home. Five passes have produced ten underdetermined rules points, SOLVE offered on a face-down problem and always inert, GR-A13's wasted SOLVE, GR-E01 unreachable below 5 seats, six provisional scenario defaults, and two scoring modes never played to the end -- every one found by BUILDING the simulator rather than by playing it. A simulator rigorous enough to refuse ambiguity is a design instrument, because it cannot proceed past a rule that does not decide. All of it has been carried in prose in six places and one sat unread in an inbox for four days. The load-bearing rule: a design finding is not admissible without its reproduction. A register that collects opinions would reproduce this project's standing failure -- unexecuted verification -- in a new medium. Recorded as a judgment for the adversarial review rather than assumed: the engine-evolution meta the maintainer also asked about should NOT be built, because evidence/, decisions/, gates.toml and workplans already carry nineteen passes of it with dates, costs and falsifiers. A second register for the same subject is ceremony. The asymmetry is the point -- engine evolution has a home and game design does not. T05 backfills the six known findings as the TEST of the register: one that cannot express findings the project already has is the wrong register, and discovering that after designing it is why the order is survey, review, decide, specify, build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 20:48:08 +02:00
```
Tier L requires it. Give the reviewer the survey **and** the §judgment
above, and require an attempt at:
- **that the engine-evolution register is redundant** — the strongest
counter is that ADRs record *decisions* and evidence records *findings*,
but nothing records *what we learned about building engines*, which is a
third thing;
- **that "carries its reproduction" is affordable** — if half the real
findings cannot be reproduced cheaply, the rule will be quietly dropped
CB-WP-0022 T05/T06/T07: the register, and what its first run found T05. tools/design.py, make design, and the register in GroundRules.md -- 14 rows, no new file, because ADR-0012 D2 made §Underdetermined the register rather than building one beside it. Backfill was the test and it caught two things the ADR did not have. First, a `role` column. The first report alarmed on U2 and was wrong to: U2's scenario is green BECAUSE the provisional default it documents is implemented, which says nothing about whether ground-game agrees. GR-E01's was a counterexample that went green. Same colour, opposite meaning -- a register that cannot tell them apart either alarms constantly or never. Only a green counterexample alarms. Folded back into GameDesign §1.3. Second, it contradicted the survey. CB-RES-0007 said "six of the ten already have provisional scenarios." Measured -- grep -lE "\bU<n>\b" over scenarios/ground -- exactly ONE U-item names itself. Five provisional scenarios exist and four probably encode U-item defaults, but the mapping is not written down, so it is not checkable. Same defect class as the wrong premises, found inside the survey that proposed the fix. Now a reported debt: open, lacking a reproduction: 9, target 0. design.py carries the control design-baseline.py never had, asserted directly: a row citing a nonexistent file must not count as reproduced, using the exact path 2da19a4 deleted -- which the old tool called green. design-baseline.py is marked superseded rather than deleted; it is the evidence for how a wrong number got into a survey. T06. The report is a FILE in ground-game under GROUND-WP-0002, committed there, with a hub message that only points at it. It asks for no ruling: it carries GR-E01's withdrawal, our reproduction debt, and two notes that are explicitly not findings. And it had to acknowledge something nobody anticipated. GROUND-WP-0002 is finished -- all ten U-items were RULED 2026-08-03, every one confirmed as the default we simulate, plus five of six provisional scenarios. The survey said "0 of 10 ruled" two days later and this register was built saying `reported`. That is the unread-inbox failure running in the opposite direction: they answered and we did not collect it. The instrument's first run surfaced it. They are `ruled`, not `applied` -- lifting the now-settled provisional flags is owed and is not done, and make design shows them open until it is. T07. evidence/CB-EV-0021. Two of six catches in this pass came from execution rather than process (the role distinction from building it, the ten uncollected rulings from running it), which is InnerLoop §Design goal's prediction holding. make self-tests, facts-check, loop-lint: clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:22:33 +02:00
and the register becomes a suggestion box anyway. *(Hardened before the review ran: two findings had already reached
ground-game on wrong premises, so the reviewer was told to press whether
the rule is **sufficient**, not whether it is affordable.)*
Declare CB-WP-0022: the design instrument, tier L The maintainer named a new aspect: clay-borg as a game design tool, with a register for design flaws, questions, results and trial protocols. Structural L on the maintainer-named-high-leverage trigger, and it amends INTENT. Chaos d8=6, no override. The insight is that this is already happening with no home. Five passes have produced ten underdetermined rules points, SOLVE offered on a face-down problem and always inert, GR-A13's wasted SOLVE, GR-E01 unreachable below 5 seats, six provisional scenario defaults, and two scoring modes never played to the end -- every one found by BUILDING the simulator rather than by playing it. A simulator rigorous enough to refuse ambiguity is a design instrument, because it cannot proceed past a rule that does not decide. All of it has been carried in prose in six places and one sat unread in an inbox for four days. The load-bearing rule: a design finding is not admissible without its reproduction. A register that collects opinions would reproduce this project's standing failure -- unexecuted verification -- in a new medium. Recorded as a judgment for the adversarial review rather than assumed: the engine-evolution meta the maintainer also asked about should NOT be built, because evidence/, decisions/, gates.toml and workplans already carry nineteen passes of it with dates, costs and falsifiers. A second register for the same subject is ceremony. The asymmetry is the point -- engine evolution has a home and game design does not. T05 backfills the six known findings as the TEST of the register: one that cannot express findings the project already has is the wrong register, and discovering that after designing it is why the order is survey, review, decide, specify, build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 20:48:08 +02:00
- **that a register is needed at all**, rather than one more section in
`GroundRules.md §Underdetermined`, which already exists and already
works.
Record the trail in `history/`, unpolished.
CB-WP-0022 T02: the separate reviewer found the showcase finding was false First adversarial review in this repo run by a genuinely separate agent. CB-RES-0006's reviewer opened by conceding it could not be, and called its own findings "a lower bound on what a genuinely separate reviewer would find." This is the measurement: the separate reviewer ran git log against the survey's central example and found 2da19a4 had falsified it four days earlier, while the author -- who wrote that commit -- quoted the dead number twice. Seven challenges: four conceded, two conceded in part, one answered. C1 changes the design. "GR-E01 is admissible because 4/6/9 against 5/7/9 is a computation anyone can rerun" was a computation already rerun: the edition import measured 6/9/12, the conclusion inverted, and the scenario was renamed -unreachable- to -reachable-. That finding was one of the TWO that passed the reproduction rule. So three wrong premises have now reached ground-game and the third satisfied an existence test -- existence is not the property that was missing. The rule gains shape (ground-game's row-level deal table, promoted from a T04 addendum) and a clause the survey never contemplated: a reproduction must be able to fail. Ours went green and stayed admissible. C1 also caught a defect in flight. T06's payload, status todo, still named 4/6/9 and was queued to send it to ground-game as "no dataset reconciles them." Withdrawn before sending -- the fourth wrong premise, and the only one stopped. C2 withdraws the baseline's precision: design-baseline.py is a hand-maintained dict counting itself, has_reproduction never checks the file exists (its YES-control is green against a deleted path), and Makefile:127 runs only --self-test so the reporting path has no CI. The direction survives; 33% is not a measured rate and T05 must not build on it. C3: "six provisional defaults" is five, GR-E01 double-counted. C4: GroundRules §Underdetermined was never evaluated as a candidate and already delivers four of five benchmarks -- T03's burden flips to arguing extension over replacement. Survived: the rule's affordability, and reuse of the provisional machinery. loop-lint: no findings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:03:12 +02:00
**Done 2026-08-05.** Trail:
[challenge](../history/260805-design-instrument-challenge.md),
[response](../history/260805-design-instrument-response.md).
**Run by a separate agent** — the first in this repo that was. CB-RES-0006's
review opened by conceding it could not be, and called its own findings
*"a lower bound on what a genuinely separate reviewer would find."* That
was measurable, and this is the measurement: the separate reviewer ran
`git log` against the survey's central example and found our own commit
had falsified it four days earlier, while the author — who wrote that
commit — quoted the dead number twice.
**Seven challenges: four conceded, two conceded in part, one answered.**
CB-WP-0022 T05/T06/T07: the register, and what its first run found T05. tools/design.py, make design, and the register in GroundRules.md -- 14 rows, no new file, because ADR-0012 D2 made §Underdetermined the register rather than building one beside it. Backfill was the test and it caught two things the ADR did not have. First, a `role` column. The first report alarmed on U2 and was wrong to: U2's scenario is green BECAUSE the provisional default it documents is implemented, which says nothing about whether ground-game agrees. GR-E01's was a counterexample that went green. Same colour, opposite meaning -- a register that cannot tell them apart either alarms constantly or never. Only a green counterexample alarms. Folded back into GameDesign §1.3. Second, it contradicted the survey. CB-RES-0007 said "six of the ten already have provisional scenarios." Measured -- grep -lE "\bU<n>\b" over scenarios/ground -- exactly ONE U-item names itself. Five provisional scenarios exist and four probably encode U-item defaults, but the mapping is not written down, so it is not checkable. Same defect class as the wrong premises, found inside the survey that proposed the fix. Now a reported debt: open, lacking a reproduction: 9, target 0. design.py carries the control design-baseline.py never had, asserted directly: a row citing a nonexistent file must not count as reproduced, using the exact path 2da19a4 deleted -- which the old tool called green. design-baseline.py is marked superseded rather than deleted; it is the evidence for how a wrong number got into a survey. T06. The report is a FILE in ground-game under GROUND-WP-0002, committed there, with a hub message that only points at it. It asks for no ruling: it carries GR-E01's withdrawal, our reproduction debt, and two notes that are explicitly not findings. And it had to acknowledge something nobody anticipated. GROUND-WP-0002 is finished -- all ten U-items were RULED 2026-08-03, every one confirmed as the default we simulate, plus five of six provisional scenarios. The survey said "0 of 10 ruled" two days later and this register was built saying `reported`. That is the unread-inbox failure running in the opposite direction: they answered and we did not collect it. The instrument's first run surfaced it. They are `ruled`, not `applied` -- lifting the now-settled provisional flags is owed and is not done, and make design shows them open until it is. T07. evidence/CB-EV-0021. Two of six catches in this pass came from execution rather than process (the role distinction from building it, the ten uncollected rulings from running it), which is InnerLoop §Design goal's prediction holding. make self-tests, facts-check, loop-lint: clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:22:33 +02:00
**C1 changed the design** — the rule's showcase finding was false and had
*passed* the rule, so existence is not what was missing — and **caught a
defect in flight**, T06's payload still naming the dead number. C2
withdrew the baseline's precision, C3 its arithmetic, C4 flipped T03's
burden toward extending `§Underdetermined`, C5 corrected the redundancy
inventory. Survived: affordability, and reuse of the provisional
machinery. Full account:
[CB-EV-0021](../evidence/CB-EV-0021-the-design-instrument.md) §1§3.
CB-WP-0022 T02: the separate reviewer found the showcase finding was false First adversarial review in this repo run by a genuinely separate agent. CB-RES-0006's reviewer opened by conceding it could not be, and called its own findings "a lower bound on what a genuinely separate reviewer would find." This is the measurement: the separate reviewer ran git log against the survey's central example and found 2da19a4 had falsified it four days earlier, while the author -- who wrote that commit -- quoted the dead number twice. Seven challenges: four conceded, two conceded in part, one answered. C1 changes the design. "GR-E01 is admissible because 4/6/9 against 5/7/9 is a computation anyone can rerun" was a computation already rerun: the edition import measured 6/9/12, the conclusion inverted, and the scenario was renamed -unreachable- to -reachable-. That finding was one of the TWO that passed the reproduction rule. So three wrong premises have now reached ground-game and the third satisfied an existence test -- existence is not the property that was missing. The rule gains shape (ground-game's row-level deal table, promoted from a T04 addendum) and a clause the survey never contemplated: a reproduction must be able to fail. Ours went green and stayed admissible. C1 also caught a defect in flight. T06's payload, status todo, still named 4/6/9 and was queued to send it to ground-game as "no dataset reconciles them." Withdrawn before sending -- the fourth wrong premise, and the only one stopped. C2 withdraws the baseline's precision: design-baseline.py is a hand-maintained dict counting itself, has_reproduction never checks the file exists (its YES-control is green against a deleted path), and Makefile:127 runs only --self-test so the reporting path has no CI. The direction survives; 33% is not a measured rate and T05 must not build on it. C3: "six provisional defaults" is five, GR-E01 double-counted. C4: GroundRules §Underdetermined was never evaluated as a candidate and already delivers four of five benchmarks -- T03's burden flips to arguing extension over replacement. Survived: the rule's affordability, and reuse of the provisional machinery. loop-lint: no findings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:03:12 +02:00
Declare CB-WP-0022: the design instrument, tier L The maintainer named a new aspect: clay-borg as a game design tool, with a register for design flaws, questions, results and trial protocols. Structural L on the maintainer-named-high-leverage trigger, and it amends INTENT. Chaos d8=6, no override. The insight is that this is already happening with no home. Five passes have produced ten underdetermined rules points, SOLVE offered on a face-down problem and always inert, GR-A13's wasted SOLVE, GR-E01 unreachable below 5 seats, six provisional scenario defaults, and two scoring modes never played to the end -- every one found by BUILDING the simulator rather than by playing it. A simulator rigorous enough to refuse ambiguity is a design instrument, because it cannot proceed past a rule that does not decide. All of it has been carried in prose in six places and one sat unread in an inbox for four days. The load-bearing rule: a design finding is not admissible without its reproduction. A register that collects opinions would reproduce this project's standing failure -- unexecuted verification -- in a new medium. Recorded as a judgment for the adversarial review rather than assumed: the engine-evolution meta the maintainer also asked about should NOT be built, because evidence/, decisions/, gates.toml and workplans already carry nineteen passes of it with dates, costs and falsifiers. A second register for the same subject is ceremony. The asymmetry is the point -- engine evolution has a home and game design does not. T05 backfills the six known findings as the TEST of the register: one that cannot express findings the project already has is the wrong register, and discovering that after designing it is why the order is survey, review, decide, specify, build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 20:48:08 +02:00
## Task: decide
```task
id: CB-WP-0022-T03
CB-WP-0022 T03: ADR-0012 -- the register already existed, and the rule was one clause short Nine decisions. Two were not on T03's list; both are the review's. D2: specs/GroundRules.md §Underdetermined IS the register. C4 pointed out it was never evaluated as a candidate, and against the survey's own five benchmarks it already delivers four -- including the Magic Oracle property ("a ruling flips the scenario, not the kernel", :231-233) that the survey travelled to Magic to discover and we had written down ourselves eight days earlier. What it lacks is reproductions. So this pass extends a section rather than building a register: no new file, no new schema, and no second mechanism to disagree with the first. D3: admissibility is three clauses. It exists; it has the ruled shape (GROUND-WP-0004 T02's row-level table, never a sum -- promoted from a T04 addendum because two of three wrong premises were sums without tables); and it CAN FAIL. The third is C1's. GR-E01's scenario went green when the edition landed, and the finding stayed admissible and stayed queued for transmission, because nothing in the rule said a passing artifact was a signal. A green reproduction is an alarm, not a reassurance. D1 applied: INTENT gains a fourth property, Instrument, worded as a mechanism rather than an ambition and carrying its own falsifier -- if a pass tolerates an undecided rule by quietly picking a default, the property is false. D4 five kinds, each forced by an existing finding; a sixth during backfill means the taxonomy was invented. D5 lifecycle where `applied` means the source changed, the queue empties while the log accumulates, and withdrawals are reported rather than deleted -- GR-E01 is why. D6 notes admitted but never reportable, 30-day expiry on the existing age machinery; refusing them would discard the only class of finding the engine cannot produce itself, which is CB-WP-0025's whole input. D7 no engine-evolution register, on an inventory C5 corrected -- narrowed, not settled. D8 design-baseline.py retired, kept as a dated snapshot because deleting it erases the evidence for how 33% got in. D9 the artifact stays here, ground-game gets a generated file under its own workplan. loop-lint: no findings. facts-check: no findings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:05:51 +02:00
status: done
Declare CB-WP-0022: the design instrument, tier L The maintainer named a new aspect: clay-borg as a game design tool, with a register for design flaws, questions, results and trial protocols. Structural L on the maintainer-named-high-leverage trigger, and it amends INTENT. Chaos d8=6, no override. The insight is that this is already happening with no home. Five passes have produced ten underdetermined rules points, SOLVE offered on a face-down problem and always inert, GR-A13's wasted SOLVE, GR-E01 unreachable below 5 seats, six provisional scenario defaults, and two scoring modes never played to the end -- every one found by BUILDING the simulator rather than by playing it. A simulator rigorous enough to refuse ambiguity is a design instrument, because it cannot proceed past a rule that does not decide. All of it has been carried in prose in six places and one sat unread in an inbox for four days. The load-bearing rule: a design finding is not admissible without its reproduction. A register that collects opinions would reproduce this project's standing failure -- unexecuted verification -- in a new medium. Recorded as a judgment for the adversarial review rather than assumed: the engine-evolution meta the maintainer also asked about should NOT be built, because evidence/, decisions/, gates.toml and workplans already carry nineteen passes of it with dates, costs and falsifiers. A second register for the same subject is ceremony. The asymmetry is the point -- engine evolution has a home and game design does not. T05 backfills the six known findings as the TEST of the register: one that cannot express findings the project already has is the wrong register, and discovering that after designing it is why the order is survey, review, decide, specify, build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 20:48:08 +02:00
priority: high
2026-08-04 00:21:21 +02:00
state_hub_task_id: "a2e85810-e949-4ea0-81fc-36b912af326c"
Declare CB-WP-0022: the design instrument, tier L The maintainer named a new aspect: clay-borg as a game design tool, with a register for design flaws, questions, results and trial protocols. Structural L on the maintainer-named-high-leverage trigger, and it amends INTENT. Chaos d8=6, no override. The insight is that this is already happening with no home. Five passes have produced ten underdetermined rules points, SOLVE offered on a face-down problem and always inert, GR-A13's wasted SOLVE, GR-E01 unreachable below 5 seats, six provisional scenario defaults, and two scoring modes never played to the end -- every one found by BUILDING the simulator rather than by playing it. A simulator rigorous enough to refuse ambiguity is a design instrument, because it cannot proceed past a rule that does not decide. All of it has been carried in prose in six places and one sat unread in an inbox for four days. The load-bearing rule: a design finding is not admissible without its reproduction. A register that collects opinions would reproduce this project's standing failure -- unexecuted verification -- in a new medium. Recorded as a judgment for the adversarial review rather than assumed: the engine-evolution meta the maintainer also asked about should NOT be built, because evidence/, decisions/, gates.toml and workplans already carry nineteen passes of it with dates, costs and falsifiers. A second register for the same subject is ceremony. The asymmetry is the point -- engine evolution has a home and game design does not. T05 backfills the six known findings as the TEST of the register: one that cannot express findings the project already has is the wrong register, and discovering that after designing it is why the order is survey, review, decide, specify, build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 20:48:08 +02:00
```
`decisions/ADR-0012-*.md`. At minimum:
- whether clay-borg's **INTENT gains a stated aspect** as a design
instrument, and in what words — this is the change with the longest
half-life in the pass;
- the finding **taxonomy**, and it should be grounded in the six findings
above rather than invented: *underdetermined* (rules do not say),
*inconsistent* (rules disagree with each other or the data), *inert* (a
rule that cannot fire), *degenerate* (fires, but collapses play),
*unplayed* (implemented, never played);
- the **lifecycle** and who owns each state: raised → reported → ruled →
applied, or withdrawn;
- whether a finding without a reproduction is **rejected** or **admitted
as a note** — and if admitted, how it is prevented from aging into an
apparent finding.
CB-WP-0022 T03: ADR-0012 -- the register already existed, and the rule was one clause short Nine decisions. Two were not on T03's list; both are the review's. D2: specs/GroundRules.md §Underdetermined IS the register. C4 pointed out it was never evaluated as a candidate, and against the survey's own five benchmarks it already delivers four -- including the Magic Oracle property ("a ruling flips the scenario, not the kernel", :231-233) that the survey travelled to Magic to discover and we had written down ourselves eight days earlier. What it lacks is reproductions. So this pass extends a section rather than building a register: no new file, no new schema, and no second mechanism to disagree with the first. D3: admissibility is three clauses. It exists; it has the ruled shape (GROUND-WP-0004 T02's row-level table, never a sum -- promoted from a T04 addendum because two of three wrong premises were sums without tables); and it CAN FAIL. The third is C1's. GR-E01's scenario went green when the edition landed, and the finding stayed admissible and stayed queued for transmission, because nothing in the rule said a passing artifact was a signal. A green reproduction is an alarm, not a reassurance. D1 applied: INTENT gains a fourth property, Instrument, worded as a mechanism rather than an ambition and carrying its own falsifier -- if a pass tolerates an undecided rule by quietly picking a default, the property is false. D4 five kinds, each forced by an existing finding; a sixth during backfill means the taxonomy was invented. D5 lifecycle where `applied` means the source changed, the queue empties while the log accumulates, and withdrawals are reported rather than deleted -- GR-E01 is why. D6 notes admitted but never reportable, 30-day expiry on the existing age machinery; refusing them would discard the only class of finding the engine cannot produce itself, which is CB-WP-0025's whole input. D7 no engine-evolution register, on an inventory C5 corrected -- narrowed, not settled. D8 design-baseline.py retired, kept as a dated snapshot because deleting it erases the evidence for how 33% got in. D9 the artifact stays here, ground-game gets a generated file under its own workplan. loop-lint: no findings. facts-check: no findings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:05:51 +02:00
**Done 2026-08-05.**
[ADR-0012](../decisions/ADR-0012-the-design-instrument.md), nine
decisions. The two not on this list are the two the review forced:
CB-WP-0022 T03: ADR-0012 -- the register already existed, and the rule was one clause short Nine decisions. Two were not on T03's list; both are the review's. D2: specs/GroundRules.md §Underdetermined IS the register. C4 pointed out it was never evaluated as a candidate, and against the survey's own five benchmarks it already delivers four -- including the Magic Oracle property ("a ruling flips the scenario, not the kernel", :231-233) that the survey travelled to Magic to discover and we had written down ourselves eight days earlier. What it lacks is reproductions. So this pass extends a section rather than building a register: no new file, no new schema, and no second mechanism to disagree with the first. D3: admissibility is three clauses. It exists; it has the ruled shape (GROUND-WP-0004 T02's row-level table, never a sum -- promoted from a T04 addendum because two of three wrong premises were sums without tables); and it CAN FAIL. The third is C1's. GR-E01's scenario went green when the edition landed, and the finding stayed admissible and stayed queued for transmission, because nothing in the rule said a passing artifact was a signal. A green reproduction is an alarm, not a reassurance. D1 applied: INTENT gains a fourth property, Instrument, worded as a mechanism rather than an ambition and carrying its own falsifier -- if a pass tolerates an undecided rule by quietly picking a default, the property is false. D4 five kinds, each forced by an existing finding; a sixth during backfill means the taxonomy was invented. D5 lifecycle where `applied` means the source changed, the queue empties while the log accumulates, and withdrawals are reported rather than deleted -- GR-E01 is why. D6 notes admitted but never reportable, 30-day expiry on the existing age machinery; refusing them would discard the only class of finding the engine cannot produce itself, which is CB-WP-0025's whole input. D7 no engine-evolution register, on an inventory C5 corrected -- narrowed, not settled. D8 design-baseline.py retired, kept as a dated snapshot because deleting it erases the evidence for how 33% got in. D9 the artifact stays here, ground-game gets a generated file under its own workplan. loop-lint: no findings. facts-check: no findings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:05:51 +02:00
- **D2 — `§Underdetermined` *is* the register; nothing parallel is built.**
Against the survey's own five benchmarks the incumbent already delivers
four, including the Oracle property the survey went to Magic to find and
we had written ourselves eight days earlier (`GroundRules.md:231-233`).
What it lacks is reproductions. So this pass **extends** a section — no
new file, no new schema.
- **D3 — admissibility is three clauses.** Exists, has the ruled shape
(row-level table, never a sum), **and can fail.** GR-E01's artifact went
green and the finding stayed admissible and stayed queued, because
nothing said a passing artifact was a signal. **A green reproduction is
an alarm.**
The rest, in one line each: **D1** INTENT gains property 4, *Instrument*,
applied with its falsifier. **D4** five kinds, each forced by an existing
finding. **D5** `applied` means the source changed; withdrawals are
reported, not deleted. **D6** notes admitted but never reportable, 30-day
expiry. **D7** no engine-evolution register, on an inventory C5 corrected.
**D8** `design-baseline.py` retired. **D9** the artifact stays here,
ground-game gets a generated file under its own workplan.
CB-WP-0022 T03: ADR-0012 -- the register already existed, and the rule was one clause short Nine decisions. Two were not on T03's list; both are the review's. D2: specs/GroundRules.md §Underdetermined IS the register. C4 pointed out it was never evaluated as a candidate, and against the survey's own five benchmarks it already delivers four -- including the Magic Oracle property ("a ruling flips the scenario, not the kernel", :231-233) that the survey travelled to Magic to discover and we had written down ourselves eight days earlier. What it lacks is reproductions. So this pass extends a section rather than building a register: no new file, no new schema, and no second mechanism to disagree with the first. D3: admissibility is three clauses. It exists; it has the ruled shape (GROUND-WP-0004 T02's row-level table, never a sum -- promoted from a T04 addendum because two of three wrong premises were sums without tables); and it CAN FAIL. The third is C1's. GR-E01's scenario went green when the edition landed, and the finding stayed admissible and stayed queued for transmission, because nothing in the rule said a passing artifact was a signal. A green reproduction is an alarm, not a reassurance. D1 applied: INTENT gains a fourth property, Instrument, worded as a mechanism rather than an ambition and carrying its own falsifier -- if a pass tolerates an undecided rule by quietly picking a default, the property is false. D4 five kinds, each forced by an existing finding; a sixth during backfill means the taxonomy was invented. D5 lifecycle where `applied` means the source changed, the queue empties while the log accumulates, and withdrawals are reported rather than deleted -- GR-E01 is why. D6 notes admitted but never reportable, 30-day expiry on the existing age machinery; refusing them would discard the only class of finding the engine cannot produce itself, which is CB-WP-0025's whole input. D7 no engine-evolution register, on an inventory C5 corrected -- narrowed, not settled. D8 design-baseline.py retired, kept as a dated snapshot because deleting it erases the evidence for how 33% got in. D9 the artifact stays here, ground-game gets a generated file under its own workplan. loop-lint: no findings. facts-check: no findings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:05:51 +02:00
Declare CB-WP-0022: the design instrument, tier L The maintainer named a new aspect: clay-borg as a game design tool, with a register for design flaws, questions, results and trial protocols. Structural L on the maintainer-named-high-leverage trigger, and it amends INTENT. Chaos d8=6, no override. The insight is that this is already happening with no home. Five passes have produced ten underdetermined rules points, SOLVE offered on a face-down problem and always inert, GR-A13's wasted SOLVE, GR-E01 unreachable below 5 seats, six provisional scenario defaults, and two scoring modes never played to the end -- every one found by BUILDING the simulator rather than by playing it. A simulator rigorous enough to refuse ambiguity is a design instrument, because it cannot proceed past a rule that does not decide. All of it has been carried in prose in six places and one sat unread in an inbox for four days. The load-bearing rule: a design finding is not admissible without its reproduction. A register that collects opinions would reproduce this project's standing failure -- unexecuted verification -- in a new medium. Recorded as a judgment for the adversarial review rather than assumed: the engine-evolution meta the maintainer also asked about should NOT be built, because evidence/, decisions/, gates.toml and workplans already carry nineteen passes of it with dates, costs and falsifiers. A second register for the same subject is ceremony. The asymmetry is the point -- engine evolution has a home and game design does not. T05 backfills the six known findings as the TEST of the register: one that cannot express findings the project already has is the wrong register, and discovering that after designing it is why the order is survey, review, decide, specify, build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 20:48:08 +02:00
## Task: specify
```task
id: CB-WP-0022-T04
CB-WP-0022 T04: specs/GameDesign.md -- what a reproduction must show Not a register; ADR-0012 D2 put that in GroundRules §Underdetermined. This spec says what may go in it, what a reproduction must show, how a finding dies, and how a trial game is run. §1.2 is written against evidence rather than principle. A finding must print the rows behind any number it claims, and the spec carries the table of what shipped instead: a sum ("12 in the file"), a green scenario ("4/6/9 against 5/7/9"), and a condition named without checking which one fired ("SOLVE on a face-down Problem"). "12" was arithmetically defensible and still wrong about the game -- that sentence is the requirement. §1.3's target is 0 reproductions that have gone green while open. GR-E01 would have tripped it four days before a human caught it by hand. No baseline rate is quoted. The 33% was withdrawn by C2 and the first honest denominator is T05's backfill; quoting a new number from a discredited instrument is how the first one got in. The trial protocol costs one flag: cb-play --record already writes a finished game as a scenario, so a trial is that plus a sibling .md in the player's own words. An observation is a NOTE until it has a reproduction, and notes may not cross the repo boundary and expire at 30 days on the existing provisional-age machinery. The maintainer's "I felt it was too easy but then we lost" is the case the protocol is shaped around -- forcing it into a schema at the moment of observation would lose it. loop-lint: no findings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:07:38 +02:00
status: done
Declare CB-WP-0022: the design instrument, tier L The maintainer named a new aspect: clay-borg as a game design tool, with a register for design flaws, questions, results and trial protocols. Structural L on the maintainer-named-high-leverage trigger, and it amends INTENT. Chaos d8=6, no override. The insight is that this is already happening with no home. Five passes have produced ten underdetermined rules points, SOLVE offered on a face-down problem and always inert, GR-A13's wasted SOLVE, GR-E01 unreachable below 5 seats, six provisional scenario defaults, and two scoring modes never played to the end -- every one found by BUILDING the simulator rather than by playing it. A simulator rigorous enough to refuse ambiguity is a design instrument, because it cannot proceed past a rule that does not decide. All of it has been carried in prose in six places and one sat unread in an inbox for four days. The load-bearing rule: a design finding is not admissible without its reproduction. A register that collects opinions would reproduce this project's standing failure -- unexecuted verification -- in a new medium. Recorded as a judgment for the adversarial review rather than assumed: the engine-evolution meta the maintainer also asked about should NOT be built, because evidence/, decisions/, gates.toml and workplans already carry nineteen passes of it with dates, costs and falsifiers. A second register for the same subject is ceremony. The asymmetry is the point -- engine evolution has a home and game design does not. T05 backfills the six known findings as the TEST of the register: one that cannot express findings the project already has is the wrong register, and discovering that after designing it is why the order is survey, review, decide, specify, build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 20:48:08 +02:00
priority: high
2026-08-04 00:21:21 +02:00
state_hub_task_id: "60ddfeab-81b1-45e4-9ec9-91cc8ea7fb72"
Declare CB-WP-0022: the design instrument, tier L The maintainer named a new aspect: clay-borg as a game design tool, with a register for design flaws, questions, results and trial protocols. Structural L on the maintainer-named-high-leverage trigger, and it amends INTENT. Chaos d8=6, no override. The insight is that this is already happening with no home. Five passes have produced ten underdetermined rules points, SOLVE offered on a face-down problem and always inert, GR-A13's wasted SOLVE, GR-E01 unreachable below 5 seats, six provisional scenario defaults, and two scoring modes never played to the end -- every one found by BUILDING the simulator rather than by playing it. A simulator rigorous enough to refuse ambiguity is a design instrument, because it cannot proceed past a rule that does not decide. All of it has been carried in prose in six places and one sat unread in an inbox for four days. The load-bearing rule: a design finding is not admissible without its reproduction. A register that collects opinions would reproduce this project's standing failure -- unexecuted verification -- in a new medium. Recorded as a judgment for the adversarial review rather than assumed: the engine-evolution meta the maintainer also asked about should NOT be built, because evidence/, decisions/, gates.toml and workplans already carry nineteen passes of it with dates, costs and falsifiers. A second register for the same subject is ceremony. The asymmetry is the point -- engine evolution has a home and game design does not. T05 backfills the six known findings as the TEST of the register: one that cannot express findings the project already has is the wrong register, and discovering that after designing it is why the order is survey, review, decide, specify, build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 20:48:08 +02:00
```
`specs/GameDesign.md`, with metrics, because a spec without them is prose.
Candidate measures, to be argued not adopted:
- **findings with a runnable reproduction** — target 100%, and the
denominator includes withdrawn ones;
- **time from raised to reported** — the U-items took four days to be
*read*; that is the number this exists to fix;
- **findings closed by a ruling** vs **findings still open**, with age.
Adapt CB-WP-0021 and CB-WP-0022 to ground-game's rulings GR-S01 was ruled 2026-08-04: Surface always, union hidden priorities 1..k, k = 2/3/4. So the deal is 3/4/5 Problems, not 2/3/4; available points 6/9/12; thresholds 5/7/9 stand; SHARED GROUND is 2-6p as printed. The question CB-WP-0021 was going to send has been answered, and it was the deal. CB-WP-0021 is re-scoped. The import is now LOAD-BEARING rather than merely correct: the ruled 6/9/12 holds only with Problems.csv values, and the same deal with the stand-in gives 6/10/15 -- a different game that happens to also be winnable. Fixing the deal without importing the data would produce numbers nobody ruled on, so T05 (deal) and T02 (import) must land together. gd0001 is to be INVERTED, not deleted: it is the record of why this changed. gr-e01 is rewritten as a non-provisional import check, per the ruling's own wording, and loses its provisional owner because ground-game has now ruled. CB-WP-0022 absorbs ground-game's process ruling, which is stricter than this pass proposed: arithmetic findings need a runnable reproduction AND a row-level deal table listing Surface and each hidden priority separately, never only a sum or a deal depth. That is a direct consequence of both premises we got wrong. So the reproduction rule gains a SHAPE requirement, not just an existence one -- a finding that ships a passing test but describes the wrong quantity is still a bad finding, and that is what happened twice. T02's review brief is flipped accordingly: press whether the rule is SUFFICIENT, not whether it is affordable. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 00:29:45 +02:00
**ground-game has ruled on what a finding must carry** (GROUND-WP-0004
T02, 2026-08-04), and it is stricter than this pass proposed. Adopt it:
CB-WP-0022 T05/T06/T07: the register, and what its first run found T05. tools/design.py, make design, and the register in GroundRules.md -- 14 rows, no new file, because ADR-0012 D2 made §Underdetermined the register rather than building one beside it. Backfill was the test and it caught two things the ADR did not have. First, a `role` column. The first report alarmed on U2 and was wrong to: U2's scenario is green BECAUSE the provisional default it documents is implemented, which says nothing about whether ground-game agrees. GR-E01's was a counterexample that went green. Same colour, opposite meaning -- a register that cannot tell them apart either alarms constantly or never. Only a green counterexample alarms. Folded back into GameDesign §1.3. Second, it contradicted the survey. CB-RES-0007 said "six of the ten already have provisional scenarios." Measured -- grep -lE "\bU<n>\b" over scenarios/ground -- exactly ONE U-item names itself. Five provisional scenarios exist and four probably encode U-item defaults, but the mapping is not written down, so it is not checkable. Same defect class as the wrong premises, found inside the survey that proposed the fix. Now a reported debt: open, lacking a reproduction: 9, target 0. design.py carries the control design-baseline.py never had, asserted directly: a row citing a nonexistent file must not count as reproduced, using the exact path 2da19a4 deleted -- which the old tool called green. design-baseline.py is marked superseded rather than deleted; it is the evidence for how a wrong number got into a survey. T06. The report is a FILE in ground-game under GROUND-WP-0002, committed there, with a hub message that only points at it. It asks for no ruling: it carries GR-E01's withdrawal, our reproduction debt, and two notes that are explicitly not findings. And it had to acknowledge something nobody anticipated. GROUND-WP-0002 is finished -- all ten U-items were RULED 2026-08-03, every one confirmed as the default we simulate, plus five of six provisional scenarios. The survey said "0 of 10 ruled" two days later and this register was built saying `reported`. That is the unread-inbox failure running in the opposite direction: they answered and we did not collect it. The instrument's first run surfaced it. They are `ruled`, not `applied` -- lifting the now-settled provisional flags is owed and is not done, and make design shows them open until it is. T07. evidence/CB-EV-0021. Two of six catches in this pass came from execution rather than process (the role distinction from building it, the ten uncollected rulings from running it), which is InnerLoop §Design goal's prediction holding. make self-tests, facts-check, loop-lint: clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:22:33 +02:00
> Arithmetic findings ship a **runnable reproduction** *and* a
> **row-level deal table** — never only "sum of file" or "deal depth N";
> and ground-game's arithmetic rulings cite that reproduction by path.
Adapt CB-WP-0021 and CB-WP-0022 to ground-game's rulings GR-S01 was ruled 2026-08-04: Surface always, union hidden priorities 1..k, k = 2/3/4. So the deal is 3/4/5 Problems, not 2/3/4; available points 6/9/12; thresholds 5/7/9 stand; SHARED GROUND is 2-6p as printed. The question CB-WP-0021 was going to send has been answered, and it was the deal. CB-WP-0021 is re-scoped. The import is now LOAD-BEARING rather than merely correct: the ruled 6/9/12 holds only with Problems.csv values, and the same deal with the stand-in gives 6/10/15 -- a different game that happens to also be winnable. Fixing the deal without importing the data would produce numbers nobody ruled on, so T05 (deal) and T02 (import) must land together. gd0001 is to be INVERTED, not deleted: it is the record of why this changed. gr-e01 is rewritten as a non-provisional import check, per the ruling's own wording, and loses its provisional owner because ground-game has now ruled. CB-WP-0022 absorbs ground-game's process ruling, which is stricter than this pass proposed: arithmetic findings need a runnable reproduction AND a row-level deal table listing Surface and each hidden priority separately, never only a sum or a deal depth. That is a direct consequence of both premises we got wrong. So the reproduction rule gains a SHAPE requirement, not just an existence one -- a finding that ships a passing test but describes the wrong quantity is still a bad finding, and that is what happened twice. T02's review brief is flipped accordingly: press whether the rule is SUFFICIENT, not whether it is affordable. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 00:29:45 +02:00
CB-WP-0022 T05/T06/T07: the register, and what its first run found T05. tools/design.py, make design, and the register in GroundRules.md -- 14 rows, no new file, because ADR-0012 D2 made §Underdetermined the register rather than building one beside it. Backfill was the test and it caught two things the ADR did not have. First, a `role` column. The first report alarmed on U2 and was wrong to: U2's scenario is green BECAUSE the provisional default it documents is implemented, which says nothing about whether ground-game agrees. GR-E01's was a counterexample that went green. Same colour, opposite meaning -- a register that cannot tell them apart either alarms constantly or never. Only a green counterexample alarms. Folded back into GameDesign §1.3. Second, it contradicted the survey. CB-RES-0007 said "six of the ten already have provisional scenarios." Measured -- grep -lE "\bU<n>\b" over scenarios/ground -- exactly ONE U-item names itself. Five provisional scenarios exist and four probably encode U-item defaults, but the mapping is not written down, so it is not checkable. Same defect class as the wrong premises, found inside the survey that proposed the fix. Now a reported debt: open, lacking a reproduction: 9, target 0. design.py carries the control design-baseline.py never had, asserted directly: a row citing a nonexistent file must not count as reproduced, using the exact path 2da19a4 deleted -- which the old tool called green. design-baseline.py is marked superseded rather than deleted; it is the evidence for how a wrong number got into a survey. T06. The report is a FILE in ground-game under GROUND-WP-0002, committed there, with a hub message that only points at it. It asks for no ruling: it carries GR-E01's withdrawal, our reproduction debt, and two notes that are explicitly not findings. And it had to acknowledge something nobody anticipated. GROUND-WP-0002 is finished -- all ten U-items were RULED 2026-08-03, every one confirmed as the default we simulate, plus five of six provisional scenarios. The survey said "0 of 10 ruled" two days later and this register was built saying `reported`. That is the unread-inbox failure running in the opposite direction: they answered and we did not collect it. The instrument's first run surfaced it. They are `ruled`, not `applied` -- lifting the now-settled provisional flags is owed and is not done, and make design shows them open until it is. T07. evidence/CB-EV-0021. Two of six catches in this pass came from execution rather than process (the role distinction from building it, the ten uncollected rulings from running it), which is InnerLoop §Design goal's prediction holding. make self-tests, facts-check, loop-lint: clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:22:33 +02:00
The second half is theirs to keep. **So the reproduction rule gains a
shape requirement, not just an existence one** — a finding that ships a
passing test but describes the wrong quantity is still a bad finding,
which is exactly what happened twice.
Adapt CB-WP-0021 and CB-WP-0022 to ground-game's rulings GR-S01 was ruled 2026-08-04: Surface always, union hidden priorities 1..k, k = 2/3/4. So the deal is 3/4/5 Problems, not 2/3/4; available points 6/9/12; thresholds 5/7/9 stand; SHARED GROUND is 2-6p as printed. The question CB-WP-0021 was going to send has been answered, and it was the deal. CB-WP-0021 is re-scoped. The import is now LOAD-BEARING rather than merely correct: the ruled 6/9/12 holds only with Problems.csv values, and the same deal with the stand-in gives 6/10/15 -- a different game that happens to also be winnable. Fixing the deal without importing the data would produce numbers nobody ruled on, so T05 (deal) and T02 (import) must land together. gd0001 is to be INVERTED, not deleted: it is the record of why this changed. gr-e01 is rewritten as a non-provisional import check, per the ruling's own wording, and loses its provisional owner because ground-game has now ruled. CB-WP-0022 absorbs ground-game's process ruling, which is stricter than this pass proposed: arithmetic findings need a runnable reproduction AND a row-level deal table listing Surface and each hidden priority separately, never only a sum or a deal depth. That is a direct consequence of both premises we got wrong. So the reproduction rule gains a SHAPE requirement, not just an existence one -- a finding that ships a passing test but describes the wrong quantity is still a bad finding, and that is what happened twice. T02's review brief is flipped accordingly: press whether the rule is SUFFICIENT, not whether it is affordable. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 00:29:45 +02:00
CB-WP-0022 T05/T06/T07: the register, and what its first run found T05. tools/design.py, make design, and the register in GroundRules.md -- 14 rows, no new file, because ADR-0012 D2 made §Underdetermined the register rather than building one beside it. Backfill was the test and it caught two things the ADR did not have. First, a `role` column. The first report alarmed on U2 and was wrong to: U2's scenario is green BECAUSE the provisional default it documents is implemented, which says nothing about whether ground-game agrees. GR-E01's was a counterexample that went green. Same colour, opposite meaning -- a register that cannot tell them apart either alarms constantly or never. Only a green counterexample alarms. Folded back into GameDesign §1.3. Second, it contradicted the survey. CB-RES-0007 said "six of the ten already have provisional scenarios." Measured -- grep -lE "\bU<n>\b" over scenarios/ground -- exactly ONE U-item names itself. Five provisional scenarios exist and four probably encode U-item defaults, but the mapping is not written down, so it is not checkable. Same defect class as the wrong premises, found inside the survey that proposed the fix. Now a reported debt: open, lacking a reproduction: 9, target 0. design.py carries the control design-baseline.py never had, asserted directly: a row citing a nonexistent file must not count as reproduced, using the exact path 2da19a4 deleted -- which the old tool called green. design-baseline.py is marked superseded rather than deleted; it is the evidence for how a wrong number got into a survey. T06. The report is a FILE in ground-game under GROUND-WP-0002, committed there, with a hub message that only points at it. It asks for no ruling: it carries GR-E01's withdrawal, our reproduction debt, and two notes that are explicitly not findings. And it had to acknowledge something nobody anticipated. GROUND-WP-0002 is finished -- all ten U-items were RULED 2026-08-03, every one confirmed as the default we simulate, plus five of six provisional scenarios. The survey said "0 of 10 ruled" two days later and this register was built saying `reported`. That is the unread-inbox failure running in the opposite direction: they answered and we did not collect it. The instrument's first run surfaced it. They are `ruled`, not `applied` -- lifting the now-settled provisional flags is owed and is not done, and make design shows them open until it is. T07. evidence/CB-EV-0021. Two of six catches in this pass came from execution rather than process (the role distinction from building it, the ten uncollected rulings from running it), which is InnerLoop §Design goal's prediction holding. make self-tests, facts-check, loop-lint: clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:22:33 +02:00
Also specify the **trial protocol**: a trial game is a `--record`ed
session plus an observation log, so *"we played it and X happened"* is
replayable rather than remembered. It must cost almost nothing or it will
not be done.
Declare CB-WP-0022: the design instrument, tier L The maintainer named a new aspect: clay-borg as a game design tool, with a register for design flaws, questions, results and trial protocols. Structural L on the maintainer-named-high-leverage trigger, and it amends INTENT. Chaos d8=6, no override. The insight is that this is already happening with no home. Five passes have produced ten underdetermined rules points, SOLVE offered on a face-down problem and always inert, GR-A13's wasted SOLVE, GR-E01 unreachable below 5 seats, six provisional scenario defaults, and two scoring modes never played to the end -- every one found by BUILDING the simulator rather than by playing it. A simulator rigorous enough to refuse ambiguity is a design instrument, because it cannot proceed past a rule that does not decide. All of it has been carried in prose in six places and one sat unread in an inbox for four days. The load-bearing rule: a design finding is not admissible without its reproduction. A register that collects opinions would reproduce this project's standing failure -- unexecuted verification -- in a new medium. Recorded as a judgment for the adversarial review rather than assumed: the engine-evolution meta the maintainer also asked about should NOT be built, because evidence/, decisions/, gates.toml and workplans already carry nineteen passes of it with dates, costs and falsifiers. A second register for the same subject is ceremony. The asymmetry is the point -- engine evolution has a home and game design does not. T05 backfills the six known findings as the TEST of the register: one that cannot express findings the project already has is the wrong register, and discovering that after designing it is why the order is survey, review, decide, specify, build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 20:48:08 +02:00
CB-WP-0022 T05/T06/T07: the register, and what its first run found T05. tools/design.py, make design, and the register in GroundRules.md -- 14 rows, no new file, because ADR-0012 D2 made §Underdetermined the register rather than building one beside it. Backfill was the test and it caught two things the ADR did not have. First, a `role` column. The first report alarmed on U2 and was wrong to: U2's scenario is green BECAUSE the provisional default it documents is implemented, which says nothing about whether ground-game agrees. GR-E01's was a counterexample that went green. Same colour, opposite meaning -- a register that cannot tell them apart either alarms constantly or never. Only a green counterexample alarms. Folded back into GameDesign §1.3. Second, it contradicted the survey. CB-RES-0007 said "six of the ten already have provisional scenarios." Measured -- grep -lE "\bU<n>\b" over scenarios/ground -- exactly ONE U-item names itself. Five provisional scenarios exist and four probably encode U-item defaults, but the mapping is not written down, so it is not checkable. Same defect class as the wrong premises, found inside the survey that proposed the fix. Now a reported debt: open, lacking a reproduction: 9, target 0. design.py carries the control design-baseline.py never had, asserted directly: a row citing a nonexistent file must not count as reproduced, using the exact path 2da19a4 deleted -- which the old tool called green. design-baseline.py is marked superseded rather than deleted; it is the evidence for how a wrong number got into a survey. T06. The report is a FILE in ground-game under GROUND-WP-0002, committed there, with a hub message that only points at it. It asks for no ruling: it carries GR-E01's withdrawal, our reproduction debt, and two notes that are explicitly not findings. And it had to acknowledge something nobody anticipated. GROUND-WP-0002 is finished -- all ten U-items were RULED 2026-08-03, every one confirmed as the default we simulate, plus five of six provisional scenarios. The survey said "0 of 10 ruled" two days later and this register was built saying `reported`. That is the unread-inbox failure running in the opposite direction: they answered and we did not collect it. The instrument's first run surfaced it. They are `ruled`, not `applied` -- lifting the now-settled provisional flags is owed and is not done, and make design shows them open until it is. T07. evidence/CB-EV-0021. Two of six catches in this pass came from execution rather than process (the role distinction from building it, the ten uncollected rulings from running it), which is InnerLoop §Design goal's prediction holding. make self-tests, facts-check, loop-lint: clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:22:33 +02:00
**Done 2026-08-05.** [specs/GameDesign.md](../specs/GameDesign.md) v1.0 —
not a register (ADR-0012 D2 put that in `§Underdetermined`).
CB-WP-0022 T04: specs/GameDesign.md -- what a reproduction must show Not a register; ADR-0012 D2 put that in GroundRules §Underdetermined. This spec says what may go in it, what a reproduction must show, how a finding dies, and how a trial game is run. §1.2 is written against evidence rather than principle. A finding must print the rows behind any number it claims, and the spec carries the table of what shipped instead: a sum ("12 in the file"), a green scenario ("4/6/9 against 5/7/9"), and a condition named without checking which one fired ("SOLVE on a face-down Problem"). "12" was arithmetically defensible and still wrong about the game -- that sentence is the requirement. §1.3's target is 0 reproductions that have gone green while open. GR-E01 would have tripped it four days before a human caught it by hand. No baseline rate is quoted. The 33% was withdrawn by C2 and the first honest denominator is T05's backfill; quoting a new number from a discredited instrument is how the first one got in. The trial protocol costs one flag: cb-play --record already writes a finished game as a scenario, so a trial is that plus a sibling .md in the player's own words. An observation is a NOTE until it has a reproduction, and notes may not cross the repo boundary and expire at 30 days on the existing provisional-age machinery. The maintainer's "I felt it was too easy but then we lost" is the case the protocol is shaped around -- forcing it into a schema at the moment of observation would lose it. loop-lint: no findings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:07:38 +02:00
**§1.2 is written against evidence rather than principle**: a finding must
CB-WP-0022 T05/T06/T07: the register, and what its first run found T05. tools/design.py, make design, and the register in GroundRules.md -- 14 rows, no new file, because ADR-0012 D2 made §Underdetermined the register rather than building one beside it. Backfill was the test and it caught two things the ADR did not have. First, a `role` column. The first report alarmed on U2 and was wrong to: U2's scenario is green BECAUSE the provisional default it documents is implemented, which says nothing about whether ground-game agrees. GR-E01's was a counterexample that went green. Same colour, opposite meaning -- a register that cannot tell them apart either alarms constantly or never. Only a green counterexample alarms. Folded back into GameDesign §1.3. Second, it contradicted the survey. CB-RES-0007 said "six of the ten already have provisional scenarios." Measured -- grep -lE "\bU<n>\b" over scenarios/ground -- exactly ONE U-item names itself. Five provisional scenarios exist and four probably encode U-item defaults, but the mapping is not written down, so it is not checkable. Same defect class as the wrong premises, found inside the survey that proposed the fix. Now a reported debt: open, lacking a reproduction: 9, target 0. design.py carries the control design-baseline.py never had, asserted directly: a row citing a nonexistent file must not count as reproduced, using the exact path 2da19a4 deleted -- which the old tool called green. design-baseline.py is marked superseded rather than deleted; it is the evidence for how a wrong number got into a survey. T06. The report is a FILE in ground-game under GROUND-WP-0002, committed there, with a hub message that only points at it. It asks for no ruling: it carries GR-E01's withdrawal, our reproduction debt, and two notes that are explicitly not findings. And it had to acknowledge something nobody anticipated. GROUND-WP-0002 is finished -- all ten U-items were RULED 2026-08-03, every one confirmed as the default we simulate, plus five of six provisional scenarios. The survey said "0 of 10 ruled" two days later and this register was built saying `reported`. That is the unread-inbox failure running in the opposite direction: they answered and we did not collect it. The instrument's first run surfaced it. They are `ruled`, not `applied` -- lifting the now-settled provisional flags is owed and is not done, and make design shows them open until it is. T07. evidence/CB-EV-0021. Two of six catches in this pass came from execution rather than process (the role distinction from building it, the ten uncollected rulings from running it), which is InnerLoop §Design goal's prediction holding. make self-tests, facts-check, loop-lint: clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:22:33 +02:00
print the rows behind any number it claims. *"12" was arithmetically
defensible and still wrong about the game.* **§1.3's target is `0`
reproductions gone green while open** — what GR-E01 would have tripped
four days before a human caught it. **No baseline rate is quoted.**
**The trial protocol costs one flag**: `cb-play --record` plus a sibling
`.md` in the player's own words. An observation is a **note** until it has
a reproduction — *"I felt it was too easy but then we lost"* is the case
it is shaped around, and a schema at the moment of observation would lose
it.
CB-WP-0022 T04: specs/GameDesign.md -- what a reproduction must show Not a register; ADR-0012 D2 put that in GroundRules §Underdetermined. This spec says what may go in it, what a reproduction must show, how a finding dies, and how a trial game is run. §1.2 is written against evidence rather than principle. A finding must print the rows behind any number it claims, and the spec carries the table of what shipped instead: a sum ("12 in the file"), a green scenario ("4/6/9 against 5/7/9"), and a condition named without checking which one fired ("SOLVE on a face-down Problem"). "12" was arithmetically defensible and still wrong about the game -- that sentence is the requirement. §1.3's target is 0 reproductions that have gone green while open. GR-E01 would have tripped it four days before a human caught it by hand. No baseline rate is quoted. The 33% was withdrawn by C2 and the first honest denominator is T05's backfill; quoting a new number from a discredited instrument is how the first one got in. The trial protocol costs one flag: cb-play --record already writes a finished game as a scenario, so a trial is that plus a sibling .md in the player's own words. An observation is a NOTE until it has a reproduction, and notes may not cross the repo boundary and expire at 30 days on the existing provisional-age machinery. The maintainer's "I felt it was too easy but then we lost" is the case the protocol is shaped around -- forcing it into a schema at the moment of observation would lose it. loop-lint: no findings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:07:38 +02:00
Declare CB-WP-0022: the design instrument, tier L The maintainer named a new aspect: clay-borg as a game design tool, with a register for design flaws, questions, results and trial protocols. Structural L on the maintainer-named-high-leverage trigger, and it amends INTENT. Chaos d8=6, no override. The insight is that this is already happening with no home. Five passes have produced ten underdetermined rules points, SOLVE offered on a face-down problem and always inert, GR-A13's wasted SOLVE, GR-E01 unreachable below 5 seats, six provisional scenario defaults, and two scoring modes never played to the end -- every one found by BUILDING the simulator rather than by playing it. A simulator rigorous enough to refuse ambiguity is a design instrument, because it cannot proceed past a rule that does not decide. All of it has been carried in prose in six places and one sat unread in an inbox for four days. The load-bearing rule: a design finding is not admissible without its reproduction. A register that collects opinions would reproduce this project's standing failure -- unexecuted verification -- in a new medium. Recorded as a judgment for the adversarial review rather than assumed: the engine-evolution meta the maintainer also asked about should NOT be built, because evidence/, decisions/, gates.toml and workplans already carry nineteen passes of it with dates, costs and falsifiers. A second register for the same subject is ceremony. The asymmetry is the point -- engine evolution has a home and game design does not. T05 backfills the six known findings as the TEST of the register: one that cannot express findings the project already has is the wrong register, and discovering that after designing it is why the order is survey, review, decide, specify, build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 20:48:08 +02:00
## Task: build it, and backfill what is already known
```task
id: CB-WP-0022-T05
CB-WP-0022 T05/T06/T07: the register, and what its first run found T05. tools/design.py, make design, and the register in GroundRules.md -- 14 rows, no new file, because ADR-0012 D2 made §Underdetermined the register rather than building one beside it. Backfill was the test and it caught two things the ADR did not have. First, a `role` column. The first report alarmed on U2 and was wrong to: U2's scenario is green BECAUSE the provisional default it documents is implemented, which says nothing about whether ground-game agrees. GR-E01's was a counterexample that went green. Same colour, opposite meaning -- a register that cannot tell them apart either alarms constantly or never. Only a green counterexample alarms. Folded back into GameDesign §1.3. Second, it contradicted the survey. CB-RES-0007 said "six of the ten already have provisional scenarios." Measured -- grep -lE "\bU<n>\b" over scenarios/ground -- exactly ONE U-item names itself. Five provisional scenarios exist and four probably encode U-item defaults, but the mapping is not written down, so it is not checkable. Same defect class as the wrong premises, found inside the survey that proposed the fix. Now a reported debt: open, lacking a reproduction: 9, target 0. design.py carries the control design-baseline.py never had, asserted directly: a row citing a nonexistent file must not count as reproduced, using the exact path 2da19a4 deleted -- which the old tool called green. design-baseline.py is marked superseded rather than deleted; it is the evidence for how a wrong number got into a survey. T06. The report is a FILE in ground-game under GROUND-WP-0002, committed there, with a hub message that only points at it. It asks for no ruling: it carries GR-E01's withdrawal, our reproduction debt, and two notes that are explicitly not findings. And it had to acknowledge something nobody anticipated. GROUND-WP-0002 is finished -- all ten U-items were RULED 2026-08-03, every one confirmed as the default we simulate, plus five of six provisional scenarios. The survey said "0 of 10 ruled" two days later and this register was built saying `reported`. That is the unread-inbox failure running in the opposite direction: they answered and we did not collect it. The instrument's first run surfaced it. They are `ruled`, not `applied` -- lifting the now-settled provisional flags is owed and is not done, and make design shows them open until it is. T07. evidence/CB-EV-0021. Two of six catches in this pass came from execution rather than process (the role distinction from building it, the ten uncollected rulings from running it), which is InnerLoop §Design goal's prediction holding. make self-tests, facts-check, loop-lint: clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:22:33 +02:00
status: done
Declare CB-WP-0022: the design instrument, tier L The maintainer named a new aspect: clay-borg as a game design tool, with a register for design flaws, questions, results and trial protocols. Structural L on the maintainer-named-high-leverage trigger, and it amends INTENT. Chaos d8=6, no override. The insight is that this is already happening with no home. Five passes have produced ten underdetermined rules points, SOLVE offered on a face-down problem and always inert, GR-A13's wasted SOLVE, GR-E01 unreachable below 5 seats, six provisional scenario defaults, and two scoring modes never played to the end -- every one found by BUILDING the simulator rather than by playing it. A simulator rigorous enough to refuse ambiguity is a design instrument, because it cannot proceed past a rule that does not decide. All of it has been carried in prose in six places and one sat unread in an inbox for four days. The load-bearing rule: a design finding is not admissible without its reproduction. A register that collects opinions would reproduce this project's standing failure -- unexecuted verification -- in a new medium. Recorded as a judgment for the adversarial review rather than assumed: the engine-evolution meta the maintainer also asked about should NOT be built, because evidence/, decisions/, gates.toml and workplans already carry nineteen passes of it with dates, costs and falsifiers. A second register for the same subject is ceremony. The asymmetry is the point -- engine evolution has a home and game design does not. T05 backfills the six known findings as the TEST of the register: one that cannot express findings the project already has is the wrong register, and discovering that after designing it is why the order is survey, review, decide, specify, build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 20:48:08 +02:00
priority: high
2026-08-04 00:21:21 +02:00
state_hub_task_id: "469413ba-ed47-4840-a60e-ee7f95dc06ff"
Declare CB-WP-0022: the design instrument, tier L The maintainer named a new aspect: clay-borg as a game design tool, with a register for design flaws, questions, results and trial protocols. Structural L on the maintainer-named-high-leverage trigger, and it amends INTENT. Chaos d8=6, no override. The insight is that this is already happening with no home. Five passes have produced ten underdetermined rules points, SOLVE offered on a face-down problem and always inert, GR-A13's wasted SOLVE, GR-E01 unreachable below 5 seats, six provisional scenario defaults, and two scoring modes never played to the end -- every one found by BUILDING the simulator rather than by playing it. A simulator rigorous enough to refuse ambiguity is a design instrument, because it cannot proceed past a rule that does not decide. All of it has been carried in prose in six places and one sat unread in an inbox for four days. The load-bearing rule: a design finding is not admissible without its reproduction. A register that collects opinions would reproduce this project's standing failure -- unexecuted verification -- in a new medium. Recorded as a judgment for the adversarial review rather than assumed: the engine-evolution meta the maintainer also asked about should NOT be built, because evidence/, decisions/, gates.toml and workplans already carry nineteen passes of it with dates, costs and falsifiers. A second register for the same subject is ceremony. The asymmetry is the point -- engine evolution has a home and game design does not. T05 backfills the six known findings as the TEST of the register: one that cannot express findings the project already has is the wrong register, and discovering that after designing it is why the order is survey, review, decide, specify, build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 20:48:08 +02:00
```
The register, the tool, and then **the six findings above entered into
it** — backfilling is the test. A register that cannot express findings
the project already has is the wrong register, and discovering that after
designing it is the point of doing it in this order.
`make design` (or equivalent) must report: open findings by kind, those
without a reproduction, and those never reported to their owner.
CB-WP-0022 T05/T06/T07: the register, and what its first run found T05. tools/design.py, make design, and the register in GroundRules.md -- 14 rows, no new file, because ADR-0012 D2 made §Underdetermined the register rather than building one beside it. Backfill was the test and it caught two things the ADR did not have. First, a `role` column. The first report alarmed on U2 and was wrong to: U2's scenario is green BECAUSE the provisional default it documents is implemented, which says nothing about whether ground-game agrees. GR-E01's was a counterexample that went green. Same colour, opposite meaning -- a register that cannot tell them apart either alarms constantly or never. Only a green counterexample alarms. Folded back into GameDesign §1.3. Second, it contradicted the survey. CB-RES-0007 said "six of the ten already have provisional scenarios." Measured -- grep -lE "\bU<n>\b" over scenarios/ground -- exactly ONE U-item names itself. Five provisional scenarios exist and four probably encode U-item defaults, but the mapping is not written down, so it is not checkable. Same defect class as the wrong premises, found inside the survey that proposed the fix. Now a reported debt: open, lacking a reproduction: 9, target 0. design.py carries the control design-baseline.py never had, asserted directly: a row citing a nonexistent file must not count as reproduced, using the exact path 2da19a4 deleted -- which the old tool called green. design-baseline.py is marked superseded rather than deleted; it is the evidence for how a wrong number got into a survey. T06. The report is a FILE in ground-game under GROUND-WP-0002, committed there, with a hub message that only points at it. It asks for no ruling: it carries GR-E01's withdrawal, our reproduction debt, and two notes that are explicitly not findings. And it had to acknowledge something nobody anticipated. GROUND-WP-0002 is finished -- all ten U-items were RULED 2026-08-03, every one confirmed as the default we simulate, plus five of six provisional scenarios. The survey said "0 of 10 ruled" two days later and this register was built saying `reported`. That is the unread-inbox failure running in the opposite direction: they answered and we did not collect it. The instrument's first run surfaced it. They are `ruled`, not `applied` -- lifting the now-settled provisional flags is owed and is not done, and make design shows them open until it is. T07. evidence/CB-EV-0021. Two of six catches in this pass came from execution rather than process (the role distinction from building it, the ten uncollected rulings from running it), which is InnerLoop §Design goal's prediction holding. make self-tests, facts-check, loop-lint: clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:22:33 +02:00
**Done 2026-08-05.** `tools/design.py`, `make design`, and the register in
[`GroundRules.md`](../specs/GroundRules.md) — **14 rows, no new file.**
Backfill was the test. The taxonomy held (five kinds, no sixth), and it
**produced a `role` column ADR-0012 does not have**: the first report
alarmed on U2, wrongly — a green *default* is expected, a green
*counterexample* is the alarm. Folded into GameDesign §1.3. It also
contradicted the survey: **one** U-item names itself in a scenario, not
six. Detail and figures:
[CB-EV-0021](../evidence/CB-EV-0021-the-design-instrument.md) §4, §6.
Declare CB-WP-0022: the design instrument, tier L The maintainer named a new aspect: clay-borg as a game design tool, with a register for design flaws, questions, results and trial protocols. Structural L on the maintainer-named-high-leverage trigger, and it amends INTENT. Chaos d8=6, no override. The insight is that this is already happening with no home. Five passes have produced ten underdetermined rules points, SOLVE offered on a face-down problem and always inert, GR-A13's wasted SOLVE, GR-E01 unreachable below 5 seats, six provisional scenario defaults, and two scoring modes never played to the end -- every one found by BUILDING the simulator rather than by playing it. A simulator rigorous enough to refuse ambiguity is a design instrument, because it cannot proceed past a rule that does not decide. All of it has been carried in prose in six places and one sat unread in an inbox for four days. The load-bearing rule: a design finding is not admissible without its reproduction. A register that collects opinions would reproduce this project's standing failure -- unexecuted verification -- in a new medium. Recorded as a judgment for the adversarial review rather than assumed: the engine-evolution meta the maintainer also asked about should NOT be built, because evidence/, decisions/, gates.toml and workplans already carry nineteen passes of it with dates, costs and falsifiers. A second register for the same subject is ceremony. The asymmetry is the point -- engine evolution has a home and game design does not. T05 backfills the six known findings as the TEST of the register: one that cannot express findings the project already has is the wrong register, and discovering that after designing it is why the order is survey, review, decide, specify, build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 20:48:08 +02:00
## Task: report to ground-game, mechanically
```task
id: CB-WP-0022-T06
CB-WP-0022 T05/T06/T07: the register, and what its first run found T05. tools/design.py, make design, and the register in GroundRules.md -- 14 rows, no new file, because ADR-0012 D2 made §Underdetermined the register rather than building one beside it. Backfill was the test and it caught two things the ADR did not have. First, a `role` column. The first report alarmed on U2 and was wrong to: U2's scenario is green BECAUSE the provisional default it documents is implemented, which says nothing about whether ground-game agrees. GR-E01's was a counterexample that went green. Same colour, opposite meaning -- a register that cannot tell them apart either alarms constantly or never. Only a green counterexample alarms. Folded back into GameDesign §1.3. Second, it contradicted the survey. CB-RES-0007 said "six of the ten already have provisional scenarios." Measured -- grep -lE "\bU<n>\b" over scenarios/ground -- exactly ONE U-item names itself. Five provisional scenarios exist and four probably encode U-item defaults, but the mapping is not written down, so it is not checkable. Same defect class as the wrong premises, found inside the survey that proposed the fix. Now a reported debt: open, lacking a reproduction: 9, target 0. design.py carries the control design-baseline.py never had, asserted directly: a row citing a nonexistent file must not count as reproduced, using the exact path 2da19a4 deleted -- which the old tool called green. design-baseline.py is marked superseded rather than deleted; it is the evidence for how a wrong number got into a survey. T06. The report is a FILE in ground-game under GROUND-WP-0002, committed there, with a hub message that only points at it. It asks for no ruling: it carries GR-E01's withdrawal, our reproduction debt, and two notes that are explicitly not findings. And it had to acknowledge something nobody anticipated. GROUND-WP-0002 is finished -- all ten U-items were RULED 2026-08-03, every one confirmed as the default we simulate, plus five of six provisional scenarios. The survey said "0 of 10 ruled" two days later and this register was built saying `reported`. That is the unread-inbox failure running in the opposite direction: they answered and we did not collect it. The instrument's first run surfaced it. They are `ruled`, not `applied` -- lifting the now-settled provisional flags is owed and is not done, and make design shows them open until it is. T07. evidence/CB-EV-0021. Two of six catches in this pass came from execution rather than process (the role distinction from building it, the ten uncollected rulings from running it), which is InnerLoop §Design goal's prediction holding. make self-tests, facts-check, loop-lint: clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:22:33 +02:00
status: done
Declare CB-WP-0022: the design instrument, tier L The maintainer named a new aspect: clay-borg as a game design tool, with a register for design flaws, questions, results and trial protocols. Structural L on the maintainer-named-high-leverage trigger, and it amends INTENT. Chaos d8=6, no override. The insight is that this is already happening with no home. Five passes have produced ten underdetermined rules points, SOLVE offered on a face-down problem and always inert, GR-A13's wasted SOLVE, GR-E01 unreachable below 5 seats, six provisional scenario defaults, and two scoring modes never played to the end -- every one found by BUILDING the simulator rather than by playing it. A simulator rigorous enough to refuse ambiguity is a design instrument, because it cannot proceed past a rule that does not decide. All of it has been carried in prose in six places and one sat unread in an inbox for four days. The load-bearing rule: a design finding is not admissible without its reproduction. A register that collects opinions would reproduce this project's standing failure -- unexecuted verification -- in a new medium. Recorded as a judgment for the adversarial review rather than assumed: the engine-evolution meta the maintainer also asked about should NOT be built, because evidence/, decisions/, gates.toml and workplans already carry nineteen passes of it with dates, costs and falsifiers. A second register for the same subject is ceremony. The asymmetry is the point -- engine evolution has a home and game design does not. T05 backfills the six known findings as the TEST of the register: one that cannot express findings the project already has is the wrong register, and discovering that after designing it is why the order is survey, review, decide, specify, build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 20:48:08 +02:00
priority: high
2026-08-04 00:21:21 +02:00
state_hub_task_id: "c6a659bd-c495-4754-9332-baadad51012a"
Declare CB-WP-0022: the design instrument, tier L The maintainer named a new aspect: clay-borg as a game design tool, with a register for design flaws, questions, results and trial protocols. Structural L on the maintainer-named-high-leverage trigger, and it amends INTENT. Chaos d8=6, no override. The insight is that this is already happening with no home. Five passes have produced ten underdetermined rules points, SOLVE offered on a face-down problem and always inert, GR-A13's wasted SOLVE, GR-E01 unreachable below 5 seats, six provisional scenario defaults, and two scoring modes never played to the end -- every one found by BUILDING the simulator rather than by playing it. A simulator rigorous enough to refuse ambiguity is a design instrument, because it cannot proceed past a rule that does not decide. All of it has been carried in prose in six places and one sat unread in an inbox for four days. The load-bearing rule: a design finding is not admissible without its reproduction. A register that collects opinions would reproduce this project's standing failure -- unexecuted verification -- in a new medium. Recorded as a judgment for the adversarial review rather than assumed: the engine-evolution meta the maintainer also asked about should NOT be built, because evidence/, decisions/, gates.toml and workplans already carry nineteen passes of it with dates, costs and falsifiers. A second register for the same subject is ceremony. The asymmetry is the point -- engine evolution has a home and game design does not. T05 backfills the six known findings as the TEST of the register: one that cannot express findings the project already has is the wrong register, and discovering that after designing it is why the order is survey, review, decide, specify, build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 20:48:08 +02:00
```
Generate the report and send it. **The message that sat unread for four
CB-WP-0022 T05/T06/T07: the register, and what its first run found T05. tools/design.py, make design, and the register in GroundRules.md -- 14 rows, no new file, because ADR-0012 D2 made §Underdetermined the register rather than building one beside it. Backfill was the test and it caught two things the ADR did not have. First, a `role` column. The first report alarmed on U2 and was wrong to: U2's scenario is green BECAUSE the provisional default it documents is implemented, which says nothing about whether ground-game agrees. GR-E01's was a counterexample that went green. Same colour, opposite meaning -- a register that cannot tell them apart either alarms constantly or never. Only a green counterexample alarms. Folded back into GameDesign §1.3. Second, it contradicted the survey. CB-RES-0007 said "six of the ten already have provisional scenarios." Measured -- grep -lE "\bU<n>\b" over scenarios/ground -- exactly ONE U-item names itself. Five provisional scenarios exist and four probably encode U-item defaults, but the mapping is not written down, so it is not checkable. Same defect class as the wrong premises, found inside the survey that proposed the fix. Now a reported debt: open, lacking a reproduction: 9, target 0. design.py carries the control design-baseline.py never had, asserted directly: a row citing a nonexistent file must not count as reproduced, using the exact path 2da19a4 deleted -- which the old tool called green. design-baseline.py is marked superseded rather than deleted; it is the evidence for how a wrong number got into a survey. T06. The report is a FILE in ground-game under GROUND-WP-0002, committed there, with a hub message that only points at it. It asks for no ruling: it carries GR-E01's withdrawal, our reproduction debt, and two notes that are explicitly not findings. And it had to acknowledge something nobody anticipated. GROUND-WP-0002 is finished -- all ten U-items were RULED 2026-08-03, every one confirmed as the default we simulate, plus five of six provisional scenarios. The survey said "0 of 10 ruled" two days later and this register was built saying `reported`. That is the unread-inbox failure running in the opposite direction: they answered and we did not collect it. The instrument's first run surfaced it. They are `ruled`, not `applied` -- lifting the now-settled provisional flags is owed and is not done, and make design shows them open until it is. T07. evidence/CB-EV-0021. Two of six catches in this pass came from execution rather than process (the role distinction from building it, the ten uncollected rulings from running it), which is InnerLoop §Design goal's prediction holding. make self-tests, facts-check, loop-lint: clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:22:33 +02:00
days is the baseline to beat** — the failure was not the message, it was
that nothing pointed at it. So the report lands as a file in `ground-game`
under its own workplan, extending GROUND-WP-0002 rather than duplicating
it.
Declare CB-WP-0022: the design instrument, tier L The maintainer named a new aspect: clay-borg as a game design tool, with a register for design flaws, questions, results and trial protocols. Structural L on the maintainer-named-high-leverage trigger, and it amends INTENT. Chaos d8=6, no override. The insight is that this is already happening with no home. Five passes have produced ten underdetermined rules points, SOLVE offered on a face-down problem and always inert, GR-A13's wasted SOLVE, GR-E01 unreachable below 5 seats, six provisional scenario defaults, and two scoring modes never played to the end -- every one found by BUILDING the simulator rather than by playing it. A simulator rigorous enough to refuse ambiguity is a design instrument, because it cannot proceed past a rule that does not decide. All of it has been carried in prose in six places and one sat unread in an inbox for four days. The load-bearing rule: a design finding is not admissible without its reproduction. A register that collects opinions would reproduce this project's standing failure -- unexecuted verification -- in a new medium. Recorded as a judgment for the adversarial review rather than assumed: the engine-evolution meta the maintainer also asked about should NOT be built, because evidence/, decisions/, gates.toml and workplans already carry nineteen passes of it with dates, costs and falsifiers. A second register for the same subject is ceremony. The asymmetry is the point -- engine evolution has a home and game design does not. T05 backfills the six known findings as the TEST of the register: one that cannot express findings the project already has is the wrong register, and discovering that after designing it is why the order is survey, review, decide, specify, build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 20:48:08 +02:00
CB-WP-0022 T02: the separate reviewer found the showcase finding was false First adversarial review in this repo run by a genuinely separate agent. CB-RES-0006's reviewer opened by conceding it could not be, and called its own findings "a lower bound on what a genuinely separate reviewer would find." This is the measurement: the separate reviewer ran git log against the survey's central example and found 2da19a4 had falsified it four days earlier, while the author -- who wrote that commit -- quoted the dead number twice. Seven challenges: four conceded, two conceded in part, one answered. C1 changes the design. "GR-E01 is admissible because 4/6/9 against 5/7/9 is a computation anyone can rerun" was a computation already rerun: the edition import measured 6/9/12, the conclusion inverted, and the scenario was renamed -unreachable- to -reachable-. That finding was one of the TWO that passed the reproduction rule. So three wrong premises have now reached ground-game and the third satisfied an existence test -- existence is not the property that was missing. The rule gains shape (ground-game's row-level deal table, promoted from a T04 addendum) and a clause the survey never contemplated: a reproduction must be able to fail. Ours went green and stayed admissible. C1 also caught a defect in flight. T06's payload, status todo, still named 4/6/9 and was queued to send it to ground-game as "no dataset reconciles them." Withdrawn before sending -- the fourth wrong premise, and the only one stopped. C2 withdraws the baseline's precision: design-baseline.py is a hand-maintained dict counting itself, has_reproduction never checks the file exists (its YES-control is green against a deleted path), and Makefile:127 runs only --self-test so the reporting path has no CI. The direction survives; 33% is not a measured rate and T05 must not build on it. C3: "six provisional defaults" is five, GR-E01 double-counted. C4: GroundRules §Underdetermined was never evaluated as a candidate and already delivers four of five benchmarks -- T03's burden flips to arguing extension over replacement. Survived: the rule's affordability, and reuse of the provisional machinery. loop-lint: no findings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:03:12 +02:00
Include the findings this pass has sharpened:
- **SOLVE's legality** against a face-down Problem or an unmatchable suit —
and note that the case we *reported* was not the case that fired
(CB-WP-0023 T01).
CB-WP-0022 T05/T06/T07: the register, and what its first run found T05. tools/design.py, make design, and the register in GroundRules.md -- 14 rows, no new file, because ADR-0012 D2 made §Underdetermined the register rather than building one beside it. Backfill was the test and it caught two things the ADR did not have. First, a `role` column. The first report alarmed on U2 and was wrong to: U2's scenario is green BECAUSE the provisional default it documents is implemented, which says nothing about whether ground-game agrees. GR-E01's was a counterexample that went green. Same colour, opposite meaning -- a register that cannot tell them apart either alarms constantly or never. Only a green counterexample alarms. Folded back into GameDesign §1.3. Second, it contradicted the survey. CB-RES-0007 said "six of the ten already have provisional scenarios." Measured -- grep -lE "\bU<n>\b" over scenarios/ground -- exactly ONE U-item names itself. Five provisional scenarios exist and four probably encode U-item defaults, but the mapping is not written down, so it is not checkable. Same defect class as the wrong premises, found inside the survey that proposed the fix. Now a reported debt: open, lacking a reproduction: 9, target 0. design.py carries the control design-baseline.py never had, asserted directly: a row citing a nonexistent file must not count as reproduced, using the exact path 2da19a4 deleted -- which the old tool called green. design-baseline.py is marked superseded rather than deleted; it is the evidence for how a wrong number got into a survey. T06. The report is a FILE in ground-game under GROUND-WP-0002, committed there, with a hub message that only points at it. It asks for no ruling: it carries GR-E01's withdrawal, our reproduction debt, and two notes that are explicitly not findings. And it had to acknowledge something nobody anticipated. GROUND-WP-0002 is finished -- all ten U-items were RULED 2026-08-03, every one confirmed as the default we simulate, plus five of six provisional scenarios. The survey said "0 of 10 ruled" two days later and this register was built saying `reported`. That is the unread-inbox failure running in the opposite direction: they answered and we did not collect it. The instrument's first run surfaced it. They are `ruled`, not `applied` -- lifting the now-settled provisional flags is owed and is not done, and make design shows them open until it is. T07. evidence/CB-EV-0021. Two of six catches in this pass came from execution rather than process (the role distinction from building it, the ten uncollected rulings from running it), which is InnerLoop §Design goal's prediction holding. make self-tests, facts-check, loop-lint: clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:22:33 +02:00
- ~~**GR-E01 vs GR-S01** — 4/6/9 against 5/7/9~~ — **withdrawn
2026-08-05, before sending.** C1 found `2da19a4` had measured **6/9/12**:
the dataset reconciles them. It would have been the **fourth** wrong
premise to reach `ground-game` and is the only one caught before
transmission. **Report the withdrawal** — a claim retracted silently is
how the first three survived.
**Done 2026-08-05.**
[`ground-game/workplans/GROUND-WP-0002-clay-borg-report-260805.md`](../../ground-game/workplans/GROUND-WP-0002-clay-borg-report-260805.md),
committed there, with a hub message that only *points at* the file.
**The report asks for no ruling.** It carries GR-E01's withdrawal, our own
reproduction debt, and two notes that are explicitly not findings.
**And it acknowledged something the pass did not expect.**
GROUND-WP-0002 is `finished`: **all ten U-items were ruled 2026-08-03**,
every one confirmed as the default we simulate. CB-RES-0007 said *"0 of 10
ruled"* two days later. **The unread-inbox failure running in the opposite
direction** — they answered and we did not collect it. The instrument's
first run surfaced it.
Declare CB-WP-0022: the design instrument, tier L The maintainer named a new aspect: clay-borg as a game design tool, with a register for design flaws, questions, results and trial protocols. Structural L on the maintainer-named-high-leverage trigger, and it amends INTENT. Chaos d8=6, no override. The insight is that this is already happening with no home. Five passes have produced ten underdetermined rules points, SOLVE offered on a face-down problem and always inert, GR-A13's wasted SOLVE, GR-E01 unreachable below 5 seats, six provisional scenario defaults, and two scoring modes never played to the end -- every one found by BUILDING the simulator rather than by playing it. A simulator rigorous enough to refuse ambiguity is a design instrument, because it cannot proceed past a rule that does not decide. All of it has been carried in prose in six places and one sat unread in an inbox for four days. The load-bearing rule: a design finding is not admissible without its reproduction. A register that collects opinions would reproduce this project's standing failure -- unexecuted verification -- in a new medium. Recorded as a judgment for the adversarial review rather than assumed: the engine-evolution meta the maintainer also asked about should NOT be built, because evidence/, decisions/, gates.toml and workplans already carry nineteen passes of it with dates, costs and falsifiers. A second register for the same subject is ceremony. The asymmetry is the point -- engine evolution has a home and game design does not. T05 backfills the six known findings as the TEST of the register: one that cannot express findings the project already has is the wrong register, and discovering that after designing it is why the order is survey, review, decide, specify, build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 20:48:08 +02:00
## Task: evidence
```task
2026-08-04 00:21:21 +02:00
id: CB-WP-0022-T07
CB-WP-0022 T05/T06/T07: the register, and what its first run found T05. tools/design.py, make design, and the register in GroundRules.md -- 14 rows, no new file, because ADR-0012 D2 made §Underdetermined the register rather than building one beside it. Backfill was the test and it caught two things the ADR did not have. First, a `role` column. The first report alarmed on U2 and was wrong to: U2's scenario is green BECAUSE the provisional default it documents is implemented, which says nothing about whether ground-game agrees. GR-E01's was a counterexample that went green. Same colour, opposite meaning -- a register that cannot tell them apart either alarms constantly or never. Only a green counterexample alarms. Folded back into GameDesign §1.3. Second, it contradicted the survey. CB-RES-0007 said "six of the ten already have provisional scenarios." Measured -- grep -lE "\bU<n>\b" over scenarios/ground -- exactly ONE U-item names itself. Five provisional scenarios exist and four probably encode U-item defaults, but the mapping is not written down, so it is not checkable. Same defect class as the wrong premises, found inside the survey that proposed the fix. Now a reported debt: open, lacking a reproduction: 9, target 0. design.py carries the control design-baseline.py never had, asserted directly: a row citing a nonexistent file must not count as reproduced, using the exact path 2da19a4 deleted -- which the old tool called green. design-baseline.py is marked superseded rather than deleted; it is the evidence for how a wrong number got into a survey. T06. The report is a FILE in ground-game under GROUND-WP-0002, committed there, with a hub message that only points at it. It asks for no ruling: it carries GR-E01's withdrawal, our reproduction debt, and two notes that are explicitly not findings. And it had to acknowledge something nobody anticipated. GROUND-WP-0002 is finished -- all ten U-items were RULED 2026-08-03, every one confirmed as the default we simulate, plus five of six provisional scenarios. The survey said "0 of 10 ruled" two days later and this register was built saying `reported`. That is the unread-inbox failure running in the opposite direction: they answered and we did not collect it. The instrument's first run surfaced it. They are `ruled`, not `applied` -- lifting the now-settled provisional flags is owed and is not done, and make design shows them open until it is. T07. evidence/CB-EV-0021. Two of six catches in this pass came from execution rather than process (the role distinction from building it, the ten uncollected rulings from running it), which is InnerLoop §Design goal's prediction holding. make self-tests, facts-check, loop-lint: clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:22:33 +02:00
status: done
Declare CB-WP-0022: the design instrument, tier L The maintainer named a new aspect: clay-borg as a game design tool, with a register for design flaws, questions, results and trial protocols. Structural L on the maintainer-named-high-leverage trigger, and it amends INTENT. Chaos d8=6, no override. The insight is that this is already happening with no home. Five passes have produced ten underdetermined rules points, SOLVE offered on a face-down problem and always inert, GR-A13's wasted SOLVE, GR-E01 unreachable below 5 seats, six provisional scenario defaults, and two scoring modes never played to the end -- every one found by BUILDING the simulator rather than by playing it. A simulator rigorous enough to refuse ambiguity is a design instrument, because it cannot proceed past a rule that does not decide. All of it has been carried in prose in six places and one sat unread in an inbox for four days. The load-bearing rule: a design finding is not admissible without its reproduction. A register that collects opinions would reproduce this project's standing failure -- unexecuted verification -- in a new medium. Recorded as a judgment for the adversarial review rather than assumed: the engine-evolution meta the maintainer also asked about should NOT be built, because evidence/, decisions/, gates.toml and workplans already carry nineteen passes of it with dates, costs and falsifiers. A second register for the same subject is ceremony. The asymmetry is the point -- engine evolution has a home and game design does not. T05 backfills the six known findings as the TEST of the register: one that cannot express findings the project already has is the wrong register, and discovering that after designing it is why the order is survey, review, decide, specify, build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 20:48:08 +02:00
priority: high
2026-08-04 00:21:21 +02:00
state_hub_task_id: "4fb77911-486d-4d9e-a53d-c4e486c74ce2"
Declare CB-WP-0022: the design instrument, tier L The maintainer named a new aspect: clay-borg as a game design tool, with a register for design flaws, questions, results and trial protocols. Structural L on the maintainer-named-high-leverage trigger, and it amends INTENT. Chaos d8=6, no override. The insight is that this is already happening with no home. Five passes have produced ten underdetermined rules points, SOLVE offered on a face-down problem and always inert, GR-A13's wasted SOLVE, GR-E01 unreachable below 5 seats, six provisional scenario defaults, and two scoring modes never played to the end -- every one found by BUILDING the simulator rather than by playing it. A simulator rigorous enough to refuse ambiguity is a design instrument, because it cannot proceed past a rule that does not decide. All of it has been carried in prose in six places and one sat unread in an inbox for four days. The load-bearing rule: a design finding is not admissible without its reproduction. A register that collects opinions would reproduce this project's standing failure -- unexecuted verification -- in a new medium. Recorded as a judgment for the adversarial review rather than assumed: the engine-evolution meta the maintainer also asked about should NOT be built, because evidence/, decisions/, gates.toml and workplans already carry nineteen passes of it with dates, costs and falsifiers. A second register for the same subject is ceremony. The asymmetry is the point -- engine evolution has a home and game design does not. T05 backfills the six known findings as the TEST of the register: one that cannot express findings the project already has is the wrong register, and discovering that after designing it is why the order is survey, review, decide, specify, build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 20:48:08 +02:00
```
CB-WP-0024/0025: what play reported, split into a renderer and a search Five remarks from the maintainer's test games, checked against the code before being written down — two had already reached ground-game on wrong premises, so a claim now names the line that makes it true. Three of the five turned out to be data the projection already carries, drawn as text: solution_deck_len, solution_discard, and OutcomeView's personal/mastery/winners. One control (`close — I have read this`) is labelled as a reading but shuts the server down, and leaves a live-looking page pointing at a dead port. One number does not exist at all: table.rs loops run_game and keeps only the last summary. CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the other seats' plays as objects on the table, the ending control saying what it does, and a tally that survives "play again". CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how hard is this" are the same search asked twice. Tier L because the information boundary is the whole design problem: a solver reading GroundState sees the deck the rules hide, and would tell the maintainer he could have won by playing a card he had no way to know was there. Also unblocks GROUND-WP-0005, active with both tasks waiting on a measured difficulty baseline. Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023 shipped CB-EV-0020 first. loop-lint: no findings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 12:59:08 +02:00
`evidence/CB-EV-0021-*.md`. *(Was CB-EV-0020 when written; CB-WP-0023
shipped that number first — `evidence/CB-EV-0020-solve-legality.md` — so
this one moves rather than collides.)*
Declare CB-WP-0022: the design instrument, tier L The maintainer named a new aspect: clay-borg as a game design tool, with a register for design flaws, questions, results and trial protocols. Structural L on the maintainer-named-high-leverage trigger, and it amends INTENT. Chaos d8=6, no override. The insight is that this is already happening with no home. Five passes have produced ten underdetermined rules points, SOLVE offered on a face-down problem and always inert, GR-A13's wasted SOLVE, GR-E01 unreachable below 5 seats, six provisional scenario defaults, and two scoring modes never played to the end -- every one found by BUILDING the simulator rather than by playing it. A simulator rigorous enough to refuse ambiguity is a design instrument, because it cannot proceed past a rule that does not decide. All of it has been carried in prose in six places and one sat unread in an inbox for four days. The load-bearing rule: a design finding is not admissible without its reproduction. A register that collects opinions would reproduce this project's standing failure -- unexecuted verification -- in a new medium. Recorded as a judgment for the adversarial review rather than assumed: the engine-evolution meta the maintainer also asked about should NOT be built, because evidence/, decisions/, gates.toml and workplans already carry nineteen passes of it with dates, costs and falsifiers. A second register for the same subject is ceremony. The asymmetry is the point -- engine evolution has a home and game design does not. T05 backfills the six known findings as the TEST of the register: one that cannot express findings the project already has is the wrong register, and discovering that after designing it is why the order is survey, review, decide, specify, build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 20:48:08 +02:00
- **Whether backfilling changed the design** — if all six findings fit the
first taxonomy, say so and be suspicious of it.
- **What tier L cost against what it caught**, since this is the second
full-weight L pass and CB-WP-0012's deleted its own structural trigger.
- **The engine-evolution question**, as the review left it.
- **Quote CB-WP-0021's cost by re-running the instrument.**
CB-WP-0022 T05/T06/T07: the register, and what its first run found T05. tools/design.py, make design, and the register in GroundRules.md -- 14 rows, no new file, because ADR-0012 D2 made §Underdetermined the register rather than building one beside it. Backfill was the test and it caught two things the ADR did not have. First, a `role` column. The first report alarmed on U2 and was wrong to: U2's scenario is green BECAUSE the provisional default it documents is implemented, which says nothing about whether ground-game agrees. GR-E01's was a counterexample that went green. Same colour, opposite meaning -- a register that cannot tell them apart either alarms constantly or never. Only a green counterexample alarms. Folded back into GameDesign §1.3. Second, it contradicted the survey. CB-RES-0007 said "six of the ten already have provisional scenarios." Measured -- grep -lE "\bU<n>\b" over scenarios/ground -- exactly ONE U-item names itself. Five provisional scenarios exist and four probably encode U-item defaults, but the mapping is not written down, so it is not checkable. Same defect class as the wrong premises, found inside the survey that proposed the fix. Now a reported debt: open, lacking a reproduction: 9, target 0. design.py carries the control design-baseline.py never had, asserted directly: a row citing a nonexistent file must not count as reproduced, using the exact path 2da19a4 deleted -- which the old tool called green. design-baseline.py is marked superseded rather than deleted; it is the evidence for how a wrong number got into a survey. T06. The report is a FILE in ground-game under GROUND-WP-0002, committed there, with a hub message that only points at it. It asks for no ruling: it carries GR-E01's withdrawal, our reproduction debt, and two notes that are explicitly not findings. And it had to acknowledge something nobody anticipated. GROUND-WP-0002 is finished -- all ten U-items were RULED 2026-08-03, every one confirmed as the default we simulate, plus five of six provisional scenarios. The survey said "0 of 10 ruled" two days later and this register was built saying `reported`. That is the unread-inbox failure running in the opposite direction: they answered and we did not collect it. The instrument's first run surfaced it. They are `ruled`, not `applied` -- lifting the now-settled provisional flags is owed and is not done, and make design shows them open until it is. T07. evidence/CB-EV-0021. Two of six catches in this pass came from execution rather than process (the role distinction from building it, the ten uncollected rulings from running it), which is InnerLoop §Design goal's prediction holding. make self-tests, facts-check, loop-lint: clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:22:33 +02:00
**Done 2026-08-05.**
[CB-EV-0021](../evidence/CB-EV-0021-the-design-instrument.md).
**Backfill did change the design** — and the honest answer to *"be
suspicious if all six fit"* is that only **five** were entered (one was a
double-count), so fitting them is close to circular. The taxonomy's real
test is the seventh finding.
**Tier L's cost against what it caught**: four of six catches came only
from the separate reviewer, and **two came from execution rather than
process** — the `role` distinction from building it, the ten uncollected
rulings from running it. That is InnerLoop §Design goal's prediction
holding, and an argument against front-loading more review rather than
less.