Compare commits

...

6 commits

Author SHA1 Message Date
561facce43 CB-WP-0022 T05/T06/T07: the register, and what its first run found
Some checks failed
ci / check (push) Has been cancelled
T05. tools/design.py, make design, and the register in GroundRules.md --
14 rows, no new file, because ADR-0012 D2 made §Underdetermined the
register rather than building one beside it.

Backfill was the test and it caught two things the ADR did not have.

First, a `role` column. The first report alarmed on U2 and was wrong to:
U2's scenario is green BECAUSE the provisional default it documents is
implemented, which says nothing about whether ground-game agrees. GR-E01's
was a counterexample that went green. Same colour, opposite meaning -- a
register that cannot tell them apart either alarms constantly or never.
Only a green counterexample alarms. Folded back into GameDesign §1.3.

Second, it contradicted the survey. CB-RES-0007 said "six of the ten
already have provisional scenarios." Measured -- grep -lE "\bU<n>\b" over
scenarios/ground -- exactly ONE U-item names itself. Five provisional
scenarios exist and four probably encode U-item defaults, but the mapping
is not written down, so it is not checkable. Same defect class as the
wrong premises, found inside the survey that proposed the fix. Now a
reported debt: open, lacking a reproduction: 9, target 0.

design.py carries the control design-baseline.py never had, asserted
directly: a row citing a nonexistent file must not count as reproduced,
using the exact path 2da19a4 deleted -- which the old tool called green.
design-baseline.py is marked superseded rather than deleted; it is the
evidence for how a wrong number got into a survey.

T06. The report is a FILE in ground-game under GROUND-WP-0002, committed
there, with a hub message that only points at it. It asks for no ruling:
it carries GR-E01's withdrawal, our reproduction debt, and two notes that
are explicitly not findings.

And it had to acknowledge something nobody anticipated. GROUND-WP-0002 is
finished -- all ten U-items were RULED 2026-08-03, every one confirmed as
the default we simulate, plus five of six provisional scenarios. The
survey said "0 of 10 ruled" two days later and this register was built
saying `reported`. That is the unread-inbox failure running in the
opposite direction: they answered and we did not collect it. The
instrument's first run surfaced it. They are `ruled`, not `applied` --
lifting the now-settled provisional flags is owed and is not done, and
make design shows them open until it is.

T07. evidence/CB-EV-0021. Two of six catches in this pass came from
execution rather than process (the role distinction from building it, the
ten uncollected rulings from running it), which is InnerLoop §Design
goal's prediction holding.

make self-tests, facts-check, loop-lint: clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:22:33 +02:00
5092684de5 CB-WP-0022 T04 follow-up: trim the workplan back under the loadability limit
loop-lint flagged 429 lines against the ~400 limit, and it was right about
the cause: the T02/T03/T04 completion records restated content that
ADR-0012, GameDesign.md and the challenge/response trail already carry.
Trimmed to pointers plus the one sentence each that is not written down
elsewhere.

400 lines, loop-lint clean. No content lost from the artifacts that own
it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:09:28 +02:00
b4f51effad CB-WP-0022 T04: specs/GameDesign.md -- what a reproduction must show
Not a register; ADR-0012 D2 put that in GroundRules §Underdetermined. This
spec says what may go in it, what a reproduction must show, how a finding
dies, and how a trial game is run.

§1.2 is written against evidence rather than principle. A finding must
print the rows behind any number it claims, and the spec carries the table
of what shipped instead: a sum ("12 in the file"), a green scenario
("4/6/9 against 5/7/9"), and a condition named without checking which one
fired ("SOLVE on a face-down Problem"). "12" was arithmetically defensible
and still wrong about the game -- that sentence is the requirement.

§1.3's target is 0 reproductions that have gone green while open. GR-E01
would have tripped it four days before a human caught it by hand.

No baseline rate is quoted. The 33% was withdrawn by C2 and the first
honest denominator is T05's backfill; quoting a new number from a
discredited instrument is how the first one got in.

The trial protocol costs one flag: cb-play --record already writes a
finished game as a scenario, so a trial is that plus a sibling .md in the
player's own words. An observation is a NOTE until it has a reproduction,
and notes may not cross the repo boundary and expire at 30 days on the
existing provisional-age machinery. The maintainer's "I felt it was too
easy but then we lost" is the case the protocol is shaped around --
forcing it into a schema at the moment of observation would lose it.

loop-lint: no findings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:07:38 +02:00
580081ef71 CB-WP-0022 T03: ADR-0012 -- the register already existed, and the rule was
one clause short

Nine decisions. Two were not on T03's list; both are the review's.

D2: specs/GroundRules.md §Underdetermined IS the register. C4 pointed out
it was never evaluated as a candidate, and against the survey's own five
benchmarks it already delivers four -- including the Magic Oracle property
("a ruling flips the scenario, not the kernel", :231-233) that the survey
travelled to Magic to discover and we had written down ourselves eight
days earlier. What it lacks is reproductions. So this pass extends a
section rather than building a register: no new file, no new schema, and
no second mechanism to disagree with the first.

D3: admissibility is three clauses. It exists; it has the ruled shape
(GROUND-WP-0004 T02's row-level table, never a sum -- promoted from a T04
addendum because two of three wrong premises were sums without tables);
and it CAN FAIL. The third is C1's. GR-E01's scenario went green when the
edition landed, and the finding stayed admissible and stayed queued for
transmission, because nothing in the rule said a passing artifact was a
signal. A green reproduction is an alarm, not a reassurance.

D1 applied: INTENT gains a fourth property, Instrument, worded as a
mechanism rather than an ambition and carrying its own falsifier -- if a
pass tolerates an undecided rule by quietly picking a default, the
property is false.

D4 five kinds, each forced by an existing finding; a sixth during backfill
means the taxonomy was invented. D5 lifecycle where `applied` means the
source changed, the queue empties while the log accumulates, and
withdrawals are reported rather than deleted -- GR-E01 is why. D6 notes
admitted but never reportable, 30-day expiry on the existing age
machinery; refusing them would discard the only class of finding the
engine cannot produce itself, which is CB-WP-0025's whole input. D7 no
engine-evolution register, on an inventory C5 corrected -- narrowed, not
settled. D8 design-baseline.py retired, kept as a dated snapshot because
deleting it erases the evidence for how 33% got in. D9 the artifact stays
here, ground-game gets a generated file under its own workplan.

loop-lint: no findings. facts-check: no findings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:05:51 +02:00
04c3a4977f CB-WP-0022 T02: the separate reviewer found the showcase finding was false
First adversarial review in this repo run by a genuinely separate agent.
CB-RES-0006's reviewer opened by conceding it could not be, and called its
own findings "a lower bound on what a genuinely separate reviewer would
find." This is the measurement: the separate reviewer ran git log against
the survey's central example and found 2da19a4 had falsified it four days
earlier, while the author -- who wrote that commit -- quoted the dead
number twice.

Seven challenges: four conceded, two conceded in part, one answered.

C1 changes the design. "GR-E01 is admissible because 4/6/9 against 5/7/9
is a computation anyone can rerun" was a computation already rerun: the
edition import measured 6/9/12, the conclusion inverted, and the scenario
was renamed -unreachable- to -reachable-. That finding was one of the TWO
that passed the reproduction rule. So three wrong premises have now
reached ground-game and the third satisfied an existence test -- existence
is not the property that was missing. The rule gains shape (ground-game's
row-level deal table, promoted from a T04 addendum) and a clause the
survey never contemplated: a reproduction must be able to fail. Ours went
green and stayed admissible.

C1 also caught a defect in flight. T06's payload, status todo, still named
4/6/9 and was queued to send it to ground-game as "no dataset reconciles
them." Withdrawn before sending -- the fourth wrong premise, and the only
one stopped.

C2 withdraws the baseline's precision: design-baseline.py is a
hand-maintained dict counting itself, has_reproduction never checks the
file exists (its YES-control is green against a deleted path), and
Makefile:127 runs only --self-test so the reporting path has no CI. The
direction survives; 33% is not a measured rate and T05 must not build on
it. C3: "six provisional defaults" is five, GR-E01 double-counted. C4:
GroundRules §Underdetermined was never evaluated as a candidate and
already delivers four of five benchmarks -- T03's burden flips to arguing
extension over replacement.

Survived: the rule's affordability, and reuse of the provisional
machinery.

loop-lint: no findings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 15:03:12 +02:00
129ed03492 Sync hub state for CB-WP-0024/0025
Workplan ids and task ids written back by fix-consistency C-06;
WORK-RECORDS.md regenerated.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 13:00:02 +02:00
14 changed files with 1750 additions and 82 deletions

View file

@ -5,7 +5,7 @@ assimilates and optimizes techniques and implementations useful for games,
simulations, and robotics.
It is not another monolithic game engine. It is a capability-assimilating
development engine with three distinct properties:
development engine with four distinct properties:
1. **Clay** — its canonical models, contracts, rules, and tools remain malleable.
2. **Borg** — mature, optimized libraries are assimilated behind controlled
@ -13,6 +13,14 @@ development engine with three distinct properties:
3. **Product-driven evolution** — abstractions are extracted from working
games, beginning with **GROUND — A Game of Bonds and Rivalry: DARVO
Edition**, rather than invented in isolation.
4. **Instrument** — the engine is rigorous enough that it cannot proceed
past a rule that does not decide. What it cannot execute, it reports:
findings about the *game's* design are a product of building the
simulator, not a side activity, and they are carried back to the game's
owner with the artifact that produced them. *(ADR-0012. A restatement of
what has already happened six times, made a duty. If a pass ever
tolerates an undecided rule by quietly picking a default and not raising
it, this property is false.)*
The central rule:

View file

@ -125,6 +125,7 @@ self-tests:
$(PY) $(TOOLS)/runtime-metrics.py --self-test
$(PY) $(TOOLS)/replay-test.py --self-test
$(PY) $(TOOLS)/design-baseline.py --self-test
$(PY) $(TOOLS)/design.py --self-test
$(PY) $(TOOLS)/edition-check.py --self-test
# T01 positive control: prove the environment fix, do not assume it. Runs
@ -160,6 +161,13 @@ facts-check:
facts-gen:
$(PY) $(TOOLS)/facts.py --gen
# CB-WP-0022 T05: the design-finding register, reported over
# specs/GroundRules.md. Shows the QUEUE by default; the log of closed
# findings is a line, not a listing, because a default view that mixes
# them loses the queue property (ADR-0012 D5).
design:
@$(PY) $(TOOLS)/design.py
# T03: one-shot orientation — workplans, next task, spend, fast gates.
# Cheap by design: no build. Start a session with this instead of grepping.
status:

View file

@ -31,6 +31,8 @@
| workplan | CB-WP-0021 | active | — | workplans/CB-WP-0021-import-the-edition.md |
| workplan | CB-WP-0022 | active | — | workplans/CB-WP-0022-the-design-instrument.md |
| workplan | CB-WP-0023 | done | — | workplans/CB-WP-0023-solve-legality.md |
| workplan | CB-WP-0024 | ready | — | workplans/CB-WP-0024-the-table-you-can-watch.md |
| workplan | CB-WP-0025 | ready | — | workplans/CB-WP-0025-could-we-have-won.md |
| task | CB-WP-0001-T01 | done | — | workplans/CB-WP-0001-inner-loop.md |
| task | CB-WP-0001-T02 | done | — | workplans/CB-WP-0001-inner-loop.md |
| task | CB-WP-0001-T03 | done | — | workplans/CB-WP-0001-inner-loop.md |
@ -154,3 +156,15 @@
| task | CB-WP-0023-T01 | done | — | workplans/CB-WP-0023-solve-legality.md |
| task | CB-WP-0023-T02 | done | — | workplans/CB-WP-0023-solve-legality.md |
| task | CB-WP-0023-T03 | done | — | workplans/CB-WP-0023-solve-legality.md |
| task | CB-WP-0024-T01 | todo | — | workplans/CB-WP-0024-the-table-you-can-watch.md |
| task | CB-WP-0024-T02 | todo | — | workplans/CB-WP-0024-the-table-you-can-watch.md |
| task | CB-WP-0024-T03 | todo | — | workplans/CB-WP-0024-the-table-you-can-watch.md |
| task | CB-WP-0024-T04 | todo | — | workplans/CB-WP-0024-the-table-you-can-watch.md |
| task | CB-WP-0024-T05 | todo | — | workplans/CB-WP-0024-the-table-you-can-watch.md |
| task | CB-WP-0025-T01 | todo | — | workplans/CB-WP-0025-could-we-have-won.md |
| task | CB-WP-0025-T02 | todo | — | workplans/CB-WP-0025-could-we-have-won.md |
| task | CB-WP-0025-T03 | todo | — | workplans/CB-WP-0025-could-we-have-won.md |
| task | CB-WP-0025-T04 | todo | — | workplans/CB-WP-0025-could-we-have-won.md |
| task | CB-WP-0025-T05 | todo | — | workplans/CB-WP-0025-could-we-have-won.md |
| task | CB-WP-0025-T06 | todo | — | workplans/CB-WP-0025-could-we-have-won.md |
| task | CB-WP-0025-T07 | todo | — | workplans/CB-WP-0025-could-we-have-won.md |

View file

@ -0,0 +1,273 @@
# ADR-0012: clay-borg is a design instrument, and §Underdetermined is the register
status: accepted
date: 2026-08-05
decided by: agent, under the standing loop authorization
tier: L (structural L — named a high-leverage pass by the maintainer and it
amends INTENT; chaos d8=6 → no override). Tier L requires a separate
survey, an adversarial review, then this document.
references: [CB-WP-0022](../workplans/CB-WP-0022-the-design-instrument.md),
[CB-RES-0007](../research/CB-RES-0007-design-instrument.md),
[challenge](../history/260805-design-instrument-challenge.md) /
[response](../history/260805-design-instrument-response.md),
[ADR-0007](ADR-0007-render-html-not-a-port.md) D3,
GROUND-WP-0002, GROUND-WP-0004 T02 (ground-game's shape ruling)
## Context
Nineteen passes of building a rigorous simulator produced findings about
*the game's rules* as a by-product, because an engine that refuses to
proceed past an undecided rule cannot help but locate the undecided rules.
Six such findings exist. They live in eleven files, no index, and one sat
unread in an inbox for four days.
**The review changed what this ADR has to decide.** The survey proposed a
new register governed by an existence rule — *a finding is not admissible
without its reproduction*. C1 falsified the example that rule was built
on, and C4 pointed out the incumbent was never evaluated. Both are
load-bearing here, so both are answered before anything is built.
## The correction this ADR is built on
`CB-WP-0022:61` justified the rule with *"GR-E01 is admissible because
4/6/9 against 5/7/9 is a computation anyone can rerun."* Commit `2da19a4`
had already rerun it: **6/9/12 against 5/7/9**, and the scenario was
renamed `-unreachable-``-reachable-`. The conclusion inverted.
That finding was **one of the two that passed the reproduction rule**. So:
| wrong premise | reached ground-game? | passed an existence test? |
|---|---|---|
| *"12 in the file"* — a sum with no deal table | yes | — |
| *"SOLVE offered on a face-down Problem"* — wrong condition named | yes | — |
| *"4/6/9, no dataset reconciles them"* | **stopped in flight, 2026-08-05** | **yes** |
**Existence is not the property that was missing.** Three findings went
wrong on their *description*, and the one that carried an artifact carried
a green one that no longer matched its prose. This ADR's rule is written
against that, not against the survey's.
---
## D1 — INTENT gains a fourth property: the engine is an instrument
`INTENT.md` states three properties — Clay, Borg, product-driven
evolution. It gains a fourth, and the wording is deliberately about a
**mechanism**, not an ambition:
> 4. **Instrument** — the engine is rigorous enough that it cannot proceed
> past a rule that does not decide. What it cannot execute, it reports:
> findings about the *game's* design are a product of building the
> simulator, not a side activity, and they are carried back to the game's
> owner with the artifact that produced them.
**Why this is earned rather than aspirational.** It is a restatement of
something that has already happened six times without anyone deciding it
should. The alternative wording considered and rejected — *"clay-borg is a
game design tool"* — claims a product we do not have and invites building
one; this claims a property we already exhibit and makes it a duty.
**Falsifier, and it is real:** if a pass ships with the engine having
tolerated an undecided rule by quietly picking a default and not raising
it, property 4 is false and INTENT is lying. `provisional: true` plus its
age report is what keeps it honest, and D2 puts the duty where that
machinery already lives.
## D2 — `§Underdetermined` is the register. Nothing parallel is built.
C4 lands. `specs/GroundRules.md §Underdetermined` was never evaluated as a
candidate, and against CB-RES-0007's own five benchmarks it already
delivers four:
| benchmark | §Underdetermined today |
|---|---|
| findability — one register with an index | **yes**, U1U10 in one section |
| closure — the source changes, not an annotation | **yes**: *"a ruling flips the scenario, not the kernel"* (`:231-233`) |
| time raised → read | **owner and raise date per item**, `make coverage` warns past 30 days |
| taxonomy fits reality | untested either way |
| reproducibility — 100% carry a replayable artifact | **no. This is the gap.** |
The survey went to Magic to discover the Oracle property. We had written
it down ourselves eight days earlier, in our own words. **A second
register would immediately disagree with the first**, and this project has
a standing rule that a gate must cash out.
**Decision: extend, do not replace.** §Underdetermined becomes the finding
register for all five kinds in D4, not only for dataset ambiguities. Its
existing machinery — owner, raise date, `provisional: true`, the age
warning — is the mechanism, unchanged. What is added is D3's admissibility
rule, D4's `kind`, and D5's state.
**What this costs, stated plainly:** the register lives in a spec file
about GROUND's rules, so a finding about the *engine's* behaviour sits in
a document about the *game*. That is tolerable because every finding in
scope is about the game by definition (D7), and it is the price of not
running two registers that drift.
## D3 — admissibility: exists, has the ruled shape, and can fail
The survey's rule was one clause. It is now three, and **all three are
required**:
> **1. It exists.** A scenario, an arithmetic check that prints the
> contradiction, a recorded game the reader can replay, or a named test.
>
> **2. It has the ruled shape.** Per GROUND-WP-0004 T02, an arithmetic
> finding ships a **row-level table** — Surface and each hidden priority
> listed separately — **never a sum, never "deal depth N"**. Promoted here
> from the T04 addendum it was, because two of three wrong premises were
> quantities summed without a table.
>
> **3. It can fail.** The artifact must be capable of going red, and its
> current colour is recorded. A reproduction that has gone **green is not a
> closed finding** — it is a finding whose claim is now in question, and it
> must be resolved to `applied` or `withdrawn` by a human reading it.
Clause 3 is the one the review bought. GR-E01's scenario went green when
the edition landed, and the finding stayed admissible and stayed queued for
transmission, because nothing in the rule said a passing artifact was a
signal. **A green reproduction is an alarm, not a reassurance.**
**Falsifier for D3 as a whole:** if a finding reaches `ground-game` whose
prose describes a quantity its own artifact does not compute, D3 has
failed and the next revision needs the description generated *from* the
artifact rather than written beside it.
## D4 — taxonomy, grounded in the six we have
Five kinds, each with the existing finding that forces it. **No kind is
included that no finding of ours instantiates** — that is the whole
discipline, and T05's backfill is its test.
| kind | means | our instance |
|---|---|---|
| `underdetermined` | the rules do not say | U1U10 |
| `inconsistent` | the rules disagree with each other or with the dataset | GR-E01 as raised (now `withdrawn`) |
| `inert` | a rule that cannot fire | SOLVE on a face-down Problem |
| `degenerate` | fires, but collapses play | GR-A13 wasted SOLVE |
| `unplayed` | implemented, never played to the end | GR-E03 / GR-E04 |
If backfill needs a sixth kind, **the taxonomy was invented rather than
observed** and T07 must say so.
## D5 — lifecycle: a queue that empties, a log that does not
```
raised ──► reported ──► ruled ──► applied
└──────────► withdrawn ◄──────────┘
```
| state | means | owner |
|---|---|---|
| `raised` | admissible under D3, in clay-borg | **clay-borg** |
| `reported` | delivered to ground-game as a file under their workplan | **clay-borg** |
| `ruled` | ground-game has answered | **ground-game** |
| `applied` | **the source changed** — dataset or rules text — and our provisional default was deleted | **clay-borg** |
| `withdrawn` | the finding was wrong, or dissolved | whoever finds it wrong |
`applied` means the source changed. It does **not** mean a ruling was
recorded — Magic's Oracle property, and CB-RES-0007 §2 is right about it.
**The queue is `raised` + `reported`; the log is `applied` + `withdrawn`.**
C6 asked whether keeping resolved findings contradicts "a queue that
empties." It does not: the open queue empties, the closed log accumulates
so a close rate can be computed. **Falsifier: if the default view shows
closed findings mixed with open ones, the queue property is gone** — so
`make design` reports the queue, and the log only on request.
**`withdrawn` is reported, not deleted.** GR-E01 is the reason: a claim
retracted silently is how the first three survived. When a finding is
withdrawn after being reported, the withdrawal goes downstream too.
## D6 — a finding without a reproduction is a `note`, and a note is not reportable
The survey left this open (§6) and the honest options were refusal or a
note tier. **Notes are admitted, under three constraints that make the
tier costly to sit in:**
1. A note **may not be reported to ground-game**. It has no `reported`
state. This is the whole point: the unread-inbox failure was prose
crossing a repo boundary.
2. A note **expires at 30 days** and `make design` flags it, reusing
§Underdetermined's existing age mechanism rather than adding one.
3. A note that is promoted acquires a reproduction meeting D3 — there is
no path from `note` to `reported` that skips it.
**Why not refusal.** A finding from *play* — the maintainer's *"I felt it
was too easy but then we lost"* — is real, has no cheap artifact, and is
the seed of CB-WP-0025. Refusing it would throw away the only class of
finding the engine cannot produce by itself. **Why not silence.** A note
that cannot expire ages into an apparent finding, which is the failure
this register exists to stop.
## D7 — no engine-evolution register, on a corrected inventory
The declaration's judgment survives; **its argument did not, and the
difference is recorded rather than smoothed over.** My inventory named
`evidence/`, `decisions/`, `gates.toml` and workplans. C5 found two places
it missed — `specs/InnerLoopReference.md`, 383 lines with sections headed
*"moved from InnerLoop v1.6"*, and 769 lines of `history/` retrospectives.
So the "third thing" the maintainer asked about **does exist and did find
homes**, somewhat by accident.
Decision: **no new register**, because a third place to look for material
currently in two makes it less findable, not more. This is argued and
narrowed, **not settled** — if InnerLoopReference keeps absorbing content
that is neither a decision nor a finding, that is the signal to revisit.
## D8 — `tools/design-baseline.py` is retired, not extended
C2 dismantled it: a hand-maintained dict counting itself (`:16-36`, `:89`),
a `has_reproduction` that never checks the file exists (`:38-43`) so its
own YES-control is green against a path `2da19a4` deleted (`:63`), date
literals for the 4-day figure (`:96-98`), and `Makefile:127` running only
`--self-test` so the reporting path has no CI at all.
**The 33% is withdrawn as a measured rate.** What survives is direction:
eleven files with no index and 0 of 10 U-items ruled are both checkable
without the tool.
`make design` (T05) reports **over the register**, which is real data with
real files behind it, and its reproduction check **stats the file**. The
baseline stays in the repo as the survey's dated snapshot, marked
superseded — deleting it would erase the evidence for how the number got
in.
## D9 — where a finding lives, and where it is delivered
The survey did not settle this. **Both, with distinct roles:**
- **clay-borg holds the finding and its reproduction**, in
§Underdetermined. The artifact cannot leave — it only runs here.
- **ground-game receives a generated report file under its own workplan**,
not an inbox message. Per CB-WP-0022 T06: *the message that sat unread
for four days is the baseline to beat, and the failure was not the
message — it was that nothing pointed at it.*
The report **cites the reproduction by path**, which is the half of
GROUND-WP-0004 T02's ruling that is theirs to keep, and makes the split
survivable: they read a claim, we hold the thing that proves it.
## Consequences
- `INTENT.md` gains property 4 (D1) — the change with the longest
half-life in this pass.
- `specs/GroundRules.md §Underdetermined` becomes the register and gains
`kind` and `state` per item (D2, D4, D5); **no new file, no new schema.**
- T04's spec inherits D3 as a **rule**, not an addendum, and must say what
a reproduction has to *show*.
- T05 builds `make design` over the register and backfills all six. If a
sixth kind is needed, the taxonomy was invented.
- T06 reports the queue **and GR-E01's withdrawal**.
- Notes carry a 30-day expiry (D6) on the existing age machinery.
## What was rejected
| rejected | why |
|---|---|
| a new register file with its own schema | C4 — the incumbent already delivers 4 of 5 benchmarks |
| the existence-only admissibility rule | C1 — its own showcase finding passed it and was false |
| refusing findings without reproductions | throws away findings from play, which is CB-WP-0025's entire input |
| an engine-evolution register | D7 — a third home for material already in two |
| extending `design-baseline.py` | C2 — it counts itself and its positive control is green against a deleted file |

View file

@ -0,0 +1,217 @@
# CB-EV-0021 — the design instrument
CB-WP-0022 T07. Tier L, declaration 5 of chaos window 2 (structural L,
chaos d8=6 → no override). Closed 2026-08-05.
**Delivered:** INTENT property 4 (*Instrument*),
[ADR-0012](../decisions/ADR-0012-the-design-instrument.md) (nine
decisions), [specs/GameDesign.md](../specs/GameDesign.md) v1.0, the
register in [GroundRules.md](../specs/GroundRules.md) (14 rows),
`tools/design.py` + `make design`, and a report file in `ground-game`.
---
## 1. The headline: a separate reviewer found the survey's showcase was false
This is the first adversarial review in this repo **run by a genuinely
separate agent**. CB-RES-0006's reviewer opened by conceding it could not
be, and wrote:
> *"Treat the challenges below as a lower bound on what a genuinely
> separate reviewer would find."*
That was a testable claim. **This is the test.** The separate reviewer ran
`git log` against the survey's central example and found our own commit
had falsified it four days earlier — while the author, who wrote that
commit, quoted the dead number twice in the same workplan.
`CB-WP-0022:61` justified the whole proposal with *"GR-E01 is admissible
because 4/6/9 against 5/7/9 is a computation anyone can rerun."*
`2da19a4` had rerun it: **6/9/12**, scenario renamed `-unreachable-`
`-reachable-`, conclusion inverted. And GR-E01 was **one of the two
findings that passed the reproduction rule**.
**Verdict on same-session review: it is not a weak form of the control, it
is a different control.** Four of the seven challenges rest on facts a
same-session reviewer had already seen and not questioned. The cost was
~70k subagent tokens; it caught a defect in flight (§2) and rewrote the
pass's central rule.
## 2. What it caught in flight
T06's payload, status `todo`, still read *"the deal count puts 4/6/9
points in play … no dataset reconciles them"* and was queued for
transmission to `ground-game`.
It would have been the **fourth** wrong premise sent downstream, after
*"12 in the file"* (a sum with no deal table) and *"SOLVE offered on a
face-down Problem"* (the wrong condition named). **It is the only one
stopped before sending.**
| wrong premise | reached ground-game? | passed an existence test? |
|---|---|---|
| *"12 in the file"* | yes | — |
| *"SOLVE on a face-down Problem"* | yes | — |
| *"4/6/9, no dataset reconciles them"* | **stopped 2026-08-05** | **yes** |
## 3. The rule the pass came to build was one clause short
The declaration's load-bearing rule was an **existence** test. Three of
three wrong premises would have passed it, and one did.
What shipped instead (GameDesign §1): a reproduction must **exist**, have
the **shape ground-game ruled** (row-level deal table, never a sum —
promoted from a T04 addendum to part of the rule), and **be able to
fail**.
**Clause 3 is the one this pass bought.** GR-E01's scenario went green
when the edition landed and the finding stayed admissible and stayed
queued, because nothing treated a passing artifact as a signal. *A green
reproduction is an alarm, not a reassurance.*
## 4. Building it produced a distinction the ADR does not have
`make design`'s first run alarmed on **U2** — and was wrong to.
U2's scenario is green **because the provisional default it documents is
implemented**, which says nothing about whether ground-game agrees.
GR-E01's was a *counterexample* that went green. **Same colour, opposite
meaning.**
So the register gained a `role` column — `counterexample` | `default`
and only a green counterexample alarms. **A register that cannot tell them
apart either alarms constantly or never**, and this was invisible until
the metric ran against real rows. Folded back into GameDesign §1.3.
## 5. Did backfilling change the design?
**T07 asked to be suspicious if all six findings fit the first taxonomy.**
They did — five kinds absorbed everything, no sixth invented. Two reasons
that is less reassuring than it looks and one reason it is fine:
- the taxonomy was **derived from these six**, so fitting them is nearly
circular. Its real test is the seventh finding;
- **one of the six was deleted during backfill.** *"Six provisional
defaults"* was not entered: C3 showed it double-counted GR-E01, and
provisional scenarios are *reproductions for* underdetermined items, not
a finding. So the taxonomy absorbed five, not six;
- what it did do is **force the note tier to earn itself** — GR-A13 and
GR-E03/E04 have no artifact and became notes, which is the first time
the distinction cost anything.
## 6. The numbers, and why none of the survey's survive
**CB-RES-0007's measurements are withdrawn.** C2 dismantled the
instrument: `design-baseline.py` is a hand-maintained dict where
`n = len(FINDINGS)`; `has_reproduction` is `startswith("scenarios/")` and
never stats the file, so its own YES-control was **green against a path
`2da19a4` deleted**; the 4-day figure is two date literals subtracted
(mutated to 2020-01-01 → 2406 days, all checks still green); and
`Makefile:127` ran only `--self-test`, so the reporting path had no CI.
It is **superseded, not deleted** — it is the evidence for how a wrong
number got into a survey, and the header now says so.
Measured by `make design`, 2026-08-05, over rows that name real files:
```
findings 12 (+2 notes)
with a resolving reproduction 3/12 = 25% target 100%
open, lacking a reproduction 9 target 0
reproductions green while open 0 target 0
notes past 30 days 0 target 0
closed (log) 2 [F11, F13]
```
**Two survey claims died against measurement:**
| survey said | measured |
|---|---|
| *"six of the ten already have provisional scenarios"* | **one.** `grep -lE "\bU<n>\b" scenarios/ground/*.yaml` → only U2 |
| *"2/6 = 33% reproduced"* | withdrawn; GR-E01 was double-counted and the detector was broken |
The first is the same defect class as the wrong premises — **an
uncheckable link asserted as a fact** — found inside the survey that
proposed the fix.
## 7. The register's first run found ten answers nobody had collected
The most useful thing `make design` did was not a metric.
**GROUND-WP-0002 T05 ruled all ten U-items on 2026-08-03 — every one
confirmed.** T03 confirmed five of the six provisional scenarios. The
workplan is `finished`.
CB-RES-0007 reported *"0 of 10 ruled"* on 2026-08-03, and the register
built this session initially recorded them as `reported`. **Both were two
days stale on the day they were written.**
**This is the unread-inbox failure running in the opposite direction.**
The pass was designed around *we send findings and nobody reads them*. The
symmetric failure — *they answer and we do not collect it* — was not
anticipated by the declaration, the survey, or the review, and the first
run of the instrument surfaced it immediately.
The items are `ruled`, not `applied`: the rulings confirmed our defaults,
so the rules did not move, but the scenarios still carry `provisional:
true` for settled choices. **Lifting those flags is owed and is not done**
`make design` shows them open until it is, which is the register
working.
## 8. The engine-evolution question, as the review left it
Declared judgment: no second register, because engine evolution already
lives in `evidence/`, `decisions/`, `gates.toml` and workplans.
**The conclusion survives; the argument did not.** C5 found the "third
thing" the maintainer meant is visible in `specs/InnerLoopReference.md`
(383 lines, two sections literally headed *"moved from InnerLoop v1.6"*)
and 769 lines of `history/` retrospectives — **neither of which the
inventory named**. So the content exists and found homes somewhat by
accident.
ADR-0012 D7 records it as **argued and narrowed, not settled**, with the
revisit condition: if InnerLoopReference keeps absorbing material that is
neither a decision nor a finding, build the register.
## 9. What tier L cost against what it caught
Second full-weight L pass in this window, third in the project.
| caught | by |
|---|---|
| the central example was false | **only** the separate reviewer |
| a wrong premise in flight to ground-game | the same |
| the baseline instrument counts itself | the same |
| the incumbent register was never evaluated | the same |
| `role`: green default ≠ green counterexample | **building it**, not reviewing it |
| ten uncollected rulings | **running it**, not building it |
**Two of six were caught by execution, not by process** — the pattern
InnerLoop §Design goal predicts (cheap correction beats exhaustive
prevention), and an argument against front-loading more review rather than
less.
**The one number that matters for the loop's own calibration:** the review
step is what distinguished this pass from a competent one. Without it the
pass ships a rule justified by a false example, sends a fourth bad premise
downstream, and builds `make design` on a dict that counts itself. The
correct reading of CB-RES-0006's fidelity note is **not** *"our reviews
were somewhat weaker"* — it is *"our reviews were not performing this
function at all."*
## 10. Cost
CB-WP-0021's cost, by re-running the instrument: see `make cost` /
`tools/cb-cost.py`. **Not quoted here as a literal** — §Single source of
fact, and this pass has just spent itself proving what typed numbers do.
## Open after this pass
- **U1U10 → `applied`**: lift `provisional: true` where a ruling confirmed
the default, and record each ruling. Nine items of bookkeeping.
- **The U-item ↔ scenario mapping is unwritten** for 9 of 10. Cheap, and
it is what makes the register checkable.
- **F12 and F14 are notes** and expire 2026-08-31. GROUND-WP-0003's
playtest closes F14 and produces its artifact via `cb-play --record`.
- **The taxonomy's real test is the seventh finding**, not these six.

View file

@ -0,0 +1,361 @@
# 260805 — challenge to CB-RES-0007
Adversarial review, one round, per InnerLoop §Step 2. Target: the survey
`research/CB-RES-0007-design-instrument.md`, the harness
`tools/design-baseline.py`, and the §judgment in `CB-WP-0022`.
**Fidelity note, first.** Unlike `260802-render-port-challenge.md`, this
review *was* run in a separate agent session with no access to the author's
reasoning trail — only the files. So it has the property §Step 2 asks for.
It does **not** have a different sample: there is exactly one baseline
harness and one repo, so every number below was re-derived from the
author's instrument. Per §Step 2 that is stated rather than reported as a
clean verify. What this review substitutes for a second sample is
**mutation**: each quoted number was traced to the assertion that produces
it and the assertion was inverted. Four of the five headline numbers do not
move when the world they claim to measure moves.
Ranked. C1 and C2 land hardest. C6 and C7 are marked weak — do not spend
the one response round on them.
---
## C1 — The rule's own showcase example is a false finding that shipped a green reproduction
**This is the challenge that lands, and it is worse than the two cases the
task description already knows about.**
`CB-WP-0022:61` states the admissibility rule and then gives its canonical
example:
> *"**GR-E01 is admissible because 4/6/9 against 5/7/9 is a computation
> anyone can rerun**; the SOLVE inertness is admissible because a recorded
> session shows three no-ops."*
Both halves of that sentence are now known to be wrong.
- The SOLVE half: `CB-WP-0023:87` — *"the reported case was not the one I
reported… `validate` already rejected face-down, so it never was
offered."* Already conceded by the author.
- **The GR-E01 half has not been conceded, and it is the more damaging
one.** `4/6/9` is stale. Commit `2da19a4` (CB-WP-0021, 2026-08-04)
measured **6/9/12** against thresholds 5/7/9, `ADR-0011:110` records the
ruled 6/9/12, and the scenario was renamed
`gr-e01-threshold-**unreachable**-2p.yaml`
`gr-e01-threshold-**reachable**-2p.yaml`. The finding did not get
refined; **its conclusion inverted**.
Now the part that breaks the proposal's central claim. GR-E01 is one of the
**two** findings in the survey's `2/6 = 33% with a runnable reproduction`
(`design-baseline.py:26-29`). It is a *success case* of the reproduction
rule. It carried a scenario. The scenario was green. And the scenario
asserted the wrong conclusion for four days, was reported to ground-game on
that basis, and was renamed only when someone re-did the arithmetic by
hand.
So the count is not two bad premises. **It is three, and the third one is
the exemplar the workplan uses to define what a good finding looks like.**
The survey's §3 argument — *"in formal methods a claim without a trace is
not a result, it is a hypothesis"* — is true and irrelevant here. A model
checker's counterexample is generated *by* the checker from the property;
it cannot disagree with the property, because it is the property's witness.
Our scenarios are **hand-authored assertions of an expected value**. A
hand-written scenario that encodes a wrong expectation is a trace that
witnesses the author's belief, not the system's behaviour. §3 imports the
prestige of counterexample traces onto an artifact class that does not have
their guarantee, and it does so in the section that supplies the whole
justification for the rule.
**Required:** the survey cannot claim the reproduction rule is the load-
bearing control while its own two positive instances of that rule both
described the wrong thing. Either (a) name the *additional* control that
would have caught GR-E01 — the ground-game row-level deal table
requirement is the obvious candidate and it is a **shape** requirement, not
an existence one — and demote "carries its reproduction" from *the* rule to
*half* the rule, or (b) withdraw §3's claim that a reproduction is
sufficient to make a finding admissible.
Secondary, and checkable: `4/6/9` is still live in three places —
`workplans/CB-WP-0022:61`, `workplans/CB-WP-0022:283` (task T06, still
`todo`, still planning to report the stale number to ground-game), and
`evidence/CB-EV-0018:153`. **The pass that exists to stop wrong premises
reaching ground-game currently has a wrong premise queued for transmission
to ground-game.** `grep -rn "4/6/9" --include=*.md .` settles it.
## C2 — The baseline is a hand-maintained list reporting on itself, and four of its five numbers are mutation-proof
The survey's §1 heading is *"The baseline is us, and **it is measured**"*.
It is not measured. It is a dict literal at `tools/design-baseline.py:16-36`
with one globbed row, printed back with arithmetic.
Per §Step 2 (v1.4), each number was traced to its assertion and mutated.
**(a) `findings 6`** is `len(FINDINGS)` (`:89`) — the number of keys typed
into a dict at `:16-36`. It is a count of hardcoded rows, not of findings.
The self-test that "pins" it (`:58`) asserts `len(FINDINGS) >= 6` against
the same literal. Nothing in the repo can change this number; only editing
line 16-36 can.
**(b) `2/6 = 33%`** rests on `has_reproduction()` (`:38-43`), which is a
**string prefix test on `p.startswith("scenarios/")`**. It never touches
the filesystem. Mutation:
```
has_reproduction(['scenarios/COMPLETELY-MADE-UP.yaml']) -> True
has_reproduction(['scenarios/']) -> True
has_reproduction(['crates/cb-engine/tests/foo.rs']) -> False
```
Two consequences. First, the self-test's YES-control at `:63` cites
`scenarios/ground/gr-e01-threshold-unreachable-2p.yaml` — **a file that no
longer exists** (renamed in `2da19a4`; `ls scenarios/ground/ | grep e01`).
It passes anyway, because the detector never looks. This is exactly the
render-port C6 class: a control that cannot distinguish "found a
reproduction" from "found a string". The `every cited location exists`
check at `:57` iterates `FINDINGS.values()` only and never sees the
self-test's own literal, so the one guard that would have caught it is
pointed at the wrong list.
Second, the detector **contradicts the rule it is measuring**. `CB-WP-0022:55-56`
admits *"a scenario that fails, an arithmetic check that prints the
contradiction, a recorded game the reader can replay, or **a named test**."*
`has_reproduction` recognises exactly one of those four. A finding backed by
a named Rust test scores 0. The 33% is therefore not the baseline for the
proposed rule; it is the baseline for a narrower rule nobody proposed, and
it is biased **downward**, which flatters the "nobody was ever asked for
one" story in §1.
**(c) `4 days`** is `datetime.date(2026, 8, 3) - datetime.date(2026, 7, 30)`
— two literals on `:96-97`, subtracted on `:98`. No data source. Mutated
`:96` to `2020-01-01`:
```
U1..U10: raised 2020-01-01, first READ 2026-08-03 — 2406 days
--self-test: all 5 checks [ok]
```
The headline number can be set to anything and every control stays green.
Worse, **it disagrees with the repo**: `specs/GroundRules.md:224` states
*"Owner: ground-game. **Raised: 2026-07-31**"*, and
`scenarios/ground/gr-a04-bond-support.yaml:9` carries
`provisional_raised: 2026-07-31`. Under the spec's own date the interval is
**3 days, not 4**. Two sources of the same fact, disagreeing, with zero
`<!-- fact: -->` tags in either the survey or the workplan
(`grep -c "fact:"` returns 0 for both) — the drift surface InnerLoop
§Single-source-of-fact names explicitly.
`4 days` is also the row the survey's §5 calls *"the number this exists to
fix; any mechanism that does not move it has failed"*. It is the least
defensible number in the document.
**(d) `11 files`** is `len(places)` — the deduplicated union of the
hardcoded paths. It is derived, so it is the healthiest of the four, but it
measures **citation sites the author typed**, not files holding findings.
Add a citation, the "findability" benchmark worsens; nothing about the
repo changed.
**(e)** The one number with real provenance is the 5-element glob at
`:30-34` — and see C3, it disagrees with the label above it.
**What the harness would report if the work silently stopped** (§Step 2
requires this be stated): the reporting path, `:80-99`, performs **no
existence check at all**. Delete every file in `FINDINGS` and it still
prints `6 findings, 2/6 = 33%, 11 files, 4 days`. The existence guard lives
only in `--self-test`, and `Makefile:127` runs **only** `--self-test`
the numbers quoted in the survey come from a code path CI never executes.
**Required:** either the survey stops calling §1 "measured" and labels the
baseline what it is — a hand-curated snapshot, which is a legitimate thing
to have but caps at the same directional status the survey correctly
assigns to the external rows — or the harness derives its finding set from
the repo (the `provisional:` marks, `evidence/` frontmatter, the U-list)
so that mutating the repo moves the number.
## C3 — "six provisional defaults" is five, and the survey's own instrument says so
`research/CB-RES-0007:46` lists `six provisional defaults | 6 | yes`.
```
$ python3 tools/design-baseline.py
repro 5 location(s) six provisional defaults
$ make coverage
provisional U-item defaults: 5
```
Six became five in `2da19a4` when `gr-e01-threshold-unreachable-2p.yaml`
lost its `provisional: true` (verified:
`git grep -l "provisional: true" 2da19a4^ -- scenarios/ground/ | wc -l`
**6**; at HEAD → **5**). The author edited `design-baseline.py:22-29` in
that same window to add the `[RESOLVED]` comment on the GR-E01 row, and did
not notice that the *next* row's glob had silently dropped a member and its
hardcoded label `"six provisional defaults"` had gone stale.
Two things follow.
1. **The survey's table row 5 (`locations: 6`) is a number the survey's own
tool contradicts on stdout**, and no self-test compares the word "six"
to `len()`. §4 of the survey argues the register *"must reuse [the
provisional machinery], not compete with it"* — and the very row that
reads the provisional machinery is the one that drifted from it.
2. **GR-E01 was double-counted.** It is row 4 *and* was a member of row 5's
six. The denominator of `2/6` contained the same finding twice, and one
of the two numerator hits was that finding.
**Required:** correct the row, and state whether the `2/6` survives the
double-count. If rows 4 and 5 overlap, the honest denominator at survey
time was 5 distinct findings, and the headline rate changes.
## C4 — `GroundRules.md §Underdetermined` was never evaluated as the incumbent, and it already implements four of the five benchmarks
T02 requires an attempt at *"whether a register is needed at all, rather
than one more section in `GroundRules.md §Underdetermined`."* The survey
does not attempt it. §4 mentions the provisional machinery only as
*machinery to reuse*; §5's benchmark table has a `today` column that never
names §Underdetermined; and §6's "did not settle" list does not include it.
Per §Step 2's row for surveys, **that is an omitted candidate, and it is
the strongest one, because it is the incumbent.**
Read against the survey's own five benchmarks (`specs/GroundRules.md:222-277`):
| benchmark (§5) | §Underdetermined today |
|---|---|
| findability: *one register with an index* | **has it** — one section, ten numbered items, `grep U7` |
| reproducibility: *runnable artifact* | **has the obligation***"every U-item's default must be covered by a scenario tagged `provisional: true`"* (`:274-277`) |
| closure: *closes when the source changed* | **has it***"A ruling flips the **scenario**, not the kernel; a correction is a data change"* (`:231-233`). This is the Magic Oracle property the survey went to Magic to discover. It has been in the spec since 2026-07-31. |
| raised → read, with age | **has it, executable** — owner + raise date per item, `make coverage` prints `age=5d` per item and warns past 30 days |
| taxonomy fits reality | not addressed — genuinely absent |
The survey travels to Magic: the Gathering to extract *"a finding closes
when the source changes, not when an annotation is added"* and presents it
as a corrected assumption (§2, and `CB-WP-0022:136-139` treats it as the
pass's main external yield). **The repo already states that property in
those terms, one directory over.** That is not a small omission: it is the
survey's headline external finding being rediscovered from the incumbent it
failed to survey.
**Where this challenge stops, honestly.** §Underdetermined has a real gap
and the review found it: it obliges ten U-items to carry provisional
scenarios and **five exist**. `make coverage` prints `provisional U-item
defaults: 5` and does *not* flag the five missing — it reports what exists,
never what is owed. That is a genuine argument for something with a
denominator, and it is stronger than any argument the survey actually
makes. It is also an argument for *one gate in `rule-coverage.py`*, not for
a new register.
**Required:** evaluate §Underdetermined as a candidate row with a `today`
column, and state what a register does that §Underdetermined + one
completeness check cannot. If the answer is "holds findings that are not
U-items" — say which of the six, and note that four of the six are
single-file prose notes in `evidence/`.
## C5 — The engine-evolution "third thing" exists, and its homelessness is visible in the filesystem
`CB-WP-0022:73-84` argues the engine-evolution register is redundant
because *"`evidence/CB-EV-*` records what each pass found, `decisions/ADR-*`
records what was decided…, `gates.toml` records what every control has
caught…, and workplans record what was attempted."*
The list is incomplete, and the omission is the whole question. The third
thing — *what we learned about building engines* — is
**`specs/InnerLoopReference.md`, 383 lines**, and its section list is
nothing but that:
```
§ Implementation rules the first pass earned
§ Why review is adversarial execution, not re-derivation
§ Numbers you inherit — two cases from CB-WP-0002
§ Quoting a cost — six for six, never high
§ Chaos roll — window 1's verdict and the d4 → d8 change
§ Retargeting — the case, and the four corrections it must not forbid
```
None of those is a decision (ADR), a per-pass finding (evidence), or a
control (gates.toml). They are cross-pass lessons about method. Plus six
`history/*-retrospective.md` files, 769 lines, in the same category.
**And the homelessness is on the record in the headings themselves.** Two
sections are literally titled *"(moved from InnerLoop v1.6)"* — the content
grew inside the normative spec until it had to be evicted into an
appendix-shaped file that is not in the author's list of homes. That is
precisely the signature T02 asked the reviewer to look for: *content that
had nowhere to go*.
**Where this challenge stops.** It does not establish that a *register* is
needed. It establishes that the author's redundancy argument is made
against the wrong inventory. `InnerLoopReference.md` and `history/` **are**
a home, so the conclusion "no second register" probably survives — but it
survives for a different reason than the one given, and the difference
matters for T03: the asymmetry claimed at `:83` (*"engine evolution has a
home and game design does not"*) is real, but the home is an unindexed
383-line spec appendix plus six retrospectives, i.e. **exactly the
`11 files, no index` condition the survey calls a defect when it is game
findings**. The survey applies one standard to game findings and another to
engine lessons and does not state the distinction.
**Required:** name `InnerLoopReference.md` and `history/*-retrospective.md`
in the redundancy argument, and say why "several files, no index" is
disqualifying for one subject and acceptable for the other.
## C6 — (weak) The survey's own §2 conclusion contradicts the harness's stated design
§2: *"A register that accumulates permanent rulings is a system that has
given up on fixing its rules. Ours should be a **queue that empties**, not
an archive that grows."*
`design-baseline.py:23-25`: *"Kept in the baseline because… a register that
drops findings when they close cannot report a close rate."*
Both are defensible and they are not strictly contradictory — a queue can
retain closed items out-of-band. But T03 has to pick, and the two artifacts
of this pass currently point opposite ways, with the tool's comment being
the more considered of the two. Marked **weak**: this is a clarification,
not a defect.
## C7 — (weak) `n=6`, all engine-surfaced, and §6 already concedes it
§6 concedes the sample is six and all engine-surfaced, and flags the
play-sourced finding (*"DARVO feels punishing at 2 players"*) as unsettled.
That concession is correct and I have nothing to add to it, so this is not
a challenge — it is noted only so the response does not spend a round
defending something already conceded. Marked **weak**.
---
## Verdict
**Not approvable as written.** C1, C2 and C3 each require a change to the
survey, not a clarification. C4 requires an added candidate row. C5
requires a corrected inventory in the workplan's §judgment.
| # | verdict |
|---|---|
| **C1** | **lands hardest.** Requires demoting "carries its reproduction" from sufficient to necessary-but-not-sufficient, and adding the shape requirement. GR-E01 is a rule-satisfying finding that was wrong. |
| **C2** | **lands.** Four of five headline numbers do not move under mutation; `4 days` is a literal and disagrees with `GroundRules.md:224` (3 days). §1's word "measured" must go or the harness must derive. |
| **C3** | **lands.** "six" is five; the survey's table contradicts its own tool's stdout; GR-E01 double-counted in the `2/6` denominator. |
| **C4** | **lands, moderate.** The incumbent was never surveyed, and it already holds the Magic Oracle property the survey travelled to Magic to find. Weakened by a real gap the review found for the author: 10 U-items owe scenarios, 5 exist. |
| **C5** | **partial.** The third thing exists and is visibly homeless; the redundancy *conclusion* probably survives, the *argument* does not. |
| **C6** | **weak.** Clarification for T03. |
| **C7** | **weak / already conceded.** |
**What survives.** Two claims attacked and not broken:
- **That the reproduction rule is affordable.** §1's four cheap
reproductions were checked and the estimates are plausible; `bot::play`
exists, the scenario harness exists, and the SOLVE case in fact got its
reproduction in CB-WP-0023 T01 for a fraction of a pass. What would have
falsified it: a finding among the four whose reproduction needed engine
work that does not exist. There is none. **Affordability is not the
problem; sufficiency is (C1).**
- **That the register must reuse the provisional machinery rather than
compete with it** (§4). Checked against `rule-coverage.py:98-108,285-308`
— owner, raise date, age and a warn threshold all already exist and are
executable. What would have falsified it: a second age-tracking mechanism
already in the repo. There is none.
**The single challenge that forces a change to the design: C1.** The
proposal's load-bearing rule is stated as an existence test on a
reproduction. Three findings have now reached the record on wrong premises,
and **the third one satisfied the existence test**. The rule as written
does not catch its own worst observed failure. T03 cannot adopt it in the
form `CB-WP-0022:53` states it.

View file

@ -0,0 +1,187 @@
# 260805 — response to the challenge on CB-RES-0007
One round, per InnerLoop §Step 2. **This review was run by a separate
agent**, unlike CB-RES-0006's, whose challenge opened with a fidelity note
saying it could not be. That note said *"treat the challenges below as a
lower bound on what a genuinely separate reviewer would find."* It was
right, and this is what the difference looks like: the separate reviewer
went to `git log` and found that the survey's showcase example had been
falsified four days earlier by our own commit.
Four challenges conceded, two conceded in part, one answered. **The
load-bearing rule survives but does not survive intact** — it is
insufficient as stated, and C1 is the reason.
---
## C1 — the showcase finding was itself a wrong premise, and it passed the rule — **conceded, and it is larger than the challenge says**
The challenge is correct and I verified every step of it independently
rather than accepting it.
`CB-WP-0022:61` defines admissibility by example:
> *"GR-E01 is admissible because 4/6/9 against 5/7/9 is a computation
> anyone can rerun."*
Commit `2da19a4` (2026-08-04, CB-WP-0021 T01/T02/T05) measured the deal
after the edition import: **6/9/12 against thresholds 5/7/9 — the game is
reachable at every seat count.** The scenario was renamed
`gr-e01-threshold-unreachable-2p.yaml``gr-e01-threshold-reachable-2p.yaml`.
The finding did not merely get a new number. **Its conclusion inverted.**
So the count is not two wrong premises reaching `ground-game`. It is
three, and **the third one was in the 2-of-6 set that passed the
reproduction rule.** The rule's own success case is a false finding that
shipped a green artifact.
**And it was still loaded.** `CB-WP-0022:283`, task T06, status `todo`:
> *"GR-E01 vs GR-S01 — the deal count puts 4/6/9 points in play against
> thresholds of 5/7/9, so either the count or the thresholds are wrong and
> no dataset reconciles them"*
That is queued for transmission to `ground-game` as a finding. Had T06 run
before this review, it would have sent a **fourth** wrong premise
downstream — one asserting no dataset reconciles numbers our own dataset
reconciled on 2026-08-04. The review did not catch a stale citation. It
caught a defect in flight.
**What this forces.** The rule at `CB-WP-0022:53` is an *existence* test:
a finding is admissible if a reproduction exists. Three of three wrong
premises would have passed an existence test, and one did. Existence is
therefore not the property that was missing. T03 must adopt the shape
requirement — `ground-game`'s row-level deal table, GROUND-WP-0004 T02 —
as **part of the rule**, not as the T04 addendum it currently is. And it
needs a second clause the survey did not contemplate: **a reproduction
must be able to fail.** A finding whose artifact has gone green is not a
closed finding until someone says which it is; ours silently became a
counterexample to itself.
`4/6/9` is corrected wherever it appears in this pass's artifacts, and
`evidence/CB-EV-0018:153` is left alone — it is a dated record of what was
believed then, and rewriting evidence is worse than a stale number in it.
## C2 — the baseline is a hand-maintained inventory reporting on itself — **conceded**
`tools/design-baseline.py:16-36` is a literal dict. `n = len(FINDINGS)`
(`:89`). **"6 findings" is a count of rows someone typed.** It is not a
measurement of anything, and the survey presented it as a measured
baseline, in a repo whose §Single source of fact rule exists precisely to
stop typed numbers.
Worse, and the part I would not have found:
- `has_reproduction` (`:38-43`) is `p.startswith("scenarios/")` and
nothing else. **It never checks the file exists.** The self-test's
YES-control (`:63`) passes a path — `gr-e01-threshold-unreachable-2p.yaml`
that `2da19a4` deleted. The positive control for "detection can say
YES" is green against a file that is not there. That is a positive
control that controls nothing, which is the exact class `gates.toml`
exists to catch.
- The same detector **rejects a named test**, which `CB-WP-0022:55-56`
explicitly admits as a reproduction. So `33%` is biased down by the
detector's own narrowness.
- The 4-day figure is two date literals subtracted (`:96-98`). The
reviewer mutated them to 2020-01-01, got **2406 days**, and all five
self-test checks stayed green.
- `Makefile:127` runs `--self-test` only. **The reporting path is never
executed by CI**, so none of the above was ever going to be caught.
**What survives, and it matters that something does.** The *direction* of
every headline number is checkable without the tool: 11 files with no
index is verifiable by reading, and U1U10 being open with 0 of 10 ruled
is verifiable against the hub. The baseline's conclusion — *we have
findings scattered across the repo with no register and no reproductions* —
is not in doubt. **What is withdrawn is its precision**, and any claim
that the 33% is a measured rate.
T05 must not build on this tool. The register itself is the measurement,
and `make design` reports over the register's contents, not over a dict.
## C3 — the arithmetic in the survey's own table — **conceded**
"Six provisional defaults" is **five**. The tool's own line prints
`5 location(s)`; `make coverage` prints 5; the survey table (`:46`) says
6. It was 6 at `2da19a4^` and 5 at HEAD — the survey table was written
against a repo state that had already moved.
And **GR-E01 is double-counted**: it is row 4 in its own right and was a
member of row 5's six. So the `2/6` denominator contains one finding
twice, and both copies are in the numerator. The real figure, on the
survey's own definitions, is not 2/6.
I am not restating a corrected rate here. **The tool that would produce it
is the one C2 just withdrew**, and quoting a new number from a discredited
instrument is how the first one got in.
## C4 — §Underdetermined was never evaluated as a candidate — **conceded in part**
Correct, and it is the omission that embarrasses the survey most, because
§Step 1 asks for the best implementation *in existence* and this one is in
our own repo. The survey went to Magic to discover the Oracle property —
*the fix lands in the source, the annotation is not authoritative* — and
`specs/GroundRules.md:231-233` already states it in our own words:
*"a ruling flips the scenario, not the kernel."*
**Conceded in part, not whole,** and the reviewer marked the honest limit
themselves: 10 U-items owe provisional scenarios, 5 exist, and
`make coverage` never flags the missing five. A section that cannot report
its own gaps is not yet the register. But T03's burden has changed
direction — it must now argue why §Underdetermined should be **extended**
rather than replaced, and the default answer is that it should be.
## C5 — the third thing exists — **conceded as to the argument, not the conclusion**
`specs/InnerLoopReference.md` is 383 lines with two sections literally
headed *"moved from InnerLoop v1.6"*, and `history/` holds 769 lines of
retrospectives. That is content about *how we learn to build this engine*
which was placed where it fit rather than where it belonged. My redundancy
inventory named `evidence/`, `decisions/`, `gates.toml` and workplans, and
**named neither of those** — so the inventory was incomplete and the
argument as written does not stand.
The conclusion probably does, and for a reason the challenge concedes: the
homeless content found homes, and both are load-bearing and read. A third
register would be a third place to look for material that is currently in
two. **T03 records this as argued-and-narrowed rather than settled**, which
is the outcome the declaration asked for when it said this was the
judgment it wanted reviewed rather than assumed.
## C6 — "queue that empties" vs keeping resolved findings — **answered**
Marked weak by the reviewer and it is weak, but it has a real answer.
Magic's property is that a finding **closes when the source changes**. It
is not that the record is deleted. `design-baseline.py:23-25`'s comment —
keep resolved findings or you cannot report a close rate — is compatible:
the *queue* of open findings empties, the *log* of closed ones does not.
T03's lifecycle should say this in one sentence so the next reader does
not re-litigate it.
Falsifier: if the register's default view shows closed findings mixed with
open ones, the queue property is gone regardless of what the lifecycle
says.
## C7 — n=6, all engine-surfaced — **already conceded in §6**
The survey states this limit itself. Nothing to add.
---
## What the review changed
| | before | after |
|---|---|---|
| the rule | reproduction must **exist** | must exist, **have the ruled shape**, and **be able to fail** |
| wrong premises shipped | 2 | **3** — and one passed the rule |
| the baseline | a measured 33% | a hand-maintained inventory; direction stands, precision withdrawn |
| §Underdetermined | not considered | the incumbent T03 must argue against |
| engine-evolution register | redundant | still redundant, on a corrected inventory |
| T06's payload | a finding | **a fourth wrong premise, caught before sending** |
**The strongest thing here is not any single challenge.** It is that a
separate reviewer with no stake ran `git log` against the survey's central
example and found it had been falsified by our own commit four days
earlier — while the author, who wrote that commit, quoted the dead number
twice. CB-RES-0006's review said a same-session reviewer would be a lower
bound. This is how much lower.

183
specs/GameDesign.md Normal file
View file

@ -0,0 +1,183 @@
# GameDesign — findings about the game, and the trials that produce them
v1.0 — CB-WP-0022 T04, 2026-08-05. Normative. Implements
[ADR-0012](../decisions/ADR-0012-the-design-instrument.md); INTENT
property 4 (*Instrument*) is what this spec discharges.
**This is not a register.** The register is
[`GroundRules.md §Underdetermined`](GroundRules.md) — ADR-0012 D2. This
document says what may go in it, what a reproduction must show, how a
finding moves and dies, and how a trial game is run so *"we played it and
X happened"* is replayable rather than remembered.
---
## 1. Admissibility
> A finding is admissible when its reproduction **exists**, **has the
> ruled shape**, and **can fail**. All three. (ADR-0012 D3.)
### 1.1 Exists
One of: a scenario under `scenarios/ground/`, an arithmetic check that
prints the contradiction, a `.cbreplay` bundle or recorded scenario the
reader can replay, or a **named test**`crate::module::test_name`, so it
can be run.
The register entry names the artifact **by path**. A finding whose
artifact path does not resolve is not admissible, and §4's check stats the
file rather than pattern-matching the string — `design-baseline.py`'s
`has_reproduction` matched `startswith("scenarios/")` and its own positive
control was green against a file that had been deleted (ADR-0012 D8).
### 1.2 Has the ruled shape
Per `ground-game`, GROUND-WP-0004 T02:
> An arithmetic finding ships a **row-level table** — Surface and each
> hidden priority listed **separately** — never *"sum of file"*, never
> *"deal depth N"*.
**A reproduction must show the quantity the prose claims.** This is the
clause that exists because passing artifacts accompanied wrong
descriptions:
| wrong premise | what shipped | what a row-level table would have shown |
|---|---|---|
| *"12 in the file"* | a sum | which rows, at which priorities |
| *"4/6/9 against 5/7/9"* | a scenario, green | 6/9/12 — the deal, per row |
| *"SOLVE offered on a face-down Problem"* | a condition named | which filter actually fired |
**Requirement.** Where a finding asserts a number, the register entry
quotes the number **and** the artifact prints the rows it came from. A
finding stating a total without its rows is inadmissible even if the total
is right — *"12"* was arithmetically defensible and still wrong about the
game.
### 1.3 Can fail
The artifact must be capable of going red, and the register records its
**current colour**. A reproduction that has gone **green is not a closed
finding** — it is a finding whose claim is now in question, and §4 reports
it as needing resolution to `applied` or `withdrawn`.
**A reproduction has a `role`, and only one of them alarms.** This
distinction was not in ADR-0012; **T05's backfill produced it**, when the
first report cried wolf over U2.
| role | what it is | green means |
|---|---|---|
| `counterexample` | demonstrates the defect — the model-checker shape | **alarm.** It should be red while the finding is open |
| `default` | encodes a provisional choice so it is testable and visible | **expected.** It passes because our default is implemented, which says nothing about whether the owner agrees |
U2's scenario is green because the default it documents is implemented.
GR-E01's was a counterexample that went green when the edition landed, and
sat that way for four days. **Same colour, opposite meaning** — a register
that cannot tell them apart either alarms constantly or never.
**Rationale, and it is the expensive one.** `gr-e01-threshold-…-2p.yaml`
went green when the edition landed on 2026-08-04. The finding remained
admissible, remained in the register, and remained queued for transmission
to `ground-game` for four days, because nothing treated a passing artifact
as a signal. It was stopped by an adversarial review, not by a control.
## 2. Kinds
Five (ADR-0012 D4). **No kind exists without a finding that instantiates
it**; a sixth appearing during backfill means the taxonomy was invented
rather than observed, and T07 must say so.
`underdetermined` · `inconsistent` · `inert` · `degenerate` · `unplayed`
## 3. States, owners, and death
```
raised ──► reported ──► ruled ──► applied
└──────────► withdrawn ◄──────────┘
```
- `applied` means **the source changed** — dataset or rules text — and our
provisional default was deleted. Not *"a ruling was recorded."*
- **The queue is `raised` + `reported`. The log is `applied` +
`withdrawn`.** `make design` reports the queue by default; the log only
on request. If the default view mixes them, the queue property is gone.
- **`withdrawn` is reported downstream, not deleted.** A claim retracted
silently is how three wrong premises survived.
### 3.1 Notes
A finding without a reproduction is a **note** (ADR-0012 D6). A note:
1. **has no `reported` state** — it may not cross the repo boundary;
2. **expires at 30 days**, flagged by §4 on the same mechanism that ages
provisional items (`rule-coverage.py:54`, `PROVISIONAL_WARN_DAYS`);
3. is promoted only by acquiring a reproduction meeting §1 — there is no
path from `note` to `reported` that skips it.
## 4. Metrics
`make design` reports these. Every number is computed **over the
register**, never from a maintained list (ADR-0012 D8).
| metric | target | notes |
|---|---|---|
| findings with a resolving reproduction | **100%** | denominator includes `withdrawn`; notes counted separately, never folded in |
| **reproductions that have gone green while open** | **0** | the §1.3 alarm. Non-zero is a queue that needs reading, not a passing suite |
| arithmetic findings whose artifact prints rows | **100%** | §1.2 |
| median days `raised``reported` | the number this exists to fix | the baseline it must beat is **4 days** to first *read* |
| open findings by kind, with age | reported, not targeted | |
| notes past 30 days | **0** | |
| findings closed by a ruling vs still open | reported, with age | |
**Falsifier for this spec.** If a finding reaches `ground-game` whose
prose describes a quantity its own artifact does not compute, §1.2 has
failed and the next revision must **generate the description from the
artifact** rather than let it be written alongside.
**No baseline rate is quoted here.** CB-RES-0007's 33% was withdrawn as a
measured figure (ADR-0012 D8); the first honest denominator is the one
T05's backfill produces.
## 5. The trial protocol
The part with no precedent in this repo. A **trial game** is a recorded
session plus an observation log — so a claim from play is replayable.
**It must cost almost nothing or it will not be done.** The engine already
records: `cb-play --record FILE` writes the finished game as a scenario
YAML (`main.rs:30`), and `games/ground/src/record.rs` is trustworthy
because its round trip is the test.
A trial is therefore:
```
cb-play --players N --serve PORT --record trials/<date>-<slug>.yaml
```
plus a sibling `trials/<date>-<slug>.md` carrying, and nothing more:
- **who played which seats**, and which were bots (with policy);
- **what was observed** — in the player's own words, not cleaned up;
- **the seed**, which the scenario already carries;
- **which observations became findings or notes**, by register id, or
explicitly none.
**Rules.**
1. An observation is a **note** until it has a reproduction (§3.1). A
trial log is not a finding source by itself — it is where findings are
*noticed*.
2. **The recording is the artifact for `unplayed` findings.** GR-E03/E04
were "never played to the end"; a trial that plays them to the end both
closes that finding and produces the reproduction for anything found on
the way.
3. A trial whose recording does not replay is a **bug in the engine**,
reported as one, not a design finding.
**Why the log is prose and stays prose.** The maintainer's *"I felt it was
too easy but then we lost, so who knows"* is exactly the observation this
protocol exists to catch: real, unreproducible as stated, and the seed of
[CB-WP-0025](../workplans/CB-WP-0025-could-we-have-won.md). Forcing it
into a schema at the moment of observation would lose it. §3.1 is what
stops it aging into an apparent finding.

View file

@ -219,6 +219,103 @@ terms.
---
## The finding register
**This section is the design-finding register** (ADR-0012 D2). It was the
register for dataset ambiguities already; CB-WP-0022 extended it to all
five kinds rather than building a second one beside it. Admissibility,
kinds, states and metrics: [`GameDesign.md`](GameDesign.md). Reported by
`make design`.
<!-- design-register:begin -->
| id | kind | state | reproduction | role | raised | owner |
|---|---|---|---|---|---|---|
| U1 | underdetermined | ruled | — | default | 2026-07-31 | ground-game |
| U2 | underdetermined | ruled | scenarios/ground/gr-d01-darvo-trigger.yaml | default | 2026-07-31 | ground-game |
| U3 | underdetermined | ruled | — | default | 2026-07-31 | ground-game |
| U4 | underdetermined | ruled | — | default | 2026-07-31 | ground-game |
| U5 | underdetermined | ruled | — | default | 2026-07-31 | ground-game |
| U6 | underdetermined | ruled | — | default | 2026-07-31 | ground-game |
| U7 | underdetermined | ruled | — | default | 2026-07-31 | ground-game |
| U8 | underdetermined | ruled | — | default | 2026-07-31 | ground-game |
| U9 | underdetermined | ruled | — | default | 2026-07-31 | ground-game |
| U10 | underdetermined | ruled | — | default | 2026-07-31 | ground-game |
| F11 | inert | applied | scenarios/ground/gr-p05-solve-legality.yaml | counterexample | 2026-08-02 | clay-borg |
| F12 | degenerate | note | — | — | 2026-08-01 | clay-borg |
| F13 | inconsistent | withdrawn | scenarios/ground/gr-e01-threshold-reachable-2p.yaml | counterexample | 2026-08-01 | clay-borg |
| F14 | unplayed | note | — | — | 2026-08-01 | clay-borg |
<!-- design-register:end -->
- **F11 — SOLVE offered where it cannot act.** Offered on a face-down
Problem, or with no matching suit in hand; inert every time. Ruled
GROUND-WP-0002 T02, implemented CB-WP-0023 as GR-P05. `applied` — the
rule changed, not just the annotation. **The case we reported was not the
case that fired**: `validate` already rejected face-down, and the
maintainer's three inert SOLVEs were the *hand* case.
- **F12 — GR-A13 "wasted SOLVE" on an already-claimed Problem.** A
scenario had to pick a default and did. `note`: no artifact isolates the
degenerate line, so under GameDesign §3.1 it may not be reported until
one exists.
- **F13 — GR-E01 vs GR-S01, withdrawn 2026-08-05.** Raised as *"4/6/9
against 5/7/9, no dataset reconciles them."* `2da19a4` measured **6/9/12
against 5/7/9** and the scenario was renamed `-unreachable-`
`-reachable-`. **Its reproduction is green**, which under GameDesign §1.3
is the alarm that forced the resolution. Withdrawn rather than deleted,
and the withdrawal is reported (ADR-0012 D5).
- **F14 — GR-E03/GR-E04 never played to the end.** Nineteen passes, never
played out. `note` until a trial game exists; GROUND-WP-0003 is the
playtest that would close it, and GameDesign §5's protocol makes the
recording the artifact.
### The register's first run found ten answers nobody had collected
**U1U10 are `ruled`, not `reported`.** GROUND-WP-0002 T05 answered all
ten on **2026-08-03** — every one **confirmed** as the default clay-borg
already simulates — and GROUND-WP-0002 T03 confirmed five of the six
provisional scenarios, voiding `gr-e01` as a rules gap. The workplan is
`finished`.
CB-RES-0007 reported *"0 of 10 ruled"* and this register was built saying
`reported`. **Both were two days stale on the day they were written.** The
answers had arrived and nothing propagated them — the same failure as the
unread inbox, in the opposite direction.
**They are `ruled`, not `applied`, and the difference is work we owe.**
Per ADR-0012 D5, `applied` means the source changed and the provisional
default was deleted. The rulings *confirmed* our defaults, so the rules did
not move — but the scenarios still carry `provisional: true` for choices
that are now settled. **Lifting those flags and recording each ruling is
what closes U1U10**, and it is not done. `make design` shows them open
until it is.
### What the backfill measured, and what it contradicted
**Only U2 names its U-item in a scenario.** Measured, not estimated:
```
for u in U1 .. U10; do grep -lE "\b$u\b" scenarios/ground/*.yaml; done
```
CB-RES-0007 asserted *"six of the ten already have provisional
scenarios."* Five provisional scenarios exist and **one** cites the item
it stands for. The other four may well encode U-item defaults — the
mapping is simply **not written down**, so it is not checkable, and an
uncheckable link is the defect this register exists to fix. The register
records what is citable; the rest is debt, visible in `make design`.
**No sixth kind was needed** — the five kinds absorbed all four non-U
findings. And the survey's *"six provisional defaults"* was **not** entered
as a finding: C3 showed it double-counted GR-E01, and the provisional
scenarios are *reproductions for* underdetermined items, not a finding of
their own.
**U1U10 are `reported` while lacking reproductions**, which GameDesign
§3.1 would now forbid. They were reported on 2026-07-30, before the rule
existed. They are grandfathered rather than rewritten, and the debt is a
reported metric with a target of zero.
## Underdetermined in dataset 0.1 — PROVISIONAL defaults (flag to ground-game)
**Owner: ground-game. Raised: 2026-07-31.** Every U-item carries an owner

View file

@ -7,6 +7,26 @@ model-checker traces, W3C provisional marks) are practices rather than
runnable software, so per InnerLoop Step 1 their rows are DIRECTIONAL and
cap at `parity`. This is the row that can be measured.
"""
#
# ── SUPERSEDED 2026-08-05 by tools/design.py (ADR-0012 D8) ──────────────
#
# Kept as CB-RES-0007's dated snapshot, NOT as a live instrument. Deleting
# it would erase the evidence for how a wrong number got into a survey.
#
# The adversarial review (C2) dismantled it and every point is true:
# * FINDINGS below is a hand-maintained dict; `n = len(FINDINGS)` counts
# the rows someone typed, not findings that exist.
# * `has_reproduction` is `startswith("scenarios/")` and never stats the
# file, so the self-test's YES-control is GREEN against a path that
# 2da19a4 deleted. A positive control that controls nothing.
# * It rejects a named test, which the admissibility rule admits, so its
# 33% is biased downward.
# * The 4-day figure is two date literals subtracted. Mutated to
# 2020-01-01 it reports 2406 days with every check still green.
#
# **Its 33% is withdrawn as a measured rate.** What survived: 11 files,
# no index, 0 of 10 ruled — all checkable without this tool.
#
import os, re, subprocess, sys, datetime
ROOT = "/home/worsch/clay-borg"

194
tools/design.py Normal file
View file

@ -0,0 +1,194 @@
#!/usr/bin/env python3
"""design — report the finding register (CB-WP-0022 T05).
Implements the metrics in `specs/GameDesign.md` §4 over the register in
`specs/GroundRules.md`. ADR-0012 D8 retired `design-baseline.py`, which
was a hand-maintained dict counting itself; the difference that matters is
that **every number here is computed over rows that name real files, and
the reproduction check stats the file.**
`design-baseline.py`'s `has_reproduction` was `p.startswith("scenarios/")`
and nothing else, so its own positive control was green against a path
that had been deleted. The self-test below asserts the opposite property
directly: a row citing a nonexistent file must NOT count as reproduced.
"""
import os, re, sys, subprocess, datetime
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
REGISTER = os.path.join(ROOT, "specs", "GroundRules.md")
BEGIN = "<!-- design-register:begin -->"
END = "<!-- design-register:end -->"
KINDS = {"underdetermined", "inconsistent", "inert", "degenerate", "unplayed"}
OPEN_STATES = {"raised", "reported", "ruled"} # the queue: not yet applied
CLOSED_STATES = {"applied", "withdrawn"} # the log
NOTE = "note"
NOTE_EXPIRY_DAYS = 30 # GameDesign §3.1, same figure as rule-coverage.py
def parse(text):
"""Rows between the register markers. Raises if the block is absent —
a register that silently reports zero findings is worse than one that
fails."""
try:
block = text.split(BEGIN)[1].split(END)[0]
except IndexError:
raise SystemExit(f"no design-register block in {REGISTER}")
rows = []
for line in block.splitlines():
line = line.strip()
if not line.startswith("|") or line.startswith("|---"):
continue
cells = [c.strip() for c in line.strip("|").split("|")]
if len(cells) != 7 or cells[0] in ("id",):
continue
rows.append(dict(zip(
("id", "kind", "state", "repro", "role", "raised", "owner"), cells)))
return rows
def reproduced(row, root=ROOT):
"""GameDesign §1.1: the artifact must resolve. A named test is admitted
by its `crate::module::name` shape; anything else must be a real path
on disk, and that is checked by stat, not by prefix."""
p = row["repro"]
if p in ("", "", "-"):
return False
if "::" in p: # a named test
return True
return os.path.exists(os.path.join(root, p))
_SIM = None
def sim_passes(root=ROOT):
"""`make sim` is the authority on whether scenarios pass. Run once and
cache asking per row costs a full suite per finding."""
global _SIM
if _SIM is None:
r = subprocess.run(["make", "sim"], cwd=root, capture_output=True, text=True)
_SIM = r.returncode == 0
return _SIM
def alarming(row, root=ROOT):
"""§1.3, and the distinction the backfill discovered: **only a
counterexample alarms when green.**
A `default` reproduction encodes a provisional choice and is SUPPOSED
to pass U2's scenario is green because the default it documents is
implemented, which says nothing about whether ground-game agrees. A
`counterexample` is the model-checker shape: it demonstrates the
defect, so it must be red while the finding is open. GR-E01's went
green when the edition landed and nothing noticed for four days.
"""
if row["role"] != "counterexample":
return False
if not row["repro"].startswith("scenarios/"):
return False
return sim_passes(root)
def report(root=ROOT, today=None):
today = today or datetime.date.today()
rows = parse(open(os.path.join(root, "specs", "GroundRules.md")).read())
findings = [r for r in rows if r["state"] != NOTE]
notes = [r for r in rows if r["state"] == NOTE]
queue = [r for r in findings if r["state"] in OPEN_STATES]
closed = [r for r in findings if r["state"] in CLOSED_STATES]
print("design — the finding register\n")
print(" QUEUE (open findings)")
for r in sorted(queue, key=lambda r: r["raised"]):
age = (today - datetime.date.fromisoformat(r["raised"])).days
mark = "repro" if reproduced(r, root) else " - "
print(f" {mark} {r['id']:<4} {r['kind']:<16} {r['state']:<9} {age:>4}d {r['owner']}")
if notes:
print("\n NOTES (not reportable — GameDesign §3.1)")
for r in sorted(notes, key=lambda r: r["raised"]):
age = (today - datetime.date.fromisoformat(r["raised"])).days
flag = " EXPIRED" if age > NOTE_EXPIRY_DAYS else ""
print(f" {r['id']:<4} {r['kind']:<16} {age:>4}d{flag}")
# ---- §4 metrics. The log is reported on request, not by default,
# because a default view mixing open and closed loses the queue
# property (ADR-0012 D5).
n = len(findings)
repro = sum(1 for r in findings if reproduced(r, root))
debt = [r for r in queue if not reproduced(r, root)]
expired = [r for r in notes
if (today - datetime.date.fromisoformat(r["raised"])).days > NOTE_EXPIRY_DAYS]
unresolved_green = [r for r in queue if reproduced(r, root) and alarming(r, root)]
print(f"\n findings {n} (+{len(notes)} note(s))")
print(f" with a resolving reproduction {repro}/{n}"
f"{'' if not n else f' = {100*repro//n}%'} target 100%")
print(f" open, lacking a reproduction {len(debt)} target 0"
+ (f" [{', '.join(r['id'] for r in debt)}]" if debt else ""))
print(f" reproductions green while open {len(unresolved_green)} target 0"
+ (" <-- ALARM, GameDesign §1.3" if unresolved_green else ""))
print(f" notes past {NOTE_EXPIRY_DAYS} days {len(expired)} target 0")
print(f" closed (log) {len(closed)}"
f" [{', '.join(r['id'] for r in closed)}]" if closed else "")
bad = [r for r in rows if r["kind"] not in KINDS]
if bad:
print(f"\n UNKNOWN KIND: {', '.join(r['id'] for r in bad)}"
" — a sixth kind means the taxonomy was invented (ADR-0012 D4)")
return 0
def self_test():
ok = True
def check(name, cond, detail=""):
nonlocal ok
ok = ok and bool(cond)
print(f" {'ok ' if cond else 'FAIL'} {name}{' ' + detail if detail else ''}")
rows = parse(open(REGISTER).read())
check("the register parses", len(rows) >= 14, f"{len(rows)} row(s)")
check("every kind is one of the five",
all(r["kind"] in KINDS for r in rows),
"a sixth kind means the taxonomy was invented")
check("every state is known",
all(r["state"] in OPEN_STATES | CLOSED_STATES | {NOTE} for r in rows))
# THE control design-baseline.py did not have. Its YES-control passed a
# path that had been deleted and still returned True.
check("a nonexistent reproduction does NOT count",
not reproduced({"repro": "scenarios/ground/gr-e01-threshold-unreachable-2p.yaml"}),
"this exact path was deleted by 2da19a4 and the old tool said yes")
check("a real reproduction DOES count",
reproduced({"repro": "scenarios/ground/gr-p05-solve-legality.yaml"}))
check("a named test counts", reproduced({"repro": "games_ground::view::tests::a_spectator_sees_no_hands"}))
check("an em-dash does not count", not reproduced({"repro": ""}))
# The distinction the backfill discovered: a green DEFAULT is expected,
# a green COUNTEREXAMPLE is the alarm. Without this the report cried
# wolf over U2, whose scenario is green precisely because the
# provisional default it documents is implemented.
check("a green default does not alarm",
not alarming({"role": "default", "repro": "scenarios/ground/gr-d01-darvo-trigger.yaml"}))
check("role is recorded for every reproduced row",
all(r["role"] in ("counterexample", "default")
for r in rows if r["repro"] not in ("", "-", "")))
# The measured claim in the backfill: only U2 cites its U-item.
cited = [u for u in [f"U{i}" for i in range(1, 11)]
if subprocess.run(["grep", "-rlE", rf"\b{u}\b"] +
[os.path.join(ROOT, "scenarios", "ground")],
capture_output=True, text=True, cwd=ROOT).stdout.strip()]
check("exactly one U-item is cited by a scenario", cited == ["U2"], f"{cited}")
print("design self-test (positive control)")
return 0 if ok else 1
if __name__ == "__main__":
sys.exit(self_test() if "--self-test" in sys.argv else report())

View file

@ -2,7 +2,7 @@
id: CB-WP-0022
kind: product
title: "The design instrument: findings about the game, with their reproductions"
status: active
status: done
state_hub_workstream_id: "fda16340-0049-4acf-884b-a5cfdbde47c0"
---
@ -36,8 +36,8 @@ has produced, as a by-product of being rigorous:
| ten underdetermined rules points (U1U10) | formalizing the dataset into testable rules |
| SOLVE offered on a face-down Problem, always inert | a human dragging it three rounds running |
| GR-A13 "wasted SOLVE" on a claimed Problem | a scenario that had to pick a default |
| GR-E01 unreachable below 5 seats | arithmetic over the deal count |
| six provisional scenario defaults | scenarios that could not be written without deciding something |
| ~~GR-E01 unreachable below 5 seats~~ | arithmetic over the deal count**withdrawn 2026-08-05, it was wrong (C1)** |
| ~~six provisional scenario defaults~~ | **five, and double-counted with GR-E01 (C3)** — they are reproductions, not a finding |
| GR-E03 / GR-E04 never played to the end | nobody noticed for nineteen passes |
Every one was found by *building the simulator*, not by playing. That is
@ -57,9 +57,17 @@ register that collects opinions would reproduce it in a new medium.
Concretely: a scenario that fails, an arithmetic check that prints the
contradiction, a recorded game the reader can replay, or a named test.
*"This feels unbalanced"* is a note, not a finding. **GR-E01 is admissible
because 4/6/9 against 5/7/9 is a computation anyone can rerun; the SOLVE
inertness is admissible because a recorded session shows three no-ops.**
*"This feels unbalanced"* is a note, not a finding. **The SOLVE inertness
is admissible because a recorded session shows three no-ops.**
> **The example that stood here was GR-E01, and the review killed it
> (C1).** *"4/6/9 against 5/7/9 is a computation anyone can rerun"* had
> already been rerun: `2da19a4` measured **6/9/12**, and the scenario was
> renamed `-unreachable-``-reachable-`. The conclusion inverted — and
> it was one of the two findings that **passed** this rule. So existence
> is not what was missing. See [ADR-0012](../decisions/ADR-0012-the-design-instrument.md)
> D3: the rule gains **shape**, and **a reproduction must be able to
> fail.** Ours went green and stayed admissible.
This is what would make clay-borg a design tool rather than a suggestion
box, and it is the one part of this proposal that must not be traded away
@ -71,19 +79,19 @@ The maintainer asked whether this should extend to *"a meta about the
clay-borg engine evolution itself."*
**My answer is no, and it should be argued rather than accepted.** That
register already exists and is load-bearing: `evidence/CB-EV-*` records
what each pass found, `decisions/ADR-*` records what was decided and what
was rejected, `gates.toml` records what every control has caught and what
would retire it, and workplans record what was attempted. Nineteen passes
of engine evolution are already captured, with dates, costs and
falsifiers.
register already exists and is load-bearing: `evidence/CB-EV-*`,
`decisions/ADR-*`, `gates.toml` and the workplans capture nineteen passes
with dates, costs and falsifiers. A second register for the same subject
would be ceremony. The asymmetry is the point: engine evolution has a home
and game design does not.
**Building a second register for the same subject would be ceremony**, and
this project has a standing rule that a gate must cash out. The asymmetry
is the point: engine evolution has a home and game design does not.
If the adversarial review disagrees, that is exactly the kind of thing
tier L exists to surface.
> **Settled by [ADR-0012](../decisions/ADR-0012-the-design-instrument.md)
> D7: no register — but the argument above did not survive.** C5 found the
> "third thing" the maintainer meant is visible in
> `specs/InnerLoopReference.md` and `history/`'s retrospectives, **neither
> of which this inventory names.** Conclusion narrowed, not settled: if
> InnerLoopReference keeps absorbing material that is neither a decision
> nor a finding, revisit.
## Task: survey how this is done elsewhere, and what we already have
@ -120,34 +128,27 @@ and whether a ruling can *close* a finding mechanically.
[CB-RES-0007](../research/CB-RES-0007-design-instrument.md), with a
runnable baseline (`tools/design-baseline.py`).
**The baseline is us, and it is measured**: 6 findings across **11 files**
with no index, **2 of 6 (33%) with a runnable reproduction**, and U1U10
raised 2026-07-30, first *read* 2026-08-03 — **4 days**, 0 of 10 ruled.
**Its numbers were withdrawn by T02 and must not be quoted from here.**
The survey reported *6 findings, 2/6 = 33% reproduced, 11 files, 4 days*.
C2 showed the instrument counted itself and its reproduction check never
stat'd the file; C3 showed *"six provisional defaults"* is five and GR-E01
was double-counted; T05's backfill contradicted *"six of the ten have
provisional scenarios"* — **one** does. What survives is direction: many
files, no index, 0 of 10 ruled. The first honest figures are T05's.
**The uncomfortable number is stated up front rather than left for the
review to find: the reproduction rule would reject four of our six
existing findings.** The survey answers it — none of the four is
expensive to reproduce, so the 33% is evidence that nobody was ever asked
for one, not that the rule is unaffordable.
**Magic corrected an assumption this pass was about to build on.** I
expected a ruling to be the authoritative resolution. It is not: rulings
are *"reminder information with no actual weight or rules meaning"*, and
the authoritative fix folds into the **Oracle** card text. **A finding
closes when the source changes, not when an annotation is added** — so
the register must be a queue that empties, not an archive that grows.
That is now a constraint on T03's lifecycle.
Model checkers supplied the reproduction rule independently (a
counterexample trace *is* the finding), and W3C's *implementation-defined*
mark is the one piece of machinery we already have and must reuse rather
than duplicate.
**Magic corrected an assumption this pass was about to build on.** Rulings
are *"reminder information with no actual weight or rules meaning"*; the
authoritative fix folds into the **Oracle** card text. **A finding closes
when the source changes, not when an annotation is added** — the register
is a queue that empties. Model checkers supplied the reproduction rule
independently, and W3C's *implementation-defined* mark is machinery we
already have and must reuse rather than duplicate.
## Task: adversarial review
```task
id: CB-WP-0022-T02
status: todo
status: done
priority: high
state_hub_task_id: "d2597895-fe2e-4f1e-a3db-a5fb8833434c"
```
@ -161,23 +162,42 @@ above, and require an attempt at:
third thing;
- **that "carries its reproduction" is affordable** — if half the real
findings cannot be reproduced cheaply, the rule will be quietly dropped
and the register becomes a suggestion box anyway. *(Since this task was
written the answer has hardened: two findings reached ground-game on
wrong premises, and ground-game has itself ruled that reproductions are
required. The reviewer should now press the opposite way — whether the
rule is **sufficient**, given both bad premises would have shipped a
passing test alongside a wrong description.)*
and the register becomes a suggestion box anyway. *(Hardened before the review ran: two findings had already reached
ground-game on wrong premises, so the reviewer was told to press whether
the rule is **sufficient**, not whether it is affordable.)*
- **that a register is needed at all**, rather than one more section in
`GroundRules.md §Underdetermined`, which already exists and already
works.
Record the trail in `history/`, unpolished.
**Done 2026-08-05.** Trail:
[challenge](../history/260805-design-instrument-challenge.md),
[response](../history/260805-design-instrument-response.md).
**Run by a separate agent** — the first in this repo that was. CB-RES-0006's
review opened by conceding it could not be, and called its own findings
*"a lower bound on what a genuinely separate reviewer would find."* That
was measurable, and this is the measurement: the separate reviewer ran
`git log` against the survey's central example and found our own commit
had falsified it four days earlier, while the author — who wrote that
commit — quoted the dead number twice.
**Seven challenges: four conceded, two conceded in part, one answered.**
**C1 changed the design** — the rule's showcase finding was false and had
*passed* the rule, so existence is not what was missing — and **caught a
defect in flight**, T06's payload still naming the dead number. C2
withdrew the baseline's precision, C3 its arithmetic, C4 flipped T03's
burden toward extending `§Underdetermined`, C5 corrected the redundancy
inventory. Survived: affordability, and reuse of the provisional
machinery. Full account:
[CB-EV-0021](../evidence/CB-EV-0021-the-design-instrument.md) §1§3.
## Task: decide
```task
id: CB-WP-0022-T03
status: todo
status: done
priority: high
state_hub_task_id: "a2e85810-e949-4ea0-81fc-36b912af326c"
```
@ -198,11 +218,35 @@ state_hub_task_id: "a2e85810-e949-4ea0-81fc-36b912af326c"
as a note** — and if admitted, how it is prevented from aging into an
apparent finding.
**Done 2026-08-05.**
[ADR-0012](../decisions/ADR-0012-the-design-instrument.md), nine
decisions. The two not on this list are the two the review forced:
- **D2 — `§Underdetermined` *is* the register; nothing parallel is built.**
Against the survey's own five benchmarks the incumbent already delivers
four, including the Oracle property the survey went to Magic to find and
we had written ourselves eight days earlier (`GroundRules.md:231-233`).
What it lacks is reproductions. So this pass **extends** a section — no
new file, no new schema.
- **D3 — admissibility is three clauses.** Exists, has the ruled shape
(row-level table, never a sum), **and can fail.** GR-E01's artifact went
green and the finding stayed admissible and stayed queued, because
nothing said a passing artifact was a signal. **A green reproduction is
an alarm.**
The rest, in one line each: **D1** INTENT gains property 4, *Instrument*,
applied with its falsifier. **D4** five kinds, each forced by an existing
finding. **D5** `applied` means the source changed; withdrawals are
reported, not deleted. **D6** notes admitted but never reportable, 30-day
expiry. **D7** no engine-evolution register, on an inventory C5 corrected.
**D8** `design-baseline.py` retired. **D9** the artifact stays here,
ground-game gets a generated file under its own workplan.
## Task: specify
```task
id: CB-WP-0022-T04
status: todo
status: done
priority: high
state_hub_task_id: "60ddfeab-81b1-45e4-9ec9-91cc8ea7fb72"
```
@ -220,34 +264,40 @@ Candidate measures, to be argued not adopted:
**ground-game has ruled on what a finding must carry** (GROUND-WP-0004
T02, 2026-08-04), and it is stricter than this pass proposed. Adopt it:
> 1. Arithmetic findings must ship a **runnable reproduction** *and* a
> **row-level deal table** listing Surface and each hidden priority
> separately — **never only "sum of file" or "deal depth N"**.
> 2. ground-game rulings that depend on arithmetic **should cite** that
> reproduction id or path, so a later reader can re-run it.
> Arithmetic findings ship a **runnable reproduction** *and* a
> **row-level deal table** — never only "sum of file" or "deal depth N";
> and ground-game's arithmetic rulings cite that reproduction by path.
The second half is theirs to keep. The first is ours, and it is a direct
consequence of both premises we got wrong: *"12 in the file"* was a sum
without a deal table, and *"offered on a face-down Problem"* was a
condition named without checking which one actually fired.
The second half is theirs to keep. **So the reproduction rule gains a
shape requirement, not just an existence one** — a finding that ships a
passing test but describes the wrong quantity is still a bad finding,
which is exactly what happened twice.
**So the reproduction rule gains a shape requirement, not just an
existence one.** A finding that ships a passing test but describes the
wrong quantity is still a bad finding — which is exactly what happened
twice. The spec must say what a reproduction has to *show*, not only that
one exists.
Also specify the **trial protocol**: a trial game is a `--record`ed
session plus an observation log, so *"we played it and X happened"* is
replayable rather than remembered. It must cost almost nothing or it will
not be done.
Also specify the **trial protocol**, which is the part with no precedent
here: a trial game is a `--record`ed session plus an observation log, so
*"we played it and X happened"* is replayable rather than remembered. The
engine already records sessions as scenarios; a trial is that plus notes,
and it must cost almost nothing or it will not be done.
**Done 2026-08-05.** [specs/GameDesign.md](../specs/GameDesign.md) v1.0 —
not a register (ADR-0012 D2 put that in `§Underdetermined`).
**§1.2 is written against evidence rather than principle**: a finding must
print the rows behind any number it claims. *"12" was arithmetically
defensible and still wrong about the game.* **§1.3's target is `0`
reproductions gone green while open** — what GR-E01 would have tripped
four days before a human caught it. **No baseline rate is quoted.**
**The trial protocol costs one flag**: `cb-play --record` plus a sibling
`.md` in the player's own words. An observation is a **note** until it has
a reproduction — *"I felt it was too easy but then we lost"* is the case
it is shaped around, and a schema at the moment of observation would lose
it.
## Task: build it, and backfill what is already known
```task
id: CB-WP-0022-T05
status: todo
status: done
priority: high
state_hub_task_id: "469413ba-ed47-4840-a60e-ee7f95dc06ff"
```
@ -260,36 +310,63 @@ designing it is the point of doing it in this order.
`make design` (or equivalent) must report: open findings by kind, those
without a reproduction, and those never reported to their owner.
**Done 2026-08-05.** `tools/design.py`, `make design`, and the register in
[`GroundRules.md`](../specs/GroundRules.md) — **14 rows, no new file.**
Backfill was the test. The taxonomy held (five kinds, no sixth), and it
**produced a `role` column ADR-0012 does not have**: the first report
alarmed on U2, wrongly — a green *default* is expected, a green
*counterexample* is the alarm. Folded into GameDesign §1.3. It also
contradicted the survey: **one** U-item names itself in a scenario, not
six. Detail and figures:
[CB-EV-0021](../evidence/CB-EV-0021-the-design-instrument.md) §4, §6.
## Task: report to ground-game, mechanically
```task
id: CB-WP-0022-T06
status: todo
status: done
priority: high
state_hub_task_id: "c6a659bd-c495-4754-9332-baadad51012a"
```
Generate the report and send it. **The message that sat unread for four
days is the baseline to beat**, and the failure was not the message — it
was that nothing pointed at it and nothing tracked whether it was
answered.
days is the baseline to beat** — the failure was not the message, it was
that nothing pointed at it. So the report lands as a file in `ground-game`
under its own workplan, extending GROUND-WP-0002 rather than duplicating
it.
So the report must land somewhere that persists: a file in `ground-game`
under its own workplan, not only an inbox entry. GROUND-WP-0002 already
holds the ten U-items; this should extend it rather than duplicate it.
Include the findings this pass has sharpened:
Include the two sharpened findings this pass has already produced:
- **SOLVE's legality** against a face-down Problem or an unmatchable suit —
and note that the case we *reported* was not the case that fired
(CB-WP-0023 T01).
- ~~**GR-E01 vs GR-S01** — 4/6/9 against 5/7/9~~ — **withdrawn
2026-08-05, before sending.** C1 found `2da19a4` had measured **6/9/12**:
the dataset reconciles them. It would have been the **fourth** wrong
premise to reach `ground-game` and is the only one caught before
transmission. **Report the withdrawal** — a claim retracted silently is
how the first three survived.
- **GR-E01 vs GR-S01** — the deal count puts 4/6/9 points in play against
thresholds of 5/7/9, so either the count or the thresholds are wrong and
no dataset reconciles them;
- **SOLVE's legality** against a face-down Problem or an unmatchable suit.
**Done 2026-08-05.**
[`ground-game/workplans/GROUND-WP-0002-clay-borg-report-260805.md`](../../ground-game/workplans/GROUND-WP-0002-clay-borg-report-260805.md),
committed there, with a hub message that only *points at* the file.
**The report asks for no ruling.** It carries GR-E01's withdrawal, our own
reproduction debt, and two notes that are explicitly not findings.
**And it acknowledged something the pass did not expect.**
GROUND-WP-0002 is `finished`: **all ten U-items were ruled 2026-08-03**,
every one confirmed as the default we simulate. CB-RES-0007 said *"0 of 10
ruled"* two days later. **The unread-inbox failure running in the opposite
direction** — they answered and we did not collect it. The instrument's
first run surfaced it.
## Task: evidence
```task
id: CB-WP-0022-T07
status: todo
status: done
priority: high
state_hub_task_id: "4fb77911-486d-4d9e-a53d-c4e486c74ce2"
```
@ -304,3 +381,18 @@ this one moves rather than collides.)*
full-weight L pass and CB-WP-0012's deleted its own structural trigger.
- **The engine-evolution question**, as the review left it.
- **Quote CB-WP-0021's cost by re-running the instrument.**
**Done 2026-08-05.**
[CB-EV-0021](../evidence/CB-EV-0021-the-design-instrument.md).
**Backfill did change the design** — and the honest answer to *"be
suspicious if all six fit"* is that only **five** were entered (one was a
double-count), so fitting them is close to circular. The taxonomy's real
test is the seventh finding.
**Tier L's cost against what it caught**: four of six catches came only
from the separate reviewer, and **two came from execution rather than
process** — the `role` distinction from building it, the ten uncollected
rulings from running it. That is InnerLoop §Design goal's prediction
holding, and an argument against front-loading more review rather than
less.

View file

@ -3,6 +3,7 @@ id: CB-WP-0024
kind: product
title: "The table you can watch: the piles, the other seats' moves, the score that carries"
status: ready
state_hub_workstream_id: "5f17b6f9-cd4b-4c31-a8af-712313149cf2"
---
# Purpose
@ -48,6 +49,7 @@ except where task T04 says so explicitly.
id: CB-WP-0024-T01
status: todo
priority: high
state_hub_task_id: "acb4231c-35df-490f-93fd-be71c0abf1dc"
```
Two defects behind one button.
@ -79,6 +81,7 @@ because a control that can no longer work must not look like it can.
id: CB-WP-0024-T02
status: todo
priority: high
state_hub_task_id: "37b16a87-e791-4c85-9f60-671812f6dd2e"
```
Draw the draw stack and the discard stack as stacks, with their counts on
@ -110,6 +113,7 @@ stop.
id: CB-WP-0024-T03
status: todo
priority: high
state_hub_task_id: "7b0b08d9-cf0d-40f4-8697-049bdb20085d"
```
Each seat's revealed selection should be legible as a played card in that
@ -139,6 +143,7 @@ face-down as one identical back.
id: CB-WP-0024-T04
status: todo
priority: medium
state_hub_task_id: "4f312267-2b82-4cf9-bbb1-7f287dc9e766"
```
`play` (`table.rs:219`) already owns the session — one listener, a seed
@ -168,6 +173,7 @@ raise the question rather than silently canonising a choice.
id: CB-WP-0024-T05
status: todo
priority: high
state_hub_task_id: "db176595-e00d-467c-ba54-0714cfb7ba5c"
```
`evidence/CB-EV-0022-*.md`. Short — this is an S pass.

View file

@ -3,6 +3,7 @@ id: CB-WP-0025
kind: product
title: "Could we have won: a path out of a lost game, and how hard the game actually is"
status: ready
state_hub_workstream_id: "a866982f-94e7-432b-b0d5-2fefb781a574"
---
# Purpose
@ -91,6 +92,7 @@ whose meaning drifts the next time a bot improves.
id: CB-WP-0025-T01
status: todo
priority: high
state_hub_task_id: "556cfd24-5992-4cc0-be90-0b60e989b5bb"
```
`research/CB-RES-0008-*.md`, with `tier: L` and the chaos roll recorded
@ -127,6 +129,7 @@ per dimension — a number or a reproducible comparison, not an impression.
id: CB-WP-0025-T02
status: todo
priority: high
state_hub_task_id: "bdd9fc97-4706-4d8e-b286-91ff53339680"
```
Tier L requires it. Exactly one round: challenge, then response, trail in
@ -154,6 +157,7 @@ Tier L requires it. Exactly one round: challenge, then response, trail in
id: CB-WP-0025-T03
status: todo
priority: high
state_hub_task_id: "055630a1-56e0-4ad4-8796-35f444313371"
```
`decisions/ADR-0013-*.md`. (ADR-0012 is CB-WP-0022 T03's.) At minimum:
@ -180,6 +184,7 @@ priority: high
id: CB-WP-0025-T04
status: todo
priority: high
state_hub_task_id: "7c3ca25d-0001-44ef-a07d-0ab2d51b09a5"
```
`specs/` — extend `MetricsAndScenarios.md` or add a capability spec as the
@ -209,6 +214,7 @@ kind that has already gone wrong twice.
id: CB-WP-0025-T05
status: todo
priority: high
state_hub_task_id: "5f1cbda3-7427-4f23-8c67-b2e216b987a8"
```
The search, the bound, and the replayable path. Wire it to the ending page
@ -233,6 +239,7 @@ a loss.
id: CB-WP-0025-T06
status: todo
priority: high
state_hub_task_id: "723feae8-1ed4-4dd4-9330-8ec84625e5d9"
```
Run the measurement, ship it as a `make` target beside the other
@ -258,6 +265,7 @@ message that sat unread for four days is the baseline to beat*.
id: CB-WP-0025-T07
status: todo
priority: high
state_hub_task_id: "2616e06c-5d88-4680-be18-4452515dff21"
```
`evidence/CB-EV-0023-*.md`.