T05. tools/design.py, make design, and the register in GroundRules.md --
14 rows, no new file, because ADR-0012 D2 made §Underdetermined the
register rather than building one beside it.
Backfill was the test and it caught two things the ADR did not have.
First, a `role` column. The first report alarmed on U2 and was wrong to:
U2's scenario is green BECAUSE the provisional default it documents is
implemented, which says nothing about whether ground-game agrees. GR-E01's
was a counterexample that went green. Same colour, opposite meaning -- a
register that cannot tell them apart either alarms constantly or never.
Only a green counterexample alarms. Folded back into GameDesign §1.3.
Second, it contradicted the survey. CB-RES-0007 said "six of the ten
already have provisional scenarios." Measured -- grep -lE "\bU<n>\b" over
scenarios/ground -- exactly ONE U-item names itself. Five provisional
scenarios exist and four probably encode U-item defaults, but the mapping
is not written down, so it is not checkable. Same defect class as the
wrong premises, found inside the survey that proposed the fix. Now a
reported debt: open, lacking a reproduction: 9, target 0.
design.py carries the control design-baseline.py never had, asserted
directly: a row citing a nonexistent file must not count as reproduced,
using the exact path 2da19a4 deleted -- which the old tool called green.
design-baseline.py is marked superseded rather than deleted; it is the
evidence for how a wrong number got into a survey.
T06. The report is a FILE in ground-game under GROUND-WP-0002, committed
there, with a hub message that only points at it. It asks for no ruling:
it carries GR-E01's withdrawal, our reproduction debt, and two notes that
are explicitly not findings.
And it had to acknowledge something nobody anticipated. GROUND-WP-0002 is
finished -- all ten U-items were RULED 2026-08-03, every one confirmed as
the default we simulate, plus five of six provisional scenarios. The
survey said "0 of 10 ruled" two days later and this register was built
saying `reported`. That is the unread-inbox failure running in the
opposite direction: they answered and we did not collect it. The
instrument's first run surfaced it. They are `ruled`, not `applied` --
lifting the now-settled provisional flags is owed and is not done, and
make design shows them open until it is.
T07. evidence/CB-EV-0021. Two of six catches in this pass came from
execution rather than process (the role distinction from building it, the
ten uncollected rulings from running it), which is InnerLoop §Design
goal's prediction holding.
make self-tests, facts-check, loop-lint: clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
loop-lint flagged 429 lines against the ~400 limit, and it was right about
the cause: the T02/T03/T04 completion records restated content that
ADR-0012, GameDesign.md and the challenge/response trail already carry.
Trimmed to pointers plus the one sentence each that is not written down
elsewhere.
400 lines, loop-lint clean. No content lost from the artifacts that own
it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Not a register; ADR-0012 D2 put that in GroundRules §Underdetermined. This
spec says what may go in it, what a reproduction must show, how a finding
dies, and how a trial game is run.
§1.2 is written against evidence rather than principle. A finding must
print the rows behind any number it claims, and the spec carries the table
of what shipped instead: a sum ("12 in the file"), a green scenario
("4/6/9 against 5/7/9"), and a condition named without checking which one
fired ("SOLVE on a face-down Problem"). "12" was arithmetically defensible
and still wrong about the game -- that sentence is the requirement.
§1.3's target is 0 reproductions that have gone green while open. GR-E01
would have tripped it four days before a human caught it by hand.
No baseline rate is quoted. The 33% was withdrawn by C2 and the first
honest denominator is T05's backfill; quoting a new number from a
discredited instrument is how the first one got in.
The trial protocol costs one flag: cb-play --record already writes a
finished game as a scenario, so a trial is that plus a sibling .md in the
player's own words. An observation is a NOTE until it has a reproduction,
and notes may not cross the repo boundary and expire at 30 days on the
existing provisional-age machinery. The maintainer's "I felt it was too
easy but then we lost" is the case the protocol is shaped around --
forcing it into a schema at the moment of observation would lose it.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
one clause short
Nine decisions. Two were not on T03's list; both are the review's.
D2: specs/GroundRules.md §Underdetermined IS the register. C4 pointed out
it was never evaluated as a candidate, and against the survey's own five
benchmarks it already delivers four -- including the Magic Oracle property
("a ruling flips the scenario, not the kernel", :231-233) that the survey
travelled to Magic to discover and we had written down ourselves eight
days earlier. What it lacks is reproductions. So this pass extends a
section rather than building a register: no new file, no new schema, and
no second mechanism to disagree with the first.
D3: admissibility is three clauses. It exists; it has the ruled shape
(GROUND-WP-0004 T02's row-level table, never a sum -- promoted from a T04
addendum because two of three wrong premises were sums without tables);
and it CAN FAIL. The third is C1's. GR-E01's scenario went green when the
edition landed, and the finding stayed admissible and stayed queued for
transmission, because nothing in the rule said a passing artifact was a
signal. A green reproduction is an alarm, not a reassurance.
D1 applied: INTENT gains a fourth property, Instrument, worded as a
mechanism rather than an ambition and carrying its own falsifier -- if a
pass tolerates an undecided rule by quietly picking a default, the
property is false.
D4 five kinds, each forced by an existing finding; a sixth during backfill
means the taxonomy was invented. D5 lifecycle where `applied` means the
source changed, the queue empties while the log accumulates, and
withdrawals are reported rather than deleted -- GR-E01 is why. D6 notes
admitted but never reportable, 30-day expiry on the existing age
machinery; refusing them would discard the only class of finding the
engine cannot produce itself, which is CB-WP-0025's whole input. D7 no
engine-evolution register, on an inventory C5 corrected -- narrowed, not
settled. D8 design-baseline.py retired, kept as a dated snapshot because
deleting it erases the evidence for how 33% got in. D9 the artifact stays
here, ground-game gets a generated file under its own workplan.
loop-lint: no findings. facts-check: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
First adversarial review in this repo run by a genuinely separate agent.
CB-RES-0006's reviewer opened by conceding it could not be, and called its
own findings "a lower bound on what a genuinely separate reviewer would
find." This is the measurement: the separate reviewer ran git log against
the survey's central example and found 2da19a4 had falsified it four days
earlier, while the author -- who wrote that commit -- quoted the dead
number twice.
Seven challenges: four conceded, two conceded in part, one answered.
C1 changes the design. "GR-E01 is admissible because 4/6/9 against 5/7/9
is a computation anyone can rerun" was a computation already rerun: the
edition import measured 6/9/12, the conclusion inverted, and the scenario
was renamed -unreachable- to -reachable-. That finding was one of the TWO
that passed the reproduction rule. So three wrong premises have now
reached ground-game and the third satisfied an existence test -- existence
is not the property that was missing. The rule gains shape (ground-game's
row-level deal table, promoted from a T04 addendum) and a clause the
survey never contemplated: a reproduction must be able to fail. Ours went
green and stayed admissible.
C1 also caught a defect in flight. T06's payload, status todo, still named
4/6/9 and was queued to send it to ground-game as "no dataset reconciles
them." Withdrawn before sending -- the fourth wrong premise, and the only
one stopped.
C2 withdraws the baseline's precision: design-baseline.py is a
hand-maintained dict counting itself, has_reproduction never checks the
file exists (its YES-control is green against a deleted path), and
Makefile:127 runs only --self-test so the reporting path has no CI. The
direction survives; 33% is not a measured rate and T05 must not build on
it. C3: "six provisional defaults" is five, GR-E01 double-counted. C4:
GroundRules §Underdetermined was never evaluated as a candidate and
already delivers four of five benchmarks -- T03's burden flips to arguing
extension over replacement.
Survived: the rule's affordability, and reuse of the provisional
machinery.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Five remarks from the maintainer's test games, checked against the code
before being written down — two had already reached ground-game on wrong
premises, so a claim now names the line that makes it true.
Three of the five turned out to be data the projection already carries,
drawn as text: solution_deck_len, solution_discard, and OutcomeView's
personal/mastery/winners. One control (`close — I have read this`) is
labelled as a reading but shuts the server down, and leaves a live-looking
page pointing at a dead port. One number does not exist at all: table.rs
loops run_game and keeps only the last summary.
CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the
other seats' plays as objects on the table, the ending control saying what
it does, and a tally that survives "play again".
CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how
hard is this" are the same search asked twice. Tier L because the
information boundary is the whole design problem: a solver reading
GroundState sees the deck the rules hide, and would tell the maintainer he
could have won by playing a card he had no way to know was there. Also
unblocks GROUND-WP-0005, active with both tasks waiting on a measured
difficulty baseline.
Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023
shipped CB-EV-0020 first.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
GR-S01 was ruled 2026-08-04: Surface always, union hidden priorities
1..k, k = 2/3/4. So the deal is 3/4/5 Problems, not 2/3/4; available
points 6/9/12; thresholds 5/7/9 stand; SHARED GROUND is 2-6p as printed.
The question CB-WP-0021 was going to send has been answered, and it was
the deal.
CB-WP-0021 is re-scoped. The import is now LOAD-BEARING rather than
merely correct: the ruled 6/9/12 holds only with Problems.csv values, and
the same deal with the stand-in gives 6/10/15 -- a different game that
happens to also be winnable. Fixing the deal without importing the data
would produce numbers nobody ruled on, so T05 (deal) and T02 (import)
must land together. gd0001 is to be INVERTED, not deleted: it is the
record of why this changed. gr-e01 is rewritten as a non-provisional
import check, per the ruling's own wording, and loses its provisional
owner because ground-game has now ruled.
CB-WP-0022 absorbs ground-game's process ruling, which is stricter than
this pass proposed: arithmetic findings need a runnable reproduction AND
a row-level deal table listing Surface and each hidden priority
separately, never only a sum or a deal depth. That is a direct
consequence of both premises we got wrong. So the reproduction rule gains
a SHAPE requirement, not just an existence one -- a finding that ships a
passing test but describes the wrong quantity is still a bad finding, and
that is what happened twice. T02's review brief is flipped accordingly:
press whether the rule is SUFFICIENT, not whether it is affordable.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The AM-1 coverage gate failed the build on GR-P05 being uncovered, which
is what showed the rule was in the offer layer rather than in validate.
A rule enforced only by the offer is enforced only for clients that ask
what is legal. The gate did not catch a bug, it caught a design error.
And the reported case was not the one reported. CB-WP-0018, CB-EV-0016
and the message to ground-game all described SOLVE offered on a
face-down Problem; validate already rejected face-down, so it never was.
Problem 1 is the Surface Problem, face-up from the deal, so the three
inert SOLVEs were the HAND case. The ruling covers both so nothing is
invalidated, but a ruling was requested on a wrong description -- the
second time in three passes that a premise reached ground-game
unchecked, after the '12 points available' that voided GR-E01.
Two of two. The pattern is not careless analysis; it is that a claim gets
SENT the moment it is interesting and checked afterwards. Unexecuted
verification, one step further out: not a belief acted on, but a belief
published. CB-WP-0022's reproduction rule would have caught both.
An earlier mutation run reported three survivors and was wrong -- the
replacement strings did not match, so nothing was mutated. It proved
nothing and looked like a result.
Also renames CB-WP-0022-T06B to T07; the hub flagged it as an
unregistered species.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CB-RES-0007 plus a runnable baseline harness. Tier L invokes the
runnable-baseline option; the external candidates are practices rather
than software, so their rows are directional and cap at parity, and the
row that CAN be run is our own.
Measured: 6 findings across 11 files with no index, 2 of 6 (33%) with a
runnable reproduction, U1-U10 raised 2026-07-30 and first READ
2026-08-03 -- 4 days, 0 of 10 ruled.
The uncomfortable number is stated before the review can find it: the
proposed 'no finding without its reproduction' rule would reject four of
our six existing findings. The survey answers rather than routes around
it -- none of the four is expensive to reproduce, so 33% is evidence
nobody was ever asked for one.
Magic corrected an assumption this pass was about to build on. Rulings
are NOT authoritative -- they are 'reminder information with no actual
weight or rules meaning' -- and the authoritative fix folds into the
Oracle card text. So a finding closes when the SOURCE changes, not when
an annotation is added, and the register must be a queue that empties
rather than an archive that grows. That is now a constraint on the ADR's
lifecycle.
Model checkers supply the reproduction rule independently: a
counterexample trace IS the finding. W3C's implementation-defined mark is
the machinery we already have in provisional: scenarios and must reuse.
The loop-lint gate caught the new tool with no --self-test; it has one,
pinning the 2-of-6 baseline so a later edit cannot move it silently.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The maintainer named a new aspect: clay-borg as a game design tool, with
a register for design flaws, questions, results and trial protocols.
Structural L on the maintainer-named-high-leverage trigger, and it amends
INTENT. Chaos d8=6, no override.
The insight is that this is already happening with no home. Five passes
have produced ten underdetermined rules points, SOLVE offered on a
face-down problem and always inert, GR-A13's wasted SOLVE, GR-E01
unreachable below 5 seats, six provisional scenario defaults, and two
scoring modes never played to the end -- every one found by BUILDING the
simulator rather than by playing it. A simulator rigorous enough to
refuse ambiguity is a design instrument, because it cannot proceed past a
rule that does not decide. All of it has been carried in prose in six
places and one sat unread in an inbox for four days.
The load-bearing rule: a design finding is not admissible without its
reproduction. A register that collects opinions would reproduce this
project's standing failure -- unexecuted verification -- in a new medium.
Recorded as a judgment for the adversarial review rather than assumed:
the engine-evolution meta the maintainer also asked about should NOT be
built, because evidence/, decisions/, gates.toml and workplans already
carry nineteen passes of it with dates, costs and falsifiers. A second
register for the same subject is ceremony. The asymmetry is the point --
engine evolution has a home and game design does not.
T05 backfills the six known findings as the TEST of the register: one
that cannot express findings the project already has is the wrong
register, and discovering that after designing it is why the order is
survey, review, decide, specify, build.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>