Five remarks from the maintainer's test games, checked against the code
before being written down — two had already reached ground-game on wrong
premises, so a claim now names the line that makes it true.
Three of the five turned out to be data the projection already carries,
drawn as text: solution_deck_len, solution_discard, and OutcomeView's
personal/mastery/winners. One control (`close — I have read this`) is
labelled as a reading but shuts the server down, and leaves a live-looking
page pointing at a dead port. One number does not exist at all: table.rs
loops run_game and keeps only the last summary.
CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the
other seats' plays as objects on the table, the ending control saying what
it does, and a tally that survives "play again".
CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how
hard is this" are the same search asked twice. Tier L because the
information boundary is the whole design problem: a solver reading
GroundState sees the deck the rules hide, and would tell the maintainer he
could have won by playing a card he had no way to know was there. Also
unblocks GROUND-WP-0005, active with both tasks waiting on a measured
difficulty baseline.
Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023
shipped CB-EV-0020 first.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The row folded a 5,000-event log against a 100,000-event log and compared
throughputs, which confounds 'does cost per event grow with history'
(the property it claims) with 'does streaming a 20x longer Vec cost more
per element' (a memory-hierarchy fact true of any program). It measured
the second and reported it as the first: importing the edition enlarged
the aggregate and the ratio fell to 0.845 with the state bounded.
Corrected to time the SAME 5,000 events on a state at depth 0 and on a
state at depth 100,000. Equal windows, equal event mix, so the only
difference left is history depth.
corrected: clean 1.004, mutated 0.589 (red)
old: clean 0.845 (red on healthy code), mutated 0.751
It also runs in 8.5s instead of timing out: the first version re-walked
the 100k prefix every repetition, 200M untimed folds per sample, which
under the mutation never finished. A control that cannot be run is not a
control. It now advances to depth once per sample and clones.
Two of my own measurements here were wrong and both were caught by
measuring again. A 2-minute timeout killed the shell line before its
restoring cp ran, so three readings were taken on MUTATED code -- I
diagnosed an event-mix confound that did not exist and 'fixed' it. The
fix is kept on its merits; the justification was fiction. And the probe
that proved state was bounded had checked four of eleven collections.
make all exits 0.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
ADR-0011 decided it: vendor the CSV with a checked digest, read it with
a ~50-line reader, and let the hashes move.
The declaration's constraint was measured against the WRONG BUDGET. It
said a CSV crate costs 21,613 against AM-4a's 3,798 of headroom, '5.7x
over, settled by measurement'. But setup and problem_priorities are
cfg(scenarios) and are not in the shipped runtime at all, so AM-4a never
sees them. Against AM-4b, csv costs 17,651 against 19,742 -- it FITS,
with 2,091 to spare. It is refused anyway, on proportion: 89% of the
budget's remaining capacity to read 20 rows. The revisit condition is
stated (nested quoting, embedded newlines, multiple dialects).
GR-S01 now deals Surface + hidden 1..=k as ruled, with edition values and
suits. Measured: 6/9/12 available against thresholds 5/7/9 -- the game is
winnable at every seat count, which is what the maintainer could not do.
gd0001 is INVERTED, not deleted, and now also asserts the 6/9/12 so a
deal that is reachable for the wrong reason still fails.
Blast radius was scenario expectations, exactly as the ADR predicted: no
scenario pinned a hash and no bundle is committed. Six scenarios and two
unit tests updated, each with a note. gr-e01-threshold-unreachable-2p is
RENAMED to -reachable- and rewritten as the non-provisional import check
ground-game asked for by name. gr-e03's setup was restructured, not just
renumbered: with values 2,2,2 its personal-edge test would have tied
three ways and asserted nothing.
BLOCKING: AM-7 fails at median 0.845 against its 0.9 floor. Isolated
across three runs -- 3 problems + stand-in 0.97, 3 problems + edition
0.909, 4 problems + edition 0.845. State is BOUNDED (proven: identical
after 5k and 100k events), so this is not the unbounded-growth defect
AM-7 exists to catch; it is a bigger working set streaming a long log.
Whether AM-7's floor is still right for a larger aggregate is a spec
question and lowering it requires an ADR, so it is not being tuned here.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
GR-S01 was ruled 2026-08-04: Surface always, union hidden priorities
1..k, k = 2/3/4. So the deal is 3/4/5 Problems, not 2/3/4; available
points 6/9/12; thresholds 5/7/9 stand; SHARED GROUND is 2-6p as printed.
The question CB-WP-0021 was going to send has been answered, and it was
the deal.
CB-WP-0021 is re-scoped. The import is now LOAD-BEARING rather than
merely correct: the ruled 6/9/12 holds only with Problems.csv values, and
the same deal with the stand-in gives 6/10/15 -- a different game that
happens to also be winnable. Fixing the deal without importing the data
would produce numbers nobody ruled on, so T05 (deal) and T02 (import)
must land together. gd0001 is to be INVERTED, not deleted: it is the
record of why this changed. gr-e01 is rewritten as a non-provisional
import check, per the ruling's own wording, and loses its provisional
owner because ground-game has now ruled.
CB-WP-0022 absorbs ground-game's process ruling, which is stricter than
this pass proposed: arithmetic findings need a runnable reproduction AND
a row-level deal table listing Surface and each hidden priority
separately, never only a sum or a deal depth. That is a direct
consequence of both premises we got wrong. So the reproduction rule gains
a SHAPE requirement, not just an existence one -- a finding that ships a
passing test but describes the wrong quantity is still a bad finding, and
that is what happened twice. T02's review brief is flipped accordingly:
press whether the rule is SUFFICIENT, not whether it is affordable.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The AM-1 coverage gate failed the build on GR-P05 being uncovered, which
is what showed the rule was in the offer layer rather than in validate.
A rule enforced only by the offer is enforced only for clients that ask
what is legal. The gate did not catch a bug, it caught a design error.
And the reported case was not the one reported. CB-WP-0018, CB-EV-0016
and the message to ground-game all described SOLVE offered on a
face-down Problem; validate already rejected face-down, so it never was.
Problem 1 is the Surface Problem, face-up from the deal, so the three
inert SOLVEs were the HAND case. The ruling covers both so nothing is
invalidated, but a ruling was requested on a wrong description -- the
second time in three passes that a premise reached ground-game
unchecked, after the '12 points available' that voided GR-E01.
Two of two. The pattern is not careless analysis; it is that a claim gets
SENT the moment it is interesting and checked afterwards. Unexecuted
verification, one step further out: not a belief acted on, but a belief
published. CB-WP-0022's reproduction rule would have caught both.
An earlier mutation run reported three survivors and was wrong -- the
replacement strings did not match, so nothing was mutated. It proved
nothing and looked like a result.
Also renames CB-WP-0022-T06B to T07; the hub flagged it as an
unregistered species.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Implements ground-game's ruling of 2026-08-03. make all exits 0, 26
scenarios, rule coverage 59/59, and no scenario encoded the bug.
The rule ended up somewhere other than where I put it, and a gate moved
it. It went into legal_commands first; the AM-1 coverage gate then
demanded a scenario for the new GR-P05, and scenarios drive validate, not
the offer layer. A rule enforced only by the offer is enforced only for
clients that ask what is legal -- the browser would be filtered and a
scenario file would walk straight past it. Once GR-P05 moved into
validate, every condition in legal_commands was dead code, and the
layering test said so in those words.
And the reported case was not the one I reported. CB-WP-0018 and the
message to ground-game described SOLVE offered on a FACE-DOWN Problem.
Measured: validate already rejected face-down, so it never was offered.
Problem 1 is the Surface Problem, face-up from the deal -- the
maintainer's three inert SOLVEs were the HAND case, holding no Clarify
for a Clarify Problem. The ruling covers both so nothing is invalidated,
but the record was wrong.
Four conditions asserted separately, because one 'SOLVE is filtered' test
would pass with three of four implemented.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
ground-game ruled SOLVE's legality on 2026-08-03: not offered on a
face-down or Denied Problem, not offered without a matching Solution in
hand, and not offered on a Problem claimed in a prior round. The bluff
reading CB-WP-0018 raised is dead -- it was a filter bug, and the engine
has been offering an inert move since legal_commands was written.
Only SOLVE is ruled on, so only SOLVE is touched. Implementing more than
was ruled would be inventing rules, which is what this exchange exists
to stop.
Chaos d8=6, no override.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The maintainer played several 3-player games on 2026-08-03 and could not
win any. This says why, from the engine's own constants rather than my
arithmetic: claim every Problem the deal puts in play, concede nothing,
and the total still falls short of GR-E01's threshold at 2, 3 and 4
seats.
2p: 2 problems worth 3 vs threshold 5 — UNREACHABLE
3p: 3 problems worth 6 vs threshold 7 — UNREACHABLE
4p: 3 problems worth 6 vs threshold 7 — UNREACHABLE
5p: 4 problems worth 10 vs threshold 9 — reachable
6p: 4 problems worth 10 vs threshold 9 — reachable
The test reads problem_priorities (GR-S01's deal) and threshold (GR-E01)
out of the engine, so it cannot drift from the rules it tests, and it
holds for either dataset -- the stand-in gives 3/6/10 and Problems.csv
gives 4/6/9 against the same 5/7/9.
It ASSERTS THE DEFECT and is expected to keep passing until ground-game
rules, then be inverted. This is CB-WP-0022's reproduction rule applied
before the register exists, because delivering feedback should not block
on building the tool that tracks it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CB-RES-0007 plus a runnable baseline harness. Tier L invokes the
runnable-baseline option; the external candidates are practices rather
than software, so their rows are directional and cap at parity, and the
row that CAN be run is our own.
Measured: 6 findings across 11 files with no index, 2 of 6 (33%) with a
runnable reproduction, U1-U10 raised 2026-07-30 and first READ
2026-08-03 -- 4 days, 0 of 10 ruled.
The uncomfortable number is stated before the review can find it: the
proposed 'no finding without its reproduction' rule would reject four of
our six existing findings. The survey answers rather than routes around
it -- none of the four is expensive to reproduce, so 33% is evidence
nobody was ever asked for one.
Magic corrected an assumption this pass was about to build on. Rulings
are NOT authoritative -- they are 'reminder information with no actual
weight or rules meaning' -- and the authoritative fix folds into the
Oracle card text. So a finding closes when the SOURCE changes, not when
an annotation is added, and the register must be a queue that empties
rather than an archive that grows. That is now a constraint on the ADR's
lifecycle.
Model checkers supply the reproduction rule independently: a
counterexample trace IS the finding. W3C's implementation-defined mark is
the machinery we already have in provisional: scenarios and must reuse.
The loop-lint gate caught the new tool with no --self-test; it has one,
pinning the 2-of-6 baseline so a later edit cannot move it silently.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The maintainer named a new aspect: clay-borg as a game design tool, with
a register for design flaws, questions, results and trial protocols.
Structural L on the maintainer-named-high-leverage trigger, and it amends
INTENT. Chaos d8=6, no override.
The insight is that this is already happening with no home. Five passes
have produced ten underdetermined rules points, SOLVE offered on a
face-down problem and always inert, GR-A13's wasted SOLVE, GR-E01
unreachable below 5 seats, six provisional scenario defaults, and two
scoring modes never played to the end -- every one found by BUILDING the
simulator rather than by playing it. A simulator rigorous enough to
refuse ambiguity is a design instrument, because it cannot proceed past a
rule that does not decide. All of it has been carried in prose in six
places and one sat unread in an inbox for four days.
The load-bearing rule: a design finding is not admissible without its
reproduction. A register that collects opinions would reproduce this
project's standing failure -- unexecuted verification -- in a new medium.
Recorded as a judgment for the adversarial review rather than assumed:
the engine-evolution meta the maintainer also asked about should NOT be
built, because evidence/, decisions/, gates.toml and workplans already
carry nineteen passes of it with dates, costs and falsifiers. A second
register for the same subject is ceremony. The asymmetry is the point --
engine evolution has a home and game design does not.
T05 backfills the six known findings as the TEST of the register: one
that cannot express findings the project already has is the wrong
register, and discovering that after designing it is why the order is
survey, review, decide, specify, build.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The declaration claimed importing the edition data would resolve GR-E01.
Measured across all four scenarios: GR-S01 deals 2/3/4 problems by player
count, not all five, so 4/6/9 points are in play against thresholds of
5/7/9 -- unreachable at 2p and 3-4p with the REAL data, in the same shape
as the stand-in's 3/6/10. So GR-E01 unreachable below 5 seats is a real
property of the game and gr-e01-threshold-unreachable-2p asserts
something true.
The error was the cheap kind: 12 points exist in the file, so I assumed
12 are in play. One command over the CSV settled it and was not run until
after the declaration was committed -- this project's characteristic
error, in the pass that followed a ruling obtained because of it.
CB-EV-0018 corrected too: 'confirmed as the stand-in's doing' was
overstated. The zero came from no Problem being claimed at all.
T03 now owes ground-game a sharper question than a retirement: either
GR-S01's deal count is wrong or GR-E01's thresholds are, and no dataset
can reconcile them.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
GROUND-WP-0002 T01 ruled the edition dataset authoritative, so the engine
must stop inventing Problem values and suits. Problems.csv carries 5
problems per scenario worth 2,2,2,3,3 (total 12) with a required_solution
each; the stand-in deals 3 worth 1,2,3 (total 6). GR-E01's thresholds of
5/7/9 are ordinary against 12 and unreachable against 6 -- which is why
the maintainer's last game ended 0 scores and winners nobody, and why
'GR-E01 unreachable below 5 seats' was carried as a rules gap. It was
never a rules gap.
Measured before declaring: AM-4a has 3,798 lines of headroom and a CSV
crate costs 21,613 marginal (csv 14,291 + csv-core 3,360 + ryu 3,962;
itoa/memchr/serde/serde_core are already present and free). 5.7x over, so
the shipped runtime cannot gain a CSV parser and that is settled by
measurement rather than preference.
The ADR's third question is the one that bites: Problem values and suits
become part of GroundState, which is hashed, so every recorded state hash
changes. A content import that quietly invalidates every hash in a
project whose central invariant is replay determinism is not a
data-loading change.
Chaos d8=7, no override.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Six of seven perceptual defects fixed; item 1 already passed.
T01, at the maintainer's instruction: a legal target restyles its
EXISTING border rather than drawing a new box. outline + outline-offset
drew a second rectangle, which an SVG viewport clips (the missing top and
left edges) and which made a seat's highlight card-sized. A border
already in the layout cannot move the layout.
T02: the ghost was a textContent copy of the card, which is why the line
break collapsed and it read as a second card, and why showing the
explanation destroyed the label. It is now a pill, the explanation is
appended beside the label, and the left-behind element is dimmed and
dashed. The stub grew innerHTML so a test can assert BOTH are present --
it could previously only see that something was displayed.
T03: NOT reproduced and recorded as not reproduced. The likeliest cause
is which element the browser reports -- for touch and pen the pointer is
captured to the pointerdown target, making every drop look like a
drop-on-itself, which is the other half of the report. elementFromPoint
is correct under both explanations. Separately the refusal was written in
element ids on the one surface a player reads when something goes wrong;
it now speaks the game's words and a test forbids id leakage.
T04: seat selections rendered as Debug. The coverage gate then failed my
first fix for dropping a field when target and problem were both set --
the aggregate does not produce that shape and the gate was right not to
care.
T05: the headline reads from group_success. 'Play again' is real, and its
first version was useless: run_game bound a fresh listener per game, so a
second game moved to a new port and left the tab pointing at a dead one.
One listener per session now, and the test asserts the second game is a
DIFFERENT deal.
Chaos d8=8 fired the first override at the new rate and drew S, changing
nothing -- one half of window 2's retirement condition.
CB-WP-0019 settled at $38.54/117 against $34.80/107. Eight for eight,
and the first under 20%.
make all exits 0.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The perceptual check found seven items; item 1 passes. Chaos d8=8 fired
the FIRST override at the new rate, on the second roll, and drew S --
which is what the structural derivation said, so it changed nothing.
That is one half of window 2's retirement condition (retire if an
override changes nothing twice running).
The maintainer's design instruction is adopted directly: a legal drop
target should change its EXISTING border to dashed rather than draw a
new outline. That explains the hidden top/left edges (an outline on an
SVG <g> is clipped by the viewport) and the oversized seat highlight.
The ghost is a textContent copy, which is why the linebreak collapses and
it reads as a second card, and why showing the explanation destroys the
label -- one cause, two reports.
T03 carries an explicit instruction not to fix a message that already
works: the 'nothing droppable' path may simply be unreachable because
almost every part of the page is a card. Reproduce before changing.
T05 records that 0 scores and no winner is very likely the stand-in
dataset rather than a scoring bug, now that GROUND-WP-0002 T01 has ruled
the edition data authoritative. The import is its own pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
T03: InnerLoop v1.7 plus loop-lint's own-cost check. Six passes
under-reported themselves by 30-45%, never once high, and the rule lived
only in evidence files having been re-derived three times. The READING
is load bearing, not the boundary: CB-WP-0018 T04 applied 're-run the
instrument at the moment of quoting' alone and its figure was correct.
So the operative instruction is re-run when you quote, and loop-lint
fails an evidence file naming its own workplan beside a dollar amount
without marking it provisional.
It binds forward from this pass. The check fires on seven historical
files which ARE the evidence for the rule; making them comply would edit
the record to remove the thing it proves -- the same category error as a
live fact: tag on a dated measurement, which this pass also hit.
Lifecycle, at the maintainer's instruction: ready -> active -> done,
where ready means declared and not started. loop-lint fails a workplan
that has started and still says ready, one that is active with
everything closed, and one that is done with an open task. The first
version of that check was WRONG and its own self-test caught it: it
stripped the leading status: assuming frontmatter, which silently
dropped a real task once the frontmatter said ready or active.
Both new checks then fired on this pass's own artifacts and both were
right to.
T04: CB-EV-0017. The new meta budget's first reading is a breach it
caused -- 27% against the 20% line, because this pass cost $31.18
against product passes averaging ~$21. Reported rather than exempted:
ADR-0006 D2 covers the instrument repairs but not the rule-writing, and
the honest reading is that this should have been two passes.
CB-WP-0018 settled at $36.53/95 against $28.08/82 last reported, 30%
higher. Seven for seven.
make all exits 0.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The two AM-4 budgets had the SAME scope -- one package, no dev edges --
while claiming to bound different things. AM-4b now measures the
workspace with dev edges: 57 crates / 725,258 lines where it read 29 /
317,021, having been blind to 28 crates and 408,237 lines, more source
than its own target.
Target 745,000, ~2.7% of room -- the same margin ADR-0008 D3 gave AM-4a,
applied to a number that grew because the instrument was repaired, not
because anything was added. The target moved to fit the measurement.
T02: proc-macros are COUNTED here and excluded from AM-4a, on purpose.
AM-4a asks what ships and a proc-macro never ships. AM-4b asks what is
acquired, and ADR-0007 D3's acquisition rule counts what the build
fetches -- 'it does not ship' is no answer to 'we downloaded it'. When
the rules disagree, the question each budget asks decides. Measured
share 109,585 lines / 15.1% against AM-4a's 36.2%, so ADR-0008 D2's
refusal to borrow the ratio was right by more than a factor of two.
Caught by this project's own earlier work twice: the mutation
find-string went stale and --self-test reported it BUILD-FREE (the check
CB-WP-0015 added after AM-4a's rotted for two passes), then the DFD gate
caught facts.toml carrying the old numbers.
CB-EV-0001 and ADR-0004 carried live fact: tags on historical readings.
A dated record asserting a CURRENT value is a category error, so those
occurrences are marked as-measured instead of retro-edited, and ADR-0004
gains a supersession note.
make all exits 0.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CB-WP-0005 read 5/8 and CB-WP-0007 read 2/6 in every status run since
they closed, because only 'done' counted and their remaining tasks are
'cancel'. Two permanently-wrong numbers teach the reader to skip the
column. Both now collapse into 'closed and complete' -- 18 of them --
while an open workplan with cancellations still shows the count
separately, so a cancellation is visible rather than laundered into
completion.
The workplan block is extracted as render_workplans so the control is
stated over what the tool PRINTS rather than over a literal the test
wrote: a done+cancelled fixture must collapse, and one with a real todo
must not. The first version of that control asserted arithmetic on its
own input, which is the tautological shape ADR-0010 D2 demoted.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
InnerLoop v1.7. The purpose is written first and the number follows from
it: most spend on the task at hand, some on control, review and
improving the process. make status prints it above the figure, because a
threshold with no stated purpose is what let this number be argued three
times.
Soft 20% over a trailing 5, and the self-test enforces that the ratio and
the window are a PAIR: META_SOFT_PCT == 100 / TRAILING_PASSES. One meta
pass among n at parity cost reads 1/n, so 80/20 is one pass in five at
normal cost -- a five-pass window. The same 20% over three would have
silently also demanded the meta pass be half-price, which makes meta work
rushed rather than rare. Moving the ratio without the window goes red.
The phase setting is declared, argued and expiring in gates.toml, and
reverts on review_by unless re-argued. Verified live at 35%. One with no
reason or no expiry is refused rather than honoured, because a threshold
anyone may move is not a threshold.
Measured: the last five passes read 7% against the new line.
InnerLoop.md crossed the 400-line limit three times while this was
written and was fixed structurally each time -- the arithmetic, the
cost-per-response basis and the two review case studies moved to
InnerLoopReference.md. The limit was not raised.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The maintainer asked for a rule for what the budget is FOR: main spend on
the task at hand, some on control, review and improvement, 80/20 to
start, adjustable by phase.
META-25 has a threshold and no stated purpose, which is why the number
has been argued three times. The purpose goes first.
Recorded in the task: the ratio and the window are a pair. Over a
trailing 3-pass window one meta pass at parity cost is already 33%, so a
20% line there means 'one in five AND half price' rather than 'one in
five'. Over trailing 5, 20% is exactly one pass in five at normal cost,
which is the literal reading of the instruction.
The phase adjustment must be declared, argued and expiring in the shape
gates.toml already uses -- a threshold anyone may move is not a
threshold, and this project fixes limits structurally rather than
raising them.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Both owed numbers measured BEFORE declaring, so the work is scoped
against facts. AM-4b's scope: 29 crates / 317,021 lines instrumented
against 57 / 725,258 real, so 28 crates and 408,237 lines are uncounted
-- more than its own 350,000 target. AM-4b's proc-macro share: 109,585
lines, 15.1%.
That 15.1% vindicates ADR-0008 D2, which refused to correct AM-4b using
AM-4a's measured 36.2% because 'correcting a second instrument on the
strength of the first one's ratio is the error this change exists to
fix'. Borrowing would have been wrong by more than a factor of two.
Also carries the self-quoting rule, which is six-for-six under-reported
by never less than 30% with both causes diagnosed, and still lives only
in evidence files.
Structural tier M: changes a budget's scope and target and a reporting
rule. Chaos d8=5, no override -- the first roll at the new rate.
Declaration 2 of window 2. Meta budget 0%; ADR-0006 D2 exempts
instrument repair regardless.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
T03: input::describe writes a sentence per legal command; data-descs
carries them in step with data-targets; the ghost already following the
pointer shows the one for whatever legal target is under it, so the
explanation lands beside the target with no overlay layer to keep
aligned. ADR-0010 D1 binds -- the page renders it, never composes it.
Both mutations INITIALLY SURVIVED because the fixture's Attack card had
exactly one target, where an off-by-one shift and a truncation are both
no-ops. CB-EV-0014's lesson one level in: a fixture too thin to express
a failure is how the failure survives. Two attack targets now, both red.
T04: chaos rate d4 -> d8, window 2 open at 12 declarations, retiring if
an override changes nothing twice running. Window 1's condition was NOT
met -- both overrides changed the outcome -- so the mechanism is kept.
The weakest part of the decision is that it is a rate change argued from
n=2, so window 2 carries a falsifier: no override at all is evidence the
rate went too far, not that the mechanism is healthy.
InnerLoop.md hit 401 lines and the loadability gate fired; the rationale
moved to InnerLoopReference.md, structurally, per the standing precedent
that limits are not raised.
CB-WP-0017 settled at $9.48/40 against $5.19/23 reported mid-flight,
83% higher. Six for six, always low -- read by re-running the instrument
at the moment of quoting, which is CB-EV-0015's correction applied for
the first time.
make all exits 0.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
bot::Journal -- a shared list of Applied { actor, command, events } the
driver appends to via play_journaled; play delegates with None so
nothing existing changed. BotGame.events only appears after play
returns, which is no use to a page rendered mid-game.
Phrased with record::to_step, the recorder's vocabulary, so what the
player reads is what the scenario file will say, and all 29 GroundEvent
variants now render in words instead of Debug.
A command that produced no events says 'no effect'; the mutation
dropping that branch goes red. Honest limitation recorded: the reported
SOLVE case is resolved inside the system's resolve command, which does
produce events for other seats, so it shows as a selection with no claim
following rather than an explicit 'no effect'. Making it explicit would
mean the renderer deciding why a rule did nothing -- a second
implementation of the rules, which this task's control forbids.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Server::serve_end plus doc::ending, wired into both of run_game's exits.
Where the browser used to get Connection refused it now gets the ending
page with the result and the final table. It serves until the page posts
'done' (the page carries a close control), with a 600s linger so an
abandoned tab cannot hold the process open.
document() split into body() and move_section() so the ending shows the
same table rather than a second rendering of it.
The control had to be built twice and the first was worthless:
the_end_of_the_game_reaches_the_browser calls serve_end directly, and
deleting the call from run_game left it GREEN -- it tested the link and
not the chain, which is CB-EV-0012's finding recurring.
a_real_game_played_to_its_end_leaves_the_ending_on_screen runs the real
play() with a browser seat, drives a real game to its end over a real
socket, and goes red under that mutation printing an empty page -- the
reported symptom exactly.
A weak assertion of mine caught by itself: the first draft grepped the
page for location.reload, which would have forced a second script to
satisfy a test rather than a requirement. It now asserts the ending
endpoint cannot answer 'ok'.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A discard pile already exists -- solution_discard on GroundState,
SolutionDiscarded removes from hand, and the page renders it. Building
one would have been building a thing that is there. What was dragged are
ACTION cards, which are not cards and are correctly never consumed;
solution cards leave the hand at Resolve because a selection is a
face-down commit.
But the report points at something real. Measured live: Investigate
draws correctly (2 -> 3 -> 4 cards), and Solve was then played three
times on problem-1 with the hand unchanged at 4 and discard empty --
because problem-1 was face down and GR-A02's resolver silently
continues. legal_commands offers Solve on every face-up problem without
consulting the hand.
Whether SOLVE should be selectable against a face-down problem or an
unmatchable suit is a game-semantics question and INTENT defers those to
ground-game. What is ours is that a provably-inert move is offered,
accepted and never accounted for. Folded into T02 as the case the log
must handle: a command that produced NO events is the one the player
needs to see.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Provenance: the maintainer reported 'after some time i get an empty page
back. I guess the game crashes or ends but that is unclear as the ui
disappears.' Reproduced by driving a real game to completion over HTTP:
move 5 accepted, then GET / -> Connection refused. The game ENDED
normally, 5 rounds and 30 commands, and its whole result -- coalitions,
scores, winners, hash -- went to the terminal. next_choice only accepts
connections inside a human decision point, so when play() returns the
listener dies and the post-ok reload is refused. A crash and a win
render identically: nothing. Same class as CB-WP-0016's silent drop.
Also carries the chaos rate change CB-EV-0015 owed to the next
declaration (d4 -> d8, second window of 12), which is what makes this
structurally M. Rolled at the old d4=3, no override, because a rate
changes when the decision lands and not retroactively.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Provenance (tier M, structural S, chaos d4=4 -> OVERRIDE drawn M):
the maintainer could drag after CB-WP-0016 but could not tell what was
pickable, held, or droppable. Underneath that, the page was WRONG about
which moves exist: 9 legal commands rendered as 5 cards each claiming
all three target kinds, from a const string in the emitter. Investigate
is legal on problems 2 and 3 but not 1; Solve on 1 but not 2 or 3. The
live page now says 'Solve onto problem 1'.
ADR-0010 restates control 5, which this work would otherwise have
outgrown in silence: every game fact the page acts on must arrive from
Rust as data; the script may read, match and render it, never compute,
infer, filter or default one. The survey's real finding is that the
permitted and forbidden designs are indistinguishable from outside, so
the vocabulary grep is demoted to a cheap first line and two behavioural
properties become the controls -- the highlighted set EQUALS the set
Rust emitted, and anything the page marks legal must resolve. Both
mutation-proven; the derive-legality mutation produces a plausible
highlight (seat-0,1,2 where only seat-1 is legal) and is caught.
Visible now: .pick resting shadow, .held on the grabbed element, .dropok
on every legal target including BOTH drawings of a seat, and a ghost
following the pointer. Nothing perceptual is verified and ADR-0010 D5
says so.
The DOM stub now models classList/querySelectorAll/createElement and
builds its node set from the real emitted page. Trap recorded: QuickJS
fixes its stack limit at Context creation relative to that frame, so a
helper returning a Context makes every later eval report
'SyntaxError: stack overflow'.
CHAOS WINDOW CLOSED, 12 declarations, 2 overrides, one each way. Both
changed the outcome, so the retirement condition is not met. Verdict:
keep, and recommend d4 -> d8 with a second window of 12 -- that is a
change to the loop's own constraints and is owed to the next declaration
as tier-M work, not made here.
CB-EV-0014 corrected: it quoted CB-WP-0015 at $15.14/136 and called it
the first settled figure quoted. Now $22.70/166. The number had been
read during CB-WP-0015 itself, so there are two defects -- the boundary,
and quoting from memory instead of re-running the instrument.
make all exits 0.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Provenance: the maintainer ran stage 1's check after CB-WP-0016, could
drag, and reported that the UI gives no way to tell what can be picked
up, what is being dragged, or where it may be dropped. Combined with the
prior finding that the page hides which moves are legal -- measured, 9
legal commands rendered as 5 cards each claiming all three target kinds,
with Investigate legal on problems 2 and 3 but not 1 -- highlighting drop
targets is not decoration, it is the first time the page tells the truth.
Structural tier S: presentation work inside an existing capability.
CHAOS ROLLED 4 -> OVERRIDE, drawn tier M. Second override in twelve
declarations and it rolls the opposite way from the first (CB-WP-0011 was
L rolled down to S), so the calibration window closes with one of each,
which is the minimum that makes its evaluation possible.
Declaration 12 of 12 -- the window closes here and T04 owes the verdict.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Provenance (tier S, one paragraph in lieu of survey and ADR): the human
check that kept INTENT stage 1 open was run and the drag was broken.
Root cause, worth more than the instance: drop targets were ids, and an
id must be unique, so exactly one element could ever be seat-0. The
relationship-graph circle took it and the seat card that every action
card's own text points at -- 'drag Attack onto a seat' -- silently had
none. A seat is drawn twice and both drawings are the seat; the document
model could not express that.
Drop keys are now data-drop. Any number of elements may carry the same
key, so a seat is droppable on its card and on its graph node. Measured
on a live server: seat-0/1/2 each appear twice, id survives only on
cb-status which is the one element the script looks up, and
down=action-attack&up=seat-1 returns ok.
Second defect: a drop on nothing returned without posting and without
touching the status line, so a broken target was indistinguishable from
a working page. resolve already refuses rather than defaulting, which is
right; refusing SILENTLY is not. The page now reports the raw fact --
'took action-attack, let go over nothing droppable' -- which names
elements, not moves, so ADR-0007 control 5 holds.
And the honest part: the general check added here -- every offered
affordance names a key that exists, driven through Policy::choose over
four real bot games -- does NOT catch the reported defect. seat-0 did
exist, on the graph circle. It is kept because a wholly absent target is
a real class, and paired with a targeted regression test that does catch
it. Three mutations, each red for its stated reason, including the
reported defect reintroduced; only the targeted test fires on that one.
A cb-play assertion matched id="action-ground" as a substring while
describing itself as checking the page; rewritten through drop_keys.
make all exits 0. Stage 1 stays open: verified by tests, mutation and a
live server, not by a human dragging.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Provenance (tier S, one paragraph in lieu of survey and ADR): the human
check CB-EV-0012 kept stage 1 open on was run by the maintainer and
found the drag broken. Diagnosed against the live server first:
down=action-attack&up=seat-1 returns ok, so socket, guard, resolve and
dispatch are correct. seat-{n} ids exist only on the SVG circles in the
relationship graph; player_card emits the visible seat cards with no id,
so the target every action card names is inert. jsrun feeds element ids
straight in and never hit-tests, which is why every test passed.
Structural tier S: a defect fix inside an existing capability, and the
check it adds is a product test rather than a control gate. Chaos d4=3,
no override. Declaration 11 of 12.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The rule adopted in CB-EV-0012 -- quote the previous pass's final cost,
never your own -- was applied here for the first time and did not hold.
This file opened quoting CB-WP-0014 at $7.47/34, which is what
make status reported then; by the close it read $8.56/48. A pass's
window runs to the next pass's first commit, so the previous pass is
not final until the pass after it starts. The rule fixed the wrong
boundary. Recorded as owed rather than changed silently.
Also: CB-EV-0009's prediction now has a point on each side. CB-WP-0014
opened above the SH-1 hard line at 0.220 $/response; CB-WP-0015 opened
below it, after a compaction, at 0.111. Both on the predicted side, and
the second is the control the last report said was missing -- but n=2,
different tiers and subjects, and the compaction that supplied the
control is also what makes the passes differ.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Provenance (tier S, one paragraph in lieu of survey and ADR): the two
clauses mutation-check has reported inert since CB-WP-0005. AM-7's
scaling ratio was held up by a test literally named
replay_100k_events_is_linear_and_fast that computed both throughputs,
printed both, and never divided one by the other. AM-8's N=10 was held
up by a runner that does two.
Both are now red. AM-7 3/3, AM-8 2/2, M-D1-MUT 10/14, and ADR-0005's
>=10-of-14 prediction MET for the first time. Neither was closed by
amending the question away, which was the live risk: the denominator is
unchanged and the four unenforced rows are the four already
unenforceable.
AM-7 needed three estimators. Best-of-N per leg then divide (AM-6's,
correct for a floor on one number) gave 0.581-1.085 on an unchanged
binary; legs back-to-back gave medians 0.931-1.004; legs interleaved at
fold granularity give 0.987/0.991/0.989, and 0.989 under 8-way CPU
contention while absolute throughput fell 4x. The INDETERMINATE guard
demanded unanimity and failed a good measurement over one sample
0.001 under the floor; it now requires a two-thirds majority. The
control that matters: AM-6's constant-cost mutation halves throughput
and leaves this ratio at 0.999x green, so AM-7 is not a second AM-6.
AM-8 kept N=10 because the measurement said so. Perturbing the RNG only
from its fourth construction on: --runs 2 PASSES, --runs 10 fails. A
late-onset divergence is deterministic, not flaky, so it is a control
rather than a coin flip. Ten runs live on one scenario (make am8, ~2s)
rather than all 25 (47s a build). GameKernel 5b records it.
The full run also found AM-4a's own mutation stale since ADR-0008 D3
moved the target 250,000 -> 161,000 in CB-WP-0013 -- reported
HARNESS-BROKEN, no score published. The build-free half of that check
is now a --self-test assertion, so make all catches the next one.
mutation-check clauses may now carry their own verify and mutation, and
then the enforced flag is measured rather than declared; a declaration
disagreeing with its measurement is refused.
make all exits 0.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Provenance (tier S, one paragraph in lieu of survey and ADR): the two
clauses tools/mutation-check.py has reported inert since CB-WP-0005 —
AM-7's scaling ratio (nothing relates the two throughput numbers
Criterion prints) and AM-8's N=10 (the runner does two). They are the
last two PARTIAL rows in the acceptance table. Structural tier S:
acceptance rows measure the product, and bench-test is in gates.toml's
not_control_gates list, so the M trigger about the loop's own
constraints does not fire. Chaos d4=2, no override. Declaration 10 of 12.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CB-EV-0012. Stage 1, deliverable by deliverable:
relationship-graph visualization emitted and gated, NEVER SEEN
drag-to-propose evidenced end to end
debug inspector evidenced (CB-WP-0011)
hot-seat play evidenced here
Hot-seat was the one closest to being claimed on the strength of the code
path existing. SeatPolicy hands every human seat a handle on one shared
Server, so turn-taking "obviously" worked — and nothing drove more than
one seat until now. The property that matters is not that two turns
happen but that the same tab, asked twice, shows two different hands.
Mutating the projection to serve P1's view to every seat turns it red.
The stage stays open on ONE named blocker rather than a vague
reservation: no browser is available to this loop, so the visualization
is evidenced only as correctly emitted. Everything testable from here has
been tested. What remains is `cb-play --serve 0`, open the URL, confirm
the table reads and a drag works. INTENT carries that note now.
The self-quoting rule from CB-EV-0011 §4 is ADOPTED: an evidence file
quotes the previous pass's final cost and never its own. CB-WP-0013
reported itself at $5.78/34 mid-flight; final is $8.26/47, under by 43%.
Four for four, always low.
Meta budget 29% [OVER] soft 25%, driven by CB-WP-0013 in a trailing three
with two cheap product passes; it was an instrument repair, which
ADR-0006 D2 exempts.
SH-1 at 347,720 [HARD] against a 300,000 ceiling. Compaction is the
remedy and this session cannot do it for itself. CB-EV-0009's standing
prediction is now live and testable for the first time in three passes.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
ADR-0009: embed quick-js; node is refused. Measured marginal cost against
the dev-toolchain graph, under the positive control:
boa_engine 896,410
rquickjs 69,985
quick-js 11,434
node 0 <- and that zero is the problem
ADR-0007 D3's acquisition rule biting its author. CI runs on rust:1.97,
which has no node, so the test would make our build fetch a JS runtime of
tens of millions of unaudited lines while scoring zero on the only
instrument that governs dependencies. A browser is exempt because a
developer has one regardless of us; a CI-installed runtime is not.
The loop is now closed: the real server serves the real page, QuickJS
runs that page's own scripts, the gesture goes over a real socket, and
the seat's Choice comes back. Before this, every link was tested and the
chain was not — a page whose JavaScript sent something else entirely
would have passed everything.
Three controls, each red for its stated reason: the JS posting a command
name instead of ids, the gesture not being delivered (EXPECT-VACUOUS),
and the token stripped from the endpoint.
A wrong assertion worth keeping: the first draft required the body not to
contain "attack". It legitimately does — action-attack is the id of an
element a finger landed on. An element may name an action; that is not
the page deciding. The real test is the shape: exactly two fields, down
and up, carrying two ids and nothing derived from them.
AND the ADR's own cost argument was wrong. It claimed 35% of AM-4b's
headroom; after landing AM-4b did not move at all. It measures
games-ground --edges normal — one package, no dev edges. Measured, the
workspace including dev edges is 725,258 lines against AM-4b's 317,021:
408,237 uncounted, MORE THAN THE TARGET ITSELF (criterion, clap,
ciborium, quick-js). The decision stands on the acquisition rule; the
affordability argument is withdrawn. Third defect in the AM-4 family.
Also fixed structurally rather than by raising a limit: `make status` had
grown past its 40-line readability gate as workplans accumulated. Closed
workplans now collapse to one line, so the report is fixed-size.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>