Every workplan is done and nothing is in flight -- the session stopped at
a boundary, not mid-change -- so the risk on pick-up is not lost context
but MISREAD context.
Records the three findings most likely to be misread (attack_relief is
unreachable-then-real, not null; H1's rejection covers half of H1; the
scope module is inert in round one), what is waiting on ground-game
versus what is ours, the conventions a new session trips over (hub is a
read model; workplan status is active not in_progress; make vendor;
replays/ is gitignored), and the two habits that paid for themselves
repeatedly -- mutate every control, and test the path a player takes.
Linked from the README above the gates, since it should be read before
design — the finding register
QUEUE (open findings)
repro F17 degenerate raised 6d ground-game
repro F18 inert raised 6d clay-borg
repro F24 inert raised 5d clay-borg
repro F25 inert raised 4d clay-borg
repro F26 inert ruled 4d ground-game
repro F27 unplayed reported 4d clay-borg
repro F31 unplayed reported 4d clay-borg
repro F30 degenerate ruled 4d ground-game
NOTES (not reportable — GameDesign §3.1)
F12 degenerate 11d
F15 underdetermined 7d
F21 degenerate 6d
findings 28 (+3 note(s))
with a resolving reproduction 19/28 = 67% target 100%
open, lacking a reproduction 0 target 0
reproductions green while open 0 target 0
notes past 30 days 0 target 0
closed (log) 20 [U1, U2, U3, U4, U5, U6, U7, U8, U9, U10, F11, F13, F14, F16, F19, F20, F23, F22, F29, F28].
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Reported to ground-game as a CORRECTION to GROUND-RPT-0004, not a new
result. Half of H1 was never measured: every game behind that verdict
was played by bots that never attacked, and H1-B acts only on an
uncancelled ATTACK at the stress gate. The rejection stands on H1-A, but
'H1-B does nothing' was never established and is false.
Also carries the structural finding -- a module can be unreachable
because of an aspect it does not name -- with an ask: should a module
declare its reachability preconditions, so a consumer reports 'not
reachable in this configuration' rather than a number that reads as a
null result?
Repeats the two unanswered items from RPT-0006: the competitive modes
are weakly tested (no bot models a rival), and problem_stress.scoped
cannot reach a decision in round one.
Left uncommitted in their tree, as with every prior report.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
GateAttackPolicy ranks ATTACK above GROUND at the stress gate and
delegates everything else. A PROBE, not a better bot: it wins 12/100
where greedy wins 60/100 under the same module, and the panel's banner
says so, because a column that looks like a policy comparison will be
read as one.
(a) THE MODULE IS UNREACHABLE IN THE PRINTED GAME. Selected alone it
still never fires: peak Stress never exceeds 2, so the gate never bites
and no ATTACK is ever gated. A module can be unreachable because of an
aspect it does not name -- attack_relief needs a problem_stress module
before it can act at all, and nothing in its own declaration says so.
(b) ONCE REACHABLE IT IS REAL. Holding the probe fixed and varying only
the module: group wins 12->21, 19->28, 50->57 at 3/4/6p, and DARVO arms
299->221, 294->213, 334->164. Under flat pressure the same shape appears
in DARVO alone: h1 arms 200 against flat_any_open's 300/400/600.
Stated as the conditional it is: GIVEN seats that attack at the gate.
Not advice to attack, and not evidence against F17.
AND IT CHANGES WHAT WE TOLD GROUND-GAME. CB-EV-0030 rejected H1 with
policies that never attacked, so H1-B was inert for every game behind
that verdict. The rejection stands on flat pressure alone -- wins are 0
either way -- but "H1-B does nothing" was never established and is now
known to be false. We owe them that correction.
Sensitivity: vary only the ATTACK rank. At 10 (greedy) and +55
(module-aware) the module never fires; at 110 it fires in every game.
Nothing about the module changed -- only whether any seat gave it a turn.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
module-panel sweeps points in aspect space the way scenario-panel sweeps
boards: one module at a time with every other aspect held, then the
profiles where an interaction is the hypothesis.
F31: attack_relief.self_soothe_ge4 HAS NEVER FIRED. Zero ATTACKs in
every cell, every seat band, both bots. The module acts only on an
uncancelled ATTACK by a seat at Stress >= 4, and greedy ranks Ground 100
at the gate against Attack's 10 -- so at exactly the position where
self-soothe would pay, GROUND wins. ModuleAwarePolicy adds 55 and that
is still not enough. `unplayed`, not `inert`: the kernel implements it
correctly and nothing has ever given it an opportunity.
This reaches backwards: CB-EV-0030's H1 verdict rests on its
flat-pressure half alone, because H1's other half never had an
opportunity in those games either.
Also: problem_stress.scoped is the only module with a measured effect
(3p 85->60 greedy, 86->63 module-aware, peak Stress 2->4);
flat_any_open drives group success to 0 at every band, reproducing the
H1 rejection from the module side; and scoped_plus_attack_soothe is
EXACTLY scoped alone -- necessarily, given F31 -- so the catalog's first
intentional multi-aspect combination cannot currently be evaluated as a
combination.
THE PANEL'S OWN DEFECT, first run: it printed 85/86 for a module that
never fired -- real-looking numbers inviting "measured, no effect" when
no seat ever created the precondition. Its docstring already said it
would not do that; the claim was written before the behaviour was.
First fix marked whole rows unmeasured, which threw away h1's real
flat-pressure result; the shipped fix names the specific module and
keeps the row's numbers, which are real for the modules that did fire.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
ModuleAwarePolicy reads the resolved Rules, so 'a bot that does not
attend to a mechanism cannot test claims about it' no longer holds for
modules. The MODE half stands: no policy models a rival playing their
objective, and that is the whole of what keeps F27 open.
Recorded the second limit found while closing the first: the scope term
is inert at round-one positions, so any measurement of
problem_stress.scoped weighted toward early rounds is measuring a
mechanism that has not started.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
ModuleAwarePolicy reads the RESOLVED Rules -- one policy, not one per
module. H2AwarePolicy would have been the blob schema 2 exists to
retire, rebuilt a layer up; reading Rules means a two-module
configuration gets both terms and a future module is one arm here rather
than a new policy per combination.
Two terms, both reading the table: prefer the Problem whose Stress falls
on me (problem_stress.scoped), and treat ATTACK as a Stress tool at the
gate (attack_relief.self_soothe_ge4). stress_scope and the owner marker
are public regardless of the card's face -- the H2 rules place the
marker ON the card -- so it is blind by construction.
THE CONTROL I CLAIMED WAS REAL WAS VACUOUS, AND MUTATION SAID SO.
"Under the baseline the two must be identical" swept fresh deals across
36 combinations, and forcing the scope term to fire regardless of
configuration left it GREEN. At a fresh deal only the Surface Problem is
face up, so exactly one SOLVE is legal and no ranking term can move the
argmax. The control could not distinguish the property from its
negation.
Both tests now use a BUILT position -- two face-up, unclaimed, equally
valuable Problems differing only in scope -- under H2 for divergence and
under the baseline for the control. The mutation goes red there.
That is also a finding about the module: the scope term is INERT at
round-one positions, so anything measuring H2 with round-one-heavy play
is measuring a mechanism that has not started. Recorded, not tuned away.
Also: CB-WP-0048 was `active` with every task closed; loop-lint's
lifecycle rule caught it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The three-armed enum could not express a configuration carrying two
modules. It does now: --variant h1 expands to problem_stress.flat_any_open
AND attack_relief.self_soothe_ge4, and scoped_plus_attack_soothe plays.
Legacy names alias forever via serde(alias="variant") plus a
scalar-or-map deserialiser. serde(default) alone would have been a silent
migration bug -- every H1/H2 recording would have come back as baseline.
All 26 scenarios pass unchanged; none pins a state hash.
Three things shipped broken first, all caught by gates rather than by
reading:
1. `profile` was in the state hash. with_config(select("ground-darvo-r0"))
sets profile=Some("baseline") where setup alone leaves None, so two
states at THE SAME POINT IN ASPECT SPACE hashed differently and a
replay bundle stopped reproducing its own initial state. Now
serde(skip): a hash covers what determines play. This was the open
judgement from T01 and it did not survive contact with the replay path.
2. A YAML parse inside the event loop. rules() -> resolve() -> catalog()
re-parsed catalog.yaml per rule check; AM-6 fell to 9,345 events/s
against a 100,000 target. OnceLock, and aspect validation moved to
where a configuration is BUILT.
Then I nearly optimised a phantom: 470k still looked like a 3x
regression against the "~1.7M on bnt-lap001" reference in the gate's
own message. Making rules() free measured 491k -- this machine's
ceiling. Before optimising against a reference, measure the ceiling
with the suspect code removed.
3. The refusal did not fire on the path a player takes.
`--module problem_deal.pressure_deck` played a full baseline game and
reported success, because with_config is a builder and fell back to
the printed rules -- the silent no-op ADR-0022 exists to refuse. Every
unit test of resolve() passed. The helper was tested and the driver
was not, which is CB-WP-0033's finding verbatim. The new test asserts
refusal BY NAME and that no game was played.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
vendor tool that covers what the gate checks
They ruled on all seven items the same day. Two were actionable here.
F28 RULED: points. Modes.csv MODE_COOP clarified upstream to say
"penalties apply to points, not card count"; mastery is now
total - blame - denied. A recorded scenario went red on it --
gr-e02-shared-ground pinned 0 (2 claimed CARDS - 1 - 1) and now expects
2 (4 POINTS - 1 - 1). The number moved because the rule was decided, not
because the engine drifted, and the scenario records both rulings; its
schema has no field for a second one, so both live in ruled_note with
`ruled` carrying the LATEST date.
F29 RULED not-intended and APPLIED upstream: SCN_02's suits re-tuned the
same day. The characterisation test is how we found out -- it pinned the
duplication, went red on the re-tune, and that red WAS the notification.
It now asserts every pair distinct, the stronger statement the
duplication had made unavailable. SCN_02 re-measures at 73 at 2p, not
67: its own board now.
F26/F30 ruled and recorded. F30's ruling incidentally confirms our
reading -- they name priority-2's suit as the first lever, which is the
difference we identified without having measured causation.
vendor-editions grew twice, both times because it covered less than the
gate it exists to satisfy:
- It refused to touch ground-darvo-r0/ on the reasoning that the
baseline is "a separate record". That was wrong within the hour:
ground-game clarified Modes.csv and `make vendor` reported a clean
sync while edition-check went red. A sync tool that covers less than
its check reports success into a red gate.
- Its two-block rewrite DETECTED which fence held which set and
preserved the arrangement -- faithfully preserving a swap an earlier
write had introduced, leaving each fence under a heading describing
the other. edition-check reads every sha256 line flat and passed
throughout: a document can be self-consistently wrong and green.
Order is now asserted, with a control that goes red on a swap.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The spec is explicit: a WORKPLAN is active|done|paused; in_progress is a
TASK status. I used the task vocabulary on the workplan frontmatter, and
the hub rejected both with a 422 on every sync.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Five things worth their time, four of them asks rather than statements:
F29 SCN_01 and SCN_02 are the SAME BOARD -- identical suit and value at
every priority, every cell matching exactly. Not called a defect (a
reskin is legitimate) but "four scenarios" is three boards. Asked
whether it is intended.
F30 SCN_04 is materially harder at 2p: 52/100 against 67 and 73, with
every other parameter held by the edition itself -- same deal shape,
same 6 available points, same threshold, same starting Stress. The lone
difference is that it is the only 2p deal needing two of one suit.
Causation explicitly NOT claimed; the falsifier is stated. Sensitivity:
at 4p SCN_04 is 97 against 94/99, so it is not the hard board there.
The modes finding, corrected in their favour: we previously reported the
three modes produced identical play. That was OUR INSTRUMENT, not their
game -- the bot never read the Mode card. With a mode-aware bot, group
success is unchanged in 34 of 36 cells but coalition size moves by half
again at 4p. So the modes decide the distribution of the win and the
threshold decides survival independently of it.
F28 (mastery counts cards where the shared score counts points) and F26
(a package that adds a FILE is invisible to a consumer) promoted to
reported -- both need their ruling, neither changed on our side.
F17 and F26 move raised -> reported now that they are in a delivered
report.
The register crossed the ~400-line loadability limit, so prose for
CLOSED findings moved to FindingRegister-closed.md. The rows are
untouched and `make design` still reads one file -- ADR-0012 D5's
reasoning about mixing open and closed applies to files too.
Report left UNCOMMITTED in ground-game: their tree has live uncommitted
work from their own agent.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
objective() reads GroundState::score (now public) rather than restating
what winning is; a copy in the bot would disagree with the kernel the
first time ground-game rules on F28.
Working out WHERE the modes can differ was most of the task and it
bounds the result: SOLVE always claims for the actor, so own-score and
group-score want the same SOLVE nearly everywhere. That is a fact about
GROUND's action set, not a shortcoming of the bot. Two real divergences,
both readable off the table: SUPPORT regulates someone else (worth less
against a rival, worth MORE under coalitions where a Bond merges them
into my side), and SOLVE's value is the card's value, which greedy
ignores entirely.
THE RESULT — F27 splits in two:
group success UNCHANGED in 34 of 36 cells
who wins MOVES: BONDED COALITIONS at 4p goes 2.04 -> 2.98,
2.12 -> 3.29, 2.05 -> 3.01 winning seats per game
So "the competitive modes are scoring lenses over cooperative play" was
too strong and is withdrawn. The sharper claim: GROUND's scoring modes
change WHO WINS, not WHETHER THE GROUP SUCCEEDS. And the effect is
seat-band dependent -- 2p none, 4p largest, 6p none under coalitions;
two relation slots capping network growth is a candidate explanation and
is untested.
The panel now prints BOTH policies side by side. That was a correction
mid-task: the first version printed only the new one and I compared it
against a figure remembered from CB-WP-0047 -- a comparison against a
board nobody re-ran.
Control that makes the numbers mean anything: under SHARED GROUND the
two policies agree at all but <=2 decision points across 12 boards, so a
moving column is mode-awareness and not simply a different bot.
Also: two T01 tests keyed on `status: proposed`, which ground-game
renamed to `ready-for-implement` mid-session. They now find the module
by asking resolve() -- the structural property is ours and does not move
when another repo edits its vocabulary.
Also: `make vendor` replaces three hand re-vendors with a tool that
regenerates digests by walking editions/, and reports one-sided files
rather than resolving them.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Policy::choose takes the whole GroundState -- every face-down Problem's
suit and value, every seat's hand -- which is exactly what project()
exists to withhold. No shipped policy reads it, but the first thing a
competitive policy must do is VALUE a Problem, and value is the hidden
field. The trap goes live on the first line of the F27 work.
Established behaviourally rather than by narrowing the trait: vary only
what the seat cannot see, and the choice must not move. That binds every
policy including ones written later and outside this crate, without
their cooperation. The mirror of ADR-0013 D1 -- same kernel, two
searches, opposite permissions, discriminated by WHEN the question is
asked; a policy plays from inside an information set, so retrospective
permission would be strategy fusion.
Running the control found two defects IN THE CONTROL:
1. The rearrangements rotated hidden values 2<->3 together, leaving max
invariant -- so the deliberate peeker, which ranks by the largest
hidden value, was not caught. A control whose variation is invariant
under the statistic a violator reads is not a control.
2. It accused `random` of peeking, because it reused one policy instance
and compared a first call against a fourth. It takes a constructor
now, so every variant is judged from identical policy state.
Both are a difference in output read as evidence about hidden state --
the wrong-subject family, found twice inside a control written to detect
wrong subjects.
Three mutations, three red. The control is proven against a deliberate
violator before being trusted about compliant policies.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
catalog.rs reads ground-game's schema 2 and refuses an unknown schema by
name: a reader that silently accepted schema 1 would answer questions
about aspects over a file that has none.
config.rs holds the two types ADR-0022 chose. Configuration is identity
and round-trips anything the catalog names, including modules with no
kernel path; Rules is behaviour, exhaustive, no catch-all. resolve() is
the boundary and emits the two distinct errors -- "known module with no
kernel path (status: proposed)" vs "not a module the catalog has" --
which is the entire reason this shape was chosen over a per-aspect enum.
Nothing about aspects, defaults or module status is written in our
source; all of it is read. The tests find the proposed module by
searching for status: proposed rather than naming one, so implementing
it upstream makes the test look elsewhere instead of going stale.
Four mutations, four red. scoped_plus_attack_soothe now resolves -- the
combination the three-armed enum could not express.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
status.py resolves `## Task: <text>` immediately above a task block;
`## Task T00:` resolved to nothing and the gate said so.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The maintainer's observation decided the design: aspects partition the
GAME, strata partition our apparatus, and they are orthogonal. A module
is one coordinate change in aspect space with an obligation in every
stratum. So aspect identity must NOT be Rust types -- an aspect
ground-game adds would make clay-borg fail to parse a configuration
rather than fail to run it, welding the two coordinate systems at the
one place they must stay independent.
Chosen: identity as data (Configuration round-trips anything the catalog
names), behaviour exhaustive (Rules, no catch-all), resolve() between.
Decisive argument: the catalog ALREADY ships modules with a rules_delta
and status: proposed, so a per-aspect enum would report them as "unknown
module" -- indistinguishable from a typo, a false statement about the
edition, and this project's signature failure shape. Two facts need two
errors. Federating design authority is permanent, so the representation
must outlive the implementation.
Legacy ids alias forever through the catalog's own legacy_experiment_id,
on the standard-Np precedent: 26 recordings name them and the expansion
is exact, so there is nothing to deprecate.
T00 done: the schema-2 mirror had arrived with no digests (19 files) and
edition-check was red. Digests are now generated by WALKING editions/,
not typed -- two reviews already found hand-written lists that made
their own controls vacuous, and a mirror that grows a directory is what
breaks a maintained list. PROVENANCE-catalog.md was a file inside the
mirrored tree that upstream does not have; folded into our own
PROVENANCE.md, since provenance about the mirror does not belong inside
the thing it describes.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The modes were already implemented; nothing had ever COMPARED them. The
scenarios were not implemented at all: edition::deal has taken a
scenario_id since it was written and the only caller passed the literal
"SCN_01", so 15 of 20 Problem cards had never been dealt by anything.
The seam was the whole mechanism and it sat unused, with nothing red
because nothing asked.
Scenario is now state (serde default SCN_01, so all 26 recordings replay
unchanged), selected by preset `scn-03-4p` with `standard-Np` still
meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or
titles, validated against the edition rather than a pattern.
The threshold now comes off the Scenario card, closing F25's hardcoded
5/7/9. The first version of that control was worthless and mutation said
so: all four scenarios print 5/7/9, so reverting to the bands left it
green. Split threshold_from() so it can be handed a card that disagrees.
The header read `scoring CommonProblem` where the Mode card is titled
COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the
move buttons, still standing on the line that says what winning means.
The coverage probe was matching that Debug output and went red when it
was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the
premise, the mode's rules text, and the tiebreak.
scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same
board (identical cells, pinned by a characterisation test); SCN_04 is
the hard board at 2p (52% vs 67/73%, the only deck needing two Repair);
and group success is EXACTLY equal across all three modes in all 36
cells, because greedy never reads state.mode -- filed F27, the two
competitive modes are scoring lenses over cooperative play.
F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT
where the mode card's shared score is claimed VALUE. Raised, not fixed;
scoring is ground-game's to rule on.
Also fixes design.py reporting a backticked path as no reproduction.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CB-WP-0045 left "nothing explains what a scope does" open and gave a
FALSE reason: that scoping is our Variant so no printed sentence exists.
The H2 package ships Rules_Text.csv -- twenty-two passages of
player-facing rules -- and nothing in clay-borg had ever read the file.
A wrong reason for an open item is worse than an open item; it retires
the question. Filed as F26 against ground-game: a package that adds a
FILE is invisible where one that adds a column is not.
The scope rule now renders under the table in the edition's own words,
only when a non-global scope is in play, matched by heading rather than
row number, and absent (never paraphrased) if the edition drops it.
The trial log stamps its variant on the begin marker -- a session
property, not an eighth column -- read off state.variant rather than the
--variant flag, because a bare `state.variant = v` leaves H2 inert and a
flag-stamped log would put false provenance on real player words. An
unstamped log reports `unrecorded`, never `ground-darvo-r0`.
Six mutations, six red. The sixth is the finding: every trials.py fixture
built its marker out of BEGIN, so nine checks followed BEGIN away from
what hotseat.rs writes and stayed green while real logs broke. A fixture
built from the constant under test cannot test the constant -- the
control is now a literal, asserted from both sides.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three reports from the 2026-08-08 sessions. H2 lands ("that seemed more
interesting"), two legibility gaps do not.
The number was only a FALLBACK for a missing title, so it showed exactly
when it was least useful. That was survivable while Problems sat in one
ordered row -- position WAS the number. CB-WP-0043 scattered them by
scope and took the implicit index with it, with nothing going red because
no test named the number.
DARVO.csv's mandatory_effect and Actions.csv's GROUND rules_text were
both vendored, both used only as tripwires, and neither ever reached the
player -- F18's shape again. The modes now explain themselves at the
point of choice, and NOT when no mode is on offer.
Tests assert the edition's own words verbatim (ADR-0015), the number over
every key in view.problems, and both halves of the mode explanation.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Reported as "I did a session and noted no changes" — correct, and the
fault was ours twice.
make ground passes no --variant, so the session was baseline, and
CB-WP-0043 deliberately leaves the baseline layout untouched because a
variant that redraws the baseline invalidates every prior look at it.
But the deeper defect is that nothing said so. The page named the round,
the step, the Lead, the scoring mode and the viewer, and never which rules
it was playing — GroundView did not even carry the variant. There was no
way from the screen to tell baseline from H1 from H2. A session that
cannot say which rules it is playing cannot report a rules change; the
player did the right thing and the instrument had nothing to tell them.
Now: variant on GroundView, `rules <id>` in the page header beside the
scoring mode, `rules <id>` in the inspector because a replay that cannot
say which rules produced it is the same defect in the other tool, and
make ground VARIANT=h2 so the capability is reachable.
Both coverage probes caught the new field independently — the render
crate's and cb-play's — the second time in two passes that they have
turned an addition into a legibility requirement instead of letting it be
silent state.
Verified by fetching the served page rather than by reading the code:
make ground VARIANT=h2 prints "rules: h2" and the page carries
"rules h2-scoped-problem-stress" with the scope labels; the baseline says
"rules ground-darvo-r0" and keeps its row.
Still open: nothing explains what a scope DOES, and the trial log header
does not record the variant either — the same defect one artifact along.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The picture had become a lie. A row across the middle says "these are the
table's" — true under the baseline, false under H2, where three of five
cards fall on particular seats.
Global goes to the middle, personal beside its owner, and bond on the mean
midpoint of the owner's Bond edges. A bond Problem whose owner has no Bond
is drawn as personal, because that is the rule — bond_network falls back
to the owner at degree 0 — and the picture must agree with the arithmetic.
The baseline page is untouched: the scoped layout is taken only when a
non-global scope exists, so every prior look at the baseline still holds.
The scope reaches the view as a separate marker rather than a field on
ProblemView, because the delta says "place owner marker on the card" — a
token beside a card — and it is public while the card is face down, which
a ProblemView variant could not express.
The coverage probe forced a real improvement. It demanded a text token for
the new fields and a POSITION is not a token — which is the probe being
right: position alone is invisible to text_of and to a screen reader, and
illegible when two anchors coincide. So each scoped Problem now says whose
it is: everyone's, P1's alone, P2's Bond network, and "P1's alone — no
Bond to share it" at degree 0.
The fixture gained a third Problem. With two, one of the three placement
rules was unexercised and the probe unsatisfiable — a fixture that cannot
reach a branch is how a rule ships untested.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The H2 panel measured wins, variance and DARVO but not ATTACK selection —
which is F17's actual question. attack-value now runs H2 too.
The shape is the finding. Under H1 a seat that sometimes attacks loses
everything at 3p and above. Under H2 rank-75 wins 112/173/199 against
greedy's 120/175/199, while attacking and arming DARVO. So H2 makes
occasional ATTACK affordable — it does not make it pay. rank-75 never
beats greedy in any cell, and rank-95 (always attack) still wins 0
everywhere in all three variants, so "not always-attack-optimal" holds.
F17 therefore stands: ATTACK earns its place in no mode. What changed is
that choosing it is no longer catastrophic. Whether affordable is what the
design wants is ground-game's judgement.
And a hazard: with_variant() exists because H2 assigns Problem owners at
setup and `state.variant = v` leaves them unassigned, so scoped pressure
ticks nobody and H2 measures as INERT. Three call sites had the bare
write, including cb-play's driver.
No published figure is affected, and that was checked rather than assumed:
h2-panel used the builder, and the two harnesses with the bare write had
only ever run baseline and H1, neither of which has a setup step; the
driver has never played H2. All three fixed, and
a_bare_variant_write_leaves_h2_inert now states the difference so a
regression is caught by a named test rather than by a reader wondering why
H2 did nothing.
The builder was not enough — the field is public, so the old form still
compiles. Worth knowing before the next variant.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Measured against ground-game's own §3 criteria, read from their design
note rather than reused from H1.
Criterion 1 met with room: greedy SHARED wins 120 at 3p and 175 at 4p,
against H1's 0 and 0, restoring 73% and 92% of baseline.
Criterion 3 met, and it was H1's clearest failure. Under H1 the
unregulated seat armed DARVO constantly and never won; under H2 it arms
and wins 13/13/57. "Non-zero for some policy that still sometimes wins" is
exactly the shape H1 could not produce.
Criterion 2 met at 2p/4p/6p and missed at 3p — 1.50 against baseline's
1.57 — reported as a miss because that is what this sample says. The
mechanism is visible: H1-greedy's spread is 0.00 at 3p+, because a flat
tax on every seat creates no variance at all. That is the clearest
statement of why scoping was the right correction.
Criterion 5 is the best evidence in the pass. Forcing every scope to
global and changing nothing else reproduces H1's collapse exactly — 120 to
0 at 3p, 175 to 0 at 4p — so the scoping is what saves it, not any other
difference between the packages.
Criterion 4 came out backwards and the prediction held. The workplan said
this panel might be unable to test it, because no policy here models
another seat or knows what a scope is; bond claim rates are LOWER than
personal at 3p and 4p, driven by suit availability rather than incentive.
Reported as untested with an incidental figure pointing the wrong way,
not as a refutation.
Wired into make panels. First pass declared after ADR-0021, so no chaos
roll is recorded.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
H2 is ground-game's answer to our H1 reading — that a flat +1 to every
seat is a solve-rate tax scaling with the number of Problems. Unclaimed
Problems now tick only the seats in scope: global (all), personal (the
owner), bond (the owner's Bond network over Bond edges only, degree 0
falling back to personal), assigned by hidden priority so 2p never has the
bond card in play.
T01: the package is vendored with digests, and H2's Problems.csv is r0's
with one column added and NOTHING else changed — checked, not assumed,
because the delta claims deal_and_thresholds unchanged and a silent
difference would make every H2-vs-baseline comparison a comparison of two
boards as well as two rule sets. Scopes are read from the column, not
derived from the priority in Rust: F25 exists because we hardcoded numbers
the edition already carried.
T02: owner and scope are new ProblemState fields, both Option and both
skipped when None, so a baseline state serialises without them and every
recorded scenario's hash is untouched — asserted on the JSON, not assumed.
with_variant() replaces the bare field write, because state.variant = v
would leave owners unassigned: a silently wrong game rather than a failing
one.
T03: every named defect is mutation-proven — traversing Rivalry edges,
applying stacking once, a degree-0 owner ticking everyone, personal
hitting everyone. The degree-0 mutation MISSED first: the fallback lives
inside bond_network and the mutation broke the None-owner arm instead, a
different branch. It stayed green until aimed at the path the test
exercises. A mutation that misses is not evidence the test works.
T04: ownership is not a permission. Filtering SOLVE to the owner turns it
red, which is the regression this task exists for — the engine had no
owner concept before T02 added one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The condition is met and the tally was verified rather than recalled.
ADR-0017 D2 named window 3 as the decider; windows 2 and 3 each produced
exactly one override and each changed nothing. Every roll in window 3 was
cross-checked against the workplan that made it, because the last time
this project tallied chaos rolls from memory it was wrong and asserted the
wrong figure four times (F23). All twelve agree.
But meeting the condition is not evidence. P(an override changes nothing)
is 1/3, and each window had exactly one override, so the condition fires
on a 1/9 coincidence. ADR-0017 restated it to be REACHABLE and made it
weak in the process; reachability was checked and discriminating power was
not.
Worse, it measures the wrong subject. It asks whether an override changed
the tier; the question is whether changing the tier helped. Under it, a
die that always changed the tier could never be retired however useless
its changes were.
The real ground is stronger. Four overrides across roughly forty
declarations, and the mechanism's value has never once been demonstrated.
The one substantive intervention dropped CB-WP-0011 from a structural L to
S, and that work then needed CB-WP-0016 and CB-WP-0017 to fix defects a
human found by playing. Not offered as causation — a tier is process
weight, not a guarantee — but it is the only evidence we have about an
override's consequences and it points the wrong way.
And the purpose has no live evidence of need: 17 M, 11 S, 4 L across every
workplan, with the only two structural/declared mismatches being the
window-1 overrides themselves. Tier declaration has not ossified.
So: retired, with nothing replacing it. Adding a successor to guard
against ossification that has not occurred would invent a gate for an
instance we do not have. The revival trigger is stated: tiers collapsing
toward one value, or a pass declaring below its structural tier to dodge a
review.
InnerLoop loses the chaos paragraph, loop-lint loses the chaos-recorded
check — a check that outlives its rule becomes an obstruction — and its
self-test now asserts the opposite: a note with no roll must pass.
ChaosRollHistory is closed. Window 4 ends incomplete at four declarations,
the last of which is this one, rolled because the rule was still in force.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
T02 — all chance derives from one root seed. Three chance points, all
reading it: the setup deck shuffle, the setup Lead draw, and the reshuffle
permutation. The Problems deal is not chance at all. So in extensive-form
terms the tree has a single chance node at the root.
That test was wrong first, and the mutation caught it. It compared state
hashes — and GroundState carries `seed` as a field, so "different seeds
differ" was true by construction. Mutating the shuffle away left it green.
It now compares the dealt configuration, and the same mutation fails it: a
wrong-subject error inside the control written for T02.
The reshuffle is a pure function of (seed, round) because K5 requires
deterministic replay, where a real table reshuffles independently. That is
a modelling restriction, not a defect, and it is now pinned.
T03 — commit/reveal checked in both directions: before Reveal each seat
sees its own selection and no other; after Reveal the information sets
merge, because an encoding that hides forever is not commit/reveal either.
T04 — ADR-0020 refuses the EFG port, and the blocker is T02 rather than
T01, which inverts what the workplan expected. Perfect recall looked like
the risk and is a constraint with a known answer: key on observation
histories. Making chance explicit is the expensive one — the reshuffle
would become a real chance node and break the K5 purity that every
recording, replay bundle and trial-note hash depends on. A port would
trade the property this project is built on for one it has never needed.
Track B's first move is therefore a question, not a build: take "is
exploitability meaningful for a co-operative game with a shared threshold"
to OpenSpiel on a toy model, where answering it costs nothing. D4 states
what being wrong looks like — OpenSpiel settling on a toy what three
rounds of policy sweeps could not — and makes watching for it the next
action.
Taxonomy §4.1 records the EFG correspondence with the test that checks
each row, so a later pass starts from a specification rather than a memory.
Chaos window 4 at three declarations. Window 3's verdict is now two
windows behind and should be evaluated rather than restated.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>