ModuleAwarePolicy reads the resolved Rules, so 'a bot that does not attend to a mechanism cannot test claims about it' no longer holds for modules. The MODE half stands: no policy models a rival playing their objective, and that is the whole of what keeps F27 open. Recorded the second limit found while closing the first: the scope term is inert at round-one positions, so any measurement of problem_stress.scoped weighted toward early rounds is measuring a mechanism that has not started. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
14 KiB
| id | kind | title | status | state_hub_workstream_id |
|---|---|---|---|---|
| CB-WP-0049 | product | A seat that plays its own objective | done | 92b98076-65c2-4db3-8ab0-6f0139d4f893 |
Purpose
structural tier M (adds a measurement instrument whose output changes
what published panels mean)
declared tier M
Closes the gap CB-WP-0048 opened: configurations are now selectable and nothing can judge them.
The finding this exists to answer
F27. scenario-panel measured 4 scenarios × 3 modes × 3 seat bands
and found group success exactly equal across all three scoring modes in
all 36 cells. The arithmetic is right and the reason is simple:
GreedyPolicy never reads state.mode. The same games are played and
only the winner set is carved differently.
So the two competitive modes are scoring lenses over cooperative play, and the same will be true of every aspect module: a bot that does not attend to a mechanism cannot test a claim about that mechanism's incentives. CB-RES-0010 §5.4's last bullet is F27 restated from the design side.
This is now the binding constraint on the whole programme. scoped_plus_attack_soothe
is selectable today (CB-WP-0048 T01) and no instrument in this repo can
say whether it is any good.
The prerequisite nobody had noticed
Policy::choose takes &GroundState, which carries every face-down
Problem's suit and value and every seat's hand. project exists to
withhold exactly those.
No shipped policy reads them. But the first thing a competitive policy must do is value a Problem, and value is the field the projection hides — so the trap becomes live on the first line of the work this workplan is about. ADR-0023.
Task: a policy is bound by what its seat can see
id: CB-WP-0049-T01
status: done
priority: high
state_hub_task_id: "2b84e2d1-a26a-4cfa-bc96-21b40cb487df"
Per ADR-0023: established behaviourally, not by changing the trait — vary only what the seat cannot see and the choice must not move. That binds every policy, including ones written later and outside this crate, without their cooperation.
Done 2026-08-08. bot::blindness provides hidden_rearrangements
and is_blind; GreedyPolicy and RandomPolicy are proven blind across
three seeds × three seat bands × every seat.
Two defects in the control itself, both caught by running it
1. The first rearrangement was invariant under what a violator reads.
It rotated every hidden value 2↔3 together, which leaves max
unchanged and (at equal counts) sum too — so the deliberate peeker,
which ranks by the largest hidden value, was not caught. A control
whose variation is invariant under the statistic a violator reads is not
a control. The family now forces all-2, all-3 and two alternating splits,
so any symmetric function of the hidden values moves across it.
2. It accused random of peeking. is_blind reused one policy
instance, so it compared a first call against a fourth — and
RandomPolicy advances its own stream. It takes a constructor now,
so every variant is judged from an identical starting policy and the only
difference between two runs is the state.
Both are the same mistake in different clothes: a difference in the output was read as evidence about the hidden state, when the first was insensitive to it and the second was caused by something else entirely. That is the wrong-subject family, found twice inside one control written to detect wrong subjects.
| mutation | what went red |
|---|---|
| rearrangements leave the multiset invariant | "a policy reading face-down values must be caught" |
| rearrangements vary nothing at all | "no rearrangement changed a hidden card" + the peeker escapes |
| one policy instance reused across variants | "random chose Command(5) and then Command(2)" |
The control is proven against a deliberate violator before being trusted about compliant policies (ADR-0023 D3) — a peek control nothing can fail is decoration.
Task: an objective, taken from the game's own scoring
id: CB-WP-0049-T02
status: done
priority: high
state_hub_task_id: "7e3611bf-4cfb-412c-8e33-695778c7980b"
What a seat is trying to maximise, per ScoringMode:
| mode | objective |
|---|---|
| SHARED GROUND | the group total |
| COMMON PROBLEM | own claimed value − own Blame |
| BONDED COALITIONS | own coalition's summed personal score |
Derived from GroundState::score, never reimplemented. All three are
already computed there for the outcome; a second copy in the bot would be
a second definition of winning, and the two would disagree the first time
ground-game rules on F28.
Controls:
- the objective agrees with
Outcomeon a finished game, for every mode — asserted against the real scoring rather than a fixture; - a seat's objective is not the group's under the competitive modes, shown on a concrete board rather than argued;
- the objective is blind (T01's control applies to anything that reaches a policy).
Done 2026-08-08. GroundState::score is now public and
bot::objective reads it. Nothing about winning is restated in the bot.
Because score is a pure function of the state it answers mid-game too —
"what would this position score if it stopped now" — which is what a
policy needs and what a second implementation would have had to invent.
Task: a seat that plays it, and the F27 re-measurement
id: CB-WP-0049-T03
status: done
priority: high
state_hub_task_id: "a842c0a3-9374-4f08-8e2b-1571b8f241ef"
A policy that maximises its own objective, delegating to
GreedyPolicy::rank wherever the objective is silent — delegating, not
copying, which CB-WP-0039 learned the hard way when a re-typed "greedy
with one preference changed" differed in five places and burned the
Freedom token in round one of every game.
Then re-run scenario-panel.
The result is not predicted here. Either the modes now diverge — F27 resolves, and the competitive modes are real — or they still do not, and that is a much more interesting finding about GROUND: that the game's scoring modes do not reach its decisions. Both outcomes are publishable and the workplan must not be written as though one is expected.
Controls:
- the panel's group-success column moves, or it is reported as not moving — a null result stated as a null result;
- the competitive policy is blind (T01);
- greedy's numbers are unchanged, so the new policy is an addition and not a silent edit of the published baseline.
Done 2026-08-08. ObjectivePolicy delegates to GreedyPolicy::rank
and overrides two arms.
Where the modes can differ at all
Working this out was most of the task, and it bounds the result: SOLVE always claims for the actor, so a seat maximising its own score and one maximising the group's want the same SOLVE in nearly every position. That is a fact about GROUND's action set, not a shortcoming of the bot. Two places the objective genuinely diverges, both readable off the table:
- SUPPORT regulates someone else — worth less when the beneficiary is a rival (COMMON PROBLEM), worth more when a Bond merges them into my coalition and my score is the coalition's sum (BONDED COALITIONS).
- SOLVE's value is the card's value — greedy ranks every legal SOLVE at 90 regardless, and under a competitive objective the difference between a 2 and a 3 is the margin.
The result — F27 splits in two
The panel now prints both policies side by side. That was a
correction mid-task: the first version printed only the new policy's
win/g and I compared it against a figure remembered from CB-WP-0047,
which is a comparison against a board nobody re-ran.
| group success | unchanged — 34 of 36 cells identical; SCN_03 at 4p goes 99→100 in all three modes, which is the SOLVE-by-value refinement, not a mode effect |
| who wins | moves — BONDED COALITIONS at 4p: 2.04 → 2.98, 2.12 → 3.29, 2.05 → 3.01 winning seats per game. COMMON PROBLEM at 6p: 1.10 → 1.17 |
So the original claim was too strong. The modes are not "scoring lenses over cooperative play" — they reach decisions. The sharper statement: GROUND's scoring modes change who wins, not whether the group succeeds.
And the effect is seat-band dependent: 2p shows no divergence in any mode, 4p the largest, 6p none under BONDED COALITIONS. A candidate explanation is that two relation slots per seat cap network growth, so at 6p the incentive exists and cannot be acted on. That is untested and is the next variable to move.
Under SHARED GROUND the two policies agree at all but ≤2 decision points across 12 boards — the control that separates "attends to the objective" from "is simply a different bot". Without it, every number above could have been the new policy being different rather than being mode-aware.
Task: module-aware evaluation
id: CB-WP-0049-T04
status: done
priority: normal
state_hub_task_id: "1b26c7e9-e049-49d0-aecd-63830c7c8a37"
A seat under problem_stress.scoped should prefer clearing a Problem
whose scope includes it; under attack_relief.self_soothe_ge4, ATTACK
becomes a Stress tool at the gate. Both are visible to the seat
(ProblemMarker is in the projection), so neither needs hidden state.
The composition risk is the same one schema 2 exists to retire. One
policy per module is H2AwarePolicy — the blob at a new layer. Evaluation
should compose per aspect the way a configuration does, and if that is not
achievable cheaply, the honest move is one policy that reads the resolved
Rules rather than several that each assume one.
Controls:
- a module-aware seat differs from greedy ONLY where the module is live — under the baseline the two must be identical, which is the control that separates "attends to the module" from "plays differently";
- blind (T01).
Done 2026-08-08. ModuleAwarePolicy reads the resolved Rules —
one policy, not one per module. H2AwarePolicy would have been the
blob schema 2 exists to retire, rebuilt a layer up; reading Rules means
a configuration carrying two modules gets both terms and a module the
kernel gains later is one arm here rather than a new policy for every
combination it appears in.
It delegates to ObjectivePolicy::rank, which delegates to
GreedyPolicy::rank, so it differs in exactly two arms:
| module | term |
|---|---|
problem_stress.scoped |
prefer the Problem whose Stress falls on me — two equally valuable cards are not equally urgent, which is the module's whole point |
attack_relief.self_soothe_ge4 |
at the gate, ATTACK is a Stress tool; greedy ranks it 10 in every position because under the printed rules it does nothing for the attacker |
Both read the table: stress_scope and the owner marker are public
regardless of the card's face, because the H2 rules say to place the
marker on the card and at a table everyone can see it. So it is blind
(ADR-0023) by construction, and the T01 sweep now covers it.
The control I claimed was real was vacuous, and mutation said so
"Under the baseline the two must be identical" swept fresh deals across four seeds × three seat bands × three modes — and forcing the scope term to fire regardless of configuration left it green.
At a fresh deal only the Surface Problem is face up, so exactly one SOLVE is legal and no ranking term can move the argmax. The control could not distinguish the property from its negation, which is the definition of decoration (ADR-0006 D3).
Both tests now use a built position — two face-up, unclaimed, equally valuable Problems the seat can solve either of, differing only in scope — played under H2 for the divergence and under the baseline for the control. The mutation goes red there.
This is also a finding about the module, not only about the test. The scope term is inert at round-one positions: the module cannot reach a decision until more than one Problem is solvable, so anything measuring H2 with round-one-heavy play is measuring a mechanism that has not started. Recorded rather than tuned away.
Two things this pass fixed that were not the task
A test keyed on a word another repo owns. T01's controls found their
example module by searching the catalog for status: proposed — and
ground-game renamed those to ready-for-implement mid-session, so two
tests failed over a string neither was about. The property they actually
need is "a module the catalog has that this kernel cannot run", which
resolve() determines and which does not move when someone edits their
vocabulary. Same lesson as matching a rules passage by heading rather
than row number (CB-WP-0046) and reading a module's aspect: field
rather than splitting its id on the dot.
The mirror was re-vendored by hand three times in one session.
make vendor now does it and regenerates the digest block by walking
editions/, never by typing it. It reports files only one side has
rather than resolving them — deleting a file upstream dropped, or copying
one it added, are decisions (ADR-0011), not a sync.
Its self-test caught a defect in itself on first run: it read both
digest blocks in PROVENANCE.md and reported ten edition CSVs as
"recorded but absent", when they are absent only from the sibling walk
— a different question from absent. Mutation-proved: touch a vendored
file without re-running and "every recorded digest is the file's current
bytes" goes red.
Not done here
- The trait still takes
&GroundState. ADR-0023 D2 chose the behavioural control first, on the grounds that it binds policies the type system cannot reach. Narrowingchooseto&GroundViewremains worth doing and is not refused — it is simply not what establishes the property. - Opponent modelling is out of scope. These policies play their own objective; none of them models another seat playing theirs. A competitive mode where nobody anticipates a rival is still a weak test of that mode, and saying so is part of T03's result.
- No lookahead.
search.rsis retrospective and may not be reused here: ADR-0013 D1 permits it to see everything precisely because after the game there is one world. A policy plays from inside an information set, so the same permission would be strategy fusion. Same kernel, two searches, opposite permissions, and the discriminator is when the question is asked.