197 lines
7.9 KiB
Markdown
197 lines
7.9 KiB
Markdown
|
|
---
|
|||
|
|
id: CB-WP-0049
|
|||
|
|
kind: product
|
|||
|
|
title: "A seat that plays its own objective"
|
|||
|
|
status: in_progress
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
# Purpose
|
|||
|
|
|
|||
|
|
```
|
|||
|
|
structural tier M (adds a measurement instrument whose output changes
|
|||
|
|
what published panels mean)
|
|||
|
|
declared tier M
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Closes the gap [CB-WP-0048](CB-WP-0048-a-configuration-not-a-variant.md)
|
|||
|
|
opened: configurations are now selectable and **nothing can judge them**.
|
|||
|
|
|
|||
|
|
## The finding this exists to answer
|
|||
|
|
|
|||
|
|
**F27.** `scenario-panel` measured 4 scenarios × 3 modes × 3 seat bands
|
|||
|
|
and found group success *exactly* equal across all three scoring modes in
|
|||
|
|
all 36 cells. The arithmetic is right and the reason is simple:
|
|||
|
|
`GreedyPolicy` never reads `state.mode`. The same games are played and
|
|||
|
|
only the winner set is carved differently.
|
|||
|
|
|
|||
|
|
So **the two competitive modes are scoring lenses over cooperative
|
|||
|
|
play**, and the same will be true of every aspect module: a bot that does
|
|||
|
|
not attend to a mechanism cannot test a claim about that mechanism's
|
|||
|
|
incentives. CB-RES-0010 §5.4's last bullet is F27 restated from the
|
|||
|
|
design side.
|
|||
|
|
|
|||
|
|
This is now the binding constraint on the whole programme. `scoped_plus_attack_soothe`
|
|||
|
|
is selectable today (CB-WP-0048 T01) and no instrument in this repo can
|
|||
|
|
say whether it is any good.
|
|||
|
|
|
|||
|
|
## The prerequisite nobody had noticed
|
|||
|
|
|
|||
|
|
`Policy::choose` takes `&GroundState`, which carries **every face-down
|
|||
|
|
Problem's suit and value and every seat's hand**. `project` exists to
|
|||
|
|
withhold exactly those.
|
|||
|
|
|
|||
|
|
No shipped policy reads them. But the first thing a competitive policy
|
|||
|
|
must do is **value a Problem**, and value is the field the projection
|
|||
|
|
hides — so the trap becomes live on the first line of the work this
|
|||
|
|
workplan is about. [ADR-0023](../decisions/ADR-0023-a-policy-is-bound-by-what-its-seat-can-see.md).
|
|||
|
|
|
|||
|
|
## Task: a policy is bound by what its seat can see
|
|||
|
|
|
|||
|
|
```task
|
|||
|
|
id: CB-WP-0049-T01
|
|||
|
|
status: done
|
|||
|
|
priority: high
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Per ADR-0023: established **behaviourally**, not by changing the trait —
|
|||
|
|
vary only what the seat cannot see and the choice must not move. That
|
|||
|
|
binds every policy, including ones written later and outside this crate,
|
|||
|
|
without their cooperation.
|
|||
|
|
|
|||
|
|
**Done 2026-08-08.** `bot::blindness` provides `hidden_rearrangements`
|
|||
|
|
and `is_blind`; `GreedyPolicy` and `RandomPolicy` are proven blind across
|
|||
|
|
three seeds × three seat bands × every seat.
|
|||
|
|
|
|||
|
|
### Two defects in the control itself, both caught by running it
|
|||
|
|
|
|||
|
|
**1. The first rearrangement was invariant under what a violator reads.**
|
|||
|
|
It rotated every hidden value 2↔3 *together*, which leaves `max`
|
|||
|
|
unchanged and (at equal counts) `sum` too — so the deliberate peeker,
|
|||
|
|
which ranks by the largest hidden value, **was not caught**. A control
|
|||
|
|
whose variation is invariant under the statistic a violator reads is not
|
|||
|
|
a control. The family now forces all-2, all-3 and two alternating splits,
|
|||
|
|
so any symmetric function of the hidden values moves across it.
|
|||
|
|
|
|||
|
|
**2. It accused `random` of peeking.** `is_blind` reused one policy
|
|||
|
|
instance, so it compared a first call against a fourth — and
|
|||
|
|
`RandomPolicy` advances its own stream. It takes a **constructor** now,
|
|||
|
|
so every variant is judged from an identical starting policy and the only
|
|||
|
|
difference between two runs is the state.
|
|||
|
|
|
|||
|
|
Both are the same mistake in different clothes: **a difference in the
|
|||
|
|
output was read as evidence about the hidden state**, when the first was
|
|||
|
|
insensitive to it and the second was caused by something else entirely.
|
|||
|
|
That is the wrong-subject family, found twice inside one control written
|
|||
|
|
to detect wrong subjects.
|
|||
|
|
|
|||
|
|
| mutation | what went red |
|
|||
|
|
|---|---|
|
|||
|
|
| rearrangements leave the multiset invariant | *"a policy reading face-down values must be caught"* |
|
|||
|
|
| rearrangements vary nothing at all | *"no rearrangement changed a hidden card"* + the peeker escapes |
|
|||
|
|
| one policy instance reused across variants | *"random chose Command(5) and then Command(2)"* |
|
|||
|
|
|
|||
|
|
**The control is proven against a deliberate violator** before being
|
|||
|
|
trusted about compliant policies (ADR-0023 D3) — a peek control nothing
|
|||
|
|
can fail is decoration.
|
|||
|
|
|
|||
|
|
## Task: an objective, taken from the game's own scoring
|
|||
|
|
|
|||
|
|
```task
|
|||
|
|
id: CB-WP-0049-T02
|
|||
|
|
status: todo
|
|||
|
|
priority: high
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
What a seat is trying to maximise, per `ScoringMode`:
|
|||
|
|
|
|||
|
|
| mode | objective |
|
|||
|
|
|---|---|
|
|||
|
|
| SHARED GROUND | the group total |
|
|||
|
|
| COMMON PROBLEM | own claimed value − own Blame |
|
|||
|
|
| BONDED COALITIONS | own coalition's summed personal score |
|
|||
|
|
|
|||
|
|
**Derived from `GroundState::score`, never reimplemented.** All three are
|
|||
|
|
already computed there for the outcome; a second copy in the bot would be
|
|||
|
|
a second definition of winning, and the two would disagree the first time
|
|||
|
|
ground-game rules on F28.
|
|||
|
|
|
|||
|
|
**Controls:**
|
|||
|
|
- **the objective agrees with `Outcome` on a finished game**, for every
|
|||
|
|
mode — asserted against the real scoring rather than a fixture;
|
|||
|
|
- **a seat's objective is not the group's** under the competitive modes,
|
|||
|
|
shown on a concrete board rather than argued;
|
|||
|
|
- **the objective is blind** (T01's control applies to anything that
|
|||
|
|
reaches a policy).
|
|||
|
|
|
|||
|
|
## Task: a seat that plays it, and the F27 re-measurement
|
|||
|
|
|
|||
|
|
```task
|
|||
|
|
id: CB-WP-0049-T03
|
|||
|
|
status: todo
|
|||
|
|
priority: high
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
A policy that maximises its own objective, delegating to
|
|||
|
|
`GreedyPolicy::rank` wherever the objective is silent — **delegating, not
|
|||
|
|
copying**, which CB-WP-0039 learned the hard way when a re-typed "greedy
|
|||
|
|
with one preference changed" differed in five places and burned the
|
|||
|
|
Freedom token in round one of every game.
|
|||
|
|
|
|||
|
|
Then re-run `scenario-panel`.
|
|||
|
|
|
|||
|
|
**The result is not predicted here.** Either the modes now diverge — F27
|
|||
|
|
resolves, and the competitive modes are real — or they still do not, and
|
|||
|
|
that is a much more interesting finding about GROUND: that the game's
|
|||
|
|
scoring modes do not reach its decisions. Both outcomes are publishable
|
|||
|
|
and the workplan must not be written as though one is expected.
|
|||
|
|
|
|||
|
|
**Controls:**
|
|||
|
|
- **the panel's group-success column moves, or it is reported as not
|
|||
|
|
moving** — a null result stated as a null result;
|
|||
|
|
- **the competitive policy is blind** (T01);
|
|||
|
|
- **greedy's numbers are unchanged**, so the new policy is an addition
|
|||
|
|
and not a silent edit of the published baseline.
|
|||
|
|
|
|||
|
|
## Task: module-aware evaluation
|
|||
|
|
|
|||
|
|
```task
|
|||
|
|
id: CB-WP-0049-T04
|
|||
|
|
status: todo
|
|||
|
|
priority: normal
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
A seat under `problem_stress.scoped` should prefer clearing a Problem
|
|||
|
|
whose scope includes it; under `attack_relief.self_soothe_ge4`, ATTACK
|
|||
|
|
becomes a Stress tool at the gate. Both are **visible** to the seat
|
|||
|
|
(`ProblemMarker` is in the projection), so neither needs hidden state.
|
|||
|
|
|
|||
|
|
**The composition risk is the same one schema 2 exists to retire.** One
|
|||
|
|
policy per module is `H2AwarePolicy` — the blob at a new layer. Evaluation
|
|||
|
|
should compose per aspect the way a configuration does, and if that is not
|
|||
|
|
achievable cheaply, the honest move is one policy that reads the resolved
|
|||
|
|
`Rules` rather than several that each assume one.
|
|||
|
|
|
|||
|
|
**Controls:**
|
|||
|
|
- **a module-aware seat differs from greedy ONLY where the module is
|
|||
|
|
live** — under the baseline the two must be identical, which is the
|
|||
|
|
control that separates "attends to the module" from "plays differently";
|
|||
|
|
- **blind** (T01).
|
|||
|
|
|
|||
|
|
## Not done here
|
|||
|
|
|
|||
|
|
- **The trait still takes `&GroundState`.** ADR-0023 D2 chose the
|
|||
|
|
behavioural control first, on the grounds that it binds policies the
|
|||
|
|
type system cannot reach. Narrowing `choose` to `&GroundView` remains
|
|||
|
|
worth doing and is not refused — it is simply not what establishes the
|
|||
|
|
property.
|
|||
|
|
- **Opponent modelling is out of scope.** These policies play their own
|
|||
|
|
objective; none of them models another seat playing theirs. A
|
|||
|
|
competitive mode where nobody anticipates a rival is still a weak test
|
|||
|
|
of that mode, and saying so is part of T03's result.
|
|||
|
|
- **No lookahead.** `search.rs` is retrospective and **may not be reused
|
|||
|
|
here**: ADR-0013 D1 permits it to see everything precisely because
|
|||
|
|
after the game there is one world. A policy plays from inside an
|
|||
|
|
information set, so the same permission would be strategy fusion. Same
|
|||
|
|
kernel, two searches, opposite permissions, and the discriminator is
|
|||
|
|
when the question is asked.
|