Declare CB-WP-0041 (extensive-form foundation) and CB-WP-0042 (H2)

0041 answers what is true of the engine as a game-theoretic object before
anything is built on it. CB-RES-0009 found the EFG is the interchange
format between describing a game and analysing it, and that we already
have most of one — the journal is the history, Outcome the payoff, and
project(Viewer::Player(seat)) the information partition. Three gaps
remain, and perfect recall is first because CFR and exploitability both
assume it and nobody has checked ours. The workplan deliberately builds no
port: creating a capability port is a tier-L trigger and this is M, so
T04 decides whether to build one and declares it separately. T01's control
includes that the answer may be NO, which would make Track B's adoption
unsound as it stands.

0042 implements H2, which is ground-game's direct answer to our H1
reading. Unclaimed Problems now tick only the seats in scope — global,
personal (the owner), or bond (the owner's Bond network over Bond edges
only, degree 0 falling back to personal) — assigned by hidden priority, so
2p never has the bond card in play. It explicitly does not stack with H1.

H2 is bigger than H1 was: it needs variant-scoped edition data (H2
overrides Problems.csv with a stress_scope column, and ours is an
include_str! constant), per-Problem ownership which is new state reaching
the hash and every recording, and Bond-network reachability. The controls
name the likely defects in advance: traversing Rivalry edges, applying
stacking once, forgetting the degree-0 fallback, and ownership silently
becoming a permission to SOLVE.

Chaos window 4 opens: d8 = 5 and d8 = 4, no overrides. Window 3's verdict
is still owed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
tegwick 2026-08-08 15:11:09 +02:00
parent 55a475b1a9
commit d4d25b903e
2 changed files with 299 additions and 0 deletions

View file

@ -0,0 +1,138 @@
---
id: CB-WP-0041
kind: product
title: "The extensive-form foundation"
status: active
---
# Purpose
```
structural tier M (states what the kernel claims about itself as a
game-theoretic object, and prepares — but does not
build — a capability port)
chaos d8 = 5 → no override
declared tier M
```
**Declaration 1 of chaos window 4.** Window 3 closed at 12 declarations
with one override that changed nothing, and **its verdict is still owed**
([`ChaosRollHistory.md`](../specs/ChaosRollHistory.md)).
## Why
[CB-RES-0009](../research/CB-RES-0009-extensive-form-is-the-lingua-franca.md)
found that the **extensive-form game is the interchange format** between
describing a game and analysing it — Ludii's universality result grounds
its language in EFGs, and OpenSpiel's CFR, best-response and
exploitability consume them.
**And that we already have most of one**, under other names:
| EFG component | ours |
|---|---|
| histories | the journal — `Applied { actor, command, events }` in order |
| actions | `bot::legal_commands(state, seat)` |
| **information sets** | **`state.project(Viewer::Player(seat))`** |
| payoffs | `Outcome` / `score()` |
| terminal | `outcome.is_some()` |
**This workplan does not build the port.** It answers what is true of the
engine today, because a port built on an unchecked assumption is worse
than no port — and one of the three gaps invalidates every equilibrium
concept if it turns out badly.
## Task: does the engine satisfy perfect recall?
```task
id: CB-WP-0041-T01
status: todo
priority: high
```
**Perfect recall** — every player remembers their own past actions and
observations — is the assumption **CFR and exploitability both rest on**.
Nobody has checked whether ours holds.
Formally, for a seat `i` and any two histories `h`, `h'` in the same
information set of `i`: the *sequence of `i`'s own actions and
information sets* along `h` and `h'` must be identical.
**Controls:**
- **derived from the journal and the projection**, not asserted — the
check must be able to say NO;
- **a positive control**: a deliberately forgetful projection must fail
it, or the check proves nothing;
- **stated per seat count**, since the partition depends on how many
seats there are to hide from;
- **the answer may be that we do NOT have perfect recall**, and that is
reported as plainly as the other outcome. It would be a real finding
and would make Track B's adoption unsound as it stands.
## Task: say precisely what our chance is
```task
id: CB-WP-0041-T02
status: todo
priority: high
```
`setup` shuffles, deals and draws Lead from a seeded RNG, and the
mid-game reshuffle derives from seed and round. **So a clay-borg game is
one chance realisation, not a game with chance nodes**, and the panels
approximate the distribution by sampling seeds.
**Controls:**
- **state it, do not fix it.** Monte Carlo over seeds is legitimate and
is what we do; the defect would be calling it an EFG;
- **name every point where chance enters**, from the code, not from
memory;
- **say what an explicit chance player would cost** — that is the input to
T04's decision, and guessing it is how a port gets built on a hope.
## Task: state the simultaneity encoding, and check it
```task
id: CB-WP-0041-T03
status: todo
priority: medium
```
Commit/reveal **is** the textbook EFG encoding of simultaneous moves:
sequence them, and hide the earlier move in an information set.
**Control:** it is not enough to say so. **The projection must actually
hide another seat's selection before Reveal**, and a test must fail if it
stops doing that. That property is load-bearing for every claim in §1 of
the research note and is currently only implied by
`SelectionView::Hidden`.
## Task: decide whether to build the port at all
```task
id: CB-WP-0041-T04
status: todo
priority: medium
```
`decisions/ADR-*.md`, written **after** T01T03 and not before.
**It must be able to conclude "no".** Options include exporting an EFG,
adopting OpenSpiel's API directly for analysis only, or deciding the gaps
are too expensive and Track B borrows vocabulary rather than machinery.
**Controls:**
- **the decision cites T01's answer**, because if perfect recall fails,
most of the option space closes;
- **it states what it would cost to be wrong**;
- **a port is declared separately, at its own tier.** Creating a capability
port is a tier-L trigger and this workplan is M — it may not smuggle one
in.
## Not in this workplan
- **No EFG export, no OpenSpiel integration, no equilibrium computation.**
- **No answer to "is exploitability meaningful for a co-operative game
with a shared threshold"** — that is the other open question from
CB-RES-0009 §6, it is a question about game theory rather than about our
engine, and it wants its own pass.

View file

@ -0,0 +1,161 @@
---
id: CB-WP-0042
kind: product
title: "H2 — scoped problem stress"
status: ready
---
# Purpose
```
structural tier M (touches a canonical interface: edition loading
becomes variant-parameterised, and the kernel gains
per-Problem state)
chaos d8 = 4 → no override
declared tier M
```
**Declaration 2 of chaos window 4.**
## Why, and it is a direct answer to ours
We reported that H1-A behaves as a **solve-rate tax**: a flat +1 to every
seat while any Problem is unclaimed, so the pressure scales with the
number of Problems while the intended effect does not.
`ground-game` responded with **H2 — scoped problem stress**. Unclaimed
Problems now tick **only the seats in their scope**:
| scope | who takes the +1 |
|---|---|
| `global` | all seats |
| `personal` | the Problem's **owner** |
| `bond` | the owner's **Bond network** — owner plus every seat reachable over Bond edges only, never Rivalry; degree 0 falls back to personal |
Assigned by hidden priority: **0 (Surface) global, 1 personal, 2 personal,
3 bond, 4 personal** — which is a seat-band dial without a difficulty
card, since 2p never has the bond card in play.
**It does not stack with H1.** `replaces_experiments: [h1-problem-stress]`,
base is `ground-darvo-r0`, and there is **no ATTACK self-soothe**.
## Why this is bigger than H1 was
H1 was two arithmetic deltas on existing state. H2 needs three things we
do not have:
1. **Variant-scoped edition data.** H2 overrides `Problems.csv` with a
new `stress_scope` column. `edition.rs` `include_str!`s r0's copy at
compile time — **the data is currently a constant, not a parameter.**
2. **Per-Problem ownership**, assigned at setup. That is **new state**,
so it reaches the state hash and every recording.
3. **Bond-network reachability** — a graph traversal over Bond edges
only, with a degree-0 fallback.
## Task: variant-scoped edition data
```task
id: CB-WP-0042-T01
status: todo
priority: high
```
**Controls:**
- **the baseline's data is bit-for-bit what it was**, and its state hashes
do not move — the control that protected CB-WP-0038 and the one that
invalidates every prior measurement if it fails;
- the H2 package is **vendored with digests** like every other borrowed
file, and `edition-check` covers it — sibling discovery walks the disk
now, so an undigested file fails (CB-REV-0003 #8);
- **`stress_scope`'s assignment is read from the file, not hardcoded.**
The delta states the priority→scope mapping; **it is the edition's to
state and ours to read**, and F25 exists because we hardcoded numbers
the edition already carried.
## Task: ownership at setup
```task
id: CB-WP-0042-T02
status: todo
priority: high
```
`H2-OWN`: ascending hidden priority among in-play non-global Problems,
starting at Lead, stepping clockwise. Global Problems have no owner.
**Controls:**
- **deterministic**, and asserted so — the delta says "deterministic for
sims" and a seeded-but-unstated order would be untestable;
- **the baseline carries no owners and its hash does not move.** New
state that is `None` under baseline must serialise as it did before, or
every recorded scenario breaks;
- **exactly one owner per eligible Problem**, checked at every seat count,
because the assignment walks two sequences at once and off-by-one is
the obvious failure.
## Task: scoped pressure
```task
id: CB-WP-0042-T03
status: todo
priority: high
```
`H2-A`, at Round End, before the clamp and the DARVO arm check — the same
ordering H1-A needed, and the same trap.
**Controls:**
- **each scope fails on its own**, by mutation: global hitting only the
owner, personal hitting everyone, bond stopping at the owner;
- **`stacking: true` is tested** — two open bond Problems must tick the
network **twice**, and a careless implementation applies it once;
- **Rivalry edges are not traversed.** The delta says Bond only, and a
traversal that follows any relation is the single most likely defect;
- **degree 0 falls back to personal**, which is the branch a test forgets;
- ordering asserted as in CB-WP-0038: pressure before the arm check.
## Task: SOLVE is not restricted by scope
```task
id: CB-WP-0042-T04
status: todo
priority: medium
```
`H2-SOLVE`: any seat with a matching suit may claim; the owner need not be
the solver, and altruistic clearing is **intended**.
**Control:** our engine has no owner concept today, so this is already
true — which makes it a **regression test, not a feature**. T02 adds
ownership, and the risk is that ownership silently becomes a permission.
The test must fail if it does.
## Task: measure, and report what fails
```task
id: CB-WP-0042-T05
status: todo
priority: high
```
Same instrument as H1, same panel, both variants in one run.
**Controls:**
- **greedy and `reactive` both**, because H1's whole verdict turned on
which policy was asked (CB-EV-0031);
- **report against ground-game's own criteria** for H2, from their design
note — **read it, do not reuse H1's §3.2 from memory**;
- **the bond-scope hypothesis is the interesting one and must be measured
directly**: their claim is that a shared tick makes a Bond network
*jointly motivated* to solve that card. Whether a policy that does not
model other seats can express that at all is an open question, and the
honest answer may be "our panel cannot test this hypothesis";
- **`make panels` runs it**, or the figures come from an ungated binary
again (CB-REV-0002 #7).
## Not in this workplan
- **No stacking with H1.** The package forbids it explicitly.
- **No felt-play.** H2's central claim is about *motivation* — a bonded
pair caring about each other's card — and 200-game aggregates measure
dynamics, not that ([`Taxonomy.md`](../specs/Taxonomy.md) §4).