--- id: CB-WP-0041 kind: product title: "The extensive-form foundation" status: done state_hub_workstream_id: "dacfa1fe-81c3-4c97-b81e-ec66d95a611c" --- # Purpose ``` structural tier M (states what the kernel claims about itself as a game-theoretic object, and prepares — but does not build — a capability port) chaos d8 = 5 → no override declared tier M ``` **Declaration 1 of chaos window 4.** Window 3 closed at 12 declarations with one override that changed nothing, and **its verdict is still owed** ([`ChaosRollHistory.md`](../specs/ChaosRollHistory.md)). ## Why [CB-RES-0009](../research/CB-RES-0009-extensive-form-is-the-lingua-franca.md) found that the **extensive-form game is the interchange format** between describing a game and analysing it — Ludii's universality result grounds its language in EFGs, and OpenSpiel's CFR, best-response and exploitability consume them. **And that we already have most of one**, under other names: | EFG component | ours | |---|---| | histories | the journal — `Applied { actor, command, events }` in order | | actions | `bot::legal_commands(state, seat)` | | **information sets** | **`state.project(Viewer::Player(seat))`** | | payoffs | `Outcome` / `score()` | | terminal | `outcome.is_some()` | **This workplan does not build the port.** It answers what is true of the engine today, because a port built on an unchecked assumption is worse than no port — and one of the three gaps invalidates every equilibrium concept if it turns out badly. ## Task: does the engine satisfy perfect recall? ```task id: CB-WP-0041-T01 status: done priority: high state_hub_task_id: "a9438152-1556-4cde-9eae-af75066a485a" ``` **Perfect recall** — every player remembers their own past actions and observations — is the assumption **CFR and exploitability both rest on**. Nobody has checked whether ours holds. Formally, for a seat `i` and any two histories `h`, `h'` in the same information set of `i`: the *sequence of `i`'s own actions and information sets* along `h` and `h'` must be identical. **Controls:** - **derived from the journal and the projection**, not asserted — the check must be able to say NO; - **a positive control**: a deliberately forgetful projection must fail it, or the check proves nothing; - **stated per seat count**, since the partition depends on how many seats there are to hide from; - **the answer may be that we do NOT have perfect recall**, and that is reported as plainly as the other outcome. It would be a real finding and would make Track B's adoption unsound as it stands. **Done 2026-08-08. The answer is "it depends what you call an information set", and the distinction is the result.** 44,938 decision points, random play, 2/3/4/6 seats. | reading | information set is… | violations | |---|---|---| | **A** | the seat's **current projection** — what `project(Viewer::Player(seat))` returns and the page renders | **22** | | **B** | the seat's **observation history** — every view seen and action taken, in order | **0** | **Reading A fails, and the witness is concrete**: two histories reach a byte-identical view — round 3, Select, same hand, same claimed Problem — where the seat had played `SOLVE, GROUND—OU(protect)` in one and `SUPPORT, SOLVE` in the other. **The view does not tell the seat what it did.** **The mechanism is that our state is a snapshot, not a history.** Selections clear each round and effects coincide, so a player cannot reconstruct their own past from the present. In a real game the player's memory supplies it; in the state, nothing does. **This is precisely OpenSpiel's `ObservationString` vs `InformationStateString` split**, arrived at here by measurement rather than by reading it off. `project()` is an *observation*. **So Track B is not closed — it is constrained**, and usefully: > **An extensive-form game built from this engine must key information > sets on observation histories, never on `project()`.** **Both directions are asserted.** Reading B empty, *and* Reading A non-empty — because if the sample stops finding Reading A violations the conclusion is unsupported and must be re-derived, not quietly kept. **What this cannot say.** It samples; it can falsify perfect recall and cannot establish it. Reading B's zero means *no counterexample was drawn*, which is weaker than "the property holds" and is printed as such. ## Task: say precisely what our chance is ```task id: CB-WP-0041-T02 status: done priority: high state_hub_task_id: "e223bf7d-6e4e-4377-8eac-a63e83a60d35" ``` `setup` shuffles, deals and draws Lead from a seeded RNG, and the mid-game reshuffle derives from seed and round. **So a clay-borg game is one chance realisation, not a game with chance nodes**, and the panels approximate the distribution by sampling seeds. **Controls:** - **state it, do not fix it.** Monte Carlo over seeds is legitimate and is what we do; the defect would be calling it an EFG; - **name every point where chance enters**, from the code, not from memory; - **say what an explicit chance player would cost** — that is the input to T04's decision, and guessing it is how a port gets built on a hope. **Done 2026-08-08.** Three chance points, all reading the same root seed: | where | what | |---|---| | `setup` | shuffles the Solution deck (`ChaChaRng::from_seed(seed)`), hands dealt off the top | | `setup` | draws the **Lead** (`rng.draw(seats)`) from the same stream | | `draw_solution` | reshuffles the discard when the deck empties, with `ChaChaRng::from_seed(seed ^ round)` | **The Problems deal is not chance at all** — `edition::deal` is a pure function of the vendored CSV, so every game gets the same board. **So in extensive-form terms the tree has a single chance node at the root.** `a_game_is_determined_by_its_seed` pins it. **And that test was wrong first.** It compared *state hashes*, and `GroundState` carries `seed` as a field — so "different seeds differ" was true by construction. Mutating the shuffle away left it green. It now compares the **dealt configuration** (hands, Lead, deck order), and the same mutation fails it. **A wrong-subject error inside T02's own control**, caught by the mutation rather than by reading. **The reshuffle is correlated with the root seed, and a real table's is not.** The permutation is a pure function of `(seed, round)` — deliberate, so replay never re-derives it (GameKernel K5), and the code says so. `the_reshuffle_permutation_is_a_function_of_seed_and_round` pins that it moves with **both** inputs. **The consequence is a modelling one, not a defect**: at a table the reshuffle is an independent random event. An EFG built from this engine inherits the correlation and models a *restriction* of the game as played. **The cost of an explicit chance player, stated for T04:** - the root node is **not enumerable** — 24 Solution cards give 24! orderings — so any EFG over this must **sample** chance, which is what external-sampling MCCFR does and what our seed sweeps already do by hand; - making the reshuffle a genuine chance node **breaks the K5 purity that makes replay deterministic**. That is the real cost and it is a tension, not a line of code: our recordings replay *because* chance is a function of state. ## Task: state the simultaneity encoding, and check it ```task id: CB-WP-0041-T03 status: done priority: medium state_hub_task_id: "c95b32a1-5006-4acd-abc9-537543119344" ``` Commit/reveal **is** the textbook EFG encoding of simultaneous moves: sequence them, and hide the earlier move in an information set. **Control:** it is not enough to say so. **The projection must actually hide another seat's selection before Reveal**, and a test must fail if it stops doing that. That property is load-bearing for every claim in §1 of the research note and is currently only implied by `SelectionView::Hidden`. **Done 2026-08-08.** `a_pending_selection_is_hidden_until_reveal` checks **both halves** at four seats: before Reveal each seat sees its own selection and no other, and after Reveal the information sets **merge** — because an encoding that hides forever is not commit/reveal either. Mutation-proven: make selections public and it fails naming the seats. ## Task: decide whether to build the port at all ```task id: CB-WP-0041-T04 status: done priority: medium state_hub_task_id: "3010fb0d-84cc-4fa8-8c23-726ee946cd79" ``` `decisions/ADR-*.md`, written **after** T01–T03 and not before. **It must be able to conclude "no".** Options include exporting an EFG, adopting OpenSpiel's API directly for analysis only, or deciding the gaps are too expensive and Track B borrows vocabulary rather than machinery. **Controls:** - **the decision cites T01's answer**, because if perfect recall fails, most of the option space closes; - **it states what it would cost to be wrong**; - **a port is declared separately, at its own tier.** Creating a capability port is a tier-L trigger and this workplan is M — it may not smuggle one in. **Done 2026-08-08.** [ADR-0020](../decisions/ADR-0020-we-do-not-build-the-port.md) — **and the answer is no.** **The blocker is T02, not T01**, which inverts what the workplan expected. Perfect recall looked like the risk and turned out to be a *constraint with a known answer*: key on observation histories. **Making chance explicit is the expensive one** — the reshuffle would become a real chance node, and that breaks the K5 purity every recording, replay bundle and trial-note hash in this repo depends on. > A port would trade the property this project is built on for one it has > never needed. **Track B's first move is a question, not a build**: take "is exploitability meaningful for a co-operative game with a shared threshold" to OpenSpiel on a toy model, where answering it costs nothing. **D4 states what being wrong looks like** — OpenSpiel settling, on a toy, a question three rounds of policy sweeps could not — and makes watching for it the next action rather than a hope. ## Not in this workplan - **No EFG export, no OpenSpiel integration, no equilibrium computation.** - **No answer to "is exploitability meaningful for a co-operative game with a shared threshold"** — that is the other open question from CB-RES-0009 §6, it is a question about game theory rather than about our engine, and it wants its own pass.