clay-borg/simulators/openspiel.md

93 lines
3.9 KiB
Markdown
Raw Normal View History

simulators/: persist the survey, and let it shrink two of the three tracks Eight profiles on a common schema, each marking what was checked against a source this session and what is background recollection. Three are marked unverified in full — Machinations, the play substrates, most of RBG — and say so rather than reading as evaluations. Written straight after three review rounds whose entire yield was claims outrunning what had been checked, so the confidence rule is the first thing in the README. The survey changed the plan, which is what a survey is for. Track C was described in Positioning as open ground. It is not: Browne published 57 criteria for game quality, and Ai Ai already computes designer-facing measures — drama, lead changes, branching factor, completion, duration — from played games. The track becomes adopt, credit and find the gap. The gap looks real: those measures presume a leader, and SHARED GROUND has none — Modes.csv gives its tiebreak as "Not applicable". Track B probably adopts rather than builds. OpenSpiel implements CFR, best-response and exploitability over games that are simultaneous-move, imperfect-information and co-operative, which is all four of GROUND's awkward properties. "Does ATTACK ever pay" is a best-response question, and we spent three review rounds refining a two-policy sweep for it. The first Track B task is now one question — is exploitability meaningful for a co-operative game with a shared threshold — not a build. The cost of not surveying earlier is therefore measurable, and is recorded rather than glossed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 11:56:27 +02:00
# OpenSpiel
Google DeepMind · C++ core with Python bindings · Apache-2.0 `unverified` ·
`checked` for scope and algorithms
## What it optimises for
**Research in reinforcement learning and search/planning in games**, with
a strong game-theoretic algorithm library.
Supports n-player single- and multi-agent, **zero-sum, cooperative and
general-sum**, one-shot and sequential, turn-taking **and
simultaneous-move**, perfect and imperfect information — plus grid worlds
and **social dilemmas**. `checked`
## How a game is defined
Games are implemented against a C++ (or Python) API. `checked` — this is
a *programming* interface, not a description language, which is the axis
where Ludii and Ai Ai differ from it.
CB-RES-0009: extensive form is the lingua franca Two questions from the maintainer — is there a game-theory mapping to Ludii's language, and is that language formal enough to derive one from. Yes, no, and the no does not matter. The mapping is proven, not to be invented: "The Ludii Game Description Language is Universal" shows the language can represent an equivalent game for any finite, non-deterministic, imperfect-information game, extending earlier work limited to finite deterministic fully-observable extensive-form games. EFG is also OpenSpiel's object, so the same formalism connects description to analysis: Ludii -> EFG <- OpenSpiel. Ludii's syntax is formal and unusually so — a class grammar derived automatically from its source. Its semantics are its Java: a ludeme means what its class does, and Ludii effectively makes Java the game description language. So there is no independent calculus to extract. The formality lives in the universality RESULT, not in a definition of meaning. GDL has the semantics and pays for it in speed — six times on Gomoku, twenty on Amazons and Hex, over two hundred on Chess. Conclusion: do not derive a language from Ludii; target the EFG directly. And we are closer than the tracks assumed. The journal is the history, Outcome is the payoff, legal_commands gives the actions — and project(Viewer::Player(seat)) IS the information partition, built so a player is not shown another's hand and unremarked as exactly the machinery imperfect information needs. Three gaps: chance is folded into a seed so a game is one realisation rather than a game with chance nodes; perfect recall is unasserted, which CFR and exploitability both assume; and commit/reveal is the standard EFG encoding of simultaneity but is never stated as such. Perfect recall is checkable from the journal today and is now Track B's first task — if it fails, every equilibrium concept we might quote is unsound here. Also re-vendored the catalog twice: ground-game added H2 — scoped problem stress, applying End Stress by personal/bond/global scope instead of flat to everyone, which is a direct response to our reading that H1's tax scales with the Problems while its intended effect does not. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 14:57:25 +02:00
**But the API is an extensive-form game interface**, and that is the
connection: Ludii's universality proof grounds its language in EFGs, and
OpenSpiel's algorithms consume EFGs. **EFG is the interchange format**
([`../research/CB-RES-0009-extensive-form-is-the-lingua-franca.md`](../research/CB-RES-0009-extensive-form-is-the-lingua-franca.md)).
simulators/: persist the survey, and let it shrink two of the three tracks Eight profiles on a common schema, each marking what was checked against a source this session and what is background recollection. Three are marked unverified in full — Machinations, the play substrates, most of RBG — and say so rather than reading as evaluations. Written straight after three review rounds whose entire yield was claims outrunning what had been checked, so the confidence rule is the first thing in the README. The survey changed the plan, which is what a survey is for. Track C was described in Positioning as open ground. It is not: Browne published 57 criteria for game quality, and Ai Ai already computes designer-facing measures — drama, lead changes, branching factor, completion, duration — from played games. The track becomes adopt, credit and find the gap. The gap looks real: those measures presume a leader, and SHARED GROUND has none — Modes.csv gives its tiebreak as "Not applicable". Track B probably adopts rather than builds. OpenSpiel implements CFR, best-response and exploitability over games that are simultaneous-move, imperfect-information and co-operative, which is all four of GROUND's awkward properties. "Does ATTACK ever pay" is a best-response question, and we spent three review rounds refining a two-policy sweep for it. The first Track B task is now one question — is exploitability meaningful for a co-operative game with a shared threshold — not a build. The cost of not surveying earlier is therefore measurable, and is recorded rather than glossed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 11:56:27 +02:00
## What it can analyse
The reason this profile matters for Track B. Implemented algorithms
include **CFR** and its Monte-Carlo variants, **Deep CFR**, **exploitability**
and best-response computation, fictitious play (XFP, NFSP), **PSRO**,
**α-Rank**, replicator/evolutionary dynamics, minimax, MCTS, and
value iteration. `checked`
**Exploitability** is the key one: how much an opponent gains by
deviating to a best response — a direct, computable measure of *how far
from equilibrium a strategy is*. `checked`
## What it does not do
- Nothing design-facing. It answers "how strong is this strategy" and
"how far from equilibrium", not "is this rule doing its job".
- No game description language, so no cheap path for a designer.
- No evidence discipline: outputs are research data.
## Relevance to clay-borg
**Track B, and it is very likely the answer rather than a competitor.**
Our current instrument for "does ATTACK ever pay" is two hand-written
policies and a rank parameter. **Exploitability and best-response are the
principled versions of that question**, and they are implemented,
tested and published here.
GROUND is simultaneous-move, imperfect-information, cooperative with a
defection mechanic — **all four are inside OpenSpiel's stated scope**.
## What to steal
- **Exploitability as the replacement for "we tried three policies".**
F17's question — does ATTACK earn its place — is a best-response
question wearing a sweep's clothes.
- The distinction between **cooperative, general-sum and zero-sum** as a
first-class property of a game, which our kernel does not represent.
## What to avoid
**Quoting a solution concept without its assumptions.** An equilibrium
computed over a policy class we chose is a statement about that class.
This is the wrong-subject error in mathematical dress, and it is the
specific risk `Positioning.md` §4 flags for Track B.
## Open questions
- Cost of expressing a GROUND-like game against its API versus our kernel.
CB-RES-0009: extensive form is the lingua franca Two questions from the maintainer — is there a game-theory mapping to Ludii's language, and is that language formal enough to derive one from. Yes, no, and the no does not matter. The mapping is proven, not to be invented: "The Ludii Game Description Language is Universal" shows the language can represent an equivalent game for any finite, non-deterministic, imperfect-information game, extending earlier work limited to finite deterministic fully-observable extensive-form games. EFG is also OpenSpiel's object, so the same formalism connects description to analysis: Ludii -> EFG <- OpenSpiel. Ludii's syntax is formal and unusually so — a class grammar derived automatically from its source. Its semantics are its Java: a ludeme means what its class does, and Ludii effectively makes Java the game description language. So there is no independent calculus to extract. The formality lives in the universality RESULT, not in a definition of meaning. GDL has the semantics and pays for it in speed — six times on Gomoku, twenty on Amazons and Hex, over two hundred on Chess. Conclusion: do not derive a language from Ludii; target the EFG directly. And we are closer than the tracks assumed. The journal is the history, Outcome is the payoff, legal_commands gives the actions — and project(Viewer::Player(seat)) IS the information partition, built so a player is not shown another's hand and unremarked as exactly the machinery imperfect information needs. Three gaps: chance is folded into a seed so a game is one realisation rather than a game with chance nodes; perfect recall is unasserted, which CFR and exploitability both assume; and commit/reveal is the standard EFG encoding of simultaneity but is never stated as such. Perfect recall is checkable from the journal today and is now Track B's first task — if it fails, every equilibrium concept we might quote is unsound here. Also re-vendored the catalog twice: ground-game added H2 — scoped problem stress, applying End Stress by personal/bond/global scope instead of flat to everyone, which is a direct response to our reading that H1's tax scales with the Problems while its intended effect does not. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 14:57:25 +02:00
- **Whether our engine satisfies perfect recall.** CFR and exploitability
assume it, our per-seat projection has never been checked for it, and
it is checkable from the journal. **This is the first Track B task**
if it fails, every equilibrium concept is unsound here.
- **Our chance is folded into a seed, not an explicit chance player.** A
clay-borg game is one chance realisation; the panels sample seeds to
approximate the distribution. CFR needs chance nodes.
simulators/: persist the survey, and let it shrink two of the three tracks Eight profiles on a common schema, each marking what was checked against a source this session and what is background recollection. Three are marked unverified in full — Machinations, the play substrates, most of RBG — and say so rather than reading as evaluations. Written straight after three review rounds whose entire yield was claims outrunning what had been checked, so the confidence rule is the first thing in the README. The survey changed the plan, which is what a survey is for. Track C was described in Positioning as open ground. It is not: Browne published 57 criteria for game quality, and Ai Ai already computes designer-facing measures — drama, lead changes, branching factor, completion, duration — from played games. The track becomes adopt, credit and find the gap. The gap looks real: those measures presume a leader, and SHARED GROUND has none — Modes.csv gives its tiebreak as "Not applicable". Track B probably adopts rather than builds. OpenSpiel implements CFR, best-response and exploitability over games that are simultaneous-move, imperfect-information and co-operative, which is all four of GROUND's awkward properties. "Does ATTACK ever pay" is a best-response question, and we spent three review rounds refining a two-policy sweep for it. The first Track B task is now one question — is exploitability meaningful for a co-operative game with a shared threshold — not a build. The cost of not surveying earlier is therefore measurable, and is recorded rather than glossed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 11:56:27 +02:00
- Whether exploitability is meaningful for a co-operative game with a
shared threshold — `open`, and it is the question that decides whether
Track B adopts this or only borrows its vocabulary.
- Licence.
## Sources
- OpenSpiel repository (https://github.com/google-deepmind/open_spiel)
- *OpenSpiel: A Framework for Reinforcement Learning in Games*
(https://www.researchgate.net/publication/335419770)