Some checks failed
ci / check (push) Failing after 4s
Five remarks from the maintainer's test games, checked against the code before being written down — two had already reached ground-game on wrong premises, so a claim now names the line that makes it true. Three of the five turned out to be data the projection already carries, drawn as text: solution_deck_len, solution_discard, and OutcomeView's personal/mastery/winners. One control (`close — I have read this`) is labelled as a reading but shuts the server down, and leaves a live-looking page pointing at a dead port. One number does not exist at all: table.rs loops run_game and keeps only the last summary. CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the other seats' plays as objects on the table, the ending control saying what it does, and a tally that survives "play again". CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how hard is this" are the same search asked twice. Tier L because the information boundary is the whole design problem: a solver reading GroundState sees the deck the rules hide, and would tell the maintainer he could have won by playing a card he had no way to know was there. Also unblocks GROUND-WP-0005, active with both tasks waiting on a measured difficulty baseline. Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023 shipped CB-EV-0020 first. loop-lint: no findings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
272 lines
11 KiB
Markdown
272 lines
11 KiB
Markdown
---
|
|
id: CB-WP-0025
|
|
kind: product
|
|
title: "Could we have won: a path out of a lost game, and how hard the game actually is"
|
|
status: ready
|
|
---
|
|
|
|
# Purpose
|
|
|
|
```
|
|
structural tier L (creates a new capability — a search over game state,
|
|
and a measurement the engine does not currently take;
|
|
both produce numbers ground-game will tune against)
|
|
chaos d8 = 6 → no override
|
|
declared tier L
|
|
```
|
|
|
|
Declaration 8 of chaos window 2. Tier L: separate survey, **adversarial
|
|
review**, ADR, then spec, then code.
|
|
|
|
## Two remarks, and why they are one pass
|
|
|
|
> *"I had a game where we lost and in this case I would have liked to know
|
|
> if and how we could have won… the best path is not computable I guess so
|
|
> a path to win is fine."*
|
|
|
|
> *"Do we have difficulty estimations? If so we should show them. It will
|
|
> help tuning the game. I felt it was too easy but then we lost, so who
|
|
> knows."*
|
|
|
|
They are the same machine asked two questions. *Was this game winnable?*
|
|
is a search from a recorded state. *How hard is this game?* is that search
|
|
run over many deals and counted. Building the second without the first
|
|
gives a win-rate with no witness; building the first without the second
|
|
gives one anecdote per game.
|
|
|
|
**"I felt it was too easy but then we lost, so who knows" is the finding.**
|
|
The maintainer cannot calibrate the game from play, and that is precisely
|
|
the gap `ground-game` is currently blocked in: **GROUND-WP-0005**
|
|
*(Difficulty tiers (threshold ratio) and optional Pressure dial)* is active
|
|
with **both its tasks in `wait`**. Tiers cannot be set without a measured
|
|
baseline, and clay-borg is the thing that can measure. This pass is what
|
|
unblocks it — which is the design-instrument aspect CB-WP-0022 is in the
|
|
middle of stating, arriving with a concrete demand.
|
|
|
|
## What already exists, so the survey does not re-find it
|
|
|
|
- **The state is replayable.** `cb-game-runtime` records sessions as
|
|
scenarios; `replay.rs` and `make replay-test` already re-run them.
|
|
A search does not need new persistence.
|
|
- **The move space is enumerable.** `legal_commands` exists and, since
|
|
CB-WP-0023, is narrow enough to be worth trusting — SOLVE is offered
|
|
only where it can act, so the branching factor is real rather than
|
|
inflated by inert moves.
|
|
- **Bots exist.** `games/ground/src/bot.rs` has `GreedyPolicy` and
|
|
`RandomPolicy`, wired through `bot_policy` (`table.rs:208`). A win rate
|
|
over N seeds is reachable with what is already there — the question is
|
|
whether that number *means* anything, which is the survey's problem.
|
|
- **The threshold is public.** `OutcomeView.total` / `.threshold` /
|
|
`.group_success`. Difficulty has a denominator already.
|
|
|
|
## What makes this hard, and must not be waved through
|
|
|
|
**The game is not perfect-information and the search must respect that.**
|
|
A path computed with the deck known is a path the players could never have
|
|
found. `view.rs` hides the deck, other seats' hands, and face-down
|
|
selections *by rule* (GR-S02/S04, GR-R02/R04). A retrospective solver
|
|
running on `GroundState` sees all of it. So the ADR must decide, in
|
|
words, **which of these three the tool answers**:
|
|
|
|
- *was this deal winnable by an omniscient player* — cheap, honest,
|
|
and answers a question nobody asked;
|
|
- *was it winnable from what the seats could see* — the question actually
|
|
asked, and the expensive one;
|
|
- *did a reasonable line exist* — a bounded search from the losing seat's
|
|
information, which may be the only affordable honest answer.
|
|
|
|
Getting this wrong produces a feature that tells the maintainer he could
|
|
have won by playing a card he had no way to know was there. **That is
|
|
worse than not shipping it.**
|
|
|
|
**And a difficulty number is a claim about a distribution.** One win rate
|
|
over one bot policy over N seeds is not "the difficulty"; it is that
|
|
policy's win rate. Whatever the spec adopts must name its policy, its N,
|
|
and its seed range, or `ground-game` will tune tiers against a number
|
|
whose meaning drifts the next time a bot improves.
|
|
|
|
## Task: survey
|
|
|
|
```task
|
|
id: CB-WP-0025-T01
|
|
status: todo
|
|
priority: high
|
|
```
|
|
|
|
`research/CB-RES-0008-*.md`, with `tier: L` and the chaos roll recorded
|
|
(`loop-lint` checks both).
|
|
|
|
Per §Step 1 the survey is done when it can name a **benchmark-to-beat**
|
|
per dimension — a number or a reproducible comparison, not an impression.
|
|
|
|
- **Retrospective solvers in games with hidden information.** The prior art
|
|
is real and should be named: determinized search (perfect-information
|
|
Monte Carlo) and its known failure — *strategy fusion*, where a
|
|
determinizing solver claims lines that require knowing which world it is
|
|
in. That failure is exactly the trap in §What makes this hard. Bridge
|
|
and Skat post-mortem tools are the closest analogues; poker solvers are
|
|
the well-studied case and the wrong shape.
|
|
- **"A path to win" as a product, not a proof.** The maintainer already
|
|
conceded optimality (*"the best path is not computable I guess"*). So
|
|
the target is a **witness**: one concrete line of play that reaches
|
|
`group_success`, or a defensible *no line found within bound B*. Name
|
|
what a witness must carry to be checkable.
|
|
- **Difficulty as a measured quantity in co-operative games.** Pandemic and
|
|
its relatives set difficulty by a dial with a published win rate. The
|
|
benchmark-to-beat is: can we produce a win rate whose confidence
|
|
interval is tight enough to distinguish two threshold settings?
|
|
- **Cost.** Search over an event-sourced aggregate with full `validate` on
|
|
every branch has a per-node price. Measure it on our machine, on our
|
|
scenarios — the runnable-baseline option applies here, since a search
|
|
that cannot finish while the player is still looking at the page is a
|
|
different feature.
|
|
|
|
## Task: adversarial review
|
|
|
|
```task
|
|
id: CB-WP-0025-T02
|
|
status: todo
|
|
priority: high
|
|
```
|
|
|
|
Tier L requires it. Exactly one round: challenge, then response, trail in
|
|
`history/`, unpolished. Require an attempt at:
|
|
|
|
- **that the honest version is unaffordable** — that a search respecting
|
|
the information rule is too expensive or too weak to find anything, so
|
|
the shipped tool will quietly become the omniscient one with a
|
|
reassuring label;
|
|
- **that a witness misleads more than it helps** — being shown a line that
|
|
needed a card you could not know about teaches a wrong lesson about the
|
|
game, and the tool would be better refusing to answer;
|
|
- **that the difficulty number is a bot benchmark wearing a difficulty
|
|
costume**, and `ground-game` will tune the game against our bot rather
|
|
than against play;
|
|
- **that this is CB-WP-0022's job** — the design instrument is being built
|
|
right now, and a difficulty measurement is a finding-producing tool. The
|
|
strongest counter is that the register records findings and this
|
|
*produces* them, but the reviewer should press whether that is a
|
|
distinction worth a separate capability.
|
|
|
|
## Task: decide
|
|
|
|
```task
|
|
id: CB-WP-0025-T03
|
|
status: todo
|
|
priority: high
|
|
```
|
|
|
|
`decisions/ADR-0013-*.md`. (ADR-0012 is CB-WP-0022 T03's.) At minimum:
|
|
|
|
- **which question the solver answers**, from the three in §What makes
|
|
this hard, and what it is called in the UI — the name must not overclaim;
|
|
- **the information boundary**: whether the search runs on `GroundState`
|
|
or on a `GroundView`, and if on state, what stops it using what the view
|
|
hides. Note that running on the view makes the rule structural rather
|
|
than a promise, and that this is the cheapest guarantee available;
|
|
- **the bound**: depth, node budget, or wall clock, and what *no path
|
|
found* means against it — a bounded search that says "unwinnable" is
|
|
lying, and the wording must say "none found within B";
|
|
- **whether difficulty ships as one number or a small table**, and what it
|
|
is a function of: policy, seat count, threshold, seed range;
|
|
- **where it lives** — a new crate, a mode of `cb-play`, or a tool under
|
|
`tools/`. The tier was declared L on the assumption of a new capability;
|
|
if the ADR concludes it is a mode of an existing one, say so, and the
|
|
over-declaration is a chaos-window data point worth recording.
|
|
|
|
## Task: specify
|
|
|
|
```task
|
|
id: CB-WP-0025-T04
|
|
status: todo
|
|
priority: high
|
|
```
|
|
|
|
`specs/` — extend `MetricsAndScenarios.md` or add a capability spec as the
|
|
ADR directs, with metrics, because a spec without them is prose.
|
|
|
|
Candidates, to be argued not adopted:
|
|
|
|
- **witness checkability** — every path the tool emits replays through the
|
|
existing scenario runner and ends in `group_success`. Target 100%, and it
|
|
is a hard gate, not a metric: a path that does not replay is a bug that
|
|
says the opposite of the truth;
|
|
- **search cost** — nodes and wall clock at the chosen bound, on the
|
|
recorded games we have;
|
|
- **difficulty resolution** — the smallest threshold difference the
|
|
measurement can distinguish, with its N. This is the number
|
|
`ground-game` needs, and stating it as *"we can tell 5 from 7 but not 7
|
|
from 8"* is more useful than a win rate with no error bar.
|
|
|
|
Per `ground-game`'s ruling (GROUND-WP-0004 T02), **any arithmetic finding
|
|
this produces ships a runnable reproduction and a row-level table** — never
|
|
a summed figure. A difficulty number is arithmetic, and it is exactly the
|
|
kind that has already gone wrong twice.
|
|
|
|
## Task: build the witness
|
|
|
|
```task
|
|
id: CB-WP-0025-T05
|
|
status: todo
|
|
priority: high
|
|
```
|
|
|
|
The search, the bound, and the replayable path. Wire it to the ending page
|
|
so a lost game can be asked the question — the page CB-WP-0024 T01 is
|
|
already reworking, so land that first or expect a conflict.
|
|
|
|
**A game that was won is not asked the question.** The feature exists for
|
|
a loss.
|
|
|
|
**Controls:**
|
|
- every emitted witness replays to `group_success` through the existing
|
|
runner — asserted, not spot-checked;
|
|
- a deal constructed to be unwinnable returns *none found*, and the test
|
|
says which construction makes it so;
|
|
- the information boundary is mutation-provable: relax it, and a test
|
|
naming *that* boundary goes red. If it cannot be mutated, it was a
|
|
comment rather than a rule.
|
|
|
|
## Task: measure the difficulty, and hand it to ground-game
|
|
|
|
```task
|
|
id: CB-WP-0025-T06
|
|
status: todo
|
|
priority: high
|
|
```
|
|
|
|
Run the measurement, ship it as a `make` target beside the other
|
|
instruments, and show the result in the game — the maintainer asked for it
|
|
to be visible, and a number in a file will not calibrate anything.
|
|
|
|
Then send it to `ground-game` **against GROUND-WP-0005**, which is active
|
|
with both tasks waiting on exactly this. Per CB-WP-0022 T06, it lands as a
|
|
file in their repo under their workplan, not only an inbox entry — *the
|
|
message that sat unread for four days is the baseline to beat*.
|
|
|
|
**Controls:**
|
|
- the number regenerates from a single command, and `facts.toml` carries
|
|
it if anything else quotes it (§Single source of fact — `make
|
|
facts-check`);
|
|
- the report carries the row-level table the ruling requires;
|
|
- **the seed range and policy are in the number's name**, not in a
|
|
footnote.
|
|
|
|
## Task: evidence
|
|
|
|
```task
|
|
id: CB-WP-0025-T07
|
|
status: todo
|
|
priority: high
|
|
```
|
|
|
|
`evidence/CB-EV-0023-*.md`.
|
|
|
|
- **Was the game winnable**, for the maintainer's actual lost game. That is
|
|
the acceptance test with a face on it.
|
|
- **What the honest search cost against the omniscient one**, since the
|
|
review will have pressed hardest there.
|
|
- **Whether the difficulty measurement moved ground-game**, or sat.
|
|
- **What tier L cost against what it caught** — third full-weight L pass in
|
|
the project, and the second in this chaos window.
|
|
- **Quote CB-WP-0024's cost by re-running the instrument.**
|