clay-borg/workplans/CB-WP-0025-could-we-have-won.md
tegwick 8e9b3c19b7
Some checks failed
ci / check (push) Failing after 4s
CB-WP-0024/0025: what play reported, split into a renderer and a search
Five remarks from the maintainer's test games, checked against the code
before being written down — two had already reached ground-game on wrong
premises, so a claim now names the line that makes it true.

Three of the five turned out to be data the projection already carries,
drawn as text: solution_deck_len, solution_discard, and OutcomeView's
personal/mastery/winners. One control (`close — I have read this`) is
labelled as a reading but shuts the server down, and leaves a live-looking
page pointing at a dead port. One number does not exist at all: table.rs
loops run_game and keeps only the last summary.

CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the
other seats' plays as objects on the table, the ending control saying what
it does, and a tally that survives "play again".

CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how
hard is this" are the same search asked twice. Tier L because the
information boundary is the whole design problem: a solver reading
GroundState sees the deck the rules hide, and would tell the maintainer he
could have won by playing a card he had no way to know was there. Also
unblocks GROUND-WP-0005, active with both tasks waiting on a measured
difficulty baseline.

Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023
shipped CB-EV-0020 first.

loop-lint: no findings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 12:59:08 +02:00

272 lines
11 KiB
Markdown

---
id: CB-WP-0025
kind: product
title: "Could we have won: a path out of a lost game, and how hard the game actually is"
status: ready
---
# Purpose
```
structural tier L (creates a new capability — a search over game state,
and a measurement the engine does not currently take;
both produce numbers ground-game will tune against)
chaos d8 = 6 → no override
declared tier L
```
Declaration 8 of chaos window 2. Tier L: separate survey, **adversarial
review**, ADR, then spec, then code.
## Two remarks, and why they are one pass
> *"I had a game where we lost and in this case I would have liked to know
> if and how we could have won… the best path is not computable I guess so
> a path to win is fine."*
> *"Do we have difficulty estimations? If so we should show them. It will
> help tuning the game. I felt it was too easy but then we lost, so who
> knows."*
They are the same machine asked two questions. *Was this game winnable?*
is a search from a recorded state. *How hard is this game?* is that search
run over many deals and counted. Building the second without the first
gives a win-rate with no witness; building the first without the second
gives one anecdote per game.
**"I felt it was too easy but then we lost, so who knows" is the finding.**
The maintainer cannot calibrate the game from play, and that is precisely
the gap `ground-game` is currently blocked in: **GROUND-WP-0005**
*(Difficulty tiers (threshold ratio) and optional Pressure dial)* is active
with **both its tasks in `wait`**. Tiers cannot be set without a measured
baseline, and clay-borg is the thing that can measure. This pass is what
unblocks it — which is the design-instrument aspect CB-WP-0022 is in the
middle of stating, arriving with a concrete demand.
## What already exists, so the survey does not re-find it
- **The state is replayable.** `cb-game-runtime` records sessions as
scenarios; `replay.rs` and `make replay-test` already re-run them.
A search does not need new persistence.
- **The move space is enumerable.** `legal_commands` exists and, since
CB-WP-0023, is narrow enough to be worth trusting — SOLVE is offered
only where it can act, so the branching factor is real rather than
inflated by inert moves.
- **Bots exist.** `games/ground/src/bot.rs` has `GreedyPolicy` and
`RandomPolicy`, wired through `bot_policy` (`table.rs:208`). A win rate
over N seeds is reachable with what is already there — the question is
whether that number *means* anything, which is the survey's problem.
- **The threshold is public.** `OutcomeView.total` / `.threshold` /
`.group_success`. Difficulty has a denominator already.
## What makes this hard, and must not be waved through
**The game is not perfect-information and the search must respect that.**
A path computed with the deck known is a path the players could never have
found. `view.rs` hides the deck, other seats' hands, and face-down
selections *by rule* (GR-S02/S04, GR-R02/R04). A retrospective solver
running on `GroundState` sees all of it. So the ADR must decide, in
words, **which of these three the tool answers**:
- *was this deal winnable by an omniscient player* — cheap, honest,
and answers a question nobody asked;
- *was it winnable from what the seats could see* — the question actually
asked, and the expensive one;
- *did a reasonable line exist* — a bounded search from the losing seat's
information, which may be the only affordable honest answer.
Getting this wrong produces a feature that tells the maintainer he could
have won by playing a card he had no way to know was there. **That is
worse than not shipping it.**
**And a difficulty number is a claim about a distribution.** One win rate
over one bot policy over N seeds is not "the difficulty"; it is that
policy's win rate. Whatever the spec adopts must name its policy, its N,
and its seed range, or `ground-game` will tune tiers against a number
whose meaning drifts the next time a bot improves.
## Task: survey
```task
id: CB-WP-0025-T01
status: todo
priority: high
```
`research/CB-RES-0008-*.md`, with `tier: L` and the chaos roll recorded
(`loop-lint` checks both).
Per §Step 1 the survey is done when it can name a **benchmark-to-beat**
per dimension — a number or a reproducible comparison, not an impression.
- **Retrospective solvers in games with hidden information.** The prior art
is real and should be named: determinized search (perfect-information
Monte Carlo) and its known failure — *strategy fusion*, where a
determinizing solver claims lines that require knowing which world it is
in. That failure is exactly the trap in §What makes this hard. Bridge
and Skat post-mortem tools are the closest analogues; poker solvers are
the well-studied case and the wrong shape.
- **"A path to win" as a product, not a proof.** The maintainer already
conceded optimality (*"the best path is not computable I guess"*). So
the target is a **witness**: one concrete line of play that reaches
`group_success`, or a defensible *no line found within bound B*. Name
what a witness must carry to be checkable.
- **Difficulty as a measured quantity in co-operative games.** Pandemic and
its relatives set difficulty by a dial with a published win rate. The
benchmark-to-beat is: can we produce a win rate whose confidence
interval is tight enough to distinguish two threshold settings?
- **Cost.** Search over an event-sourced aggregate with full `validate` on
every branch has a per-node price. Measure it on our machine, on our
scenarios — the runnable-baseline option applies here, since a search
that cannot finish while the player is still looking at the page is a
different feature.
## Task: adversarial review
```task
id: CB-WP-0025-T02
status: todo
priority: high
```
Tier L requires it. Exactly one round: challenge, then response, trail in
`history/`, unpolished. Require an attempt at:
- **that the honest version is unaffordable** — that a search respecting
the information rule is too expensive or too weak to find anything, so
the shipped tool will quietly become the omniscient one with a
reassuring label;
- **that a witness misleads more than it helps** — being shown a line that
needed a card you could not know about teaches a wrong lesson about the
game, and the tool would be better refusing to answer;
- **that the difficulty number is a bot benchmark wearing a difficulty
costume**, and `ground-game` will tune the game against our bot rather
than against play;
- **that this is CB-WP-0022's job** — the design instrument is being built
right now, and a difficulty measurement is a finding-producing tool. The
strongest counter is that the register records findings and this
*produces* them, but the reviewer should press whether that is a
distinction worth a separate capability.
## Task: decide
```task
id: CB-WP-0025-T03
status: todo
priority: high
```
`decisions/ADR-0013-*.md`. (ADR-0012 is CB-WP-0022 T03's.) At minimum:
- **which question the solver answers**, from the three in §What makes
this hard, and what it is called in the UI — the name must not overclaim;
- **the information boundary**: whether the search runs on `GroundState`
or on a `GroundView`, and if on state, what stops it using what the view
hides. Note that running on the view makes the rule structural rather
than a promise, and that this is the cheapest guarantee available;
- **the bound**: depth, node budget, or wall clock, and what *no path
found* means against it — a bounded search that says "unwinnable" is
lying, and the wording must say "none found within B";
- **whether difficulty ships as one number or a small table**, and what it
is a function of: policy, seat count, threshold, seed range;
- **where it lives** — a new crate, a mode of `cb-play`, or a tool under
`tools/`. The tier was declared L on the assumption of a new capability;
if the ADR concludes it is a mode of an existing one, say so, and the
over-declaration is a chaos-window data point worth recording.
## Task: specify
```task
id: CB-WP-0025-T04
status: todo
priority: high
```
`specs/` — extend `MetricsAndScenarios.md` or add a capability spec as the
ADR directs, with metrics, because a spec without them is prose.
Candidates, to be argued not adopted:
- **witness checkability** — every path the tool emits replays through the
existing scenario runner and ends in `group_success`. Target 100%, and it
is a hard gate, not a metric: a path that does not replay is a bug that
says the opposite of the truth;
- **search cost** — nodes and wall clock at the chosen bound, on the
recorded games we have;
- **difficulty resolution** — the smallest threshold difference the
measurement can distinguish, with its N. This is the number
`ground-game` needs, and stating it as *"we can tell 5 from 7 but not 7
from 8"* is more useful than a win rate with no error bar.
Per `ground-game`'s ruling (GROUND-WP-0004 T02), **any arithmetic finding
this produces ships a runnable reproduction and a row-level table** — never
a summed figure. A difficulty number is arithmetic, and it is exactly the
kind that has already gone wrong twice.
## Task: build the witness
```task
id: CB-WP-0025-T05
status: todo
priority: high
```
The search, the bound, and the replayable path. Wire it to the ending page
so a lost game can be asked the question — the page CB-WP-0024 T01 is
already reworking, so land that first or expect a conflict.
**A game that was won is not asked the question.** The feature exists for
a loss.
**Controls:**
- every emitted witness replays to `group_success` through the existing
runner — asserted, not spot-checked;
- a deal constructed to be unwinnable returns *none found*, and the test
says which construction makes it so;
- the information boundary is mutation-provable: relax it, and a test
naming *that* boundary goes red. If it cannot be mutated, it was a
comment rather than a rule.
## Task: measure the difficulty, and hand it to ground-game
```task
id: CB-WP-0025-T06
status: todo
priority: high
```
Run the measurement, ship it as a `make` target beside the other
instruments, and show the result in the game — the maintainer asked for it
to be visible, and a number in a file will not calibrate anything.
Then send it to `ground-game` **against GROUND-WP-0005**, which is active
with both tasks waiting on exactly this. Per CB-WP-0022 T06, it lands as a
file in their repo under their workplan, not only an inbox entry — *the
message that sat unread for four days is the baseline to beat*.
**Controls:**
- the number regenerates from a single command, and `facts.toml` carries
it if anything else quotes it (§Single source of fact — `make
facts-check`);
- the report carries the row-level table the ruling requires;
- **the seed range and policy are in the number's name**, not in a
footnote.
## Task: evidence
```task
id: CB-WP-0025-T07
status: todo
priority: high
```
`evidence/CB-EV-0023-*.md`.
- **Was the game winnable**, for the maintainer's actual lost game. That is
the acceptance test with a face on it.
- **What the honest search cost against the omniscient one**, since the
review will have pressed hardest there.
- **Whether the difficulty measurement moved ground-game**, or sat.
- **What tier L cost against what it caught** — third full-weight L pass in
the project, and the second in this chaos window.
- **Quote CB-WP-0024's cost by re-running the instrument.**