--- id: CB-WP-0025 kind: product title: "Could we have won: a path out of a lost game, and how hard the game actually is" status: active state_hub_workstream_id: "a866982f-94e7-432b-b0d5-2fefb781a574" --- # Purpose ``` structural tier L (creates a new capability — a search over game state, and a measurement the engine does not currently take; both produce numbers ground-game will tune against) chaos d8 = 6 → no override declared tier L ``` Declaration 8 of chaos window 2. Tier L: separate survey, **adversarial review**, ADR, then spec, then code. ## Two remarks, and why they are one pass > *"I had a game where we lost and in this case I would have liked to know > if and how we could have won… the best path is not computable I guess so > a path to win is fine."* > *"Do we have difficulty estimations? If so we should show them. It will > help tuning the game. I felt it was too easy but then we lost, so who > knows."* They are the same machine asked two questions. *Was this game winnable?* is a search from a recorded state. *How hard is this game?* is that search run over many deals and counted. Building the second without the first gives a win-rate with no witness; building the first without the second gives one anecdote per game. **"I felt it was too easy but then we lost, so who knows" is the finding.** The maintainer cannot calibrate the game from play, and that is precisely the gap `ground-game` is currently blocked in: **GROUND-WP-0005** *(Difficulty tiers (threshold ratio) and optional Pressure dial)* is active with **both its tasks in `wait`**. Tiers cannot be set without a measured baseline, and clay-borg is the thing that can measure. This pass is what unblocks it — which is the design-instrument aspect CB-WP-0022 is in the middle of stating, arriving with a concrete demand. ## What already exists, so the survey does not re-find it - **The state is replayable.** `cb-game-runtime` records sessions as scenarios; `replay.rs` and `make replay-test` already re-run them. A search does not need new persistence. - **The move space is enumerable.** `legal_commands` exists and, since CB-WP-0023, is narrow enough to be worth trusting — SOLVE is offered only where it can act, so the branching factor is real rather than inflated by inert moves. - **Bots exist.** `games/ground/src/bot.rs` has `GreedyPolicy` and `RandomPolicy`, wired through `bot_policy` (`table.rs:208`). A win rate over N seeds is reachable with what is already there — the question is whether that number *means* anything, which is the survey's problem. - **The threshold is public.** `OutcomeView.total` / `.threshold` / `.group_success`. Difficulty has a denominator already. ## What makes this hard, and must not be waved through **The game is not perfect-information and the search must respect that.** A path computed with the deck known is a path the players could never have found. `view.rs` hides the deck, other seats' hands, and face-down selections *by rule* (GR-S02/S04, GR-R02/R04). A retrospective solver running on `GroundState` sees all of it. So the ADR must decide, in words, **which of these three the tool answers**: - *was this deal winnable by an omniscient player* — cheap, honest, and answers a question nobody asked; - *was it winnable from what the seats could see* — the question actually asked, and the expensive one; - *did a reasonable line exist* — a bounded search from the losing seat's information, which may be the only affordable honest answer. Getting this wrong produces a feature that tells the maintainer he could have won by playing a card he had no way to know was there. **That is worse than not shipping it.** **And a difficulty number is a claim about a distribution.** One win rate over one bot policy over N seeds is not "the difficulty"; it is that policy's win rate. Whatever the spec adopts must name its policy, its N, and its seed range, or `ground-game` will tune tiers against a number whose meaning drifts the next time a bot improves. ## Task: survey ```task id: CB-WP-0025-T01 status: done priority: high state_hub_task_id: "556cfd24-5992-4cc0-be90-0b60e989b5bb" ``` `research/CB-RES-0008-*.md`, with `tier: L` and the chaos roll recorded (`loop-lint` checks both). Per §Step 1 the survey is done when it can name a **benchmark-to-beat** per dimension — a number or a reproducible comparison, not an impression. - **Retrospective solvers in games with hidden information.** The prior art is real and should be named: determinized search (perfect-information Monte Carlo) and its known failure — *strategy fusion*, where a determinizing solver claims lines that require knowing which world it is in. That failure is exactly the trap in §What makes this hard. Bridge and Skat post-mortem tools are the closest analogues; poker solvers are the well-studied case and the wrong shape. - **"A path to win" as a product, not a proof.** The maintainer already conceded optimality (*"the best path is not computable I guess"*). So the target is a **witness**: one concrete line of play that reaches `group_success`, or a defensible *no line found within bound B*. Name what a witness must carry to be checkable. - **Difficulty as a measured quantity in co-operative games.** Pandemic and its relatives set difficulty by a dial with a published win rate. The benchmark-to-beat is: can we produce a win rate whose confidence interval is tight enough to distinguish two threshold settings? - **Cost.** Search over an event-sourced aggregate with full `validate` on every branch has a per-node price. Measure it on our machine, on our scenarios — the runnable-baseline option applies here, since a search that cannot finish while the player is still looking at the page is a different feature. **Done 2026-08-05.** [CB-RES-0008](../research/CB-RES-0008-could-we-have-won.md), with a runnable baseline (`games/ground/examples/difficulty-baseline.rs`). **The baseline produced a finding before any solver exists, and it is the biggest thing in this pass.** A greedy bot wins **200 of 200** games at five and six seats — with a median margin of +3 and 11.8–12.0 points available against a threshold of 9. The curve is 66% / 82.5% / 95% / 100% / 100% across 2/3/4/5/6 seats. The row-level table shows why: available points go **6 / 9 / 12** against thresholds **5 / 7 / 9**, so the ratio *rises* with seat count (1.20 → 1.29 → 1.33) while the table also gains actions per round. **Three multipliers pointing the same way.** **It also explains the maintainer's report without a solver.** *"I felt it was too easy but then we lost"* — 66% at two seats is a real game; 100% at six is not. Both halves are true of different seat counts. **Cost measured, and it rules out the obvious approach.** Branching is small (mean 4.7–9.1) but `legal_commands` costs **112–161 µs** per call, because it filters candidates through full `validate`. Exhaustive search is out at every seat count; ~10⁴–10⁵ nodes is 1.4–14 s, which is the budget the ADR must design inside. **Prior art names the trap.** Determinized search (PIMC) suffers *strategy fusion* — Frank, Basin & Matsubara 1998 — where the search picks different actions in states a real player cannot tell apart. A witness built that way may require knowing what was on top of the deck. **And it would still replay green**, so the checkability benchmark does not catch it. Honesty and checkability are different properties, stated so T03 cannot conflate them. ## Task: adversarial review ```task id: CB-WP-0025-T02 status: done priority: high state_hub_task_id: "bdd9fc97-4706-4d8e-b286-91ff53339680" ``` Tier L requires it. Exactly one round: challenge, then response, trail in `history/`, unpolished. Require an attempt at: - **that the honest version is unaffordable** — that a search respecting the information rule is too expensive or too weak to find anything, so the shipped tool will quietly become the omniscient one with a reassuring label; - **that a witness misleads more than it helps** — being shown a line that needed a card you could not know about teaches a wrong lesson about the game, and the tool would be better refusing to answer; - **that the difficulty number is a bot benchmark wearing a difficulty costume**, and `ground-game` will tune the game against our bot rather than against play; - **that this is CB-WP-0022's job** — the design instrument is being built right now, and a difficulty measurement is a finding-producing tool. The strongest counter is that the register records findings and this *produces* them, but the reviewer should press whether that is a distinction worth a separate capability. **Done 2026-08-05.** Trail: [challenge](../history/260805-could-we-have-won-challenge.md), [response](../history/260805-could-we-have-won-response.md). Separate agent, as in CB-WP-0022. **Six of seven conceded, and the survey's headline finding is withdrawn — not softened.** **C4 is the one that kills it, and the reviewer ranked it fourth.** Measured: a `FirstLegal` policy — take `legal[0]`, no heuristic — scores **0% at five and six seats** where greedy scores 100%, and **77.5% at two seats** where greedy scores 66%. **Two unsophisticated agents span the entire range at the same seat count**, so *"the game is too easy at 5–6 seats"* is a statement about `GreedyPolicy`, not about GROUND. Their offered rescue — greedy hits the 12-point ceiling every time, so it is a rules claim — dies on the same data. **C1: the per-node cost was wrong by 30–50×.** The timer bracketed two `setup`s, a whole greedy game and a validate+fold replay, then divided by player decisions. **The tell was in my own published output**: the figure *fell* as branching *rose*, which no per-enumeration cost can do. Re-measured at **3.0–4.1 µs**; the reviewer got 15.6–20.4 by a different isolation and **that discrepancy is unsettled** — T04 benchmarks it. **C6: exhaustive search is not out**, which changes T03's premise. With C1's correction the budget is ~10⁵–10⁶ nodes and bounded endgame search fits — so the ADR cannot open with *"exhaustive is impossible, therefore determinized sampling"*, especially since sampling carries strategy fusion that exhaustive search does not. **C3: the finding failed the admissibility rule this project wrote nine hours earlier** — sums where GROUND-WP-0004 T02 requires per-priority rows, and a harness with no assertions, no `--self-test` and no `make` target, so nothing can turn it red. A `default` artifact wearing a `counterexample` label. **C2: the ratio explains nothing** — 3p and 4p share deal, threshold and ratio, and differ by 12.5 points of win rate. **C5: the "explains the maintainer's report" claim is contradicted by `lib.rs:2487`**, which records his losses as 3-player on the pre-ruling deal, unwinnable at 6 against 7. **T06 was pointed at GROUND-WP-0005, which is blocked on this number.** Sending it would have invited threshold changes on one bot's behaviour — the fifth wrong premise to reach ground-game. **Both tier-L reviews in this project have now caught a false headline that every gate passed.** ## Task: decide ```task id: CB-WP-0025-T03 status: done priority: high state_hub_task_id: "055630a1-56e0-4ad4-8796-35f444313371" ``` `decisions/ADR-0013-*.md`. (ADR-0012 is CB-WP-0022 T03's.) At minimum: - **which question the solver answers**, from the three in §What makes this hard, and what it is called in the UI — the name must not overclaim; - **the information boundary**: whether the search runs on `GroundState` or on a `GroundView`, and if on state, what stops it using what the view hides. Note that running on the view makes the rule structural rather than a promise, and that this is the cheapest guarantee available; - **the bound**: depth, node budget, or wall clock, and what *no path found* means against it — a bounded search that says "unwinnable" is lying, and the wording must say "none found within B"; - **whether difficulty ships as one number or a small table**, and what it is a function of: policy, seat count, threshold, seed range; - **where it lives** — a new crate, a mode of `cb-play`, or a tool under `tools/`. The tier was declared L on the assumption of a new capability; if the ADR concludes it is a mode of an existing one, say so, and the over-declaration is a chaos-window data point worth recording. **Done 2026-08-05.** [ADR-0013](../decisions/ADR-0013-could-we-have-won.md), seven decisions. **Two of them are not what T03 was written expecting**, because the review moved the ground under both. - **D1 — strategy fusion does not apply, and that is why this is affordable.** Fusion is a defect of *aggregating over determinizations to choose a move*. **After the game there is one world**: the deal is known, so a search over it produces a line executable in the only world there is. The survey treated fusion as this pass's central obstacle; it is an obstacle to a *playing* engine, which we are not building. - **D4 — difficulty is the WINNABLE FRACTION, not any bot's win rate.** C4 killed the bot rate: two trivial policies span 0–100% on the same deals, and improving the bot would make the game "easier" without a rule changing. The solver supplies a policy-independent measure — *over N deals, in what proportion does a winning line exist* — which is a property of the deal distribution and the threshold, and is what GROUND-WP-0005 actually needs. **The bot rate never was.** **D2** searches `GroundState` (the survey's view-only boundary is not implementable — a view cannot fold events) and moves the guarantee to a checkable per-move `visible`/`hidden` marking computed from `project()`. A witness reads *"you could have won, but two of these six moves needed a card you had no way to know was coming."* **D3** is bounded exhaustive over the last K rounds — measured affordable at 2–4 seats once C6 corrected the premise — with normative wording: *"no winning line found in the last K rounds"*, never *"unwinnable"*. **D5** makes the harness an instrument before any figure is quoted (C3). **D6**: no new crate, no port — **the L declaration was an over-declaration and is recorded as one**. **D7**: the node cost is disputed 5× and T04 must benchmark it; neither figure may be cited, including by the ADR. ## Task: specify ```task id: CB-WP-0025-T04 status: todo priority: high state_hub_task_id: "7c3ca25d-0001-44ef-a07d-0ab2d51b09a5" ``` `specs/` — extend `MetricsAndScenarios.md` or add a capability spec as the ADR directs, with metrics, because a spec without them is prose. Candidates, to be argued not adopted: - **witness checkability** — every path the tool emits replays through the existing scenario runner and ends in `group_success`. Target 100%, and it is a hard gate, not a metric: a path that does not replay is a bug that says the opposite of the truth; - **search cost** — nodes and wall clock at the chosen bound, on the recorded games we have; - **difficulty resolution** — the smallest threshold difference the measurement can distinguish, with its N. This is the number `ground-game` needs, and stating it as *"we can tell 5 from 7 but not 7 from 8"* is more useful than a win rate with no error bar. Per `ground-game`'s ruling (GROUND-WP-0004 T02), **any arithmetic finding this produces ships a runnable reproduction and a row-level table** — never a summed figure. A difficulty number is arithmetic, and it is exactly the kind that has already gone wrong twice. ## Task: build the witness ```task id: CB-WP-0025-T05 status: todo priority: high state_hub_task_id: "5f1cbda3-7427-4f23-8c67-b2e216b987a8" ``` The search, the bound, and the replayable path. Wire it to the ending page so a lost game can be asked the question — the page CB-WP-0024 T01 is already reworking, so land that first or expect a conflict. **A game that was won is not asked the question.** The feature exists for a loss. **Controls:** - every emitted witness replays to `group_success` through the existing runner — asserted, not spot-checked; - a deal constructed to be unwinnable returns *none found*, and the test says which construction makes it so; - the information boundary is mutation-provable: relax it, and a test naming *that* boundary goes red. If it cannot be mutated, it was a comment rather than a rule. ## Task: measure the difficulty, and hand it to ground-game ```task id: CB-WP-0025-T06 status: todo priority: high state_hub_task_id: "723feae8-1ed4-4dd4-9330-8ec84625e5d9" ``` Run the measurement, ship it as a `make` target beside the other instruments, and show the result in the game — the maintainer asked for it to be visible, and a number in a file will not calibrate anything. Then send it to `ground-game` **against GROUND-WP-0005**, which is active with both tasks waiting on exactly this. Per CB-WP-0022 T06, it lands as a file in their repo under their workplan, not only an inbox entry — *the message that sat unread for four days is the baseline to beat*. **Controls:** - the number regenerates from a single command, and `facts.toml` carries it if anything else quotes it (§Single source of fact — `make facts-check`); - the report carries the row-level table the ruling requires; - **the seed range and policy are in the number's name**, not in a footnote. ## Task: evidence ```task id: CB-WP-0025-T07 status: todo priority: high state_hub_task_id: "2616e06c-5d88-4680-be18-4452515dff21" ``` `evidence/CB-EV-0023-*.md`. - **Was the game winnable**, for the maintainer's actual lost game. That is the acceptance test with a face on it. - **What the honest search cost against the omniscient one**, since the review will have pressed hardest there. - **Whether the difficulty measurement moved ground-game**, or sat. - **What tier L cost against what it caught** — third full-weight L pass in the project, and the second in this chaos window. - **Quote CB-WP-0024's cost by re-running the instrument.**