Five remarks from the maintainer's test games, checked against the code before being written down — two had already reached ground-game on wrong premises, so a claim now names the line that makes it true. Three of the five turned out to be data the projection already carries, drawn as text: solution_deck_len, solution_discard, and OutcomeView's personal/mastery/winners. One control (`close — I have read this`) is labelled as a reading but shuts the server down, and leaves a live-looking page pointing at a dead port. One number does not exist at all: table.rs loops run_game and keeps only the last summary. CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the other seats' plays as objects on the table, the ending control saying what it does, and a tally that survives "play again". CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how hard is this" are the same search asked twice. Tier L because the information boundary is the whole design problem: a solver reading GroundState sees the deck the rules hide, and would tell the maintainer he could have won by playing a card he had no way to know was there. Also unblocks GROUND-WP-0005, active with both tasks waiting on a measured difficulty baseline. Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023 shipped CB-EV-0020 first. loop-lint: no findings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
11 KiB
| id | kind | title | status |
|---|---|---|---|
| CB-WP-0025 | product | Could we have won: a path out of a lost game, and how hard the game actually is | ready |
Purpose
structural tier L (creates a new capability — a search over game state,
and a measurement the engine does not currently take;
both produce numbers ground-game will tune against)
chaos d8 = 6 → no override
declared tier L
Declaration 8 of chaos window 2. Tier L: separate survey, adversarial review, ADR, then spec, then code.
Two remarks, and why they are one pass
"I had a game where we lost and in this case I would have liked to know if and how we could have won… the best path is not computable I guess so a path to win is fine."
"Do we have difficulty estimations? If so we should show them. It will help tuning the game. I felt it was too easy but then we lost, so who knows."
They are the same machine asked two questions. Was this game winnable? is a search from a recorded state. How hard is this game? is that search run over many deals and counted. Building the second without the first gives a win-rate with no witness; building the first without the second gives one anecdote per game.
"I felt it was too easy but then we lost, so who knows" is the finding.
The maintainer cannot calibrate the game from play, and that is precisely
the gap ground-game is currently blocked in: GROUND-WP-0005
(Difficulty tiers (threshold ratio) and optional Pressure dial) is active
with both its tasks in wait. Tiers cannot be set without a measured
baseline, and clay-borg is the thing that can measure. This pass is what
unblocks it — which is the design-instrument aspect CB-WP-0022 is in the
middle of stating, arriving with a concrete demand.
What already exists, so the survey does not re-find it
- The state is replayable.
cb-game-runtimerecords sessions as scenarios;replay.rsandmake replay-testalready re-run them. A search does not need new persistence. - The move space is enumerable.
legal_commandsexists and, since CB-WP-0023, is narrow enough to be worth trusting — SOLVE is offered only where it can act, so the branching factor is real rather than inflated by inert moves. - Bots exist.
games/ground/src/bot.rshasGreedyPolicyandRandomPolicy, wired throughbot_policy(table.rs:208). A win rate over N seeds is reachable with what is already there — the question is whether that number means anything, which is the survey's problem. - The threshold is public.
OutcomeView.total/.threshold/.group_success. Difficulty has a denominator already.
What makes this hard, and must not be waved through
The game is not perfect-information and the search must respect that.
A path computed with the deck known is a path the players could never have
found. view.rs hides the deck, other seats' hands, and face-down
selections by rule (GR-S02/S04, GR-R02/R04). A retrospective solver
running on GroundState sees all of it. So the ADR must decide, in
words, which of these three the tool answers:
- was this deal winnable by an omniscient player — cheap, honest, and answers a question nobody asked;
- was it winnable from what the seats could see — the question actually asked, and the expensive one;
- did a reasonable line exist — a bounded search from the losing seat's information, which may be the only affordable honest answer.
Getting this wrong produces a feature that tells the maintainer he could have won by playing a card he had no way to know was there. That is worse than not shipping it.
And a difficulty number is a claim about a distribution. One win rate
over one bot policy over N seeds is not "the difficulty"; it is that
policy's win rate. Whatever the spec adopts must name its policy, its N,
and its seed range, or ground-game will tune tiers against a number
whose meaning drifts the next time a bot improves.
Task: survey
id: CB-WP-0025-T01
status: todo
priority: high
research/CB-RES-0008-*.md, with tier: L and the chaos roll recorded
(loop-lint checks both).
Per §Step 1 the survey is done when it can name a benchmark-to-beat per dimension — a number or a reproducible comparison, not an impression.
- Retrospective solvers in games with hidden information. The prior art is real and should be named: determinized search (perfect-information Monte Carlo) and its known failure — strategy fusion, where a determinizing solver claims lines that require knowing which world it is in. That failure is exactly the trap in §What makes this hard. Bridge and Skat post-mortem tools are the closest analogues; poker solvers are the well-studied case and the wrong shape.
- "A path to win" as a product, not a proof. The maintainer already
conceded optimality ("the best path is not computable I guess"). So
the target is a witness: one concrete line of play that reaches
group_success, or a defensible no line found within bound B. Name what a witness must carry to be checkable. - Difficulty as a measured quantity in co-operative games. Pandemic and its relatives set difficulty by a dial with a published win rate. The benchmark-to-beat is: can we produce a win rate whose confidence interval is tight enough to distinguish two threshold settings?
- Cost. Search over an event-sourced aggregate with full
validateon every branch has a per-node price. Measure it on our machine, on our scenarios — the runnable-baseline option applies here, since a search that cannot finish while the player is still looking at the page is a different feature.
Task: adversarial review
id: CB-WP-0025-T02
status: todo
priority: high
Tier L requires it. Exactly one round: challenge, then response, trail in
history/, unpolished. Require an attempt at:
- that the honest version is unaffordable — that a search respecting the information rule is too expensive or too weak to find anything, so the shipped tool will quietly become the omniscient one with a reassuring label;
- that a witness misleads more than it helps — being shown a line that needed a card you could not know about teaches a wrong lesson about the game, and the tool would be better refusing to answer;
- that the difficulty number is a bot benchmark wearing a difficulty
costume, and
ground-gamewill tune the game against our bot rather than against play; - that this is CB-WP-0022's job — the design instrument is being built right now, and a difficulty measurement is a finding-producing tool. The strongest counter is that the register records findings and this produces them, but the reviewer should press whether that is a distinction worth a separate capability.
Task: decide
id: CB-WP-0025-T03
status: todo
priority: high
decisions/ADR-0013-*.md. (ADR-0012 is CB-WP-0022 T03's.) At minimum:
- which question the solver answers, from the three in §What makes this hard, and what it is called in the UI — the name must not overclaim;
- the information boundary: whether the search runs on
GroundStateor on aGroundView, and if on state, what stops it using what the view hides. Note that running on the view makes the rule structural rather than a promise, and that this is the cheapest guarantee available; - the bound: depth, node budget, or wall clock, and what no path found means against it — a bounded search that says "unwinnable" is lying, and the wording must say "none found within B";
- whether difficulty ships as one number or a small table, and what it is a function of: policy, seat count, threshold, seed range;
- where it lives — a new crate, a mode of
cb-play, or a tool undertools/. The tier was declared L on the assumption of a new capability; if the ADR concludes it is a mode of an existing one, say so, and the over-declaration is a chaos-window data point worth recording.
Task: specify
id: CB-WP-0025-T04
status: todo
priority: high
specs/ — extend MetricsAndScenarios.md or add a capability spec as the
ADR directs, with metrics, because a spec without them is prose.
Candidates, to be argued not adopted:
- witness checkability — every path the tool emits replays through the
existing scenario runner and ends in
group_success. Target 100%, and it is a hard gate, not a metric: a path that does not replay is a bug that says the opposite of the truth; - search cost — nodes and wall clock at the chosen bound, on the recorded games we have;
- difficulty resolution — the smallest threshold difference the
measurement can distinguish, with its N. This is the number
ground-gameneeds, and stating it as "we can tell 5 from 7 but not 7 from 8" is more useful than a win rate with no error bar.
Per ground-game's ruling (GROUND-WP-0004 T02), any arithmetic finding
this produces ships a runnable reproduction and a row-level table — never
a summed figure. A difficulty number is arithmetic, and it is exactly the
kind that has already gone wrong twice.
Task: build the witness
id: CB-WP-0025-T05
status: todo
priority: high
The search, the bound, and the replayable path. Wire it to the ending page so a lost game can be asked the question — the page CB-WP-0024 T01 is already reworking, so land that first or expect a conflict.
A game that was won is not asked the question. The feature exists for a loss.
Controls:
- every emitted witness replays to
group_successthrough the existing runner — asserted, not spot-checked; - a deal constructed to be unwinnable returns none found, and the test says which construction makes it so;
- the information boundary is mutation-provable: relax it, and a test naming that boundary goes red. If it cannot be mutated, it was a comment rather than a rule.
Task: measure the difficulty, and hand it to ground-game
id: CB-WP-0025-T06
status: todo
priority: high
Run the measurement, ship it as a make target beside the other
instruments, and show the result in the game — the maintainer asked for it
to be visible, and a number in a file will not calibrate anything.
Then send it to ground-game against GROUND-WP-0005, which is active
with both tasks waiting on exactly this. Per CB-WP-0022 T06, it lands as a
file in their repo under their workplan, not only an inbox entry — the
message that sat unread for four days is the baseline to beat.
Controls:
- the number regenerates from a single command, and
facts.tomlcarries it if anything else quotes it (§Single source of fact —make facts-check); - the report carries the row-level table the ruling requires;
- the seed range and policy are in the number's name, not in a footnote.
Task: evidence
id: CB-WP-0025-T07
status: todo
priority: high
evidence/CB-EV-0023-*.md.
- Was the game winnable, for the maintainer's actual lost game. That is the acceptance test with a face on it.
- What the honest search cost against the omniscient one, since the review will have pressed hardest there.
- Whether the difficulty measurement moved ground-game, or sat.
- What tier L cost against what it caught — third full-weight L pass in the project, and the second in this chaos window.
- Quote CB-WP-0024's cost by re-running the instrument.