caught me repeating C1 specs/RetrospectiveAnalysis.md v1.0 plus games/ground/benches/search.rs, which exists because ADR-0013 D7 refused to let the spec quote either disputed figure. THE BENCHMARK'S OWN FIRST FIXTURE WAS DEFECTIVE, and it is the same defect class the review caught one layer up. Stopping at a fixed step 20 put 2p and 4p in states where seat 0 had NO legal commands, so it timed an empty Vec (~120 ns) and silently skipped validate_fold because there was nothing to validate. It now advances until the seat has a real branch and ASSERTS it. A clone benchmark was added too: a search must copy state per branch, and iter_batched excludes setup from timing, so without it the budget would again rest on an unmeasured span. Measured at real decision points: legal_commands 4.06-4.76 us, clone 378-639 ns, validate+fold 0.5-3.8 us. Per-child cost is NOT uniform -- some commands resolve cascades -- so budgets use the upper end (~5 us/child). That settles D3 with real numbers. Joint branching over the last two rounds is ~5x10^2 / 1.6x10^5 / 5.7x10^5 at 2/3/4 seats, so K=2 costs negligible / 0.8 s / 2.9 s and holds at two to four seats. It does NOT hold at five or six, where the tool must reduce K and say that it did rather than silently searching less. §4.1 is a normative prohibition, not a preference: a single policy's win rate MAY NOT be reported as a difficulty. The spec carries the measured reason -- greedy 100% against first-legal 0% on identical deals -- because this project already made that error and nearly exported it to a repo that is blocked waiting on the number. §2.3 makes the empty-result wording normative: "no winning line found in the last K rounds", never "unwinnable". A bounded search cannot establish unwinnability and that sentence is what a player who just lost reads. Also corrected: the T01 completion record still asserted all three withdrawn claims as fact. It now carries claimed / withdrawn / survives explicitly rather than being rewritten -- a retraction that does not propagate to every place the claim lives is how the earlier ones survived. make all: exit 0. loop-lint clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
7.8 KiB
RetrospectiveAnalysis — was this deal winnable, and how hard is the game
v1.0 — CB-WP-0025 T04, 2026-08-05. Normative. Implements ADR-0013. Admissibility of anything this produces is governed by GameDesign.md §1.
Two capabilities, one machine. Was this deal winnable? is a search over a finished game. How hard is the game? is that search run over many deals and counted — not a bot's win rate (§4.1).
1. The question, and its name
"Given the deal as it actually was, was there a line of play that reached the threshold?"
Never labelled "how you should have played." The distinction is the honest content of the feature: the tool answers a question about the deal, and a label promising advice about the player turns a true answer into a false lesson.
Strategy fusion does not apply and must not be invoked as an objection. Fusion (Frank, Basin & Matsubara 1998) is a defect of aggregating over determinizations to choose a move. After the game there is one world — the deal is known — so a line found in it is executable in the only world there is. This is why the affordable option is also the honest one.
2. The witness
A witness is a sequence of joint selections that, replayed from the
recorded initial state, ends with group_success == true.
2.1 It must replay — hard gate, not a metric
100% of emitted witnesses replay through the existing scenario runner and end in
group_success.
Not a target: a gate. A witness that does not replay asserts the opposite of the truth to a player who just lost, which is worse than emitting nothing.
2.2 Every move carries its information dependence
ADR-0013 D2. Each move in a witness is marked:
| mark | meaning |
|---|---|
visible |
everything the move depended on was in the acting seat's projection at that point — the Problem face-up, the Solution in that seat's own hand |
hidden |
it was not |
Computed from project(), which already exists and whose hiding rules are
already asserted by games_ground::view.
This replaces the structural boundary the survey wanted. Searching a
GroundView is not implementable — a view cannot fold events — so the
guarantee moved from the search cannot see it to the answer says which
moves needed it. A witness reads:
"This deal was winnable. Two of these six moves needed a card you had no way to know was coming."
Falsifier: a witness whose moves are all visible but which no seat
could have chosen means the marking is wrong. A test constructs that case.
2.3 Wording when nothing is found
"No winning line found in the last K rounds."
Never "unwinnable". A bounded search cannot establish unwinnability, and this sentence is what the player reads.
3. The bound
Exhaustive over the last K rounds, table treated as one co-operative
agent choosing joint selections. K = 2 by default.
Bounded in rounds, not nodes: "winnable from round 4" means something to a player; "winnable within 100,000 nodes" does not. A node budget is a secondary cut that aborts with a stated reason so a wide table cannot hang the page.
3.1 Measured cost, and what it permits
cargo bench -p games-ground --bench search — the single source for these
numbers (ADR-0013 D7). Mid-game states at real decision points:
| seats | branch width | legal_commands |
clone |
validate+fold |
|---|---|---|---|---|
| 2 | 5 | 4.06 µs | 378 ns | 696 ns |
| 3 | 8 | 4.13 µs | 432 ns | 508 ns |
| 4 | 11 | 4.76 µs | 639 ns | 3.76 µs |
Per-child cost is not uniform — validate+fold ranges 0.5–3.8 µs
depending on which command is taken, because some resolve cascades and
some do not. Budgets use the upper end, so ~5 µs per child
(clone + validate + fold).
Joint branching over the last two rounds, from the measured per-seat widths:
| seats | joint / 2 rounds | at ~5 µs/child |
|---|---|---|
| 2 | ~5×10² | negligible |
| 3 | ~1.6×10⁵ | ~0.8 s |
| 4 | ~5.7×10⁵ | ~2.9 s |
So K = 2 is affordable at two, three and four seats, and is not at
five or six — joint branching there exceeds 10⁶ per round. Five and six
seats require a smaller K, and the tool must reduce it and say that it
did rather than silently searching less.
The published 112–161 µs/node figure is withdrawn (CB-RES-0008 §1.2, challenge C1) and must not be quoted from anywhere.
4. Difficulty
4.1 A bot's win rate is not a difficulty
Normative prohibition, because this project already made the error and nearly exported it:
A win rate from a single policy may not be reported as a difficulty.
Measured, on identical deals: GreedyPolicy wins 100% at five and six
seats where a FirstLegal policy — take legal[0], no heuristic — wins
0%; at two seats FirstLegal (77.5%) beats greedy (66.0%). Two
unsophisticated agents span the entire range.
And a measure that improves when the measurer improves is not measuring the subject: a better bot would make the game "easier" with no rule changing.
4.2 What is reported instead
Winnable fraction — over N deals at a seat count, the proportion in which the search finds a winning line within its bound.
A property of the deal distribution and the threshold, which is what
ground-game tunes. Ships as a table, never one number:
| column | what it is |
|---|---|
| winnable fraction | can the deal be won at all — bounded, K stated |
| reference-policy win rate | what a named policy achieves |
| skill gap | the difference — how much play has to supply |
It is a lower bound and must be labelled one. A K-round search cannot
see a line that required round 1, so the figure is
"winnable-from-round-(6−K)", never "winnable".
Every rate carries its policy, N, seed range and K in the number's name, not in a footnote — GameDesign §1.2, and the reason the withdrawn finding was inadmissible.
4.3 Resolution — the number that makes it usable
The smallest threshold change the measurement can distinguish, with its N.
"We can tell a threshold of 5 from 7 but not 7 from 8" is more useful to
ground-game than any rate with no error bar, and it is what makes the
figure a tuning instrument rather than a statistic.
5. The instruments must be able to fail
ADR-0013 D5, and GameDesign §1.3. Before any figure from these tools is quoted anywhere:
- positive controls — a deal constructed to be unwinnable returns none; a deal constructed to be winnable returns a witness that replays;
--self-test, wired intomake self-testslike every other reporting tool;- one command regenerates the figure (
make difficulty); - the policy panel is plural — at least
greedy,randomandfirst-legal. The spread between them is the finding §4.1 rests on, and reporting one policy would restore the error.
games/ground/examples/difficulty-baseline.rs currently satisfies none
of the first three and is inadmissible until it does. It has no
assertions, no self-test, and no make target — nothing can turn it red,
which under CB-WP-0022 T05's role distinction makes it a default
artifact wearing a counterexample label.
6. Falsifiers for this spec
- §2.2 fails if a witness is emitted whose moves are all
visiblebut which no seat could have chosen. Then the marking must be derived from the search rather than checked after it. - §4.2 fails if the winnable fraction turns out to be ~100% or ~0% at every seat count and threshold — it would then have no resolution (§4.3) and be as useless as the bot rate it replaced.
- §3 fails if
K = 2proves unaffordable in practice at four seats; the measured 2.9 s is a projection from branch widths, not a timing of the real search.