CB-WP-0025 T04: specs/RetrospectiveAnalysis.md, and a benchmark that
caught me repeating C1
specs/RetrospectiveAnalysis.md v1.0 plus games/ground/benches/search.rs,
which exists because ADR-0013 D7 refused to let the spec quote either
disputed figure.
THE BENCHMARK'S OWN FIRST FIXTURE WAS DEFECTIVE, and it is the same defect
class the review caught one layer up. Stopping at a fixed step 20 put 2p
and 4p in states where seat 0 had NO legal commands, so it timed an empty
Vec (~120 ns) and silently skipped validate_fold because there was nothing
to validate. It now advances until the seat has a real branch and ASSERTS
it. A clone benchmark was added too: a search must copy state per branch,
and iter_batched excludes setup from timing, so without it the budget
would again rest on an unmeasured span.
Measured at real decision points: legal_commands 4.06-4.76 us, clone
378-639 ns, validate+fold 0.5-3.8 us. Per-child cost is NOT uniform --
some commands resolve cascades -- so budgets use the upper end (~5
us/child).
That settles D3 with real numbers. Joint branching over the last two
rounds is ~5x10^2 / 1.6x10^5 / 5.7x10^5 at 2/3/4 seats, so K=2 costs
negligible / 0.8 s / 2.9 s and holds at two to four seats. It does NOT
hold at five or six, where the tool must reduce K and say that it did
rather than silently searching less.
§4.1 is a normative prohibition, not a preference: a single policy's win
rate MAY NOT be reported as a difficulty. The spec carries the measured
reason -- greedy 100% against first-legal 0% on identical deals -- because
this project already made that error and nearly exported it to a repo that
is blocked waiting on the number.
§2.3 makes the empty-result wording normative: "no winning line found in
the last K rounds", never "unwinnable". A bounded search cannot establish
unwinnability and that sentence is what a player who just lost reads.
Also corrected: the T01 completion record still asserted all three
withdrawn claims as fact. It now carries claimed / withdrawn / survives
explicitly rather than being rewritten -- a retraction that does not
propagate to every place the claim lives is how the earlier ones survived.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 18:59:46 +02:00
# RetrospectiveAnalysis — was this deal winnable, and how hard is the game
v1.0 — CB-WP-0025 T04, 2026-08-05. Normative. Implements
[ADR-0013 ](../decisions/ADR-0013-could-we-have-won.md ). Admissibility of
anything this produces is governed by
[GameDesign.md ](GameDesign.md ) §1.
**Two capabilities, one machine.** *Was this deal winnable?* is a search
over a finished game. *How hard is the game?* is that search run over many
deals and counted — **not** a bot's win rate (§4.1).
---
## 1. The question, and its name
> **"Given the deal as it actually was, was there a line of play that
> reached the threshold?"**
**Never labelled "how you should have played."** The distinction is the
honest content of the feature: the tool answers a question about the
*deal*, and a label promising advice about the *player* turns a true
answer into a false lesson.
**Strategy fusion does not apply and must not be invoked as an objection.**
Fusion (Frank, Basin & Matsubara 1998) is a defect of aggregating over
determinizations to choose a move. After the game there is **one world** —
the deal is known — so a line found in it is executable in the only world
there is. This is why the affordable option is also the honest one.
## 2. The witness
A witness is a sequence of joint selections that, replayed from the
recorded initial state, ends with `group_success == true` .
### 2.1 It must replay — hard gate, not a metric
> **100% of emitted witnesses replay through the existing scenario runner
> and end in `group_success`.**
Not a target: a **gate** . A witness that does not replay asserts the
opposite of the truth to a player who just lost, which is worse than
emitting nothing.
### 2.2 Every move carries its information dependence
ADR-0013 D2. Each move in a witness is marked:
| mark | meaning |
|---|---|
| `visible` | everything the move depended on was in the acting seat's projection at that point — the Problem face-up, the Solution in that seat's own hand |
| `hidden` | it was not |
Computed from `project()` , which already exists and whose hiding rules are
already asserted by `games_ground::view` .
**This replaces the structural boundary the survey wanted.** Searching a
`GroundView` is not implementable — a view cannot `fold` events — so the
guarantee moved from *the search cannot see it* to *the answer says which
moves needed it*. A witness reads:
> *"This deal was winnable. Two of these six moves needed a card you had
> no way to know was coming."*
**Falsifier:** a witness whose moves are all `visible` but which no seat
could have chosen means the marking is wrong. A test constructs that case.
### 2.3 Wording when nothing is found
> **"No winning line found in the last K rounds."**
**Never "unwinnable".** A bounded search cannot establish unwinnability,
and this sentence is what the player reads.
## 3. The bound
**Exhaustive over the last `K` rounds**, table treated as one co-operative
CB-WP-0025 T05: the search works, and it falsified this pass's own
affordability projection
games/ground/src/search.rs, five tests. It finds real winning lines and
replays them through validate/fold to group_success.
Two bugs in my own work, found and fixed here.
THE TRAVERSAL WAS WRONG. It branched on "the first seat with any legal
command" and stopped there, so a later seat never acted if an earlier one
was already selected but still had a legal move. Restructured around what
the rules oblige: a seat without a selection MUST select (GR-R02) and
nothing else can happen first; after Reveal the optional actions branch
freely and the aggregate rejects Resolve until the obligatory ones are
done -- so the search needs no phase logic of its own.
AND MY REWIND WAS OFF BY ONE ROUND, replaying the round it was meant to
search. That is why the first run reported 3 nodes and looked like a
working search.
Measured with the real search, rewinding real games to the start of their
last K rounds:
2p K=1 exhausted, 8,103 nodes, ~29 ms
2p K=2 budget cut at 2,000,000 nodes, ~5 s
3p K=2 win found, 41 nodes, ~157 us
The spec's own falsifier said "§3 fails if K=2 proves unaffordable at four
seats". IT FAILED AT TWO. The projection assumed a joint product per
round; the search explores sequential per-seat decisions, so orderings
multiply the tree far beyond width^seats. That is the second projection
this pass published in place of a measurement -- C1's timer was the first.
THE ASYMMETRY IS THE OPERATIVE FINDING. Finding a win is cheap: DFS
stumbles onto one in tens of nodes. Proving none exists needs exhaustion.
So the witness feature is affordable now at any K a player would ask
about, and the winnable fraction (ADR-0013 D4) is NOT, because its
negative half must exhaust every deal it counts. K=1 is the honest default
for exhaustive answers today; making K=2 exhaustible needs transposition
or move-ordering, neither of which this pass built. specs §3 and §3.1
corrected accordingly, and the K=2 default withdrawn.
The negative control that makes "winnable" falsifiable: 2p seed 7 over its
last round returns NoneFound with exhausted=true in ~8k nodes -- a real
negative, not a budget cut wearing a verdict's clothes. And the visible/
hidden marking is tested both ways, since a marking that can only say YES
is decoration.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 19:10:25 +02:00
agent choosing joint selections.
**`K = 1` for an exhaustive answer; `K` may be larger when a witness is
all that is wanted.** ADR-0013 said `K = 2` by default; §3.1's measurement
overrides it, and the difference is which question is being asked:
| answer | needs | affordable `K` today |
|---|---|---|
| *"here is a winning line"* | one success | 2+ — DFS finds one in tens of nodes |
| *"there is no winning line"* | exhaustion | **1** — `K=2` exceeded 2× 10⁶ nodes at two seats |
A `K` that cannot be exhausted may still emit a witness; it may **not**
report `NoneFound { exhausted: true }` , and the type keeps those apart.
CB-WP-0025 T04: specs/RetrospectiveAnalysis.md, and a benchmark that
caught me repeating C1
specs/RetrospectiveAnalysis.md v1.0 plus games/ground/benches/search.rs,
which exists because ADR-0013 D7 refused to let the spec quote either
disputed figure.
THE BENCHMARK'S OWN FIRST FIXTURE WAS DEFECTIVE, and it is the same defect
class the review caught one layer up. Stopping at a fixed step 20 put 2p
and 4p in states where seat 0 had NO legal commands, so it timed an empty
Vec (~120 ns) and silently skipped validate_fold because there was nothing
to validate. It now advances until the seat has a real branch and ASSERTS
it. A clone benchmark was added too: a search must copy state per branch,
and iter_batched excludes setup from timing, so without it the budget
would again rest on an unmeasured span.
Measured at real decision points: legal_commands 4.06-4.76 us, clone
378-639 ns, validate+fold 0.5-3.8 us. Per-child cost is NOT uniform --
some commands resolve cascades -- so budgets use the upper end (~5
us/child).
That settles D3 with real numbers. Joint branching over the last two
rounds is ~5x10^2 / 1.6x10^5 / 5.7x10^5 at 2/3/4 seats, so K=2 costs
negligible / 0.8 s / 2.9 s and holds at two to four seats. It does NOT
hold at five or six, where the tool must reduce K and say that it did
rather than silently searching less.
§4.1 is a normative prohibition, not a preference: a single policy's win
rate MAY NOT be reported as a difficulty. The spec carries the measured
reason -- greedy 100% against first-legal 0% on identical deals -- because
this project already made that error and nearly exported it to a repo that
is blocked waiting on the number.
§2.3 makes the empty-result wording normative: "no winning line found in
the last K rounds", never "unwinnable". A bounded search cannot establish
unwinnability and that sentence is what a player who just lost reads.
Also corrected: the T01 completion record still asserted all three
withdrawn claims as fact. It now carries claimed / withdrawn / survives
explicitly rather than being rewritten -- a retraction that does not
propagate to every place the claim lives is how the earlier ones survived.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 18:59:46 +02:00
Bounded in **rounds** , not nodes: *"winnable from round 4"* means something
to a player; *"winnable within 100,000 nodes"* does not. A node budget is a
secondary cut that aborts with a stated reason so a wide table cannot hang
the page.
### 3.1 Measured cost, and what it permits
`cargo bench -p games-ground --bench search` — the single source for these
numbers (ADR-0013 D7). Mid-game states at real decision points:
| seats | branch width | `legal_commands` | `clone` | `validate+fold` |
|---|---:|---:|---:|---:|
| 2 | 5 | 4.06 µs | 378 ns | 696 ns |
| 3 | 8 | 4.13 µs | 432 ns | 508 ns |
| 4 | 11 | 4.76 µs | 639 ns | **3.76 µs** |
**Per-child cost is not uniform** — `validate+fold` ranges 0.5– 3.8 µs
depending on which command is taken, because some resolve cascades and
some do not. **Budgets use the upper end** , so ~5 µs per child
(clone + validate + fold).
Joint branching over the last two rounds, from the measured per-seat
widths:
| seats | joint / 2 rounds | at ~5 µs/child |
|---|---:|---:|
| 2 | ~5× 10² | negligible |
| 3 | ~1.6× 10⁵ | ** ~0.8 s** |
| 4 | ~5.7× 10⁵ | ** ~2.9 s** |
CB-WP-0025 T05: the search works, and it falsified this pass's own
affordability projection
games/ground/src/search.rs, five tests. It finds real winning lines and
replays them through validate/fold to group_success.
Two bugs in my own work, found and fixed here.
THE TRAVERSAL WAS WRONG. It branched on "the first seat with any legal
command" and stopped there, so a later seat never acted if an earlier one
was already selected but still had a legal move. Restructured around what
the rules oblige: a seat without a selection MUST select (GR-R02) and
nothing else can happen first; after Reveal the optional actions branch
freely and the aggregate rejects Resolve until the obligatory ones are
done -- so the search needs no phase logic of its own.
AND MY REWIND WAS OFF BY ONE ROUND, replaying the round it was meant to
search. That is why the first run reported 3 nodes and looked like a
working search.
Measured with the real search, rewinding real games to the start of their
last K rounds:
2p K=1 exhausted, 8,103 nodes, ~29 ms
2p K=2 budget cut at 2,000,000 nodes, ~5 s
3p K=2 win found, 41 nodes, ~157 us
The spec's own falsifier said "§3 fails if K=2 proves unaffordable at four
seats". IT FAILED AT TWO. The projection assumed a joint product per
round; the search explores sequential per-seat decisions, so orderings
multiply the tree far beyond width^seats. That is the second projection
this pass published in place of a measurement -- C1's timer was the first.
THE ASYMMETRY IS THE OPERATIVE FINDING. Finding a win is cheap: DFS
stumbles onto one in tens of nodes. Proving none exists needs exhaustion.
So the witness feature is affordable now at any K a player would ask
about, and the winnable fraction (ADR-0013 D4) is NOT, because its
negative half must exhaust every deal it counts. K=1 is the honest default
for exhaustive answers today; making K=2 exhaustible needs transposition
or move-ordering, neither of which this pass built. specs §3 and §3.1
corrected accordingly, and the K=2 default withdrawn.
The negative control that makes "winnable" falsifiable: 2p seed 7 over its
last round returns NoneFound with exhausted=true in ~8k nodes -- a real
negative, not a budget cut wearing a verdict's clothes. And the visible/
hidden marking is tested both ways, since a marking that can only say YES
is decoration.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 19:10:25 +02:00
> ### The projection above was wrong, and the real search falsified it
>
> **Measured 2026-08-05 with the search built in T05**, rewinding real
> games to the start of their last `K` rounds:
>
> | case | result |
> |---|---|
> | 2p, `K=1` | **exhausted** in 8,103 nodes, ~29 ms — a real negative |
> | 2p, `K=2` | **budget cut** at 2,000,000 nodes, ~5 s — not exhausted |
> | 3p, `K=2` | win found in 41 nodes, ~157 µs |
>
> §6's falsifier said *"§3 fails if K=2 proves unaffordable in practice at
> four seats"*. **It failed at two.**
>
> The projection assumed a joint product per round. The search explores
> **sequential per-seat decisions**, and the post-Reveal phase branches
> over every seat's options at every level, so orderings multiply the tree
> far beyond `width^seats`.
>
> **And the asymmetry is the operative fact:** *finding* a win is cheap —
> depth-first stumbles onto one in tens of nodes — while *proving none
> exists* is expensive, because it must exhaust the space. So:
>
> - **the witness feature (§2) is affordable now**, at any `K` a player
> would ask about;
> - **the winnable fraction (§4.2) is not**, because its "not winnable"
> half requires exhaustion on every deal it counts.
>
> `K = 1` is the honest default for exhaustive answers today. Making
> `K = 2` exhaustible needs transposition or move-ordering, neither of
> which this pass built.
CB-WP-0025 T04: specs/RetrospectiveAnalysis.md, and a benchmark that
caught me repeating C1
specs/RetrospectiveAnalysis.md v1.0 plus games/ground/benches/search.rs,
which exists because ADR-0013 D7 refused to let the spec quote either
disputed figure.
THE BENCHMARK'S OWN FIRST FIXTURE WAS DEFECTIVE, and it is the same defect
class the review caught one layer up. Stopping at a fixed step 20 put 2p
and 4p in states where seat 0 had NO legal commands, so it timed an empty
Vec (~120 ns) and silently skipped validate_fold because there was nothing
to validate. It now advances until the seat has a real branch and ASSERTS
it. A clone benchmark was added too: a search must copy state per branch,
and iter_batched excludes setup from timing, so without it the budget
would again rest on an unmeasured span.
Measured at real decision points: legal_commands 4.06-4.76 us, clone
378-639 ns, validate+fold 0.5-3.8 us. Per-child cost is NOT uniform --
some commands resolve cascades -- so budgets use the upper end (~5
us/child).
That settles D3 with real numbers. Joint branching over the last two
rounds is ~5x10^2 / 1.6x10^5 / 5.7x10^5 at 2/3/4 seats, so K=2 costs
negligible / 0.8 s / 2.9 s and holds at two to four seats. It does NOT
hold at five or six, where the tool must reduce K and say that it did
rather than silently searching less.
§4.1 is a normative prohibition, not a preference: a single policy's win
rate MAY NOT be reported as a difficulty. The spec carries the measured
reason -- greedy 100% against first-legal 0% on identical deals -- because
this project already made that error and nearly exported it to a repo that
is blocked waiting on the number.
§2.3 makes the empty-result wording normative: "no winning line found in
the last K rounds", never "unwinnable". A bounded search cannot establish
unwinnability and that sentence is what a player who just lost reads.
Also corrected: the T01 completion record still asserted all three
withdrawn claims as fact. It now carries claimed / withdrawn / survives
explicitly rather than being rewritten -- a retraction that does not
propagate to every place the claim lives is how the earlier ones survived.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 18:59:46 +02:00
**The published 112– 161 µs/node figure is withdrawn** (CB-RES-0008 §1.2,
challenge C1) and must not be quoted from anywhere.
## 4. Difficulty
### 4.1 A bot's win rate is not a difficulty
**Normative prohibition**, because this project already made the error and
nearly exported it:
> A win rate from a single policy **may not be reported as a difficulty**.
Measured, on identical deals: `GreedyPolicy` wins **100%** at five and six
seats where a `FirstLegal` policy — take `legal[0]` , no heuristic — wins
**0%**; at two seats `FirstLegal` (77.5%) *beats* greedy (66.0%). Two
unsophisticated agents span the entire range.
And a measure that improves when the *measurer* improves is not measuring
the subject: a better bot would make the game "easier" with no rule
changing.
### 4.2 What is reported instead
> **Winnable fraction** — over N deals at a seat count, the proportion in
> which the search finds a winning line within its bound.
A property of the **deal distribution and the threshold** , which is what
`ground-game` tunes. Ships as a table, never one number:
| column | what it is |
|---|---|
| winnable fraction | can the deal be won at all — bounded, `K` stated |
| reference-policy win rate | what a **named** policy achieves |
| skill gap | the difference — how much play has to supply |
**It is a lower bound and must be labelled one.** A `K` -round search cannot
see a line that required round 1, so the figure is
**"winnable-from-round-(6− K)"**, never "winnable".
**Every rate carries its policy, N, seed range and K in the number's
name**, not in a footnote — GameDesign §1.2, and the reason the withdrawn
finding was inadmissible.
### 4.3 Resolution — the number that makes it usable
> The smallest threshold change the measurement can distinguish, with its
> N.
*"We can tell a threshold of 5 from 7 but not 7 from 8"* is more useful to
`ground-game` than any rate with no error bar, and it is what makes the
figure a tuning instrument rather than a statistic.
## 5. The instruments must be able to fail
ADR-0013 D5, and GameDesign §1.3. Before **any** figure from these tools is
quoted anywhere:
- **positive controls** — a deal constructed to be unwinnable returns
none; a deal constructed to be winnable returns a witness that replays;
- **`--self-test` **, wired into `make self-tests` like every other
reporting tool;
- **one command regenerates the figure** (`make difficulty` );
- **the policy panel is plural** — at least `greedy` , `random` and
`first-legal` . The spread between them is the finding §4.1 rests on, and
reporting one policy would restore the error.
**`games/ground/examples/difficulty-baseline.rs` currently satisfies none
of the first three** and is inadmissible until it does. It has no
assertions, no self-test, and no `make` target — nothing can turn it red,
which under CB-WP-0022 T05's `role` distinction makes it a `default`
artifact wearing a `counterexample` label.
## 6. Falsifiers for this spec
- **§2.2 fails** if a witness is emitted whose moves are all `visible` but
which no seat could have chosen. Then the marking must be derived from
the search rather than checked after it.
- **§4.2 fails** if the winnable fraction turns out to be ~100% or ~0% at
every seat count and threshold — it would then have no resolution (§4.3)
and be as useless as the bot rate it replaced.
- **§3 fails** if `K = 2` proves unaffordable in practice at four seats;
the measured 2.9 s is a projection from branch widths, not a timing of
the real search.