CB-WP-0025 T03: ADR-0013 -- one world after the game, and difficulty that
does not measure the bot Seven decisions. Two are not what T03 expected, because the review moved the ground under both. D1: strategy fusion DOES NOT APPLY, and that is why the affordable option is also the honest one. Fusion is a defect of aggregating over determinizations to choose a move -- the search picking different actions in states a player cannot distinguish. After the game there is ONE WORLD: the deal is known, so a search over it yields a line executable in the only world there is. The survey treated fusion as this pass's central obstacle; it is an obstacle to a playing engine, which we are not building. The tool answers "given the deal as it actually was, was there a line that reached the threshold" and is labelled that way on screen -- never "how you should have played". D4: difficulty is the WINNABLE FRACTION, not a bot's win rate. C4 killed the bot rate -- two trivial policies span 0-100% on the same deals, and improving the bot would make the game "easier" without a rule changing. A measure that moves when the measurer improves is not measuring the thing. The solver supplies the alternative: over N deals, in what proportion does a winning line exist. That is a property of the deal distribution and the threshold, which is what ground-game tunes and what GROUND-WP-0005 is blocked on. Ships as a table -- winnable fraction, named reference policy, skill gap -- never one number, with policy/N/seed-range/K in the number's name. Stated as a LOWER BOUND, since a K-round search cannot see a line that needed round 1. D2: search GroundState. The survey's view-only structural boundary is not implementable -- a view cannot fold events, and it said so in §6. The guarantee moves to something checkable: every move in a witness is marked visible or hidden, computed from project(). A witness reads "you could have won, but two of these six moves needed a card you had no way to know was coming" -- more useful than either extreme. D3: bounded exhaustive over the last K rounds, table as one co-operative agent, K=2 default. Affordable once C6 corrected the premise: joint branching over the last two rounds is ~5x10^2 / 1.6x10^5 / 5.7x10^5 at 2/3/4 seats. Wording is normative: "no winning line found in the last K rounds", never "unwinnable". PIMC and ISMCTS rejected -- they add strategy fusion to a problem that does not have it. D5: the harness becomes an instrument before any figure is quoted (C3) -- positive controls, --self-test in make self-tests, a make target, and a PLURAL policy panel, because the spread between policies is what C4 exposed and hiding it would restore the error. D6: no new crate and no port. The L declaration was an over-declaration and is recorded as one. D7: the node cost is disputed 5x between author and reviewer; T04 benchmarks it with criterion and neither existing figure may be cited, including by this ADR. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
88e25ae7c4
commit
3a026b1e1f
2 changed files with 263 additions and 1 deletions
230
decisions/ADR-0013-could-we-have-won.md
Normal file
230
decisions/ADR-0013-could-we-have-won.md
Normal file
|
|
@ -0,0 +1,230 @@
|
|||
# ADR-0013: answer the retrospective question, and measure difficulty with the solver rather than with a bot
|
||||
|
||||
status: accepted
|
||||
date: 2026-08-05
|
||||
decided by: agent, under the standing loop authorization
|
||||
tier: L (structural L — a new analysis capability whose numbers `ground-game`
|
||||
will tune against; chaos d8=6 → no override). Tier L: separate survey,
|
||||
adversarial review, then this document.
|
||||
references: [CB-WP-0025](../workplans/CB-WP-0025-could-we-have-won.md),
|
||||
[CB-RES-0008](../research/CB-RES-0008-could-we-have-won.md),
|
||||
[challenge](../history/260805-could-we-have-won-challenge.md) /
|
||||
[response](../history/260805-could-we-have-won-response.md),
|
||||
[ADR-0012](ADR-0012-the-design-instrument.md) (admissibility),
|
||||
[GameDesign.md](../specs/GameDesign.md),
|
||||
GROUND-WP-0004 T02, GROUND-WP-0005 (blocked on D4)
|
||||
|
||||
## Context
|
||||
|
||||
The maintainer asked two things: *"we lost — could we have won, and how?"*
|
||||
and *"do we have difficulty estimations?"*
|
||||
|
||||
**The survey answered the second and was wrong.** It measured
|
||||
`GreedyPolicy` winning 200/200 at five and six seats and called the game
|
||||
too easy there. A `FirstLegal` policy — `legal[0]`, no heuristic — scores
|
||||
**0%** on the same deals. Two unsophisticated agents span the whole range,
|
||||
so the measurement was about the policy.
|
||||
|
||||
That failure is not incidental to this ADR; **it determines D4.**
|
||||
|
||||
## The premise that changed, and it changes the algorithm
|
||||
|
||||
The survey said exhaustive search was impossible and reached for
|
||||
determinized sampling, which carries strategy fusion. **Both halves were
|
||||
wrong.**
|
||||
|
||||
- Its per-node cost was **30–50× too high** (a timer bracketing whole
|
||||
games). Corrected: ~3–4 µs per `legal_commands` call, with the exact
|
||||
figure still disputed (§D7).
|
||||
- Bounded exhaustive search is **affordable**: measured ~3 s over the last
|
||||
two rounds at three seats.
|
||||
|
||||
Joint branching, treating the table as one co-operative agent — the
|
||||
product over seats of the measured per-seat branching:
|
||||
|
||||
| seats | per-seat mean | joint per round | last 2 rounds |
|
||||
|---|---:|---:|---:|
|
||||
| 2 | 4.7 | ~22 | ~5×10² |
|
||||
| 3 | 7.4 | ~405 | ~1.6×10⁵ |
|
||||
| 4 | 9.1 | ~754 | ~5.7×10⁵ |
|
||||
|
||||
Against a ~10⁵–10⁶ node budget, **the last two rounds are exhaustively
|
||||
searchable at two, three and four seats.** Five rounds is not, at any seat
|
||||
count.
|
||||
|
||||
---
|
||||
|
||||
## D1 — answer the *retrospective* question, and say so in those words
|
||||
|
||||
Three questions were on the table (CB-RES-0008 §3). The tool answers:
|
||||
|
||||
> **"Given the deal as it actually was, was there a line of play that
|
||||
> reached the threshold — and here is one."**
|
||||
|
||||
**Strategy fusion does not apply to this question, and that is the whole
|
||||
reason it is the affordable one.** Fusion is a defect of *aggregating over
|
||||
determinizations to choose a move*: the search picks different actions in
|
||||
states the player cannot distinguish. **After the game there is one
|
||||
world.** The deck is known, the deal is known, and a search over that
|
||||
single world produces a line that is executable in it — because it is the
|
||||
only world there is.
|
||||
|
||||
The survey treated fusion as an obstacle to this pass. It is an obstacle
|
||||
to a *playing* engine. We are not building one.
|
||||
|
||||
**What remains true is that the line may have been unfindable at the
|
||||
time**, and D2 handles that by annotation rather than by refusing to
|
||||
answer.
|
||||
|
||||
**On screen it is called** *"was this deal winnable?"* — never *"how you
|
||||
should have played"*. The distinction is the honest content of the
|
||||
feature, and a label that overclaims turns a true answer into a false
|
||||
lesson.
|
||||
|
||||
## D2 — run on `GroundState`, and mark each move's information dependence
|
||||
|
||||
The survey's preferred guarantee was structural: search a `GroundView` so
|
||||
the boundary cannot be crossed. **It is not implementable** — a view
|
||||
cannot `fold` events, so a search needs a state it may not see. The survey
|
||||
said so in §6 and was right to.
|
||||
|
||||
Decision: **search `GroundState`** — legitimate here, because post-game
|
||||
the deal is public (`solution_discard` already is, and the game is over) —
|
||||
and move the honesty guarantee to something checkable:
|
||||
|
||||
> **Every move in an emitted witness is marked `visible` or `hidden`.**
|
||||
> A move is `visible` if, at the point it is played, everything it depends
|
||||
> on was in the acting seat's projection: the target Problem face-up, the
|
||||
> Solution in that seat's own hand. Otherwise `hidden`.
|
||||
|
||||
So a witness reads *"you could have won — but two of these six moves
|
||||
needed a card you had no way to know was coming."* **That is a more useful
|
||||
answer than either extreme**, and it is computed from `project()`, which
|
||||
already exists and is already tested.
|
||||
|
||||
**Falsifier:** if a witness is emitted whose moves are all marked
|
||||
`visible` but which no seat could actually have chosen, the marking is
|
||||
wrong and D2 has failed. A test constructs exactly that case.
|
||||
|
||||
## D3 — bounded exhaustive over the endgame, `K` rounds, and honest wording
|
||||
|
||||
**Exhaustive search over the last `K` rounds**, with the table treated as
|
||||
one co-operative agent choosing joint selections. `K = 2` by default,
|
||||
which the measurements put inside budget at 2–4 seats.
|
||||
|
||||
- The bound is **rounds**, not nodes or seconds, because rounds are what a
|
||||
player understands: *"winnable from round 4"* means something; *"winnable
|
||||
within 100,000 nodes"* does not.
|
||||
- A node budget is a **secondary** cut that aborts with a stated reason,
|
||||
so a wide table cannot hang the page.
|
||||
- **Wording is normative.** When no line is found the tool says
|
||||
**"no winning line found in the last K rounds"** — never *"unwinnable"*.
|
||||
A bounded search that claims unwinnability is lying, and this is the
|
||||
sentence the maintainer will read.
|
||||
|
||||
**Not chosen: determinized sampling (PIMC), ISMCTS.** Both are for playing
|
||||
under uncertainty. Here there is one world (D1), so they would add strategy
|
||||
fusion to a problem that does not have it.
|
||||
|
||||
## D4 — difficulty is the **winnable fraction**, not any bot's win rate
|
||||
|
||||
**This is the decision the review forced, and it is the useful half of the
|
||||
pass.**
|
||||
|
||||
A single-policy win rate cannot be a difficulty: two trivial policies span
|
||||
0–100% on the same deals. Worse, *improving the bot would make the game
|
||||
"easier"* without a rule changing — a measure that moves when the
|
||||
measurer improves is not measuring the thing.
|
||||
|
||||
The solver supplies a policy-independent alternative:
|
||||
|
||||
> **Winnable fraction** — over N deals at a seat count, the proportion in
|
||||
> which the search finds *any* winning line within its bound.
|
||||
|
||||
That is a property of **the deal distribution and the threshold**, which
|
||||
is exactly what `ground-game` tunes. It is the number GROUND-WP-0005 is
|
||||
blocked on, and the bot rate never was.
|
||||
|
||||
Difficulty therefore ships as **a small table, never one number**:
|
||||
|
||||
| column | what it is |
|
||||
|---|---|
|
||||
| winnable fraction | can the deal be won at all (bounded, K stated) |
|
||||
| reference-policy win rate | what a stated bot achieves — **named policy** |
|
||||
| skill gap | the difference: how much play has to supply |
|
||||
|
||||
**Every rate carries its policy, its N, its seed range and its K in the
|
||||
number's name**, not in a footnote. A figure that loses them is
|
||||
inadmissible under GameDesign §1.2.
|
||||
|
||||
**Bounded-below caveat, stated because it will be quoted:** the winnable
|
||||
fraction from a K-round search is a **lower bound** on true winnability —
|
||||
a deal unwinnable in the last 2 rounds may have been winnable in round 1.
|
||||
The report says "winnable-from-round-(6−K)", never "winnable".
|
||||
|
||||
## D5 — the harness becomes an instrument before any figure is quoted
|
||||
|
||||
C3 established that `difficulty-baseline.rs` has no assertions, no
|
||||
`--self-test` and no `make` target — nothing can turn it red. Under
|
||||
CB-WP-0022 T05's own `role` distinction it is a `default` artifact wearing
|
||||
a `counterexample` label, and **GameDesign §1.3 makes it inadmissible.**
|
||||
|
||||
Required before T06 reports anything:
|
||||
|
||||
- **positive controls** — a deal constructed to be unwinnable returns
|
||||
none; a deal constructed to be winnable returns a witness that replays;
|
||||
- **`--self-test`**, wired into `make self-tests` like every other
|
||||
reporting tool;
|
||||
- **`make difficulty`** (or equivalent), so the figure regenerates from
|
||||
one command;
|
||||
- the **policy panel is plural**: at least `greedy`, `random` and
|
||||
`first-legal`, because the spread between them is what C4 exposed and
|
||||
hiding it would restore the error.
|
||||
|
||||
## D6 — it lives in `games/ground`, not a new crate
|
||||
|
||||
The search needs `validate`, `fold`, `legal_commands` and `project` —
|
||||
all of `games_ground`. A separate crate would either re-export the
|
||||
aggregate or take a dependency on it and add nothing.
|
||||
|
||||
**The tier was declared L on the assumption of a new capability port.
|
||||
There is no port**, and that over-declaration is recorded rather than
|
||||
hidden — it is a data point for the tier rules, and the L weight paid for
|
||||
itself twice over regardless (§Consequences).
|
||||
|
||||
`cb-play` gains a mode to ask the question about a finished game; the
|
||||
difficulty sweep is an example/binary, as the baseline is.
|
||||
|
||||
## D7 — the per-node cost is unsettled and T04 must benchmark it
|
||||
|
||||
The author measured **3.0–4.1 µs**, the reviewer **15.6–20.4 µs**, by
|
||||
different isolations. Both agree the published 112–161 µs was wrong by
|
||||
1–2 orders; neither has established which is right.
|
||||
|
||||
**T04 benchmarks it with `criterion`** — already a dev-dependency, already
|
||||
used by `benches/synthetic.rs` — and the spec quotes that number and no
|
||||
other. **Neither figure above may be cited**, including by this ADR.
|
||||
|
||||
## Consequences
|
||||
|
||||
- `specs/` gains the witness contract and the difficulty table's shape
|
||||
(T04), plus the benchmarked node cost.
|
||||
- T05 builds the K-round search, the `visible`/`hidden` marking, and the
|
||||
replay check.
|
||||
- **T06's payload changes completely.** It reports a winnable fraction and
|
||||
a policy panel to GROUND-WP-0005 — *not* the withdrawn "too easy at 5–6
|
||||
seats". The withdrawal itself is reported, per ADR-0012 D5.
|
||||
- The register gains the withdrawn finding as `inconsistent` /
|
||||
`withdrawn`, so it is in the log rather than forgotten.
|
||||
|
||||
## What was rejected
|
||||
|
||||
| rejected | why |
|
||||
|---|---|
|
||||
| determinized sampling / PIMC / ISMCTS | strategy fusion, added to a question that has one world (D1) |
|
||||
| a view-only search as a structural boundary | not implementable — a view cannot fold events |
|
||||
| refusing to answer unless the line was findable at the time | throws away a true and useful answer; annotate instead (D2) |
|
||||
| a single bot win rate as "difficulty" | two trivial policies span 0–100% on the same deals (C4) |
|
||||
| "unwinnable" as output wording | a bounded search cannot know it (D3) |
|
||||
| a new crate | no port exists; it would re-export the aggregate (D6) |
|
||||
| quoting either measured node cost | they disagree 5× and neither is established (D7) |
|
||||
|
|
@ -235,7 +235,7 @@ this project have now caught a false headline that every gate passed.**
|
|||
|
||||
```task
|
||||
id: CB-WP-0025-T03
|
||||
status: todo
|
||||
status: done
|
||||
priority: high
|
||||
state_hub_task_id: "055630a1-56e0-4ad4-8796-35f444313371"
|
||||
```
|
||||
|
|
@ -258,6 +258,38 @@ state_hub_task_id: "055630a1-56e0-4ad4-8796-35f444313371"
|
|||
if the ADR concludes it is a mode of an existing one, say so, and the
|
||||
over-declaration is a chaos-window data point worth recording.
|
||||
|
||||
**Done 2026-08-05.**
|
||||
[ADR-0013](../decisions/ADR-0013-could-we-have-won.md), seven decisions.
|
||||
**Two of them are not what T03 was written expecting**, because the review
|
||||
moved the ground under both.
|
||||
|
||||
- **D1 — strategy fusion does not apply, and that is why this is
|
||||
affordable.** Fusion is a defect of *aggregating over determinizations
|
||||
to choose a move*. **After the game there is one world**: the deal is
|
||||
known, so a search over it produces a line executable in the only world
|
||||
there is. The survey treated fusion as this pass's central obstacle; it
|
||||
is an obstacle to a *playing* engine, which we are not building.
|
||||
- **D4 — difficulty is the WINNABLE FRACTION, not any bot's win rate.**
|
||||
C4 killed the bot rate: two trivial policies span 0–100% on the same
|
||||
deals, and improving the bot would make the game "easier" without a rule
|
||||
changing. The solver supplies a policy-independent measure — *over N
|
||||
deals, in what proportion does a winning line exist* — which is a
|
||||
property of the deal distribution and the threshold, and is what
|
||||
GROUND-WP-0005 actually needs. **The bot rate never was.**
|
||||
|
||||
**D2** searches `GroundState` (the survey's view-only boundary is not
|
||||
implementable — a view cannot fold events) and moves the guarantee to a
|
||||
checkable per-move `visible`/`hidden` marking computed from `project()`.
|
||||
A witness reads *"you could have won, but two of these six moves needed a
|
||||
card you had no way to know was coming."* **D3** is bounded exhaustive
|
||||
over the last K rounds — measured affordable at 2–4 seats once C6
|
||||
corrected the premise — with normative wording: *"no winning line found in
|
||||
the last K rounds"*, never *"unwinnable"*. **D5** makes the harness an
|
||||
instrument before any figure is quoted (C3). **D6**: no new crate, no
|
||||
port — **the L declaration was an over-declaration and is recorded as
|
||||
one**. **D7**: the node cost is disputed 5× and T04 must benchmark it;
|
||||
neither figure may be cited, including by the ADR.
|
||||
|
||||
## Task: specify
|
||||
|
||||
```task
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue