From 3a026b1e1f32150e77f0202718a8068a48632f96 Mon Sep 17 00:00:00 2001 From: tegwick Date: Wed, 5 Aug 2026 18:47:07 +0200 Subject: [PATCH] CB-WP-0025 T03: ADR-0013 -- one world after the game, and difficulty that does not measure the bot MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Seven decisions. Two are not what T03 expected, because the review moved the ground under both. D1: strategy fusion DOES NOT APPLY, and that is why the affordable option is also the honest one. Fusion is a defect of aggregating over determinizations to choose a move -- the search picking different actions in states a player cannot distinguish. After the game there is ONE WORLD: the deal is known, so a search over it yields a line executable in the only world there is. The survey treated fusion as this pass's central obstacle; it is an obstacle to a playing engine, which we are not building. The tool answers "given the deal as it actually was, was there a line that reached the threshold" and is labelled that way on screen -- never "how you should have played". D4: difficulty is the WINNABLE FRACTION, not a bot's win rate. C4 killed the bot rate -- two trivial policies span 0-100% on the same deals, and improving the bot would make the game "easier" without a rule changing. A measure that moves when the measurer improves is not measuring the thing. The solver supplies the alternative: over N deals, in what proportion does a winning line exist. That is a property of the deal distribution and the threshold, which is what ground-game tunes and what GROUND-WP-0005 is blocked on. Ships as a table -- winnable fraction, named reference policy, skill gap -- never one number, with policy/N/seed-range/K in the number's name. Stated as a LOWER BOUND, since a K-round search cannot see a line that needed round 1. D2: search GroundState. The survey's view-only structural boundary is not implementable -- a view cannot fold events, and it said so in §6. The guarantee moves to something checkable: every move in a witness is marked visible or hidden, computed from project(). A witness reads "you could have won, but two of these six moves needed a card you had no way to know was coming" -- more useful than either extreme. D3: bounded exhaustive over the last K rounds, table as one co-operative agent, K=2 default. Affordable once C6 corrected the premise: joint branching over the last two rounds is ~5x10^2 / 1.6x10^5 / 5.7x10^5 at 2/3/4 seats. Wording is normative: "no winning line found in the last K rounds", never "unwinnable". PIMC and ISMCTS rejected -- they add strategy fusion to a problem that does not have it. D5: the harness becomes an instrument before any figure is quoted (C3) -- positive controls, --self-test in make self-tests, a make target, and a PLURAL policy panel, because the spread between policies is what C4 exposed and hiding it would restore the error. D6: no new crate and no port. The L declaration was an over-declaration and is recorded as one. D7: the node cost is disputed 5x between author and reviewer; T04 benchmarks it with criterion and neither existing figure may be cited, including by this ADR. Co-Authored-By: Claude Opus 5 --- decisions/ADR-0013-could-we-have-won.md | 230 ++++++++++++++++++++++ workplans/CB-WP-0025-could-we-have-won.md | 34 +++- 2 files changed, 263 insertions(+), 1 deletion(-) create mode 100644 decisions/ADR-0013-could-we-have-won.md diff --git a/decisions/ADR-0013-could-we-have-won.md b/decisions/ADR-0013-could-we-have-won.md new file mode 100644 index 0000000..730f8ed --- /dev/null +++ b/decisions/ADR-0013-could-we-have-won.md @@ -0,0 +1,230 @@ +# ADR-0013: answer the retrospective question, and measure difficulty with the solver rather than with a bot + +status: accepted +date: 2026-08-05 +decided by: agent, under the standing loop authorization +tier: L (structural L — a new analysis capability whose numbers `ground-game` +will tune against; chaos d8=6 → no override). Tier L: separate survey, +adversarial review, then this document. +references: [CB-WP-0025](../workplans/CB-WP-0025-could-we-have-won.md), +[CB-RES-0008](../research/CB-RES-0008-could-we-have-won.md), +[challenge](../history/260805-could-we-have-won-challenge.md) / +[response](../history/260805-could-we-have-won-response.md), +[ADR-0012](ADR-0012-the-design-instrument.md) (admissibility), +[GameDesign.md](../specs/GameDesign.md), +GROUND-WP-0004 T02, GROUND-WP-0005 (blocked on D4) + +## Context + +The maintainer asked two things: *"we lost — could we have won, and how?"* +and *"do we have difficulty estimations?"* + +**The survey answered the second and was wrong.** It measured +`GreedyPolicy` winning 200/200 at five and six seats and called the game +too easy there. A `FirstLegal` policy — `legal[0]`, no heuristic — scores +**0%** on the same deals. Two unsophisticated agents span the whole range, +so the measurement was about the policy. + +That failure is not incidental to this ADR; **it determines D4.** + +## The premise that changed, and it changes the algorithm + +The survey said exhaustive search was impossible and reached for +determinized sampling, which carries strategy fusion. **Both halves were +wrong.** + +- Its per-node cost was **30–50× too high** (a timer bracketing whole + games). Corrected: ~3–4 µs per `legal_commands` call, with the exact + figure still disputed (§D7). +- Bounded exhaustive search is **affordable**: measured ~3 s over the last + two rounds at three seats. + +Joint branching, treating the table as one co-operative agent — the +product over seats of the measured per-seat branching: + +| seats | per-seat mean | joint per round | last 2 rounds | +|---|---:|---:|---:| +| 2 | 4.7 | ~22 | ~5×10² | +| 3 | 7.4 | ~405 | ~1.6×10⁵ | +| 4 | 9.1 | ~754 | ~5.7×10⁵ | + +Against a ~10⁵–10⁶ node budget, **the last two rounds are exhaustively +searchable at two, three and four seats.** Five rounds is not, at any seat +count. + +--- + +## D1 — answer the *retrospective* question, and say so in those words + +Three questions were on the table (CB-RES-0008 §3). The tool answers: + +> **"Given the deal as it actually was, was there a line of play that +> reached the threshold — and here is one."** + +**Strategy fusion does not apply to this question, and that is the whole +reason it is the affordable one.** Fusion is a defect of *aggregating over +determinizations to choose a move*: the search picks different actions in +states the player cannot distinguish. **After the game there is one +world.** The deck is known, the deal is known, and a search over that +single world produces a line that is executable in it — because it is the +only world there is. + +The survey treated fusion as an obstacle to this pass. It is an obstacle +to a *playing* engine. We are not building one. + +**What remains true is that the line may have been unfindable at the +time**, and D2 handles that by annotation rather than by refusing to +answer. + +**On screen it is called** *"was this deal winnable?"* — never *"how you +should have played"*. The distinction is the honest content of the +feature, and a label that overclaims turns a true answer into a false +lesson. + +## D2 — run on `GroundState`, and mark each move's information dependence + +The survey's preferred guarantee was structural: search a `GroundView` so +the boundary cannot be crossed. **It is not implementable** — a view +cannot `fold` events, so a search needs a state it may not see. The survey +said so in §6 and was right to. + +Decision: **search `GroundState`** — legitimate here, because post-game +the deal is public (`solution_discard` already is, and the game is over) — +and move the honesty guarantee to something checkable: + +> **Every move in an emitted witness is marked `visible` or `hidden`.** +> A move is `visible` if, at the point it is played, everything it depends +> on was in the acting seat's projection: the target Problem face-up, the +> Solution in that seat's own hand. Otherwise `hidden`. + +So a witness reads *"you could have won — but two of these six moves +needed a card you had no way to know was coming."* **That is a more useful +answer than either extreme**, and it is computed from `project()`, which +already exists and is already tested. + +**Falsifier:** if a witness is emitted whose moves are all marked +`visible` but which no seat could actually have chosen, the marking is +wrong and D2 has failed. A test constructs exactly that case. + +## D3 — bounded exhaustive over the endgame, `K` rounds, and honest wording + +**Exhaustive search over the last `K` rounds**, with the table treated as +one co-operative agent choosing joint selections. `K = 2` by default, +which the measurements put inside budget at 2–4 seats. + +- The bound is **rounds**, not nodes or seconds, because rounds are what a + player understands: *"winnable from round 4"* means something; *"winnable + within 100,000 nodes"* does not. +- A node budget is a **secondary** cut that aborts with a stated reason, + so a wide table cannot hang the page. +- **Wording is normative.** When no line is found the tool says + **"no winning line found in the last K rounds"** — never *"unwinnable"*. + A bounded search that claims unwinnability is lying, and this is the + sentence the maintainer will read. + +**Not chosen: determinized sampling (PIMC), ISMCTS.** Both are for playing +under uncertainty. Here there is one world (D1), so they would add strategy +fusion to a problem that does not have it. + +## D4 — difficulty is the **winnable fraction**, not any bot's win rate + +**This is the decision the review forced, and it is the useful half of the +pass.** + +A single-policy win rate cannot be a difficulty: two trivial policies span +0–100% on the same deals. Worse, *improving the bot would make the game +"easier"* without a rule changing — a measure that moves when the +measurer improves is not measuring the thing. + +The solver supplies a policy-independent alternative: + +> **Winnable fraction** — over N deals at a seat count, the proportion in +> which the search finds *any* winning line within its bound. + +That is a property of **the deal distribution and the threshold**, which +is exactly what `ground-game` tunes. It is the number GROUND-WP-0005 is +blocked on, and the bot rate never was. + +Difficulty therefore ships as **a small table, never one number**: + +| column | what it is | +|---|---| +| winnable fraction | can the deal be won at all (bounded, K stated) | +| reference-policy win rate | what a stated bot achieves — **named policy** | +| skill gap | the difference: how much play has to supply | + +**Every rate carries its policy, its N, its seed range and its K in the +number's name**, not in a footnote. A figure that loses them is +inadmissible under GameDesign §1.2. + +**Bounded-below caveat, stated because it will be quoted:** the winnable +fraction from a K-round search is a **lower bound** on true winnability — +a deal unwinnable in the last 2 rounds may have been winnable in round 1. +The report says "winnable-from-round-(6−K)", never "winnable". + +## D5 — the harness becomes an instrument before any figure is quoted + +C3 established that `difficulty-baseline.rs` has no assertions, no +`--self-test` and no `make` target — nothing can turn it red. Under +CB-WP-0022 T05's own `role` distinction it is a `default` artifact wearing +a `counterexample` label, and **GameDesign §1.3 makes it inadmissible.** + +Required before T06 reports anything: + +- **positive controls** — a deal constructed to be unwinnable returns + none; a deal constructed to be winnable returns a witness that replays; +- **`--self-test`**, wired into `make self-tests` like every other + reporting tool; +- **`make difficulty`** (or equivalent), so the figure regenerates from + one command; +- the **policy panel is plural**: at least `greedy`, `random` and + `first-legal`, because the spread between them is what C4 exposed and + hiding it would restore the error. + +## D6 — it lives in `games/ground`, not a new crate + +The search needs `validate`, `fold`, `legal_commands` and `project` — +all of `games_ground`. A separate crate would either re-export the +aggregate or take a dependency on it and add nothing. + +**The tier was declared L on the assumption of a new capability port. +There is no port**, and that over-declaration is recorded rather than +hidden — it is a data point for the tier rules, and the L weight paid for +itself twice over regardless (§Consequences). + +`cb-play` gains a mode to ask the question about a finished game; the +difficulty sweep is an example/binary, as the baseline is. + +## D7 — the per-node cost is unsettled and T04 must benchmark it + +The author measured **3.0–4.1 µs**, the reviewer **15.6–20.4 µs**, by +different isolations. Both agree the published 112–161 µs was wrong by +1–2 orders; neither has established which is right. + +**T04 benchmarks it with `criterion`** — already a dev-dependency, already +used by `benches/synthetic.rs` — and the spec quotes that number and no +other. **Neither figure above may be cited**, including by this ADR. + +## Consequences + +- `specs/` gains the witness contract and the difficulty table's shape + (T04), plus the benchmarked node cost. +- T05 builds the K-round search, the `visible`/`hidden` marking, and the + replay check. +- **T06's payload changes completely.** It reports a winnable fraction and + a policy panel to GROUND-WP-0005 — *not* the withdrawn "too easy at 5–6 + seats". The withdrawal itself is reported, per ADR-0012 D5. +- The register gains the withdrawn finding as `inconsistent` / + `withdrawn`, so it is in the log rather than forgotten. + +## What was rejected + +| rejected | why | +|---|---| +| determinized sampling / PIMC / ISMCTS | strategy fusion, added to a question that has one world (D1) | +| a view-only search as a structural boundary | not implementable — a view cannot fold events | +| refusing to answer unless the line was findable at the time | throws away a true and useful answer; annotate instead (D2) | +| a single bot win rate as "difficulty" | two trivial policies span 0–100% on the same deals (C4) | +| "unwinnable" as output wording | a bounded search cannot know it (D3) | +| a new crate | no port exists; it would re-export the aggregate (D6) | +| quoting either measured node cost | they disagree 5× and neither is established (D7) | diff --git a/workplans/CB-WP-0025-could-we-have-won.md b/workplans/CB-WP-0025-could-we-have-won.md index 23b4739..cec037a 100644 --- a/workplans/CB-WP-0025-could-we-have-won.md +++ b/workplans/CB-WP-0025-could-we-have-won.md @@ -235,7 +235,7 @@ this project have now caught a false headline that every gate passed.** ```task id: CB-WP-0025-T03 -status: todo +status: done priority: high state_hub_task_id: "055630a1-56e0-4ad4-8796-35f444313371" ``` @@ -258,6 +258,38 @@ state_hub_task_id: "055630a1-56e0-4ad4-8796-35f444313371" if the ADR concludes it is a mode of an existing one, say so, and the over-declaration is a chaos-window data point worth recording. +**Done 2026-08-05.** +[ADR-0013](../decisions/ADR-0013-could-we-have-won.md), seven decisions. +**Two of them are not what T03 was written expecting**, because the review +moved the ground under both. + +- **D1 — strategy fusion does not apply, and that is why this is + affordable.** Fusion is a defect of *aggregating over determinizations + to choose a move*. **After the game there is one world**: the deal is + known, so a search over it produces a line executable in the only world + there is. The survey treated fusion as this pass's central obstacle; it + is an obstacle to a *playing* engine, which we are not building. +- **D4 — difficulty is the WINNABLE FRACTION, not any bot's win rate.** + C4 killed the bot rate: two trivial policies span 0–100% on the same + deals, and improving the bot would make the game "easier" without a rule + changing. The solver supplies a policy-independent measure — *over N + deals, in what proportion does a winning line exist* — which is a + property of the deal distribution and the threshold, and is what + GROUND-WP-0005 actually needs. **The bot rate never was.** + +**D2** searches `GroundState` (the survey's view-only boundary is not +implementable — a view cannot fold events) and moves the guarantee to a +checkable per-move `visible`/`hidden` marking computed from `project()`. +A witness reads *"you could have won, but two of these six moves needed a +card you had no way to know was coming."* **D3** is bounded exhaustive +over the last K rounds — measured affordable at 2–4 seats once C6 +corrected the premise — with normative wording: *"no winning line found in +the last K rounds"*, never *"unwinnable"*. **D5** makes the harness an +instrument before any figure is quoted (C3). **D6**: no new crate, no +port — **the L declaration was an over-declaration and is recorded as +one**. **D7**: the node cost is disputed 5× and T04 must benchmark it; +neither figure may be cited, including by the ADR. + ## Task: specify ```task