CB-WP-0024/0025: what play reported, split into a renderer and a search
Five remarks from the maintainer's test games, checked against the code
before being written down — two had already reached ground-game on wrong
premises, so a claim now names the line that makes it true.
Three of the five turned out to be data the projection already carries,
drawn as text: solution_deck_len, solution_discard, and OutcomeView's
personal/mastery/winners. One control (`close — I have read this`) is
labelled as a reading but shuts the server down, and leaves a live-looking
page pointing at a dead port. One number does not exist at all: table.rs
loops run_game and keeps only the last summary.
CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the
other seats' plays as objects on the table, the ending control saying what
it does, and a tally that survives "play again".
CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how
hard is this" are the same search asked twice. Tier L because the
information boundary is the whole design problem: a solver reading
GroundState sees the deck the rules hide, and would tell the maintainer he
could have won by playing a card he had no way to know was there. Also
unblocks GROUND-WP-0005, active with both tasks waiting on a measured
difficulty baseline.
Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023
shipped CB-EV-0020 first.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 12:59:08 +02:00
|
|
|
|
---
|
|
|
|
|
|
id: CB-WP-0025
|
|
|
|
|
|
kind: product
|
|
|
|
|
|
title: "Could we have won: a path out of a lost game, and how hard the game actually is"
|
CB-WP-0025 T06/T07: a difficulty baseline that reports its own confound
T06. games/ground/examples/difficulty.rs, make difficulty, wired into make
self-tests, and a report file in ground-game under GROUND-WP-0005 with a
hub message pointing at it.
THE REPORT OPENS WITH THE RETRACTION, because what this task was written
to send was withdrawn by T02 and GROUND-WP-0005 is blocked on exactly that
number. They are told, in the first section, that we nearly sent them "the
game is too easy at 5-6 seats" and why it was wrong.
seats winnable greedy random first-legal spread
2p 60% 60.0% 5.0% 76.7% 71.7
3p 93% 88.3% 6.7% 25.0% 81.7
4p 100% 93.3% 6.7% 30.0% 86.7
5p 100% 100.0% 3.3% 0.0% 100.0
6p 100% 100.0% 3.3% 0.0% 100.0
SPREAD justifies the whole redesign: 71.7 to 100.0 points between three
trivial policies. The table now shows why no single rate is a difficulty
rather than asserting it.
And the 5-6 rows point the OPPOSITE way from the withdrawn claim --
first-legal 0% against greedy 100% is the widest spread in the table,
which suggests play matters MORE there, not less. Neither reading is
established and the report says so.
The confound is stated in the tool's own output, not only in prose:
`winnable` is conditioned on greedy's play up to the final round, because
searching from round 1 is unaffordable. Presenting it as a property of the
deal would repeat this pass's error in a subtler form -- which is exactly
how a corrected project reintroduces a defect. NO THRESHOLD CHANGES ARE
PROPOSED.
The instrument can fail (spec §5): a witness must replay to a win, an
unwinnable position must report searched-out rather than a budget cut, a
one-node budget must not claim exhaustion, and the policy panel must
actually disagree. difficulty-baseline.rs marked superseded, kept as the
survey's dated snapshot. Registered as F16, inconsistent / withdrawn.
T07. evidence/CB-EV-0024. Five of nine defects came only from the review;
four from execution, and all four of those were in work written after it.
The wrong-denominator family now has five instances and still no control
-- facts-check catches copies that disagree, nothing catches a number
computed correctly against the wrong base. Tier L was an over-declaration
(no port, structurally M) and paid for itself anyway, because the review
is L-only. Chaos window 2 will close with zero overrides, making its
retirement condition untestable.
Named as open rather than implied done: the witness is NOT wired to the
ending page. The search works; the browser cannot ask it yet.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 19:24:49 +02:00
|
|
|
|
status: done
|
2026-08-05 13:00:02 +02:00
|
|
|
|
state_hub_workstream_id: "a866982f-94e7-432b-b0d5-2fefb781a574"
|
CB-WP-0024/0025: what play reported, split into a renderer and a search
Five remarks from the maintainer's test games, checked against the code
before being written down — two had already reached ground-game on wrong
premises, so a claim now names the line that makes it true.
Three of the five turned out to be data the projection already carries,
drawn as text: solution_deck_len, solution_discard, and OutcomeView's
personal/mastery/winners. One control (`close — I have read this`) is
labelled as a reading but shuts the server down, and leaves a live-looking
page pointing at a dead port. One number does not exist at all: table.rs
loops run_game and keeps only the last summary.
CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the
other seats' plays as objects on the table, the ending control saying what
it does, and a tally that survives "play again".
CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how
hard is this" are the same search asked twice. Tier L because the
information boundary is the whole design problem: a solver reading
GroundState sees the deck the rules hide, and would tell the maintainer he
could have won by playing a card he had no way to know was there. Also
unblocks GROUND-WP-0005, active with both tasks waiting on a measured
difficulty baseline.
Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023
shipped CB-EV-0020 first.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 12:59:08 +02:00
|
|
|
|
---
|
|
|
|
|
|
|
|
|
|
|
|
# Purpose
|
|
|
|
|
|
|
|
|
|
|
|
```
|
|
|
|
|
|
structural tier L (creates a new capability — a search over game state,
|
|
|
|
|
|
and a measurement the engine does not currently take;
|
|
|
|
|
|
both produce numbers ground-game will tune against)
|
|
|
|
|
|
chaos d8 = 6 → no override
|
|
|
|
|
|
declared tier L
|
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
Declaration 8 of chaos window 2. Tier L: separate survey, **adversarial
|
|
|
|
|
|
review**, ADR, then spec, then code.
|
|
|
|
|
|
|
|
|
|
|
|
## Two remarks, and why they are one pass
|
|
|
|
|
|
|
|
|
|
|
|
> *"I had a game where we lost and in this case I would have liked to know
|
|
|
|
|
|
> if and how we could have won… the best path is not computable I guess so
|
|
|
|
|
|
> a path to win is fine."*
|
|
|
|
|
|
|
|
|
|
|
|
> *"Do we have difficulty estimations? If so we should show them. It will
|
|
|
|
|
|
> help tuning the game. I felt it was too easy but then we lost, so who
|
|
|
|
|
|
> knows."*
|
|
|
|
|
|
|
|
|
|
|
|
They are the same machine asked two questions. *Was this game winnable?*
|
|
|
|
|
|
is a search from a recorded state. *How hard is this game?* is that search
|
|
|
|
|
|
run over many deals and counted. Building the second without the first
|
|
|
|
|
|
gives a win-rate with no witness; building the first without the second
|
|
|
|
|
|
gives one anecdote per game.
|
|
|
|
|
|
|
|
|
|
|
|
**"I felt it was too easy but then we lost, so who knows" is the finding.**
|
|
|
|
|
|
The maintainer cannot calibrate the game from play, and that is precisely
|
|
|
|
|
|
the gap `ground-game` is currently blocked in: **GROUND-WP-0005**
|
|
|
|
|
|
*(Difficulty tiers (threshold ratio) and optional Pressure dial)* is active
|
|
|
|
|
|
with **both its tasks in `wait`**. Tiers cannot be set without a measured
|
|
|
|
|
|
baseline, and clay-borg is the thing that can measure. This pass is what
|
|
|
|
|
|
unblocks it — which is the design-instrument aspect CB-WP-0022 is in the
|
|
|
|
|
|
middle of stating, arriving with a concrete demand.
|
|
|
|
|
|
|
|
|
|
|
|
## What already exists, so the survey does not re-find it
|
|
|
|
|
|
|
CB-WP-0025 T05: the search works, and it falsified this pass's own
affordability projection
games/ground/src/search.rs, five tests. It finds real winning lines and
replays them through validate/fold to group_success.
Two bugs in my own work, found and fixed here.
THE TRAVERSAL WAS WRONG. It branched on "the first seat with any legal
command" and stopped there, so a later seat never acted if an earlier one
was already selected but still had a legal move. Restructured around what
the rules oblige: a seat without a selection MUST select (GR-R02) and
nothing else can happen first; after Reveal the optional actions branch
freely and the aggregate rejects Resolve until the obligatory ones are
done -- so the search needs no phase logic of its own.
AND MY REWIND WAS OFF BY ONE ROUND, replaying the round it was meant to
search. That is why the first run reported 3 nodes and looked like a
working search.
Measured with the real search, rewinding real games to the start of their
last K rounds:
2p K=1 exhausted, 8,103 nodes, ~29 ms
2p K=2 budget cut at 2,000,000 nodes, ~5 s
3p K=2 win found, 41 nodes, ~157 us
The spec's own falsifier said "§3 fails if K=2 proves unaffordable at four
seats". IT FAILED AT TWO. The projection assumed a joint product per
round; the search explores sequential per-seat decisions, so orderings
multiply the tree far beyond width^seats. That is the second projection
this pass published in place of a measurement -- C1's timer was the first.
THE ASYMMETRY IS THE OPERATIVE FINDING. Finding a win is cheap: DFS
stumbles onto one in tens of nodes. Proving none exists needs exhaustion.
So the witness feature is affordable now at any K a player would ask
about, and the winnable fraction (ADR-0013 D4) is NOT, because its
negative half must exhaust every deal it counts. K=1 is the honest default
for exhaustive answers today; making K=2 exhaustible needs transposition
or move-ordering, neither of which this pass built. specs §3 and §3.1
corrected accordingly, and the K=2 default withdrawn.
The negative control that makes "winnable" falsifiable: 2p seed 7 over its
last round returns NoneFound with exhausted=true in ~8k nodes -- a real
negative, not a budget cut wearing a verdict's clothes. And the visible/
hidden marking is tested both ways, since a marking that can only say YES
is decoration.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 19:10:25 +02:00
|
|
|
|
The state is replayable (`replay.rs`, `make replay-test`); the move space
|
|
|
|
|
|
is enumerable (`legal_commands`, narrowed by CB-WP-0023 so branching is
|
|
|
|
|
|
real rather than inflated by inert moves); bots exist (`bot.rs`); and the
|
|
|
|
|
|
threshold is public (`OutcomeView`). **A search needs no new persistence
|
|
|
|
|
|
and no new rules** — which is why D6 put it in `games/ground` with no new
|
|
|
|
|
|
crate and no port.
|
CB-WP-0024/0025: what play reported, split into a renderer and a search
Five remarks from the maintainer's test games, checked against the code
before being written down — two had already reached ground-game on wrong
premises, so a claim now names the line that makes it true.
Three of the five turned out to be data the projection already carries,
drawn as text: solution_deck_len, solution_discard, and OutcomeView's
personal/mastery/winners. One control (`close — I have read this`) is
labelled as a reading but shuts the server down, and leaves a live-looking
page pointing at a dead port. One number does not exist at all: table.rs
loops run_game and keeps only the last summary.
CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the
other seats' plays as objects on the table, the ending control saying what
it does, and a tally that survives "play again".
CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how
hard is this" are the same search asked twice. Tier L because the
information boundary is the whole design problem: a solver reading
GroundState sees the deck the rules hide, and would tell the maintainer he
could have won by playing a card he had no way to know was there. Also
unblocks GROUND-WP-0005, active with both tasks waiting on a measured
difficulty baseline.
Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023
shipped CB-EV-0020 first.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 12:59:08 +02:00
|
|
|
|
|
|
|
|
|
|
## What makes this hard, and must not be waved through
|
|
|
|
|
|
|
CB-WP-0025 T05: the search works, and it falsified this pass's own
affordability projection
games/ground/src/search.rs, five tests. It finds real winning lines and
replays them through validate/fold to group_success.
Two bugs in my own work, found and fixed here.
THE TRAVERSAL WAS WRONG. It branched on "the first seat with any legal
command" and stopped there, so a later seat never acted if an earlier one
was already selected but still had a legal move. Restructured around what
the rules oblige: a seat without a selection MUST select (GR-R02) and
nothing else can happen first; after Reveal the optional actions branch
freely and the aggregate rejects Resolve until the obligatory ones are
done -- so the search needs no phase logic of its own.
AND MY REWIND WAS OFF BY ONE ROUND, replaying the round it was meant to
search. That is why the first run reported 3 nodes and looked like a
working search.
Measured with the real search, rewinding real games to the start of their
last K rounds:
2p K=1 exhausted, 8,103 nodes, ~29 ms
2p K=2 budget cut at 2,000,000 nodes, ~5 s
3p K=2 win found, 41 nodes, ~157 us
The spec's own falsifier said "§3 fails if K=2 proves unaffordable at four
seats". IT FAILED AT TWO. The projection assumed a joint product per
round; the search explores sequential per-seat decisions, so orderings
multiply the tree far beyond width^seats. That is the second projection
this pass published in place of a measurement -- C1's timer was the first.
THE ASYMMETRY IS THE OPERATIVE FINDING. Finding a win is cheap: DFS
stumbles onto one in tens of nodes. Proving none exists needs exhaustion.
So the witness feature is affordable now at any K a player would ask
about, and the winnable fraction (ADR-0013 D4) is NOT, because its
negative half must exhaust every deal it counts. K=1 is the honest default
for exhaustive answers today; making K=2 exhaustible needs transposition
or move-ordering, neither of which this pass built. specs §3 and §3.1
corrected accordingly, and the K=2 default withdrawn.
The negative control that makes "winnable" falsifiable: 2p seed 7 over its
last round returns NoneFound with exhausted=true in ~8k nodes -- a real
negative, not a budget cut wearing a verdict's clothes. And the visible/
hidden marking is tested both ways, since a marking that can only say YES
is decoration.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 19:10:25 +02:00
|
|
|
|
**Settled by [ADR-0013](../decisions/ADR-0013-could-we-have-won.md).** The
|
|
|
|
|
|
declaration framed the central risk as *a path computed with the deck
|
|
|
|
|
|
known is a path the players could never have found*, and asked the ADR to
|
|
|
|
|
|
choose between an omniscient, an information-respecting, and a bounded
|
|
|
|
|
|
search.
|
|
|
|
|
|
|
|
|
|
|
|
**D1 dissolved the choice**: strategy fusion is a defect of aggregating
|
|
|
|
|
|
over determinizations to *choose a move*, and after the game there is one
|
|
|
|
|
|
world — so a line found in it is executable in it. **D2** keeps the
|
|
|
|
|
|
declaration's real concern by marking each move `visible`/`hidden` rather
|
|
|
|
|
|
than by refusing to search.
|
|
|
|
|
|
|
|
|
|
|
|
The second warning here — *"one win rate over one bot policy is not the
|
|
|
|
|
|
difficulty"* — was right, and **the survey made exactly that error
|
|
|
|
|
|
anyway**; see T02.
|
CB-WP-0024/0025: what play reported, split into a renderer and a search
Five remarks from the maintainer's test games, checked against the code
before being written down — two had already reached ground-game on wrong
premises, so a claim now names the line that makes it true.
Three of the five turned out to be data the projection already carries,
drawn as text: solution_deck_len, solution_discard, and OutcomeView's
personal/mastery/winners. One control (`close — I have read this`) is
labelled as a reading but shuts the server down, and leaves a live-looking
page pointing at a dead port. One number does not exist at all: table.rs
loops run_game and keeps only the last summary.
CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the
other seats' plays as objects on the table, the ending control saying what
it does, and a tally that survives "play again".
CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how
hard is this" are the same search asked twice. Tier L because the
information boundary is the whole design problem: a solver reading
GroundState sees the deck the rules hide, and would tell the maintainer he
could have won by playing a card he had no way to know was there. Also
unblocks GROUND-WP-0005, active with both tasks waiting on a measured
difficulty baseline.
Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023
shipped CB-EV-0020 first.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 12:59:08 +02:00
|
|
|
|
|
|
|
|
|
|
## Task: survey
|
|
|
|
|
|
|
|
|
|
|
|
```task
|
|
|
|
|
|
id: CB-WP-0025-T01
|
CB-WP-0025 T01: survey -- the baseline found the game before the solver did
CB-RES-0008 with a runnable baseline
(games/ground/examples/difficulty-baseline.rs), and the measurement
produced a finding before any solver exists.
A GREEDY BOT WINS 200 OF 200 GAMES AT FIVE AND SIX SEATS. Median margin
+3, 11.8-12.0 points available against a threshold of 9. The curve is
66% / 82.5% / 95% / 100% / 100% across 2/3/4/5/6 seats.
Row-level, as GROUND-WP-0004 T02 requires: available points 6 / 9 / 12
against thresholds 5 / 7 / 9, so the ratio RISES with seat count (1.20,
1.29, 1.33) while the table also gains actions per round to clear it with.
Three multipliers pointing the same direction.
It also explains the maintainer's report without needing a solver at all:
"I felt it was too easy but then we lost" is two true statements about
different seat counts.
Cost measured and it rules out the obvious approach. Branching is small
(mean 4.7-9.1) but legal_commands costs 112-161 us per call because it
filters candidates through full validate. Exhaustive search is out at
every seat count; 10^4-10^5 nodes is 1.4-14 seconds, which is the budget
the ADR must design inside.
Prior art names the trap: determinized search (PIMC) suffers strategy
fusion (Frank, Basin & Matsubara 1998) -- the search picks different
actions in states a real player cannot distinguish, so the witness may
require knowing what was on top of the deck. Such a line still replays
green, so the checkability benchmark does not catch it. Honesty and
checkability are different properties; stated explicitly so T03 cannot
conflate them.
The survey states its own most likely killer up front (§6): a view-only
search cannot fold events, so making the information boundary structural
rather than a promise may not be affordable. Better found here than in
T05.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 17:52:46 +02:00
|
|
|
|
status: done
|
CB-WP-0024/0025: what play reported, split into a renderer and a search
Five remarks from the maintainer's test games, checked against the code
before being written down — two had already reached ground-game on wrong
premises, so a claim now names the line that makes it true.
Three of the five turned out to be data the projection already carries,
drawn as text: solution_deck_len, solution_discard, and OutcomeView's
personal/mastery/winners. One control (`close — I have read this`) is
labelled as a reading but shuts the server down, and leaves a live-looking
page pointing at a dead port. One number does not exist at all: table.rs
loops run_game and keeps only the last summary.
CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the
other seats' plays as objects on the table, the ending control saying what
it does, and a tally that survives "play again".
CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how
hard is this" are the same search asked twice. Tier L because the
information boundary is the whole design problem: a solver reading
GroundState sees the deck the rules hide, and would tell the maintainer he
could have won by playing a card he had no way to know was there. Also
unblocks GROUND-WP-0005, active with both tasks waiting on a measured
difficulty baseline.
Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023
shipped CB-EV-0020 first.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 12:59:08 +02:00
|
|
|
|
priority: high
|
2026-08-05 13:00:02 +02:00
|
|
|
|
state_hub_task_id: "556cfd24-5992-4cc0-be90-0b60e989b5bb"
|
CB-WP-0024/0025: what play reported, split into a renderer and a search
Five remarks from the maintainer's test games, checked against the code
before being written down — two had already reached ground-game on wrong
premises, so a claim now names the line that makes it true.
Three of the five turned out to be data the projection already carries,
drawn as text: solution_deck_len, solution_discard, and OutcomeView's
personal/mastery/winners. One control (`close — I have read this`) is
labelled as a reading but shuts the server down, and leaves a live-looking
page pointing at a dead port. One number does not exist at all: table.rs
loops run_game and keeps only the last summary.
CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the
other seats' plays as objects on the table, the ending control saying what
it does, and a tally that survives "play again".
CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how
hard is this" are the same search asked twice. Tier L because the
information boundary is the whole design problem: a solver reading
GroundState sees the deck the rules hide, and would tell the maintainer he
could have won by playing a card he had no way to know was there. Also
unblocks GROUND-WP-0005, active with both tasks waiting on a measured
difficulty baseline.
Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023
shipped CB-EV-0020 first.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 12:59:08 +02:00
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
`research/CB-RES-0008-*.md`, with `tier: L` and the chaos roll recorded
|
|
|
|
|
|
(`loop-lint` checks both).
|
|
|
|
|
|
|
|
|
|
|
|
Per §Step 1 the survey is done when it can name a **benchmark-to-beat**
|
CB-WP-0025 T05: the search works, and it falsified this pass's own
affordability projection
games/ground/src/search.rs, five tests. It finds real winning lines and
replays them through validate/fold to group_success.
Two bugs in my own work, found and fixed here.
THE TRAVERSAL WAS WRONG. It branched on "the first seat with any legal
command" and stopped there, so a later seat never acted if an earlier one
was already selected but still had a legal move. Restructured around what
the rules oblige: a seat without a selection MUST select (GR-R02) and
nothing else can happen first; after Reveal the optional actions branch
freely and the aggregate rejects Resolve until the obligatory ones are
done -- so the search needs no phase logic of its own.
AND MY REWIND WAS OFF BY ONE ROUND, replaying the round it was meant to
search. That is why the first run reported 3 nodes and looked like a
working search.
Measured with the real search, rewinding real games to the start of their
last K rounds:
2p K=1 exhausted, 8,103 nodes, ~29 ms
2p K=2 budget cut at 2,000,000 nodes, ~5 s
3p K=2 win found, 41 nodes, ~157 us
The spec's own falsifier said "§3 fails if K=2 proves unaffordable at four
seats". IT FAILED AT TWO. The projection assumed a joint product per
round; the search explores sequential per-seat decisions, so orderings
multiply the tree far beyond width^seats. That is the second projection
this pass published in place of a measurement -- C1's timer was the first.
THE ASYMMETRY IS THE OPERATIVE FINDING. Finding a win is cheap: DFS
stumbles onto one in tens of nodes. Proving none exists needs exhaustion.
So the witness feature is affordable now at any K a player would ask
about, and the winnable fraction (ADR-0013 D4) is NOT, because its
negative half must exhaust every deal it counts. K=1 is the honest default
for exhaustive answers today; making K=2 exhaustible needs transposition
or move-ordering, neither of which this pass built. specs §3 and §3.1
corrected accordingly, and the K=2 default withdrawn.
The negative control that makes "winnable" falsifiable: 2p seed 7 over its
last round returns NoneFound with exhausted=true in ~8k nodes -- a real
negative, not a budget cut wearing a verdict's clothes. And the visible/
hidden marking is tested both ways, since a marking that can only say YES
is decoration.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 19:10:25 +02:00
|
|
|
|
per dimension. Four were asked for: retrospective solvers in
|
|
|
|
|
|
hidden-information games (and their known failure, strategy fusion); *a
|
|
|
|
|
|
path to win* as a **witness** rather than a proof; difficulty as a measured
|
|
|
|
|
|
quantity in co-operative games; and the **per-node cost**, measured on our
|
|
|
|
|
|
machine — the runnable-baseline option applies, since a search that cannot
|
|
|
|
|
|
finish while the player is looking at the page is a different feature.
|
CB-WP-0024/0025: what play reported, split into a renderer and a search
Five remarks from the maintainer's test games, checked against the code
before being written down — two had already reached ground-game on wrong
premises, so a claim now names the line that makes it true.
Three of the five turned out to be data the projection already carries,
drawn as text: solution_deck_len, solution_discard, and OutcomeView's
personal/mastery/winners. One control (`close — I have read this`) is
labelled as a reading but shuts the server down, and leaves a live-looking
page pointing at a dead port. One number does not exist at all: table.rs
loops run_game and keeps only the last summary.
CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the
other seats' plays as objects on the table, the ending control saying what
it does, and a tally that survives "play again".
CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how
hard is this" are the same search asked twice. Tier L because the
information boundary is the whole design problem: a solver reading
GroundState sees the deck the rules hide, and would tell the maintainer he
could have won by playing a card he had no way to know was there. Also
unblocks GROUND-WP-0005, active with both tasks waiting on a measured
difficulty baseline.
Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023
shipped CB-EV-0020 first.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 12:59:08 +02:00
|
|
|
|
|
CB-WP-0025 T01: survey -- the baseline found the game before the solver did
CB-RES-0008 with a runnable baseline
(games/ground/examples/difficulty-baseline.rs), and the measurement
produced a finding before any solver exists.
A GREEDY BOT WINS 200 OF 200 GAMES AT FIVE AND SIX SEATS. Median margin
+3, 11.8-12.0 points available against a threshold of 9. The curve is
66% / 82.5% / 95% / 100% / 100% across 2/3/4/5/6 seats.
Row-level, as GROUND-WP-0004 T02 requires: available points 6 / 9 / 12
against thresholds 5 / 7 / 9, so the ratio RISES with seat count (1.20,
1.29, 1.33) while the table also gains actions per round to clear it with.
Three multipliers pointing the same direction.
It also explains the maintainer's report without needing a solver at all:
"I felt it was too easy but then we lost" is two true statements about
different seat counts.
Cost measured and it rules out the obvious approach. Branching is small
(mean 4.7-9.1) but legal_commands costs 112-161 us per call because it
filters candidates through full validate. Exhaustive search is out at
every seat count; 10^4-10^5 nodes is 1.4-14 seconds, which is the budget
the ADR must design inside.
Prior art names the trap: determinized search (PIMC) suffers strategy
fusion (Frank, Basin & Matsubara 1998) -- the search picks different
actions in states a real player cannot distinguish, so the witness may
require knowing what was on top of the deck. Such a line still replays
green, so the checkability benchmark does not catch it. Honesty and
checkability are different properties; stated explicitly so T03 cannot
conflate them.
The survey states its own most likely killer up front (§6): a view-only
search cannot fold events, so making the information boundary structural
rather than a promise may not be affordable. Better found here than in
T05.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 17:52:46 +02:00
|
|
|
|
**Done 2026-08-05.**
|
|
|
|
|
|
[CB-RES-0008](../research/CB-RES-0008-could-we-have-won.md), with a
|
|
|
|
|
|
runnable baseline (`games/ground/examples/difficulty-baseline.rs`).
|
|
|
|
|
|
|
CB-WP-0025 T04: specs/RetrospectiveAnalysis.md, and a benchmark that
caught me repeating C1
specs/RetrospectiveAnalysis.md v1.0 plus games/ground/benches/search.rs,
which exists because ADR-0013 D7 refused to let the spec quote either
disputed figure.
THE BENCHMARK'S OWN FIRST FIXTURE WAS DEFECTIVE, and it is the same defect
class the review caught one layer up. Stopping at a fixed step 20 put 2p
and 4p in states where seat 0 had NO legal commands, so it timed an empty
Vec (~120 ns) and silently skipped validate_fold because there was nothing
to validate. It now advances until the seat has a real branch and ASSERTS
it. A clone benchmark was added too: a search must copy state per branch,
and iter_batched excludes setup from timing, so without it the budget
would again rest on an unmeasured span.
Measured at real decision points: legal_commands 4.06-4.76 us, clone
378-639 ns, validate+fold 0.5-3.8 us. Per-child cost is NOT uniform --
some commands resolve cascades -- so budgets use the upper end (~5
us/child).
That settles D3 with real numbers. Joint branching over the last two
rounds is ~5x10^2 / 1.6x10^5 / 5.7x10^5 at 2/3/4 seats, so K=2 costs
negligible / 0.8 s / 2.9 s and holds at two to four seats. It does NOT
hold at five or six, where the tool must reduce K and say that it did
rather than silently searching less.
§4.1 is a normative prohibition, not a preference: a single policy's win
rate MAY NOT be reported as a difficulty. The spec carries the measured
reason -- greedy 100% against first-legal 0% on identical deals -- because
this project already made that error and nearly exported it to a repo that
is blocked waiting on the number.
§2.3 makes the empty-result wording normative: "no winning line found in
the last K rounds", never "unwinnable". A bounded search cannot establish
unwinnability and that sentence is what a player who just lost reads.
Also corrected: the T01 completion record still asserted all three
withdrawn claims as fact. It now carries claimed / withdrawn / survives
explicitly rather than being rewritten -- a retraction that does not
propagate to every place the claim lives is how the earlier ones survived.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 18:59:46 +02:00
|
|
|
|
> **Everything this record originally claimed was withdrawn by T02 the
|
|
|
|
|
|
> same day.** Kept as a pointer rather than rewritten, because a claim
|
|
|
|
|
|
> retracted silently is how three earlier wrong premises survived
|
|
|
|
|
|
> (ADR-0012 D5).
|
|
|
|
|
|
|
|
|
|
|
|
**Claimed:** a greedy bot wins 200/200 at five and six seats, so the game
|
|
|
|
|
|
is too easy there; the points-to-threshold ratio rises with seat count and
|
|
|
|
|
|
explains the curve; `legal_commands` costs 112–161 µs, so exhaustive
|
|
|
|
|
|
search is out; and all of this explains the maintainer's report.
|
|
|
|
|
|
|
|
|
|
|
|
**Withdrawn:** the win rate is `GreedyPolicy`'s, not the game's (C4 —
|
|
|
|
|
|
`FirstLegal` scores 0% on the same deals); the ratio is identical at 3p
|
|
|
|
|
|
and 4p, which differ by 12.5 points (C2); the cost figure was wrong by
|
|
|
|
|
|
30–50× and exhaustive search is affordable (C1, C6); and the maintainer's
|
|
|
|
|
|
losses were 3-player games on the pre-ruling deal (C5).
|
|
|
|
|
|
|
|
|
|
|
|
**What survives:** the harness exists and runs, the 6/9/12 arithmetic is
|
|
|
|
|
|
right against `Problems.csv`, and the branching widths (4.7–9.1) hold. See
|
|
|
|
|
|
[CB-RES-0008](../research/CB-RES-0008-could-we-have-won.md) for the
|
|
|
|
|
|
corrected text and
|
|
|
|
|
|
[the response](../history/260805-could-we-have-won-response.md) for the
|
|
|
|
|
|
full accounting.
|
|
|
|
|
|
|
CB-WP-0024/0025: what play reported, split into a renderer and a search
Five remarks from the maintainer's test games, checked against the code
before being written down — two had already reached ground-game on wrong
premises, so a claim now names the line that makes it true.
Three of the five turned out to be data the projection already carries,
drawn as text: solution_deck_len, solution_discard, and OutcomeView's
personal/mastery/winners. One control (`close — I have read this`) is
labelled as a reading but shuts the server down, and leaves a live-looking
page pointing at a dead port. One number does not exist at all: table.rs
loops run_game and keeps only the last summary.
CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the
other seats' plays as objects on the table, the ending control saying what
it does, and a tally that survives "play again".
CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how
hard is this" are the same search asked twice. Tier L because the
information boundary is the whole design problem: a solver reading
GroundState sees the deck the rules hide, and would tell the maintainer he
could have won by playing a card he had no way to know was there. Also
unblocks GROUND-WP-0005, active with both tasks waiting on a measured
difficulty baseline.
Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023
shipped CB-EV-0020 first.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 12:59:08 +02:00
|
|
|
|
## Task: adversarial review
|
|
|
|
|
|
|
|
|
|
|
|
```task
|
|
|
|
|
|
id: CB-WP-0025-T02
|
CB-WP-0025 T02: the review withdrew the finding, and the fifth wrong
premise never left the repo
Separate agent, second tier-L review in this project. Six of seven
challenges conceded. The survey's headline finding is WITHDRAWN, not
softened.
C4 kills it, and the reviewer ranked it fourth. A FirstLegal policy --
take legal[0], no heuristic at all -- scores 0% at five and six seats
where GreedyPolicy scores 100%, and 77.5% at two seats where greedy scores
66%. Two unsophisticated agents span the entire range at the same seat
count. "The game is too easy at 5-6 seats" is therefore a statement about
GreedyPolicy, not about GROUND. The rescue the reviewer offered -- greedy
hits the 12-point ceiling in 200/200 deals, so the 6p row is a rules claim
-- dies on the same data: FirstLegal reaches that ceiling never.
C1: the per-node cost was wrong by 30-50x. The timer started before the
seed loop, so "us/node" included two setups, an entire greedy game and a
full validate+fold replay, divided by player-decision count. The tell was
in my own published output and I did not look at it: the figure FELL
(161/139/112) as branching ROSE (4.7/7.4/9.1), which no per-enumeration
cost can do. Re-measured with the clock around legal_commands alone:
3.0/3.5/4.1 us, now rising with branching. The reviewer measured 15.6-20.4
by a different isolation; we disagree by ~5x and neither has established
which is right, so T04 must benchmark it with criterion rather than adopt
either number.
C6: "exhaustive search is out at any seat count" is false -- ~3 seconds
over the last two rounds at 3p. With C1's correction the budget is
~10^5-10^6 nodes and bounded endgame search fits, so ADR-0013 cannot open
with "exhaustive is impossible, therefore determinized sampling" --
especially as sampling carries strategy fusion that exhaustive search does
not.
C3: the finding failed the admissibility rule this project wrote nine
hours earlier. 6/9/12 are sums where GROUND-WP-0004 T02 requires
per-priority rows, and the harness has no assertions, no --self-test and
no make target, so nothing can turn it red -- a `default` artifact wearing
a `counterexample` label, by CB-WP-0022 T05's own distinction.
C2: the ratio story explains nothing; 3p and 4p share deal, threshold and
ratio and differ by 12.5 points of win rate. C5: "explains the
maintainer's report" is contradicted by lib.rs:2487, which records his
losses as 3-player games on the pre-ruling deal, arithmetically unwinnable
at 6 against 7.
T06 exists to report to GROUND-WP-0005, which is BLOCKED waiting on a
difficulty baseline. Had this proceeded they would have been invited to
move thresholds on the strength of one bot's behaviour. That is the fifth
wrong premise this project would have sent them, and the second stopped by
an adversarial review rather than by a control. Both tier-L reviews here
have now caught a false headline that every gate passed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 18:06:03 +02:00
|
|
|
|
status: done
|
CB-WP-0024/0025: what play reported, split into a renderer and a search
Five remarks from the maintainer's test games, checked against the code
before being written down — two had already reached ground-game on wrong
premises, so a claim now names the line that makes it true.
Three of the five turned out to be data the projection already carries,
drawn as text: solution_deck_len, solution_discard, and OutcomeView's
personal/mastery/winners. One control (`close — I have read this`) is
labelled as a reading but shuts the server down, and leaves a live-looking
page pointing at a dead port. One number does not exist at all: table.rs
loops run_game and keeps only the last summary.
CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the
other seats' plays as objects on the table, the ending control saying what
it does, and a tally that survives "play again".
CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how
hard is this" are the same search asked twice. Tier L because the
information boundary is the whole design problem: a solver reading
GroundState sees the deck the rules hide, and would tell the maintainer he
could have won by playing a card he had no way to know was there. Also
unblocks GROUND-WP-0005, active with both tasks waiting on a measured
difficulty baseline.
Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023
shipped CB-EV-0020 first.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 12:59:08 +02:00
|
|
|
|
priority: high
|
2026-08-05 13:00:02 +02:00
|
|
|
|
state_hub_task_id: "bdd9fc97-4706-4d8e-b286-91ff53339680"
|
CB-WP-0024/0025: what play reported, split into a renderer and a search
Five remarks from the maintainer's test games, checked against the code
before being written down — two had already reached ground-game on wrong
premises, so a claim now names the line that makes it true.
Three of the five turned out to be data the projection already carries,
drawn as text: solution_deck_len, solution_discard, and OutcomeView's
personal/mastery/winners. One control (`close — I have read this`) is
labelled as a reading but shuts the server down, and leaves a live-looking
page pointing at a dead port. One number does not exist at all: table.rs
loops run_game and keeps only the last summary.
CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the
other seats' plays as objects on the table, the ending control saying what
it does, and a tally that survives "play again".
CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how
hard is this" are the same search asked twice. Tier L because the
information boundary is the whole design problem: a solver reading
GroundState sees the deck the rules hide, and would tell the maintainer he
could have won by playing a card he had no way to know was there. Also
unblocks GROUND-WP-0005, active with both tasks waiting on a measured
difficulty baseline.
Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023
shipped CB-EV-0020 first.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 12:59:08 +02:00
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
Tier L requires it. Exactly one round: challenge, then response, trail in
|
CB-WP-0025 T06/T07: a difficulty baseline that reports its own confound
T06. games/ground/examples/difficulty.rs, make difficulty, wired into make
self-tests, and a report file in ground-game under GROUND-WP-0005 with a
hub message pointing at it.
THE REPORT OPENS WITH THE RETRACTION, because what this task was written
to send was withdrawn by T02 and GROUND-WP-0005 is blocked on exactly that
number. They are told, in the first section, that we nearly sent them "the
game is too easy at 5-6 seats" and why it was wrong.
seats winnable greedy random first-legal spread
2p 60% 60.0% 5.0% 76.7% 71.7
3p 93% 88.3% 6.7% 25.0% 81.7
4p 100% 93.3% 6.7% 30.0% 86.7
5p 100% 100.0% 3.3% 0.0% 100.0
6p 100% 100.0% 3.3% 0.0% 100.0
SPREAD justifies the whole redesign: 71.7 to 100.0 points between three
trivial policies. The table now shows why no single rate is a difficulty
rather than asserting it.
And the 5-6 rows point the OPPOSITE way from the withdrawn claim --
first-legal 0% against greedy 100% is the widest spread in the table,
which suggests play matters MORE there, not less. Neither reading is
established and the report says so.
The confound is stated in the tool's own output, not only in prose:
`winnable` is conditioned on greedy's play up to the final round, because
searching from round 1 is unaffordable. Presenting it as a property of the
deal would repeat this pass's error in a subtler form -- which is exactly
how a corrected project reintroduces a defect. NO THRESHOLD CHANGES ARE
PROPOSED.
The instrument can fail (spec §5): a witness must replay to a win, an
unwinnable position must report searched-out rather than a budget cut, a
one-node budget must not claim exhaustion, and the policy panel must
actually disagree. difficulty-baseline.rs marked superseded, kept as the
survey's dated snapshot. Registered as F16, inconsistent / withdrawn.
T07. evidence/CB-EV-0024. Five of nine defects came only from the review;
four from execution, and all four of those were in work written after it.
The wrong-denominator family now has five instances and still no control
-- facts-check catches copies that disagree, nothing catches a number
computed correctly against the wrong base. Tier L was an over-declaration
(no port, structurally M) and paid for itself anyway, because the review
is L-only. Chaos window 2 will close with zero overrides, making its
retirement condition untestable.
Named as open rather than implied done: the witness is NOT wired to the
ending page. The search works; the browser cannot ask it yet.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 19:24:49 +02:00
|
|
|
|
`history/`, unpolished. Four attempts required: that the honest version is
|
|
|
|
|
|
unaffordable; that a witness misleads more than it helps; that the
|
|
|
|
|
|
difficulty number is a bot benchmark in a difficulty costume; and that
|
|
|
|
|
|
this is CB-WP-0022's job rather than a separate capability.
|
CB-WP-0024/0025: what play reported, split into a renderer and a search
Five remarks from the maintainer's test games, checked against the code
before being written down — two had already reached ground-game on wrong
premises, so a claim now names the line that makes it true.
Three of the five turned out to be data the projection already carries,
drawn as text: solution_deck_len, solution_discard, and OutcomeView's
personal/mastery/winners. One control (`close — I have read this`) is
labelled as a reading but shuts the server down, and leaves a live-looking
page pointing at a dead port. One number does not exist at all: table.rs
loops run_game and keeps only the last summary.
CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the
other seats' plays as objects on the table, the ending control saying what
it does, and a tally that survives "play again".
CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how
hard is this" are the same search asked twice. Tier L because the
information boundary is the whole design problem: a solver reading
GroundState sees the deck the rules hide, and would tell the maintainer he
could have won by playing a card he had no way to know was there. Also
unblocks GROUND-WP-0005, active with both tasks waiting on a measured
difficulty baseline.
Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023
shipped CB-EV-0020 first.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 12:59:08 +02:00
|
|
|
|
|
CB-WP-0025 T02: the review withdrew the finding, and the fifth wrong
premise never left the repo
Separate agent, second tier-L review in this project. Six of seven
challenges conceded. The survey's headline finding is WITHDRAWN, not
softened.
C4 kills it, and the reviewer ranked it fourth. A FirstLegal policy --
take legal[0], no heuristic at all -- scores 0% at five and six seats
where GreedyPolicy scores 100%, and 77.5% at two seats where greedy scores
66%. Two unsophisticated agents span the entire range at the same seat
count. "The game is too easy at 5-6 seats" is therefore a statement about
GreedyPolicy, not about GROUND. The rescue the reviewer offered -- greedy
hits the 12-point ceiling in 200/200 deals, so the 6p row is a rules claim
-- dies on the same data: FirstLegal reaches that ceiling never.
C1: the per-node cost was wrong by 30-50x. The timer started before the
seed loop, so "us/node" included two setups, an entire greedy game and a
full validate+fold replay, divided by player-decision count. The tell was
in my own published output and I did not look at it: the figure FELL
(161/139/112) as branching ROSE (4.7/7.4/9.1), which no per-enumeration
cost can do. Re-measured with the clock around legal_commands alone:
3.0/3.5/4.1 us, now rising with branching. The reviewer measured 15.6-20.4
by a different isolation; we disagree by ~5x and neither has established
which is right, so T04 must benchmark it with criterion rather than adopt
either number.
C6: "exhaustive search is out at any seat count" is false -- ~3 seconds
over the last two rounds at 3p. With C1's correction the budget is
~10^5-10^6 nodes and bounded endgame search fits, so ADR-0013 cannot open
with "exhaustive is impossible, therefore determinized sampling" --
especially as sampling carries strategy fusion that exhaustive search does
not.
C3: the finding failed the admissibility rule this project wrote nine
hours earlier. 6/9/12 are sums where GROUND-WP-0004 T02 requires
per-priority rows, and the harness has no assertions, no --self-test and
no make target, so nothing can turn it red -- a `default` artifact wearing
a `counterexample` label, by CB-WP-0022 T05's own distinction.
C2: the ratio story explains nothing; 3p and 4p share deal, threshold and
ratio and differ by 12.5 points of win rate. C5: "explains the
maintainer's report" is contradicted by lib.rs:2487, which records his
losses as 3-player games on the pre-ruling deal, arithmetically unwinnable
at 6 against 7.
T06 exists to report to GROUND-WP-0005, which is BLOCKED waiting on a
difficulty baseline. Had this proceeded they would have been invited to
move thresholds on the strength of one bot's behaviour. That is the fifth
wrong premise this project would have sent them, and the second stopped by
an adversarial review rather than by a control. Both tier-L reviews here
have now caught a false headline that every gate passed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 18:06:03 +02:00
|
|
|
|
**Done 2026-08-05.** Trail:
|
|
|
|
|
|
[challenge](../history/260805-could-we-have-won-challenge.md),
|
|
|
|
|
|
[response](../history/260805-could-we-have-won-response.md).
|
|
|
|
|
|
Separate agent, as in CB-WP-0022.
|
|
|
|
|
|
|
CB-WP-0025 T04: specs/RetrospectiveAnalysis.md, and a benchmark that
caught me repeating C1
specs/RetrospectiveAnalysis.md v1.0 plus games/ground/benches/search.rs,
which exists because ADR-0013 D7 refused to let the spec quote either
disputed figure.
THE BENCHMARK'S OWN FIRST FIXTURE WAS DEFECTIVE, and it is the same defect
class the review caught one layer up. Stopping at a fixed step 20 put 2p
and 4p in states where seat 0 had NO legal commands, so it timed an empty
Vec (~120 ns) and silently skipped validate_fold because there was nothing
to validate. It now advances until the seat has a real branch and ASSERTS
it. A clone benchmark was added too: a search must copy state per branch,
and iter_batched excludes setup from timing, so without it the budget
would again rest on an unmeasured span.
Measured at real decision points: legal_commands 4.06-4.76 us, clone
378-639 ns, validate+fold 0.5-3.8 us. Per-child cost is NOT uniform --
some commands resolve cascades -- so budgets use the upper end (~5
us/child).
That settles D3 with real numbers. Joint branching over the last two
rounds is ~5x10^2 / 1.6x10^5 / 5.7x10^5 at 2/3/4 seats, so K=2 costs
negligible / 0.8 s / 2.9 s and holds at two to four seats. It does NOT
hold at five or six, where the tool must reduce K and say that it did
rather than silently searching less.
§4.1 is a normative prohibition, not a preference: a single policy's win
rate MAY NOT be reported as a difficulty. The spec carries the measured
reason -- greedy 100% against first-legal 0% on identical deals -- because
this project already made that error and nearly exported it to a repo that
is blocked waiting on the number.
§2.3 makes the empty-result wording normative: "no winning line found in
the last K rounds", never "unwinnable". A bounded search cannot establish
unwinnability and that sentence is what a player who just lost reads.
Also corrected: the T01 completion record still asserted all three
withdrawn claims as fact. It now carries claimed / withdrawn / survives
explicitly rather than being rewritten -- a retraction that does not
propagate to every place the claim lives is how the earlier ones survived.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 18:59:46 +02:00
|
|
|
|
**Six of seven conceded. The survey's headline finding is WITHDRAWN.**
|
|
|
|
|
|
|
|
|
|
|
|
**C4 kills it and the reviewer ranked it fourth.** A `FirstLegal` policy —
|
|
|
|
|
|
`legal[0]`, no heuristic — scores **0% at five and six seats** where greedy
|
|
|
|
|
|
scores 100%, and **77.5% at two** where greedy scores 66%. Two
|
|
|
|
|
|
unsophisticated agents span the whole range, so *"too easy at 5–6 seats"*
|
|
|
|
|
|
was about `GreedyPolicy`, not GROUND.
|
|
|
|
|
|
|
|
|
|
|
|
**C1**: the node cost was wrong by 30–50× — the timer bracketed whole
|
|
|
|
|
|
games — and **the tell was in my own output**, falling as branching rose.
|
|
|
|
|
|
**C6**: exhaustive search is *not* out, which changes T03's premise.
|
|
|
|
|
|
**C3**: the finding failed the admissibility rule this project wrote nine
|
|
|
|
|
|
hours earlier. **C2**: 3p and 4p share deal, threshold and ratio and differ
|
|
|
|
|
|
by 12.5 points. **C5**: `lib.rs:2487` contradicts the "explains the
|
|
|
|
|
|
maintainer" claim — his losses were 3-player on the pre-ruling deal,
|
|
|
|
|
|
unwinnable at 6 against 7.
|
CB-WP-0025 T02: the review withdrew the finding, and the fifth wrong
premise never left the repo
Separate agent, second tier-L review in this project. Six of seven
challenges conceded. The survey's headline finding is WITHDRAWN, not
softened.
C4 kills it, and the reviewer ranked it fourth. A FirstLegal policy --
take legal[0], no heuristic at all -- scores 0% at five and six seats
where GreedyPolicy scores 100%, and 77.5% at two seats where greedy scores
66%. Two unsophisticated agents span the entire range at the same seat
count. "The game is too easy at 5-6 seats" is therefore a statement about
GreedyPolicy, not about GROUND. The rescue the reviewer offered -- greedy
hits the 12-point ceiling in 200/200 deals, so the 6p row is a rules claim
-- dies on the same data: FirstLegal reaches that ceiling never.
C1: the per-node cost was wrong by 30-50x. The timer started before the
seed loop, so "us/node" included two setups, an entire greedy game and a
full validate+fold replay, divided by player-decision count. The tell was
in my own published output and I did not look at it: the figure FELL
(161/139/112) as branching ROSE (4.7/7.4/9.1), which no per-enumeration
cost can do. Re-measured with the clock around legal_commands alone:
3.0/3.5/4.1 us, now rising with branching. The reviewer measured 15.6-20.4
by a different isolation; we disagree by ~5x and neither has established
which is right, so T04 must benchmark it with criterion rather than adopt
either number.
C6: "exhaustive search is out at any seat count" is false -- ~3 seconds
over the last two rounds at 3p. With C1's correction the budget is
~10^5-10^6 nodes and bounded endgame search fits, so ADR-0013 cannot open
with "exhaustive is impossible, therefore determinized sampling" --
especially as sampling carries strategy fusion that exhaustive search does
not.
C3: the finding failed the admissibility rule this project wrote nine
hours earlier. 6/9/12 are sums where GROUND-WP-0004 T02 requires
per-priority rows, and the harness has no assertions, no --self-test and
no make target, so nothing can turn it red -- a `default` artifact wearing
a `counterexample` label, by CB-WP-0022 T05's own distinction.
C2: the ratio story explains nothing; 3p and 4p share deal, threshold and
ratio and differ by 12.5 points of win rate. C5: "explains the
maintainer's report" is contradicted by lib.rs:2487, which records his
losses as 3-player games on the pre-ruling deal, arithmetically unwinnable
at 6 against 7.
T06 exists to report to GROUND-WP-0005, which is BLOCKED waiting on a
difficulty baseline. Had this proceeded they would have been invited to
move thresholds on the strength of one bot's behaviour. That is the fifth
wrong premise this project would have sent them, and the second stopped by
an adversarial review rather than by a control. Both tier-L reviews here
have now caught a false headline that every gate passed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 18:06:03 +02:00
|
|
|
|
|
|
|
|
|
|
**T06 was pointed at GROUND-WP-0005, which is blocked on this number.**
|
CB-WP-0025 T04: specs/RetrospectiveAnalysis.md, and a benchmark that
caught me repeating C1
specs/RetrospectiveAnalysis.md v1.0 plus games/ground/benches/search.rs,
which exists because ADR-0013 D7 refused to let the spec quote either
disputed figure.
THE BENCHMARK'S OWN FIRST FIXTURE WAS DEFECTIVE, and it is the same defect
class the review caught one layer up. Stopping at a fixed step 20 put 2p
and 4p in states where seat 0 had NO legal commands, so it timed an empty
Vec (~120 ns) and silently skipped validate_fold because there was nothing
to validate. It now advances until the seat has a real branch and ASSERTS
it. A clone benchmark was added too: a search must copy state per branch,
and iter_batched excludes setup from timing, so without it the budget
would again rest on an unmeasured span.
Measured at real decision points: legal_commands 4.06-4.76 us, clone
378-639 ns, validate+fold 0.5-3.8 us. Per-child cost is NOT uniform --
some commands resolve cascades -- so budgets use the upper end (~5
us/child).
That settles D3 with real numbers. Joint branching over the last two
rounds is ~5x10^2 / 1.6x10^5 / 5.7x10^5 at 2/3/4 seats, so K=2 costs
negligible / 0.8 s / 2.9 s and holds at two to four seats. It does NOT
hold at five or six, where the tool must reduce K and say that it did
rather than silently searching less.
§4.1 is a normative prohibition, not a preference: a single policy's win
rate MAY NOT be reported as a difficulty. The spec carries the measured
reason -- greedy 100% against first-legal 0% on identical deals -- because
this project already made that error and nearly exported it to a repo that
is blocked waiting on the number.
§2.3 makes the empty-result wording normative: "no winning line found in
the last K rounds", never "unwinnable". A bounded search cannot establish
unwinnability and that sentence is what a player who just lost reads.
Also corrected: the T01 completion record still asserted all three
withdrawn claims as fact. It now carries claimed / withdrawn / survives
explicitly rather than being rewritten -- a retraction that does not
propagate to every place the claim lives is how the earlier ones survived.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 18:59:46 +02:00
|
|
|
|
Sending it would have been the fifth wrong premise to reach ground-game.
|
|
|
|
|
|
**Both tier-L reviews here have now caught a false headline that every
|
|
|
|
|
|
gate passed.**
|
CB-WP-0025 T02: the review withdrew the finding, and the fifth wrong
premise never left the repo
Separate agent, second tier-L review in this project. Six of seven
challenges conceded. The survey's headline finding is WITHDRAWN, not
softened.
C4 kills it, and the reviewer ranked it fourth. A FirstLegal policy --
take legal[0], no heuristic at all -- scores 0% at five and six seats
where GreedyPolicy scores 100%, and 77.5% at two seats where greedy scores
66%. Two unsophisticated agents span the entire range at the same seat
count. "The game is too easy at 5-6 seats" is therefore a statement about
GreedyPolicy, not about GROUND. The rescue the reviewer offered -- greedy
hits the 12-point ceiling in 200/200 deals, so the 6p row is a rules claim
-- dies on the same data: FirstLegal reaches that ceiling never.
C1: the per-node cost was wrong by 30-50x. The timer started before the
seed loop, so "us/node" included two setups, an entire greedy game and a
full validate+fold replay, divided by player-decision count. The tell was
in my own published output and I did not look at it: the figure FELL
(161/139/112) as branching ROSE (4.7/7.4/9.1), which no per-enumeration
cost can do. Re-measured with the clock around legal_commands alone:
3.0/3.5/4.1 us, now rising with branching. The reviewer measured 15.6-20.4
by a different isolation; we disagree by ~5x and neither has established
which is right, so T04 must benchmark it with criterion rather than adopt
either number.
C6: "exhaustive search is out at any seat count" is false -- ~3 seconds
over the last two rounds at 3p. With C1's correction the budget is
~10^5-10^6 nodes and bounded endgame search fits, so ADR-0013 cannot open
with "exhaustive is impossible, therefore determinized sampling" --
especially as sampling carries strategy fusion that exhaustive search does
not.
C3: the finding failed the admissibility rule this project wrote nine
hours earlier. 6/9/12 are sums where GROUND-WP-0004 T02 requires
per-priority rows, and the harness has no assertions, no --self-test and
no make target, so nothing can turn it red -- a `default` artifact wearing
a `counterexample` label, by CB-WP-0022 T05's own distinction.
C2: the ratio story explains nothing; 3p and 4p share deal, threshold and
ratio and differ by 12.5 points of win rate. C5: "explains the
maintainer's report" is contradicted by lib.rs:2487, which records his
losses as 3-player games on the pre-ruling deal, arithmetically unwinnable
at 6 against 7.
T06 exists to report to GROUND-WP-0005, which is BLOCKED waiting on a
difficulty baseline. Had this proceeded they would have been invited to
move thresholds on the strength of one bot's behaviour. That is the fifth
wrong premise this project would have sent them, and the second stopped by
an adversarial review rather than by a control. Both tier-L reviews here
have now caught a false headline that every gate passed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 18:06:03 +02:00
|
|
|
|
|
CB-WP-0024/0025: what play reported, split into a renderer and a search
Five remarks from the maintainer's test games, checked against the code
before being written down — two had already reached ground-game on wrong
premises, so a claim now names the line that makes it true.
Three of the five turned out to be data the projection already carries,
drawn as text: solution_deck_len, solution_discard, and OutcomeView's
personal/mastery/winners. One control (`close — I have read this`) is
labelled as a reading but shuts the server down, and leaves a live-looking
page pointing at a dead port. One number does not exist at all: table.rs
loops run_game and keeps only the last summary.
CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the
other seats' plays as objects on the table, the ending control saying what
it does, and a tally that survives "play again".
CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how
hard is this" are the same search asked twice. Tier L because the
information boundary is the whole design problem: a solver reading
GroundState sees the deck the rules hide, and would tell the maintainer he
could have won by playing a card he had no way to know was there. Also
unblocks GROUND-WP-0005, active with both tasks waiting on a measured
difficulty baseline.
Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023
shipped CB-EV-0020 first.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 12:59:08 +02:00
|
|
|
|
## Task: decide
|
|
|
|
|
|
|
|
|
|
|
|
```task
|
|
|
|
|
|
id: CB-WP-0025-T03
|
CB-WP-0025 T03: ADR-0013 -- one world after the game, and difficulty that
does not measure the bot
Seven decisions. Two are not what T03 expected, because the review moved
the ground under both.
D1: strategy fusion DOES NOT APPLY, and that is why the affordable option
is also the honest one. Fusion is a defect of aggregating over
determinizations to choose a move -- the search picking different actions
in states a player cannot distinguish. After the game there is ONE WORLD:
the deal is known, so a search over it yields a line executable in the
only world there is. The survey treated fusion as this pass's central
obstacle; it is an obstacle to a playing engine, which we are not
building. The tool answers "given the deal as it actually was, was there a
line that reached the threshold" and is labelled that way on screen --
never "how you should have played".
D4: difficulty is the WINNABLE FRACTION, not a bot's win rate. C4 killed
the bot rate -- two trivial policies span 0-100% on the same deals, and
improving the bot would make the game "easier" without a rule changing. A
measure that moves when the measurer improves is not measuring the thing.
The solver supplies the alternative: over N deals, in what proportion does
a winning line exist. That is a property of the deal distribution and the
threshold, which is what ground-game tunes and what GROUND-WP-0005 is
blocked on. Ships as a table -- winnable fraction, named reference policy,
skill gap -- never one number, with policy/N/seed-range/K in the number's
name. Stated as a LOWER BOUND, since a K-round search cannot see a line
that needed round 1.
D2: search GroundState. The survey's view-only structural boundary is not
implementable -- a view cannot fold events, and it said so in §6. The
guarantee moves to something checkable: every move in a witness is marked
visible or hidden, computed from project(). A witness reads "you could
have won, but two of these six moves needed a card you had no way to know
was coming" -- more useful than either extreme.
D3: bounded exhaustive over the last K rounds, table as one co-operative
agent, K=2 default. Affordable once C6 corrected the premise: joint
branching over the last two rounds is ~5x10^2 / 1.6x10^5 / 5.7x10^5 at
2/3/4 seats. Wording is normative: "no winning line found in the last K
rounds", never "unwinnable". PIMC and ISMCTS rejected -- they add strategy
fusion to a problem that does not have it.
D5: the harness becomes an instrument before any figure is quoted (C3) --
positive controls, --self-test in make self-tests, a make target, and a
PLURAL policy panel, because the spread between policies is what C4
exposed and hiding it would restore the error.
D6: no new crate and no port. The L declaration was an over-declaration
and is recorded as one. D7: the node cost is disputed 5x between author
and reviewer; T04 benchmarks it with criterion and neither existing figure
may be cited, including by this ADR.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 18:47:07 +02:00
|
|
|
|
status: done
|
CB-WP-0024/0025: what play reported, split into a renderer and a search
Five remarks from the maintainer's test games, checked against the code
before being written down — two had already reached ground-game on wrong
premises, so a claim now names the line that makes it true.
Three of the five turned out to be data the projection already carries,
drawn as text: solution_deck_len, solution_discard, and OutcomeView's
personal/mastery/winners. One control (`close — I have read this`) is
labelled as a reading but shuts the server down, and leaves a live-looking
page pointing at a dead port. One number does not exist at all: table.rs
loops run_game and keeps only the last summary.
CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the
other seats' plays as objects on the table, the ending control saying what
it does, and a tally that survives "play again".
CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how
hard is this" are the same search asked twice. Tier L because the
information boundary is the whole design problem: a solver reading
GroundState sees the deck the rules hide, and would tell the maintainer he
could have won by playing a card he had no way to know was there. Also
unblocks GROUND-WP-0005, active with both tasks waiting on a measured
difficulty baseline.
Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023
shipped CB-EV-0020 first.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 12:59:08 +02:00
|
|
|
|
priority: high
|
2026-08-05 13:00:02 +02:00
|
|
|
|
state_hub_task_id: "055630a1-56e0-4ad4-8796-35f444313371"
|
CB-WP-0024/0025: what play reported, split into a renderer and a search
Five remarks from the maintainer's test games, checked against the code
before being written down — two had already reached ground-game on wrong
premises, so a claim now names the line that makes it true.
Three of the five turned out to be data the projection already carries,
drawn as text: solution_deck_len, solution_discard, and OutcomeView's
personal/mastery/winners. One control (`close — I have read this`) is
labelled as a reading but shuts the server down, and leaves a live-looking
page pointing at a dead port. One number does not exist at all: table.rs
loops run_game and keeps only the last summary.
CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the
other seats' plays as objects on the table, the ending control saying what
it does, and a tally that survives "play again".
CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how
hard is this" are the same search asked twice. Tier L because the
information boundary is the whole design problem: a solver reading
GroundState sees the deck the rules hide, and would tell the maintainer he
could have won by playing a card he had no way to know was there. Also
unblocks GROUND-WP-0005, active with both tasks waiting on a measured
difficulty baseline.
Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023
shipped CB-EV-0020 first.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 12:59:08 +02:00
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
`decisions/ADR-0013-*.md`. (ADR-0012 is CB-WP-0022 T03's.) At minimum:
|
CB-WP-0025 T06/T07: a difficulty baseline that reports its own confound
T06. games/ground/examples/difficulty.rs, make difficulty, wired into make
self-tests, and a report file in ground-game under GROUND-WP-0005 with a
hub message pointing at it.
THE REPORT OPENS WITH THE RETRACTION, because what this task was written
to send was withdrawn by T02 and GROUND-WP-0005 is blocked on exactly that
number. They are told, in the first section, that we nearly sent them "the
game is too easy at 5-6 seats" and why it was wrong.
seats winnable greedy random first-legal spread
2p 60% 60.0% 5.0% 76.7% 71.7
3p 93% 88.3% 6.7% 25.0% 81.7
4p 100% 93.3% 6.7% 30.0% 86.7
5p 100% 100.0% 3.3% 0.0% 100.0
6p 100% 100.0% 3.3% 0.0% 100.0
SPREAD justifies the whole redesign: 71.7 to 100.0 points between three
trivial policies. The table now shows why no single rate is a difficulty
rather than asserting it.
And the 5-6 rows point the OPPOSITE way from the withdrawn claim --
first-legal 0% against greedy 100% is the widest spread in the table,
which suggests play matters MORE there, not less. Neither reading is
established and the report says so.
The confound is stated in the tool's own output, not only in prose:
`winnable` is conditioned on greedy's play up to the final round, because
searching from round 1 is unaffordable. Presenting it as a property of the
deal would repeat this pass's error in a subtler form -- which is exactly
how a corrected project reintroduces a defect. NO THRESHOLD CHANGES ARE
PROPOSED.
The instrument can fail (spec §5): a witness must replay to a win, an
unwinnable position must report searched-out rather than a budget cut, a
one-node budget must not claim exhaustion, and the policy panel must
actually disagree. difficulty-baseline.rs marked superseded, kept as the
survey's dated snapshot. Registered as F16, inconsistent / withdrawn.
T07. evidence/CB-EV-0024. Five of nine defects came only from the review;
four from execution, and all four of those were in work written after it.
The wrong-denominator family now has five instances and still no control
-- facts-check catches copies that disagree, nothing catches a number
computed correctly against the wrong base. Tier L was an over-declaration
(no port, structurally M) and paid for itself anyway, because the review
is L-only. Chaos window 2 will close with zero overrides, making its
retirement condition untestable.
Named as open rather than implied done: the witness is NOT wired to the
ending page. The search works; the browser cannot ask it yet.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 19:24:49 +02:00
|
|
|
|
which question the solver answers and what it is called in the UI; the
|
|
|
|
|
|
**information boundary** and whether it is structural or promised; the
|
|
|
|
|
|
**bound**, and what *no path found* means against it; whether difficulty
|
|
|
|
|
|
is one number or a table, and of what; and **where it lives** — the tier
|
|
|
|
|
|
was declared L assuming a new capability port, and if there is none, say
|
|
|
|
|
|
so.
|
CB-WP-0024/0025: what play reported, split into a renderer and a search
Five remarks from the maintainer's test games, checked against the code
before being written down — two had already reached ground-game on wrong
premises, so a claim now names the line that makes it true.
Three of the five turned out to be data the projection already carries,
drawn as text: solution_deck_len, solution_discard, and OutcomeView's
personal/mastery/winners. One control (`close — I have read this`) is
labelled as a reading but shuts the server down, and leaves a live-looking
page pointing at a dead port. One number does not exist at all: table.rs
loops run_game and keeps only the last summary.
CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the
other seats' plays as objects on the table, the ending control saying what
it does, and a tally that survives "play again".
CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how
hard is this" are the same search asked twice. Tier L because the
information boundary is the whole design problem: a solver reading
GroundState sees the deck the rules hide, and would tell the maintainer he
could have won by playing a card he had no way to know was there. Also
unblocks GROUND-WP-0005, active with both tasks waiting on a measured
difficulty baseline.
Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023
shipped CB-EV-0020 first.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 12:59:08 +02:00
|
|
|
|
|
CB-WP-0025 T03: ADR-0013 -- one world after the game, and difficulty that
does not measure the bot
Seven decisions. Two are not what T03 expected, because the review moved
the ground under both.
D1: strategy fusion DOES NOT APPLY, and that is why the affordable option
is also the honest one. Fusion is a defect of aggregating over
determinizations to choose a move -- the search picking different actions
in states a player cannot distinguish. After the game there is ONE WORLD:
the deal is known, so a search over it yields a line executable in the
only world there is. The survey treated fusion as this pass's central
obstacle; it is an obstacle to a playing engine, which we are not
building. The tool answers "given the deal as it actually was, was there a
line that reached the threshold" and is labelled that way on screen --
never "how you should have played".
D4: difficulty is the WINNABLE FRACTION, not a bot's win rate. C4 killed
the bot rate -- two trivial policies span 0-100% on the same deals, and
improving the bot would make the game "easier" without a rule changing. A
measure that moves when the measurer improves is not measuring the thing.
The solver supplies the alternative: over N deals, in what proportion does
a winning line exist. That is a property of the deal distribution and the
threshold, which is what ground-game tunes and what GROUND-WP-0005 is
blocked on. Ships as a table -- winnable fraction, named reference policy,
skill gap -- never one number, with policy/N/seed-range/K in the number's
name. Stated as a LOWER BOUND, since a K-round search cannot see a line
that needed round 1.
D2: search GroundState. The survey's view-only structural boundary is not
implementable -- a view cannot fold events, and it said so in §6. The
guarantee moves to something checkable: every move in a witness is marked
visible or hidden, computed from project(). A witness reads "you could
have won, but two of these six moves needed a card you had no way to know
was coming" -- more useful than either extreme.
D3: bounded exhaustive over the last K rounds, table as one co-operative
agent, K=2 default. Affordable once C6 corrected the premise: joint
branching over the last two rounds is ~5x10^2 / 1.6x10^5 / 5.7x10^5 at
2/3/4 seats. Wording is normative: "no winning line found in the last K
rounds", never "unwinnable". PIMC and ISMCTS rejected -- they add strategy
fusion to a problem that does not have it.
D5: the harness becomes an instrument before any figure is quoted (C3) --
positive controls, --self-test in make self-tests, a make target, and a
PLURAL policy panel, because the spread between policies is what C4
exposed and hiding it would restore the error.
D6: no new crate and no port. The L declaration was an over-declaration
and is recorded as one. D7: the node cost is disputed 5x between author
and reviewer; T04 benchmarks it with criterion and neither existing figure
may be cited, including by this ADR.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 18:47:07 +02:00
|
|
|
|
**Done 2026-08-05.**
|
|
|
|
|
|
[ADR-0013](../decisions/ADR-0013-could-we-have-won.md), seven decisions.
|
|
|
|
|
|
**Two of them are not what T03 was written expecting**, because the review
|
|
|
|
|
|
moved the ground under both.
|
|
|
|
|
|
|
|
|
|
|
|
- **D1 — strategy fusion does not apply, and that is why this is
|
|
|
|
|
|
affordable.** Fusion is a defect of *aggregating over determinizations
|
|
|
|
|
|
to choose a move*. **After the game there is one world**: the deal is
|
|
|
|
|
|
known, so a search over it produces a line executable in the only world
|
|
|
|
|
|
there is. The survey treated fusion as this pass's central obstacle; it
|
|
|
|
|
|
is an obstacle to a *playing* engine, which we are not building.
|
|
|
|
|
|
- **D4 — difficulty is the WINNABLE FRACTION, not any bot's win rate.**
|
|
|
|
|
|
C4 killed the bot rate: two trivial policies span 0–100% on the same
|
|
|
|
|
|
deals, and improving the bot would make the game "easier" without a rule
|
|
|
|
|
|
changing. The solver supplies a policy-independent measure — *over N
|
|
|
|
|
|
deals, in what proportion does a winning line exist* — which is a
|
|
|
|
|
|
property of the deal distribution and the threshold, and is what
|
|
|
|
|
|
GROUND-WP-0005 actually needs. **The bot rate never was.**
|
|
|
|
|
|
|
CB-WP-0025 T04: specs/RetrospectiveAnalysis.md, and a benchmark that
caught me repeating C1
specs/RetrospectiveAnalysis.md v1.0 plus games/ground/benches/search.rs,
which exists because ADR-0013 D7 refused to let the spec quote either
disputed figure.
THE BENCHMARK'S OWN FIRST FIXTURE WAS DEFECTIVE, and it is the same defect
class the review caught one layer up. Stopping at a fixed step 20 put 2p
and 4p in states where seat 0 had NO legal commands, so it timed an empty
Vec (~120 ns) and silently skipped validate_fold because there was nothing
to validate. It now advances until the seat has a real branch and ASSERTS
it. A clone benchmark was added too: a search must copy state per branch,
and iter_batched excludes setup from timing, so without it the budget
would again rest on an unmeasured span.
Measured at real decision points: legal_commands 4.06-4.76 us, clone
378-639 ns, validate+fold 0.5-3.8 us. Per-child cost is NOT uniform --
some commands resolve cascades -- so budgets use the upper end (~5
us/child).
That settles D3 with real numbers. Joint branching over the last two
rounds is ~5x10^2 / 1.6x10^5 / 5.7x10^5 at 2/3/4 seats, so K=2 costs
negligible / 0.8 s / 2.9 s and holds at two to four seats. It does NOT
hold at five or six, where the tool must reduce K and say that it did
rather than silently searching less.
§4.1 is a normative prohibition, not a preference: a single policy's win
rate MAY NOT be reported as a difficulty. The spec carries the measured
reason -- greedy 100% against first-legal 0% on identical deals -- because
this project already made that error and nearly exported it to a repo that
is blocked waiting on the number.
§2.3 makes the empty-result wording normative: "no winning line found in
the last K rounds", never "unwinnable". A bounded search cannot establish
unwinnability and that sentence is what a player who just lost reads.
Also corrected: the T01 completion record still asserted all three
withdrawn claims as fact. It now carries claimed / withdrawn / survives
explicitly rather than being rewritten -- a retraction that does not
propagate to every place the claim lives is how the earlier ones survived.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 18:59:46 +02:00
|
|
|
|
**D2** searches `GroundState` — the survey's view-only boundary is not
|
|
|
|
|
|
implementable, since a view cannot fold events — and moves the guarantee to
|
|
|
|
|
|
a checkable per-move `visible`/`hidden` marking from `project()`. **D3**
|
|
|
|
|
|
bounded exhaustive over the last K rounds, with *"no winning line found in
|
|
|
|
|
|
the last K rounds"* normative, never *"unwinnable"*. **D5** makes the
|
|
|
|
|
|
harness an instrument before any figure is quoted (C3). **D6** no new
|
|
|
|
|
|
crate and no port — **the L declaration was an over-declaration, recorded
|
|
|
|
|
|
as one**. **D7** the node cost is disputed 5×; T04 benchmarks it and
|
|
|
|
|
|
neither existing figure may be cited.
|
CB-WP-0025 T03: ADR-0013 -- one world after the game, and difficulty that
does not measure the bot
Seven decisions. Two are not what T03 expected, because the review moved
the ground under both.
D1: strategy fusion DOES NOT APPLY, and that is why the affordable option
is also the honest one. Fusion is a defect of aggregating over
determinizations to choose a move -- the search picking different actions
in states a player cannot distinguish. After the game there is ONE WORLD:
the deal is known, so a search over it yields a line executable in the
only world there is. The survey treated fusion as this pass's central
obstacle; it is an obstacle to a playing engine, which we are not
building. The tool answers "given the deal as it actually was, was there a
line that reached the threshold" and is labelled that way on screen --
never "how you should have played".
D4: difficulty is the WINNABLE FRACTION, not a bot's win rate. C4 killed
the bot rate -- two trivial policies span 0-100% on the same deals, and
improving the bot would make the game "easier" without a rule changing. A
measure that moves when the measurer improves is not measuring the thing.
The solver supplies the alternative: over N deals, in what proportion does
a winning line exist. That is a property of the deal distribution and the
threshold, which is what ground-game tunes and what GROUND-WP-0005 is
blocked on. Ships as a table -- winnable fraction, named reference policy,
skill gap -- never one number, with policy/N/seed-range/K in the number's
name. Stated as a LOWER BOUND, since a K-round search cannot see a line
that needed round 1.
D2: search GroundState. The survey's view-only structural boundary is not
implementable -- a view cannot fold events, and it said so in §6. The
guarantee moves to something checkable: every move in a witness is marked
visible or hidden, computed from project(). A witness reads "you could
have won, but two of these six moves needed a card you had no way to know
was coming" -- more useful than either extreme.
D3: bounded exhaustive over the last K rounds, table as one co-operative
agent, K=2 default. Affordable once C6 corrected the premise: joint
branching over the last two rounds is ~5x10^2 / 1.6x10^5 / 5.7x10^5 at
2/3/4 seats. Wording is normative: "no winning line found in the last K
rounds", never "unwinnable". PIMC and ISMCTS rejected -- they add strategy
fusion to a problem that does not have it.
D5: the harness becomes an instrument before any figure is quoted (C3) --
positive controls, --self-test in make self-tests, a make target, and a
PLURAL policy panel, because the spread between policies is what C4
exposed and hiding it would restore the error.
D6: no new crate and no port. The L declaration was an over-declaration
and is recorded as one. D7: the node cost is disputed 5x between author
and reviewer; T04 benchmarks it with criterion and neither existing figure
may be cited, including by this ADR.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 18:47:07 +02:00
|
|
|
|
|
CB-WP-0024/0025: what play reported, split into a renderer and a search
Five remarks from the maintainer's test games, checked against the code
before being written down — two had already reached ground-game on wrong
premises, so a claim now names the line that makes it true.
Three of the five turned out to be data the projection already carries,
drawn as text: solution_deck_len, solution_discard, and OutcomeView's
personal/mastery/winners. One control (`close — I have read this`) is
labelled as a reading but shuts the server down, and leaves a live-looking
page pointing at a dead port. One number does not exist at all: table.rs
loops run_game and keeps only the last summary.
CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the
other seats' plays as objects on the table, the ending control saying what
it does, and a tally that survives "play again".
CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how
hard is this" are the same search asked twice. Tier L because the
information boundary is the whole design problem: a solver reading
GroundState sees the deck the rules hide, and would tell the maintainer he
could have won by playing a card he had no way to know was there. Also
unblocks GROUND-WP-0005, active with both tasks waiting on a measured
difficulty baseline.
Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023
shipped CB-EV-0020 first.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 12:59:08 +02:00
|
|
|
|
## Task: specify
|
|
|
|
|
|
|
|
|
|
|
|
```task
|
|
|
|
|
|
id: CB-WP-0025-T04
|
CB-WP-0025 T04: specs/RetrospectiveAnalysis.md, and a benchmark that
caught me repeating C1
specs/RetrospectiveAnalysis.md v1.0 plus games/ground/benches/search.rs,
which exists because ADR-0013 D7 refused to let the spec quote either
disputed figure.
THE BENCHMARK'S OWN FIRST FIXTURE WAS DEFECTIVE, and it is the same defect
class the review caught one layer up. Stopping at a fixed step 20 put 2p
and 4p in states where seat 0 had NO legal commands, so it timed an empty
Vec (~120 ns) and silently skipped validate_fold because there was nothing
to validate. It now advances until the seat has a real branch and ASSERTS
it. A clone benchmark was added too: a search must copy state per branch,
and iter_batched excludes setup from timing, so without it the budget
would again rest on an unmeasured span.
Measured at real decision points: legal_commands 4.06-4.76 us, clone
378-639 ns, validate+fold 0.5-3.8 us. Per-child cost is NOT uniform --
some commands resolve cascades -- so budgets use the upper end (~5
us/child).
That settles D3 with real numbers. Joint branching over the last two
rounds is ~5x10^2 / 1.6x10^5 / 5.7x10^5 at 2/3/4 seats, so K=2 costs
negligible / 0.8 s / 2.9 s and holds at two to four seats. It does NOT
hold at five or six, where the tool must reduce K and say that it did
rather than silently searching less.
§4.1 is a normative prohibition, not a preference: a single policy's win
rate MAY NOT be reported as a difficulty. The spec carries the measured
reason -- greedy 100% against first-legal 0% on identical deals -- because
this project already made that error and nearly exported it to a repo that
is blocked waiting on the number.
§2.3 makes the empty-result wording normative: "no winning line found in
the last K rounds", never "unwinnable". A bounded search cannot establish
unwinnability and that sentence is what a player who just lost reads.
Also corrected: the T01 completion record still asserted all three
withdrawn claims as fact. It now carries claimed / withdrawn / survives
explicitly rather than being rewritten -- a retraction that does not
propagate to every place the claim lives is how the earlier ones survived.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 18:59:46 +02:00
|
|
|
|
status: done
|
CB-WP-0024/0025: what play reported, split into a renderer and a search
Five remarks from the maintainer's test games, checked against the code
before being written down — two had already reached ground-game on wrong
premises, so a claim now names the line that makes it true.
Three of the five turned out to be data the projection already carries,
drawn as text: solution_deck_len, solution_discard, and OutcomeView's
personal/mastery/winners. One control (`close — I have read this`) is
labelled as a reading but shuts the server down, and leaves a live-looking
page pointing at a dead port. One number does not exist at all: table.rs
loops run_game and keeps only the last summary.
CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the
other seats' plays as objects on the table, the ending control saying what
it does, and a tally that survives "play again".
CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how
hard is this" are the same search asked twice. Tier L because the
information boundary is the whole design problem: a solver reading
GroundState sees the deck the rules hide, and would tell the maintainer he
could have won by playing a card he had no way to know was there. Also
unblocks GROUND-WP-0005, active with both tasks waiting on a measured
difficulty baseline.
Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023
shipped CB-EV-0020 first.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 12:59:08 +02:00
|
|
|
|
priority: high
|
2026-08-05 13:00:02 +02:00
|
|
|
|
state_hub_task_id: "7c3ca25d-0001-44ef-a07d-0ab2d51b09a5"
|
CB-WP-0024/0025: what play reported, split into a renderer and a search
Five remarks from the maintainer's test games, checked against the code
before being written down — two had already reached ground-game on wrong
premises, so a claim now names the line that makes it true.
Three of the five turned out to be data the projection already carries,
drawn as text: solution_deck_len, solution_discard, and OutcomeView's
personal/mastery/winners. One control (`close — I have read this`) is
labelled as a reading but shuts the server down, and leaves a live-looking
page pointing at a dead port. One number does not exist at all: table.rs
loops run_game and keeps only the last summary.
CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the
other seats' plays as objects on the table, the ending control saying what
it does, and a tally that survives "play again".
CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how
hard is this" are the same search asked twice. Tier L because the
information boundary is the whole design problem: a solver reading
GroundState sees the deck the rules hide, and would tell the maintainer he
could have won by playing a card he had no way to know was there. Also
unblocks GROUND-WP-0005, active with both tasks waiting on a measured
difficulty baseline.
Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023
shipped CB-EV-0020 first.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 12:59:08 +02:00
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
`specs/` — extend `MetricsAndScenarios.md` or add a capability spec as the
|
|
|
|
|
|
ADR directs, with metrics, because a spec without them is prose.
|
|
|
|
|
|
|
CB-WP-0025 T06/T07: a difficulty baseline that reports its own confound
T06. games/ground/examples/difficulty.rs, make difficulty, wired into make
self-tests, and a report file in ground-game under GROUND-WP-0005 with a
hub message pointing at it.
THE REPORT OPENS WITH THE RETRACTION, because what this task was written
to send was withdrawn by T02 and GROUND-WP-0005 is blocked on exactly that
number. They are told, in the first section, that we nearly sent them "the
game is too easy at 5-6 seats" and why it was wrong.
seats winnable greedy random first-legal spread
2p 60% 60.0% 5.0% 76.7% 71.7
3p 93% 88.3% 6.7% 25.0% 81.7
4p 100% 93.3% 6.7% 30.0% 86.7
5p 100% 100.0% 3.3% 0.0% 100.0
6p 100% 100.0% 3.3% 0.0% 100.0
SPREAD justifies the whole redesign: 71.7 to 100.0 points between three
trivial policies. The table now shows why no single rate is a difficulty
rather than asserting it.
And the 5-6 rows point the OPPOSITE way from the withdrawn claim --
first-legal 0% against greedy 100% is the widest spread in the table,
which suggests play matters MORE there, not less. Neither reading is
established and the report says so.
The confound is stated in the tool's own output, not only in prose:
`winnable` is conditioned on greedy's play up to the final round, because
searching from round 1 is unaffordable. Presenting it as a property of the
deal would repeat this pass's error in a subtler form -- which is exactly
how a corrected project reintroduces a defect. NO THRESHOLD CHANGES ARE
PROPOSED.
The instrument can fail (spec §5): a witness must replay to a win, an
unwinnable position must report searched-out rather than a budget cut, a
one-node budget must not claim exhaustion, and the policy panel must
actually disagree. difficulty-baseline.rs marked superseded, kept as the
survey's dated snapshot. Registered as F16, inconsistent / withdrawn.
T07. evidence/CB-EV-0024. Five of nine defects came only from the review;
four from execution, and all four of those were in work written after it.
The wrong-denominator family now has five instances and still no control
-- facts-check catches copies that disagree, nothing catches a number
computed correctly against the wrong base. Tier L was an over-declaration
(no port, structurally M) and paid for itself anyway, because the review
is L-only. Chaos window 2 will close with zero overrides, making its
retirement condition untestable.
Named as open rather than implied done: the witness is NOT wired to the
ending page. The search works; the browser cannot ask it yet.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 19:24:49 +02:00
|
|
|
|
Candidates, to be argued not adopted: **witness checkability** (every
|
|
|
|
|
|
emitted path replays to `group_success` — a hard gate, not a metric, since
|
|
|
|
|
|
a path that does not replay says the opposite of the truth); **search
|
|
|
|
|
|
cost** at the chosen bound; and **difficulty resolution** — the smallest
|
|
|
|
|
|
threshold difference the measurement can distinguish, with its N, which is
|
|
|
|
|
|
more useful to `ground-game` than any rate with no error bar.
|
CB-WP-0024/0025: what play reported, split into a renderer and a search
Five remarks from the maintainer's test games, checked against the code
before being written down — two had already reached ground-game on wrong
premises, so a claim now names the line that makes it true.
Three of the five turned out to be data the projection already carries,
drawn as text: solution_deck_len, solution_discard, and OutcomeView's
personal/mastery/winners. One control (`close — I have read this`) is
labelled as a reading but shuts the server down, and leaves a live-looking
page pointing at a dead port. One number does not exist at all: table.rs
loops run_game and keeps only the last summary.
CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the
other seats' plays as objects on the table, the ending control saying what
it does, and a tally that survives "play again".
CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how
hard is this" are the same search asked twice. Tier L because the
information boundary is the whole design problem: a solver reading
GroundState sees the deck the rules hide, and would tell the maintainer he
could have won by playing a card he had no way to know was there. Also
unblocks GROUND-WP-0005, active with both tasks waiting on a measured
difficulty baseline.
Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023
shipped CB-EV-0020 first.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 12:59:08 +02:00
|
|
|
|
|
|
|
|
|
|
Per `ground-game`'s ruling (GROUND-WP-0004 T02), **any arithmetic finding
|
|
|
|
|
|
this produces ships a runnable reproduction and a row-level table** — never
|
|
|
|
|
|
a summed figure. A difficulty number is arithmetic, and it is exactly the
|
|
|
|
|
|
kind that has already gone wrong twice.
|
|
|
|
|
|
|
CB-WP-0025 T04: specs/RetrospectiveAnalysis.md, and a benchmark that
caught me repeating C1
specs/RetrospectiveAnalysis.md v1.0 plus games/ground/benches/search.rs,
which exists because ADR-0013 D7 refused to let the spec quote either
disputed figure.
THE BENCHMARK'S OWN FIRST FIXTURE WAS DEFECTIVE, and it is the same defect
class the review caught one layer up. Stopping at a fixed step 20 put 2p
and 4p in states where seat 0 had NO legal commands, so it timed an empty
Vec (~120 ns) and silently skipped validate_fold because there was nothing
to validate. It now advances until the seat has a real branch and ASSERTS
it. A clone benchmark was added too: a search must copy state per branch,
and iter_batched excludes setup from timing, so without it the budget
would again rest on an unmeasured span.
Measured at real decision points: legal_commands 4.06-4.76 us, clone
378-639 ns, validate+fold 0.5-3.8 us. Per-child cost is NOT uniform --
some commands resolve cascades -- so budgets use the upper end (~5
us/child).
That settles D3 with real numbers. Joint branching over the last two
rounds is ~5x10^2 / 1.6x10^5 / 5.7x10^5 at 2/3/4 seats, so K=2 costs
negligible / 0.8 s / 2.9 s and holds at two to four seats. It does NOT
hold at five or six, where the tool must reduce K and say that it did
rather than silently searching less.
§4.1 is a normative prohibition, not a preference: a single policy's win
rate MAY NOT be reported as a difficulty. The spec carries the measured
reason -- greedy 100% against first-legal 0% on identical deals -- because
this project already made that error and nearly exported it to a repo that
is blocked waiting on the number.
§2.3 makes the empty-result wording normative: "no winning line found in
the last K rounds", never "unwinnable". A bounded search cannot establish
unwinnability and that sentence is what a player who just lost reads.
Also corrected: the T01 completion record still asserted all three
withdrawn claims as fact. It now carries claimed / withdrawn / survives
explicitly rather than being rewritten -- a retraction that does not
propagate to every place the claim lives is how the earlier ones survived.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 18:59:46 +02:00
|
|
|
|
**Done 2026-08-05.**
|
|
|
|
|
|
[specs/RetrospectiveAnalysis.md](../specs/RetrospectiveAnalysis.md) v1.0,
|
|
|
|
|
|
and `games/ground/benches/search.rs` for D7's disputed number.
|
|
|
|
|
|
|
|
|
|
|
|
**The benchmark's own first fixture was defective — the same defect class,
|
|
|
|
|
|
one layer down.** A fixed stop at step 20 put 2p and 4p where seat 0 had
|
|
|
|
|
|
**no legal commands**, so it timed an empty `Vec` (~120 ns) and silently
|
|
|
|
|
|
skipped `validate_fold`. It now advances to a real branch **and asserts
|
|
|
|
|
|
it**.
|
|
|
|
|
|
|
|
|
|
|
|
Measured at real decision points (table in
|
|
|
|
|
|
[RetrospectiveAnalysis §3.1](../specs/RetrospectiveAnalysis.md)):
|
|
|
|
|
|
`legal_commands` 4.06–4.76 µs, `clone` 378–639 ns, `validate+fold`
|
|
|
|
|
|
**0.5–3.8 µs**.
|
|
|
|
|
|
|
|
|
|
|
|
**`clone` is in there because a search must copy state per branch**, and
|
CB-WP-0025 T05: the search works, and it falsified this pass's own
affordability projection
games/ground/src/search.rs, five tests. It finds real winning lines and
replays them through validate/fold to group_success.
Two bugs in my own work, found and fixed here.
THE TRAVERSAL WAS WRONG. It branched on "the first seat with any legal
command" and stopped there, so a later seat never acted if an earlier one
was already selected but still had a legal move. Restructured around what
the rules oblige: a seat without a selection MUST select (GR-R02) and
nothing else can happen first; after Reveal the optional actions branch
freely and the aggregate rejects Resolve until the obligatory ones are
done -- so the search needs no phase logic of its own.
AND MY REWIND WAS OFF BY ONE ROUND, replaying the round it was meant to
search. That is why the first run reported 3 nodes and looked like a
working search.
Measured with the real search, rewinding real games to the start of their
last K rounds:
2p K=1 exhausted, 8,103 nodes, ~29 ms
2p K=2 budget cut at 2,000,000 nodes, ~5 s
3p K=2 win found, 41 nodes, ~157 us
The spec's own falsifier said "§3 fails if K=2 proves unaffordable at four
seats". IT FAILED AT TWO. The projection assumed a joint product per
round; the search explores sequential per-seat decisions, so orderings
multiply the tree far beyond width^seats. That is the second projection
this pass published in place of a measurement -- C1's timer was the first.
THE ASYMMETRY IS THE OPERATIVE FINDING. Finding a win is cheap: DFS
stumbles onto one in tens of nodes. Proving none exists needs exhaustion.
So the witness feature is affordable now at any K a player would ask
about, and the winnable fraction (ADR-0013 D4) is NOT, because its
negative half must exhaust every deal it counts. K=1 is the honest default
for exhaustive answers today; making K=2 exhaustible needs transposition
or move-ordering, neither of which this pass built. specs §3 and §3.1
corrected accordingly, and the K=2 default withdrawn.
The negative control that makes "winnable" falsifiable: 2p seed 7 over its
last round returns NoneFound with exhausted=true in ~8k nodes -- a real
negative, not a budget cut wearing a verdict's clothes. And the visible/
hidden marking is tested both ways, since a marking that can only say YES
is decoration.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 19:10:25 +02:00
|
|
|
|
`iter_batched` excludes setup from timing — leaving the budget on an
|
|
|
|
|
|
unmeasured span, which is precisely C1's mistake.
|
CB-WP-0025 T04: specs/RetrospectiveAnalysis.md, and a benchmark that
caught me repeating C1
specs/RetrospectiveAnalysis.md v1.0 plus games/ground/benches/search.rs,
which exists because ADR-0013 D7 refused to let the spec quote either
disputed figure.
THE BENCHMARK'S OWN FIRST FIXTURE WAS DEFECTIVE, and it is the same defect
class the review caught one layer up. Stopping at a fixed step 20 put 2p
and 4p in states where seat 0 had NO legal commands, so it timed an empty
Vec (~120 ns) and silently skipped validate_fold because there was nothing
to validate. It now advances until the seat has a real branch and ASSERTS
it. A clone benchmark was added too: a search must copy state per branch,
and iter_batched excludes setup from timing, so without it the budget
would again rest on an unmeasured span.
Measured at real decision points: legal_commands 4.06-4.76 us, clone
378-639 ns, validate+fold 0.5-3.8 us. Per-child cost is NOT uniform --
some commands resolve cascades -- so budgets use the upper end (~5
us/child).
That settles D3 with real numbers. Joint branching over the last two
rounds is ~5x10^2 / 1.6x10^5 / 5.7x10^5 at 2/3/4 seats, so K=2 costs
negligible / 0.8 s / 2.9 s and holds at two to four seats. It does NOT
hold at five or six, where the tool must reduce K and say that it did
rather than silently searching less.
§4.1 is a normative prohibition, not a preference: a single policy's win
rate MAY NOT be reported as a difficulty. The spec carries the measured
reason -- greedy 100% against first-legal 0% on identical deals -- because
this project already made that error and nearly exported it to a repo that
is blocked waiting on the number.
§2.3 makes the empty-result wording normative: "no winning line found in
the last K rounds", never "unwinnable". A bounded search cannot establish
unwinnability and that sentence is what a player who just lost reads.
Also corrected: the T01 completion record still asserted all three
withdrawn claims as fact. It now carries claimed / withdrawn / survives
explicitly rather than being rewritten -- a retraction that does not
propagate to every place the claim lives is how the earlier ones survived.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 18:59:46 +02:00
|
|
|
|
|
CB-WP-0025 T05: the search works, and it falsified this pass's own
affordability projection
games/ground/src/search.rs, five tests. It finds real winning lines and
replays them through validate/fold to group_success.
Two bugs in my own work, found and fixed here.
THE TRAVERSAL WAS WRONG. It branched on "the first seat with any legal
command" and stopped there, so a later seat never acted if an earlier one
was already selected but still had a legal move. Restructured around what
the rules oblige: a seat without a selection MUST select (GR-R02) and
nothing else can happen first; after Reveal the optional actions branch
freely and the aggregate rejects Resolve until the obligatory ones are
done -- so the search needs no phase logic of its own.
AND MY REWIND WAS OFF BY ONE ROUND, replaying the round it was meant to
search. That is why the first run reported 3 nodes and looked like a
working search.
Measured with the real search, rewinding real games to the start of their
last K rounds:
2p K=1 exhausted, 8,103 nodes, ~29 ms
2p K=2 budget cut at 2,000,000 nodes, ~5 s
3p K=2 win found, 41 nodes, ~157 us
The spec's own falsifier said "§3 fails if K=2 proves unaffordable at four
seats". IT FAILED AT TWO. The projection assumed a joint product per
round; the search explores sequential per-seat decisions, so orderings
multiply the tree far beyond width^seats. That is the second projection
this pass published in place of a measurement -- C1's timer was the first.
THE ASYMMETRY IS THE OPERATIVE FINDING. Finding a win is cheap: DFS
stumbles onto one in tens of nodes. Proving none exists needs exhaustion.
So the witness feature is affordable now at any K a player would ask
about, and the winnable fraction (ADR-0013 D4) is NOT, because its
negative half must exhaust every deal it counts. K=1 is the honest default
for exhaustive answers today; making K=2 exhaustible needs transposition
or move-ordering, neither of which this pass built. specs §3 and §3.1
corrected accordingly, and the K=2 default withdrawn.
The negative control that makes "winnable" falsifiable: 2p seed 7 over its
last round returns NoneFound with exhausted=true in ~8k nodes -- a real
negative, not a budget cut wearing a verdict's clothes. And the visible/
hidden marking is tested both ways, since a marking that can only say YES
is decoration.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 19:10:25 +02:00
|
|
|
|
**The affordability conclusion drawn here was itself falsified by T05.**
|
|
|
|
|
|
It projected joint branching and concluded `K = 2` holds at two to four
|
|
|
|
|
|
seats. The real search exceeded 2×10⁶ nodes at **two**. See T05's record
|
|
|
|
|
|
and [RetrospectiveAnalysis §3.1](../specs/RetrospectiveAnalysis.md) —
|
|
|
|
|
|
a projection from branch widths is not a timing of a search, and this
|
|
|
|
|
|
pass has now made that mistake twice.
|
CB-WP-0025 T04: specs/RetrospectiveAnalysis.md, and a benchmark that
caught me repeating C1
specs/RetrospectiveAnalysis.md v1.0 plus games/ground/benches/search.rs,
which exists because ADR-0013 D7 refused to let the spec quote either
disputed figure.
THE BENCHMARK'S OWN FIRST FIXTURE WAS DEFECTIVE, and it is the same defect
class the review caught one layer up. Stopping at a fixed step 20 put 2p
and 4p in states where seat 0 had NO legal commands, so it timed an empty
Vec (~120 ns) and silently skipped validate_fold because there was nothing
to validate. It now advances until the seat has a real branch and ASSERTS
it. A clone benchmark was added too: a search must copy state per branch,
and iter_batched excludes setup from timing, so without it the budget
would again rest on an unmeasured span.
Measured at real decision points: legal_commands 4.06-4.76 us, clone
378-639 ns, validate+fold 0.5-3.8 us. Per-child cost is NOT uniform --
some commands resolve cascades -- so budgets use the upper end (~5
us/child).
That settles D3 with real numbers. Joint branching over the last two
rounds is ~5x10^2 / 1.6x10^5 / 5.7x10^5 at 2/3/4 seats, so K=2 costs
negligible / 0.8 s / 2.9 s and holds at two to four seats. It does NOT
hold at five or six, where the tool must reduce K and say that it did
rather than silently searching less.
§4.1 is a normative prohibition, not a preference: a single policy's win
rate MAY NOT be reported as a difficulty. The spec carries the measured
reason -- greedy 100% against first-legal 0% on identical deals -- because
this project already made that error and nearly exported it to a repo that
is blocked waiting on the number.
§2.3 makes the empty-result wording normative: "no winning line found in
the last K rounds", never "unwinnable". A bounded search cannot establish
unwinnability and that sentence is what a player who just lost reads.
Also corrected: the T01 completion record still asserted all three
withdrawn claims as fact. It now carries claimed / withdrawn / survives
explicitly rather than being rewritten -- a retraction that does not
propagate to every place the claim lives is how the earlier ones survived.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 18:59:46 +02:00
|
|
|
|
|
|
|
|
|
|
**§4.1 is a normative prohibition**, not a preference: a single-policy win
|
|
|
|
|
|
rate may not be reported as a difficulty. The spec carries the measured
|
|
|
|
|
|
reason — greedy 100% vs first-legal 0% on identical deals.
|
|
|
|
|
|
|
CB-WP-0024/0025: what play reported, split into a renderer and a search
Five remarks from the maintainer's test games, checked against the code
before being written down — two had already reached ground-game on wrong
premises, so a claim now names the line that makes it true.
Three of the five turned out to be data the projection already carries,
drawn as text: solution_deck_len, solution_discard, and OutcomeView's
personal/mastery/winners. One control (`close — I have read this`) is
labelled as a reading but shuts the server down, and leaves a live-looking
page pointing at a dead port. One number does not exist at all: table.rs
loops run_game and keeps only the last summary.
CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the
other seats' plays as objects on the table, the ending control saying what
it does, and a tally that survives "play again".
CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how
hard is this" are the same search asked twice. Tier L because the
information boundary is the whole design problem: a solver reading
GroundState sees the deck the rules hide, and would tell the maintainer he
could have won by playing a card he had no way to know was there. Also
unblocks GROUND-WP-0005, active with both tasks waiting on a measured
difficulty baseline.
Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023
shipped CB-EV-0020 first.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 12:59:08 +02:00
|
|
|
|
## Task: build the witness
|
|
|
|
|
|
|
|
|
|
|
|
```task
|
|
|
|
|
|
id: CB-WP-0025-T05
|
CB-WP-0025 T05: the search works, and it falsified this pass's own
affordability projection
games/ground/src/search.rs, five tests. It finds real winning lines and
replays them through validate/fold to group_success.
Two bugs in my own work, found and fixed here.
THE TRAVERSAL WAS WRONG. It branched on "the first seat with any legal
command" and stopped there, so a later seat never acted if an earlier one
was already selected but still had a legal move. Restructured around what
the rules oblige: a seat without a selection MUST select (GR-R02) and
nothing else can happen first; after Reveal the optional actions branch
freely and the aggregate rejects Resolve until the obligatory ones are
done -- so the search needs no phase logic of its own.
AND MY REWIND WAS OFF BY ONE ROUND, replaying the round it was meant to
search. That is why the first run reported 3 nodes and looked like a
working search.
Measured with the real search, rewinding real games to the start of their
last K rounds:
2p K=1 exhausted, 8,103 nodes, ~29 ms
2p K=2 budget cut at 2,000,000 nodes, ~5 s
3p K=2 win found, 41 nodes, ~157 us
The spec's own falsifier said "§3 fails if K=2 proves unaffordable at four
seats". IT FAILED AT TWO. The projection assumed a joint product per
round; the search explores sequential per-seat decisions, so orderings
multiply the tree far beyond width^seats. That is the second projection
this pass published in place of a measurement -- C1's timer was the first.
THE ASYMMETRY IS THE OPERATIVE FINDING. Finding a win is cheap: DFS
stumbles onto one in tens of nodes. Proving none exists needs exhaustion.
So the witness feature is affordable now at any K a player would ask
about, and the winnable fraction (ADR-0013 D4) is NOT, because its
negative half must exhaust every deal it counts. K=1 is the honest default
for exhaustive answers today; making K=2 exhaustible needs transposition
or move-ordering, neither of which this pass built. specs §3 and §3.1
corrected accordingly, and the K=2 default withdrawn.
The negative control that makes "winnable" falsifiable: 2p seed 7 over its
last round returns NoneFound with exhausted=true in ~8k nodes -- a real
negative, not a budget cut wearing a verdict's clothes. And the visible/
hidden marking is tested both ways, since a marking that can only say YES
is decoration.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 19:10:25 +02:00
|
|
|
|
status: done
|
CB-WP-0024/0025: what play reported, split into a renderer and a search
Five remarks from the maintainer's test games, checked against the code
before being written down — two had already reached ground-game on wrong
premises, so a claim now names the line that makes it true.
Three of the five turned out to be data the projection already carries,
drawn as text: solution_deck_len, solution_discard, and OutcomeView's
personal/mastery/winners. One control (`close — I have read this`) is
labelled as a reading but shuts the server down, and leaves a live-looking
page pointing at a dead port. One number does not exist at all: table.rs
loops run_game and keeps only the last summary.
CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the
other seats' plays as objects on the table, the ending control saying what
it does, and a tally that survives "play again".
CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how
hard is this" are the same search asked twice. Tier L because the
information boundary is the whole design problem: a solver reading
GroundState sees the deck the rules hide, and would tell the maintainer he
could have won by playing a card he had no way to know was there. Also
unblocks GROUND-WP-0005, active with both tasks waiting on a measured
difficulty baseline.
Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023
shipped CB-EV-0020 first.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 12:59:08 +02:00
|
|
|
|
priority: high
|
2026-08-05 13:00:02 +02:00
|
|
|
|
state_hub_task_id: "5f1cbda3-7427-4f23-8c67-b2e216b987a8"
|
CB-WP-0024/0025: what play reported, split into a renderer and a search
Five remarks from the maintainer's test games, checked against the code
before being written down — two had already reached ground-game on wrong
premises, so a claim now names the line that makes it true.
Three of the five turned out to be data the projection already carries,
drawn as text: solution_deck_len, solution_discard, and OutcomeView's
personal/mastery/winners. One control (`close — I have read this`) is
labelled as a reading but shuts the server down, and leaves a live-looking
page pointing at a dead port. One number does not exist at all: table.rs
loops run_game and keeps only the last summary.
CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the
other seats' plays as objects on the table, the ending control saying what
it does, and a tally that survives "play again".
CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how
hard is this" are the same search asked twice. Tier L because the
information boundary is the whole design problem: a solver reading
GroundState sees the deck the rules hide, and would tell the maintainer he
could have won by playing a card he had no way to know was there. Also
unblocks GROUND-WP-0005, active with both tasks waiting on a measured
difficulty baseline.
Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023
shipped CB-EV-0020 first.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 12:59:08 +02:00
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
The search, the bound, and the replayable path. Wire it to the ending page
|
|
|
|
|
|
so a lost game can be asked the question — the page CB-WP-0024 T01 is
|
|
|
|
|
|
already reworking, so land that first or expect a conflict.
|
|
|
|
|
|
|
|
|
|
|
|
**A game that was won is not asked the question.** The feature exists for
|
|
|
|
|
|
a loss.
|
|
|
|
|
|
|
|
|
|
|
|
**Controls:**
|
|
|
|
|
|
- every emitted witness replays to `group_success` through the existing
|
|
|
|
|
|
runner — asserted, not spot-checked;
|
|
|
|
|
|
- a deal constructed to be unwinnable returns *none found*, and the test
|
|
|
|
|
|
says which construction makes it so;
|
|
|
|
|
|
- the information boundary is mutation-provable: relax it, and a test
|
|
|
|
|
|
naming *that* boundary goes red. If it cannot be mutated, it was a
|
|
|
|
|
|
comment rather than a rule.
|
|
|
|
|
|
|
CB-WP-0025 T05: the search works, and it falsified this pass's own
affordability projection
games/ground/src/search.rs, five tests. It finds real winning lines and
replays them through validate/fold to group_success.
Two bugs in my own work, found and fixed here.
THE TRAVERSAL WAS WRONG. It branched on "the first seat with any legal
command" and stopped there, so a later seat never acted if an earlier one
was already selected but still had a legal move. Restructured around what
the rules oblige: a seat without a selection MUST select (GR-R02) and
nothing else can happen first; after Reveal the optional actions branch
freely and the aggregate rejects Resolve until the obligatory ones are
done -- so the search needs no phase logic of its own.
AND MY REWIND WAS OFF BY ONE ROUND, replaying the round it was meant to
search. That is why the first run reported 3 nodes and looked like a
working search.
Measured with the real search, rewinding real games to the start of their
last K rounds:
2p K=1 exhausted, 8,103 nodes, ~29 ms
2p K=2 budget cut at 2,000,000 nodes, ~5 s
3p K=2 win found, 41 nodes, ~157 us
The spec's own falsifier said "§3 fails if K=2 proves unaffordable at four
seats". IT FAILED AT TWO. The projection assumed a joint product per
round; the search explores sequential per-seat decisions, so orderings
multiply the tree far beyond width^seats. That is the second projection
this pass published in place of a measurement -- C1's timer was the first.
THE ASYMMETRY IS THE OPERATIVE FINDING. Finding a win is cheap: DFS
stumbles onto one in tens of nodes. Proving none exists needs exhaustion.
So the witness feature is affordable now at any K a player would ask
about, and the winnable fraction (ADR-0013 D4) is NOT, because its
negative half must exhaust every deal it counts. K=1 is the honest default
for exhaustive answers today; making K=2 exhaustible needs transposition
or move-ordering, neither of which this pass built. specs §3 and §3.1
corrected accordingly, and the K=2 default withdrawn.
The negative control that makes "winnable" falsifiable: 2p seed 7 over its
last round returns NoneFound with exhausted=true in ~8k nodes -- a real
negative, not a budget cut wearing a verdict's clothes. And the visible/
hidden marking is tested both ways, since a marking that can only say YES
is decoration.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 19:10:25 +02:00
|
|
|
|
**Done 2026-08-05.** `games/ground/src/search.rs`, five tests.
|
|
|
|
|
|
|
|
|
|
|
|
**The first traversal was wrong and the diagnostic hid it.** It branched
|
|
|
|
|
|
on *the first seat with any legal command* and stopped there, so a later
|
|
|
|
|
|
seat never acted if an earlier one was already selected. Restructured
|
|
|
|
|
|
around what the rules oblige: a seat without a selection **must** select
|
|
|
|
|
|
(GR-R02) and nothing else can happen first; after Reveal the optional
|
|
|
|
|
|
actions branch freely, and the aggregate rejects `Resolve` until the
|
|
|
|
|
|
obligatory ones are done — **so the search needs no phase logic of its
|
|
|
|
|
|
own.**
|
|
|
|
|
|
|
|
|
|
|
|
**And my rewind was off by one round**, replaying the round it was meant
|
CB-WP-0025 T06/T07: a difficulty baseline that reports its own confound
T06. games/ground/examples/difficulty.rs, make difficulty, wired into make
self-tests, and a report file in ground-game under GROUND-WP-0005 with a
hub message pointing at it.
THE REPORT OPENS WITH THE RETRACTION, because what this task was written
to send was withdrawn by T02 and GROUND-WP-0005 is blocked on exactly that
number. They are told, in the first section, that we nearly sent them "the
game is too easy at 5-6 seats" and why it was wrong.
seats winnable greedy random first-legal spread
2p 60% 60.0% 5.0% 76.7% 71.7
3p 93% 88.3% 6.7% 25.0% 81.7
4p 100% 93.3% 6.7% 30.0% 86.7
5p 100% 100.0% 3.3% 0.0% 100.0
6p 100% 100.0% 3.3% 0.0% 100.0
SPREAD justifies the whole redesign: 71.7 to 100.0 points between three
trivial policies. The table now shows why no single rate is a difficulty
rather than asserting it.
And the 5-6 rows point the OPPOSITE way from the withdrawn claim --
first-legal 0% against greedy 100% is the widest spread in the table,
which suggests play matters MORE there, not less. Neither reading is
established and the report says so.
The confound is stated in the tool's own output, not only in prose:
`winnable` is conditioned on greedy's play up to the final round, because
searching from round 1 is unaffordable. Presenting it as a property of the
deal would repeat this pass's error in a subtler form -- which is exactly
how a corrected project reintroduces a defect. NO THRESHOLD CHANGES ARE
PROPOSED.
The instrument can fail (spec §5): a witness must replay to a win, an
unwinnable position must report searched-out rather than a budget cut, a
one-node budget must not claim exhaustion, and the policy panel must
actually disagree. difficulty-baseline.rs marked superseded, kept as the
survey's dated snapshot. Registered as F16, inconsistent / withdrawn.
T07. evidence/CB-EV-0024. Five of nine defects came only from the review;
four from execution, and all four of those were in work written after it.
The wrong-denominator family now has five instances and still no control
-- facts-check catches copies that disagree, nothing catches a number
computed correctly against the wrong base. Tier L was an over-declaration
(no port, structurally M) and paid for itself anyway, because the review
is L-only. Chaos window 2 will close with zero overrides, making its
retirement condition untestable.
Named as open rather than implied done: the witness is NOT wired to the
ending page. The search works; the browser cannot ask it yet.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 19:24:49 +02:00
|
|
|
|
to search — which is why the first run reported 3 nodes and looked fine.
|
CB-WP-0025 T05: the search works, and it falsified this pass's own
affordability projection
games/ground/src/search.rs, five tests. It finds real winning lines and
replays them through validate/fold to group_success.
Two bugs in my own work, found and fixed here.
THE TRAVERSAL WAS WRONG. It branched on "the first seat with any legal
command" and stopped there, so a later seat never acted if an earlier one
was already selected but still had a legal move. Restructured around what
the rules oblige: a seat without a selection MUST select (GR-R02) and
nothing else can happen first; after Reveal the optional actions branch
freely and the aggregate rejects Resolve until the obligatory ones are
done -- so the search needs no phase logic of its own.
AND MY REWIND WAS OFF BY ONE ROUND, replaying the round it was meant to
search. That is why the first run reported 3 nodes and looked like a
working search.
Measured with the real search, rewinding real games to the start of their
last K rounds:
2p K=1 exhausted, 8,103 nodes, ~29 ms
2p K=2 budget cut at 2,000,000 nodes, ~5 s
3p K=2 win found, 41 nodes, ~157 us
The spec's own falsifier said "§3 fails if K=2 proves unaffordable at four
seats". IT FAILED AT TWO. The projection assumed a joint product per
round; the search explores sequential per-seat decisions, so orderings
multiply the tree far beyond width^seats. That is the second projection
this pass published in place of a measurement -- C1's timer was the first.
THE ASYMMETRY IS THE OPERATIVE FINDING. Finding a win is cheap: DFS
stumbles onto one in tens of nodes. Proving none exists needs exhaustion.
So the witness feature is affordable now at any K a player would ask
about, and the winnable fraction (ADR-0013 D4) is NOT, because its
negative half must exhaust every deal it counts. K=1 is the honest default
for exhaustive answers today; making K=2 exhaustible needs transposition
or move-ordering, neither of which this pass built. specs §3 and §3.1
corrected accordingly, and the K=2 default withdrawn.
The negative control that makes "winnable" falsifiable: 2p seed 7 over its
last round returns NoneFound with exhausted=true in ~8k nodes -- a real
negative, not a budget cut wearing a verdict's clothes. And the visible/
hidden marking is tested both ways, since a marking that can only say YES
is decoration.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 19:10:25 +02:00
|
|
|
|
|
CB-WP-0025 T06/T07: a difficulty baseline that reports its own confound
T06. games/ground/examples/difficulty.rs, make difficulty, wired into make
self-tests, and a report file in ground-game under GROUND-WP-0005 with a
hub message pointing at it.
THE REPORT OPENS WITH THE RETRACTION, because what this task was written
to send was withdrawn by T02 and GROUND-WP-0005 is blocked on exactly that
number. They are told, in the first section, that we nearly sent them "the
game is too easy at 5-6 seats" and why it was wrong.
seats winnable greedy random first-legal spread
2p 60% 60.0% 5.0% 76.7% 71.7
3p 93% 88.3% 6.7% 25.0% 81.7
4p 100% 93.3% 6.7% 30.0% 86.7
5p 100% 100.0% 3.3% 0.0% 100.0
6p 100% 100.0% 3.3% 0.0% 100.0
SPREAD justifies the whole redesign: 71.7 to 100.0 points between three
trivial policies. The table now shows why no single rate is a difficulty
rather than asserting it.
And the 5-6 rows point the OPPOSITE way from the withdrawn claim --
first-legal 0% against greedy 100% is the widest spread in the table,
which suggests play matters MORE there, not less. Neither reading is
established and the report says so.
The confound is stated in the tool's own output, not only in prose:
`winnable` is conditioned on greedy's play up to the final round, because
searching from round 1 is unaffordable. Presenting it as a property of the
deal would repeat this pass's error in a subtler form -- which is exactly
how a corrected project reintroduces a defect. NO THRESHOLD CHANGES ARE
PROPOSED.
The instrument can fail (spec §5): a witness must replay to a win, an
unwinnable position must report searched-out rather than a budget cut, a
one-node budget must not claim exhaustion, and the policy panel must
actually disagree. difficulty-baseline.rs marked superseded, kept as the
survey's dated snapshot. Registered as F16, inconsistent / withdrawn.
T07. evidence/CB-EV-0024. Five of nine defects came only from the review;
four from execution, and all four of those were in work written after it.
The wrong-denominator family now has five instances and still no control
-- facts-check catches copies that disagree, nothing catches a number
computed correctly against the wrong base. Tier L was an over-declaration
(no port, structurally M) and paid for itself anyway, because the review
is L-only. Chaos window 2 will close with zero overrides, making its
retirement condition untestable.
Named as open rather than implied done: the witness is NOT wired to the
ending page. The search works; the browser cannot ask it yet.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 19:24:49 +02:00
|
|
|
|
**The measurement falsified the spec's own projection at two seats, where
|
|
|
|
|
|
§6 predicted four**: 2p `K=1` exhausted in 8,103 nodes (~29 ms), 2p `K=2`
|
|
|
|
|
|
cut at 2,000,000 nodes (~5 s), 3p `K=2` won in 41 nodes.
|
CB-WP-0025 T05: the search works, and it falsified this pass's own
affordability projection
games/ground/src/search.rs, five tests. It finds real winning lines and
replays them through validate/fold to group_success.
Two bugs in my own work, found and fixed here.
THE TRAVERSAL WAS WRONG. It branched on "the first seat with any legal
command" and stopped there, so a later seat never acted if an earlier one
was already selected but still had a legal move. Restructured around what
the rules oblige: a seat without a selection MUST select (GR-R02) and
nothing else can happen first; after Reveal the optional actions branch
freely and the aggregate rejects Resolve until the obligatory ones are
done -- so the search needs no phase logic of its own.
AND MY REWIND WAS OFF BY ONE ROUND, replaying the round it was meant to
search. That is why the first run reported 3 nodes and looked like a
working search.
Measured with the real search, rewinding real games to the start of their
last K rounds:
2p K=1 exhausted, 8,103 nodes, ~29 ms
2p K=2 budget cut at 2,000,000 nodes, ~5 s
3p K=2 win found, 41 nodes, ~157 us
The spec's own falsifier said "§3 fails if K=2 proves unaffordable at four
seats". IT FAILED AT TWO. The projection assumed a joint product per
round; the search explores sequential per-seat decisions, so orderings
multiply the tree far beyond width^seats. That is the second projection
this pass published in place of a measurement -- C1's timer was the first.
THE ASYMMETRY IS THE OPERATIVE FINDING. Finding a win is cheap: DFS
stumbles onto one in tens of nodes. Proving none exists needs exhaustion.
So the witness feature is affordable now at any K a player would ask
about, and the winnable fraction (ADR-0013 D4) is NOT, because its
negative half must exhaust every deal it counts. K=1 is the honest default
for exhaustive answers today; making K=2 exhaustible needs transposition
or move-ordering, neither of which this pass built. specs §3 and §3.1
corrected accordingly, and the K=2 default withdrawn.
The negative control that makes "winnable" falsifiable: 2p seed 7 over its
last round returns NoneFound with exhausted=true in ~8k nodes -- a real
negative, not a budget cut wearing a verdict's clothes. And the visible/
hidden marking is tested both ways, since a marking that can only say YES
is decoration.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 19:10:25 +02:00
|
|
|
|
|
|
|
|
|
|
The projection assumed a joint product per round; the search explores
|
|
|
|
|
|
sequential per-seat decisions, so orderings multiply the tree far beyond
|
|
|
|
|
|
`width^seats`.
|
|
|
|
|
|
|
CB-WP-0025 T06/T07: a difficulty baseline that reports its own confound
T06. games/ground/examples/difficulty.rs, make difficulty, wired into make
self-tests, and a report file in ground-game under GROUND-WP-0005 with a
hub message pointing at it.
THE REPORT OPENS WITH THE RETRACTION, because what this task was written
to send was withdrawn by T02 and GROUND-WP-0005 is blocked on exactly that
number. They are told, in the first section, that we nearly sent them "the
game is too easy at 5-6 seats" and why it was wrong.
seats winnable greedy random first-legal spread
2p 60% 60.0% 5.0% 76.7% 71.7
3p 93% 88.3% 6.7% 25.0% 81.7
4p 100% 93.3% 6.7% 30.0% 86.7
5p 100% 100.0% 3.3% 0.0% 100.0
6p 100% 100.0% 3.3% 0.0% 100.0
SPREAD justifies the whole redesign: 71.7 to 100.0 points between three
trivial policies. The table now shows why no single rate is a difficulty
rather than asserting it.
And the 5-6 rows point the OPPOSITE way from the withdrawn claim --
first-legal 0% against greedy 100% is the widest spread in the table,
which suggests play matters MORE there, not less. Neither reading is
established and the report says so.
The confound is stated in the tool's own output, not only in prose:
`winnable` is conditioned on greedy's play up to the final round, because
searching from round 1 is unaffordable. Presenting it as a property of the
deal would repeat this pass's error in a subtler form -- which is exactly
how a corrected project reintroduces a defect. NO THRESHOLD CHANGES ARE
PROPOSED.
The instrument can fail (spec §5): a witness must replay to a win, an
unwinnable position must report searched-out rather than a budget cut, a
one-node budget must not claim exhaustion, and the policy panel must
actually disagree. difficulty-baseline.rs marked superseded, kept as the
survey's dated snapshot. Registered as F16, inconsistent / withdrawn.
T07. evidence/CB-EV-0024. Five of nine defects came only from the review;
four from execution, and all four of those were in work written after it.
The wrong-denominator family now has five instances and still no control
-- facts-check catches copies that disagree, nothing catches a number
computed correctly against the wrong base. Tier L was an over-declaration
(no port, structurally M) and paid for itself anyway, because the review
is L-only. Chaos window 2 will close with zero overrides, making its
retirement condition untestable.
Named as open rather than implied done: the witness is NOT wired to the
ending page. The search works; the browser cannot ask it yet.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 19:24:49 +02:00
|
|
|
|
**The asymmetry is the operative finding.** *Finding* a win is cheap;
|
|
|
|
|
|
*proving none exists* needs exhaustion. So the **witness feature is
|
|
|
|
|
|
affordable now** and the **winnable fraction is not**, because its
|
|
|
|
|
|
negative half must exhaust every deal it counts.
|
CB-WP-0025 T05: the search works, and it falsified this pass's own
affordability projection
games/ground/src/search.rs, five tests. It finds real winning lines and
replays them through validate/fold to group_success.
Two bugs in my own work, found and fixed here.
THE TRAVERSAL WAS WRONG. It branched on "the first seat with any legal
command" and stopped there, so a later seat never acted if an earlier one
was already selected but still had a legal move. Restructured around what
the rules oblige: a seat without a selection MUST select (GR-R02) and
nothing else can happen first; after Reveal the optional actions branch
freely and the aggregate rejects Resolve until the obligatory ones are
done -- so the search needs no phase logic of its own.
AND MY REWIND WAS OFF BY ONE ROUND, replaying the round it was meant to
search. That is why the first run reported 3 nodes and looked like a
working search.
Measured with the real search, rewinding real games to the start of their
last K rounds:
2p K=1 exhausted, 8,103 nodes, ~29 ms
2p K=2 budget cut at 2,000,000 nodes, ~5 s
3p K=2 win found, 41 nodes, ~157 us
The spec's own falsifier said "§3 fails if K=2 proves unaffordable at four
seats". IT FAILED AT TWO. The projection assumed a joint product per
round; the search explores sequential per-seat decisions, so orderings
multiply the tree far beyond width^seats. That is the second projection
this pass published in place of a measurement -- C1's timer was the first.
THE ASYMMETRY IS THE OPERATIVE FINDING. Finding a win is cheap: DFS
stumbles onto one in tens of nodes. Proving none exists needs exhaustion.
So the witness feature is affordable now at any K a player would ask
about, and the winnable fraction (ADR-0013 D4) is NOT, because its
negative half must exhaust every deal it counts. K=1 is the honest default
for exhaustive answers today; making K=2 exhaustible needs transposition
or move-ordering, neither of which this pass built. specs §3 and §3.1
corrected accordingly, and the K=2 default withdrawn.
The negative control that makes "winnable" falsifiable: 2p seed 7 over its
last round returns NoneFound with exhausted=true in ~8k nodes -- a real
negative, not a budget cut wearing a verdict's clothes. And the visible/
hidden marking is tested both ways, since a marking that can only say YES
is decoration.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 19:10:25 +02:00
|
|
|
|
|
CB-WP-0024/0025: what play reported, split into a renderer and a search
Five remarks from the maintainer's test games, checked against the code
before being written down — two had already reached ground-game on wrong
premises, so a claim now names the line that makes it true.
Three of the five turned out to be data the projection already carries,
drawn as text: solution_deck_len, solution_discard, and OutcomeView's
personal/mastery/winners. One control (`close — I have read this`) is
labelled as a reading but shuts the server down, and leaves a live-looking
page pointing at a dead port. One number does not exist at all: table.rs
loops run_game and keeps only the last summary.
CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the
other seats' plays as objects on the table, the ending control saying what
it does, and a tally that survives "play again".
CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how
hard is this" are the same search asked twice. Tier L because the
information boundary is the whole design problem: a solver reading
GroundState sees the deck the rules hide, and would tell the maintainer he
could have won by playing a card he had no way to know was there. Also
unblocks GROUND-WP-0005, active with both tasks waiting on a measured
difficulty baseline.
Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023
shipped CB-EV-0020 first.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 12:59:08 +02:00
|
|
|
|
## Task: measure the difficulty, and hand it to ground-game
|
|
|
|
|
|
|
|
|
|
|
|
```task
|
|
|
|
|
|
id: CB-WP-0025-T06
|
CB-WP-0025 T06/T07: a difficulty baseline that reports its own confound
T06. games/ground/examples/difficulty.rs, make difficulty, wired into make
self-tests, and a report file in ground-game under GROUND-WP-0005 with a
hub message pointing at it.
THE REPORT OPENS WITH THE RETRACTION, because what this task was written
to send was withdrawn by T02 and GROUND-WP-0005 is blocked on exactly that
number. They are told, in the first section, that we nearly sent them "the
game is too easy at 5-6 seats" and why it was wrong.
seats winnable greedy random first-legal spread
2p 60% 60.0% 5.0% 76.7% 71.7
3p 93% 88.3% 6.7% 25.0% 81.7
4p 100% 93.3% 6.7% 30.0% 86.7
5p 100% 100.0% 3.3% 0.0% 100.0
6p 100% 100.0% 3.3% 0.0% 100.0
SPREAD justifies the whole redesign: 71.7 to 100.0 points between three
trivial policies. The table now shows why no single rate is a difficulty
rather than asserting it.
And the 5-6 rows point the OPPOSITE way from the withdrawn claim --
first-legal 0% against greedy 100% is the widest spread in the table,
which suggests play matters MORE there, not less. Neither reading is
established and the report says so.
The confound is stated in the tool's own output, not only in prose:
`winnable` is conditioned on greedy's play up to the final round, because
searching from round 1 is unaffordable. Presenting it as a property of the
deal would repeat this pass's error in a subtler form -- which is exactly
how a corrected project reintroduces a defect. NO THRESHOLD CHANGES ARE
PROPOSED.
The instrument can fail (spec §5): a witness must replay to a win, an
unwinnable position must report searched-out rather than a budget cut, a
one-node budget must not claim exhaustion, and the policy panel must
actually disagree. difficulty-baseline.rs marked superseded, kept as the
survey's dated snapshot. Registered as F16, inconsistent / withdrawn.
T07. evidence/CB-EV-0024. Five of nine defects came only from the review;
four from execution, and all four of those were in work written after it.
The wrong-denominator family now has five instances and still no control
-- facts-check catches copies that disagree, nothing catches a number
computed correctly against the wrong base. Tier L was an over-declaration
(no port, structurally M) and paid for itself anyway, because the review
is L-only. Chaos window 2 will close with zero overrides, making its
retirement condition untestable.
Named as open rather than implied done: the witness is NOT wired to the
ending page. The search works; the browser cannot ask it yet.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 19:24:49 +02:00
|
|
|
|
status: done
|
CB-WP-0024/0025: what play reported, split into a renderer and a search
Five remarks from the maintainer's test games, checked against the code
before being written down — two had already reached ground-game on wrong
premises, so a claim now names the line that makes it true.
Three of the five turned out to be data the projection already carries,
drawn as text: solution_deck_len, solution_discard, and OutcomeView's
personal/mastery/winners. One control (`close — I have read this`) is
labelled as a reading but shuts the server down, and leaves a live-looking
page pointing at a dead port. One number does not exist at all: table.rs
loops run_game and keeps only the last summary.
CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the
other seats' plays as objects on the table, the ending control saying what
it does, and a tally that survives "play again".
CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how
hard is this" are the same search asked twice. Tier L because the
information boundary is the whole design problem: a solver reading
GroundState sees the deck the rules hide, and would tell the maintainer he
could have won by playing a card he had no way to know was there. Also
unblocks GROUND-WP-0005, active with both tasks waiting on a measured
difficulty baseline.
Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023
shipped CB-EV-0020 first.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 12:59:08 +02:00
|
|
|
|
priority: high
|
2026-08-05 13:00:02 +02:00
|
|
|
|
state_hub_task_id: "723feae8-1ed4-4dd4-9330-8ec84625e5d9"
|
CB-WP-0024/0025: what play reported, split into a renderer and a search
Five remarks from the maintainer's test games, checked against the code
before being written down — two had already reached ground-game on wrong
premises, so a claim now names the line that makes it true.
Three of the five turned out to be data the projection already carries,
drawn as text: solution_deck_len, solution_discard, and OutcomeView's
personal/mastery/winners. One control (`close — I have read this`) is
labelled as a reading but shuts the server down, and leaves a live-looking
page pointing at a dead port. One number does not exist at all: table.rs
loops run_game and keeps only the last summary.
CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the
other seats' plays as objects on the table, the ending control saying what
it does, and a tally that survives "play again".
CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how
hard is this" are the same search asked twice. Tier L because the
information boundary is the whole design problem: a solver reading
GroundState sees the deck the rules hide, and would tell the maintainer he
could have won by playing a card he had no way to know was there. Also
unblocks GROUND-WP-0005, active with both tasks waiting on a measured
difficulty baseline.
Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023
shipped CB-EV-0020 first.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 12:59:08 +02:00
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
Run the measurement, ship it as a `make` target beside the other
|
CB-WP-0025 T06/T07: a difficulty baseline that reports its own confound
T06. games/ground/examples/difficulty.rs, make difficulty, wired into make
self-tests, and a report file in ground-game under GROUND-WP-0005 with a
hub message pointing at it.
THE REPORT OPENS WITH THE RETRACTION, because what this task was written
to send was withdrawn by T02 and GROUND-WP-0005 is blocked on exactly that
number. They are told, in the first section, that we nearly sent them "the
game is too easy at 5-6 seats" and why it was wrong.
seats winnable greedy random first-legal spread
2p 60% 60.0% 5.0% 76.7% 71.7
3p 93% 88.3% 6.7% 25.0% 81.7
4p 100% 93.3% 6.7% 30.0% 86.7
5p 100% 100.0% 3.3% 0.0% 100.0
6p 100% 100.0% 3.3% 0.0% 100.0
SPREAD justifies the whole redesign: 71.7 to 100.0 points between three
trivial policies. The table now shows why no single rate is a difficulty
rather than asserting it.
And the 5-6 rows point the OPPOSITE way from the withdrawn claim --
first-legal 0% against greedy 100% is the widest spread in the table,
which suggests play matters MORE there, not less. Neither reading is
established and the report says so.
The confound is stated in the tool's own output, not only in prose:
`winnable` is conditioned on greedy's play up to the final round, because
searching from round 1 is unaffordable. Presenting it as a property of the
deal would repeat this pass's error in a subtler form -- which is exactly
how a corrected project reintroduces a defect. NO THRESHOLD CHANGES ARE
PROPOSED.
The instrument can fail (spec §5): a witness must replay to a win, an
unwinnable position must report searched-out rather than a budget cut, a
one-node budget must not claim exhaustion, and the policy panel must
actually disagree. difficulty-baseline.rs marked superseded, kept as the
survey's dated snapshot. Registered as F16, inconsistent / withdrawn.
T07. evidence/CB-EV-0024. Five of nine defects came only from the review;
four from execution, and all four of those were in work written after it.
The wrong-denominator family now has five instances and still no control
-- facts-check catches copies that disagree, nothing catches a number
computed correctly against the wrong base. Tier L was an over-declaration
(no port, structurally M) and paid for itself anyway, because the review
is L-only. Chaos window 2 will close with zero overrides, making its
retirement condition untestable.
Named as open rather than implied done: the witness is NOT wired to the
ending page. The search works; the browser cannot ask it yet.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 19:24:49 +02:00
|
|
|
|
instruments, and send it to `ground-game` against **GROUND-WP-0005**,
|
|
|
|
|
|
which is active with both tasks waiting on exactly this. Per CB-WP-0022
|
|
|
|
|
|
T06 it lands as a **file in their repo under their workplan**, not only an
|
|
|
|
|
|
inbox entry — *the message that sat unread for four days is the baseline
|
|
|
|
|
|
to beat*.
|
CB-WP-0024/0025: what play reported, split into a renderer and a search
Five remarks from the maintainer's test games, checked against the code
before being written down — two had already reached ground-game on wrong
premises, so a claim now names the line that makes it true.
Three of the five turned out to be data the projection already carries,
drawn as text: solution_deck_len, solution_discard, and OutcomeView's
personal/mastery/winners. One control (`close — I have read this`) is
labelled as a reading but shuts the server down, and leaves a live-looking
page pointing at a dead port. One number does not exist at all: table.rs
loops run_game and keeps only the last summary.
CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the
other seats' plays as objects on the table, the ending control saying what
it does, and a tally that survives "play again".
CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how
hard is this" are the same search asked twice. Tier L because the
information boundary is the whole design problem: a solver reading
GroundState sees the deck the rules hide, and would tell the maintainer he
could have won by playing a card he had no way to know was there. Also
unblocks GROUND-WP-0005, active with both tasks waiting on a measured
difficulty baseline.
Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023
shipped CB-EV-0020 first.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 12:59:08 +02:00
|
|
|
|
|
|
|
|
|
|
**Controls:**
|
|
|
|
|
|
- the number regenerates from a single command, and `facts.toml` carries
|
|
|
|
|
|
it if anything else quotes it (§Single source of fact — `make
|
|
|
|
|
|
facts-check`);
|
|
|
|
|
|
- the report carries the row-level table the ruling requires;
|
|
|
|
|
|
- **the seed range and policy are in the number's name**, not in a
|
|
|
|
|
|
footnote.
|
|
|
|
|
|
|
CB-WP-0025 T06/T07: a difficulty baseline that reports its own confound
T06. games/ground/examples/difficulty.rs, make difficulty, wired into make
self-tests, and a report file in ground-game under GROUND-WP-0005 with a
hub message pointing at it.
THE REPORT OPENS WITH THE RETRACTION, because what this task was written
to send was withdrawn by T02 and GROUND-WP-0005 is blocked on exactly that
number. They are told, in the first section, that we nearly sent them "the
game is too easy at 5-6 seats" and why it was wrong.
seats winnable greedy random first-legal spread
2p 60% 60.0% 5.0% 76.7% 71.7
3p 93% 88.3% 6.7% 25.0% 81.7
4p 100% 93.3% 6.7% 30.0% 86.7
5p 100% 100.0% 3.3% 0.0% 100.0
6p 100% 100.0% 3.3% 0.0% 100.0
SPREAD justifies the whole redesign: 71.7 to 100.0 points between three
trivial policies. The table now shows why no single rate is a difficulty
rather than asserting it.
And the 5-6 rows point the OPPOSITE way from the withdrawn claim --
first-legal 0% against greedy 100% is the widest spread in the table,
which suggests play matters MORE there, not less. Neither reading is
established and the report says so.
The confound is stated in the tool's own output, not only in prose:
`winnable` is conditioned on greedy's play up to the final round, because
searching from round 1 is unaffordable. Presenting it as a property of the
deal would repeat this pass's error in a subtler form -- which is exactly
how a corrected project reintroduces a defect. NO THRESHOLD CHANGES ARE
PROPOSED.
The instrument can fail (spec §5): a witness must replay to a win, an
unwinnable position must report searched-out rather than a budget cut, a
one-node budget must not claim exhaustion, and the policy panel must
actually disagree. difficulty-baseline.rs marked superseded, kept as the
survey's dated snapshot. Registered as F16, inconsistent / withdrawn.
T07. evidence/CB-EV-0024. Five of nine defects came only from the review;
four from execution, and all four of those were in work written after it.
The wrong-denominator family now has five instances and still no control
-- facts-check catches copies that disagree, nothing catches a number
computed correctly against the wrong base. Tier L was an over-declaration
(no port, structurally M) and paid for itself anyway, because the review
is L-only. Chaos window 2 will close with zero overrides, making its
retirement condition untestable.
Named as open rather than implied done: the witness is NOT wired to the
ending page. The search works; the browser cannot ask it yet.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 19:24:49 +02:00
|
|
|
|
**Done 2026-08-05.** `games/ground/examples/difficulty.rs`, `make
|
|
|
|
|
|
difficulty`, wired into `make self-tests`, and
|
|
|
|
|
|
[a report file in ground-game](../../ground-game/workplans/GROUND-WP-0005-clay-borg-difficulty-260805.md)
|
|
|
|
|
|
under GROUND-WP-0005 with a hub message pointing at it.
|
|
|
|
|
|
|
|
|
|
|
|
**The report opens with the retraction**, because what this task was
|
|
|
|
|
|
written to send was withdrawn by T02 and GROUND-WP-0005 is blocked on
|
|
|
|
|
|
exactly that number.
|
|
|
|
|
|
|
|
|
|
|
|
Winnable-from-round-5, 60 seeds: **60% / 93% / 100% / 100% / 100%** at
|
|
|
|
|
|
2/3/4/5/6 seats, beside greedy / random / first-legal rates and the spread
|
|
|
|
|
|
between them. Full table in the report and in `make difficulty`.
|
|
|
|
|
|
|
|
|
|
|
|
**`spread` justifies the whole redesign** — 71.7 to 100.0 points between
|
|
|
|
|
|
three trivial policies. No single rate says anything about the game, and
|
|
|
|
|
|
the table now shows that rather than asserting it.
|
|
|
|
|
|
|
|
|
|
|
|
**The confound is stated in the tool's own output**: `winnable` is
|
|
|
|
|
|
conditioned on greedy's play up to the final round. Presenting it as a
|
|
|
|
|
|
property of the deal would repeat this pass's error in a subtler form.
|
|
|
|
|
|
**No threshold changes are proposed.** The instrument can fail (spec §5),
|
|
|
|
|
|
and `difficulty-baseline.rs` is marked superseded.
|
|
|
|
|
|
|
|
|
|
|
|
Registered as **F16, `inconsistent`, `withdrawn`**, with `difficulty.rs`
|
|
|
|
|
|
as its reproduction.
|
|
|
|
|
|
|
CB-WP-0024/0025: what play reported, split into a renderer and a search
Five remarks from the maintainer's test games, checked against the code
before being written down — two had already reached ground-game on wrong
premises, so a claim now names the line that makes it true.
Three of the five turned out to be data the projection already carries,
drawn as text: solution_deck_len, solution_discard, and OutcomeView's
personal/mastery/winners. One control (`close — I have read this`) is
labelled as a reading but shuts the server down, and leaves a live-looking
page pointing at a dead port. One number does not exist at all: table.rs
loops run_game and keeps only the last summary.
CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the
other seats' plays as objects on the table, the ending control saying what
it does, and a tally that survives "play again".
CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how
hard is this" are the same search asked twice. Tier L because the
information boundary is the whole design problem: a solver reading
GroundState sees the deck the rules hide, and would tell the maintainer he
could have won by playing a card he had no way to know was there. Also
unblocks GROUND-WP-0005, active with both tasks waiting on a measured
difficulty baseline.
Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023
shipped CB-EV-0020 first.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 12:59:08 +02:00
|
|
|
|
## Task: evidence
|
|
|
|
|
|
|
|
|
|
|
|
```task
|
|
|
|
|
|
id: CB-WP-0025-T07
|
CB-WP-0025 T06/T07: a difficulty baseline that reports its own confound
T06. games/ground/examples/difficulty.rs, make difficulty, wired into make
self-tests, and a report file in ground-game under GROUND-WP-0005 with a
hub message pointing at it.
THE REPORT OPENS WITH THE RETRACTION, because what this task was written
to send was withdrawn by T02 and GROUND-WP-0005 is blocked on exactly that
number. They are told, in the first section, that we nearly sent them "the
game is too easy at 5-6 seats" and why it was wrong.
seats winnable greedy random first-legal spread
2p 60% 60.0% 5.0% 76.7% 71.7
3p 93% 88.3% 6.7% 25.0% 81.7
4p 100% 93.3% 6.7% 30.0% 86.7
5p 100% 100.0% 3.3% 0.0% 100.0
6p 100% 100.0% 3.3% 0.0% 100.0
SPREAD justifies the whole redesign: 71.7 to 100.0 points between three
trivial policies. The table now shows why no single rate is a difficulty
rather than asserting it.
And the 5-6 rows point the OPPOSITE way from the withdrawn claim --
first-legal 0% against greedy 100% is the widest spread in the table,
which suggests play matters MORE there, not less. Neither reading is
established and the report says so.
The confound is stated in the tool's own output, not only in prose:
`winnable` is conditioned on greedy's play up to the final round, because
searching from round 1 is unaffordable. Presenting it as a property of the
deal would repeat this pass's error in a subtler form -- which is exactly
how a corrected project reintroduces a defect. NO THRESHOLD CHANGES ARE
PROPOSED.
The instrument can fail (spec §5): a witness must replay to a win, an
unwinnable position must report searched-out rather than a budget cut, a
one-node budget must not claim exhaustion, and the policy panel must
actually disagree. difficulty-baseline.rs marked superseded, kept as the
survey's dated snapshot. Registered as F16, inconsistent / withdrawn.
T07. evidence/CB-EV-0024. Five of nine defects came only from the review;
four from execution, and all four of those were in work written after it.
The wrong-denominator family now has five instances and still no control
-- facts-check catches copies that disagree, nothing catches a number
computed correctly against the wrong base. Tier L was an over-declaration
(no port, structurally M) and paid for itself anyway, because the review
is L-only. Chaos window 2 will close with zero overrides, making its
retirement condition untestable.
Named as open rather than implied done: the witness is NOT wired to the
ending page. The search works; the browser cannot ask it yet.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 19:24:49 +02:00
|
|
|
|
status: done
|
CB-WP-0024/0025: what play reported, split into a renderer and a search
Five remarks from the maintainer's test games, checked against the code
before being written down — two had already reached ground-game on wrong
premises, so a claim now names the line that makes it true.
Three of the five turned out to be data the projection already carries,
drawn as text: solution_deck_len, solution_discard, and OutcomeView's
personal/mastery/winners. One control (`close — I have read this`) is
labelled as a reading but shuts the server down, and leaves a live-looking
page pointing at a dead port. One number does not exist at all: table.rs
loops run_game and keeps only the last summary.
CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the
other seats' plays as objects on the table, the ending control saying what
it does, and a tally that survives "play again".
CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how
hard is this" are the same search asked twice. Tier L because the
information boundary is the whole design problem: a solver reading
GroundState sees the deck the rules hide, and would tell the maintainer he
could have won by playing a card he had no way to know was there. Also
unblocks GROUND-WP-0005, active with both tasks waiting on a measured
difficulty baseline.
Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023
shipped CB-EV-0020 first.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 12:59:08 +02:00
|
|
|
|
priority: high
|
2026-08-05 13:00:02 +02:00
|
|
|
|
state_hub_task_id: "2616e06c-5d88-4680-be18-4452515dff21"
|
CB-WP-0024/0025: what play reported, split into a renderer and a search
Five remarks from the maintainer's test games, checked against the code
before being written down — two had already reached ground-game on wrong
premises, so a claim now names the line that makes it true.
Three of the five turned out to be data the projection already carries,
drawn as text: solution_deck_len, solution_discard, and OutcomeView's
personal/mastery/winners. One control (`close — I have read this`) is
labelled as a reading but shuts the server down, and leaves a live-looking
page pointing at a dead port. One number does not exist at all: table.rs
loops run_game and keeps only the last summary.
CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the
other seats' plays as objects on the table, the ending control saying what
it does, and a tally that survives "play again".
CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how
hard is this" are the same search asked twice. Tier L because the
information boundary is the whole design problem: a solver reading
GroundState sees the deck the rules hide, and would tell the maintainer he
could have won by playing a card he had no way to know was there. Also
unblocks GROUND-WP-0005, active with both tasks waiting on a measured
difficulty baseline.
Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023
shipped CB-EV-0020 first.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 12:59:08 +02:00
|
|
|
|
```
|
|
|
|
|
|
|
CB-WP-0025 T06/T07: a difficulty baseline that reports its own confound
T06. games/ground/examples/difficulty.rs, make difficulty, wired into make
self-tests, and a report file in ground-game under GROUND-WP-0005 with a
hub message pointing at it.
THE REPORT OPENS WITH THE RETRACTION, because what this task was written
to send was withdrawn by T02 and GROUND-WP-0005 is blocked on exactly that
number. They are told, in the first section, that we nearly sent them "the
game is too easy at 5-6 seats" and why it was wrong.
seats winnable greedy random first-legal spread
2p 60% 60.0% 5.0% 76.7% 71.7
3p 93% 88.3% 6.7% 25.0% 81.7
4p 100% 93.3% 6.7% 30.0% 86.7
5p 100% 100.0% 3.3% 0.0% 100.0
6p 100% 100.0% 3.3% 0.0% 100.0
SPREAD justifies the whole redesign: 71.7 to 100.0 points between three
trivial policies. The table now shows why no single rate is a difficulty
rather than asserting it.
And the 5-6 rows point the OPPOSITE way from the withdrawn claim --
first-legal 0% against greedy 100% is the widest spread in the table,
which suggests play matters MORE there, not less. Neither reading is
established and the report says so.
The confound is stated in the tool's own output, not only in prose:
`winnable` is conditioned on greedy's play up to the final round, because
searching from round 1 is unaffordable. Presenting it as a property of the
deal would repeat this pass's error in a subtler form -- which is exactly
how a corrected project reintroduces a defect. NO THRESHOLD CHANGES ARE
PROPOSED.
The instrument can fail (spec §5): a witness must replay to a win, an
unwinnable position must report searched-out rather than a budget cut, a
one-node budget must not claim exhaustion, and the policy panel must
actually disagree. difficulty-baseline.rs marked superseded, kept as the
survey's dated snapshot. Registered as F16, inconsistent / withdrawn.
T07. evidence/CB-EV-0024. Five of nine defects came only from the review;
four from execution, and all four of those were in work written after it.
The wrong-denominator family now has five instances and still no control
-- facts-check catches copies that disagree, nothing catches a number
computed correctly against the wrong base. Tier L was an over-declaration
(no port, structurally M) and paid for itself anyway, because the review
is L-only. Chaos window 2 will close with zero overrides, making its
retirement condition untestable.
Named as open rather than implied done: the witness is NOT wired to the
ending page. The search works; the browser cannot ask it yet.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 19:24:49 +02:00
|
|
|
|
`evidence/CB-EV-*.md` (**CB-EV-0024** — CB-EV-0023 went to CB-WP-0024).
|
|
|
|
|
|
Was the maintainer's game winnable; what the
|
|
|
|
|
|
honest search cost against the omniscient one; whether the difficulty
|
|
|
|
|
|
measurement moved `ground-game`; what tier L cost against what it caught;
|
|
|
|
|
|
and CB-WP-0024's cost by re-running the instrument.
|
CB-WP-0024/0025: what play reported, split into a renderer and a search
Five remarks from the maintainer's test games, checked against the code
before being written down — two had already reached ground-game on wrong
premises, so a claim now names the line that makes it true.
Three of the five turned out to be data the projection already carries,
drawn as text: solution_deck_len, solution_discard, and OutcomeView's
personal/mastery/winners. One control (`close — I have read this`) is
labelled as a reading but shuts the server down, and leaves a live-looking
page pointing at a dead port. One number does not exist at all: table.rs
loops run_game and keeps only the last summary.
CB-WP-0024 (S, chaos d8=6, declaration 7 of window 2) — the piles and the
other seats' plays as objects on the table, the ending control saying what
it does, and a tally that survives "play again".
CB-WP-0025 (L, chaos d8=6, declaration 8) — "could we have won" and "how
hard is this" are the same search asked twice. Tier L because the
information boundary is the whole design problem: a solver reading
GroundState sees the deck the rules hide, and would tell the maintainer he
could have won by playing a card he had no way to know was there. Also
unblocks GROUND-WP-0005, active with both tasks waiting on a measured
difficulty baseline.
Also: CB-WP-0022 T07's evidence file moves to CB-EV-0021 — CB-WP-0023
shipped CB-EV-0020 first.
loop-lint: no findings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 12:59:08 +02:00
|
|
|
|
|
CB-WP-0025 T06/T07: a difficulty baseline that reports its own confound
T06. games/ground/examples/difficulty.rs, make difficulty, wired into make
self-tests, and a report file in ground-game under GROUND-WP-0005 with a
hub message pointing at it.
THE REPORT OPENS WITH THE RETRACTION, because what this task was written
to send was withdrawn by T02 and GROUND-WP-0005 is blocked on exactly that
number. They are told, in the first section, that we nearly sent them "the
game is too easy at 5-6 seats" and why it was wrong.
seats winnable greedy random first-legal spread
2p 60% 60.0% 5.0% 76.7% 71.7
3p 93% 88.3% 6.7% 25.0% 81.7
4p 100% 93.3% 6.7% 30.0% 86.7
5p 100% 100.0% 3.3% 0.0% 100.0
6p 100% 100.0% 3.3% 0.0% 100.0
SPREAD justifies the whole redesign: 71.7 to 100.0 points between three
trivial policies. The table now shows why no single rate is a difficulty
rather than asserting it.
And the 5-6 rows point the OPPOSITE way from the withdrawn claim --
first-legal 0% against greedy 100% is the widest spread in the table,
which suggests play matters MORE there, not less. Neither reading is
established and the report says so.
The confound is stated in the tool's own output, not only in prose:
`winnable` is conditioned on greedy's play up to the final round, because
searching from round 1 is unaffordable. Presenting it as a property of the
deal would repeat this pass's error in a subtler form -- which is exactly
how a corrected project reintroduces a defect. NO THRESHOLD CHANGES ARE
PROPOSED.
The instrument can fail (spec §5): a witness must replay to a win, an
unwinnable position must report searched-out rather than a budget cut, a
one-node budget must not claim exhaustion, and the policy panel must
actually disagree. difficulty-baseline.rs marked superseded, kept as the
survey's dated snapshot. Registered as F16, inconsistent / withdrawn.
T07. evidence/CB-EV-0024. Five of nine defects came only from the review;
four from execution, and all four of those were in work written after it.
The wrong-denominator family now has five instances and still no control
-- facts-check catches copies that disagree, nothing catches a number
computed correctly against the wrong base. Tier L was an over-declaration
(no port, structurally M) and paid for itself anyway, because the review
is L-only. Chaos window 2 will close with zero overrides, making its
retirement condition untestable.
Named as open rather than implied done: the witness is NOT wired to the
ending page. The search works; the browser cannot ask it yet.
make all: exit 0. loop-lint clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 19:24:49 +02:00
|
|
|
|
**Done 2026-08-05.**
|
|
|
|
|
|
[CB-EV-0024](../evidence/CB-EV-0024-could-we-have-won.md).
|
|
|
|
|
|
|
|
|
|
|
|
- **The maintainer's question is answered.** 3p seed 7 was winnable — a
|
|
|
|
|
|
10-move line, found in 41 nodes, replays to `group_success`. 2p seed 7
|
|
|
|
|
|
was **not**, exhaustively, in 8,103 nodes. **The witness feature works;
|
|
|
|
|
|
the difficulty half is where the trouble was.**
|
|
|
|
|
|
- **Five of nine defects came only from the review**, four from execution
|
|
|
|
|
|
— and all four of those were in work written *after* it.
|
|
|
|
|
|
- **The wrong-denominator family now has five instances and no control.**
|
|
|
|
|
|
- **The 5–6 seat rows point the opposite way from the withdrawn claim**;
|
|
|
|
|
|
neither reading is established, and the evidence says so.
|
|
|
|
|
|
- **Tier L was an over-declaration** (no port) **and paid for itself
|
|
|
|
|
|
anyway** — the review is L-only, and without it a false number reaches a
|
|
|
|
|
|
blocked repo.
|
|
|
|
|
|
- **Chaos window 2 will close with zero overrides**, making its retirement
|
|
|
|
|
|
condition untestable.
|
|
|
|
|
|
|
|
|
|
|
|
**Not done, and named rather than quietly dropped:** the witness is **not
|
|
|
|
|
|
wired to the ending page**, so the browser cannot yet ask the question.
|
|
|
|
|
|
That is the maintainer-facing half of remark 2.
|