clay-borg/workplans
tegwick 1f0f652920
Some checks are pending
ci / check (push) Waiting to run
CB-WP-0025 T02: the review withdrew the finding, and the fifth wrong
premise never left the repo

Separate agent, second tier-L review in this project. Six of seven
challenges conceded. The survey's headline finding is WITHDRAWN, not
softened.

C4 kills it, and the reviewer ranked it fourth. A FirstLegal policy --
take legal[0], no heuristic at all -- scores 0% at five and six seats
where GreedyPolicy scores 100%, and 77.5% at two seats where greedy scores
66%. Two unsophisticated agents span the entire range at the same seat
count. "The game is too easy at 5-6 seats" is therefore a statement about
GreedyPolicy, not about GROUND. The rescue the reviewer offered -- greedy
hits the 12-point ceiling in 200/200 deals, so the 6p row is a rules claim
-- dies on the same data: FirstLegal reaches that ceiling never.

C1: the per-node cost was wrong by 30-50x. The timer started before the
seed loop, so "us/node" included two setups, an entire greedy game and a
full validate+fold replay, divided by player-decision count. The tell was
in my own published output and I did not look at it: the figure FELL
(161/139/112) as branching ROSE (4.7/7.4/9.1), which no per-enumeration
cost can do. Re-measured with the clock around legal_commands alone:
3.0/3.5/4.1 us, now rising with branching. The reviewer measured 15.6-20.4
by a different isolation; we disagree by ~5x and neither has established
which is right, so T04 must benchmark it with criterion rather than adopt
either number.

C6: "exhaustive search is out at any seat count" is false -- ~3 seconds
over the last two rounds at 3p. With C1's correction the budget is
~10^5-10^6 nodes and bounded endgame search fits, so ADR-0013 cannot open
with "exhaustive is impossible, therefore determinized sampling" --
especially as sampling carries strategy fusion that exhaustive search does
not.

C3: the finding failed the admissibility rule this project wrote nine
hours earlier. 6/9/12 are sums where GROUND-WP-0004 T02 requires
per-priority rows, and the harness has no assertions, no --self-test and
no make target, so nothing can turn it red -- a `default` artifact wearing
a `counterexample` label, by CB-WP-0022 T05's own distinction.

C2: the ratio story explains nothing; 3p and 4p share deal, threshold and
ratio and differ by 12.5 points of win rate. C5: "explains the
maintainer's report" is contradicted by lib.rs:2487, which records his
losses as 3-player games on the pre-ruling deal, arithmetically unwinnable
at 6 against 7.

T06 exists to report to GROUND-WP-0005, which is BLOCKED waiting on a
difficulty baseline. Had this proceeded they would have been invited to
move thresholds on the strength of one bot's behaviour. That is the fifth
wrong premise this project would have sent them, and the second stopped by
an adversarial review rather than by a control. Both tier-L reviews here
have now caught a false headline that every gate passed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 18:06:03 +02:00
..
CB-WP-0001-inner-loop.md CB-WP-0007 T01+T03: window the metric, budget it, cap meta at 25% 2026-08-01 14:12:07 +02:00
CB-WP-0002-cost-accounting.md CB-WP-0007 T01+T03: window the metric, budget it, cap meta at 25% 2026-08-01 14:12:07 +02:00
CB-WP-0003-loop-hardening.md CB-WP-0007 T01+T03: window the metric, budget it, cap meta at 25% 2026-08-01 14:12:07 +02:00
CB-WP-0004-mechanical-work.md CB-WP-0007 T01+T03: window the metric, budget it, cap meta at 25% 2026-08-01 14:12:07 +02:00
CB-WP-0005-assertion-coverage.md CB-WP-0007 T01+T03: window the metric, budget it, cap meta at 25% 2026-08-01 14:12:07 +02:00
CB-WP-0006-instrument-the-table.md CB-WP-0007 T01+T03: window the metric, budget it, cap meta at 25% 2026-08-01 14:12:07 +02:00
CB-WP-0007-session-shape.md CB-WP-0010-T01: close CB-WP-0007 2026-08-02 02:15:08 +02:00
CB-WP-0008-ship-stage-0.md CB-WP-0008-T04: CB-EV-0007 — stage 0 is shipped, and the spend curve is a V 2026-08-01 15:25:19 +02:00
CB-WP-0009-adaptive-gates.md CB-WP-0009-T04: CB-EV-0008 — the gate changes, measured 2026-08-01 15:45:08 +02:00
CB-WP-0010-consolidation.md CB-WP-0010-T03: record CommitWindow's second failed second-use 2026-08-02 02:19:23 +02:00
CB-WP-0011-inspectable-table.md CB-WP-0011-T03: evidence — the chaos roll fired, and it was right 2026-08-02 02:48:51 +02:00
CB-WP-0012-render-port.md Sync hub IDs and work-record index for CB-WP-0012 2026-08-02 04:30:45 +02:00
CB-WP-0013-instrument-corrections.md Sync hub IDs and work-record index for CB-WP-0013 2026-08-02 07:32:33 +02:00
CB-WP-0014-execute-the-javascript.md Sync hub IDs and work-record index for CB-WP-0014 2026-08-02 07:59:28 +02:00
CB-WP-0015-the-inert-clauses.md Sync hub IDs and work-record index for CB-WP-0015 2026-08-02 14:08:30 +02:00
CB-WP-0016-the-drop-target.md Sync hub IDs and work-record index for CB-WP-0016 2026-08-02 20:49:21 +02:00
CB-WP-0017-legible-interaction.md Sync hub IDs and work-record index for CB-WP-0017 2026-08-02 22:43:40 +02:00
CB-WP-0018-the-browser-is-a-client.md Sync hub IDs and work-record index for CB-WP-0018 2026-08-03 02:24:57 +02:00
CB-WP-0019-the-am4-family.md CB-WP-0019 T03/T04: the cost rule written down, and the lifecycle 2026-08-03 19:25:18 +02:00
CB-WP-0020-the-table-you-can-read.md Sync hub state for CB-WP-0020 2026-08-03 20:21:18 +02:00
CB-WP-0021-import-the-edition.md CB-WP-0021 T04: evidence -- the endings are tight, and the cost chain snapped 2026-08-05 15:56:31 +02:00
CB-WP-0022-the-design-instrument.md CB-WP-0022 T05/T06/T07: the register, and what its first run found 2026-08-05 15:22:33 +02:00
CB-WP-0023-solve-legality.md CB-EV-0020: a gate moved the rule, and the report was wrong 2026-08-04 00:21:21 +02:00
CB-WP-0024-the-table-you-can-watch.md CB-WP-0024: the table you can watch 2026-08-05 17:32:48 +02:00
CB-WP-0025-could-we-have-won.md CB-WP-0025 T02: the review withdrew the finding, and the fifth wrong 2026-08-05 18:06:03 +02:00
CB-WP-0026-collect-the-rulings.md Sync hub state for CB-WP-0026 2026-08-05 16:15:22 +02:00