GROUND-WP-0005 has both tasks blocked on a measured difficulty baseline.
This is that baseline, with its confound stated rather than hidden.
It opens with a retraction. Earlier today clay-borg was going to report
that a greedy bot wins 200 of 200 games at five and six seats and the game
is too easy there -- which would have invited threshold changes. A
FirstLegal policy scores 0% on the identical deals where greedy scores
100%, and at two seats it beats greedy. Two unsophisticated agents span
the entire range, so a single policy's win rate is a statement about the
policy. clay-borg's adversarial review caught it before it left the repo.
Fifth wrong premise avoided, second stopped before sending.
What is reported instead is the WINNABLE FRACTION -- the proportion of
deals in which an exhaustive search finds a winning line in the final
round -- alongside a plural policy panel and the SPREAD between policies.
The spread is 71.7 to 100.0 percentage points, which is the direct
evidence for why the single-policy figure was meaningless.
Stated limits: the winnable figure is conditioned on greedy's play up to
the final round (searching from round 1 exceeded 2x10^6 nodes at two
seats), it is a lower bound, and budget-cut deals are excluded rather than
counted as losses.
No threshold changes are proposed. The confound is large enough that
clay-borg would rather this repo saw it than acted on it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>