clay-borg/evidence/CB-EV-0033-boards-and-modes.md
tegwick 704b99975b
Some checks failed
ci / check (push) Failing after 4s
Apply ground-game's rulings: mastery in points, four boards, and a
vendor tool that covers what the gate checks

They ruled on all seven items the same day. Two were actionable here.

F28 RULED: points. Modes.csv MODE_COOP clarified upstream to say
"penalties apply to points, not card count"; mastery is now
total - blame - denied. A recorded scenario went red on it --
gr-e02-shared-ground pinned 0 (2 claimed CARDS - 1 - 1) and now expects
2 (4 POINTS - 1 - 1). The number moved because the rule was decided, not
because the engine drifted, and the scenario records both rulings; its
schema has no field for a second one, so both live in ruled_note with
`ruled` carrying the LATEST date.

F29 RULED not-intended and APPLIED upstream: SCN_02's suits re-tuned the
same day. The characterisation test is how we found out -- it pinned the
duplication, went red on the re-tune, and that red WAS the notification.
It now asserts every pair distinct, the stronger statement the
duplication had made unavailable. SCN_02 re-measures at 73 at 2p, not
67: its own board now.

F26/F30 ruled and recorded. F30's ruling incidentally confirms our
reading -- they name priority-2's suit as the first lever, which is the
difference we identified without having measured causation.

vendor-editions grew twice, both times because it covered less than the
gate it exists to satisfy:

  - It refused to touch ground-darvo-r0/ on the reasoning that the
    baseline is "a separate record". That was wrong within the hour:
    ground-game clarified Modes.csv and `make vendor` reported a clean
    sync while edition-check went red. A sync tool that covers less than
    its check reports success into a red gate.
  - Its two-block rewrite DETECTED which fence held which set and
    preserved the arrangement -- faithfully preserving a swap an earlier
    write had introduced, leaving each fence under a heading describing
    the other. edition-check reads every sha256 line flat and passed
    throughout: a document can be self-consistently wrong and green.
    Order is now asserted, with a control that goes red on a swap.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-09 00:15:14 +02:00

6.6 KiB
Raw Permalink Blame History

CB-EV-0033 — the four boards and the three modes, measured

date: 2026-08-08 produced by: CB-WP-0047, CB-WP-0049 instrument: games/ground/examples/scenario-panel.rs (make panels) reported to ground-game as: reports/260808-clay-borg-boards-and-modes.md

What was run

4 scenarios × 3 modes × 3 seat bands × 2 policies, 100 games per cell, seeds 0..100, every seat filled by the same bot.

policy what it is
greedy GreedyPolicy, which never reads state.mode
objective ObjectivePolicy, which plays its seat's own objective, read off GroundState::score

Every cell asserts games == 100 and played == 100 — a game that ran is not a game that was played (CB-REV-0003 #2), and a short column is a sample nobody chose (CB-REV-0002 #1).

Both policies are proven blind to what their seat cannot see (ADR-0023): vary only face-down Problems and other seats' hands, and neither changes its move. The control is proven against a deliberate peeker first.

0 — what happened after this was reported

ground-game ruled on all seven items the same day (GROUND-RPT-0006 §"disposition"). Two changed the engine or the edition, and the numbers below are therefore superseded where noted:

item ruling
SCN_01 ≡ SCN_02 not intended — SCN_02 re-tuned upstream, §1 below is now history
SCN_04 harder at 2p acceptable variety; first lever would be priority-2's suit, not thresholds — which confirms our reading of the operative difference without our having measured it
6p formality known (F17 line)
modes change who wins accepted as a design property, not a defect
2 relation slots intentional cap; a third slot stays an extension, so our proposed sensitivity run would measure an extension and not the shipped game
mastery: points or cards RULED: points. Applied — see below
package file manifest accepted process ask, tracked as GROUND-WP-0008

Mastery is now total blame denied. Modes.csv was clarified upstream to say "penalties apply to points, not card count" and the engine follows. gr-e02-shared-ground went red on it — it pinned 0 (2 claimed cards 1 1) and now expects 2 (4 points 1 1). The number moved because the rule was decided, not because the engine drifted, and the scenario records both rulings.

SCN_02 re-measured after the re-tune: 2p group success 73, not 67 — it is now its own board and matches SCN_03's difficulty rather than SCN_01's. §1 below describes the state that prompted the report.

1 — SCN_01 and SCN_02 were the same board (superseded)

Identical suit and value at every priority. Every cell matches exactly, in both policies, at all three seat bands.

Not a defect — a reskin is a legitimate design choice — but "four scenarios" buys three boards, and a panel that treated them as four independent samples would be counting one of them twice. Pinned by which_scenarios_are_mechanically_distinct, so a future divergence is a decision rather than a drift.

2 — SCN_04 is the hard board at 2 players

board 2p group success (greedy, /100)
SCN_01 / SCN_02 67
SCN_03 73
SCN_04 52

The 2p deals are identical in every respect except one. All four scenarios deal Surface + priorities 12, all four give values 2+2+2 = 6 available against a threshold of 5, and all four start at Stress 2 over five rounds. The single difference is the suit multiset:

board 2p suits
SCN_01 / SCN_02 Repair, Clarify, Boundary
SCN_03 Boundary, Clarify, Repair
SCN_04 Repair, Clarify, Repair

SCN_04 is the only deck that needs two of one suit in the 2p deal. Because every other parameter is held fixed by the edition itself, this is close to a controlled comparison — but the causal claim is not measured: nothing here demonstrates that the second Repair is what costs the 15 points. The falsifier is a deck with a doubled suit that does not lose group success.

3 — every board is a formality at 6 players

100/100 group success in all twelve 6p cells, both policies, every mode. Consistent with F17's shape. Reported, not acted on.

4 — the modes change who wins, not whether the group succeeds

greedy vs objective, group success:

Unchanged in 34 of 36 cells. SCN_03 at 4p moves 99 → 100 in all three modes, which is the SOLVE-by-card-value refinement and not a mode effect (it moves under SHARED GROUND too).

Winning seats per game, BONDED COALITIONS:

board 4p greedy 4p objective
SCN_01 / SCN_02 2.04 2.98
SCN_03 2.12 3.29
SCN_04 2.05 3.01

COMMON PROBLEM at 6p: 1.10 → 1.17. Everything else flat.

The control that makes this mean anything: under SHARED GROUND the two policies agree at all but ≤2 decision points across twelve boards. A moving column is therefore mode-awareness and not simply a different bot.

Where the modes can differ

SOLVE always claims for the actor, so a seat maximising its own score and one maximising the group's want the same SOLVE in nearly every position. That is a fact about GROUND's action set and it bounds how far apart any two policies can get. The two real divergences:

  • SUPPORT regulates someone else — worth less when the beneficiary is a rival, worth more when a Bond merges them into my coalition and my score is the coalition's sum;
  • SOLVE's value is the card's value, which greedy ignores entirely.

The seat-band pattern, and one untested explanation

2p: nothing moves in any mode. 4p: the largest effect. 6p: nothing under BONDED COALITIONS.

Sensitivity: vary only the seat band and the effect appears and disappears with every other parameter held. A candidate explanation is that two relation slots per seat cap network growth, so at 6p the incentive exists and cannot be acted on. Untested. The falsifier is a run with a third relation slot: if coalition size at 6p then moves the way it does at 4p, the cap is the cause.

What this does not establish

  • No policy models a rival playing their objective. A competitive mode in which nobody anticipates an opponent is a weak test of that mode. F27 stays reported, not resolved.
  • No felt play. Every number here is bots.
  • Bond-scoped joint SOLVE remains untested (CB-EV-0032 criterion 4, unchanged): these policies read the mode, not the module.