clay-borg/evidence/CB-EV-0033-boards-and-modes.md
tegwick 704b99975b
Some checks failed
ci / check (push) Failing after 4s
Apply ground-game's rulings: mastery in points, four boards, and a
vendor tool that covers what the gate checks

They ruled on all seven items the same day. Two were actionable here.

F28 RULED: points. Modes.csv MODE_COOP clarified upstream to say
"penalties apply to points, not card count"; mastery is now
total - blame - denied. A recorded scenario went red on it --
gr-e02-shared-ground pinned 0 (2 claimed CARDS - 1 - 1) and now expects
2 (4 POINTS - 1 - 1). The number moved because the rule was decided, not
because the engine drifted, and the scenario records both rulings; its
schema has no field for a second one, so both live in ruled_note with
`ruled` carrying the LATEST date.

F29 RULED not-intended and APPLIED upstream: SCN_02's suits re-tuned the
same day. The characterisation test is how we found out -- it pinned the
duplication, went red on the re-tune, and that red WAS the notification.
It now asserts every pair distinct, the stronger statement the
duplication had made unavailable. SCN_02 re-measures at 73 at 2p, not
67: its own board now.

F26/F30 ruled and recorded. F30's ruling incidentally confirms our
reading -- they name priority-2's suit as the first lever, which is the
difference we identified without having measured causation.

vendor-editions grew twice, both times because it covered less than the
gate it exists to satisfy:

  - It refused to touch ground-darvo-r0/ on the reasoning that the
    baseline is "a separate record". That was wrong within the hour:
    ground-game clarified Modes.csv and `make vendor` reported a clean
    sync while edition-check went red. A sync tool that covers less than
    its check reports success into a red gate.
  - Its two-block rewrite DETECTED which fence held which set and
    preserved the arrangement -- faithfully preserving a swap an earlier
    write had introduced, leaving each fence under a heading describing
    the other. edition-check reads every sha256 line flat and passed
    throughout: a document can be self-consistently wrong and green.
    Order is now asserted, with a control that goes red on a swap.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-09 00:15:14 +02:00

150 lines
6.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# CB-EV-0033 — the four boards and the three modes, measured
date: 2026-08-08
produced by: [CB-WP-0047](../workplans/CB-WP-0047-all-four-boards-and-all-three-modes.md),
[CB-WP-0049](../workplans/CB-WP-0049-a-seat-that-plays-its-own-objective.md)
instrument: `games/ground/examples/scenario-panel.rs` (`make panels`)
reported to ground-game as: `reports/260808-clay-borg-boards-and-modes.md`
## What was run
4 scenarios × 3 modes × 3 seat bands × 2 policies, **100 games per cell**,
seeds 0..100, every seat filled by the same bot.
| policy | what it is |
|---|---|
| `greedy` | `GreedyPolicy`, which **never reads `state.mode`** |
| `objective` | `ObjectivePolicy`, which plays its seat's own objective, read off `GroundState::score` |
**Every cell asserts `games == 100` and `played == 100`** — a game that
ran is not a game that was played (CB-REV-0003 #2), and a short column is
a sample nobody chose (CB-REV-0002 #1).
**Both policies are proven blind** to what their seat cannot see
([ADR-0023](../decisions/ADR-0023-a-policy-is-bound-by-what-its-seat-can-see.md)):
vary only face-down Problems and other seats' hands, and neither changes
its move. The control is proven against a deliberate peeker first.
## 0 — what happened after this was reported
**ground-game ruled on all seven items the same day** (GROUND-RPT-0006
§"disposition"). Two changed the engine or the edition, and the numbers
below are therefore **superseded where noted**:
| item | ruling |
|---|---|
| SCN_01 ≡ SCN_02 | **not intended** — SCN_02 re-tuned upstream, §1 below is now history |
| SCN_04 harder at 2p | acceptable variety; first lever would be priority-2's suit, **not** thresholds — which confirms our reading of the operative difference without our having measured it |
| 6p formality | known (F17 line) |
| modes change who wins | accepted as a design property, not a defect |
| 2 relation slots | **intentional cap**; a third slot stays an extension, so our proposed sensitivity run would measure an extension and not the shipped game |
| mastery: points or cards | **RULED: points.** Applied — see below |
| package file manifest | accepted process ask, tracked as GROUND-WP-0008 |
**Mastery is now `total blame denied`.** `Modes.csv` was clarified
upstream to say *"penalties apply to points, not card count"* and the
engine follows. `gr-e02-shared-ground` went red on it — it pinned 0
(2 claimed cards 1 1) and now expects 2 (4 points 1 1). The
number moved because the rule was decided, not because the engine
drifted, and the scenario records both rulings.
**SCN_02 re-measured after the re-tune:** 2p group success **73**, not 67
— it is now its own board and matches SCN_03's difficulty rather than
SCN_01's. §1 below describes the state that prompted the report.
## 1 — SCN_01 and SCN_02 were the same board (superseded)
Identical suit and value at **every** priority. Every cell matches
exactly, in both policies, at all three seat bands.
Not a defect — a reskin is a legitimate design choice — but **"four
scenarios" buys three boards**, and a panel that treated them as four
independent samples would be counting one of them twice. Pinned by
`which_scenarios_are_mechanically_distinct`, so a future divergence is a
decision rather than a drift.
## 2 — SCN_04 is the hard board at 2 players
| board | 2p group success (greedy, /100) |
|---|---|
| SCN_01 / SCN_02 | 67 |
| SCN_03 | 73 |
| **SCN_04** | **52** |
**The 2p deals are identical in every respect except one.** All four
scenarios deal Surface + priorities 12, all four give values 2+2+2 = 6
available against a threshold of 5, and all four start at Stress 2 over
five rounds. The single difference is the **suit multiset**:
| board | 2p suits |
|---|---|
| SCN_01 / SCN_02 | Repair, Clarify, Boundary |
| SCN_03 | Boundary, Clarify, Repair |
| **SCN_04** | **Repair, Clarify, Repair** |
SCN_04 is the only deck that needs **two of one suit** in the 2p deal.
Because every other parameter is held fixed by the edition itself, this
is close to a controlled comparison — but **the causal claim is not
measured**: nothing here demonstrates that the second Repair is what
costs the 15 points. The falsifier is a deck with a doubled suit that
does *not* lose group success.
## 3 — every board is a formality at 6 players
100/100 group success in all twelve 6p cells, both policies, every mode.
Consistent with F17's shape. Reported, not acted on.
## 4 — the modes change who wins, not whether the group succeeds
`greedy` vs `objective`, group success:
**Unchanged in 34 of 36 cells.** SCN_03 at 4p moves 99 → 100 in all three
modes, which is the SOLVE-by-card-value refinement and not a mode effect
(it moves under SHARED GROUND too).
Winning **seats** per game, BONDED COALITIONS:
| board | 4p greedy | 4p objective |
|---|---|---|
| SCN_01 / SCN_02 | 2.04 | **2.98** |
| SCN_03 | 2.12 | **3.29** |
| SCN_04 | 2.05 | **3.01** |
COMMON PROBLEM at 6p: 1.10 → 1.17. Everything else flat.
**The control that makes this mean anything:** under SHARED GROUND the
two policies agree at all but ≤2 decision points across twelve boards. A
moving column is therefore mode-awareness and not simply a different bot.
### Where the modes *can* differ
**SOLVE always claims for the actor**, so a seat maximising its own score
and one maximising the group's want the same SOLVE in nearly every
position. That is a fact about GROUND's action set and it bounds how far
apart any two policies can get. The two real divergences:
- **SUPPORT regulates someone else** — worth less when the beneficiary is
a rival, worth *more* when a Bond merges them into my coalition and my
score is the coalition's sum;
- **SOLVE's value is the card's value**, which greedy ignores entirely.
### The seat-band pattern, and one untested explanation
2p: nothing moves in any mode. 4p: the largest effect. 6p: **nothing**
under BONDED COALITIONS.
**Sensitivity:** vary only the seat band and the effect appears and
disappears with every other parameter held. A candidate explanation is
that **two relation slots per seat cap network growth**, so at 6p the
incentive exists and cannot be acted on. **Untested.** The falsifier is a
run with a third relation slot: if coalition size at 6p then moves the
way it does at 4p, the cap is the cause.
## What this does not establish
- **No policy models a rival playing their objective.** A competitive
mode in which nobody anticipates an opponent is a weak test of that
mode. F27 stays `reported`, not resolved.
- **No felt play.** Every number here is bots.
- **Bond-scoped joint SOLVE remains untested** (CB-EV-0032 criterion 4,
unchanged): these policies read the mode, not the module.