Some checks failed
ci / check (push) Failing after 4s
vendor tool that covers what the gate checks
They ruled on all seven items the same day. Two were actionable here.
F28 RULED: points. Modes.csv MODE_COOP clarified upstream to say
"penalties apply to points, not card count"; mastery is now
total - blame - denied. A recorded scenario went red on it --
gr-e02-shared-ground pinned 0 (2 claimed CARDS - 1 - 1) and now expects
2 (4 POINTS - 1 - 1). The number moved because the rule was decided, not
because the engine drifted, and the scenario records both rulings; its
schema has no field for a second one, so both live in ruled_note with
`ruled` carrying the LATEST date.
F29 RULED not-intended and APPLIED upstream: SCN_02's suits re-tuned the
same day. The characterisation test is how we found out -- it pinned the
duplication, went red on the re-tune, and that red WAS the notification.
It now asserts every pair distinct, the stronger statement the
duplication had made unavailable. SCN_02 re-measures at 73 at 2p, not
67: its own board now.
F26/F30 ruled and recorded. F30's ruling incidentally confirms our
reading -- they name priority-2's suit as the first lever, which is the
difference we identified without having measured causation.
vendor-editions grew twice, both times because it covered less than the
gate it exists to satisfy:
- It refused to touch ground-darvo-r0/ on the reasoning that the
baseline is "a separate record". That was wrong within the hour:
ground-game clarified Modes.csv and `make vendor` reported a clean
sync while edition-check went red. A sync tool that covers less than
its check reports success into a red gate.
- Its two-block rewrite DETECTED which fence held which set and
preserved the arrangement -- faithfully preserving a swap an earlier
write had introduced, leaving each fence under a heading describing
the other. edition-check reads every sha256 line flat and passed
throughout: a document can be self-consistently wrong and green.
Order is now asserted, with a control that goes red on a swap.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
150 lines
6.6 KiB
Markdown
150 lines
6.6 KiB
Markdown
# CB-EV-0033 — the four boards and the three modes, measured
|
||
|
||
date: 2026-08-08
|
||
produced by: [CB-WP-0047](../workplans/CB-WP-0047-all-four-boards-and-all-three-modes.md),
|
||
[CB-WP-0049](../workplans/CB-WP-0049-a-seat-that-plays-its-own-objective.md)
|
||
instrument: `games/ground/examples/scenario-panel.rs` (`make panels`)
|
||
reported to ground-game as: `reports/260808-clay-borg-boards-and-modes.md`
|
||
|
||
## What was run
|
||
|
||
4 scenarios × 3 modes × 3 seat bands × 2 policies, **100 games per cell**,
|
||
seeds 0..100, every seat filled by the same bot.
|
||
|
||
| policy | what it is |
|
||
|---|---|
|
||
| `greedy` | `GreedyPolicy`, which **never reads `state.mode`** |
|
||
| `objective` | `ObjectivePolicy`, which plays its seat's own objective, read off `GroundState::score` |
|
||
|
||
**Every cell asserts `games == 100` and `played == 100`** — a game that
|
||
ran is not a game that was played (CB-REV-0003 #2), and a short column is
|
||
a sample nobody chose (CB-REV-0002 #1).
|
||
|
||
**Both policies are proven blind** to what their seat cannot see
|
||
([ADR-0023](../decisions/ADR-0023-a-policy-is-bound-by-what-its-seat-can-see.md)):
|
||
vary only face-down Problems and other seats' hands, and neither changes
|
||
its move. The control is proven against a deliberate peeker first.
|
||
|
||
## 0 — what happened after this was reported
|
||
|
||
**ground-game ruled on all seven items the same day** (GROUND-RPT-0006
|
||
§"disposition"). Two changed the engine or the edition, and the numbers
|
||
below are therefore **superseded where noted**:
|
||
|
||
| item | ruling |
|
||
|---|---|
|
||
| SCN_01 ≡ SCN_02 | **not intended** — SCN_02 re-tuned upstream, §1 below is now history |
|
||
| SCN_04 harder at 2p | acceptable variety; first lever would be priority-2's suit, **not** thresholds — which confirms our reading of the operative difference without our having measured it |
|
||
| 6p formality | known (F17 line) |
|
||
| modes change who wins | accepted as a design property, not a defect |
|
||
| 2 relation slots | **intentional cap**; a third slot stays an extension, so our proposed sensitivity run would measure an extension and not the shipped game |
|
||
| mastery: points or cards | **RULED: points.** Applied — see below |
|
||
| package file manifest | accepted process ask, tracked as GROUND-WP-0008 |
|
||
|
||
**Mastery is now `total − blame − denied`.** `Modes.csv` was clarified
|
||
upstream to say *"penalties apply to points, not card count"* and the
|
||
engine follows. `gr-e02-shared-ground` went red on it — it pinned 0
|
||
(2 claimed cards − 1 − 1) and now expects 2 (4 points − 1 − 1). The
|
||
number moved because the rule was decided, not because the engine
|
||
drifted, and the scenario records both rulings.
|
||
|
||
**SCN_02 re-measured after the re-tune:** 2p group success **73**, not 67
|
||
— it is now its own board and matches SCN_03's difficulty rather than
|
||
SCN_01's. §1 below describes the state that prompted the report.
|
||
|
||
## 1 — SCN_01 and SCN_02 were the same board (superseded)
|
||
|
||
Identical suit and value at **every** priority. Every cell matches
|
||
exactly, in both policies, at all three seat bands.
|
||
|
||
Not a defect — a reskin is a legitimate design choice — but **"four
|
||
scenarios" buys three boards**, and a panel that treated them as four
|
||
independent samples would be counting one of them twice. Pinned by
|
||
`which_scenarios_are_mechanically_distinct`, so a future divergence is a
|
||
decision rather than a drift.
|
||
|
||
## 2 — SCN_04 is the hard board at 2 players
|
||
|
||
| board | 2p group success (greedy, /100) |
|
||
|---|---|
|
||
| SCN_01 / SCN_02 | 67 |
|
||
| SCN_03 | 73 |
|
||
| **SCN_04** | **52** |
|
||
|
||
**The 2p deals are identical in every respect except one.** All four
|
||
scenarios deal Surface + priorities 1–2, all four give values 2+2+2 = 6
|
||
available against a threshold of 5, and all four start at Stress 2 over
|
||
five rounds. The single difference is the **suit multiset**:
|
||
|
||
| board | 2p suits |
|
||
|---|---|
|
||
| SCN_01 / SCN_02 | Repair, Clarify, Boundary |
|
||
| SCN_03 | Boundary, Clarify, Repair |
|
||
| **SCN_04** | **Repair, Clarify, Repair** |
|
||
|
||
SCN_04 is the only deck that needs **two of one suit** in the 2p deal.
|
||
Because every other parameter is held fixed by the edition itself, this
|
||
is close to a controlled comparison — but **the causal claim is not
|
||
measured**: nothing here demonstrates that the second Repair is what
|
||
costs the 15 points. The falsifier is a deck with a doubled suit that
|
||
does *not* lose group success.
|
||
|
||
## 3 — every board is a formality at 6 players
|
||
|
||
100/100 group success in all twelve 6p cells, both policies, every mode.
|
||
Consistent with F17's shape. Reported, not acted on.
|
||
|
||
## 4 — the modes change who wins, not whether the group succeeds
|
||
|
||
`greedy` vs `objective`, group success:
|
||
|
||
**Unchanged in 34 of 36 cells.** SCN_03 at 4p moves 99 → 100 in all three
|
||
modes, which is the SOLVE-by-card-value refinement and not a mode effect
|
||
(it moves under SHARED GROUND too).
|
||
|
||
Winning **seats** per game, BONDED COALITIONS:
|
||
|
||
| board | 4p greedy | 4p objective |
|
||
|---|---|---|
|
||
| SCN_01 / SCN_02 | 2.04 | **2.98** |
|
||
| SCN_03 | 2.12 | **3.29** |
|
||
| SCN_04 | 2.05 | **3.01** |
|
||
|
||
COMMON PROBLEM at 6p: 1.10 → 1.17. Everything else flat.
|
||
|
||
**The control that makes this mean anything:** under SHARED GROUND the
|
||
two policies agree at all but ≤2 decision points across twelve boards. A
|
||
moving column is therefore mode-awareness and not simply a different bot.
|
||
|
||
### Where the modes *can* differ
|
||
|
||
**SOLVE always claims for the actor**, so a seat maximising its own score
|
||
and one maximising the group's want the same SOLVE in nearly every
|
||
position. That is a fact about GROUND's action set and it bounds how far
|
||
apart any two policies can get. The two real divergences:
|
||
|
||
- **SUPPORT regulates someone else** — worth less when the beneficiary is
|
||
a rival, worth *more* when a Bond merges them into my coalition and my
|
||
score is the coalition's sum;
|
||
- **SOLVE's value is the card's value**, which greedy ignores entirely.
|
||
|
||
### The seat-band pattern, and one untested explanation
|
||
|
||
2p: nothing moves in any mode. 4p: the largest effect. 6p: **nothing**
|
||
under BONDED COALITIONS.
|
||
|
||
**Sensitivity:** vary only the seat band and the effect appears and
|
||
disappears with every other parameter held. A candidate explanation is
|
||
that **two relation slots per seat cap network growth**, so at 6p the
|
||
incentive exists and cannot be acted on. **Untested.** The falsifier is a
|
||
run with a third relation slot: if coalition size at 6p then moves the
|
||
way it does at 4p, the cap is the cause.
|
||
|
||
## What this does not establish
|
||
|
||
- **No policy models a rival playing their objective.** A competitive
|
||
mode in which nobody anticipates an opponent is a weak test of that
|
||
mode. F27 stays `reported`, not resolved.
|
||
- **No felt play.** Every number here is bots.
|
||
- **Bond-scoped joint SOLVE remains untested** (CB-EV-0032 criterion 4,
|
||
unchanged): these policies read the mode, not the module.
|