CB-WP-0049 T02/T03: a seat that plays its objective, and F27 splits in two
Some checks failed
ci / check (push) Has been cancelled
Some checks failed
ci / check (push) Has been cancelled
objective() reads GroundState::score (now public) rather than restating
what winning is; a copy in the bot would disagree with the kernel the
first time ground-game rules on F28.
Working out WHERE the modes can differ was most of the task and it
bounds the result: SOLVE always claims for the actor, so own-score and
group-score want the same SOLVE nearly everywhere. That is a fact about
GROUND's action set, not a shortcoming of the bot. Two real divergences,
both readable off the table: SUPPORT regulates someone else (worth less
against a rival, worth MORE under coalitions where a Bond merges them
into my side), and SOLVE's value is the card's value, which greedy
ignores entirely.
THE RESULT — F27 splits in two:
group success UNCHANGED in 34 of 36 cells
who wins MOVES: BONDED COALITIONS at 4p goes 2.04 -> 2.98,
2.12 -> 3.29, 2.05 -> 3.01 winning seats per game
So "the competitive modes are scoring lenses over cooperative play" was
too strong and is withdrawn. The sharper claim: GROUND's scoring modes
change WHO WINS, not WHETHER THE GROUP SUCCEEDS. And the effect is
seat-band dependent -- 2p none, 4p largest, 6p none under coalitions;
two relation slots capping network growth is a candidate explanation and
is untested.
The panel now prints BOTH policies side by side. That was a correction
mid-task: the first version printed only the new one and I compared it
against a figure remembered from CB-WP-0047 -- a comparison against a
board nobody re-ran.
Control that makes the numbers mean anything: under SHARED GROUND the
two policies agree at all but <=2 decision points across 12 boards, so a
moving column is mode-awareness and not simply a different bot.
Also: two T01 tests keyed on `status: proposed`, which ground-game
renamed to `ready-for-implement` mid-session. They now find the module
by asking resolve() -- the structural property is ours and does not move
when another repo edits its vocabulary.
Also: `make vendor` replaces three hand re-vendors with a tool that
regenerates digests by walking editions/, and reports one-sided files
rather than resolving them.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
82b9e7df31
commit
3045eb03f8
16 changed files with 859 additions and 94 deletions
|
|
@ -54,7 +54,7 @@ kinds, states and metrics: [`GameDesign.md`](GameDesign.md). Reported by
|
|||
| F23 | inconsistent | applied | decisions/ADR-0017-chaos-window-2-verdict.md | counterexample | 2026-08-07 | clay-borg |
|
||||
| F22 | underdetermined | withdrawn | games_ground::edition::supply_tests::play_never_exceeds_the_components_the_box_holds | counterexample | 2026-08-07 | clay-borg |
|
||||
| F26 | inert | raised | `crates/cb-render-html/src/lib.rs::a_scoped_table_says_what_a_scope_does` | counterexample | 2026-08-08 | ground-game |
|
||||
| F27 | unplayed | raised | `games/ground/examples/scenario-panel.rs` | counterexample | 2026-08-08 | clay-borg |
|
||||
| F27 | unplayed | reported | `games/ground/examples/scenario-panel.rs` | counterexample | 2026-08-08 | clay-borg |
|
||||
| F28 | underdetermined | raised | `games_ground::tests::the_mastery_rating_and_the_shared_score_count_different_things` | counterexample | 2026-08-08 | ground-game |
|
||||
|
||||
<!-- design-register:end -->
|
||||
|
|
@ -70,6 +70,35 @@ kinds, states and metrics: [`GameDesign.md`](GameDesign.md). Reported by
|
|||
score. `unplayed` rather than `inert`: the modes score correctly, they
|
||||
have simply never faced a seat that wanted to win alone.
|
||||
|
||||
**Re-measured 2026-08-08 with `ObjectivePolicy`** (CB-WP-0049 T03),
|
||||
which plays its seat's own objective. The finding **splits in two**:
|
||||
|
||||
- **Group success does not move.** 34 of 36 cells identical; SCN_03 at
|
||||
4p goes 99→100 in all three modes, which is the SOLVE-by-value
|
||||
refinement and not a mode effect. **Whether the table survives is the
|
||||
same game in all three modes.**
|
||||
- **Who wins does move.** Under BONDED COALITIONS at 4p, winning seats
|
||||
per game go 2.04 → 2.98 (SCN_01/02), 2.12 → 3.29 (SCN_03), 2.05 →
|
||||
3.01 (SCN_04): a seat whose score is its coalition's sum bonds more,
|
||||
and the coalitions get about half again as large. COMMON PROBLEM
|
||||
moves at 6p (1.10 → 1.17).
|
||||
|
||||
So the modes **do** reach decisions — the original claim that they are
|
||||
"scoring lenses over cooperative play" was too strong, and is withdrawn
|
||||
in favour of the sharper one: **GROUND's scoring modes change who wins,
|
||||
not whether the group succeeds.**
|
||||
|
||||
**Sensitivity:** vary only the seat band and the effect appears and
|
||||
disappears — 2p shows no divergence in any mode, 4p shows the largest,
|
||||
6p shows none under BONDED COALITIONS. A candidate explanation is that
|
||||
two relation slots per seat cap network growth, so at 6p the incentive
|
||||
exists and cannot be acted on; **that is untested** and is the next
|
||||
thing to vary.
|
||||
|
||||
**Still open**, because the part that motivated it is untested: no
|
||||
policy models a *rival* playing their objective, so a competitive mode
|
||||
in which nobody anticipates an opponent remains a weak test of that
|
||||
mode (CB-WP-0049 "Not done here").
|
||||
- **F28 — SHARED GROUND's mastery rating counts cards where the mode
|
||||
card counts points.** *"All claimed Problem cards form one shared score.
|
||||
… For a mastery rating, subtract 1 for each Blame token still in play
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue