2026-08-02 02:16:51 +02:00
|
|
|
|
scenario: ground/gr-e03-common-problem
|
|
|
|
|
|
description: >
|
|
|
|
|
|
COMMON PROBLEM scoring (GR-E03), the only scoring mode with no scenario
|
|
|
|
|
|
until CB-WP-0010 T02 — implemented since CB-WP-0001 and referenced by
|
|
|
|
|
|
nothing. Five players, where the threshold is actually reachable (10
|
|
|
|
|
|
available against 9). The group qualifies, personal scores are claimed
|
|
|
|
|
|
value −1 per Blame held, and the Blame is what decides the winner: P4
|
|
|
|
|
|
claimed the highest-value Problem and still loses to P3 because two
|
|
|
|
|
|
Blame tokens sit in front of them.
|
|
|
|
|
|
covers: [GR-E03, GR-E01, GR-R09, GR-T02]
|
|
|
|
|
|
seed: 42
|
|
|
|
|
|
setup:
|
|
|
|
|
|
players: 5
|
|
|
|
|
|
preset: standard-5p
|
|
|
|
|
|
patch:
|
|
|
|
|
|
"round": 5
|
|
|
|
|
|
"mode": CommonProblem
|
CB-WP-0021 T01/T02/T05: the engine plays its own data — AM-7 blocks
ADR-0011 decided it: vendor the CSV with a checked digest, read it with
a ~50-line reader, and let the hashes move.
The declaration's constraint was measured against the WRONG BUDGET. It
said a CSV crate costs 21,613 against AM-4a's 3,798 of headroom, '5.7x
over, settled by measurement'. But setup and problem_priorities are
cfg(scenarios) and are not in the shipped runtime at all, so AM-4a never
sees them. Against AM-4b, csv costs 17,651 against 19,742 -- it FITS,
with 2,091 to spare. It is refused anyway, on proportion: 89% of the
budget's remaining capacity to read 20 rows. The revisit condition is
stated (nested quoting, embedded newlines, multiple dialects).
GR-S01 now deals Surface + hidden 1..=k as ruled, with edition values and
suits. Measured: 6/9/12 available against thresholds 5/7/9 -- the game is
winnable at every seat count, which is what the maintainer could not do.
gd0001 is INVERTED, not deleted, and now also asserts the 6/9/12 so a
deal that is reachable for the wrong reason still fails.
Blast radius was scenario expectations, exactly as the ADR predicted: no
scenario pinned a hash and no bundle is committed. Six scenarios and two
unit tests updated, each with a note. gr-e01-threshold-unreachable-2p is
RENAMED to -reachable- and rewritten as the non-provisional import check
ground-game asked for by name. gr-e03's setup was restructured, not just
renumbered: with values 2,2,2 its personal-edge test would have tied
three ways and asserted nothing.
BLOCKING: AM-7 fails at median 0.845 against its 0.9 floor. Isolated
across three runs -- 3 problems + stand-in 0.97, 3 problems + edition
0.909, 4 problems + edition 0.845. State is BOUNDED (proven: identical
after 5k and 100k events), so this is not the unbounded-growth defect
AM-7 exists to catch; it is a bigger working set streaming a long log.
Whether AM-7's floor is still right for a larger aggregate is a spec
question and lowering it requires an ADR, so it is not being tuned here.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 00:47:56 +02:00
|
|
|
|
# Four of the five dealt Problems claimed: 2+2+3+2 = 9, exactly the
|
|
|
|
|
|
# 5-6p threshold. P3 takes the 3-value Problem so the personal edge
|
|
|
|
|
|
# this scenario exists to test has a unique winner — with the edition
|
|
|
|
|
|
# values (2,2,2,3,3) an even spread would tie three ways and the test
|
|
|
|
|
|
# would assert nothing about GR-E03.
|
2026-08-02 02:16:51 +02:00
|
|
|
|
"problems.1.claimed_by": 0
|
|
|
|
|
|
"problems.2.claimed_by": 1
|
CB-WP-0021 T01/T02/T05: the engine plays its own data — AM-7 blocks
ADR-0011 decided it: vendor the CSV with a checked digest, read it with
a ~50-line reader, and let the hashes move.
The declaration's constraint was measured against the WRONG BUDGET. It
said a CSV crate costs 21,613 against AM-4a's 3,798 of headroom, '5.7x
over, settled by measurement'. But setup and problem_priorities are
cfg(scenarios) and are not in the shipped runtime at all, so AM-4a never
sees them. Against AM-4b, csv costs 17,651 against 19,742 -- it FITS,
with 2,091 to spare. It is refused anyway, on proportion: 89% of the
budget's remaining capacity to read 20 rows. The revisit condition is
stated (nested quoting, embedded newlines, multiple dialects).
GR-S01 now deals Surface + hidden 1..=k as ruled, with edition values and
suits. Measured: 6/9/12 available against thresholds 5/7/9 -- the game is
winnable at every seat count, which is what the maintainer could not do.
gd0001 is INVERTED, not deleted, and now also asserts the 6/9/12 so a
deal that is reachable for the wrong reason still fails.
Blast radius was scenario expectations, exactly as the ADR predicted: no
scenario pinned a hash and no bundle is committed. Six scenarios and two
unit tests updated, each with a note. gr-e01-threshold-unreachable-2p is
RENAMED to -reachable- and rewritten as the non-provisional import check
ground-game asked for by name. gr-e03's setup was restructured, not just
renumbered: with values 2,2,2 its personal-edge test would have tied
three ways and asserted nothing.
BLOCKING: AM-7 fails at median 0.845 against its 0.9 floor. Isolated
across three runs -- 3 problems + stand-in 0.97, 3 problems + edition
0.909, 4 problems + edition 0.845. State is BOUNDED (proven: identical
after 5k and 100k events), so this is not the unbounded-growth defect
AM-7 exists to catch; it is a bigger working set streaming a long log.
Whether AM-7's floor is still right for a larger aggregate is a spec
question and lowering it requires an ADR, so it is not being tuned here.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 00:47:56 +02:00
|
|
|
|
"problems.4.claimed_by": 2
|
|
|
|
|
|
"problems.3.claimed_by": 3
|
2026-08-02 02:16:51 +02:00
|
|
|
|
"problems.2.face_up": true
|
|
|
|
|
|
"problems.3.face_up": true
|
|
|
|
|
|
"problems.4.face_up": true
|
|
|
|
|
|
# GR-T02: two Blame tokens in front of P4, −1 each.
|
|
|
|
|
|
"players.3.blame_from": [0, 1]
|
|
|
|
|
|
commands:
|
|
|
|
|
|
- actor: P1
|
|
|
|
|
|
cmd: select_action
|
|
|
|
|
|
args: { action: GROUND }
|
|
|
|
|
|
- actor: P2
|
|
|
|
|
|
cmd: select_action
|
|
|
|
|
|
args: { action: GROUND }
|
|
|
|
|
|
- actor: P3
|
|
|
|
|
|
cmd: select_action
|
|
|
|
|
|
args: { action: GROUND }
|
|
|
|
|
|
- actor: P4
|
|
|
|
|
|
cmd: select_action
|
|
|
|
|
|
args: { action: GROUND }
|
|
|
|
|
|
- actor: P5
|
|
|
|
|
|
cmd: select_action
|
|
|
|
|
|
args: { action: GROUND }
|
|
|
|
|
|
- actor: SYSTEM
|
|
|
|
|
|
cmd: reveal
|
|
|
|
|
|
- actor: P1
|
|
|
|
|
|
cmd: choose_ground_mode
|
|
|
|
|
|
args: { mode: GR }
|
|
|
|
|
|
- actor: P2
|
|
|
|
|
|
cmd: choose_ground_mode
|
|
|
|
|
|
args: { mode: GR }
|
|
|
|
|
|
- actor: P3
|
|
|
|
|
|
cmd: choose_ground_mode
|
|
|
|
|
|
args: { mode: GR }
|
|
|
|
|
|
- actor: P4
|
|
|
|
|
|
cmd: choose_ground_mode
|
|
|
|
|
|
args: { mode: GR }
|
|
|
|
|
|
- actor: P5
|
|
|
|
|
|
cmd: choose_ground_mode
|
|
|
|
|
|
args: { mode: GR }
|
|
|
|
|
|
- actor: SYSTEM
|
|
|
|
|
|
cmd: resolve
|
|
|
|
|
|
- actor: SYSTEM
|
|
|
|
|
|
cmd: end_round
|
|
|
|
|
|
expect:
|
|
|
|
|
|
events:
|
|
|
|
|
|
- kind: GameEnded
|
|
|
|
|
|
state:
|
CB-WP-0021 T01/T02/T05: the engine plays its own data — AM-7 blocks
ADR-0011 decided it: vendor the CSV with a checked digest, read it with
a ~50-line reader, and let the hashes move.
The declaration's constraint was measured against the WRONG BUDGET. It
said a CSV crate costs 21,613 against AM-4a's 3,798 of headroom, '5.7x
over, settled by measurement'. But setup and problem_priorities are
cfg(scenarios) and are not in the shipped runtime at all, so AM-4a never
sees them. Against AM-4b, csv costs 17,651 against 19,742 -- it FITS,
with 2,091 to spare. It is refused anyway, on proportion: 89% of the
budget's remaining capacity to read 20 rows. The revisit condition is
stated (nested quoting, embedded newlines, multiple dialects).
GR-S01 now deals Surface + hidden 1..=k as ruled, with edition values and
suits. Measured: 6/9/12 available against thresholds 5/7/9 -- the game is
winnable at every seat count, which is what the maintainer could not do.
gd0001 is INVERTED, not deleted, and now also asserts the 6/9/12 so a
deal that is reachable for the wrong reason still fails.
Blast radius was scenario expectations, exactly as the ADR predicted: no
scenario pinned a hash and no bundle is committed. Six scenarios and two
unit tests updated, each with a note. gr-e01-threshold-unreachable-2p is
RENAMED to -reachable- and rewritten as the non-provisional import check
ground-game asked for by name. gr-e03's setup was restructured, not just
renumbered: with values 2,2,2 its personal-edge test would have tied
three ways and asserted nothing.
BLOCKING: AM-7 fails at median 0.845 against its 0.9 floor. Isolated
across three runs -- 3 problems + stand-in 0.97, 3 problems + edition
0.909, 4 problems + edition 0.845. State is BOUNDED (proven: identical
after 5k and 100k events), so this is not the unbounded-growth defect
AM-7 exists to catch; it is a bigger working set streaming a long log.
Whether AM-7's floor is still right for a larger aggregate is a spec
question and lowering it requires an ADR, so it is not being tuned here.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 00:47:56 +02:00
|
|
|
|
"outcome.total": 9
|
2026-08-02 02:16:51 +02:00
|
|
|
|
"outcome.threshold": 9
|
|
|
|
|
|
"outcome.group_success": true
|
|
|
|
|
|
# GR-E03: claimed value −1 per Blame held.
|
CB-WP-0021 T01/T02/T05: the engine plays its own data — AM-7 blocks
ADR-0011 decided it: vendor the CSV with a checked digest, read it with
a ~50-line reader, and let the hashes move.
The declaration's constraint was measured against the WRONG BUDGET. It
said a CSV crate costs 21,613 against AM-4a's 3,798 of headroom, '5.7x
over, settled by measurement'. But setup and problem_priorities are
cfg(scenarios) and are not in the shipped runtime at all, so AM-4a never
sees them. Against AM-4b, csv costs 17,651 against 19,742 -- it FITS,
with 2,091 to spare. It is refused anyway, on proportion: 89% of the
budget's remaining capacity to read 20 rows. The revisit condition is
stated (nested quoting, embedded newlines, multiple dialects).
GR-S01 now deals Surface + hidden 1..=k as ruled, with edition values and
suits. Measured: 6/9/12 available against thresholds 5/7/9 -- the game is
winnable at every seat count, which is what the maintainer could not do.
gd0001 is INVERTED, not deleted, and now also asserts the 6/9/12 so a
deal that is reachable for the wrong reason still fails.
Blast radius was scenario expectations, exactly as the ADR predicted: no
scenario pinned a hash and no bundle is committed. Six scenarios and two
unit tests updated, each with a note. gr-e01-threshold-unreachable-2p is
RENAMED to -reachable- and rewritten as the non-provisional import check
ground-game asked for by name. gr-e03's setup was restructured, not just
renumbered: with values 2,2,2 its personal-edge test would have tied
three ways and asserted nothing.
BLOCKING: AM-7 fails at median 0.845 against its 0.9 floor. Isolated
across three runs -- 3 problems + stand-in 0.97, 3 problems + edition
0.909, 4 problems + edition 0.845. State is BOUNDED (proven: identical
after 5k and 100k events), so this is not the unbounded-growth defect
AM-7 exists to catch; it is a bigger working set streaming a long log.
Whether AM-7's floor is still right for a larger aggregate is a spec
question and lowering it requires an ADR, so it is not being tuned here.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 00:47:56 +02:00
|
|
|
|
"outcome.personal.3": 0
|
2026-08-02 02:16:51 +02:00
|
|
|
|
"outcome.personal.2": 3
|
|
|
|
|
|
# P4 claimed 4 and still loses: the Blame is load-bearing here.
|
|
|
|
|
|
"outcome.winners": [2]
|
|
|
|
|
|
# GR-E03 has no Mastery rating; that is GR-E02's.
|
|
|
|
|
|
"outcome.mastery": null
|
|
|
|
|
|
"round": 5
|
|
|
|
|
|
"step": End
|
|
|
|
|
|
rejects: []
|