clay-borg/editions/catalog.yaml

264 lines
10 KiB
YAML
Raw Normal View History

CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
# GROUND edition catalog — schema 2: composable modules
# Docs: CATALOG.md
ADR-0022 + CB-WP-0048 T00: the selector decision, and the mirror held The maintainer's observation decided the design: aspects partition the GAME, strata partition our apparatus, and they are orthogonal. A module is one coordinate change in aspect space with an obligation in every stratum. So aspect identity must NOT be Rust types -- an aspect ground-game adds would make clay-borg fail to parse a configuration rather than fail to run it, welding the two coordinate systems at the one place they must stay independent. Chosen: identity as data (Configuration round-trips anything the catalog names), behaviour exhaustive (Rules, no catch-all), resolve() between. Decisive argument: the catalog ALREADY ships modules with a rules_delta and status: proposed, so a per-aspect enum would report them as "unknown module" -- indistinguishable from a typo, a false statement about the edition, and this project's signature failure shape. Two facts need two errors. Federating design authority is permanent, so the representation must outlive the implementation. Legacy ids alias forever through the catalog's own legacy_experiment_id, on the standard-Np precedent: 26 recordings name them and the expansion is exact, so there is nothing to deprecate. T00 done: the schema-2 mirror had arrived with no digests (19 files) and edition-check was red. Digests are now generated by WALKING editions/, not typed -- two reviews already found hand-written lists that made their own controls vacuous, and a mirror that grows a directory is what breaks a maintained list. PROVENANCE-catalog.md was a file inside the mirrored tree that upstream does not have; folded into our own PROVENANCE.md, since provenance about the mirror does not belong inside the thing it describes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 22:07:57 +02:00
# clay-borg: select baseline + 0..N modules (≤1 per aspect), or a named profile.
CB-WP-0038: variant selection, H1 implemented, and H1 measured ground-game packages hypotheses as selectable rules variants — a catalog, a rules_delta.yaml, and prose — and their note is explicit that CSV text alone is not executable here. So the kernel gains a Variant in game state: in the state, therefore in the hash, therefore in the recording, because a scenario replayed under a different variant would diverge silently. Baseline is bit-for-bit what it was, asserted across seat counts and seeds. A variant system that perturbs the baseline invalidates every measurement this repo has. H1-A and H1-B implemented from rules_delta.yaml and mutation-proven on their own defects: "unclaimed" misread as face-up-and-unsolved, and the attacker's Stress read after the attack's effects. Their `unchanged:` list is asserted rather than trusted — that list is their claim about their own experiment. Measured, and three of their four criteria fail. DARVO arm rate is still 0 under greedy; ATTACK selection does not rise and falls for the rank-75 policy; group success collapses from 165/190/200 to 0 at 3/4/6 seats. The mechanism is not the assumed one: greedy answers the pressure by regulating, Stress plateaus at 3, so it never reaches the gate at 4 or the arm at 5 — H1-A acts as a solve-rate tax and H1-B is unreachable under competent play. A harness defect was caught before the claim: sweep discarded refused games silently and never reported its count, so "nobody won" and "nothing played" printed identically. Reporting H1 as unwinnable on that basis would have been the ADR-0018 family aimed at another repo's design. All 200 games ran in every cell; the zeros are real. Chaos d8 = 8 — the window's first override, redrew L against a structural L, so it changed nothing. Window 3 recorded in ChaosRollHistory. NOT REVIEWED: tier L owes a separate-agent adversarial review, and no H1 result may reach ground-game until it has run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 00:50:08 +02:00
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
schema_version: 2
ADR-0022 + CB-WP-0048 T00: the selector decision, and the mirror held The maintainer's observation decided the design: aspects partition the GAME, strata partition our apparatus, and they are orthogonal. A module is one coordinate change in aspect space with an obligation in every stratum. So aspect identity must NOT be Rust types -- an aspect ground-game adds would make clay-borg fail to parse a configuration rather than fail to run it, welding the two coordinate systems at the one place they must stay independent. Chosen: identity as data (Configuration round-trips anything the catalog names), behaviour exhaustive (Rules, no catch-all), resolve() between. Decisive argument: the catalog ALREADY ships modules with a rules_delta and status: proposed, so a per-aspect enum would report them as "unknown module" -- indistinguishable from a typo, a false statement about the edition, and this project's signature failure shape. Two facts need two errors. Federating design authority is permanent, so the representation must outlive the implementation. Legacy ids alias forever through the catalog's own legacy_experiment_id, on the standard-Np precedent: 26 recordings name them and the expansion is exact, so there is nothing to deprecate. T00 done: the schema-2 mirror had arrived with no digests (19 files) and edition-check was red. Digests are now generated by WALKING editions/, not typed -- two reviews already found hand-written lists that made their own controls vacuous, and a mirror that grows a directory is what breaks a maintained list. PROVENANCE-catalog.md was a file inside the mirrored tree that upstream does not have; folded into our own PROVENANCE.md, since provenance about the mirror does not belong inside the thing it describes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 22:07:57 +02:00
# terminology: aspect (human) = former axis; modules remain module_id
Apply ground-game's rulings: mastery in points, four boards, and a vendor tool that covers what the gate checks They ruled on all seven items the same day. Two were actionable here. F28 RULED: points. Modes.csv MODE_COOP clarified upstream to say "penalties apply to points, not card count"; mastery is now total - blame - denied. A recorded scenario went red on it -- gr-e02-shared-ground pinned 0 (2 claimed CARDS - 1 - 1) and now expects 2 (4 POINTS - 1 - 1). The number moved because the rule was decided, not because the engine drifted, and the scenario records both rulings; its schema has no field for a second one, so both live in ruled_note with `ruled` carrying the LATEST date. F29 RULED not-intended and APPLIED upstream: SCN_02's suits re-tuned the same day. The characterisation test is how we found out -- it pinned the duplication, went red on the re-tune, and that red WAS the notification. It now asserts every pair distinct, the stronger statement the duplication had made unavailable. SCN_02 re-measures at 73 at 2p, not 67: its own board now. F26/F30 ruled and recorded. F30's ruling incidentally confirms our reading -- they name priority-2's suit as the first lever, which is the difference we identified without having measured causation. vendor-editions grew twice, both times because it covered less than the gate it exists to satisfy: - It refused to touch ground-darvo-r0/ on the reasoning that the baseline is "a separate record". That was wrong within the hour: ground-game clarified Modes.csv and `make vendor` reported a clean sync while edition-check went red. A sync tool that covers less than its check reports success into a red gate. - Its two-block rewrite DETECTED which fence held which set and preserved the arrangement -- faithfully preserving a swap an earlier write had introduced, leaving each fence under a heading describing the other. edition-check reads every sha256 line flat and passed throughout: a document can be self-consistently wrong and green. Order is now asserted, with a control that goes red on a swap. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-09 00:15:14 +02:00
# packages declare files consumers must read (consumed_files / data_overlays)
updated: "2026-08-08"
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
default_baseline: ground-darvo-r0
default_profile: baseline
CB-WP-0038: variant selection, H1 implemented, and H1 measured ground-game packages hypotheses as selectable rules variants — a catalog, a rules_delta.yaml, and prose — and their note is explicit that CSV text alone is not executable here. So the kernel gains a Variant in game state: in the state, therefore in the hash, therefore in the recording, because a scenario replayed under a different variant would diverge silently. Baseline is bit-for-bit what it was, asserted across seat counts and seeds. A variant system that perturbs the baseline invalidates every measurement this repo has. H1-A and H1-B implemented from rules_delta.yaml and mutation-proven on their own defects: "unclaimed" misread as face-up-and-unsolved, and the attacker's Stress read after the attack's effects. Their `unchanged:` list is asserted rather than trusted — that list is their claim about their own experiment. Measured, and three of their four criteria fail. DARVO arm rate is still 0 under greedy; ATTACK selection does not rise and falls for the rank-75 policy; group success collapses from 165/190/200 to 0 at 3/4/6 seats. The mechanism is not the assumed one: greedy answers the pressure by regulating, Stress plateaus at 3, so it never reaches the gate at 4 or the arm at 5 — H1-A acts as a solve-rate tax and H1-B is unreachable under competent play. A harness defect was caught before the claim: sweep discarded refused games silently and never reported its count, so "nobody won" and "nothing played" printed identically. Reporting H1 as unwinnable on that basis would have been the ADR-0018 family aimed at another repo's design. All 200 games ran in every cell; the zeros are real. Chaos d8 = 8 — the window's first override, redrew L against a structural L, so it changed nothing. Window 3 recorded in ChaosRollHistory. NOT REVIEWED: tier L owes a separate-agent adversarial review, and no H1 result may reach ground-game until it has run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 00:50:08 +02:00
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
# ---------------------------------------------------------------------------
ADR-0022 + CB-WP-0048 T00: the selector decision, and the mirror held The maintainer's observation decided the design: aspects partition the GAME, strata partition our apparatus, and they are orthogonal. A module is one coordinate change in aspect space with an obligation in every stratum. So aspect identity must NOT be Rust types -- an aspect ground-game adds would make clay-borg fail to parse a configuration rather than fail to run it, welding the two coordinate systems at the one place they must stay independent. Chosen: identity as data (Configuration round-trips anything the catalog names), behaviour exhaustive (Rules, no catch-all), resolve() between. Decisive argument: the catalog ALREADY ships modules with a rules_delta and status: proposed, so a per-aspect enum would report them as "unknown module" -- indistinguishable from a typo, a false statement about the edition, and this project's signature failure shape. Two facts need two errors. Federating design authority is permanent, so the representation must outlive the implementation. Legacy ids alias forever through the catalog's own legacy_experiment_id, on the standard-Np precedent: 26 recordings name them and the expansion is exact, so there is nothing to deprecate. T00 done: the schema-2 mirror had arrived with no digests (19 files) and edition-check was red. Digests are now generated by WALKING editions/, not typed -- two reviews already found hand-written lists that made their own controls vacuous, and a mirror that grows a directory is what breaks a maintained list. PROVENANCE-catalog.md was a file inside the mirrored tree that upstream does not have; folded into our own PROVENANCE.md, since provenance about the mirror does not belong inside the thing it describes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 22:07:57 +02:00
# Aspects — orthogonal design dimensions (formerly "axes"). At most one non-default module per aspect.
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
# ---------------------------------------------------------------------------
ADR-0022 + CB-WP-0048 T00: the selector decision, and the mirror held The maintainer's observation decided the design: aspects partition the GAME, strata partition our apparatus, and they are orthogonal. A module is one coordinate change in aspect space with an obligation in every stratum. So aspect identity must NOT be Rust types -- an aspect ground-game adds would make clay-borg fail to parse a configuration rather than fail to run it, welding the two coordinate systems at the one place they must stay independent. Chosen: identity as data (Configuration round-trips anything the catalog names), behaviour exhaustive (Rules, no catch-all), resolve() between. Decisive argument: the catalog ALREADY ships modules with a rules_delta and status: proposed, so a per-aspect enum would report them as "unknown module" -- indistinguishable from a typo, a false statement about the edition, and this project's signature failure shape. Two facts need two errors. Federating design authority is permanent, so the representation must outlive the implementation. Legacy ids alias forever through the catalog's own legacy_experiment_id, on the standard-Np precedent: 26 recordings name them and the expansion is exact, so there is nothing to deprecate. T00 done: the schema-2 mirror had arrived with no digests (19 files) and edition-check was red. Digests are now generated by WALKING editions/, not typed -- two reviews already found hand-written lists that made their own controls vacuous, and a mirror that grows a directory is what breaks a maintained list. PROVENANCE-catalog.md was a file inside the mirrored tree that upstream does not have; folded into our own PROVENANCE.md, since provenance about the mirror does not belong inside the thing it describes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 22:07:57 +02:00
aspects:
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
- id: problem_stress
ADR-0022 + CB-WP-0048 T00: the selector decision, and the mirror held The maintainer's observation decided the design: aspects partition the GAME, strata partition our apparatus, and they are orthogonal. A module is one coordinate change in aspect space with an obligation in every stratum. So aspect identity must NOT be Rust types -- an aspect ground-game adds would make clay-borg fail to parse a configuration rather than fail to run it, welding the two coordinate systems at the one place they must stay independent. Chosen: identity as data (Configuration round-trips anything the catalog names), behaviour exhaustive (Rules, no catch-all), resolve() between. Decisive argument: the catalog ALREADY ships modules with a rules_delta and status: proposed, so a per-aspect enum would report them as "unknown module" -- indistinguishable from a typo, a false statement about the edition, and this project's signature failure shape. Two facts need two errors. Federating design authority is permanent, so the representation must outlive the implementation. Legacy ids alias forever through the catalog's own legacy_experiment_id, on the standard-Np precedent: 26 recordings name them and the expansion is exact, so there is nothing to deprecate. T00 done: the schema-2 mirror had arrived with no digests (19 files) and edition-check was red. Digests are now generated by WALKING editions/, not typed -- two reviews already found hand-written lists that made their own controls vacuous, and a mirror that grows a directory is what breaks a maintained list. PROVENANCE-catalog.md was a file inside the mirrored tree that upstream does not have; folded into our own PROVENANCE.md, since provenance about the mirror does not belong inside the thing it describes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 22:07:57 +02:00
# legacy key: axis (same meaning)
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
title: Problem → Stress routing
default_module: problem_stress.none
summary: >
Whether and how unclaimed Problems raise Stress at Round End.
- id: attack_relief
title: ATTACK self-soothing
default_module: attack_relief.none
summary: >
Whether resolving ATTACK can reduce the attacker's Stress.
- id: end_condition
title: How the game ends
default_module: end_condition.fixed_rounds_5
summary: >
Fixed round clock vs clear-board / collapse / hybrid ends.
- id: problem_deal
title: How Problems enter play
default_module: problem_deal.fixed_setup
summary: >
Fixed setup deal only vs mid-game influx (pressure deck, etc.).
ADR-0022 + CB-WP-0048 T00: the selector decision, and the mirror held The maintainer's observation decided the design: aspects partition the GAME, strata partition our apparatus, and they are orthogonal. A module is one coordinate change in aspect space with an obligation in every stratum. So aspect identity must NOT be Rust types -- an aspect ground-game adds would make clay-borg fail to parse a configuration rather than fail to run it, welding the two coordinate systems at the one place they must stay independent. Chosen: identity as data (Configuration round-trips anything the catalog names), behaviour exhaustive (Rules, no catch-all), resolve() between. Decisive argument: the catalog ALREADY ships modules with a rules_delta and status: proposed, so a per-aspect enum would report them as "unknown module" -- indistinguishable from a typo, a false statement about the edition, and this project's signature failure shape. Two facts need two errors. Federating design authority is permanent, so the representation must outlive the implementation. Legacy ids alias forever through the catalog's own legacy_experiment_id, on the standard-Np precedent: 26 recordings name them and the expansion is exact, so there is nothing to deprecate. T00 done: the schema-2 mirror had arrived with no digests (19 files) and edition-check was red. Digests are now generated by WALKING editions/, not typed -- two reviews already found hand-written lists that made their own controls vacuous, and a mirror that grows a directory is what breaks a maintained list. PROVENANCE-catalog.md was a file inside the mirrored tree that upstream does not have; folded into our own PROVENANCE.md, since provenance about the mirror does not belong inside the thing it describes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 22:07:57 +02:00
# Future aspects (not yet registered): setup_difficulty, sequence_pacing,
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
# ground_as_sequence, status_stress, competence_track.
# ---------------------------------------------------------------------------
# Baseline content package (CSV edition data)
# ---------------------------------------------------------------------------
baselines:
- baseline_id: ground-darvo-r0
CB-WP-0038: variant selection, H1 implemented, and H1 measured ground-game packages hypotheses as selectable rules variants — a catalog, a rules_delta.yaml, and prose — and their note is explicit that CSV text alone is not executable here. So the kernel gains a Variant in game state: in the state, therefore in the hash, therefore in the recording, because a scenario replayed under a different variant would diverge silently. Baseline is bit-for-bit what it was, asserted across seat counts and seeds. A variant system that perturbs the baseline invalidates every measurement this repo has. H1-A and H1-B implemented from rules_delta.yaml and mutation-proven on their own defects: "unclaimed" misread as face-up-and-unsolved, and the attacker's Stress read after the attack's effects. Their `unchanged:` list is asserted rather than trusted — that list is their claim about their own experiment. Measured, and three of their four criteria fail. DARVO arm rate is still 0 under greedy; ATTACK selection does not rise and falls for the rank-75 policy; group success collapses from 165/190/200 to 0 at 3/4/6 seats. The mechanism is not the assumed one: greedy answers the pressure by regulating, Stress plateaus at 3, so it never reaches the gate at 4 or the arm at 5 — H1-A acts as a solve-rate tax and H1-B is unreachable under competent play. A harness defect was caught before the claim: sweep discarded refused games silently and never reported its count, so "nobody won" and "nothing played" printed identically. Reporting H1 as unwinnable on that basis would have been the ADR-0018 family aimed at another repo's design. All 200 games ran in every cell; the zeros are real. Chaos d8 = 8 — the window's first override, redrew L against a structural L, so it changed nothing. Window 3 recorded in ChaosRollHistory. NOT REVIEWED: tier L owes a separate-agent adversarial review, and no H1 result may reach ground-game until it has run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 00:50:08 +02:00
path: editions/ground-darvo-r0
selectable: true
status: baseline
dataset_id: GROUND-DARVO-CORE-0.1
title: "GROUND DARVO Edition — core r0"
summary: >
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
Print/playtest content. Modes, deal 6/9/12, thresholds 5/7/9.
ADR-0022 + CB-WP-0048 T00: the selector decision, and the mirror held The maintainer's observation decided the design: aspects partition the GAME, strata partition our apparatus, and they are orthogonal. A module is one coordinate change in aspect space with an obligation in every stratum. So aspect identity must NOT be Rust types -- an aspect ground-game adds would make clay-borg fail to parse a configuration rather than fail to run it, welding the two coordinate systems at the one place they must stay independent. Chosen: identity as data (Configuration round-trips anything the catalog names), behaviour exhaustive (Rules, no catch-all), resolve() between. Decisive argument: the catalog ALREADY ships modules with a rules_delta and status: proposed, so a per-aspect enum would report them as "unknown module" -- indistinguishable from a typo, a false statement about the edition, and this project's signature failure shape. Two facts need two errors. Federating design authority is permanent, so the representation must outlive the implementation. Legacy ids alias forever through the catalog's own legacy_experiment_id, on the standard-Np precedent: 26 recordings name them and the expansion is exact, so there is nothing to deprecate. T00 done: the schema-2 mirror had arrived with no digests (19 files) and edition-check was red. Digests are now generated by WALKING editions/, not typed -- two reviews already found hand-written lists that made their own controls vacuous, and a mirror that grows a directory is what breaks a maintained list. PROVENANCE-catalog.md was a file inside the mirrored tree that upstream does not have; folded into our own PROVENANCE.md, since provenance about the mirror does not belong inside the thing it describes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 22:07:57 +02:00
Default modules on all aspects = r0 printed behaviour.
CB-WP-0038: variant selection, H1 implemented, and H1 measured ground-game packages hypotheses as selectable rules variants — a catalog, a rules_delta.yaml, and prose — and their note is explicit that CSV text alone is not executable here. So the kernel gains a Variant in game state: in the state, therefore in the hash, therefore in the recording, because a scenario replayed under a different variant would diverge silently. Baseline is bit-for-bit what it was, asserted across seat counts and seeds. A variant system that perturbs the baseline invalidates every measurement this repo has. H1-A and H1-B implemented from rules_delta.yaml and mutation-proven on their own defects: "unclaimed" misread as face-up-and-unsolved, and the attacker's Stress read after the attack's effects. Their `unchanged:` list is asserted rather than trusted — that list is their claim about their own experiment. Measured, and three of their four criteria fail. DARVO arm rate is still 0 under greedy; ATTACK selection does not rise and falls for the rank-75 policy; group success collapses from 165/190/200 to 0 at 3/4/6 seats. The mechanism is not the assumed one: greedy answers the pressure by regulating, Stress plateaus at 3, so it never reaches the gate at 4 or the arm at 5 — H1-A acts as a solve-rate tax and H1-B is unreachable under competent play. A harness defect was caught before the claim: sweep discarded refused games silently and never reported its count, so "nobody won" and "nothing played" printed identically. Reporting H1 as unwinnable on that basis would have been the ADR-0018 family aimed at another repo's design. All 200 games ran in every cell; the zeros are real. Chaos d8 = 8 — the window's first override, redrew L against a structural L, so it changed nothing. Window 3 recorded in ChaosRollHistory. NOT REVIEWED: tier L owes a separate-agent adversarial review, and no H1 result may reach ground-game until it has run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 00:50:08 +02:00
utility_estimate: >
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
Ship-default content. Stress has no problem pressure until a
problem_stress module is selected.
CB-WP-0038: variant selection, H1 implemented, and H1 measured ground-game packages hypotheses as selectable rules variants — a catalog, a rules_delta.yaml, and prose — and their note is explicit that CSV text alone is not executable here. So the kernel gains a Variant in game state: in the state, therefore in the hash, therefore in the recording, because a scenario replayed under a different variant would diverge silently. Baseline is bit-for-bit what it was, asserted across seat counts and seeds. A variant system that perturbs the baseline invalidates every measurement this repo has. H1-A and H1-B implemented from rules_delta.yaml and mutation-proven on their own defects: "unclaimed" misread as face-up-and-unsolved, and the attacker's Stress read after the attack's effects. Their `unchanged:` list is asserted rather than trusted — that list is their claim about their own experiment. Measured, and three of their four criteria fail. DARVO arm rate is still 0 under greedy; ATTACK selection does not rise and falls for the rank-75 policy; group success collapses from 165/190/200 to 0 at 3/4/6 seats. The mechanism is not the assumed one: greedy answers the pressure by regulating, Stress plateaus at 3, so it never reaches the gate at 4 or the arm at 5 — H1-A acts as a solve-rate tax and H1-B is unreachable under competent play. A harness defect was caught before the claim: sweep discarded refused games silently and never reported its count, so "nobody won" and "nothing played" printed identically. Reporting H1 as unwinnable on that basis would have been the ADR-0018 family aimed at another repo's design. All 200 games ran in every cell; the zeros are real. Chaos d8 = 8 — the window's first override, redrew L against a structural L, so it changed nothing. Window 3 recorded in ChaosRollHistory. NOT REVIEWED: tier L owes a separate-agent adversarial review, and no H1 result may reach ground-game until it has run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 00:50:08 +02:00
decision: none
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
# ---------------------------------------------------------------------------
# Modules — independent variations (one directory each)
# ---------------------------------------------------------------------------
modules:
# --- problem_stress ---
- module_id: problem_stress.none
ADR-0022 + CB-WP-0048 T00: the selector decision, and the mirror held The maintainer's observation decided the design: aspects partition the GAME, strata partition our apparatus, and they are orthogonal. A module is one coordinate change in aspect space with an obligation in every stratum. So aspect identity must NOT be Rust types -- an aspect ground-game adds would make clay-borg fail to parse a configuration rather than fail to run it, welding the two coordinate systems at the one place they must stay independent. Chosen: identity as data (Configuration round-trips anything the catalog names), behaviour exhaustive (Rules, no catch-all), resolve() between. Decisive argument: the catalog ALREADY ships modules with a rules_delta and status: proposed, so a per-aspect enum would report them as "unknown module" -- indistinguishable from a typo, a false statement about the edition, and this project's signature failure shape. Two facts need two errors. Federating design authority is permanent, so the representation must outlive the implementation. Legacy ids alias forever through the catalog's own legacy_experiment_id, on the standard-Np precedent: 26 recordings name them and the expansion is exact, so there is nothing to deprecate. T00 done: the schema-2 mirror had arrived with no digests (19 files) and edition-check was red. Digests are now generated by WALKING editions/, not typed -- two reviews already found hand-written lists that made their own controls vacuous, and a mirror that grows a directory is what breaks a maintained list. PROVENANCE-catalog.md was a file inside the mirrored tree that upstream does not have; folded into our own PROVENANCE.md, since provenance about the mirror does not belong inside the thing it describes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 22:07:57 +02:00
aspect: problem_stress
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
path: editions/modules/problem_stress/none
is_default: true
selectable: true
status: baseline-default
rules_delta: null
summary: Unclaimed Problems do not raise Stress (r0).
decision: none
- module_id: problem_stress.flat_any_open
ADR-0022 + CB-WP-0048 T00: the selector decision, and the mirror held The maintainer's observation decided the design: aspects partition the GAME, strata partition our apparatus, and they are orthogonal. A module is one coordinate change in aspect space with an obligation in every stratum. So aspect identity must NOT be Rust types -- an aspect ground-game adds would make clay-borg fail to parse a configuration rather than fail to run it, welding the two coordinate systems at the one place they must stay independent. Chosen: identity as data (Configuration round-trips anything the catalog names), behaviour exhaustive (Rules, no catch-all), resolve() between. Decisive argument: the catalog ALREADY ships modules with a rules_delta and status: proposed, so a per-aspect enum would report them as "unknown module" -- indistinguishable from a typo, a false statement about the edition, and this project's signature failure shape. Two facts need two errors. Federating design authority is permanent, so the representation must outlive the implementation. Legacy ids alias forever through the catalog's own legacy_experiment_id, on the standard-Np precedent: 26 recordings name them and the expansion is exact, so there is nothing to deprecate. T00 done: the schema-2 mirror had arrived with no digests (19 files) and edition-check was red. Digests are now generated by WALKING editions/, not typed -- two reviews already found hand-written lists that made their own controls vacuous, and a mirror that grows a directory is what breaks a maintained list. PROVENANCE-catalog.md was a file inside the mirrored tree that upstream does not have; folded into our own PROVENANCE.md, since provenance about the mirror does not belong inside the thing it describes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 22:07:57 +02:00
aspect: problem_stress
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
path: editions/modules/problem_stress/flat_any_open
is_default: false
CB-WP-0038: variant selection, H1 implemented, and H1 measured ground-game packages hypotheses as selectable rules variants — a catalog, a rules_delta.yaml, and prose — and their note is explicit that CSV text alone is not executable here. So the kernel gains a Variant in game state: in the state, therefore in the hash, therefore in the recording, because a scenario replayed under a different variant would diverge silently. Baseline is bit-for-bit what it was, asserted across seat counts and seeds. A variant system that perturbs the baseline invalidates every measurement this repo has. H1-A and H1-B implemented from rules_delta.yaml and mutation-proven on their own defects: "unclaimed" misread as face-up-and-unsolved, and the attacker's Stress read after the attack's effects. Their `unchanged:` list is asserted rather than trusted — that list is their claim about their own experiment. Measured, and three of their four criteria fail. DARVO arm rate is still 0 under greedy; ATTACK selection does not rise and falls for the rank-75 policy; group success collapses from 165/190/200 to 0 at 3/4/6 seats. The mechanism is not the assumed one: greedy answers the pressure by regulating, Stress plateaus at 3, so it never reaches the gate at 4 or the arm at 5 — H1-A acts as a solve-rate tax and H1-B is unreachable under competent play. A harness defect was caught before the claim: sweep discarded refused games silently and never reported its count, so "nobody won" and "nothing played" printed identically. Reporting H1 as unwinnable on that basis would have been the ADR-0018 family aimed at another repo's design. All 200 games ran in every cell; the zeros are real. Chaos d8 = 8 — the window's first override, redrew L against a structural L, so it changed nothing. Window 3 recorded in ChaosRollHistory. NOT REVIEWED: tier L owes a separate-agent adversarial review, and no H1 result may reach ground-game until it has run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 00:50:08 +02:00
selectable: true
status: measured
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
rules_delta: editions/modules/problem_stress/flat_any_open/rules_delta.yaml
legacy_experiment_ids: [h1-problem-stress]
measurement_ref: reports/260808-clay-borg-h1-measured.md
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
summary: >
+1 Stress to every seat if any Problem unclaimed (former H1-A).
CB-WP-0038: variant selection, H1 implemented, and H1 measured ground-game packages hypotheses as selectable rules variants — a catalog, a rules_delta.yaml, and prose — and their note is explicit that CSV text alone is not executable here. So the kernel gains a Variant in game state: in the state, therefore in the hash, therefore in the recording, because a scenario replayed under a different variant would diverge silently. Baseline is bit-for-bit what it was, asserted across seat counts and seeds. A variant system that perturbs the baseline invalidates every measurement this repo has. H1-A and H1-B implemented from rules_delta.yaml and mutation-proven on their own defects: "unclaimed" misread as face-up-and-unsolved, and the attacker's Stress read after the attack's effects. Their `unchanged:` list is asserted rather than trusted — that list is their claim about their own experiment. Measured, and three of their four criteria fail. DARVO arm rate is still 0 under greedy; ATTACK selection does not rise and falls for the rank-75 policy; group success collapses from 165/190/200 to 0 at 3/4/6 seats. The mechanism is not the assumed one: greedy answers the pressure by regulating, Stress plateaus at 3, so it never reaches the gate at 4 or the arm at 5 — H1-A acts as a solve-rate tax and H1-B is unreachable under competent play. A harness defect was caught before the claim: sweep discarded refused games silently and never reported its count, so "nobody won" and "nothing played" printed identically. Reporting H1 as unwinnable on that basis would have been the ADR-0018 family aimed at another repo's design. All 200 games ran in every cell; the zeros are real. Chaos d8 = 8 — the window's first override, redrew L against a structural L, so it changed nothing. Window 3 recorded in ChaosRollHistory. NOT REVIEWED: tier L owes a separate-agent adversarial review, and no H1 result may reach ground-game until it has run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 00:50:08 +02:00
utility_estimate: >
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
Reject as sole pressure: greedy 34p wins → 0. Keep for A/B control.
decision: reject-as-baseline
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
clay_borg_notes: Former H1-A; implement without H1-B unless attack_relief also selected.
CB-RES-0009: extensive form is the lingua franca Two questions from the maintainer — is there a game-theory mapping to Ludii's language, and is that language formal enough to derive one from. Yes, no, and the no does not matter. The mapping is proven, not to be invented: "The Ludii Game Description Language is Universal" shows the language can represent an equivalent game for any finite, non-deterministic, imperfect-information game, extending earlier work limited to finite deterministic fully-observable extensive-form games. EFG is also OpenSpiel's object, so the same formalism connects description to analysis: Ludii -> EFG <- OpenSpiel. Ludii's syntax is formal and unusually so — a class grammar derived automatically from its source. Its semantics are its Java: a ludeme means what its class does, and Ludii effectively makes Java the game description language. So there is no independent calculus to extract. The formality lives in the universality RESULT, not in a definition of meaning. GDL has the semantics and pays for it in speed — six times on Gomoku, twenty on Amazons and Hex, over two hundred on Chess. Conclusion: do not derive a language from Ludii; target the EFG directly. And we are closer than the tracks assumed. The journal is the history, Outcome is the payoff, legal_commands gives the actions — and project(Viewer::Player(seat)) IS the information partition, built so a player is not shown another's hand and unremarked as exactly the machinery imperfect information needs. Three gaps: chance is folded into a seed so a game is one realisation rather than a game with chance nodes; perfect recall is unasserted, which CFR and exploitability both assume; and commit/reveal is the standard EFG encoding of simultaneity but is never stated as such. Perfect recall is checkable from the journal today and is now Track B's first task — if it fails, every equilibrium concept we might quote is unsound here. Also re-vendored the catalog twice: ground-game added H2 — scoped problem stress, applying End Stress by personal/bond/global scope instead of flat to everyone, which is a direct response to our reading that H1's tax scales with the Problems while its intended effect does not. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 14:57:25 +02:00
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
- module_id: problem_stress.scoped
ADR-0022 + CB-WP-0048 T00: the selector decision, and the mirror held The maintainer's observation decided the design: aspects partition the GAME, strata partition our apparatus, and they are orthogonal. A module is one coordinate change in aspect space with an obligation in every stratum. So aspect identity must NOT be Rust types -- an aspect ground-game adds would make clay-borg fail to parse a configuration rather than fail to run it, welding the two coordinate systems at the one place they must stay independent. Chosen: identity as data (Configuration round-trips anything the catalog names), behaviour exhaustive (Rules, no catch-all), resolve() between. Decisive argument: the catalog ALREADY ships modules with a rules_delta and status: proposed, so a per-aspect enum would report them as "unknown module" -- indistinguishable from a typo, a false statement about the edition, and this project's signature failure shape. Two facts need two errors. Federating design authority is permanent, so the representation must outlive the implementation. Legacy ids alias forever through the catalog's own legacy_experiment_id, on the standard-Np precedent: 26 recordings name them and the expansion is exact, so there is nothing to deprecate. T00 done: the schema-2 mirror had arrived with no digests (19 files) and edition-check was red. Digests are now generated by WALKING editions/, not typed -- two reviews already found hand-written lists that made their own controls vacuous, and a mirror that grows a directory is what breaks a maintained list. PROVENANCE-catalog.md was a file inside the mirrored tree that upstream does not have; folded into our own PROVENANCE.md, since provenance about the mirror does not belong inside the thing it describes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 22:07:57 +02:00
aspect: problem_stress
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
path: editions/modules/problem_stress/scoped
is_default: false
CB-RES-0009: extensive form is the lingua franca Two questions from the maintainer — is there a game-theory mapping to Ludii's language, and is that language formal enough to derive one from. Yes, no, and the no does not matter. The mapping is proven, not to be invented: "The Ludii Game Description Language is Universal" shows the language can represent an equivalent game for any finite, non-deterministic, imperfect-information game, extending earlier work limited to finite deterministic fully-observable extensive-form games. EFG is also OpenSpiel's object, so the same formalism connects description to analysis: Ludii -> EFG <- OpenSpiel. Ludii's syntax is formal and unusually so — a class grammar derived automatically from its source. Its semantics are its Java: a ludeme means what its class does, and Ludii effectively makes Java the game description language. So there is no independent calculus to extract. The formality lives in the universality RESULT, not in a definition of meaning. GDL has the semantics and pays for it in speed — six times on Gomoku, twenty on Amazons and Hex, over two hundred on Chess. Conclusion: do not derive a language from Ludii; target the EFG directly. And we are closer than the tracks assumed. The journal is the history, Outcome is the payoff, legal_commands gives the actions — and project(Viewer::Player(seat)) IS the information partition, built so a player is not shown another's hand and unremarked as exactly the machinery imperfect information needs. Three gaps: chance is folded into a seed so a game is one realisation rather than a game with chance nodes; perfect recall is unasserted, which CFR and exploitability both assume; and commit/reveal is the standard EFG encoding of simultaneity but is never stated as such. Perfect recall is checkable from the journal today and is now Track B's first task — if it fails, every equilibrium concept we might quote is unsound here. Also re-vendored the catalog twice: ground-game added H2 — scoped problem stress, applying End Stress by personal/bond/global scope instead of flat to everyone, which is a direct response to our reading that H1's tax scales with the Problems while its intended effect does not. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 14:57:25 +02:00
selectable: true
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
status: measured
rules_delta: editions/modules/problem_stress/scoped/rules_delta.yaml
data_overlays: [Problems.csv, Rules_Text.csv]
Apply ground-game's rulings: mastery in points, four boards, and a vendor tool that covers what the gate checks They ruled on all seven items the same day. Two were actionable here. F28 RULED: points. Modes.csv MODE_COOP clarified upstream to say "penalties apply to points, not card count"; mastery is now total - blame - denied. A recorded scenario went red on it -- gr-e02-shared-ground pinned 0 (2 claimed CARDS - 1 - 1) and now expects 2 (4 POINTS - 1 - 1). The number moved because the rule was decided, not because the engine drifted, and the scenario records both rulings; its schema has no field for a second one, so both live in ruled_note with `ruled` carrying the LATEST date. F29 RULED not-intended and APPLIED upstream: SCN_02's suits re-tuned the same day. The characterisation test is how we found out -- it pinned the duplication, went red on the re-tune, and that red WAS the notification. It now asserts every pair distinct, the stronger statement the duplication had made unavailable. SCN_02 re-measures at 73 at 2p, not 67: its own board now. F26/F30 ruled and recorded. F30's ruling incidentally confirms our reading -- they name priority-2's suit as the first lever, which is the difference we identified without having measured causation. vendor-editions grew twice, both times because it covered less than the gate it exists to satisfy: - It refused to touch ground-darvo-r0/ on the reasoning that the baseline is "a separate record". That was wrong within the hour: ground-game clarified Modes.csv and `make vendor` reported a clean sync while edition-check went red. A sync tool that covers less than its check reports success into a red gate. - Its two-block rewrite DETECTED which fence held which set and preserved the arrangement -- faithfully preserving a swap an earlier write had introduced, leaving each fence under a heading describing the other. edition-check reads every sha256 line flat and passed throughout: a document can be self-consistently wrong and green. Order is now asserted, with a control that goes red on a swap. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-09 00:15:14 +02:00
# files a consumer must open to implement this module (RPT-0006 §5b)
consumed_files:
- editions/modules/problem_stress/scoped/rules_delta.yaml
- editions/modules/problem_stress/scoped/Problems.csv
- editions/modules/problem_stress/scoped/Rules_Text.csv
- editions/modules/problem_stress/scoped/MODULE.md
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
legacy_experiment_ids: [h2-scoped-problem-stress]
measurement_ref: reports/260808-clay-borg-h2-measured.md
CB-RES-0009: extensive form is the lingua franca Two questions from the maintainer — is there a game-theory mapping to Ludii's language, and is that language formal enough to derive one from. Yes, no, and the no does not matter. The mapping is proven, not to be invented: "The Ludii Game Description Language is Universal" shows the language can represent an equivalent game for any finite, non-deterministic, imperfect-information game, extending earlier work limited to finite deterministic fully-observable extensive-form games. EFG is also OpenSpiel's object, so the same formalism connects description to analysis: Ludii -> EFG <- OpenSpiel. Ludii's syntax is formal and unusually so — a class grammar derived automatically from its source. Its semantics are its Java: a ludeme means what its class does, and Ludii effectively makes Java the game description language. So there is no independent calculus to extract. The formality lives in the universality RESULT, not in a definition of meaning. GDL has the semantics and pays for it in speed — six times on Gomoku, twenty on Amazons and Hex, over two hundred on Chess. Conclusion: do not derive a language from Ludii; target the EFG directly. And we are closer than the tracks assumed. The journal is the history, Outcome is the payoff, legal_commands gives the actions — and project(Viewer::Player(seat)) IS the information partition, built so a player is not shown another's hand and unremarked as exactly the machinery imperfect information needs. Three gaps: chance is folded into a seed so a game is one realisation rather than a game with chance nodes; perfect recall is unasserted, which CFR and exploitability both assume; and commit/reveal is the standard EFG encoding of simultaneity but is never stated as such. Perfect recall is checkable from the journal today and is now Track B's first task — if it fails, every equilibrium concept we might quote is unsound here. Also re-vendored the catalog twice: ground-game added H2 — scoped problem stress, applying End Stress by personal/bond/global scope instead of flat to everyone, which is a direct response to our reading that H1's tax scales with the Problems while its intended effect does not. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 14:57:25 +02:00
hypothesis_ref: history/260808-h2-scoped-problem-stress.md
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
summary: >
+1 Stress per unclaimed Problem to stress_scope (personal/bond/global).
CB-RES-0009: extensive form is the lingua franca Two questions from the maintainer — is there a game-theory mapping to Ludii's language, and is that language formal enough to derive one from. Yes, no, and the no does not matter. The mapping is proven, not to be invented: "The Ludii Game Description Language is Universal" shows the language can represent an equivalent game for any finite, non-deterministic, imperfect-information game, extending earlier work limited to finite deterministic fully-observable extensive-form games. EFG is also OpenSpiel's object, so the same formalism connects description to analysis: Ludii -> EFG <- OpenSpiel. Ludii's syntax is formal and unusually so — a class grammar derived automatically from its source. Its semantics are its Java: a ludeme means what its class does, and Ludii effectively makes Java the game description language. So there is no independent calculus to extract. The formality lives in the universality RESULT, not in a definition of meaning. GDL has the semantics and pays for it in speed — six times on Gomoku, twenty on Amazons and Hex, over two hundred on Chess. Conclusion: do not derive a language from Ludii; target the EFG directly. And we are closer than the tracks assumed. The journal is the history, Outcome is the payoff, legal_commands gives the actions — and project(Viewer::Player(seat)) IS the information partition, built so a player is not shown another's hand and unremarked as exactly the machinery imperfect information needs. Three gaps: chance is folded into a seed so a game is one realisation rather than a game with chance nodes; perfect recall is unasserted, which CFR and exploitability both assume; and commit/reveal is the standard EFG encoding of simultaneity but is never stated as such. Perfect recall is checkable from the journal today and is now Track B's first task — if it fails, every equilibrium concept we might quote is unsound here. Also re-vendored the catalog twice: ground-game added H2 — scoped problem stress, applying End Stress by personal/bond/global scope instead of flat to everyone, which is a direct response to our reading that H1's tax scales with the Problems while its intended effect does not. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 14:57:25 +02:00
utility_estimate: >
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
Scoping works (RPT-0005). Keep; candidate promote with other axes later.
decision: keep-as-experiment
clay_borg_notes: >
Prefer path editions/modules/problem_stress/scoped. Use with_variant()/owners.
legacy_experiment_id h2-scoped-problem-stress remains an alias profile.
# --- attack_relief ---
- module_id: attack_relief.none
ADR-0022 + CB-WP-0048 T00: the selector decision, and the mirror held The maintainer's observation decided the design: aspects partition the GAME, strata partition our apparatus, and they are orthogonal. A module is one coordinate change in aspect space with an obligation in every stratum. So aspect identity must NOT be Rust types -- an aspect ground-game adds would make clay-borg fail to parse a configuration rather than fail to run it, welding the two coordinate systems at the one place they must stay independent. Chosen: identity as data (Configuration round-trips anything the catalog names), behaviour exhaustive (Rules, no catch-all), resolve() between. Decisive argument: the catalog ALREADY ships modules with a rules_delta and status: proposed, so a per-aspect enum would report them as "unknown module" -- indistinguishable from a typo, a false statement about the edition, and this project's signature failure shape. Two facts need two errors. Federating design authority is permanent, so the representation must outlive the implementation. Legacy ids alias forever through the catalog's own legacy_experiment_id, on the standard-Np precedent: 26 recordings name them and the expansion is exact, so there is nothing to deprecate. T00 done: the schema-2 mirror had arrived with no digests (19 files) and edition-check was red. Digests are now generated by WALKING editions/, not typed -- two reviews already found hand-written lists that made their own controls vacuous, and a mirror that grows a directory is what breaks a maintained list. PROVENANCE-catalog.md was a file inside the mirrored tree that upstream does not have; folded into our own PROVENANCE.md, since provenance about the mirror does not belong inside the thing it describes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 22:07:57 +02:00
aspect: attack_relief
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
path: editions/modules/attack_relief/none
is_default: true
selectable: true
status: baseline-default
rules_delta: null
summary: ATTACK does not self-soothe (r0).
decision: none
- module_id: attack_relief.self_soothe_ge4
ADR-0022 + CB-WP-0048 T00: the selector decision, and the mirror held The maintainer's observation decided the design: aspects partition the GAME, strata partition our apparatus, and they are orthogonal. A module is one coordinate change in aspect space with an obligation in every stratum. So aspect identity must NOT be Rust types -- an aspect ground-game adds would make clay-borg fail to parse a configuration rather than fail to run it, welding the two coordinate systems at the one place they must stay independent. Chosen: identity as data (Configuration round-trips anything the catalog names), behaviour exhaustive (Rules, no catch-all), resolve() between. Decisive argument: the catalog ALREADY ships modules with a rules_delta and status: proposed, so a per-aspect enum would report them as "unknown module" -- indistinguishable from a typo, a false statement about the edition, and this project's signature failure shape. Two facts need two errors. Federating design authority is permanent, so the representation must outlive the implementation. Legacy ids alias forever through the catalog's own legacy_experiment_id, on the standard-Np precedent: 26 recordings name them and the expansion is exact, so there is nothing to deprecate. T00 done: the schema-2 mirror had arrived with no digests (19 files) and edition-check was red. Digests are now generated by WALKING editions/, not typed -- two reviews already found hand-written lists that made their own controls vacuous, and a mirror that grows a directory is what breaks a maintained list. PROVENANCE-catalog.md was a file inside the mirrored tree that upstream does not have; folded into our own PROVENANCE.md, since provenance about the mirror does not belong inside the thing it describes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 22:07:57 +02:00
aspect: attack_relief
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
path: editions/modules/attack_relief/self_soothe_ge4
is_default: false
selectable: true
status: measured-as-combo
rules_delta: editions/modules/attack_relief/self_soothe_ge4/rules_delta.yaml
legacy_experiment_ids: [h1-problem-stress]
measurement_ref: reports/260808-clay-borg-h1-measured.md
summary: Uncancelled ATTACK at Stress ≥4 → attacker 1 Stress (former H1-B).
utility_estimate: >
Alone unmeasured. With flat problem stress, dead for competent play.
Re-measure with problem_stress.scoped via profile scoped_plus_attack_soothe.
decision: none
clay_borg_notes: Independent of problem_stress; compose explicitly.
# --- end_condition ---
- module_id: end_condition.fixed_rounds_5
ADR-0022 + CB-WP-0048 T00: the selector decision, and the mirror held The maintainer's observation decided the design: aspects partition the GAME, strata partition our apparatus, and they are orthogonal. A module is one coordinate change in aspect space with an obligation in every stratum. So aspect identity must NOT be Rust types -- an aspect ground-game adds would make clay-borg fail to parse a configuration rather than fail to run it, welding the two coordinate systems at the one place they must stay independent. Chosen: identity as data (Configuration round-trips anything the catalog names), behaviour exhaustive (Rules, no catch-all), resolve() between. Decisive argument: the catalog ALREADY ships modules with a rules_delta and status: proposed, so a per-aspect enum would report them as "unknown module" -- indistinguishable from a typo, a false statement about the edition, and this project's signature failure shape. Two facts need two errors. Federating design authority is permanent, so the representation must outlive the implementation. Legacy ids alias forever through the catalog's own legacy_experiment_id, on the standard-Np precedent: 26 recordings name them and the expansion is exact, so there is nothing to deprecate. T00 done: the schema-2 mirror had arrived with no digests (19 files) and edition-check was red. Digests are now generated by WALKING editions/, not typed -- two reviews already found hand-written lists that made their own controls vacuous, and a mirror that grows a directory is what breaks a maintained list. PROVENANCE-catalog.md was a file inside the mirrored tree that upstream does not have; folded into our own PROVENANCE.md, since provenance about the mirror does not belong inside the thing it describes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 22:07:57 +02:00
aspect: end_condition
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
path: editions/modules/end_condition/fixed_rounds_5
is_default: true
selectable: true
status: baseline-default
rules_delta: null
summary: Always 5 rounds then threshold scoring (r0).
decision: none
- module_id: end_condition.hybrid_clear_collapse
ADR-0022 + CB-WP-0048 T00: the selector decision, and the mirror held The maintainer's observation decided the design: aspects partition the GAME, strata partition our apparatus, and they are orthogonal. A module is one coordinate change in aspect space with an obligation in every stratum. So aspect identity must NOT be Rust types -- an aspect ground-game adds would make clay-borg fail to parse a configuration rather than fail to run it, welding the two coordinate systems at the one place they must stay independent. Chosen: identity as data (Configuration round-trips anything the catalog names), behaviour exhaustive (Rules, no catch-all), resolve() between. Decisive argument: the catalog ALREADY ships modules with a rules_delta and status: proposed, so a per-aspect enum would report them as "unknown module" -- indistinguishable from a typo, a false statement about the edition, and this project's signature failure shape. Two facts need two errors. Federating design authority is permanent, so the representation must outlive the implementation. Legacy ids alias forever through the catalog's own legacy_experiment_id, on the standard-Np precedent: 26 recordings name them and the expansion is exact, so there is nothing to deprecate. T00 done: the schema-2 mirror had arrived with no digests (19 files) and edition-check was red. Digests are now generated by WALKING editions/, not typed -- two reviews already found hand-written lists that made their own controls vacuous, and a mirror that grows a directory is what breaks a maintained list. PROVENANCE-catalog.md was a file inside the mirrored tree that upstream does not have; folded into our own PROVENANCE.md, since provenance about the mirror does not belong inside the thing it describes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 22:07:57 +02:00
aspect: end_condition
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
path: editions/modules/end_condition/hybrid_clear_collapse
is_default: false
selectable: true
CB-WP-0049 T02/T03: a seat that plays its objective, and F27 splits in two objective() reads GroundState::score (now public) rather than restating what winning is; a copy in the bot would disagree with the kernel the first time ground-game rules on F28. Working out WHERE the modes can differ was most of the task and it bounds the result: SOLVE always claims for the actor, so own-score and group-score want the same SOLVE nearly everywhere. That is a fact about GROUND's action set, not a shortcoming of the bot. Two real divergences, both readable off the table: SUPPORT regulates someone else (worth less against a rival, worth MORE under coalitions where a Bond merges them into my side), and SOLVE's value is the card's value, which greedy ignores entirely. THE RESULT — F27 splits in two: group success UNCHANGED in 34 of 36 cells who wins MOVES: BONDED COALITIONS at 4p goes 2.04 -> 2.98, 2.12 -> 3.29, 2.05 -> 3.01 winning seats per game So "the competitive modes are scoring lenses over cooperative play" was too strong and is withdrawn. The sharper claim: GROUND's scoring modes change WHO WINS, not WHETHER THE GROUP SUCCEEDS. And the effect is seat-band dependent -- 2p none, 4p largest, 6p none under coalitions; two relation slots capping network growth is a candidate explanation and is untested. The panel now prints BOTH policies side by side. That was a correction mid-task: the first version printed only the new one and I compared it against a figure remembered from CB-WP-0047 -- a comparison against a board nobody re-ran. Control that makes the numbers mean anything: under SHARED GROUND the two policies agree at all but <=2 decision points across 12 boards, so a moving column is mode-awareness and not simply a different bot. Also: two T01 tests keyed on `status: proposed`, which ground-game renamed to `ready-for-implement` mid-session. They now find the module by asking resolve() -- the structural property is ours and does not move when another repo edits its vocabulary. Also: `make vendor` replaces three hand re-vendors with a tool that regenerates digests by walking editions/, and reports one-sided files rather than resolving them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 23:31:11 +02:00
status: ready-for-implement
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
rules_delta: editions/modules/end_condition/hybrid_clear_collapse/rules_delta.yaml
hypothesis_ref: history/260808-deal-end-sequences-design.md
summary: >
CB-WP-0049 T02/T03: a seat that plays its objective, and F27 splits in two objective() reads GroundState::score (now public) rather than restating what winning is; a copy in the bot would disagree with the kernel the first time ground-game rules on F28. Working out WHERE the modes can differ was most of the task and it bounds the result: SOLVE always claims for the actor, so own-score and group-score want the same SOLVE nearly everywhere. That is a fact about GROUND's action set, not a shortcoming of the bot. Two real divergences, both readable off the table: SUPPORT regulates someone else (worth less against a rival, worth MORE under coalitions where a Bond merges them into my side), and SOLVE's value is the card's value, which greedy ignores entirely. THE RESULT — F27 splits in two: group success UNCHANGED in 34 of 36 cells who wins MOVES: BONDED COALITIONS at 4p goes 2.04 -> 2.98, 2.12 -> 3.29, 2.05 -> 3.01 winning seats per game So "the competitive modes are scoring lenses over cooperative play" was too strong and is withdrawn. The sharper claim: GROUND's scoring modes change WHO WINS, not WHETHER THE GROUP SUCCEEDS. And the effect is seat-band dependent -- 2p none, 4p largest, 6p none under coalitions; two relation slots capping network growth is a candidate explanation and is untested. The panel now prints BOTH policies side by side. That was a correction mid-task: the first version printed only the new one and I compared it against a figure remembered from CB-WP-0047 -- a comparison against a board nobody re-ran. Control that makes the numbers mean anything: under SHARED GROUND the two policies agree at all but <=2 decision points across 12 boards, so a moving column is mode-awareness and not simply a different bot. Also: two T01 tests keyed on `status: proposed`, which ground-game renamed to `ready-for-implement` mid-session. They now find the module by asking resolve() -- the structural property is ours and does not move when another repo edits its vocabulary. Also: `make vendor` replaces three hand re-vendors with a tool that regenerates digests by walking editions/, and reports one-sided files rather than resolving them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 23:31:11 +02:00
End on board clear, group collapse, or round-5 ceiling (v0 frozen).
utility_estimate: Unimplemented in kernel — package ready for clay-borg.
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
decision: none
CB-WP-0049 T02/T03: a seat that plays its objective, and F27 splits in two objective() reads GroundState::score (now public) rather than restating what winning is; a copy in the bot would disagree with the kernel the first time ground-game rules on F28. Working out WHERE the modes can differ was most of the task and it bounds the result: SOLVE always claims for the actor, so own-score and group-score want the same SOLVE nearly everywhere. That is a fact about GROUND's action set, not a shortcoming of the bot. Two real divergences, both readable off the table: SUPPORT regulates someone else (worth less against a rival, worth MORE under coalitions where a Bond merges them into my side), and SOLVE's value is the card's value, which greedy ignores entirely. THE RESULT — F27 splits in two: group success UNCHANGED in 34 of 36 cells who wins MOVES: BONDED COALITIONS at 4p goes 2.04 -> 2.98, 2.12 -> 3.29, 2.05 -> 3.01 winning seats per game So "the competitive modes are scoring lenses over cooperative play" was too strong and is withdrawn. The sharper claim: GROUND's scoring modes change WHO WINS, not WHETHER THE GROUP SUCCEEDS. And the effect is seat-band dependent -- 2p none, 4p largest, 6p none under coalitions; two relation slots capping network growth is a candidate explanation and is untested. The panel now prints BOTH policies side by side. That was a correction mid-task: the first version printed only the new one and I compared it against a figure remembered from CB-WP-0047 -- a comparison against a board nobody re-ran. Control that makes the numbers mean anything: under SHARED GROUND the two policies agree at all but <=2 decision points across 12 boards, so a moving column is mode-awareness and not simply a different bot. Also: two T01 tests keyed on `status: proposed`, which ground-game renamed to `ready-for-implement` mid-session. They now find the module by asking resolve() -- the structural property is ours and does not move when another repo edits its vocabulary. Also: `make vendor` replaces three hand re-vendors with a tool that regenerates digests by walking editions/, and reports one-sided files rather than resolving them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 23:31:11 +02:00
clay_borg_notes: >
v0 frozen in MODULE.md. Implement alone then profile scoped_plus_hybrid_end.
allow_proposed/ready-for-implement: refuse measure-as-green until kernel ships.
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
# --- problem_deal ---
- module_id: problem_deal.fixed_setup
ADR-0022 + CB-WP-0048 T00: the selector decision, and the mirror held The maintainer's observation decided the design: aspects partition the GAME, strata partition our apparatus, and they are orthogonal. A module is one coordinate change in aspect space with an obligation in every stratum. So aspect identity must NOT be Rust types -- an aspect ground-game adds would make clay-borg fail to parse a configuration rather than fail to run it, welding the two coordinate systems at the one place they must stay independent. Chosen: identity as data (Configuration round-trips anything the catalog names), behaviour exhaustive (Rules, no catch-all), resolve() between. Decisive argument: the catalog ALREADY ships modules with a rules_delta and status: proposed, so a per-aspect enum would report them as "unknown module" -- indistinguishable from a typo, a false statement about the edition, and this project's signature failure shape. Two facts need two errors. Federating design authority is permanent, so the representation must outlive the implementation. Legacy ids alias forever through the catalog's own legacy_experiment_id, on the standard-Np precedent: 26 recordings name them and the expansion is exact, so there is nothing to deprecate. T00 done: the schema-2 mirror had arrived with no digests (19 files) and edition-check was red. Digests are now generated by WALKING editions/, not typed -- two reviews already found hand-written lists that made their own controls vacuous, and a mirror that grows a directory is what breaks a maintained list. PROVENANCE-catalog.md was a file inside the mirrored tree that upstream does not have; folded into our own PROVENANCE.md, since provenance about the mirror does not belong inside the thing it describes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 22:07:57 +02:00
aspect: problem_deal
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
path: editions/modules/problem_deal/fixed_setup
is_default: true
selectable: true
status: baseline-default
rules_delta: null
summary: Surface + hidden 1..k at setup only (r0).
decision: none
- module_id: problem_deal.pressure_deck
ADR-0022 + CB-WP-0048 T00: the selector decision, and the mirror held The maintainer's observation decided the design: aspects partition the GAME, strata partition our apparatus, and they are orthogonal. A module is one coordinate change in aspect space with an obligation in every stratum. So aspect identity must NOT be Rust types -- an aspect ground-game adds would make clay-borg fail to parse a configuration rather than fail to run it, welding the two coordinate systems at the one place they must stay independent. Chosen: identity as data (Configuration round-trips anything the catalog names), behaviour exhaustive (Rules, no catch-all), resolve() between. Decisive argument: the catalog ALREADY ships modules with a rules_delta and status: proposed, so a per-aspect enum would report them as "unknown module" -- indistinguishable from a typo, a false statement about the edition, and this project's signature failure shape. Two facts need two errors. Federating design authority is permanent, so the representation must outlive the implementation. Legacy ids alias forever through the catalog's own legacy_experiment_id, on the standard-Np precedent: 26 recordings name them and the expansion is exact, so there is nothing to deprecate. T00 done: the schema-2 mirror had arrived with no digests (19 files) and edition-check was red. Digests are now generated by WALKING editions/, not typed -- two reviews already found hand-written lists that made their own controls vacuous, and a mirror that grows a directory is what breaks a maintained list. PROVENANCE-catalog.md was a file inside the mirrored tree that upstream does not have; folded into our own PROVENANCE.md, since provenance about the mirror does not belong inside the thing it describes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 22:07:57 +02:00
aspect: problem_deal
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
path: editions/modules/problem_deal/pressure_deck
is_default: false
selectable: true
CB-WP-0049 T02/T03: a seat that plays its objective, and F27 splits in two objective() reads GroundState::score (now public) rather than restating what winning is; a copy in the bot would disagree with the kernel the first time ground-game rules on F28. Working out WHERE the modes can differ was most of the task and it bounds the result: SOLVE always claims for the actor, so own-score and group-score want the same SOLVE nearly everywhere. That is a fact about GROUND's action set, not a shortcoming of the bot. Two real divergences, both readable off the table: SUPPORT regulates someone else (worth less against a rival, worth MORE under coalitions where a Bond merges them into my side), and SOLVE's value is the card's value, which greedy ignores entirely. THE RESULT — F27 splits in two: group success UNCHANGED in 34 of 36 cells who wins MOVES: BONDED COALITIONS at 4p goes 2.04 -> 2.98, 2.12 -> 3.29, 2.05 -> 3.01 winning seats per game So "the competitive modes are scoring lenses over cooperative play" was too strong and is withdrawn. The sharper claim: GROUND's scoring modes change WHO WINS, not WHETHER THE GROUP SUCCEEDS. And the effect is seat-band dependent -- 2p none, 4p largest, 6p none under coalitions; two relation slots capping network growth is a candidate explanation and is untested. The panel now prints BOTH policies side by side. That was a correction mid-task: the first version printed only the new one and I compared it against a figure remembered from CB-WP-0047 -- a comparison against a board nobody re-ran. Control that makes the numbers mean anything: under SHARED GROUND the two policies agree at all but <=2 decision points across 12 boards, so a moving column is mode-awareness and not simply a different bot. Also: two T01 tests keyed on `status: proposed`, which ground-game renamed to `ready-for-implement` mid-session. They now find the module by asking resolve() -- the structural property is ours and does not move when another repo edits its vocabulary. Also: `make vendor` replaces three hand re-vendors with a tool that regenerates digests by walking editions/, and reports one-sided files rather than resolving them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 23:31:11 +02:00
status: ready-for-implement
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
rules_delta: editions/modules/problem_deal/pressure_deck/rules_delta.yaml
hypothesis_ref: history/260808-deal-end-sequences-design.md
summary: >
CB-WP-0049 T02/T03: a seat that plays its objective, and F27 splits in two objective() reads GroundState::score (now public) rather than restating what winning is; a copy in the bot would disagree with the kernel the first time ground-game rules on F28. Working out WHERE the modes can differ was most of the task and it bounds the result: SOLVE always claims for the actor, so own-score and group-score want the same SOLVE nearly everywhere. That is a fact about GROUND's action set, not a shortcoming of the bot. Two real divergences, both readable off the table: SUPPORT regulates someone else (worth less against a rival, worth MORE under coalitions where a Bond merges them into my side), and SOLVE's value is the card's value, which greedy ignores entirely. THE RESULT — F27 splits in two: group success UNCHANGED in 34 of 36 cells who wins MOVES: BONDED COALITIONS at 4p goes 2.04 -> 2.98, 2.12 -> 3.29, 2.05 -> 3.01 winning seats per game So "the competitive modes are scoring lenses over cooperative play" was too strong and is withdrawn. The sharper claim: GROUND's scoring modes change WHO WINS, not WHETHER THE GROUP SUCCEEDS. And the effect is seat-band dependent -- 2p none, 4p largest, 6p none under coalitions; two relation slots capping network growth is a candidate explanation and is untested. The panel now prints BOTH policies side by side. That was a correction mid-task: the first version printed only the new one and I compared it against a figure remembered from CB-WP-0047 -- a comparison against a board nobody re-ran. Control that makes the numbers mean anything: under SHARED GROUND the two policies agree at all but <=2 decision points across 12 boards, so a moving column is mode-awareness and not simply a different bot. Also: two T01 tests keyed on `status: proposed`, which ground-game renamed to `ready-for-implement` mid-session. They now find the module by asking resolve() -- the structural property is ours and does not move when another repo edits its vocabulary. Also: `make vendor` replaces three hand re-vendors with a tool that regenerates digests by walking editions/, and reports one-sided files rather than resolving them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 23:31:11 +02:00
Reduced starters + Pressure deck draws; drawn cards 0 points (v0 frozen).
utility_estimate: Unimplemented in kernel — package ready for clay-borg.
CB-RES-0009: extensive form is the lingua franca Two questions from the maintainer — is there a game-theory mapping to Ludii's language, and is that language formal enough to derive one from. Yes, no, and the no does not matter. The mapping is proven, not to be invented: "The Ludii Game Description Language is Universal" shows the language can represent an equivalent game for any finite, non-deterministic, imperfect-information game, extending earlier work limited to finite deterministic fully-observable extensive-form games. EFG is also OpenSpiel's object, so the same formalism connects description to analysis: Ludii -> EFG <- OpenSpiel. Ludii's syntax is formal and unusually so — a class grammar derived automatically from its source. Its semantics are its Java: a ludeme means what its class does, and Ludii effectively makes Java the game description language. So there is no independent calculus to extract. The formality lives in the universality RESULT, not in a definition of meaning. GDL has the semantics and pays for it in speed — six times on Gomoku, twenty on Amazons and Hex, over two hundred on Chess. Conclusion: do not derive a language from Ludii; target the EFG directly. And we are closer than the tracks assumed. The journal is the history, Outcome is the payoff, legal_commands gives the actions — and project(Viewer::Player(seat)) IS the information partition, built so a player is not shown another's hand and unremarked as exactly the machinery imperfect information needs. Three gaps: chance is folded into a seed so a game is one realisation rather than a game with chance nodes; perfect recall is unasserted, which CFR and exploitability both assume; and commit/reveal is the standard EFG encoding of simultaneity but is never stated as such. Perfect recall is checkable from the journal today and is now Track B's first task — if it fails, every equilibrium concept we might quote is unsound here. Also re-vendored the catalog twice: ground-game added H2 — scoped problem stress, applying End Stress by personal/bond/global scope instead of flat to everyone, which is a direct response to our reading that H1's tax scales with the Problems while its intended effect does not. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 14:57:25 +02:00
decision: none
clay_borg_notes: >
CB-WP-0049 T02/T03: a seat that plays its objective, and F27 splits in two objective() reads GroundState::score (now public) rather than restating what winning is; a copy in the bot would disagree with the kernel the first time ground-game rules on F28. Working out WHERE the modes can differ was most of the task and it bounds the result: SOLVE always claims for the actor, so own-score and group-score want the same SOLVE nearly everywhere. That is a fact about GROUND's action set, not a shortcoming of the bot. Two real divergences, both readable off the table: SUPPORT regulates someone else (worth less against a rival, worth MORE under coalitions where a Bond merges them into my side), and SOLVE's value is the card's value, which greedy ignores entirely. THE RESULT — F27 splits in two: group success UNCHANGED in 34 of 36 cells who wins MOVES: BONDED COALITIONS at 4p goes 2.04 -> 2.98, 2.12 -> 3.29, 2.05 -> 3.01 winning seats per game So "the competitive modes are scoring lenses over cooperative play" was too strong and is withdrawn. The sharper claim: GROUND's scoring modes change WHO WINS, not WHETHER THE GROUP SUCCEEDS. And the effect is seat-band dependent -- 2p none, 4p largest, 6p none under coalitions; two relation slots capping network growth is a candidate explanation and is untested. The panel now prints BOTH policies side by side. That was a correction mid-task: the first version printed only the new one and I compared it against a figure remembered from CB-WP-0047 -- a comparison against a board nobody re-ran. Control that makes the numbers mean anything: under SHARED GROUND the two policies agree at all but <=2 decision points across 12 boards, so a moving column is mode-awareness and not simply a different bot. Also: two T01 tests keyed on `status: proposed`, which ground-game renamed to `ready-for-implement` mid-session. They now find the module by asking resolve() -- the structural property is ours and does not move when another repo edits its vocabulary. Also: `make vendor` replaces three hand re-vendors with a tool that regenerates digests by walking editions/, and reports one-sided files rather than resolving them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 23:31:11 +02:00
v0 frozen in MODULE.md. Drawn point_value 0; thresholds unchanged.
Compose with problem_stress.scoped via profile scoped_plus_pressure_deck.
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
# ---------------------------------------------------------------------------
# Profiles — named compositions (convenience; not a second rules source)
# Selection = baseline + modules list. Defaults fill missing axes.
# ---------------------------------------------------------------------------
profiles:
- profile_id: baseline
title: Pure r0
modules: []
ADR-0022 + CB-WP-0048 T00: the selector decision, and the mirror held The maintainer's observation decided the design: aspects partition the GAME, strata partition our apparatus, and they are orthogonal. A module is one coordinate change in aspect space with an obligation in every stratum. So aspect identity must NOT be Rust types -- an aspect ground-game adds would make clay-borg fail to parse a configuration rather than fail to run it, welding the two coordinate systems at the one place they must stay independent. Chosen: identity as data (Configuration round-trips anything the catalog names), behaviour exhaustive (Rules, no catch-all), resolve() between. Decisive argument: the catalog ALREADY ships modules with a rules_delta and status: proposed, so a per-aspect enum would report them as "unknown module" -- indistinguishable from a typo, a false statement about the edition, and this project's signature failure shape. Two facts need two errors. Federating design authority is permanent, so the representation must outlive the implementation. Legacy ids alias forever through the catalog's own legacy_experiment_id, on the standard-Np precedent: 26 recordings name them and the expansion is exact, so there is nothing to deprecate. T00 done: the schema-2 mirror had arrived with no digests (19 files) and edition-check was red. Digests are now generated by WALKING editions/, not typed -- two reviews already found hand-written lists that made their own controls vacuous, and a mirror that grows a directory is what breaks a maintained list. PROVENANCE-catalog.md was a file inside the mirrored tree that upstream does not have; folded into our own PROVENANCE.md, since provenance about the mirror does not belong inside the thing it describes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 22:07:57 +02:00
summary: All aspect defaults — printed ground-darvo-r0 behaviour.
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
- profile_id: h1
title: Legacy H1 (flat problem stress + attack soothe)
modules:
- problem_stress.flat_any_open
- attack_relief.self_soothe_ge4
legacy_experiment_id: h1-problem-stress
summary: Equivalent to old monolithic experiment h1-problem-stress.
decision: reject-as-baseline
- profile_id: h2
title: Legacy H2 (scoped problem stress only)
modules:
- problem_stress.scoped
legacy_experiment_id: h2-scoped-problem-stress
summary: Equivalent to old monolithic experiment h2-scoped-problem-stress.
decision: keep-as-experiment
- profile_id: scoped_plus_attack_soothe
title: Scoped stress + ATTACK self-soothe
modules:
- problem_stress.scoped
- attack_relief.self_soothe_ge4
ADR-0022 + CB-WP-0048 T00: the selector decision, and the mirror held The maintainer's observation decided the design: aspects partition the GAME, strata partition our apparatus, and they are orthogonal. A module is one coordinate change in aspect space with an obligation in every stratum. So aspect identity must NOT be Rust types -- an aspect ground-game adds would make clay-borg fail to parse a configuration rather than fail to run it, welding the two coordinate systems at the one place they must stay independent. Chosen: identity as data (Configuration round-trips anything the catalog names), behaviour exhaustive (Rules, no catch-all), resolve() between. Decisive argument: the catalog ALREADY ships modules with a rules_delta and status: proposed, so a per-aspect enum would report them as "unknown module" -- indistinguishable from a typo, a false statement about the edition, and this project's signature failure shape. Two facts need two errors. Federating design authority is permanent, so the representation must outlive the implementation. Legacy ids alias forever through the catalog's own legacy_experiment_id, on the standard-Np precedent: 26 recordings name them and the expansion is exact, so there is nothing to deprecate. T00 done: the schema-2 mirror had arrived with no digests (19 files) and edition-check was red. Digests are now generated by WALKING editions/, not typed -- two reviews already found hand-written lists that made their own controls vacuous, and a mirror that grows a directory is what breaks a maintained list. PROVENANCE-catalog.md was a file inside the mirrored tree that upstream does not have; folded into our own PROVENANCE.md, since provenance about the mirror does not belong inside the thing it describes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 22:07:57 +02:00
summary: First intentional multi-aspect combo after modular catalog.
CB-WP-0047: all four boards, and every mode named on the page The modes were already implemented; nothing had ever COMPARED them. The scenarios were not implemented at all: edition::deal has taken a scenario_id since it was written and the only caller passed the literal "SCN_01", so 15 of 20 Problem cards had never been dealt by anything. The seam was the whole mechanism and it sat unused, with nothing red because nothing asked. Scenario is now state (serde default SCN_01, so all 26 recordings replay unchanged), selected by preset `scn-03-4p` with `standard-Np` still meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or titles, validated against the edition rather than a pattern. The threshold now comes off the Scenario card, closing F25's hardcoded 5/7/9. The first version of that control was worthless and mutation said so: all four scenarios print 5/7/9, so reverting to the bands left it green. Split threshold_from() so it can be handed a card that disagrees. The header read `scoring CommonProblem` where the Mode card is titled COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the move buttons, still standing on the line that says what winning means. The coverage probe was matching that Debug output and went red when it was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the premise, the mode's rules text, and the tiebreak. scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same board (identical cells, pinned by a characterisation test); SCN_04 is the hard board at 2p (52% vs 67/73%, the only deck needing two Repair); and group success is EXACTLY equal across all three modes in all 36 cells, because greedy never reads state.mode -- filed F27, the two competitive modes are scoring lenses over cooperative play. F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT where the mode card's shared score is claimed VALUE. Raised, not fixed; scoring is ground-game's to rule on. Also fixes design.py reporting a backticked path as no reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00
status: unmeasured
- profile_id: scoped_plus_hybrid_end
title: Scoped stress + hybrid end (when end module ships)
modules:
- problem_stress.scoped
- end_condition.hybrid_clear_collapse
summary: Requires end_condition.hybrid_clear_collapse kernel support.
status: proposed
- profile_id: scoped_plus_pressure_deck
title: Scoped stress + pressure deck (when deal module ships)
modules:
- problem_stress.scoped
- problem_deal.pressure_deck
summary: Requires problem_deal.pressure_deck kernel support.
status: proposed
# ---------------------------------------------------------------------------
# Legacy experiment paths (still on disk; prefer modules + profiles)
# ---------------------------------------------------------------------------
legacy_experiments:
- experiment_id: h1-problem-stress
path: editions/experiments/h1-problem-stress
equivalent_profile: h1
note: Prefer profile h1 or modules problem_stress.flat_any_open + attack_relief.self_soothe_ge4
- experiment_id: h2-scoped-problem-stress
path: editions/experiments/h2-scoped-problem-stress
equivalent_profile: h2
note: Prefer profile h2 or module problem_stress.scoped