CB-WP-0038: variant selection, H1 implemented, and H1 measured
ground-game packages hypotheses as selectable rules variants — a catalog,
a rules_delta.yaml, and prose — and their note is explicit that CSV text
alone is not executable here. So the kernel gains a Variant in game state:
in the state, therefore in the hash, therefore in the recording, because a
scenario replayed under a different variant would diverge silently.
Baseline is bit-for-bit what it was, asserted across seat counts and
seeds. A variant system that perturbs the baseline invalidates every
measurement this repo has.
H1-A and H1-B implemented from rules_delta.yaml and mutation-proven on
their own defects: "unclaimed" misread as face-up-and-unsolved, and the
attacker's Stress read after the attack's effects. Their `unchanged:` list
is asserted rather than trusted — that list is their claim about their own
experiment.
Measured, and three of their four criteria fail. DARVO arm rate is still
0 under greedy; ATTACK selection does not rise and falls for the rank-75
policy; group success collapses from 165/190/200 to 0 at 3/4/6 seats.
The mechanism is not the assumed one: greedy answers the pressure by
regulating, Stress plateaus at 3, so it never reaches the gate at 4 or the
arm at 5 — H1-A acts as a solve-rate tax and H1-B is unreachable under
competent play.
A harness defect was caught before the claim: sweep discarded refused
games silently and never reported its count, so "nobody won" and "nothing
played" printed identically. Reporting H1 as unwinnable on that basis
would have been the ADR-0018 family aimed at another repo's design. All
200 games ran in every cell; the zeros are real.
Chaos d8 = 8 — the window's first override, redrew L against a structural
L, so it changed nothing. Window 3 recorded in ChaosRollHistory.
NOT REVIEWED: tier L owes a separate-agent adversarial review, and no H1
result may reach ground-game until it has run.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 00:50:08 +02:00
|
|
|
|
# Selectable GROUND edition / rules packages.
|
|
|
|
|
|
# Schema docs: CATALOG.md
|
|
|
|
|
|
# clay-borg: pin by variant_id → path (+ optional git_pin).
|
|
|
|
|
|
|
|
|
|
|
|
schema_version: 1
|
2026-08-08 11:43:13 +02:00
|
|
|
|
updated: "2026-08-08"
|
CB-WP-0038: variant selection, H1 implemented, and H1 measured
ground-game packages hypotheses as selectable rules variants — a catalog,
a rules_delta.yaml, and prose — and their note is explicit that CSV text
alone is not executable here. So the kernel gains a Variant in game state:
in the state, therefore in the hash, therefore in the recording, because a
scenario replayed under a different variant would diverge silently.
Baseline is bit-for-bit what it was, asserted across seat counts and
seeds. A variant system that perturbs the baseline invalidates every
measurement this repo has.
H1-A and H1-B implemented from rules_delta.yaml and mutation-proven on
their own defects: "unclaimed" misread as face-up-and-unsolved, and the
attacker's Stress read after the attack's effects. Their `unchanged:` list
is asserted rather than trusted — that list is their claim about their own
experiment.
Measured, and three of their four criteria fail. DARVO arm rate is still
0 under greedy; ATTACK selection does not rise and falls for the rank-75
policy; group success collapses from 165/190/200 to 0 at 3/4/6 seats.
The mechanism is not the assumed one: greedy answers the pressure by
regulating, Stress plateaus at 3, so it never reaches the gate at 4 or the
arm at 5 — H1-A acts as a solve-rate tax and H1-B is unreachable under
competent play.
A harness defect was caught before the claim: sweep discarded refused
games silently and never reported its count, so "nobody won" and "nothing
played" printed identically. Reporting H1 as unwinnable on that basis
would have been the ADR-0018 family aimed at another repo's design. All
200 games ran in every cell; the zeros are real.
Chaos d8 = 8 — the window's first override, redrew L against a structural
L, so it changed nothing. Window 3 recorded in ChaosRollHistory.
NOT REVIEWED: tier L owes a separate-agent adversarial review, and no H1
result may reach ground-game until it has run.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 00:50:08 +02:00
|
|
|
|
default_variant: ground-darvo-r0
|
CB-RES-0009: extensive form is the lingua franca
Two questions from the maintainer — is there a game-theory mapping to
Ludii's language, and is that language formal enough to derive one from.
Yes, no, and the no does not matter.
The mapping is proven, not to be invented: "The Ludii Game Description
Language is Universal" shows the language can represent an equivalent game
for any finite, non-deterministic, imperfect-information game, extending
earlier work limited to finite deterministic fully-observable
extensive-form games. EFG is also OpenSpiel's object, so the same
formalism connects description to analysis: Ludii -> EFG <- OpenSpiel.
Ludii's syntax is formal and unusually so — a class grammar derived
automatically from its source. Its semantics are its Java: a ludeme means
what its class does, and Ludii effectively makes Java the game description
language. So there is no independent calculus to extract. The formality
lives in the universality RESULT, not in a definition of meaning. GDL has
the semantics and pays for it in speed — six times on Gomoku, twenty on
Amazons and Hex, over two hundred on Chess.
Conclusion: do not derive a language from Ludii; target the EFG directly.
And we are closer than the tracks assumed. The journal is the history,
Outcome is the payoff, legal_commands gives the actions — and
project(Viewer::Player(seat)) IS the information partition, built so a
player is not shown another's hand and unremarked as exactly the machinery
imperfect information needs.
Three gaps: chance is folded into a seed so a game is one realisation
rather than a game with chance nodes; perfect recall is unasserted, which
CFR and exploitability both assume; and commit/reveal is the standard EFG
encoding of simultaneity but is never stated as such. Perfect recall is
checkable from the journal today and is now Track B's first task — if it
fails, every equilibrium concept we might quote is unsound here.
Also re-vendored the catalog twice: ground-game added H2 — scoped problem
stress, applying End Stress by personal/bond/global scope instead of flat
to everyone, which is a direct response to our reading that H1's tax
scales with the Problems while its intended effect does not.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 14:57:25 +02:00
|
|
|
|
# H2 added; H1 remains selectable as reject-as-baseline control.
|
CB-WP-0038: variant selection, H1 implemented, and H1 measured
ground-game packages hypotheses as selectable rules variants — a catalog,
a rules_delta.yaml, and prose — and their note is explicit that CSV text
alone is not executable here. So the kernel gains a Variant in game state:
in the state, therefore in the hash, therefore in the recording, because a
scenario replayed under a different variant would diverge silently.
Baseline is bit-for-bit what it was, asserted across seat counts and
seeds. A variant system that perturbs the baseline invalidates every
measurement this repo has.
H1-A and H1-B implemented from rules_delta.yaml and mutation-proven on
their own defects: "unclaimed" misread as face-up-and-unsolved, and the
attacker's Stress read after the attack's effects. Their `unchanged:` list
is asserted rather than trusted — that list is their claim about their own
experiment.
Measured, and three of their four criteria fail. DARVO arm rate is still
0 under greedy; ATTACK selection does not rise and falls for the rank-75
policy; group success collapses from 165/190/200 to 0 at 3/4/6 seats.
The mechanism is not the assumed one: greedy answers the pressure by
regulating, Stress plateaus at 3, so it never reaches the gate at 4 or the
arm at 5 — H1-A acts as a solve-rate tax and H1-B is unreachable under
competent play.
A harness defect was caught before the claim: sweep discarded refused
games silently and never reported its count, so "nobody won" and "nothing
played" printed identically. Reporting H1 as unwinnable on that basis
would have been the ADR-0018 family aimed at another repo's design. All
200 games ran in every cell; the zeros are real.
Chaos d8 = 8 — the window's first override, redrew L against a structural
L, so it changed nothing. Window 3 recorded in ChaosRollHistory.
NOT REVIEWED: tier L owes a separate-agent adversarial review, and no H1
result may reach ground-game until it has run.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 00:50:08 +02:00
|
|
|
|
|
|
|
|
|
|
packages:
|
|
|
|
|
|
- variant_id: ground-darvo-r0
|
|
|
|
|
|
path: editions/ground-darvo-r0
|
|
|
|
|
|
kind: baseline
|
|
|
|
|
|
base: null
|
|
|
|
|
|
selectable: true
|
|
|
|
|
|
status: baseline
|
|
|
|
|
|
dataset_id: GROUND-DARVO-CORE-0.1
|
|
|
|
|
|
title: "GROUND DARVO Edition — core r0"
|
|
|
|
|
|
summary: >
|
|
|
|
|
|
Playtest baseline. Surface + hidden 1..k deal; available 6/9/12;
|
|
|
|
|
|
thresholds 5/7/9. Stress ledger has no problem pressure; ATTACK
|
|
|
|
|
|
does not self-soothe.
|
|
|
|
|
|
hypothesis_ref: null
|
|
|
|
|
|
workplan_ref: null
|
|
|
|
|
|
rules_delta: null
|
|
|
|
|
|
changed_files: []
|
|
|
|
|
|
utility_estimate: >
|
2026-08-08 11:43:13 +02:00
|
|
|
|
Ship-default. RPT-0003 + CB-EV-0031: under non-attacking competent
|
|
|
|
|
|
policies, peak Stress held never exceeds start 2 — ATTACK is the
|
|
|
|
|
|
sole inbound pressure. Keep pinned until a successor is accepted.
|
CB-WP-0038: variant selection, H1 implemented, and H1 measured
ground-game packages hypotheses as selectable rules variants — a catalog,
a rules_delta.yaml, and prose — and their note is explicit that CSV text
alone is not executable here. So the kernel gains a Variant in game state:
in the state, therefore in the hash, therefore in the recording, because a
scenario replayed under a different variant would diverge silently.
Baseline is bit-for-bit what it was, asserted across seat counts and
seeds. A variant system that perturbs the baseline invalidates every
measurement this repo has.
H1-A and H1-B implemented from rules_delta.yaml and mutation-proven on
their own defects: "unclaimed" misread as face-up-and-unsolved, and the
attacker's Stress read after the attack's effects. Their `unchanged:` list
is asserted rather than trusted — that list is their claim about their own
experiment.
Measured, and three of their four criteria fail. DARVO arm rate is still
0 under greedy; ATTACK selection does not rise and falls for the rank-75
policy; group success collapses from 165/190/200 to 0 at 3/4/6 seats.
The mechanism is not the assumed one: greedy answers the pressure by
regulating, Stress plateaus at 3, so it never reaches the gate at 4 or the
arm at 5 — H1-A acts as a solve-rate tax and H1-B is unreachable under
competent play.
A harness defect was caught before the claim: sweep discarded refused
games silently and never reported its count, so "nobody won" and "nothing
played" printed identically. Reporting H1 as unwinnable on that basis
would have been the ADR-0018 family aimed at another repo's design. All
200 games ran in every cell; the zeros are real.
Chaos d8 = 8 — the window's first override, redrew L against a structural
L, so it changed nothing. Window 3 recorded in ChaosRollHistory.
NOT REVIEWED: tier L owes a separate-agent adversarial review, and no H1
result may reach ground-game until it has run.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 00:50:08 +02:00
|
|
|
|
decision: none
|
|
|
|
|
|
git_pin: null
|
|
|
|
|
|
clay_borg_notes: >
|
|
|
|
|
|
Vendor CSVs as today (Problems, Actions, Solutions, Modes, Tokens).
|
|
|
|
|
|
Kernel implements printed r0 rules.
|
|
|
|
|
|
|
|
|
|
|
|
- variant_id: h1-problem-stress
|
|
|
|
|
|
path: editions/experiments/h1-problem-stress
|
|
|
|
|
|
kind: experiment
|
|
|
|
|
|
base: ground-darvo-r0
|
|
|
|
|
|
selectable: true
|
2026-08-08 11:43:13 +02:00
|
|
|
|
status: measured
|
CB-WP-0038: variant selection, H1 implemented, and H1 measured
ground-game packages hypotheses as selectable rules variants — a catalog,
a rules_delta.yaml, and prose — and their note is explicit that CSV text
alone is not executable here. So the kernel gains a Variant in game state:
in the state, therefore in the hash, therefore in the recording, because a
scenario replayed under a different variant would diverge silently.
Baseline is bit-for-bit what it was, asserted across seat counts and
seeds. A variant system that perturbs the baseline invalidates every
measurement this repo has.
H1-A and H1-B implemented from rules_delta.yaml and mutation-proven on
their own defects: "unclaimed" misread as face-up-and-unsolved, and the
attacker's Stress read after the attack's effects. Their `unchanged:` list
is asserted rather than trusted — that list is their claim about their own
experiment.
Measured, and three of their four criteria fail. DARVO arm rate is still
0 under greedy; ATTACK selection does not rise and falls for the rank-75
policy; group success collapses from 165/190/200 to 0 at 3/4/6 seats.
The mechanism is not the assumed one: greedy answers the pressure by
regulating, Stress plateaus at 3, so it never reaches the gate at 4 or the
arm at 5 — H1-A acts as a solve-rate tax and H1-B is unreachable under
competent play.
A harness defect was caught before the claim: sweep discarded refused
games silently and never reported its count, so "nobody won" and "nothing
played" printed identically. Reporting H1 as unwinnable on that basis
would have been the ADR-0018 family aimed at another repo's design. All
200 games ran in every cell; the zeros are real.
Chaos d8 = 8 — the window's first override, redrew L against a structural
L, so it changed nothing. Window 3 recorded in ChaosRollHistory.
NOT REVIEWED: tier L owes a separate-agent adversarial review, and no H1
result may reach ground-game until it has run.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 00:50:08 +02:00
|
|
|
|
dataset_id: GROUND-DARVO-EXP-H1-0.1
|
|
|
|
|
|
title: "H1 — problem pressure + high-stress ATTACK self-soothe"
|
|
|
|
|
|
summary: >
|
|
|
|
|
|
Two deltas only: (A) +1 Stress each player at Round End if any
|
|
|
|
|
|
Problem remains unclaimed; (B) uncancelled ACT_ATTACK by a seat
|
|
|
|
|
|
at Stress ≥4 gives that seat −1 Stress. Thresholds/deal unchanged.
|
|
|
|
|
|
hypothesis_ref: history/260807-attack-darvo-stress-design.md
|
|
|
|
|
|
workplan_ref: workplans/GROUND-WP-0006-h1-problem-stress-experiment.md
|
2026-08-08 11:43:13 +02:00
|
|
|
|
measurement_ref: reports/260808-clay-borg-h1-measured.md
|
CB-WP-0038: variant selection, H1 implemented, and H1 measured
ground-game packages hypotheses as selectable rules variants — a catalog,
a rules_delta.yaml, and prose — and their note is explicit that CSV text
alone is not executable here. So the kernel gains a Variant in game state:
in the state, therefore in the hash, therefore in the recording, because a
scenario replayed under a different variant would diverge silently.
Baseline is bit-for-bit what it was, asserted across seat counts and
seeds. A variant system that perturbs the baseline invalidates every
measurement this repo has.
H1-A and H1-B implemented from rules_delta.yaml and mutation-proven on
their own defects: "unclaimed" misread as face-up-and-unsolved, and the
attacker's Stress read after the attack's effects. Their `unchanged:` list
is asserted rather than trusted — that list is their claim about their own
experiment.
Measured, and three of their four criteria fail. DARVO arm rate is still
0 under greedy; ATTACK selection does not rise and falls for the rank-75
policy; group success collapses from 165/190/200 to 0 at 3/4/6 seats.
The mechanism is not the assumed one: greedy answers the pressure by
regulating, Stress plateaus at 3, so it never reaches the gate at 4 or the
arm at 5 — H1-A acts as a solve-rate tax and H1-B is unreachable under
competent play.
A harness defect was caught before the claim: sweep discarded refused
games silently and never reported its count, so "nobody won" and "nothing
played" printed identically. Reporting H1 as unwinnable on that basis
would have been the ADR-0018 family aimed at another repo's design. All
200 games ran in every cell; the zeros are real.
Chaos d8 = 8 — the window's first override, redrew L against a structural
L, so it changed nothing. Window 3 recorded in ChaosRollHistory.
NOT REVIEWED: tier L owes a separate-agent adversarial review, and no H1
result may reach ground-game until it has run.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 00:50:08 +02:00
|
|
|
|
rules_delta: editions/experiments/h1-problem-stress/rules_delta.yaml
|
|
|
|
|
|
changed_files:
|
|
|
|
|
|
- Actions.csv
|
|
|
|
|
|
- Rules_Text.csv
|
|
|
|
|
|
- metadata.json
|
|
|
|
|
|
utility_estimate: >
|
2026-08-08 11:43:13 +02:00
|
|
|
|
Measured 2026-08-08 (CB-EV-0030/0031; instrument caveat applies).
|
|
|
|
|
|
Direction right, magnitude wrong: H1-A is a solve-rate tax that
|
|
|
|
|
|
competent seats absorb with GROUND (Stress ~3); H1-B never fires
|
|
|
|
|
|
for them. DARVO arms for unregulated seats; group wins collapse
|
|
|
|
|
|
to 0 at 3p+ under greedy SHARED GROUND. Do not promote as-is.
|
|
|
|
|
|
decision: reject-as-baseline
|
|
|
|
|
|
# Package stays selectable for regression compare; not a ship pin.
|
CB-WP-0038: variant selection, H1 implemented, and H1 measured
ground-game packages hypotheses as selectable rules variants — a catalog,
a rules_delta.yaml, and prose — and their note is explicit that CSV text
alone is not executable here. So the kernel gains a Variant in game state:
in the state, therefore in the hash, therefore in the recording, because a
scenario replayed under a different variant would diverge silently.
Baseline is bit-for-bit what it was, asserted across seat counts and
seeds. A variant system that perturbs the baseline invalidates every
measurement this repo has.
H1-A and H1-B implemented from rules_delta.yaml and mutation-proven on
their own defects: "unclaimed" misread as face-up-and-unsolved, and the
attacker's Stress read after the attack's effects. Their `unchanged:` list
is asserted rather than trusted — that list is their claim about their own
experiment.
Measured, and three of their four criteria fail. DARVO arm rate is still
0 under greedy; ATTACK selection does not rise and falls for the rank-75
policy; group success collapses from 165/190/200 to 0 at 3/4/6 seats.
The mechanism is not the assumed one: greedy answers the pressure by
regulating, Stress plateaus at 3, so it never reaches the gate at 4 or the
arm at 5 — H1-A acts as a solve-rate tax and H1-B is unreachable under
competent play.
A harness defect was caught before the claim: sweep discarded refused
games silently and never reported its count, so "nobody won" and "nothing
played" printed identically. Reporting H1 as unwinnable on that basis
would have been the ADR-0018 family aimed at another repo's design. All
200 games ran in every cell; the zeros are real.
Chaos d8 = 8 — the window's first override, redrew L against a structural
L, so it changed nothing. Window 3 recorded in ChaosRollHistory.
NOT REVIEWED: tier L owes a separate-agent adversarial review, and no H1
result may reach ground-game until it has run.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 00:50:08 +02:00
|
|
|
|
git_pin: null
|
|
|
|
|
|
clay_borg_notes: >
|
2026-08-08 11:43:13 +02:00
|
|
|
|
Kernel implements H1-A/H1-B. Keep for A/B against successors (H2…).
|
CB-RES-0009: extensive form is the lingua franca
Two questions from the maintainer — is there a game-theory mapping to
Ludii's language, and is that language formal enough to derive one from.
Yes, no, and the no does not matter.
The mapping is proven, not to be invented: "The Ludii Game Description
Language is Universal" shows the language can represent an equivalent game
for any finite, non-deterministic, imperfect-information game, extending
earlier work limited to finite deterministic fully-observable
extensive-form games. EFG is also OpenSpiel's object, so the same
formalism connects description to analysis: Ludii -> EFG <- OpenSpiel.
Ludii's syntax is formal and unusually so — a class grammar derived
automatically from its source. Its semantics are its Java: a ludeme means
what its class does, and Ludii effectively makes Java the game description
language. So there is no independent calculus to extract. The formality
lives in the universality RESULT, not in a definition of meaning. GDL has
the semantics and pays for it in speed — six times on Gomoku, twenty on
Amazons and Hex, over two hundred on Chess.
Conclusion: do not derive a language from Ludii; target the EFG directly.
And we are closer than the tracks assumed. The journal is the history,
Outcome is the payoff, legal_commands gives the actions — and
project(Viewer::Player(seat)) IS the information partition, built so a
player is not shown another's hand and unremarked as exactly the machinery
imperfect information needs.
Three gaps: chance is folded into a seed so a game is one realisation
rather than a game with chance nodes; perfect recall is unasserted, which
CFR and exploitability both assume; and commit/reveal is the standard EFG
encoding of simultaneity but is never stated as such. Perfect recall is
checkable from the journal today and is now Track B's first task — if it
fails, every equilibrium concept we might quote is unsound here.
Also re-vendored the catalog twice: ground-game added H2 — scoped problem
stress, applying End Stress by personal/bond/global scope instead of flat
to everyone, which is a direct response to our reading that H1's tax
scales with the Problems while its intended effect does not.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 14:57:25 +02:00
|
|
|
|
|
|
|
|
|
|
- variant_id: h2-scoped-problem-stress
|
|
|
|
|
|
path: editions/experiments/h2-scoped-problem-stress
|
|
|
|
|
|
kind: experiment
|
|
|
|
|
|
base: ground-darvo-r0
|
|
|
|
|
|
selectable: true
|
|
|
|
|
|
status: experimental
|
|
|
|
|
|
dataset_id: GROUND-DARVO-EXP-H2-0.1
|
|
|
|
|
|
title: "H2 — scoped problem stress (personal / bond / global)"
|
|
|
|
|
|
summary: >
|
|
|
|
|
|
Unclaimed Problems apply +1 End Stress only to stress_scope:
|
|
|
|
|
|
personal=owner, bond=owner's Bond network, global=all. Surface
|
|
|
|
|
|
global; priority 3 is bond (in play at 3p+). Anyone may SOLVE any
|
|
|
|
|
|
card. Not stacked on H1; no ATTACK self-soothe.
|
|
|
|
|
|
hypothesis_ref: history/260808-h2-scoped-problem-stress.md
|
|
|
|
|
|
workplan_ref: workplans/GROUND-WP-0007-h2-scoped-problem-stress.md
|
|
|
|
|
|
measurement_ref: null
|
|
|
|
|
|
rules_delta: editions/experiments/h2-scoped-problem-stress/rules_delta.yaml
|
|
|
|
|
|
changed_files:
|
|
|
|
|
|
- Problems.csv
|
|
|
|
|
|
- Rules_Text.csv
|
|
|
|
|
|
- metadata.json
|
|
|
|
|
|
utility_estimate: >
|
|
|
|
|
|
Unmeasured. Expected: stress variance up; group wins at 3–4p much
|
|
|
|
|
|
better than H1; bond cards create joint SOLVE incentive in networks.
|
|
|
|
|
|
decision: none
|
|
|
|
|
|
git_pin: null
|
|
|
|
|
|
clay_borg_notes: >
|
|
|
|
|
|
Implement rules_delta H2-SCOPE, H2-OWN, H2-A, H2-SOLVE on base r0.
|
|
|
|
|
|
Do not also apply H1 deltas. Problems.csv adds stress_scope column.
|
|
|
|
|
|
Report bond-card SOLVE rates and stress variance vs r0 and h1.
|