clay-borg/workplans/CB-WP-0038-h1-variant.md
tegwick 5a55af1a21
Some checks failed
ci / check (push) Failing after 3s
Sync hub ids for CB-WP-0038
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 00:50:51 +02:00

6.3 KiB
Raw Blame History

id kind title status state_hub_workstream_id
CB-WP-0038 product Variant selection, and H1 done 0976d470-7ec4-446d-827e-3a79c95ae63d

Purpose

structural tier  L   (a new capability port: the kernel gains a selectable
                      rules variant, so what the engine plays is no longer
                      a constant)
chaos            d8 = 8  → OVERRIDE, redraw = L
declared tier    L

Declaration 10 of chaos window 3, and the window's first 8. The redraw returned the structural tier, so the override changed nothing — recorded in ChaosRollHistory.md because a roll that changes nothing is exactly the observation ADR-0017's retirement condition needs.

Tier L owes a separate-agent adversarial review (InnerLoop Step 2), and it has not been done. The maintainer asked for implementation directly, so the code proceeds — but no H1 result may go to ground-game until that review has run. Measurement that reaches the game's owner is exactly what the review exists to check.

What ground-game built, and what it asks of us

ground-game has packaged hypothesis H1 as a selectable rules variant, not a patch:

  • editions/catalog.yamlvariant_id → path, default_variant: ground-darvo-r0
  • editions/experiments/h1-problem-stress/VARIANT.md (prose), rules_delta.yaml (machine-readable), overridden Actions.csv and Rules_Text.csv
  • design note history/260807-attack-darvo-stress-design.md with success criteria

Their note is explicit about our side:

"clay-borg hardcodes much of the rules kernel and vendors selected CSVs. H1-A/B are rules-kernel changes, not Problem table changes. CSV text alone is not executable in the current engine."

They are right, and F25 is the same observation from our side.

The two deltas

id phase rule
H1-A round end if any Problem in play is unclaimed (face-up unsolved, hidden, or Denied), each player +1 Stressbefore the clamp that arms DARVO
H1-B attack resolution an Attack that resolves and is not cancelled, by an attacker whose Stress was ≥ 4 before that Attack's effects, gives the attacker 1 Stress, after target and relation effects, then clamp

H1-B applies to the DARVO extra Attack too, because it shares the resolution path.

Task: selection

id: CB-WP-0038-T01
status: done
priority: high
state_hub_task_id: "20d9c7ca-187e-44f8-96ee-7f32f948c092"

A variant is chosen, not compiled in.

Controls:

  • the variant is part of game state and reaches the recording. A scenario replayed under the wrong variant would diverge silently, and the recording is the artifact everything else rests on;
  • baseline is the default, and every existing scenario replays byte-identically — this is the load-bearing control, because a variant system that perturbs the baseline invalidates every measurement we have;
  • the catalog is vendored with a digest, like every other borrowed file.

Done 2026-08-08. Variant in GroundState, #[serde(default)] so every scenario written before variants existed loads as baseline — which is what it was. Set beside mode, before the hash, so a recording replays under the variant it was played under. --variant on cb-play.

baseline_is_bit_for_bit_what_it_was compares state hashes across 2/3/6 seats and eight seeds. Selecting the baseline explicitly must equal not selecting anything, or every measurement this repo has is invalid.

Task: the deltas in the kernel

id: CB-WP-0038-T02
status: done
priority: high
state_hub_task_id: "19c893b7-443f-48e4-b85a-61686df98343"

Controls:

  • each delta fails on its own, by mutation;
  • rules_delta.yaml's unchanged: list is asserted, not trusted — deal and thresholds, SOLVE legality, Support, GROUND modes, the DARVO stage table, starting Stress 2, two relation slots. That list is ground-game's claim about their own experiment, and it is checkable;
  • "unclaimed" includes Denied and hidden, which is the clause a careless reading drops.

Done 2026-08-08. Both mutation-proven on their own defect:

mutation what went red
"unclaimed" read as face-up and unsolved "every player takes the pressure, not just some"
attacker Stress read after the attack's effects, and soothing a cancelled Attack "a cancelled Attack still soothed the attacker"

H1-A lands before the DARVO arm check, per the delta's own ordering; applying it after would leave the pressure unable to arm anything for a round, which is the opposite of the hypothesis.

h1_changes_nothing_it_said_it_would_not asserts ground-game's unchanged: list across all five seat counts — deal, thresholds, starting Stress 2, relation slots, mode, deck. That list is their claim about their own experiment, and taking it on faith is the habit CB-WP-0037 ended.

Task: re-run the panel

id: CB-WP-0038-T03
status: done
priority: high
state_hub_task_id: "03566bc7-2000-4a39-b8f8-c02de0a12c5a"

Their §3.2 criteria, on the same harness as RPT-0003.

Controls:

  • baseline and H1 measured by the same instrument in the same run, or the comparison is between two harnesses rather than two rule sets;
  • a single policy's numbers are not a fact about the game (RetrospectiveAnalysis §4.1) — the panel is plural or there is no result;
  • report what fails to move. H1 under-delivering is a result their §6 has a branch for, and reporting only the confirming half is how a measurement becomes advocacy.

Done 2026-08-08, and the result is largely negative. See CB-EV-0030.

A harness defect was caught first, and it mattered. The first run showed won, atk and darvo all exactly 0 at 3+ seats under H1 — which is the signature of games being discarded, not lost, because sweep had let Ok(g) = play(..) else { continue } and never reported its game count. Reporting "H1 makes the game unwinnable" when the truth might have been "our harness threw every H1 game away" would have been the ADR-0018 family aimed at another repo's design decision. Instrumented: all 200 games ran in every cell. The zeros are real.