fix(workplans): adopt ADR-007 derived identifiers
Records absent from central carried random pre-ADR-007 identifiers minted by the retired local hub, which C-06 refused as stale references. Deriving from the canonical record id takes no identity from anything. Refs CUST-WP-0068-T06 Assistant: claude-code Assistant-Model: opus Assistant-Process: 2583210@bnt-lap001 Assistant-Session: f2bff2d5-e9b2-4338-92ca-10282a927006
This commit is contained in:
parent
9458d0256a
commit
b5730f280b
55 changed files with 2937 additions and 150 deletions
171
reports/260808-clay-borg-boards-and-modes.md
Normal file
171
reports/260808-clay-borg-boards-and-modes.md
Normal file
|
|
@ -0,0 +1,171 @@
|
|||
---
|
||||
id: GROUND-RPT-0006
|
||||
type: report
|
||||
title: "clay-borg: all four Scenarios and all three Modes measured — modes change who wins, not whether the group survives"
|
||||
domain: consumer
|
||||
repo: ground-game
|
||||
status: informational
|
||||
owner: bernd
|
||||
topic_slug: whynot
|
||||
created: "2026-08-08"
|
||||
extends: GROUND-WP-0003
|
||||
---
|
||||
|
||||
# Boards and Modes measurement receipt (CB-EV-0033)
|
||||
|
||||
Upstream: clay-borg `evidence/CB-EV-0033-boards-and-modes.md`.
|
||||
|
||||
**Context:** clay-borg dealt `SCN_01` and only `SCN_01` for its whole
|
||||
existence — `deal()` took a scenario id and the one caller passed a
|
||||
literal, so **15 of the 20 Problem cards had never been dealt by
|
||||
anything**. Fixed on our side; all four boards and all three modes now
|
||||
run. 4 × 3 × 3 cells, 100 games each, two bots. Everything below is bots,
|
||||
no felt play.
|
||||
|
||||
---
|
||||
|
||||
## 1. SCN_01 and SCN_02 are the same board — is that intended?
|
||||
|
||||
Identical suit **and** value at every priority. Every measured cell
|
||||
matches exactly, at every seat band, under both bots.
|
||||
|
||||
A reskin is a legitimate choice and we are not calling this a defect. But
|
||||
it means **"four scenarios" is three boards**, and any study treating them
|
||||
as four independent samples double-counts one. We have pinned the current
|
||||
state with a test so a future divergence is a decision rather than a
|
||||
drift.
|
||||
|
||||
**Ask:** intended reskin, or did one deck get copied and not re-tuned?
|
||||
|
||||
---
|
||||
|
||||
## 2. SCN_04 is materially harder at 2 players
|
||||
|
||||
| board | 2p group success (/100) |
|
||||
|---|---|
|
||||
| SCN_01 / SCN_02 | 67 |
|
||||
| SCN_03 | 73 |
|
||||
| **SCN_04** | **52** |
|
||||
|
||||
**Your own data holds everything else fixed.** All four 2p deals are
|
||||
Surface + priorities 1–2, all four give 6 available points against a
|
||||
threshold of 5, all four start at Stress 2 over five rounds. The only
|
||||
difference is the suit multiset:
|
||||
|
||||
| board | 2p suits |
|
||||
|---|---|
|
||||
| SCN_01 / SCN_02 | Repair, Clarify, Boundary |
|
||||
| SCN_03 | Boundary, Clarify, Repair |
|
||||
| **SCN_04** | **Repair, Clarify, Repair** |
|
||||
|
||||
SCN_04 is the only 2p deal needing **two of one suit**.
|
||||
|
||||
**We are not claiming causation** — nothing here demonstrates the second
|
||||
Repair is what costs the 15 points. But the seat-band thresholds (5 / 7 /
|
||||
9) are uniform across boards, and at 2p the boards are not.
|
||||
|
||||
**Ask:** is a ~15-point spread at 2p acceptable board variety, or should
|
||||
the 2p threshold or SCN_04's priority-2 suit move?
|
||||
|
||||
---
|
||||
|
||||
## 3. Every board is a formality at 6 players
|
||||
|
||||
100/100 group success in **all twelve** 6p cells, both bots, every mode.
|
||||
Same shape as F17. Reported, not acted on.
|
||||
|
||||
---
|
||||
|
||||
## 4. The three Modes change **who wins**, not **whether the group
|
||||
survives**
|
||||
|
||||
We previously reported that the modes produced identical play. That was
|
||||
**our instrument, not your game** — the bot never read the Mode card. We
|
||||
built one that plays its seat's own objective and re-ran everything.
|
||||
|
||||
**Group success: unchanged in 34 of 36 cells.** (SCN_03 at 4p moves 99 →
|
||||
100 in all three modes — that is a bot refinement, not a mode effect; it
|
||||
moves under SHARED GROUND too.)
|
||||
|
||||
**Winning seats per game, BONDED COALITIONS:**
|
||||
|
||||
| board | 4p, mode-blind bot | 4p, mode-aware bot |
|
||||
|---|---|---|
|
||||
| SCN_01 / SCN_02 | 2.04 | **2.98** |
|
||||
| SCN_03 | 2.12 | **3.29** |
|
||||
| SCN_04 | 2.05 | **3.01** |
|
||||
|
||||
A seat whose score is its coalition's sum bonds harder, and coalitions
|
||||
come out about half again as large. COMMON PROBLEM moves at 6p (1.10 →
|
||||
1.17). Nothing else moves.
|
||||
|
||||
**Control:** under SHARED GROUND the two bots agree at all but ≤2
|
||||
decision points across twelve boards — so the moving columns are
|
||||
mode-awareness, not "a different bot".
|
||||
|
||||
**Design reading, offered not asserted:** the Mode card currently decides
|
||||
the *distribution of the win*, and the threshold decides survival
|
||||
independently of it. If a Mode is meant to make the group's task feel
|
||||
different — not just the podium — nothing we can measure says it
|
||||
currently does.
|
||||
|
||||
### The seat-band pattern is the interesting part
|
||||
|
||||
2p: nothing moves in any mode. 4p: the largest effect. 6p: **nothing**
|
||||
under BONDED COALITIONS.
|
||||
|
||||
Our candidate explanation is that **two relation slots per seat cap
|
||||
network growth**, so at 6p the incentive exists and cannot be acted on.
|
||||
**This is untested.** It is cheap for us to test if you want it: run with
|
||||
a third relation slot and see whether 6p coalition size then moves the
|
||||
way 4p's does.
|
||||
|
||||
**Ask:** is the two-slot limit meant to bind this hard at 5–6 players?
|
||||
|
||||
---
|
||||
|
||||
## 5. Two things needing your ruling
|
||||
|
||||
**(a) SHARED GROUND's mastery rating counts cards where the shared score
|
||||
counts points.** The Mode card reads *"All claimed Problem cards form one
|
||||
shared score. … For a mastery rating, subtract 1 for each Blame token
|
||||
still in play and 1 for each Denied Problem."* The shared score is
|
||||
claimed **value** (what the threshold is compared against); our engine
|
||||
subtracts those penalties from the claimed **count**. Both readings fit
|
||||
the sentence and they differ on every game where a 3-point Problem is
|
||||
claimed. We have not changed it — scoring is yours.
|
||||
|
||||
**(b) A package that adds a *file* is invisible to a consumer.**
|
||||
`h2-scoped-problem-stress` ships `Rules_Text.csv` — 22 passages of
|
||||
player-facing rules, including the one that says what a `stress_scope`
|
||||
does — and clay-borg had **never read the file**. Our freshness check
|
||||
verifies files it knows about; a file no reader mentions is not stale, it
|
||||
is unseen. We have since read the passage we needed; the other 21 are
|
||||
still unread.
|
||||
|
||||
**Ask:** could a package manifest name its own files, so that a consumer
|
||||
reading none of them is a detectable state rather than a silence?
|
||||
|
||||
---
|
||||
|
||||
## Standing limits on all of the above
|
||||
|
||||
- **No bot models a rival playing their objective.** A competitive mode
|
||||
in which nobody anticipates an opponent is a weak test of that mode.
|
||||
- **Bond-scoped joint SOLVE remains untested** (CB-EV-0032 criterion 4,
|
||||
unchanged) — these bots read the Mode, not the module.
|
||||
- **No felt play.** All bots, every number.
|
||||
|
||||
---
|
||||
|
||||
## Ground-game disposition (2026-08-08 review)
|
||||
|
||||
| # | Finding | Disposition |
|
||||
|---|---------|-------------|
|
||||
| 1 | SCN_01 ≡ SCN_02 mechanically | **Not treated as intentional.** Fiction reskin only; three distinct boards, not four. Residual: re-tune SCN_02 suit and/or values (content workplan). |
|
||||
| 2 | SCN_04 harder at 2p (~52 vs ~67–73) | **Acceptable variety for now** as a harder board; monitor. If table feel is too swingy, first lever is priority-2 suit (drop double Repair), not global thresholds. |
|
||||
| 3 | 6p 100% group success | **Known pattern** (F17 / difficulty line). Not fixed here; deal/pressure modules and WP-0005 disposition (Standard thresholds stand). |
|
||||
| 4 | Modes change *who* wins, not *whether* group succeeds | **Accepted as current design property**, not a defect. Threshold = survival; Mode = distribution. Mode-aware bot was clay-borg instrument fix. |
|
||||
| 5 | 2 relation slots mute 6p coalition effect | **Intentional cap** (complexity/time). Documented; third slot stays extension, not core r0. Optional clay-borg sensitivity run welcome. |
|
||||
| 6 | Mastery: points vs card count | **RULED: points.** Shared score and mastery penalties use **point values**, not card count. Modes.csv clarified. |
|
||||
| 7 | Package file manifest | **Accepted process ask.** Catalog/modules should declare `consumed_files` / overlays so unread files are detectable. Tracked under WP-0008 residual. |
|
||||
78
reports/260808-clay-borg-h1-measured.md
Normal file
78
reports/260808-clay-borg-h1-measured.md
Normal file
|
|
@ -0,0 +1,78 @@
|
|||
---
|
||||
id: GROUND-RPT-0004
|
||||
type: report
|
||||
title: "clay-borg: H1 measured — reaches unregulated seats; group success collapses"
|
||||
domain: consumer
|
||||
repo: ground-game
|
||||
status: informational
|
||||
owner: bernd
|
||||
topic_slug: whynot
|
||||
created: "2026-08-08"
|
||||
extends: GROUND-WP-0006
|
||||
---
|
||||
|
||||
# H1 measurement receipt (clay-borg CB-EV-0030 / CB-EV-0031)
|
||||
|
||||
Upstream: clay-borg `evidence/CB-EV-0030-h1-measured.md`,
|
||||
`evidence/CB-EV-0031-a-seat-that-does-not-regulate.md`, hub message
|
||||
2026-08-08. Variant selection + H1-A/H1-B implemented in their kernel from
|
||||
`rules_delta.yaml`. Baseline `unchanged:` list asserted.
|
||||
|
||||
**Instrument caveat (their words):** three rounds of adversarial review;
|
||||
twelve defects were in *their* harness, not in H1 design. Stable finding:
|
||||
**H1 reaches the unregulated seat, and the table stops winning.** Supporting
|
||||
numbers moved across corrections — treat policy-dependent tables as
|
||||
directional, not absolute.
|
||||
|
||||
No felt-play. SHARED GROUND dominant for regulation comparison; attack-value
|
||||
covers three modes for arms.
|
||||
|
||||
---
|
||||
|
||||
## Criteria §3.2 (design note)
|
||||
|
||||
| # | criterion | verdict |
|
||||
|---|-----------|---------|
|
||||
| 1 | DARVO arm rate non-trivial | **Policy-dependent.** 0 for greedy. **Met for unregulated `reactive`** (~1 arm/seat/game at 3p+). Under rank-75, H1 *reduces* arms to 0 at 3p+ |
|
||||
| 2 | ATTACK rises for some subpopulation | **Met only for unregulated seat.** rank-75 attacks **fall** under H1 (often to 0 at 3p+) |
|
||||
| 3 | group success does not collapse | **Fails hard.** Greedy SHARED: 165/190/200 → **0** at 3/4/6p; 2p 132 → 68 |
|
||||
| 4 | Bond/GROUND better than DARVO | Vacuous for winning policies (they never arm) |
|
||||
|
||||
## Mechanism (not what H1 assumed)
|
||||
|
||||
H1-A is answered by **regulation** (GROUND), not by climbing to DARVO:
|
||||
|
||||
- Stress plateaus ~3 under greedy (below gate 4 and arm 5).
|
||||
- Actions spent on stress hygiene → fewer SOLVEs → **solve-rate tax**.
|
||||
- **H1-B never fires** under competent play (Stress never ≥4).
|
||||
|
||||
Seed-1 3p illustration: baseline clear board / success; H1 ends Stress 3,3,3 with 2 unclaimed and under threshold.
|
||||
|
||||
Tax scales with open board (harder at 3p+); intended DARVO effect does not.
|
||||
|
||||
## Baseline Stress economy (larger than H1)
|
||||
|
||||
Under greedy **and** reactive on **baseline** r0: peak Stress **held** never exceeds start **2** — Stress only falls — when those policies never select ATTACK. ATTACK appears to be the **sole inbound pressure** for those policies. That is a sharper account of RPT-0003 than “ATTACK never pays”: **the pressure ATTACK would answer never occurs** unless someone already attacks.
|
||||
|
||||
(Always-attack rank-95 *does* drive Stress under baseline — so the economy is reachable; just not by non-attacking competent play.)
|
||||
|
||||
## Invariant (unexplained)
|
||||
|
||||
Under `reactive` + H1 at 3p/4p/6p: DARVO arms **exactly `seats × games`** (one arm per seat per game). clay-borg does not explain it. H1-B disabled → arm count unchanged (withdrawn “H1-B suppresses attacker DARVO” story).
|
||||
|
||||
## Withdrawn by clay-borg
|
||||
|
||||
Story that H1-B suppresses DARVO in the attacker / concentrates it in the target — from a broken multi-preference policy.
|
||||
|
||||
## Their offered tuning directions (not findings)
|
||||
|
||||
If the aim is Stress sometimes reaching 5: pressure scaling with **count** of unclaimed Problems; delay pressure until later rounds; lower DARVO arm threshold. Each is a rules choice for ground-game.
|
||||
|
||||
---
|
||||
|
||||
## Ground-game reading (for WP-0006-T04/T05)
|
||||
|
||||
1. **Do not promote H1 as packaged** into baseline — criterion 3 fails.
|
||||
2. **Direction is right:** Problem → Stress coupling addresses the real gap (inbound pressure only via ATTACK).
|
||||
3. **Shape/magnitude is wrong:** flat +1/end is a board-clear tax that competent players absorb with GROUND before DARVO.
|
||||
4. **Next experiment** should change *how* pressure is applied (scaled/delayed/targeted), not only add status stress — though Idea 2 last-place (semi) remains a candidate **H2** per design note §6.
|
||||
30
reports/260808-clay-borg-h2-measured.md
Normal file
30
reports/260808-clay-borg-h2-measured.md
Normal file
|
|
@ -0,0 +1,30 @@
|
|||
---
|
||||
id: GROUND-RPT-0005
|
||||
type: report
|
||||
title: "clay-borg: H2 measured — scoping works; bond SOLVE untested by bots"
|
||||
domain: consumer
|
||||
repo: ground-game
|
||||
status: informational
|
||||
owner: bernd
|
||||
topic_slug: whynot
|
||||
created: "2026-08-08"
|
||||
extends: GROUND-WP-0007
|
||||
---
|
||||
|
||||
# H2 measurement receipt (CB-EV-0032)
|
||||
|
||||
Upstream: clay-borg `evidence/CB-EV-0032-h2-measured.md`, hub 2026-08-08.
|
||||
|
||||
| # | criterion | verdict |
|
||||
|---|-----------|---------|
|
||||
| 1 | greedy SHARED 3–4p ≫ H1’s 0 | **met** — 120 / 175 vs 0 / 0 (~73% / 92% of baseline) |
|
||||
| 2 | stress variance up | **met** 2p/4p/6p; miss 3p 1.50 vs baseline 1.57 |
|
||||
| 3 | DARVO + still sometimes wins | **met** (reactive 13/13/57 at 3/4/6p) |
|
||||
| 4 | bond SOLVE elevated | **untested** by panel (bots ignore scope incentives); incidental claim rate lower on bond |
|
||||
| 5 | all-global control → H1 collapse | **met** — 3p/4p → 0 |
|
||||
|
||||
**ATTACK:** under H2 occasional attack is **affordable** (rank-75 still wins near greedy) but **not profitable** (never beats greedy; always-attack still 0). F17 still open as design judgement.
|
||||
|
||||
**Instrument:** `with_variant()` required for owner assignment; bare `state.variant =` left H2 inert — fixed on their side.
|
||||
|
||||
No felt-play. Bond joint-motivation needs a table or multi-agent-aware policy.
|
||||
116
reports/260809-clay-borg-attack-relief-correction.md
Normal file
116
reports/260809-clay-borg-attack-relief-correction.md
Normal file
|
|
@ -0,0 +1,116 @@
|
|||
---
|
||||
id: GROUND-RPT-0007
|
||||
type: report
|
||||
title: "clay-borg: correction — H1-B was never measured, and it does something"
|
||||
domain: consumer
|
||||
repo: ground-game
|
||||
status: informational
|
||||
owner: bernd
|
||||
topic_slug: whynot
|
||||
created: "2026-08-09"
|
||||
extends: GROUND-WP-0006
|
||||
---
|
||||
|
||||
# Correction to GROUND-RPT-0004 (H1 measured)
|
||||
|
||||
Upstream: clay-borg `specs/FindingRegister.md` F31,
|
||||
`games/ground/examples/module-panel.rs`.
|
||||
|
||||
**This is a correction we owe you, not a new result we chose to send.**
|
||||
|
||||
---
|
||||
|
||||
## 1. What we got wrong
|
||||
|
||||
We reported H1 as measured and rejected. **Half of it was never
|
||||
measured.**
|
||||
|
||||
H1 is two modules: flat problem→stress pressure (H1-A) and ATTACK
|
||||
self-soothe at Stress ≥ 4 (H1-B). Every game behind that verdict was
|
||||
played by bots that **never attacked, in any game, at any seat count**.
|
||||
H1-B acts only on an uncancelled ATTACK by a seat at the stress gate, so
|
||||
it had no opportunity in a single one of them.
|
||||
|
||||
**The rejection itself stands** — group success is 0 with or without
|
||||
H1-B, and that is H1-A's doing. But *"H1-B does nothing"* was never
|
||||
established, and is now known to be false.
|
||||
|
||||
The cause was on our side and is not subtle: our bot ranked GROUND at
|
||||
100 when the gate bites and ATTACK at 10, so at exactly the position
|
||||
where self-soothe would pay, GROUND won the comparison every time.
|
||||
|
||||
---
|
||||
|
||||
## 2. H1-B is unreachable in the printed game
|
||||
|
||||
We built a probe that attacks at the gate. **Selected on its own,
|
||||
`attack_relief.self_soothe_ge4` still never fires.**
|
||||
|
||||
Peak Stress in the baseline never exceeds its starting 2. The gate is at
|
||||
4. So no ATTACK can ever *be* gated, however eagerly a seat attacks —
|
||||
the baseline has no inbound Stress at all, which is the same observation
|
||||
as F17 seen from the other side.
|
||||
|
||||
**A module can be unreachable because of an aspect it does not name.**
|
||||
`attack_relief` needs a `problem_stress` module before it can act, and
|
||||
nothing in its own declaration says so. That may be worth a field in the
|
||||
catalog — a module's *preconditions*, distinct from its aspect — so that
|
||||
"unreachable in this configuration" is detectable rather than discovered.
|
||||
|
||||
---
|
||||
|
||||
## 3. Once reachable, it does something
|
||||
|
||||
Same probe throughout, varying **only** the module:
|
||||
|
||||
| seats | `problem_stress.scoped` | `+ attack_relief.self_soothe_ge4` | group wins | DARVO arms |
|
||||
|---|---|---|---|---|
|
||||
| 3p | 12 | **21** | +9 | 299 → 221 |
|
||||
| 4p | 19 | **28** | +9 | 294 → 213 |
|
||||
| 6p | 50 | **57** | +7 | 334 → 164 |
|
||||
|
||||
Under flat pressure the effect shows in DARVO alone: `h1` arms **200**
|
||||
against `problem_stress.flat_any_open`'s 300 / 400 / 600, with group wins
|
||||
0 in both.
|
||||
|
||||
**Read this as the conditional it is.** *Given seats that attack at the
|
||||
gate*, self-soothe raises group success by ~7–9 per 100 and cuts DARVO
|
||||
arming by a quarter to a half. The probe is deliberately poor play — it
|
||||
wins 12/100 where our ordinary bot wins 60/100 under the same module — so
|
||||
this is **not** a claim that attacking is good, and **not** evidence
|
||||
against F17.
|
||||
|
||||
**Sensitivity:** vary only the ATTACK rank. At 10 and at +55 the module
|
||||
never fires; at 110 it fires in every game. Nothing about the module
|
||||
changed — only whether any seat gave it a turn.
|
||||
|
||||
---
|
||||
|
||||
## 4. What we are not saying
|
||||
|
||||
- **Not** that H1 should be reconsidered. H1-A drives group success to 0
|
||||
at every seat band and self-soothe does not move that.
|
||||
- **Not** that ATTACK is good play.
|
||||
- **Not** that these numbers describe a table. Every one is bots, and the
|
||||
seats that produce them play badly on purpose.
|
||||
|
||||
**Ask:** is a *reachability* declaration worth adding to a module — the
|
||||
conditions under which it can act at all — so a consumer can report "not
|
||||
reachable in this configuration" instead of a number that looks like a
|
||||
null result?
|
||||
|
||||
---
|
||||
|
||||
## 5. Two earlier items still open on your side
|
||||
|
||||
Neither is new; both were in GROUND-RPT-0006 and are unanswered.
|
||||
|
||||
**(a) The competitive modes are weakly tested.** No bot of ours models a
|
||||
*rival* playing their objective. A mode whose point is that players
|
||||
compete, measured by seats that never anticipate an opponent, is a weak
|
||||
test of that mode — and we cannot fix it by measuring harder.
|
||||
|
||||
**(b) `problem_stress.scoped` cannot reach a decision in round one.**
|
||||
Only the Surface Problem is face up, so there is never a choice between
|
||||
two Problems for scope to decide. Any read of that module weighted toward
|
||||
early rounds is measuring a mechanism that has not started.
|
||||
Loading…
Add table
Add a link
Reference in a new issue