clay-borg/specs/FindingRegister.md
tegwick 4ec165c991
Some checks failed
ci / check (push) Failing after 4s
GROUND-RPT-0007: correct our H1 report — H1-B was never measured
Reported to ground-game as a CORRECTION to GROUND-RPT-0004, not a new
result. Half of H1 was never measured: every game behind that verdict
was played by bots that never attacked, and H1-B acts only on an
uncancelled ATTACK at the stress gate. The rejection stands on H1-A, but
'H1-B does nothing' was never established and is false.

Also carries the structural finding -- a module can be unreachable
because of an aspect it does not name -- with an ask: should a module
declare its reachability preconditions, so a consumer reports 'not
reachable in this configuration' rather than a number that reads as a
null result?

Repeats the two unanswered items from RPT-0006: the competitive modes
are weakly tested (no bot models a rival), and problem_stress.scoped
cannot reach a decision in round one.

Left uncommitted in their tree, as with every prior report.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 11:28:31 +02:00

336 lines
20 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# The finding register
Design findings about **GROUND**, with their reproductions. Governed by
[`GameDesign.md`](GameDesign.md) (admissibility, kinds, states, metrics)
and [ADR-0012](../decisions/ADR-0012-the-design-instrument.md). Reported
by `make design`.
**Split out of `GroundRules.md §Underdetermined` on 2026-08-05** when that
file crossed the ~400-line loadability limit. ADR-0012 D2 said *"no new
file"* and this is a new file — but D2's substance was **one register, not
a second mechanism competing with the first**, and that holds: this *is*
§Underdetermined's register, moved, still driving off the same
`provisional`/ruling machinery. D2 also named the awkwardness this
resolves — a finding about the engine sitting in a document about the
game.
The U-items themselves, with their defaults and rulings, stay in
[`GroundRules.md §Underdetermined`](GroundRules.md); this file tracks them
*as findings*.
**This section is the design-finding register** (ADR-0012 D2). It was the
register for dataset ambiguities already; CB-WP-0022 extended it to all
five kinds rather than building a second one beside it. Admissibility,
kinds, states and metrics: [`GameDesign.md`](GameDesign.md). Reported by
`make design`.
<!-- design-register:begin -->
| id | kind | state | reproduction | role | raised | owner |
|---|---|---|---|---|---|---|
| U1 | underdetermined | applied | — | default | 2026-07-31 | ground-game |
| U2 | underdetermined | applied | scenarios/ground/gr-d01-darvo-trigger.yaml | default | 2026-07-31 | ground-game |
| U3 | underdetermined | applied | — | default | 2026-07-31 | ground-game |
| U4 | underdetermined | applied | — | default | 2026-07-31 | ground-game |
| U5 | underdetermined | applied | — | default | 2026-07-31 | ground-game |
| U6 | underdetermined | applied | — | default | 2026-07-31 | ground-game |
| U7 | underdetermined | applied | — | default | 2026-07-31 | ground-game |
| U8 | underdetermined | applied | — | default | 2026-07-31 | ground-game |
| U9 | underdetermined | applied | — | default | 2026-07-31 | ground-game |
| U10 | underdetermined | applied | — | default | 2026-07-31 | ground-game |
| F11 | inert | applied | scenarios/ground/gr-p05-solve-legality.yaml | counterexample | 2026-08-02 | clay-borg |
| F12 | degenerate | note | — | — | 2026-08-01 | clay-borg |
| F13 | inconsistent | withdrawn | scenarios/ground/gr-e01-threshold-reachable-2p.yaml | counterexample | 2026-08-01 | clay-borg |
| F14 | unplayed | applied | games/ground/examples/attack-value.rs | counterexample | 2026-08-01 | clay-borg |
| F15 | underdetermined | note | — | — | 2026-08-05 | clay-borg |
| F16 | inconsistent | withdrawn | games/ground/examples/difficulty.rs | counterexample | 2026-08-05 | clay-borg |
| F17 | degenerate | raised | games/ground/examples/attack-value.rs | counterexample | 2026-08-06 | ground-game |
| F18 | inert | raised | `games/ground/src/edition.rs::card_text_tests::the_engine_reads_only_part_of_what_it_vendored` | counterexample | 2026-08-06 | clay-borg |
| F25 | inert | raised | `games/ground/src/edition.rs::card_text_tests::the_engine_agrees_with_the_editions_own_numbers` | default | 2026-08-08 | clay-borg |
| F24 | inert | raised | `games/ground/src/edition.rs::card_text_tests::the_hardcoded_deck_still_matches_the_edition` | default | 2026-08-07 | clay-borg |
| F19 | degenerate | applied | crates/cb-render-html/src/lib.rs::overhead_table | counterexample | 2026-08-06 | clay-borg |
| F20 | inert | applied | crates/cb-render-html/src/lib.rs::ending_page | counterexample | 2026-08-06 | clay-borg |
| F21 | degenerate | note | — | — | 2026-08-06 | clay-borg |
| F23 | inconsistent | applied | decisions/ADR-0017-chaos-window-2-verdict.md | counterexample | 2026-08-07 | clay-borg |
| F22 | underdetermined | withdrawn | games_ground::edition::supply_tests::play_never_exceeds_the_components_the_box_holds | counterexample | 2026-08-07 | clay-borg |
| F26 | inert | ruled | `crates/cb-render-html/src/lib.rs::a_scoped_table_says_what_a_scope_does` | counterexample | 2026-08-08 | ground-game |
| F27 | unplayed | reported | `games/ground/examples/scenario-panel.rs` | counterexample | 2026-08-08 | clay-borg |
| F29 | inconsistent | applied | `games_ground::edition::card_text_tests::which_scenarios_are_mechanically_distinct` | counterexample | 2026-08-08 | ground-game |
| F31 | unplayed | reported | `games/ground/examples/module-panel.rs` | counterexample | 2026-08-08 | clay-borg |
| F30 | degenerate | ruled | `games/ground/examples/scenario-panel.rs` | counterexample | 2026-08-08 | ground-game |
| F28 | underdetermined | applied | `games_ground::tests::the_mastery_rating_and_the_shared_score_count_different_things` | counterexample | 2026-08-08 | ground-game |
<!-- design-register:end -->
Prose for **closed** findings (withdrawn or applied) lives in
[`FindingRegister-closed.md`](FindingRegister-closed.md); the rows above
are the whole register and `make design` reads them from here.
- **F31 — `attack_relief.self_soothe_ge4` cannot fire, so it has never
been measured.** `module-panel` records **zero ATTACKs in every cell**,
at every seat band, under both a mode-blind and a module-aware bot. The
module only acts on an uncancelled ATTACK by a seat at Stress ≥ 4, so
it has had no opportunity in any game this repo has run — including
every game behind CB-EV-0030's H1 measurement, where it was half the
package.
The mechanism is legible: `GreedyPolicy` ranks `Ground` at 100 when the
stress gate bites and `Attack` at 10, so at exactly the position where
self-soothe would pay, GROUND wins the comparison.
`ModuleAwarePolicy` raises ATTACK by 55 at the gate and **that is still
not enough**, which is a fact about the ranking rather than about the
game.
**`unplayed`, emphatically not `inert`.** The kernel implements the
module correctly; nothing has ever given it an opportunity. Reporting
it as a null result is the error the panel now refuses to make: its
row shows real numbers with `← never fired: UNMEASURED` attached,
because the numbers belong to the *other* modules in the row.
**Sensitivity:** vary only the ATTACK rank. The falsifier is a policy
that attacks at the gate — if group success then moves under
`attack_relief.self_soothe_ge4` and not under the baseline, the module
is real and only our seats were hiding it. Until such a seat exists,
**no claim about this module is supported in either direction**, and
CB-EV-0030's H1 verdict rests on its flat-pressure half alone.
**FALSIFIED, 2026-08-08, and the answer has two parts.**
`GateAttackPolicy` — a probe that ranks ATTACK above GROUND at the
stress gate, delegating everything else — was built for exactly this.
**(a) The module is unreachable in the printed game, not merely
unchosen.** With `attack_relief.self_soothe_ge4` selected *alone* the
probe still never attacks and the row still reads `never fired`:
**peak Stress never exceeds 2**, so the gate never bites and no ATTACK
is ever gated. The baseline has no inbound Stress (F17's observation
from the other side), so this module needs a `problem_stress` module
before it can act at all. **A module can be unreachable because of an
aspect it does not name.**
**(b) Once reachable it is real.** Holding the probe fixed and varying
only the module — the only comparison that isolates it:
| seats | `scoped` | `scoped_plus_attack_soothe` | group wins | DARVO arms |
|---|---|---|---|---|
| 3p | 12 | **21** | +9 | 299 → 221 |
| 4p | 19 | **28** | +9 | 294 → 213 |
| 6p | 50 | **57** | +7 | 334 → 164 |
Under flat pressure the same shape appears in DARVO alone: `h1` against
`problem_stress.flat_any_open` arms 200 against 300 / 400 / 600 at
3/4/6p, with group wins 0 in both.
**Stated as the conditional it is:** *given seats that attack at the
gate*, self-soothe raises group success by ~79 per 100 and cuts DARVO
arming by a quarter to a half. The probe is deliberately poor play — it
wins 12/100 where greedy wins 60/100 under the same module — so this is
**not** a recommendation to attack, and the row is not evidence against
F17.
**What this changes about H1.** CB-EV-0030 rejected H1 with policies
that never attacked, so H1-B was inert for every game behind that
verdict. The rejection stands on flat pressure alone — group wins are 0
with or without self-soothe — but *"H1-B does nothing"* was never
established and is now known to be false.
**Sensitivity:** vary only the ATTACK rank. At `Attack = 10` (greedy)
and at `+55` (module-aware) the module never fires; at `110` it fires
in every game. Nothing about the module changed — only whether any seat
ever gave it a turn.
**Reported to ground-game 2026-08-09** as GROUND-RPT-0007, framed as a
correction to GROUND-RPT-0004 rather than a new result. Carries one
ask: whether a module should declare its **reachability preconditions**
— the conditions under which it can act at all — so a consumer can
report *"not reachable in this configuration"* instead of a number that
reads as a null result.
- **F30 — SCN_04 is materially harder at 2 players.** 52/100 group
success against 67 and 73, with **every other parameter held by the
edition itself**: same deal shape, same 6 available points, same
threshold 5, same starting Stress. The lone difference is that SCN_04
is the only 2p deal needing **two of one suit** (Repair, Clarify,
Repair).
**Causation is not established** and the finding does not claim it. The
falsifier is a deck with a doubled suit that does not lose group
success.
**Sensitivity:** vary only the seat band and the gap closes — at 4p
SCN_04 is 97 against 94/99, so it is *not* the hard board there. The
effect is specific to the 2p deal, which is the band where the deck has
fewest cards to route around a suit it cannot match.
**RULED 2026-08-08:** acceptable variety for now; SCN_04 stands as a
harder board and is monitored. If table feel proves too swingy the
first lever is priority-2's suit (dropping the double Repair), **not**
the global thresholds — which also confirms our reading that the suit
multiset is the operative difference, without our having measured it.
- **F27 — the two competitive modes are scoring lenses over cooperative
play.** `scenario-panel` finds group success *exactly* equal across
SHARED GROUND, COMMON PROBLEM and BONDED COALITIONS in all 36 cells.
Correct arithmetic: `GreedyPolicy` maximises the group outcome and
never reads `state.mode`, so the same games are played and only the
winner set is carved differently. **Whether a mode changes how GROUND
is played is therefore untested**, and cannot be tested by any panel
built from the current policies — it needs one that plays for personal
score. `unplayed` rather than `inert`: the modes score correctly, they
have simply never faced a seat that wanted to win alone.
**Re-measured 2026-08-08 with `ObjectivePolicy`** (CB-WP-0049 T03),
which plays its seat's own objective. The finding **splits in two**:
- **Group success does not move.** 34 of 36 cells identical; SCN_03 at
4p goes 99→100 in all three modes, which is the SOLVE-by-value
refinement and not a mode effect. **Whether the table survives is the
same game in all three modes.**
- **Who wins does move.** Under BONDED COALITIONS at 4p, winning seats
per game go 2.04 → 2.98 (SCN_01/02), 2.12 → 3.29 (SCN_03), 2.05 →
3.01 (SCN_04): a seat whose score is its coalition's sum bonds more,
and the coalitions get about half again as large. COMMON PROBLEM
moves at 6p (1.10 → 1.17).
So the modes **do** reach decisions — the original claim that they are
"scoring lenses over cooperative play" was too strong, and is withdrawn
in favour of the sharper one: **GROUND's scoring modes change who wins,
not whether the group succeeds.**
**Sensitivity:** vary only the seat band and the effect appears and
disappears — 2p shows no divergence in any mode, 4p shows the largest,
6p shows none under BONDED COALITIONS. A candidate explanation is that
two relation slots per seat cap network growth, so at 6p the incentive
exists and cannot be acted on; **that is untested** and is the next
thing to vary.
**Still open**, because the part that motivated it is untested: no
policy models a *rival* playing their objective, so a competitive mode
in which nobody anticipates an opponent remains a weak test of that
mode (CB-WP-0049 "Not done here").
**The seat-band explanation is answered, 2026-08-08.** ground-game:
the two-relation-slot limit is an **intentional cap** on complexity and
play time; a third slot stays an extension rather than core r0. So the
6p flatness is designed, not a defect — and our sensitivity run is
welcome but would be measuring an extension, not the shipped game.
**The aspect half is closed, 2026-08-08** (CB-WP-0049 T04).
`ModuleAwarePolicy` reads the resolved `Rules`, so "a bot that does not
attend to a mechanism cannot test claims about it" no longer holds for
modules. **The mode half stands**: no policy models a *rival* playing
their objective, so a competitive mode in which nobody anticipates an
opponent remains weakly tested. That is the whole of what keeps this
open.
A second limit surfaced while closing the first: the scope term is
**inert at round-one positions**, because one Problem is face up and a
ranking preference needs two candidates. Any measurement of
`problem_stress.scoped` weighted toward early rounds is measuring a
mechanism that has not started.
- **F26 — a package that adds a FILE is invisible, where a package that
adds a column is not.** `h2-scoped-problem-stress` ships
`Rules_Text.csv` — twenty-two passages of player-facing rules, including
the one passage that says what a `stress_scope` does. No reader in
clay-borg mentioned the file, so `edition-check` could not call it stale
(a file nothing reads is not out of date, it is unseen) and no gate
could call it unread. The rule the table's whole layout encodes sat in
the repo, reachable by no player, for the entire measured life of H2.
**Inert rather than underdetermined**: nothing is ambiguous about the
rule and the engine implements it correctly — the text simply could not
fire. Third instance of F18's shape and the first where the unit is a
file, which is why it is filed separately instead of folded in.
Raised against `ground-game` because the actionable half is theirs: a
package's manifest should name its own files, so that a consumer reading
none of them is a detectable state. clay-borg now reads the one passage
it needed (CB-WP-0046); the other twenty-one remain unread and that is
the standing evidence.
**RULED 2026-08-08:** accepted as a process ask. Catalog and modules
will declare `consumed_files` / overlays so a consumer reading none of
a package's files is a detectable state. Tracked on their side under
GROUND-WP-0008.
- **F17 — ATTACK cannot affect whether the table succeeds.** Reported
from play as *"there is no incentive to play attacks as long as I have
positive cards"*, and **the artifact now exists**:
`games/ground/examples/attack-value.rs`, 200 games per cell, one number
varied — ATTACK's rank in an otherwise identical policy.
| ATTACK ranked | wins (2/3/4/5/6p) | attacks | DARVO armed |
|---|---|---:|---:|
| 10 (below all) | 132 / 165 / 190 / 200 / 200 | 0 | 0 |
| 75 (above SUPPORT) | **132 / 165 / 190 / 200 / 200** | 315923 | 13218 |
| 95 (above SOLVE) | **0 / 0 / 0 / 0 / 0** | 14005170 | 4001000 |
**The middle row is the finding: identical win counts at every seat
count**, while attacking hundreds of times and arming DARVO. Attacking
is not punished — it is **inert with respect to the goal**. Group success
is a function of SOLVE alone, and ATTACK only costs anything when it
ranks above SOLVE and displaces it.
**So the maintainer was right and the reason is sharper than the
phrasing.** There is no incentive because there is no *path*: ATTACK's
effects (Stress, Rivalry, DARVO) feed nothing that decides
`group_success`.
**Asked in all three modes 2026-08-07, and ATTACK earns its place in
none of them:**
| mode | never attack | attack sometimes |
|---|---|---|
| SHARED GROUND | 132/165/190/200 | **identical** — free but pointless |
| COMMON PROBLEM | 59/52/48/44 | 59/52/48/**34** — a cost at six seats |
| BONDED COALITIONS | 131/134/132/116 | **59/52/48/34** — roughly halved |
**The coalitions row has a mechanism the data confirms.** GR-A07 flips a
Bond to a Rivalry on Attack, and GR-E04 scores Bond *networks* — so
attacking destroys the thing that scores. And the attacking numbers in
E04 are **identical** to E03's, which is exactly what that predicts:
break every Bond and each seat becomes a coalition of one, so GR-E04
degenerates into GR-E03.
**Not a claim that the game is broken.** DARVO is the pattern the game
is *about* not falling into; a self-destructive ATTACK may be the
design. The question for `ground-game` is whether the namesake mechanic
being unreachable in competent co-op play is intended.
- **F18 — the engine imports one of nineteen edition files, so the cards
cannot say what they do.** Reported as *"I don't understand the GROUND
card"* — which is not a design gap: `Actions.csv` carries that card's
tagline (*"Regulate. Restore the frame. Decide."*) and full rules text,
and clay-borg never imported it. Every other card is the same: a player
sees `Clarify` where the card reads *"Ask What Happened — Invite a
concrete account before judging."* **`inert`**: the data exists and
cannot fire, because nothing reads it. **Ours, and CB-WP-0028 fixes it.**
- **F25 — the engine hardcodes numbers the edition states.**
`GroundState::threshold` is a `match` returning 5/7/9 while
`Scenarios.csv` carries `threshold_2_players`, `threshold_3_4_players`
and `threshold_5_6_players`; setup writes `stress: 2` against *"All
players start at Stress 2."*; the engine plays five rounds against a
printed 15 track. **All three agree, on all four scenarios**, and that
is the finding rather than the reassurance: these are the most contested
numbers in the project — the whole 4/6/9 versus 5/7/9 episode turned on
them — and the engine has been right by maintenance coincidence rather
than by reading the file that owns them. **`inert`**, role `default`:
green because they agree, red the moment either side moves. **Ours.**
Same shape as F24.
- **F24 — the draw pile is a Rust literal.** `solution_deck()` builds six
of each suit from an array and never opens `Solutions.csv`, whose `suit`
and `quantity` columns say the same thing. **They agree today** — 24
rows, six per suit — so nothing is wrong now, and that is the point:
the engine is right by maintenance coincidence rather than by reading.
**`inert`**: the data exists and cannot fire. Role `default`, not
`counterexample` — the reproduction is green *because* the two agree,
which is the state GameDesign §1.3 says to expect from a documented
provisional choice, and it turns red the moment either side moves.
**Ours.** CB-WP-0037 T02 deletes the literal.
- **F15 — the rules define one game, not a series.** `OutcomeView` gives
`personal` (per seat), `group_success` (per table) and `winners`. Summing
the first and counting the third answer different questions, and GROUND
says nothing about how several games combine. CB-WP-0024 T04 shows
**both, labelled**, rather than picking one and letting it become the
score by default. `note`: no artifact demonstrates that this *harms*
play — a unit test showing the two tallies point at different seats
demonstrates only that they can differ, which is arithmetic, not a design
defect. Under GameDesign §3.1 it may not be reported until one exists.