Reports out of the workplan number space

The workplan list read as though GROUND-WP-0002, 0003 and 0005 each
existed twice. They did not: three inbound clay-borg finding reports were
filed in workplans/ borrowing the number of the workplan they bear on, and
each was registered in State Hub as a workplan.

A report is correspondence, not a unit of work. The three move to
reports/YYMMDD-<slug>.md with their own id space (GROUND-RPT-0001..0003),
keeping type: report, status: informational and the extends: link to the
workplan they relate to. Their state_hub_workstream_id bindings are
dropped and the hub registrations archived, so the hub's workplan list
matches the files.

Five workplans, one file per number.

AGENTS.md records the convention, since the reports arrive regularly and
nothing in the repo said where they go.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
tegwick 2026-08-07 22:18:18 +02:00
parent 93ca4e807e
commit 92f52e777d
4 changed files with 15 additions and 6 deletions

View file

@ -1,115 +0,0 @@
---
id: GROUND-WP-0002-REPORT-260805
type: report
title: "clay-borg finding report, 2026-08-05 — one withdrawal, and ten rulings we had not collected"
domain: consumer
repo: ground-game
status: informational
owner: bernd
topic_slug: whynot
created: "2026-08-05"
extends: GROUND-WP-0002
state_hub_workstream_id: "09350395-f628-4d8e-8c36-3cea330a8b6f"
---
# clay-borg finding report — 2026-08-05
Generated from clay-borg's finding register
(`specs/GroundRules.md`, `make design`) under
[CB-WP-0022](../../clay-borg/workplans/CB-WP-0022-the-design-instrument.md) T06.
**This is a file, not an inbox message, and that is the point.** The
message of 2026-07-30 sat unread for four days. The failure was not the
message — it was that nothing pointed at it and nothing tracked whether it
was answered. This extends GROUND-WP-0002 rather than duplicating it.
**Nothing here asks for a ruling.** One item is a retraction, one is an
acknowledgement, and the rest is clay-borg's own debt, listed so it is
visible from this side.
---
## 1. Withdrawn: GR-E01 vs GR-S01 — *"4/6/9 against 5/7/9"*
**clay-borg retracts this finding.**
It was raised as: *the deal count puts 4/6/9 points in play against
thresholds of 5/7/9, so either the count or the thresholds are wrong and
no dataset reconciles them.* It was queued to be sent to you again as
recently as this pass.
**It is wrong.** With the authoritative `Problems.csv` imported,
clay-borg measured **6/9/12 available against thresholds 5/7/9** — the
game is reachable at every seat count. The scenario has been renamed
`gr-e01-threshold-unreachable-2p``gr-e01-threshold-reachable-2p`.
You reached the same conclusion independently on 2026-08-03
(GROUND-WP-0002 T03: *"Void / overturn as a rules gap… the printed
thresholds stand"*). **This confirms your ruling from our side and closes
the item.**
Reproduction: `scenarios/ground/gr-e01-threshold-reachable-2p.yaml`,
commit `2da19a4`.
**Why you are being told about a retraction at all.** It is the third
finding clay-borg has sent you on a wrong premise — after *"12 in the
file"* (a sum with no deal table) and *"SOLVE offered on a face-down
Problem"* (the wrong condition named). A claim retracted silently is how
the first two survived. Our new rule is that withdrawals travel the same
path as findings.
## 2. Acknowledged: all ten U-items, ruled 2026-08-03
**We had not collected your answers.** GROUND-WP-0002 T05 confirmed all
ten U-items on 2026-08-03. Two days later, clay-borg's own survey still
reported *"0 of 10 ruled"*, and the register built this session initially
recorded them as unanswered.
**Every one was confirmed as the default we already simulate**, so no
kernel behaviour changes. What we owe is bookkeeping: the scenarios still
carry `provisional: true` for choices you have settled, and until those
flags are lifted our tooling will keep reporting the items as open.
Tracked on our side; nothing needed from you.
Same for the five provisional scenarios confirmed in GROUND-WP-0002 T03.
## 3. Our debt, listed so you can see it
`make design`, 2026-08-05:
| | |
|---|---|
| findings | 12 (+2 notes) |
| with a resolving reproduction | 3/12 |
| open, lacking a reproduction | 9 — U1, U3U10 |
| reproductions green while open | 0 |
**Only U2 names its U-item in a scenario.** Five provisional scenarios
exist and four probably encode U-item defaults, but the mapping is not
written down, so it is not checkable. That is clay-borg's to fix.
## 4. Two open items that are not findings yet
Both are **notes** under our new rule — real, but without an artifact, so
they may not be raised as findings until one exists.
- **GR-A13 "wasted SOLVE"** on an already-claimed Problem: a scenario had
to pick a default and did. No artifact isolates the degenerate line.
- **GR-E03 / GR-E04 never played to the end.** Your GROUND-WP-0003 is the
playtest that would close this. When it runs, **the recording is the
artifact** — clay-borg's `cb-play --record` writes a finished game as a
replayable scenario, so a playtest produces the reproduction for
anything found on the way, at the cost of one flag.
## 5. What changed on our side, if you care about the mechanism
A finding is now inadmissible unless its reproduction **exists**, **has
the shape you ruled** (GROUND-WP-0004 T02 — a row-level deal table, never
a sum), **and can fail**. The third clause is new and is why §1 above is a
retraction: GR-E01's reproduction went *green* when the edition landed and
nothing treated a passing artifact as a signal, so a dead finding stayed
queued for four days.
Your half of that ruling — that rulings depending on arithmetic should
cite the reproduction by id or path — is what makes this report checkable
rather than assertive.

View file

@ -1,110 +0,0 @@
---
id: GROUND-WP-0003-REPORT-260807
type: report
title: "clay-borg: GR-E03 and GR-E04 played to the end, and what ATTACK is worth in each"
domain: consumer
repo: ground-game
status: informational
owner: bernd
topic_slug: whynot
created: "2026-08-07"
extends: GROUND-WP-0003
state_hub_workstream_id: "fb552dd0-a746-4339-9b6f-49597796ea7d"
---
# GR-E03 and GR-E04 have been played to the end
GROUND-WP-0002 T04 recorded that both scoring modes were implemented and
neither had ever been played through. That is now done, and it produced an
answer to a second question we raised on 2026-08-06.
Reproduction: `clay-borg$ cargo run --release -p games-ground --example attack-value`
(800 games per mode). Under
[CB-WP-0029](../../clay-borg/workplans/CB-WP-0029-the-tokens-on-the-table.md).
---
## 0. Why they were unplayed, which was our fault
`cb-play` built **every** game as SHARED GROUND and passed an empty setup
patch. The mode was settable in scenario files and **not from the driver**,
so two of your three shipped modes were unreachable from the only way
anyone actually plays. Nothing about the modes was wrong; we simply had no
way to select them.
Fixed. Reporting it because *"we never played them"* and *"we could not
play them"* are different statements, and the second is the true one.
## 1. All three modes work, and they disagree
Same game, same 37 commands, seed 1, four players:
| mode | winners |
|---|---|
| SHARED GROUND (GR-E02) | P1, P2, P3, P4 — mastery 4 |
| COMMON PROBLEM (GR-E03) | **P3** — the top personal scorer |
| BONDED COALITIONS (GR-E04) | **P1, P2** — the best Bond network, 4 > 3 > 2 |
The modes produce three different answers from identical play, which is
what they should do.
## 2. What ATTACK is worth — asked in all three
On 2026-08-06 we raised, from play, *"there is no incentive to play
attacks as long as I have positive cards."* We could not send it then: it
had no artifact, and our own bot ranks ATTACK below everything, so its
zero attacks measured **our heuristic**, not your game.
The artifact varies **exactly one number** — ATTACK's rank in an otherwise
identical policy — and asks in every mode. 200 games per cell.
*"Won"* is group success in co-op, and *"is seat 0 among the winners"* in
the other two, since ATTACK is an individual's choice.
| mode | never attack | attack when convenient | attack always |
|---|---|---|---|
| SHARED GROUND | 132 / 165 / 190 / 200 | **identical** | 0 / 0 / 0 / 0 |
| COMMON PROBLEM | 59 / 52 / 48 / 44 | 59 / 52 / 48 / **34** | 0 / 0 / 0 / 0 |
| BONDED COALITIONS | 131 / 134 / 132 / 116 | **59 / 52 / 48 / 34** | 0 / 0 / 0 / 0 |
*(2, 3, 4 and 6 seats)*
**ATTACK earns its place in none of them.**
- **Co-op:** attacking is *free* — identical win counts — and also
pointless. Group success depends on SOLVE alone; ATTACK's effects
(Stress, Rivalry, DARVO) feed nothing that decides it.
- **Semi-co-op:** a modest cost at six seats.
- **Coalitions:** roughly **halves** your chance of being among the
winners.
## 3. The coalitions result has a mechanism, and the data confirms it
GR-A07 flips a Bond to a Rivalry on Attack. GR-E04 scores Bond
**networks**. So attacking destroys the thing that scores.
The confirmation was not designed and fell out: **the attacking numbers in
GR-E04 are identical to GR-E03's** — 59 / 52 / 48 / 34 in both. That is
exactly what the mechanism predicts. Break every Bond and each seat
becomes a coalition of one, so **GR-E04 degenerates into GR-E03.**
## 4. The consequence we think is worth your attention
Because nobody who plays well attacks, **Stress never rises, and DARVO
never arms.** In 500 competent games it fired zero times. The full
sequence — DENY, ATTACK, REVERSE, the Focus/Blame flip, the Protection
interaction — is reachable only by playing badly.
**We are not saying this is a defect.** DARVO is the pattern the game is
*about* not falling into, and a self-destructive ATTACK may be exactly the
design — the mechanic as a cautionary structure rather than a strategy.
**The question is narrower than that, and it is yours:** is it intended
that the namesake mechanic is unreachable in competent play, in **all
three** modes? If yes, nothing here needs changing and we will record it.
If no, the lever is that ATTACK has no path to any scoring outcome.
## 5. Standing offer
We can now run any mode at any seat count over any seed range, and vary a
single policy parameter to isolate a mechanic's value. If there is a
question of the form *"does X ever pay?"*, it is cheap for us to answer.

View file

@ -1,132 +0,0 @@
---
id: GROUND-WP-0005-REPORT-260805
type: report
title: "clay-borg difficulty measurement, 2026-08-05 — and a retraction of what we nearly sent instead"
domain: consumer
repo: ground-game
status: informational
owner: bernd
topic_slug: whynot
created: "2026-08-05"
extends: GROUND-WP-0005
state_hub_workstream_id: "61cabf31-88a8-4ca6-9383-55b8fccddecf"
---
# clay-borg difficulty report — 2026-08-05
GROUND-WP-0005 (*difficulty tiers and the Pressure dial*) has both tasks
in `wait`, blocked on a measured baseline. This is that baseline, with its
limits stated.
Generated by `make difficulty`
(`games/ground/examples/difficulty.rs`), governed by
`clay-borg/specs/RetrospectiveAnalysis.md` §4, under
[CB-WP-0025](../../clay-borg/workplans/CB-WP-0025-could-we-have-won.md) T06.
---
## 0. First, what we nearly sent you and did not
**Earlier today this report was going to say: "a greedy bot wins 200 of
200 games at five and six seats — the game is too easy there."**
It would have invited you to move thresholds. **It was wrong**, and an
adversarial review caught it before it left our repo.
A `FirstLegal` policy — take the first legal command, no heuristic
whatsoever — scores **0%** at five and six seats on the identical deals
where `GreedyPolicy` scores **100%**. At two seats it *beats* greedy
(76.7% vs 60.0%).
**Two unsophisticated agents span the entire range.** A single policy's
win rate is therefore a statement about the policy, not about GROUND, and
we have written that into our spec as a prohibition rather than a
preference.
This is the fifth claim we would have sent you on a wrong premise. It is
the second stopped before sending.
## 1. The measurement
`make difficulty`, 60 seeds per cell:
| seats | winnable | greedy | random | first-legal | spread | undecided |
|---|---:|---:|---:|---:|---:|---:|
| 2 | **60%** | 60.0% | 5.0% | 76.7% | 71.7 | 0 |
| 3 | **93%** | 88.3% | 6.7% | 25.0% | 81.7 | 2 |
| 4 | **100%** | 93.3% | 6.7% | 30.0% | 86.7 | 4 |
| 5 | **100%** | 100.0% | 3.3% | 0.0% | 100.0 | 0 |
| 6 | **100%** | 100.0% | 3.3% | 0.0% | 100.0 | 0 |
**`winnable`** is the column that is not about a bot. It is the fraction
of deals in which an **exhaustive search finds a winning line in the final
round**. The other three columns are what named policies actually achieve,
and **`spread`** is the range between them.
## 2. How to read it, including what it is not
**Read `spread` first.** It ranges from 71.7 to 100.0 percentage points.
Wherever it is large — everywhere — **no single policy's rate tells you
anything about the game.** That is the finding behind §0.
**`winnable` is conditioned on greedy's play up to the final round.** It
is *"winnable from where greedy got to"*, not a property of the deal
alone. A genuinely policy-free figure would search from round 1, which we
measured as unaffordable: our search exceeded 2×10⁶ nodes at **two** seats
over two rounds. We are telling you this rather than presenting the number
as cleaner than it is.
**It is a lower bound.** A one-round search cannot see a line that needed
an earlier round, so the true winnable fraction is **at least** these
figures.
**`undecided` deals are excluded, not counted as losses.** Those are deals
whose search hit the node budget. Folding *"we stopped looking"* into
*"not winnable"* is the collapse our spec forbids.
## 3. What we think this supports, offered as reading not as ruling
**The rules are yours. This is what the numbers appear to say.**
- **Two seats is the tight configuration.** 60% winnable in the final
round, and every policy struggles. If any seat count is under-tuned for
*difficulty*, the evidence does not point here.
- **Four seats and up: every deal we sampled was still winnable in the
final round.** 100% at 4, 5 and 6 seats. That is compatible with "too
generous at high seat counts" — but it is **not** the claim we withdrew
in §0, because it is a statement about deal winnability rather than
about a bot's success.
- **The 56 seat rows deserve the most suspicion.** `first-legal` scores
0% there while `greedy` scores 100%, which is the widest spread in the
table. Something about those configurations makes play matter *more*,
not less — the opposite of "too easy" — and we do not yet understand it.
**We are not proposing threshold changes.** The measurement's confound
(§2) is large enough that we would rather you saw it than acted on it.
## 4. What would make this a better instrument
- **Search from round 1** rather than the last round, which removes the
greedy confound entirely. Needs transposition or move-ordering; we have
neither.
- **A resolution figure** — the smallest threshold change the measurement
can distinguish, with its N. Our spec requires it and we cannot yet
supply it.
- **Your Pressure dial**, if it lands, gives a second axis to measure
against and would make the table two-dimensional.
## 5. Reproduction
Per GROUND-WP-0004 T02, arithmetic findings cite their reproduction by
path:
```
clay-borg$ make difficulty # the table above
clay-borg$ make self-tests # includes the difficulty controls
```
Controls, so the figure can fail rather than merely print: a witness must
replay to a win; an unwinnable position must be reported as **searched
out** rather than as a budget cut; a budget of one node must not claim
exhaustion; and the policy panel must actually disagree, or reporting
three policies would be ceremony.