Compare commits

..

No commits in common. "dc440e82c680db779828e6baccd295952843dd19" and "25531e9d0187d43fea0e4209020cbe36868db02d" have entirely different histories.

10 changed files with 36 additions and 509 deletions

View file

@ -1,199 +0,0 @@
# CB-EV-0007 — INTENT stage 0, and whether the loop earns its keep
CB-WP-0008 T04. Measured 2026-08-01, at `HEAD` after T03.
---
## 1. Is stage 0 shipped?
> **INTENT 0. Headless GROUND** — *full authoritative state, 2–6 players,
> commit/reveal, relationships, DARVO, GROUND practice, CLI player, replay
> and scenario tests, simple bots. No rendering, no physics.*
| item | state | where |
|---|---|---|
| full authoritative state | **met** | event-sourced aggregate, K1–K3; `state_hash_hex` |
| 2–6 players | **met, with a caveat below** | CB-WP-0008 T03: all five counts play out, both policies, reproducible |
| commit/reveal | **met** | GR-R02/R04 face-down selections; hidden in the projection |
| relationships | **met** | Bond/Rivalry, GR-L01..L05, slots enforced |
| DARVO | **met** | GR-D02..D07 resolve; `gr-d01`…`gr-d06` |
| GROUND practice | **met** | GR/OU/ND with sub-choices, GR-A10..A12 |
| **CLI player** | **met (new)** | `tools/cb-play`, `make play` |
| replay and scenario tests | **met** | 24 scenarios, K10 `.cbreplay`, `make replay-test` |
| **simple bots** | **met (new)** | `games/ground/src/bot.rs`, random + greedy |
| no rendering, no physics | **met** | nothing in the tree draws |
**Stage 0 is shipped.** The two items that were missing at the start of
this pass — the CLI player and the bots — are the two the workplan was
opened for.
Three qualifications, none of which I will hide behind a checkmark:
1. **GR-E01's threshold is unreachable below five seats** with the
standard preset's placeholder Problem values. Best possible total is
3 / 6 / 6 / 10 / 10 against thresholds 5 / 7 / 7 / 9 / 9. The rules
work at every count; the *fixture* cannot be won at 2–4 players. Pinned
by `gr-e01-threshold-unreachable-2p`, an arithmetic test, and a
provisional marker owned by `ground-game`.
2. **GR-A13 admits SOLVE against an already-claimed Problem**, which then
resolves to nothing — a silently wasted action. Found by the greedy
bot in its first game.
3. **The scoring modes are not equally exercised.** Every full game ever
played — scenario, bot or CLI — scored under **SHARED GROUND**.
GR-E04 is reached only from a patched round-5 state, and **GR-E03
(COMMON PROBLEM) has no scenario at all**: it is implemented,
unreferenced by any test, and would not fail if deleted.
## 2. What did the second consumer reveal?
INTENT: *"No concept becomes canonical merely because it looks general. It
becomes canonical after surviving a second concrete use."* Before this
pass every abstraction here had survived exactly one use — the scenario
runner. The bots and the CLI are the second.
| abstraction | verdict from the second use |
|---|---|
| `Aggregate` (validate → fold) | **fits.** Bots drive it unchanged; nothing needed widening. |
| `Project` / K13 | **fits, and was untested.** Six passes with zero implementors; `GroundView` is the first. Writing it is what exposed that `seed` must not project — it is not secret content, but a seat holding it can compute the deck. |
| `ScenarioGame::parse_command` | **fits, and gained an inverse.** `record::to_step` round-trips every command shape, so a played session becomes a scenario. |
| K10 replay bundles | **fits.** `cb-play --replay` writes one with no change to the writer. |
| `Policy` (new) | one use each by bots and by the CLI's human seat — a human and a bot are the same interface, which is the strongest evidence in this table that the seam is in the right place. |
| **`CommitWindow`** | **still zero non-test users.** The bot pass was its best chance to acquire a second one and did not: GROUND implements the same contract inline. Delete-by 2026-12-31 stands. |
| `LogStore` | unchanged by this pass; still one real consumer. |
The load-bearing finding: **the abstractions that had a consumer fit; the
one that had none is still unused.** `CommitWindow` was extracted because
it looked general. It has now survived two passes that could have used it
and did not.
## 3. Meta share under InnerLoop v1.5
This pass is `kind: product` and is the first run under the soft 25% cap.
```
meta budget [OVER] 61% of $92.32 attributed (soft 25%, InnerLoop v1.5)
product $12.02 meta $32.85 mixed $47.46 (split 50/50)
```
**This pass alone is 100% product** — every dollar in it closed a stage-0
item. The 61% is cumulative and falling (68% → 61%).
> **And that is a defect in the instrument I built two tasks ago.** The
> meta budget aggregates over *every attributed task ever closed*, which
> is exactly the disease CB-RES-0005 §1 named for SH-1/SH-2: **a
> cumulative mean cannot detect the present, because the history outvotes
> it.** A pass that is 100% product reads `OVER`. A pass that is 100% meta
> would also read `OVER` and look no different.
The fix is the same one T01 of CB-WP-0007 applied to session shape: window
it. **It is not built here**, because meta is at 61% and v1.5 forbids
opening new meta work above the line — the budget's first real act was to
stop me from improving the budget. Recorded as the top candidate for the
next meta pass.
## 4. Cost per unit of product
Per-window, measured with `cb-cost --since <pass boundary>`:
| pass | kind | responses | cost | $/response |
|---|---|---|---|---|
| CB-WP-0004 | meta | 165 | $20.28 | **0.123** |
| CB-WP-0005 | meta | 98 | $22.32 | 0.228 |
| CB-WP-0006 | meta | 158 | $57.22 | **0.362** |
| CB-WP-0007 | meta | 31 | $9.24 | 0.298 |
| **CB-WP-0008** | **product** | **103** | **$12.66** | **0.123** |
Per task, as `task-done` measured them: T01 **$4.29**, T02 **$5.60**;
T03 and this file make up the balance. **Three stage-0 deliverables and
22 new tests for $12.66** — less than any single task in CB-WP-0006,
whose T04 alone cost $10.06. The window is measured at `HEAD` after T03,
so T04's own cost is not in it.
Session shape, same window, the first pass to satisfy both context targets
since they were written:
| metric | this pass | target | previous pass |
|---|---|---|---|
| SH-1 mean context | **161,994** | ≤ 200,000 | 503,464 |
| SH-2 p90 context | **215,664** | ≤ 300,000 | 607,105 |
| SH-3 batching | 6.9% | ≥ 20% | 0.0% |
## 5. The retrospective: six passes in, is the loop earning its keep?
**Answer with the spend curve.** It is a V, and both arms are informative.
```
$/response 0.123 ── 0.228 ── 0.362 ── 0.298 ── 0.123
pass 0004 0005 0006 0007 0008
kind meta meta meta meta PRODUCT
```
Three readings, and only the third survives:
1. *"Meta work is expensive, product work is cheap."* Tempting and wrong:
CB-WP-0004 was meta and matched this pass exactly.
2. *"Compaction did it."* Partly true and not sufficient. Context fell
3.1× (503,464 → 161,994) and cost fell 2.9× (0.362 → 0.123), so the
two move together here. But CB-WP-0004 reached the same $0.123 in a
stretch running at **≈209,000** mean context — derived from
CB-RES-0005 §1's cumulative rows, not measured directly — which is
above this pass and well above the target. Cost and context are
correlated, not locked.
3. **"Cost per response tracks how far the work is from a runnable
check."** CB-WP-0004 (tooling with self-tests) and CB-WP-0008 (code
with tests) both had a green/red answer within seconds of each edit.
CB-WP-0005 through CB-WP-0007 were arguing about what a number *means*
— whether a mutation is strong, whether a window is the right window,
whether a target should move — and every one of those questions is
settled by prose, in context, at the top of the price curve.
That is the honest general finding, and it is not "stop doing meta work".
The meta passes produced the mutation harness that caught six real
defects, the windowed metric, and the budget that stopped me in §3. It is:
> **Loop work should be cashed out into something runnable in the same
> pass that invents it, or it will be paid for in context.**
> `mutation-check` did this (a command, five controls, verdicts) and cost
> $0.228/response. The taxonomy of instrument failures around it did not,
> and cost $0.362.
**Is the loop earning its keep?** On this evidence, yes — but its return
is *lumpy and lagging*. CB-WP-0004's tooling is used in every pass since.
CB-WP-0005's mutation harness caught six defects in CB-WP-0006. CB-WP-0006
corrected five published verdicts. CB-WP-0007's window and budget both
fired within one pass. None of that appeared as a cost saving in the pass
that produced it, and the spend review at the start of this pass — which
scored CB-WP-0005 and CB-WP-0006 at "$74 for zero measured efficiency
gain" — was measuring too early with too narrow a definition of gain.
The correction to that review is this pass itself: the same loop, pointed
at product, ran at its best recorded cost per response and its lowest
recorded context, and shipped a stage.
## 6. Prediction discipline, fourth outing
| pass | prediction | measured | error |
|---|---|---|---|
| CB-WP-0004 | 25–30 points recovered | 6 | 4–5× |
| CB-WP-0005 | ≥10 of 14 rows | 4 | 2.5× |
| CB-WP-0006 | 9 per-task outcomes | 7 met, 2 restated | small |
| **CB-WP-0008** | *stage 0 ships; meta reads under 25%* | **stage 0 shipped; meta reads 61%** | **half right, and the miss is the finding in §3** |
The workplan said: *"the budget should read comfortably under the line; if
it does not, the evidence file says so and names what was displaced."*
Nothing was displaced. The budget is measuring the wrong window.
## 7. Open, carried forward
- **`CommitWindow`** — still zero non-test users; delete-by 2026-12-31.
- **The meta budget is cumulative** and cannot see the present (§3).
- **GR-E01 unreachable below 5 seats**; **GR-A13 wasted SOLVE** — both
need a ground-game ruling, not a kernel change.
- **GR-E03 / GR-E04 have never been played to the end.**
- **AM-3 blocked**; **AM-7/AM-8 PARTIAL**; kernel gate binds 2026-08-31 at
18/18 green.
- **SH-3 at 6.9%** against a 20% floor, unfalsified — nobody has yet tried
to hit it (CB-WP-0007 T04, cancelled).
- **A stale doc comment** claimed DARVO was unimplemented for weeks after
it shipped. `make facts-check` gates duplicated *numbers*; nothing gates
a prose claim about code. Fixed in place; the class is open.

View file

@ -64,8 +64,8 @@ fmt = "{:,}"
by = "tools/rule-coverage.py"
[gr_scenarios]
value = 24
text = "24"
value = 21
text = "21"
fmt = "{:,}"
by = "tools/rule-coverage.py"

View file

@ -820,74 +820,4 @@ mod tests {
);
}
}
/// GR-O01 says 2–6 players. Every scenario in the corpus was
/// 3-player until CB-WP-0008 T03, so the range was stated and tested
/// at one point. This plays all five counts with both policies.
#[test]
fn every_seat_count_in_gr_o01_plays_to_the_end() {
for players in 2..=6u8 {
for kind in ["random", "greedy"] {
let game = play(setup(players, 42), &mut policies(kind, players, 42))
.unwrap_or_else(|e| panic!("{players}p {kind}: {e}"));
let outcome = game
.state
.outcome
.as_ref()
.unwrap_or_else(|| panic!("{players}p {kind}: no outcome"));
assert_eq!(game.state.round, 5, "{players}p {kind}: GR-R09");
assert_eq!(
outcome.personal.len(),
usize::from(players),
"{players}p {kind}: every seat must be scored"
);
// K8 at each boundary, not only at three seats.
let again =
play(setup(players, 42), &mut policies(kind, players, 42)).expect("second run");
assert_eq!(
cb_events::state_hash_hex(&game.state),
cb_events::state_hash_hex(&again.state),
"{players}p {kind}: same seed must reproduce"
);
}
}
}
/// **A recorded finding, not a desired property.** With the standard
/// preset's placeholder Problem values (value = priority), the total
/// a game can possibly reach is below GR-E01's threshold at 2, 3 and
/// 4 players — group success is unreachable regardless of play. Only
/// 5–6p can clear its 9.
///
/// This test pins the arithmetic so the gap cannot close silently.
/// **It is expected to fail** when Problem values become scenario
/// data (GR-S01 calls the current fixture a stand-in); the failure is
/// the signal to delete it, not to re-tune it.
#[test]
fn the_standard_preset_cannot_reach_the_threshold_below_five_seats() {
let mut report = Vec::new();
for players in 2..=6u8 {
let initial = setup(players, 42);
let best: u32 = initial.problems.values().map(|p| u32::from(p.value)).sum();
let threshold = play(initial, &mut policies("greedy", players, 42))
.expect("game")
.state
.outcome
.expect("outcome")
.threshold;
report.push(format!("{players}p best {best} vs threshold {threshold}"));
if players < 5 {
assert!(
best < threshold,
"{players}p: best {best} now reaches threshold {threshold} — \
the fixture changed, delete this test"
);
} else {
assert!(
best >= threshold,
"{players}p: best {best} cannot reach threshold {threshold}"
);
}
}
println!("GR-E01 reachability: {}", report.join(", "));
}
}

View file

@ -1015,15 +1015,10 @@ impl Aggregate for GroundState {
impl GroundState {
/// GR-R06/R07: resolve in fixed step order, Lead first within a step.
///
/// Steps run: GROUND (GR-A10..A12), Support (GR-A03..A05), Attack
/// (GR-A06), DARVO (GR-D02..D07), INVESTIGATE (GR-A01), SOLVE
/// (GR-A02).
///
/// **This comment claimed DARVO and GROUND—OU/ND were "not yet
/// implemented" until 2026-08-01**, long after both landed with their
/// scenarios (`gr-d01`…`gr-d06`, `gr-a11`, `gr-a12`). Nothing checks
/// prose against code, which is the DFD class `make facts-check`
/// gates for *numbers* and cannot gate for claims like this one.
/// Step 3 (active DARVO stages) is not yet implemented — the DARVO
/// machine lands with GR-D02..D07 and no scenario claims coverage of
/// it. GROUND—OU and GROUND—ND (GR-A11/A12) each offer a three-way
/// choice that needs its own command and are likewise pending.
fn resolution_events(&self) -> Vec<GroundEvent> {
let mut events = Vec::new();
// A working copy so slot counts and Stress reflect earlier

View file

@ -100,39 +100,3 @@ This is the same class as the AM-2 `expect` that matched passing output —
an assertion that cannot distinguish the two worlds. The remedy that
worked both times was **counting what the harness examined**, not
strengthening the predicate.
## T03 — the 2–6 player range
Every scenario in the corpus was 3-player. Now: `gr-o01-two-player`,
`gr-o01-six-player`, `gr-e01-threshold-unreachable-2p`, plus all-bot
games at every count under both policies, each reproducing at the same
seed (K8 at the boundaries, not only in the middle), and the CLI
transcript test run at 2p, 3p and 6p.
**Nothing broke.** The rules are seat-count-generic and the range works.
What the boundaries exposed is arithmetic:
| seats | Problems | best possible total | GR-E01 threshold | group success |
|---|---|---|---|---|
| 2 | 2 | 3 | 5 | **unreachable** |
| 3 | 3 | 6 | 7 | **unreachable** |
| 4 | 3 | 6 | 7 | **unreachable** |
| 5 | 4 | 10 | 9 | reachable |
| 6 | 4 | 10 | 9 | reachable |
With the standard preset's placeholder values (value = priority), *no
play at all* can clear the threshold below five seats. GR-S01 calls the
fixture a stand-in for scenario Problem data, so this is evidence the
stand-in is not neutral — not that GR-E01 is wrong. It is pinned three
ways: a scenario that passes on the fact, a test that asserts the
arithmetic, and a `provisional` marker with `ground-game` as owner so it
ages in `make coverage` (now 6 provisional defaults).
The test carries its own delete-by condition: **it is expected to fail**
when Problem values become scenario data, and that failure is the signal
to delete it rather than re-tune it.
Two smaller notes, recorded and not acted on: at 2 players GR-L01's
second relation slot can never be used (there is only one possible
partner), and a human at P1 is always prompted before any other seat has
selected, so the seat order makes P1 a weaker test position than P3.

View file

@ -1,53 +0,0 @@
scenario: ground/gr-e01-threshold-unreachable-2p
description: >
GR-E01's threshold against the standard preset's Problem values, at the
2-player boundary. Both Problems are claimed — the best case available
— and the total is 3 against a threshold of 5. With the placeholder
fixture (value = priority) group success is unreachable at 2, 3 and 4
players; only 5–6p can reach its 9. Recorded here so the gap has a
failing-in-fact scenario rather than a paragraph, and marked provisional
because the fixture is explicitly a stand-in for scenario Problem data.
covers: [GR-E01, GR-E02, GR-O01]
provisional: true
provisional_owner: ground-game
provisional_raised: 2026-08-01
seed: 42
setup:
players: 2
preset: standard-2p
patch:
"round": 5
"mode": SharedGround
"problems.1.claimed_by": 0
"problems.2.claimed_by": 1
"problems.2.face_up": true
commands:
- actor: P1
cmd: select_action
args: { action: GROUND }
- actor: P2
cmd: select_action
args: { action: GROUND }
- actor: SYSTEM
cmd: reveal
- actor: P1
cmd: choose_ground_mode
args: { mode: GR }
- actor: P2
cmd: choose_ground_mode
args: { mode: GR }
- actor: SYSTEM
cmd: resolve
- actor: SYSTEM
cmd: end_round
expect:
events:
- kind: GameEnded
state:
# 1 + 2 = 3, the maximum any 2-player game of this preset can score.
"outcome.total": 3
"outcome.threshold": 5
"outcome.group_success": false
"round": 5
"step": End
rejects: []

View file

@ -1,51 +0,0 @@
scenario: ground/gr-o01-six-player
description: >
GR-O01's upper boundary, the counterpart to gr-o01-two-player. Six
seats exercise the 5–6p Problem set (GR-S01, four priorities) and the
longest resolution order (GR-R07, Lead first then clockwise with a
wrap).
covers: [GR-O01, GR-S01, GR-R07]
seed: 42
setup:
players: 6
preset: standard-6p
commands:
- actor: P1
cmd: select_action
args: { action: INVESTIGATE, problem: 2 }
- actor: P2
cmd: select_action
args: { action: INVESTIGATE, problem: 3 }
- actor: P3
cmd: select_action
args: { action: INVESTIGATE, problem: 4 }
- actor: P4
cmd: select_action
args: { action: SUPPORT, target: P5 }
- actor: P5
cmd: select_action
args: { action: GROUND }
- actor: P6
cmd: select_action
args: { action: ATTACK, target: P1 }
- actor: SYSTEM
cmd: reveal
- actor: P5
cmd: choose_ground_mode
args: { mode: GR }
- actor: P5
cmd: respond_to_support
args: { response: accept_bond }
- actor: SYSTEM
cmd: resolve
- actor: SYSTEM
cmd: end_round
expect:
events:
- kind: ProblemRevealed
- kind: RoundEnded
state:
# GR-S01: four Problems at six players.
"problems.4.face_up": true
"round": 2
rejects: []

View file

@ -1,40 +0,0 @@
scenario: ground/gr-o01-two-player
description: >
GR-O01's lower boundary. Every scenario in this corpus before
CB-WP-0008 T03 was 3-player, so a rule stated for a range (2–6) was
tested at one point. A full round at two seats: setup deals the 2p
Problem set (GR-S01), both seats select, and resolution runs with the
smallest possible seat order.
covers: [GR-O01, GR-S01, GR-R01, GR-R07]
seed: 42
setup:
players: 2
preset: standard-2p
commands:
- actor: P1
cmd: select_action
args: { action: INVESTIGATE, problem: 2 }
- actor: P2
cmd: select_action
args: { action: SUPPORT, target: P1 }
- actor: SYSTEM
cmd: reveal
- actor: P1
cmd: respond_to_support
args: { response: accept_bond }
- actor: SYSTEM
cmd: resolve
- actor: SYSTEM
cmd: end_round
expect:
events:
# GR-R06 order: Support (step 2) resolves before INVESTIGATE (step 4).
- kind: RelationFormed
- kind: ProblemRevealed
- kind: RoundEnded
state:
# GR-S01: two Problems at two players, not three.
"problems.2.face_up": true
"round": 2
"step": Select
rejects: []

View file

@ -142,42 +142,37 @@ mod tests {
list.iter().map(|s| s.to_string()).collect()
}
/// A scripted transcript plays a full game to `GameEnded`. "0"
/// always takes the first legal command, which is a real player input
/// and needs no knowledge of the board.
///
/// Run at **both GR-O01 boundaries and the middle** (T03): the CLI is
/// where a seat-count assumption would show up as a prompt nobody can
/// answer.
/// A scripted transcript plays a full 3-player game to `GameEnded`.
/// "0" always takes the first legal command, which is a real player
/// input and needs no knowledge of the board.
#[test]
fn a_scripted_transcript_plays_a_full_game() {
for players in [2u8, 3, 6] {
let config = Config {
seed: 42,
players,
human_seats: vec![0],
bot: "greedy".into(),
replay_dir: None,
record: None,
};
let script = "0\n".repeat(400);
let mut out: Vec<u8> = Vec::new();
let summary = table::play(&config, script.as_bytes(), &mut out)
.unwrap_or_else(|e| panic!("{players}p scripted game failed: {e}"));
let config = Config {
seed: 42,
players: 3,
human_seats: vec![0],
bot: "greedy".into(),
replay_dir: None,
record: None,
};
let script = "0\n".repeat(200);
let mut out: Vec<u8> = Vec::new();
let summary = table::play(&config, script.as_bytes(), &mut out)
.unwrap_or_else(|e| panic!("scripted game failed: {e}"));
assert_eq!(summary.rounds, 5, "{players}p: GR-R09 runs five rounds");
let text = String::from_utf8(out).expect("utf8");
assert!(text.contains("OUTCOME"), "{players}p: no outcome reported");
assert!(text.contains("you are P1"), "{players}p: no prompt");
assert_eq!(summary.rounds, 5, "GR-R09 runs five rounds");
let text = String::from_utf8(out).expect("utf8");
assert!(text.contains("OUTCOME"), "the game must report an outcome");
assert!(
text.contains("you are P1"),
"the human seat must be prompted"
);
// The transcript replays identically — through the scenario
// runner, not through a second call to the same driver.
match run::<GroundState>(&summary.scenario) {
RunOutcome::Passed { .. } => {}
RunOutcome::Failed { reason, .. } => {
panic!("{players}p replay failed: {reason}")
}
}
// The transcript replays identically — through the scenario
// runner, not through a second call to the same driver.
match run::<GroundState>(&summary.scenario) {
RunOutcome::Passed { .. } => {}
RunOutcome::Failed { reason, .. } => panic!("replay failed: {reason}"),
}
}

View file

@ -2,7 +2,7 @@
id: CB-WP-0008
kind: product
title: "Ship INTENT stage 0: a GROUND game you can actually play"
status: done
status: in_progress
state_hub_workstream_id: "ee960213-ffc9-4654-a962-35017a231fe2"
---
@ -97,7 +97,7 @@ records as a scenario and as a `.cbreplay` bundle. Notes in
```task
id: CB-WP-0008-T03
status: done
status: todo
priority: medium
state_hub_task_id: "438d21b7-98f3-43c0-b17f-6ae32bf176bd"
```
@ -112,17 +112,11 @@ all-bot game at each, and report what breaks. Finding that 2p or 6p does
*not* work is a legitimate and likely outcome — record it rather than
quietly narrowing GR-O01.
**Done 2026-08-01.** All five counts play to `GameEnded` under both
policies and reproduce at the same seed; three new scenarios; and the
finding is arithmetic, not a crash — GR-E01's threshold is **unreachable
at 2, 3 and 4 players** with the placeholder Problem values. Table in
`history/260801-cb-wp-0008-log.md`.
## Task: evidence and retrospective
```task
id: CB-WP-0008-T04
status: done
status: todo
priority: high
state_hub_task_id: "aa4ab660-d60c-49cc-b883-3fd20ccd0e3f"
```
@ -141,11 +135,3 @@ Commit `evidence/CB-EV-0007-stage-0.md`:
Then the retrospective the maintainer asked for: **six passes in, is the
loop earning its keep?** Answer with the spend curve, not with argument.
**Done 2026-08-01.** [CB-EV-0007](../evidence/CB-EV-0007-stage-0.md).
Stage 0 is shipped. The pass ran at **$0.123/response** — the best
recorded, tied with CB-WP-0004 — and is the first to satisfy SH-1 and
SH-2 in its own window. The meta budget read **61% OVER** and the
evidence file says why: the budget I built two tasks ago aggregates
cumulatively and cannot see a pass that is 100% product. Not fixed here;
v1.5 forbids opening meta work above the line.