AM-4: gate scenario YAML, retarget on audited source, re-measure
Some checks failed
ci / check (push) Failing after 3s

Adopts both remediations from CB-EV-0001 §4 (maintainer decision).

Option A — serde_yaml is now optional behind cb-game-runtime's
`scenarios` feature. The scenario module, the ScenarioGame impl and the
string parsers behind it are cfg-gated; cb-sim opts in explicitly. Both
configurations compile and lint clean under -D warnings.

A trap worth recording: `default-features = false` on a *member*
dependency is silently ignored when the workspace dependency does not
specify it. The first attempt gated nothing while looking correct — the
build succeeded and cargo tree still showed all six YAML crates. Fixed
by setting it on the workspace dependency. This is the positive-control
failure mode in miniature: success was not evidence the change applied.

Retarget — AM-4 now measures third-party source under audit, split by
build configuration, replacing a crate count that was unreachable
without undoing K5/K7 and that does not compare across ecosystems.

Re-measured via the new `make dep-weight`, whose own positive control
refuses to report when any crate's source cannot be located:

  shipped runtime   23 crates   246,250 lines   target <=250,000  met
  dev toolchain     29 crates   317,021 lines   target <=350,000  met
  own source                      3,408 lines

Scenario tooling costs 70,771 lines a shipped game never compiles —
the split the single number was hiding.

Targets are set at current measurement plus headroom, so they bind on
future growth rather than retroactively passing what had failed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
tegwick 2026-07-31 03:35:41 +02:00
parent 8e11fc412e
commit 4be6e020ea
12 changed files with 271 additions and 43 deletions

View file

@ -2,6 +2,7 @@
id: CB-WP-0002
title: "Make agentic cost measurable, so D2 claims are falsifiable"
status: proposed
state_hub_workstream_id: "b7c22f69-fbe9-48df-9619-007db79ae338"
---
# Purpose
@ -45,6 +46,7 @@ positive control.
id: CB-WP-0002-T01
status: todo
priority: high
state_hub_task_id: "2694c2c1-0070-4d8e-b4fc-196b582b36d5"
```
Produce `research/CB-RES-0002-cost-accounting.md` per the InnerLoop
@ -67,6 +69,7 @@ accuracy and the hub to lead on durability.
id: CB-WP-0002-T02
status: todo
priority: high
state_hub_task_id: "eae248ab-f29f-4f11-9d20-e8145b0d822d"
```
Adversarial review of T01 first (InnerLoop §Step 2), committed as
@ -92,6 +95,7 @@ Gate: no collector code before this ADR is committed.
id: CB-WP-0002-T03
status: todo
priority: high
state_hub_task_id: "00d42ed2-4391-4580-aae2-06e3e151c69b"
```
Write `specs/CostAccounting.md`: the cost model (input, output, cache
@ -114,6 +118,7 @@ replace AM-12's definition with one that is computable.
id: CB-WP-0002-T04
status: todo
priority: medium
state_hub_task_id: "9eb8329b-5f41-477b-8cf3-2cda5ba8dbe8"
```
Implement the tool chosen in T02 (expected: `tools/cb-cost`). It reads
@ -137,6 +142,7 @@ each message at its own model's rate.
id: CB-WP-0002-T05
status: todo
priority: medium
state_hub_task_id: "bea4cc0a-e4d0-4077-9dc7-df7726a48f86"
```
Run the collector over the CB-WP-0001 session and commit
@ -160,6 +166,7 @@ cleared the bar that CB-WP-0001's AM-12 failed to clear.
id: CB-WP-0002-T06
status: todo
priority: low
state_hub_task_id: "1412263b-70c1-43e4-957d-1ad6c3203ca9"
```
Make cost collection automatic rather than remembered: add the
@ -174,6 +181,7 @@ the one command surface.
id: CB-WP-0002-T07
status: todo
priority: low
state_hub_task_id: "ebe58d91-be5f-4d5b-ba40-b03275b4eefc"
```
Revise `specs/InnerLoop.md` and `specs/MetricsAndScenarios.md` from what