Normative spec for the cost model (dedup by requestId, per-model and per-TTL pricing), the attribution contract (ending-at-commit intervals scoped by sessionId), and the reported shape (composition, not a total). Every acceptance row names the command that produces its number, per InnerLoop Step 4. AC-5..AC-8 are the positive control: cb-cost --self-test must fail on a dedup violation, on zero responses, on a missing subagent tree, and on 5m cache priced at the 1h rate. The feasibility check earns its place here: AC-1's $92.87 is reachable only if CA-06 holds. The earlier $92.21 target was reachable only by a collector with the exact blind spot the survey documented. MetricsAndScenarios 1a superseded. Two of its rules struck through rather than deleted, because both were wrong in instructive ways: "if the cache split is unknown, count all input at full price" would have priced 80.5M cache reads at 10x, and "the hub already records tokens" named as a source what is only ever a lossy sink. M-D2-TOK demoted -- tokens are not comparable across models or cache states. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
240 lines
9.7 KiB
Markdown
240 lines
9.7 KiB
Markdown
# Metrics and Scenarios
|
||
|
||
Status: **v0.1 draft** — instruments for the [InnerLoop](InnerLoop.md).
|
||
Everything the loop calls "better" is measured through the three artifact
|
||
kinds defined here: **scenarios** (correctness), **benchmarks** (speed and
|
||
cost), and **evidence files** (the committed comparison record).
|
||
|
||
---
|
||
|
||
## 1. Metric selection is itself a loop pass
|
||
|
||
Metrics are not invented ad hoc. Every metric used in an acceptance table
|
||
carries **provenance** — a one-line answer to "what state-of-the-art
|
||
practice does this metric derive from, and how does ours improve on it?"
|
||
|
||
| Provenance tag | Meaning |
|
||
|---|---|
|
||
| `adopted:<source>` | taken as-is from a named practice (e.g. `adopted:criterion` regression thresholds) |
|
||
| `adapted:<source>` | derived from a named practice, with the delta stated |
|
||
| `novel` | no known precedent — requires a sentence justifying why nothing existing fits |
|
||
|
||
A metric with no provenance line is invalid. This keeps the
|
||
assimilate-and-surpass discipline applied to the measuring instruments,
|
||
not only to the measured components.
|
||
|
||
### Standing metric set (v0.1)
|
||
|
||
Selected for the four dimensions; capability specs pick from these first
|
||
and add capability-specific rows only when these don't cover the claim.
|
||
|
||
| ID | Dimension | Metric | Unit | Provenance |
|
||
|---|---|---|---|---|
|
||
| M-D1-COV | D1 | numbered spec rules covered by ≥1 passing scenario | % | adapted:requirements-traceability (per-rule, not per-feature) |
|
||
| M-D1-SPL | D1 | spec lines per numbered rule | lines | novel — proxies statement simplicity; gameable, so paired with M-D1-COV |
|
||
| M-D2-LOC | D2 | source LOC excluding tests (tokei) | lines | adopted:tokei |
|
||
| M-D2-DEP | D2 | transitive dependency count (cargo tree) | crates | adopted:cargo-deny practice |
|
||
| M-D2-BLD | D2 | clean build / incremental build time | s | adopted:cargo timing |
|
||
| M-D2-TOK | D2 | tokens consumed per completed workplan task | tokens | novel — **demoted 2026-07-31**: an input to the cost model, not comparable across models or cache states |
|
||
| M-D2-CST | D2 | **cost** per completed workplan task, attributed per CA-08 | USD | adapted:anthropic-pricing — instrument: `make cost`; normative spec [CostAccounting.md](CostAccounting.md) |
|
||
| M-D3-THR | D3 | events applied per second, headless replay | events/s | adapted:criterion (throughput mode) |
|
||
| M-D3-LAT | D3 | p99 command→state-applied latency | µs | adopted:criterion |
|
||
| M-D3-MEM | D3 | peak resident memory during benchmark scenario | MB | adopted:/usr/bin/time -v |
|
||
| M-D4-API | D4 | public API items (cargo doc item count) | items | adapted:cargo-public-api |
|
||
| M-D4-LEAK | D4 | foreign types in canonical interfaces | count | novel — must be 0; enforced by grep/deny rule, the Clay-Borg hard rule |
|
||
| M-D4-SWAP | D4 | capability has null + reference impls passing the same conformance suite | bool | adapted:hexagonal-architecture port testing |
|
||
|
||
### 1a. Token cost accounting (M-D2-CST)
|
||
|
||
> **Superseded 2026-07-31 by [CostAccounting.md](CostAccounting.md)**, which
|
||
> is normative for the cost model, attribution, and acceptance metrics.
|
||
> This section is retained for the price-sheet location and the
|
||
> quality-gate rule; where the two disagree, CostAccounting.md wins.
|
||
>
|
||
> What changed and why: the definition below named no instrument and was
|
||
> never computed, so CB-WP-0001 recorded M-D2-CST as *uncomputable* while
|
||
> the data sat in the session transcripts. Three of its rules were also
|
||
> wrong in ways that cost real money to discover — see the corrections
|
||
> inline below.
|
||
|
||
Token counts are only comparable at a single pricepoint. Since work moves
|
||
between models (Fable for demanding passes, Sonnet/Opus for routine ones),
|
||
every task's token record carries the **model** it ran on, and cost is
|
||
computed against a committed price sheet:
|
||
|
||
```text
|
||
benchmarks/baselines/model-prices.toml # the price sheet, updated when prices change
|
||
```
|
||
|
||
```toml
|
||
# USD per million tokens; source: Anthropic pricing, recorded 2026-07-31
|
||
[claude-fable-5]
|
||
input = 10.00
|
||
output = 50.00
|
||
|
||
[claude-opus-5]
|
||
input = 5.00
|
||
output = 25.00
|
||
|
||
[claude-sonnet-5]
|
||
input = 3.00 # intro 2.00 through 2026-08-31
|
||
output = 15.00 # intro 10.00 through 2026-08-31
|
||
|
||
[claude-haiku-4-5]
|
||
input = 1.00
|
||
output = 5.00
|
||
|
||
[cache] # multipliers on the input price
|
||
read = 0.1
|
||
write_5m = 1.25
|
||
write_1h = 2.0
|
||
```
|
||
|
||
Rules:
|
||
|
||
- ~~`cost = (in_tokens × input + out_tokens × output) / 1e6` … If cache
|
||
split is unknown, count all input at full price and note it — cost is then
|
||
an upper bound.~~ **Corrected:** the cache split is never unknown; it is
|
||
in every transcript. Treating it as unknown would have priced 80.5M cache
|
||
reads at 10× their rate. See CostAccounting.md §1.2 (CA-03, CA-04) — cache
|
||
writes bill at two different TTL rates and must not be aggregated.
|
||
- ~~The state-hub task close (`update_task_status`) already records tokens
|
||
and `model`.~~ **Corrected:** the hub schema has no cache fields and
|
||
cannot represent 88% of spend, and the figures it recorded for CB-WP-0001
|
||
were estimates in error by ~100%. The hub is a **sink** for numbers
|
||
computed by `make cost`, never a source. See CostAccounting.md §6.
|
||
- **Cheaper is only better at equal quality**: M-D2-CST verdicts are valid
|
||
only alongside passing scenarios/metrics from the same run — a cheap
|
||
failed pass scores nothing.
|
||
- Prices are `adopted:anthropic-pricing` with a recorded date; a stale sheet
|
||
(> 90 days or known price change) invalidates new `better` verdicts on
|
||
M-D2-CST until refreshed.
|
||
|
||
Determinism is not a metric but an **invariant**: N replays of the same
|
||
seed and command log must produce bit-identical state hashes. Invariant
|
||
violations fail the run regardless of metric values.
|
||
|
||
---
|
||
|
||
## 2. Scenario format
|
||
|
||
Scenarios are the correctness currency: executable, declarative, diffable.
|
||
One file = one scenario. Location: `scenarios/<game-or-capability>/<slug>.yaml`.
|
||
|
||
```yaml
|
||
scenario: ground/darvo-interrupted-by-ground # id = path without extension
|
||
description: GROUND practice interrupts a DARVO sequence at the Attack step.
|
||
covers: [R-041, R-052, R-053] # numbered rules from the capability spec
|
||
seed: 42
|
||
setup:
|
||
players: 3
|
||
preset: standard-3p # named setup preset from the game spec
|
||
patch: # optional explicit state overrides
|
||
relationships:
|
||
- {from: P1, to: P2, kind: rivalry, strength: 2}
|
||
commands: # ordered; actor-tagged
|
||
- {actor: P1, cmd: trigger_darvo, target: P2}
|
||
- {actor: P2, cmd: play_ground, target: P1}
|
||
expect:
|
||
events: # ordered subsequence that must occur
|
||
- {type: DarvoInterrupted, step: attack}
|
||
state: # end-state assertions, dot-path = value
|
||
darvo_sequences: []
|
||
relationships[P1->P2].strength: 1
|
||
rejects: [] # commands above that must be rejected, by index
|
||
```
|
||
|
||
Rules:
|
||
|
||
- `covers` is what feeds M-D1-COV; a scenario without `covers` counts for
|
||
nothing.
|
||
- Assertions are **partial**: only listed paths are checked. Full-state
|
||
golden comparison is opt-in via `expect.state_hash`.
|
||
- Every scenario must be deterministic given `seed`; the runner executes
|
||
each scenario twice and fails on hash divergence (cheap standing
|
||
determinism check).
|
||
- A failing run writes `replays/<scenario-id>.cbreplay` (see §4).
|
||
|
||
---
|
||
|
||
## 3. Benchmarks and baselines
|
||
|
||
Benchmarks live in `benchmarks/` as Criterion benches driving scenario
|
||
files (a benchmark is a scenario run at scale — no separate workload
|
||
format).
|
||
|
||
Baselines are **committed numbers**, recorded once per approved survey and
|
||
updated only by an explicit ADR:
|
||
|
||
```text
|
||
benchmarks/baselines/<capability>.toml
|
||
```
|
||
|
||
```toml
|
||
[M-D3-THR]
|
||
value = 120000
|
||
unit = "events/s"
|
||
source = "boardgame.io v0.50, measured locally, 3-player synthetic log"
|
||
recorded = 2026-07-31
|
||
machine = "bnt-lap001"
|
||
|
||
[M-D2-LOC]
|
||
value = 8400
|
||
unit = "lines"
|
||
source = "boardgame.io core, cloc, cited from CB-RES-0001"
|
||
```
|
||
|
||
- Comparisons are same-machine where `machine` is set; cross-machine
|
||
numbers are marked `provenance = cited` and treated as directional.
|
||
Per the runnable-baseline option in [InnerLoop.md](InnerLoop.md) §Step 1,
|
||
evidence rows compared only against cited numbers cap their verdict at
|
||
`parity`; a `better` verdict requires a locally measured baseline from a
|
||
fidelity-noted harness.
|
||
- Regression rule (adopted:criterion): a merge-blocking regression is
|
||
>3% on any D3 metric against **our own** last evidence file, independent
|
||
of the SOTA baseline.
|
||
|
||
---
|
||
|
||
## 4. Replay bundle
|
||
|
||
`*.cbreplay` is a directory (or tar) with exactly:
|
||
|
||
```text
|
||
manifest.yaml # scenario id, seed, git commit, schema versions
|
||
commands.log # the full ordered command stream (serialized events optional)
|
||
initial.snapshot # starting state
|
||
expected.yaml # the assertions that failed, with expected vs actual
|
||
```
|
||
|
||
Contract: `cb replay <bundle>` (until the CLI exists: the scenario runner's
|
||
`--replay` flag) re-executes the bundle headless and must reproduce the
|
||
failure bit-identically. A bug report without a replay bundle is
|
||
information; with one, it is work an agent can start.
|
||
|
||
---
|
||
|
||
## 5. Evidence file
|
||
|
||
`evidence/CB-EV-NNNN-<slug>.md` — the committed close-out of a loop pass:
|
||
|
||
```markdown
|
||
# CB-EV-NNNN: <capability>
|
||
research: CB-RES-NNNN adr: ADR-NNNN spec: specs/<Capability>.md
|
||
commit: <sha of measured tree>
|
||
|
||
| Metric | Baseline | Ours | Verdict |
|
||
|---|---|---|---|
|
||
| M-D3-THR | 120000 events/s (boardgame.io) | 410000 events/s | better |
|
||
| ...every acceptance row, no `unmeasured`... |
|
||
|
||
## Task cost log
|
||
| Task | Model | Tokens in/out | Cost (USD, per price sheet) | Iterations |
|
||
|
||
## Retrospective
|
||
<what the loop itself should change — one paragraph minimum>
|
||
```
|
||
|
||
Verdicts: `better / parity / worse`. A `worse` row does not necessarily
|
||
fail the pass — the ADR's declared trade governs — but an undeclared
|
||
`worse` does.
|