clay-borg/specs/MetricsAndScenarios.md
tegwick b96cd94a64 T03: specs/CostAccounting.md — 15 contracts, 8 metrics, each with a command
Normative spec for the cost model (dedup by requestId, per-model and
per-TTL pricing), the attribution contract (ending-at-commit intervals
scoped by sessionId), and the reported shape (composition, not a total).

Every acceptance row names the command that produces its number, per
InnerLoop Step 4. AC-5..AC-8 are the positive control: cb-cost
--self-test must fail on a dedup violation, on zero responses, on a
missing subagent tree, and on 5m cache priced at the 1h rate.

The feasibility check earns its place here: AC-1's $92.87 is reachable
only if CA-06 holds. The earlier $92.21 target was reachable only by a
collector with the exact blind spot the survey documented.

MetricsAndScenarios 1a superseded. Two of its rules struck through
rather than deleted, because both were wrong in instructive ways: "if
the cache split is unknown, count all input at full price" would have
priced 80.5M cache reads at 10x, and "the hub already records tokens"
named as a source what is only ever a lossy sink. M-D2-TOK demoted --
tokens are not comparable across models or cache states.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 08:46:12 +02:00

240 lines
9.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Metrics and Scenarios
Status: **v0.1 draft** — instruments for the [InnerLoop](InnerLoop.md).
Everything the loop calls "better" is measured through the three artifact
kinds defined here: **scenarios** (correctness), **benchmarks** (speed and
cost), and **evidence files** (the committed comparison record).
---
## 1. Metric selection is itself a loop pass
Metrics are not invented ad hoc. Every metric used in an acceptance table
carries **provenance** — a one-line answer to "what state-of-the-art
practice does this metric derive from, and how does ours improve on it?"
| Provenance tag | Meaning |
|---|---|
| `adopted:<source>` | taken as-is from a named practice (e.g. `adopted:criterion` regression thresholds) |
| `adapted:<source>` | derived from a named practice, with the delta stated |
| `novel` | no known precedent — requires a sentence justifying why nothing existing fits |
A metric with no provenance line is invalid. This keeps the
assimilate-and-surpass discipline applied to the measuring instruments,
not only to the measured components.
### Standing metric set (v0.1)
Selected for the four dimensions; capability specs pick from these first
and add capability-specific rows only when these don't cover the claim.
| ID | Dimension | Metric | Unit | Provenance |
|---|---|---|---|---|
| M-D1-COV | D1 | numbered spec rules covered by ≥1 passing scenario | % | adapted:requirements-traceability (per-rule, not per-feature) |
| M-D1-SPL | D1 | spec lines per numbered rule | lines | novel — proxies statement simplicity; gameable, so paired with M-D1-COV |
| M-D2-LOC | D2 | source LOC excluding tests (tokei) | lines | adopted:tokei |
| M-D2-DEP | D2 | transitive dependency count (cargo tree) | crates | adopted:cargo-deny practice |
| M-D2-BLD | D2 | clean build / incremental build time | s | adopted:cargo timing |
| M-D2-TOK | D2 | tokens consumed per completed workplan task | tokens | novel — **demoted 2026-07-31**: an input to the cost model, not comparable across models or cache states |
| M-D2-CST | D2 | **cost** per completed workplan task, attributed per CA-08 | USD | adapted:anthropic-pricing — instrument: `make cost`; normative spec [CostAccounting.md](CostAccounting.md) |
| M-D3-THR | D3 | events applied per second, headless replay | events/s | adapted:criterion (throughput mode) |
| M-D3-LAT | D3 | p99 command→state-applied latency | µs | adopted:criterion |
| M-D3-MEM | D3 | peak resident memory during benchmark scenario | MB | adopted:/usr/bin/time -v |
| M-D4-API | D4 | public API items (cargo doc item count) | items | adapted:cargo-public-api |
| M-D4-LEAK | D4 | foreign types in canonical interfaces | count | novel — must be 0; enforced by grep/deny rule, the Clay-Borg hard rule |
| M-D4-SWAP | D4 | capability has null + reference impls passing the same conformance suite | bool | adapted:hexagonal-architecture port testing |
### 1a. Token cost accounting (M-D2-CST)
> **Superseded 2026-07-31 by [CostAccounting.md](CostAccounting.md)**, which
> is normative for the cost model, attribution, and acceptance metrics.
> This section is retained for the price-sheet location and the
> quality-gate rule; where the two disagree, CostAccounting.md wins.
>
> What changed and why: the definition below named no instrument and was
> never computed, so CB-WP-0001 recorded M-D2-CST as *uncomputable* while
> the data sat in the session transcripts. Three of its rules were also
> wrong in ways that cost real money to discover — see the corrections
> inline below.
Token counts are only comparable at a single pricepoint. Since work moves
between models (Fable for demanding passes, Sonnet/Opus for routine ones),
every task's token record carries the **model** it ran on, and cost is
computed against a committed price sheet:
```text
benchmarks/baselines/model-prices.toml # the price sheet, updated when prices change
```
```toml
# USD per million tokens; source: Anthropic pricing, recorded 2026-07-31
[claude-fable-5]
input = 10.00
output = 50.00
[claude-opus-5]
input = 5.00
output = 25.00
[claude-sonnet-5]
input = 3.00 # intro 2.00 through 2026-08-31
output = 15.00 # intro 10.00 through 2026-08-31
[claude-haiku-4-5]
input = 1.00
output = 5.00
[cache] # multipliers on the input price
read = 0.1
write_5m = 1.25
write_1h = 2.0
```
Rules:
- ~~`cost = (in_tokens × input + out_tokens × output) / 1e6` … If cache
split is unknown, count all input at full price and note it — cost is then
an upper bound.~~ **Corrected:** the cache split is never unknown; it is
in every transcript. Treating it as unknown would have priced 80.5M cache
reads at 10× their rate. See CostAccounting.md §1.2 (CA-03, CA-04) — cache
writes bill at two different TTL rates and must not be aggregated.
- ~~The state-hub task close (`update_task_status`) already records tokens
and `model`.~~ **Corrected:** the hub schema has no cache fields and
cannot represent 88% of spend, and the figures it recorded for CB-WP-0001
were estimates in error by ~100%. The hub is a **sink** for numbers
computed by `make cost`, never a source. See CostAccounting.md §6.
- **Cheaper is only better at equal quality**: M-D2-CST verdicts are valid
only alongside passing scenarios/metrics from the same run — a cheap
failed pass scores nothing.
- Prices are `adopted:anthropic-pricing` with a recorded date; a stale sheet
(> 90 days or known price change) invalidates new `better` verdicts on
M-D2-CST until refreshed.
Determinism is not a metric but an **invariant**: N replays of the same
seed and command log must produce bit-identical state hashes. Invariant
violations fail the run regardless of metric values.
---
## 2. Scenario format
Scenarios are the correctness currency: executable, declarative, diffable.
One file = one scenario. Location: `scenarios/<game-or-capability>/<slug>.yaml`.
```yaml
scenario: ground/darvo-interrupted-by-ground # id = path without extension
description: GROUND practice interrupts a DARVO sequence at the Attack step.
covers: [R-041, R-052, R-053] # numbered rules from the capability spec
seed: 42
setup:
players: 3
preset: standard-3p # named setup preset from the game spec
patch: # optional explicit state overrides
relationships:
- {from: P1, to: P2, kind: rivalry, strength: 2}
commands: # ordered; actor-tagged
- {actor: P1, cmd: trigger_darvo, target: P2}
- {actor: P2, cmd: play_ground, target: P1}
expect:
events: # ordered subsequence that must occur
- {type: DarvoInterrupted, step: attack}
state: # end-state assertions, dot-path = value
darvo_sequences: []
relationships[P1->P2].strength: 1
rejects: [] # commands above that must be rejected, by index
```
Rules:
- `covers` is what feeds M-D1-COV; a scenario without `covers` counts for
nothing.
- Assertions are **partial**: only listed paths are checked. Full-state
golden comparison is opt-in via `expect.state_hash`.
- Every scenario must be deterministic given `seed`; the runner executes
each scenario twice and fails on hash divergence (cheap standing
determinism check).
- A failing run writes `replays/<scenario-id>.cbreplay` (see §4).
---
## 3. Benchmarks and baselines
Benchmarks live in `benchmarks/` as Criterion benches driving scenario
files (a benchmark is a scenario run at scale — no separate workload
format).
Baselines are **committed numbers**, recorded once per approved survey and
updated only by an explicit ADR:
```text
benchmarks/baselines/<capability>.toml
```
```toml
[M-D3-THR]
value = 120000
unit = "events/s"
source = "boardgame.io v0.50, measured locally, 3-player synthetic log"
recorded = 2026-07-31
machine = "bnt-lap001"
[M-D2-LOC]
value = 8400
unit = "lines"
source = "boardgame.io core, cloc, cited from CB-RES-0001"
```
- Comparisons are same-machine where `machine` is set; cross-machine
numbers are marked `provenance = cited` and treated as directional.
Per the runnable-baseline option in [InnerLoop.md](InnerLoop.md) §Step 1,
evidence rows compared only against cited numbers cap their verdict at
`parity`; a `better` verdict requires a locally measured baseline from a
fidelity-noted harness.
- Regression rule (adopted:criterion): a merge-blocking regression is
>3% on any D3 metric against **our own** last evidence file, independent
of the SOTA baseline.
---
## 4. Replay bundle
`*.cbreplay` is a directory (or tar) with exactly:
```text
manifest.yaml # scenario id, seed, git commit, schema versions
commands.log # the full ordered command stream (serialized events optional)
initial.snapshot # starting state
expected.yaml # the assertions that failed, with expected vs actual
```
Contract: `cb replay <bundle>` (until the CLI exists: the scenario runner's
`--replay` flag) re-executes the bundle headless and must reproduce the
failure bit-identically. A bug report without a replay bundle is
information; with one, it is work an agent can start.
---
## 5. Evidence file
`evidence/CB-EV-NNNN-<slug>.md` — the committed close-out of a loop pass:
```markdown
# CB-EV-NNNN: <capability>
research: CB-RES-NNNN adr: ADR-NNNN spec: specs/<Capability>.md
commit: <sha of measured tree>
| Metric | Baseline | Ours | Verdict |
|---|---|---|---|
| M-D3-THR | 120000 events/s (boardgame.io) | 410000 events/s | better |
| ...every acceptance row, no `unmeasured`... |
## Task cost log
| Task | Model | Tokens in/out | Cost (USD, per price sheet) | Iterations |
## Retrospective
<what the loop itself should change one paragraph minimum>
```
Verdicts: `better / parity / worse`. A `worse` row does not necessarily
fail the pass — the ADR's declared trade governs — but an undeclared
`worse` does.