clay-borg/specs/MetricsAndScenarios.md

169 lines
6.6 KiB
Markdown

# Metrics and Scenarios
Status: **v0.1 draft** — instruments for the [InnerLoop](InnerLoop.md).
Everything the loop calls "better" is measured through the three artifact
kinds defined here: **scenarios** (correctness), **benchmarks** (speed and
cost), and **evidence files** (the committed comparison record).
---
## 1. Metric selection is itself a loop pass
Metrics are not invented ad hoc. Every metric used in an acceptance table
carries **provenance** — a one-line answer to "what state-of-the-art
practice does this metric derive from, and how does ours improve on it?"
| Provenance tag | Meaning |
|---|---|
| `adopted:<source>` | taken as-is from a named practice (e.g. `adopted:criterion` regression thresholds) |
| `adapted:<source>` | derived from a named practice, with the delta stated |
| `novel` | no known precedent — requires a sentence justifying why nothing existing fits |
A metric with no provenance line is invalid. This keeps the
assimilate-and-surpass discipline applied to the measuring instruments,
not only to the measured components.
### Standing metric set (v0.1)
Selected for the four dimensions; capability specs pick from these first
and add capability-specific rows only when these don't cover the claim.
| ID | Dimension | Metric | Unit | Provenance |
|---|---|---|---|---|
| M-D1-COV | D1 | numbered spec rules covered by ≥1 passing scenario | % | adapted:requirements-traceability (per-rule, not per-feature) |
| M-D1-SPL | D1 | spec lines per numbered rule | lines | novel — proxies statement simplicity; gameable, so paired with M-D1-COV |
| M-D2-LOC | D2 | source LOC excluding tests (tokei) | lines | adopted:tokei |
| M-D2-DEP | D2 | transitive dependency count (cargo tree) | crates | adopted:cargo-deny practice |
| M-D2-BLD | D2 | clean build / incremental build time | s | adopted:cargo timing |
| M-D2-TOK | D2 | tokens consumed per completed workplan task | tokens | novel — agentic-efficiency core metric; recorded per task in evidence |
| M-D3-THR | D3 | events applied per second, headless replay | events/s | adapted:criterion (throughput mode) |
| M-D3-LAT | D3 | p99 command→state-applied latency | µs | adopted:criterion |
| M-D3-MEM | D3 | peak resident memory during benchmark scenario | MB | adopted:/usr/bin/time -v |
| M-D4-API | D4 | public API items (cargo doc item count) | items | adapted:cargo-public-api |
| M-D4-LEAK | D4 | foreign types in canonical interfaces | count | novel — must be 0; enforced by grep/deny rule, the Clay-Borg hard rule |
| M-D4-SWAP | D4 | capability has null + reference impls passing the same conformance suite | bool | adapted:hexagonal-architecture port testing |
Determinism is not a metric but an **invariant**: N replays of the same
seed and command log must produce bit-identical state hashes. Invariant
violations fail the run regardless of metric values.
---
## 2. Scenario format
Scenarios are the correctness currency: executable, declarative, diffable.
One file = one scenario. Location: `scenarios/<game-or-capability>/<slug>.yaml`.
```yaml
scenario: ground/darvo-interrupted-by-ground # id = path without extension
description: GROUND practice interrupts a DARVO sequence at the Attack step.
covers: [R-041, R-052, R-053] # numbered rules from the capability spec
seed: 42
setup:
players: 3
preset: standard-3p # named setup preset from the game spec
patch: # optional explicit state overrides
relationships:
- {from: P1, to: P2, kind: rivalry, strength: 2}
commands: # ordered; actor-tagged
- {actor: P1, cmd: trigger_darvo, target: P2}
- {actor: P2, cmd: play_ground, target: P1}
expect:
events: # ordered subsequence that must occur
- {type: DarvoInterrupted, step: attack}
state: # end-state assertions, dot-path = value
darvo_sequences: []
relationships[P1->P2].strength: 1
rejects: [] # commands above that must be rejected, by index
```
Rules:
- `covers` is what feeds M-D1-COV; a scenario without `covers` counts for
nothing.
- Assertions are **partial**: only listed paths are checked. Full-state
golden comparison is opt-in via `expect.state_hash`.
- Every scenario must be deterministic given `seed`; the runner executes
each scenario twice and fails on hash divergence (cheap standing
determinism check).
- A failing run writes `replays/<scenario-id>.cbreplay` (see §4).
---
## 3. Benchmarks and baselines
Benchmarks live in `benchmarks/` as Criterion benches driving scenario
files (a benchmark is a scenario run at scale — no separate workload
format).
Baselines are **committed numbers**, recorded once per approved survey and
updated only by an explicit ADR:
```text
benchmarks/baselines/<capability>.toml
```
```toml
[M-D3-THR]
value = 120000
unit = "events/s"
source = "boardgame.io v0.50, measured locally, 3-player synthetic log"
recorded = 2026-07-31
machine = "bnt-lap001"
[M-D2-LOC]
value = 8400
unit = "lines"
source = "boardgame.io core, cloc, cited from CB-RES-0001"
```
- Comparisons are same-machine where `machine` is set; cross-machine
numbers are marked `provenance = cited` and treated as directional.
- Regression rule (adopted:criterion): a merge-blocking regression is
>3% on any D3 metric against **our own** last evidence file, independent
of the SOTA baseline.
---
## 4. Replay bundle
`*.cbreplay` is a directory (or tar) with exactly:
```text
manifest.yaml # scenario id, seed, git commit, schema versions
commands.log # the full ordered command stream (serialized events optional)
initial.snapshot # starting state
expected.yaml # the assertions that failed, with expected vs actual
```
Contract: `cb replay <bundle>` (until the CLI exists: the scenario runner's
`--replay` flag) re-executes the bundle headless and must reproduce the
failure bit-identically. A bug report without a replay bundle is
information; with one, it is work an agent can start.
---
## 5. Evidence file
`evidence/CB-EV-NNNN-<slug>.md` — the committed close-out of a loop pass:
```markdown
# CB-EV-NNNN: <capability>
research: CB-RES-NNNN adr: ADR-NNNN spec: specs/<Capability>.md
commit: <sha of measured tree>
| Metric | Baseline | Ours | Verdict |
|---|---|---|---|
| M-D3-THR | 120000 events/s (boardgame.io) | 410000 events/s | better |
| ...every acceptance row, no `unmeasured`... |
## Task token log
| Task | Tokens (approx) | Iterations |
## Retrospective
<what the loop itself should change one paragraph minimum>
```
Verdicts: `better / parity / worse`. A `worse` row does not necessarily
fail the pass — the ADR's declared trade governs — but an undeclared
`worse` does.