T03: specs/CostAccounting.md — 15 contracts, 8 metrics, each with a command
Normative spec for the cost model (dedup by requestId, per-model and
per-TTL pricing), the attribution contract (ending-at-commit intervals
scoped by sessionId), and the reported shape (composition, not a total).
Every acceptance row names the command that produces its number, per
InnerLoop Step 4. AC-5..AC-8 are the positive control: cb-cost
--self-test must fail on a dedup violation, on zero responses, on a
missing subagent tree, and on 5m cache priced at the 1h rate.
The feasibility check earns its place here: AC-1's $92.87 is reachable
only if CA-06 holds. The earlier $92.21 target was reachable only by a
collector with the exact blind spot the survey documented.
MetricsAndScenarios 1a superseded. Two of its rules struck through
rather than deleted, because both were wrong in instructive ways: "if
the cache split is unknown, count all input at full price" would have
priced 80.5M cache reads at 10x, and "the hub already records tokens"
named as a source what is only ever a lossy sink. M-D2-TOK demoted --
tokens are not comparable across models or cache states.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 08:46:12 +02:00
|
|
|
|
# Cost Accounting
|
|
|
|
|
|
|
|
|
|
|
|
Status: **v1.0** — 2026-07-31. Derived from
|
|
|
|
|
|
[CB-RES-0002](../research/CB-RES-0002-cost-accounting.md) (approved) and
|
|
|
|
|
|
[ADR-0003](../decisions/ADR-0003-cost-accounting.md). Makes M-D2-CST
|
|
|
|
|
|
computable; supersedes its "uncomputable" disposition in CB-EV-0001.
|
|
|
|
|
|
|
|
|
|
|
|
Defines how the USD cost of agentic work is measured and attributed, so
|
|
|
|
|
|
that D2 claims about implementation efficiency are falsifiable.
|
|
|
|
|
|
|
|
|
|
|
|
---
|
|
|
|
|
|
|
|
|
|
|
|
## 1. The cost model
|
|
|
|
|
|
|
|
|
|
|
|
### 1.1 Unit of billing
|
|
|
|
|
|
|
|
|
|
|
|
The unit is one **API response**, identified by `requestId`. It is *not*
|
|
|
|
|
|
one JSONL line: a response is written as up to six lines split by content
|
|
|
|
|
|
block (`thinking`, `text`, `tool_use`), and every line repeats the same
|
|
|
|
|
|
`usage` object.
|
|
|
|
|
|
|
|
|
|
|
|
> **CA-01.** Cost is computed over responses deduplicated by `requestId`.
|
|
|
|
|
|
> Summing per line is a defect; it inflates by ≈1.9× on measured data.
|
|
|
|
|
|
|
T04: tools/cb-cost.py — and its positive control fires on first contact
Collector per ADR-0003: enumerates every transcript including the
subagents/ tree, dedups by requestId, prices per model and per cache
TTL, attributes on (prev_commit, this_commit] intervals, and reconciles
to the cent or aborts.
The positive control caught a real defect on its very first run against
real data, which is the entire argument for writing it:
CA-02 assumed usage is identical across the lines of one requestId.
True in the main transcript (206/206 groups, verified twice — by the
survey and by the adversarial reviewer). FALSE in the subagents/
tree, where output_tokens is a running count: one response reads
5, 5, 195 across its three lines. First-wins scored it at 5.
So CA-02 now splits: input-side counters are charged once and must be
identical (assertion retained); output_tokens resolves to the max
(CA-02a). AC-9 pins the exact 5,5,195 case as a regression test.
The acceptance target moved again as a result, for the third time:
$248.46 -> $92.21 -> $92.87 -> $93.32. AC-1 was computed by the same
first-wins method the tool just disproved, so the tool failing its
target was the tool being correct. Target updated, not the tool.
Measured at the pin: $92.21 main + $1.11 subagent = $93.32, residual
$0.000000. 32.5% UNATTRIBUTED, reported as its own line per CA-10.
make cost / cost-test / cost-pin wired; cost-test added to `make all`
and to CI, where it gates the collector's assertions without needing
transcripts present.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 08:50:04 +02:00
|
|
|
|
> **CA-02.** Dedup is asserted, not assumed. Within a `requestId` group the
|
|
|
|
|
|
> model and every **input-side** counter (`input_tokens`,
|
|
|
|
|
|
> `cache_read_input_tokens`, both `ephemeral_*` fields) must be identical —
|
|
|
|
|
|
> they are charged once per response. A divergence aborts the run.
|
|
|
|
|
|
|
|
|
|
|
|
> **CA-02a.** `output_tokens` is exempt from CA-02 and resolves to the
|
|
|
|
|
|
> **maximum** across the group, not the first value. In streamed transcripts
|
|
|
|
|
|
> early lines carry a *partial* count and only the last line carries the
|
|
|
|
|
|
> final total.
|
|
|
|
|
|
|
|
|
|
|
|
Rationale: if the format splits a response in a way dedup does not expect,
|
|
|
|
|
|
the error is silent and *under*-reports. The dangerous direction gets the
|
|
|
|
|
|
assertion.
|
|
|
|
|
|
|
|
|
|
|
|
*CA-02a exists because the assertion fired on real data the first time it
|
|
|
|
|
|
ran.* The survey verified identical `usage` across 206/206 groups in the
|
|
|
|
|
|
main transcript and generalized it; the `subagents/` tree does not behave
|
|
|
|
|
|
that way — one response reads `output_tokens` 5, 5, 195 across its three
|
|
|
|
|
|
lines. First-wins scored it at 5. That error moved the acceptance target
|
|
|
|
|
|
by $0.45.
|
T03: specs/CostAccounting.md — 15 contracts, 8 metrics, each with a command
Normative spec for the cost model (dedup by requestId, per-model and
per-TTL pricing), the attribution contract (ending-at-commit intervals
scoped by sessionId), and the reported shape (composition, not a total).
Every acceptance row names the command that produces its number, per
InnerLoop Step 4. AC-5..AC-8 are the positive control: cb-cost
--self-test must fail on a dedup violation, on zero responses, on a
missing subagent tree, and on 5m cache priced at the 1h rate.
The feasibility check earns its place here: AC-1's $92.87 is reachable
only if CA-06 holds. The earlier $92.21 target was reachable only by a
collector with the exact blind spot the survey documented.
MetricsAndScenarios 1a superseded. Two of its rules struck through
rather than deleted, because both were wrong in instructive ways: "if
the cache split is unknown, count all input at full price" would have
priced 80.5M cache reads at 10x, and "the hub already records tokens"
named as a source what is only ever a lossy sink. M-D2-TOK demoted --
tokens are not comparable across models or cache states.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 08:46:12 +02:00
|
|
|
|
|
|
|
|
|
|
### 1.2 Price formula
|
|
|
|
|
|
|
|
|
|
|
|
Per response, against `benchmarks/baselines/model-prices.toml`:
|
|
|
|
|
|
|
|
|
|
|
|
```text
|
|
|
|
|
|
cost = input_tokens × price.input
|
|
|
|
|
|
+ output_tokens × price.output
|
|
|
|
|
|
+ cache_read_input_tokens × price.input × cache.read (0.10)
|
|
|
|
|
|
+ ephemeral_5m_input_tokens × price.input × cache.write_5m (1.25)
|
|
|
|
|
|
+ ephemeral_1h_input_tokens × price.input × cache.write_1h (2.00)
|
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
> **CA-03.** Each response is priced at **its own** `message.model` rate. A
|
|
|
|
|
|
> session may mix models; CB-WP-0001 used three.
|
|
|
|
|
|
|
|
|
|
|
|
> **CA-04.** Cache writes are priced **per TTL**. The top-level
|
|
|
|
|
|
> `cache_creation_input_tokens` aggregate equals `ephemeral_5m +
|
|
|
|
|
|
> ephemeral_1h` and must never be priced at a single multiplier — doing so
|
|
|
|
|
|
> inflated a measured subagent transcript by 43%.
|
|
|
|
|
|
|
|
|
|
|
|
> **CA-05.** A response whose model is absent from the price sheet is
|
|
|
|
|
|
> reported as an unpriced line with its token counts, never dropped and
|
|
|
|
|
|
> never priced at a default.
|
|
|
|
|
|
|
|
|
|
|
|
### 1.3 Scope of a measurement
|
|
|
|
|
|
|
|
|
|
|
|
> **CA-06.** A measurement enumerates **every** transcript for the repo:
|
|
|
|
|
|
> `~/.claude/projects/<slug>/*.jsonl` and every
|
|
|
|
|
|
> `<session>/subagents/agent-*.jsonl`. Subagent cost is not in the main
|
|
|
|
|
|
> file and is invisible to a collector that reads one path.
|
|
|
|
|
|
|
|
|
|
|
|
> **CA-07.** Every committed number states its **pin** — a timestamp or
|
|
|
|
|
|
> commit. Transcripts are append-live: the file grows as the measuring
|
|
|
|
|
|
> session writes to it, and an unpinned total is not reproducible.
|
|
|
|
|
|
|
|
|
|
|
|
## 2. The attribution contract
|
|
|
|
|
|
|
|
|
|
|
|
> **CA-08.** A response is attributed to the task named by the next commit
|
|
|
|
|
|
> at or after it, within its own session:
|
|
|
|
|
|
> `interval := (prev_commit_time, this_commit_time]`, scoped by `sessionId`.
|
|
|
|
|
|
|
|
|
|
|
|
> **CA-09.** Timestamps are converted, never offset-subtracted. Commit
|
|
|
|
|
|
> times are parsed from `%cI` and converted to UTC; this repo carries two
|
|
|
|
|
|
> distinct offsets.
|
|
|
|
|
|
|
|
|
|
|
|
> **CA-10.** A response in a commit whose subject carries no `T##` tag is
|
|
|
|
|
|
> attributed to `UNATTRIBUTED`, which is **reported as its own line** in
|
|
|
|
|
|
> every table. On CB-WP-0001 this is 33% of cost ($30.32 of $92.21) — a
|
|
|
|
|
|
> per-task table is a view over roughly two-thirds of the money and says so
|
|
|
|
|
|
> wherever it appears.
|
|
|
|
|
|
|
|
|
|
|
|
> **CA-11.** Cost after the last commit is an **open remainder**, reported
|
|
|
|
|
|
> separately from `UNATTRIBUTED`. It is work not yet committed, not work
|
|
|
|
|
|
> without a task.
|
|
|
|
|
|
|
|
|
|
|
|
> **CA-12.** Attribution never keys on wall-clock alone. Two sessions
|
|
|
|
|
|
> overlapped 4 h 13 m on this repo carrying ~$12; only `sessionId`
|
|
|
|
|
|
> separates them.
|
|
|
|
|
|
|
|
|
|
|
|
## 3. Reported shape
|
|
|
|
|
|
|
|
|
|
|
|
> **CA-13.** Every report carries **composition** — the split across input,
|
|
|
|
|
|
> output, cache read, cache write 5m, cache write 1h — alongside the total.
|
|
|
|
|
|
> A total alone would have concealed the finding that motivated this work:
|
|
|
|
|
|
> 88.4% of spend is cache, at 256:1 cache-read to output tokens.
|
|
|
|
|
|
|
|
|
|
|
|
> **CA-14.** Reconciliation is asserted. Attributed + unattributed + open
|
|
|
|
|
|
> remainder + unpriced must equal the transcript total to the cent. A
|
|
|
|
|
|
> mismatch aborts rather than reporting.
|
|
|
|
|
|
|
|
|
|
|
|
> **CA-15.** Tables that a tool can emit are emitted by the tool. Every
|
|
|
|
|
|
> committed number in this capability's evidence is produced by a command,
|
|
|
|
|
|
> not typed. *(Adversarial review of CB-RES-0002 found every computed
|
|
|
|
|
|
> figure correct to the cent and three hand-typed markdown tables wrong.)*
|
|
|
|
|
|
|
|
|
|
|
|
## 4. Acceptance metrics
|
|
|
|
|
|
|
|
|
|
|
|
Each row names the command that produces its number, per InnerLoop §Step 4.
|
|
|
|
|
|
`cb-cost` is `tools/cb-cost` (T04).
|
|
|
|
|
|
|
|
|
|
|
|
| ID | Metric | Target | Instrument |
|
|
|
|
|
|
|---|---|---|---|
|
T04: tools/cb-cost.py — and its positive control fires on first contact
Collector per ADR-0003: enumerates every transcript including the
subagents/ tree, dedups by requestId, prices per model and per cache
TTL, attributes on (prev_commit, this_commit] intervals, and reconciles
to the cent or aborts.
The positive control caught a real defect on its very first run against
real data, which is the entire argument for writing it:
CA-02 assumed usage is identical across the lines of one requestId.
True in the main transcript (206/206 groups, verified twice — by the
survey and by the adversarial reviewer). FALSE in the subagents/
tree, where output_tokens is a running count: one response reads
5, 5, 195 across its three lines. First-wins scored it at 5.
So CA-02 now splits: input-side counters are charged once and must be
identical (assertion retained); output_tokens resolves to the max
(CA-02a). AC-9 pins the exact 5,5,195 case as a regression test.
The acceptance target moved again as a result, for the third time:
$248.46 -> $92.21 -> $92.87 -> $93.32. AC-1 was computed by the same
first-wins method the tool just disproved, so the tool failing its
target was the tool being correct. Target updated, not the tool.
Measured at the pin: $92.21 main + $1.11 subagent = $93.32, residual
$0.000000. 32.5% UNATTRIBUTED, reported as its own line per CA-10.
make cost / cost-test / cost-pin wired; cost-test added to `make all`
and to CI, where it gates the collector's assertions without needing
transcripts present.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 08:50:04 +02:00
|
|
|
|
| **AC-1** | reproduces the pinned CB-WP-0001 total | **$93.32** = $92.21 main + $1.11 subagent | `make cost-pin` |
|
T03: specs/CostAccounting.md — 15 contracts, 8 metrics, each with a command
Normative spec for the cost model (dedup by requestId, per-model and
per-TTL pricing), the attribution contract (ending-at-commit intervals
scoped by sessionId), and the reported shape (composition, not a total).
Every acceptance row names the command that produces its number, per
InnerLoop Step 4. AC-5..AC-8 are the positive control: cb-cost
--self-test must fail on a dedup violation, on zero responses, on a
missing subagent tree, and on 5m cache priced at the 1h rate.
The feasibility check earns its place here: AC-1's $92.87 is reachable
only if CA-06 holds. The earlier $92.21 target was reachable only by a
collector with the exact blind spot the survey documented.
MetricsAndScenarios 1a superseded. Two of its rules struck through
rather than deleted, because both were wrong in instructive ways: "if
the cache split is unknown, count all input at full price" would have
priced 80.5M cache reads at 10x, and "the hub already records tokens"
named as a source what is only ever a lossy sink. M-D2-TOK demoted --
tokens are not comparable across models or cache states.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 08:46:12 +02:00
|
|
|
|
| **AC-2** | reconciliation residual (CA-14) | **$0.00** exactly | same command, `reconciled: ok` line |
|
|
|
|
|
|
| **AC-3** | unattributed share reported (CA-10) | present, and **33%** on the pinned run | `cb-cost --pin fc76445 --by-task` |
|
|
|
|
|
|
| **AC-4** | composition reported (CA-13) | all five components present | `cb-cost --pin fc76445 --composition` |
|
T04: tools/cb-cost.py — and its positive control fires on first contact
Collector per ADR-0003: enumerates every transcript including the
subagents/ tree, dedups by requestId, prices per model and per cache
TTL, attributes on (prev_commit, this_commit] intervals, and reconciles
to the cent or aborts.
The positive control caught a real defect on its very first run against
real data, which is the entire argument for writing it:
CA-02 assumed usage is identical across the lines of one requestId.
True in the main transcript (206/206 groups, verified twice — by the
survey and by the adversarial reviewer). FALSE in the subagents/
tree, where output_tokens is a running count: one response reads
5, 5, 195 across its three lines. First-wins scored it at 5.
So CA-02 now splits: input-side counters are charged once and must be
identical (assertion retained); output_tokens resolves to the max
(CA-02a). AC-9 pins the exact 5,5,195 case as a regression test.
The acceptance target moved again as a result, for the third time:
$248.46 -> $92.21 -> $92.87 -> $93.32. AC-1 was computed by the same
first-wins method the tool just disproved, so the tool failing its
target was the tool being correct. Target updated, not the tool.
Measured at the pin: $92.21 main + $1.11 subagent = $93.32, residual
$0.000000. 32.5% UNATTRIBUTED, reported as its own line per CA-10.
make cost / cost-test / cost-pin wired; cost-test added to `make all`
and to CI, where it gates the collector's assertions without needing
transcripts present.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 08:50:04 +02:00
|
|
|
|
| **AC-5** | dedup invariant asserted (CA-02) | violation exits non-zero | `make cost-test` |
|
|
|
|
|
|
| **AC-6** | positive control: refuses to report on zero responses | exits non-zero | `make cost-test` |
|
|
|
|
|
|
| **AC-7** | subagent tree included (CA-06) | omitting it changes AC-1 by $1.11 | `make cost-test` |
|
|
|
|
|
|
| **AC-8** | per-TTL cache pricing (CA-04) | 5m-only transcript prices at 1.25× | `make cost-test` |
|
|
|
|
|
|
| **AC-9** | streamed partial output resolves to final (CA-02a) | 5,5,195 → 195, not 5 | `make cost-test` |
|
T03: specs/CostAccounting.md — 15 contracts, 8 metrics, each with a command
Normative spec for the cost model (dedup by requestId, per-model and
per-TTL pricing), the attribution contract (ending-at-commit intervals
scoped by sessionId), and the reported shape (composition, not a total).
Every acceptance row names the command that produces its number, per
InnerLoop Step 4. AC-5..AC-8 are the positive control: cb-cost
--self-test must fail on a dedup violation, on zero responses, on a
missing subagent tree, and on 5m cache priced at the 1h rate.
The feasibility check earns its place here: AC-1's $92.87 is reachable
only if CA-06 holds. The earlier $92.21 target was reachable only by a
collector with the exact blind spot the survey documented.
MetricsAndScenarios 1a superseded. Two of its rules struck through
rather than deleted, because both were wrong in instructive ways: "if
the cache split is unknown, count all input at full price" would have
priced 80.5M cache reads at 10x, and "the hub already records tokens"
named as a source what is only ever a lossy sink. M-D2-TOK demoted --
tokens are not comparable across models or cache states.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 08:46:12 +02:00
|
|
|
|
|
|
|
|
|
|
**AC-5 through AC-8 are the positive control.** Per InnerLoop v1.0 §Step 5,
|
|
|
|
|
|
a harness must assert it did the work it reports. `--self-test` runs each
|
|
|
|
|
|
assertion against a fixture whose expected value is known and fails loudly;
|
|
|
|
|
|
`make cost` runs it before any reported number.
|
|
|
|
|
|
|
|
|
|
|
|
## 5. Metric feasibility check
|
|
|
|
|
|
|
|
|
|
|
|
Per InnerLoop §Step 4, the acceptance table is checked against the
|
|
|
|
|
|
contracts in this same spec:
|
|
|
|
|
|
|
|
|
|
|
|
- AC-1's $92.87 is reachable only if CA-06 holds (both trees enumerated).
|
|
|
|
|
|
Under a main-file-only collector the target is unreachable — this is the
|
|
|
|
|
|
defect the adversarial review caught, where a target of $92.21 would have
|
|
|
|
|
|
been hit *only* by a broken collector.
|
|
|
|
|
|
- AC-3's 33% is a property of CB-WP-0001's commit subjects, not of the
|
|
|
|
|
|
collector. It is a regression pin on the fixture, not a quality target;
|
|
|
|
|
|
improving tagging discipline will change it, and that is expected.
|
|
|
|
|
|
- CA-07 (pinning) makes AC-1 reproducible; without it the target drifts
|
|
|
|
|
|
upward on every run and the test is meaningless.
|
|
|
|
|
|
|
|
|
|
|
|
## 6. Known limitations, stated with the numbers
|
|
|
|
|
|
|
|
|
|
|
|
- **Per-task cost covers ~67% of spend.** Structural: 19 of 33 commits
|
|
|
|
|
|
carry no task tag. Reported per CA-10, never silently dropped.
|
|
|
|
|
|
- **Work spanning a commit is assigned whole to the later task.** Bounded
|
|
|
|
|
|
by one interval: p50 6.9 min, p90 17.7 min, max 36.8 min.
|
|
|
|
|
|
- **The price sheet cannot express a time-boxed rate.** Sonnet's intro
|
|
|
|
|
|
price is a TOML comment, so the sheet is silently wrong for sonnet-priced
|
|
|
|
|
|
work ($0.17 at the pin, 0.19%). **This becomes an error, not a rounding
|
|
|
|
|
|
issue, on 2026-08-31** when the intro rate expires and the comment and
|
|
|
|
|
|
the data disagree in the other direction. Tracked as a schema defect
|
|
|
|
|
|
against the price sheet.
|
|
|
|
|
|
- **The State Hub cannot store what this spec measures.** Its token event
|
|
|
|
|
|
schema has `tokens_in`/`tokens_out` and no cache fields, so the dashboard
|
|
|
|
|
|
necessarily shows a lossy projection. This is a limitation of the sink,
|
|
|
|
|
|
not of the metric; `make cost` remains the authority.
|
|
|
|
|
|
|
|
|
|
|
|
## 7. Revisions to M-D2-CST
|
|
|
|
|
|
|
|
|
|
|
|
`specs/MetricsAndScenarios.md` §1a is superseded by this spec. M-D2-CST is
|
|
|
|
|
|
redefined from "tokens × pricepoint" — which named no instrument and was
|
|
|
|
|
|
never computed — to: **USD per completed workplan task, per CA-08
|
|
|
|
|
|
attribution, produced by `make cost`.** M-D2-TOK is retained but demoted:
|
|
|
|
|
|
tokens are the input to the cost model, not a comparable figure across
|
|
|
|
|
|
models or across cache states.
|