clay-borg/specs/CostAccounting.md
tegwick e008b1e787 T04: tools/cb-cost.py — and its positive control fires on first contact
Collector per ADR-0003: enumerates every transcript including the
subagents/ tree, dedups by requestId, prices per model and per cache
TTL, attributes on (prev_commit, this_commit] intervals, and reconciles
to the cent or aborts.

The positive control caught a real defect on its very first run against
real data, which is the entire argument for writing it:

  CA-02 assumed usage is identical across the lines of one requestId.
  True in the main transcript (206/206 groups, verified twice — by the
  survey and by the adversarial reviewer). FALSE in the subagents/
  tree, where output_tokens is a running count: one response reads
  5, 5, 195 across its three lines. First-wins scored it at 5.

So CA-02 now splits: input-side counters are charged once and must be
identical (assertion retained); output_tokens resolves to the max
(CA-02a). AC-9 pins the exact 5,5,195 case as a regression test.

The acceptance target moved again as a result, for the third time:
$248.46 -> $92.21 -> $92.87 -> $93.32. AC-1 was computed by the same
first-wins method the tool just disproved, so the tool failing its
target was the tool being correct. Target updated, not the tool.

Measured at the pin: $92.21 main + $1.11 subagent = $93.32, residual
$0.000000. 32.5% UNATTRIBUTED, reported as its own line per CA-10.

make cost / cost-test / cost-pin wired; cost-test added to `make all`
and to CI, where it gates the collector's assertions without needing
transcripts present.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 08:50:04 +02:00

8.6 KiB
Raw Blame History

Cost Accounting

Status: v1.0 — 2026-07-31. Derived from CB-RES-0002 (approved) and ADR-0003. Makes M-D2-CST computable; supersedes its "uncomputable" disposition in CB-EV-0001.

Defines how the USD cost of agentic work is measured and attributed, so that D2 claims about implementation efficiency are falsifiable.


1. The cost model

1.1 Unit of billing

The unit is one API response, identified by requestId. It is not one JSONL line: a response is written as up to six lines split by content block (thinking, text, tool_use), and every line repeats the same usage object.

CA-01. Cost is computed over responses deduplicated by requestId. Summing per line is a defect; it inflates by ≈1.9× on measured data.

CA-02. Dedup is asserted, not assumed. Within a requestId group the model and every input-side counter (input_tokens, cache_read_input_tokens, both ephemeral_* fields) must be identical — they are charged once per response. A divergence aborts the run.

CA-02a. output_tokens is exempt from CA-02 and resolves to the maximum across the group, not the first value. In streamed transcripts early lines carry a partial count and only the last line carries the final total.

Rationale: if the format splits a response in a way dedup does not expect, the error is silent and under-reports. The dangerous direction gets the assertion.

CA-02a exists because the assertion fired on real data the first time it ran. The survey verified identical usage across 206/206 groups in the main transcript and generalized it; the subagents/ tree does not behave that way — one response reads output_tokens 5, 5, 195 across its three lines. First-wins scored it at 5. That error moved the acceptance target by $0.45.

1.2 Price formula

Per response, against benchmarks/baselines/model-prices.toml:

cost = input_tokens              × price.input
     + output_tokens             × price.output
     + cache_read_input_tokens   × price.input × cache.read       (0.10)
     + ephemeral_5m_input_tokens × price.input × cache.write_5m   (1.25)
     + ephemeral_1h_input_tokens × price.input × cache.write_1h   (2.00)

CA-03. Each response is priced at its own message.model rate. A session may mix models; CB-WP-0001 used three.

CA-04. Cache writes are priced per TTL. The top-level cache_creation_input_tokens aggregate equals ephemeral_5m + ephemeral_1h and must never be priced at a single multiplier — doing so inflated a measured subagent transcript by 43%.

CA-05. A response whose model is absent from the price sheet is reported as an unpriced line with its token counts, never dropped and never priced at a default.

1.3 Scope of a measurement

CA-06. A measurement enumerates every transcript for the repo: ~/.claude/projects/<slug>/*.jsonl and every <session>/subagents/agent-*.jsonl. Subagent cost is not in the main file and is invisible to a collector that reads one path.

CA-07. Every committed number states its pin — a timestamp or commit. Transcripts are append-live: the file grows as the measuring session writes to it, and an unpinned total is not reproducible.

2. The attribution contract

CA-08. A response is attributed to the task named by the next commit at or after it, within its own session: interval := (prev_commit_time, this_commit_time], scoped by sessionId.

CA-09. Timestamps are converted, never offset-subtracted. Commit times are parsed from %cI and converted to UTC; this repo carries two distinct offsets.

CA-10. A response in a commit whose subject carries no T## tag is attributed to UNATTRIBUTED, which is reported as its own line in every table. On CB-WP-0001 this is 33% of cost ($30.32 of $92.21) — a per-task table is a view over roughly two-thirds of the money and says so wherever it appears.

CA-11. Cost after the last commit is an open remainder, reported separately from UNATTRIBUTED. It is work not yet committed, not work without a task.

CA-12. Attribution never keys on wall-clock alone. Two sessions overlapped 4 h 13 m on this repo carrying ~$12; only sessionId separates them.

3. Reported shape

CA-13. Every report carries composition — the split across input, output, cache read, cache write 5m, cache write 1h — alongside the total. A total alone would have concealed the finding that motivated this work: 88.4% of spend is cache, at 256:1 cache-read to output tokens.

CA-14. Reconciliation is asserted. Attributed + unattributed + open remainder + unpriced must equal the transcript total to the cent. A mismatch aborts rather than reporting.

CA-15. Tables that a tool can emit are emitted by the tool. Every committed number in this capability's evidence is produced by a command, not typed. (Adversarial review of CB-RES-0002 found every computed figure correct to the cent and three hand-typed markdown tables wrong.)

4. Acceptance metrics

Each row names the command that produces its number, per InnerLoop §Step 4. cb-cost is tools/cb-cost (T04).

ID Metric Target Instrument
AC-1 reproduces the pinned CB-WP-0001 total $93.32 = $92.21 main + $1.11 subagent make cost-pin
AC-2 reconciliation residual (CA-14) $0.00 exactly same command, reconciled: ok line
AC-3 unattributed share reported (CA-10) present, and 33% on the pinned run cb-cost --pin fc76445 --by-task
AC-4 composition reported (CA-13) all five components present cb-cost --pin fc76445 --composition
AC-5 dedup invariant asserted (CA-02) violation exits non-zero make cost-test
AC-6 positive control: refuses to report on zero responses exits non-zero make cost-test
AC-7 subagent tree included (CA-06) omitting it changes AC-1 by $1.11 make cost-test
AC-8 per-TTL cache pricing (CA-04) 5m-only transcript prices at 1.25× make cost-test
AC-9 streamed partial output resolves to final (CA-02a) 5,5,195 → 195, not 5 make cost-test

AC-5 through AC-8 are the positive control. Per InnerLoop v1.0 §Step 5, a harness must assert it did the work it reports. --self-test runs each assertion against a fixture whose expected value is known and fails loudly; make cost runs it before any reported number.

5. Metric feasibility check

Per InnerLoop §Step 4, the acceptance table is checked against the contracts in this same spec:

  • AC-1's $92.87 is reachable only if CA-06 holds (both trees enumerated). Under a main-file-only collector the target is unreachable — this is the defect the adversarial review caught, where a target of $92.21 would have been hit only by a broken collector.
  • AC-3's 33% is a property of CB-WP-0001's commit subjects, not of the collector. It is a regression pin on the fixture, not a quality target; improving tagging discipline will change it, and that is expected.
  • CA-07 (pinning) makes AC-1 reproducible; without it the target drifts upward on every run and the test is meaningless.

6. Known limitations, stated with the numbers

  • Per-task cost covers ~67% of spend. Structural: 19 of 33 commits carry no task tag. Reported per CA-10, never silently dropped.
  • Work spanning a commit is assigned whole to the later task. Bounded by one interval: p50 6.9 min, p90 17.7 min, max 36.8 min.
  • The price sheet cannot express a time-boxed rate. Sonnet's intro price is a TOML comment, so the sheet is silently wrong for sonnet-priced work ($0.17 at the pin, 0.19%). This becomes an error, not a rounding issue, on 2026-08-31 when the intro rate expires and the comment and the data disagree in the other direction. Tracked as a schema defect against the price sheet.
  • The State Hub cannot store what this spec measures. Its token event schema has tokens_in/tokens_out and no cache fields, so the dashboard necessarily shows a lossy projection. This is a limitation of the sink, not of the metric; make cost remains the authority.

7. Revisions to M-D2-CST

specs/MetricsAndScenarios.md §1a is superseded by this spec. M-D2-CST is redefined from "tokens × pricepoint" — which named no instrument and was never computed — to: USD per completed workplan task, per CA-08 attribution, produced by make cost. M-D2-TOK is retained but demoted: tokens are the input to the cost model, not a comparable figure across models or across cache states.