Duplicated-fact drift is the fourth error class and the only one with no executable rule. No positive control catches it (both copies are internally consistent) and re-derivation does not either (the copy reproduces whatever it was copied from). It is caught only by reading a copy against its source, which nothing in the loop required. facts.toml holds 15 facts and is GENERATED by `make facts-gen` from cb-cost, dep-weight and rule-coverage. The trap this task named — a hand-maintained registry that becomes another drifting copy — is closed by facts-check re-running the instruments and failing when the committed registry disagrees with them. A stale registry cannot certify stale artifacts. An artifact quoting a fact tags it: **$93.15** <!-- fact --> with the key. 17 occurrences across 5 artifacts are now checked. Falsified before being believed: changing CostAccounting.md line 158 from $93.15 to $92.87 — the exact historical drift — produced exit 1 naming the file, the line and the expected value. Tested against the class it exists to catch, on a real artifact, not only in its self-test. It then caught a live tag inside its own documentation example in InnerLoop.md within the hour. Third time a gate has failed on its own pass's work. What it does not close is stated rather than implied: 22 untagged literal copies remain and are reported, not failed. Tagging is opt-in, a number can legitimately recur, and a gate that fires on coincidence gets routed around. Naming the uncovered surface beats claiming the class is closed. InnerLoop single-source-of-fact moves from prose to executable — v1.3. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
13 KiB
Cost Accounting
Status: v1.0 — 2026-07-31. Derived from CB-RES-0002 (approved) and ADR-0003. Makes M-D2-CST computable; supersedes its "uncomputable" disposition in CB-EV-0001.
Defines how the USD cost of agentic work is measured and attributed, so that D2 claims about implementation efficiency are falsifiable.
1. The cost model
1.1 Unit of billing
The unit is one API response, identified by requestId. It is not
one JSONL line: a response is written as up to six lines split by content
block (thinking, text, tool_use), and every line repeats the same
usage object.
CA-01. Cost is computed over responses deduplicated by
requestId. Summing per line is a defect; it inflates by ≈1.9× on measured data.
CA-02. Dedup is asserted, not assumed. Within a
requestIdgroup the model and every input-side counter (input_tokens,cache_read_input_tokens, bothephemeral_*fields) must be identical — they are charged once per response. A divergence aborts the run.
CA-02a.
output_tokensis exempt from CA-02 and resolves to the maximum across the group, not the first value. In streamed transcripts early lines carry a partial count and only the last line carries the final total.
Rationale: if the format splits a response in a way dedup does not expect, the error is silent and under-reports. The dangerous direction gets the assertion.
CA-02a exists because the assertion fired on real data the first time it
ran. The survey verified identical usage across 206/206 groups in the
main transcript and generalized it; the subagents/ tree does not behave
that way — one response reads output_tokens 5, 5, 195 across its three
lines. First-wins scored it at 5. That error moved the acceptance target
by $0.45.
1.2 Price formula
Per response, against benchmarks/baselines/model-prices.toml:
cost = input_tokens × price.input
+ output_tokens × price.output
+ cache_read_input_tokens × price.input × cache.read (0.10)
+ ephemeral_5m_input_tokens × price.input × cache.write_5m (1.25)
+ ephemeral_1h_input_tokens × price.input × cache.write_1h (2.00)
CA-03. Each response is priced at its own
message.modelrate. A session may mix models; CB-WP-0001 used three.
CA-04. Cache writes are priced per TTL. The top-level
cache_creation_input_tokensaggregate equalsephemeral_5m + ephemeral_1hand must never be priced at a single multiplier — doing so inflated a measured subagent transcript by 43%.
CA-05. A response whose model is absent from the price sheet is reported as an unpriced line with its token counts, never dropped and never priced at a default.
CA-16. Promotional and other time-boxed rates are data, not comments. A response is priced at the rate in force at its own timestamp:
promo_input/promo_outputapply on or beforepromo_until, and the standard rate applies after.
CA-17. The price sheet carries
recordedandmax_age_days, and a sheet older thanmax_age_daysaborts the run. A stale sheet invalidatesbetterverdicts, so the staleness rule is a command rather than a sentence.
1.3 Scope of a measurement
CA-06. A measurement enumerates every transcript for the repo:
~/.claude/projects/<slug>/*.jsonland every<session>/subagents/agent-*.jsonl. Subagent cost is not in the main file and is invisible to a collector that reads one path.
CA-07. Every committed number states its pin — a timestamp or commit. Transcripts are append-live: the file grows as the measuring session writes to it, and an unpinned total is not reproducible.
2. The attribution contract
CA-08. A response is attributed to the task named by the next commit at or after it, within its own session:
interval := (prev_commit_time, this_commit_time], scoped bysessionId.
CA-09. Timestamps are converted, never offset-subtracted. Commit times are parsed from
%cIand converted to UTC; this repo carries two distinct offsets.
CA-10. A response in a commit whose subject carries no
T##tag is attributed toUNATTRIBUTED, which is reported as its own line in every table. On CB-WP-0001 this is 33% of cost ($30.32 of $92.21) — a per-task table is a view over roughly two-thirds of the money and says so wherever it appears.
CA-11. Cost after the last commit is an open remainder, reported separately from
UNATTRIBUTED. It is work not yet committed, not work without a task.
CA-12. Attribution never keys on wall-clock alone. Two sessions overlapped 4 h 13 m on this repo carrying ~$12; only
sessionIdseparates them.
3. Reported shape
CA-13. Every report carries composition — the split across input, output, cache read, cache write 5m, cache write 1h — alongside the total. A total alone would have concealed the finding that motivated this work: 88.4% of spend is cache, at 256:1 cache-read to output tokens.
CA-14. Reconciliation is asserted. Attributed + unattributed + open remainder + unpriced must equal the transcript total to the cent. A mismatch aborts rather than reporting.
CA-15. Tables that a tool can emit are emitted by the tool. Every committed number in this capability's evidence is produced by a command, not typed. (Adversarial review of CB-RES-0002 found every computed figure correct to the cent and three hand-typed markdown tables wrong.)
4. Acceptance metrics
Each row names the command that produces its number, per InnerLoop §Step 4.
cb-cost is tools/cb-cost (T04).
| ID | Metric | Target | Instrument |
|---|---|---|---|
| AC-1 | reproduces the pinned CB-WP-0001 total | $93.15 = $92.03 main + $1.11 subagent | make cost-pin |
| AC-2 | reconciliation residual (CA-14) | $0.00 exactly | same command, reconciled: ok line |
| AC-3 | unattributed share reported (CA-10) | present, and 32.4% on the pinned run | cb-cost --pin fc76445 --by-task |
| AC-4 | composition reported (CA-13) | all five components present | cb-cost --pin fc76445 --composition |
| AC-5 | dedup invariant asserted (CA-02) | violation exits non-zero | make cost-test |
| AC-6 | positive control: refuses to report on zero responses | exits non-zero | make cost-test |
| AC-7 | subagent tree included (CA-06) | omitting it changes AC-1 by $1.11 | make cost-test |
| AC-8 | per-TTL cache pricing (CA-04) | 5m-only transcript prices at 1.25× | make cost-test |
| AC-9 | streamed partial output resolves to final (CA-02a) | 5,5,195 → 195, not 5 | make cost-test |
AC-5 through AC-8 are the positive control. Per InnerLoop v1.0 §Step 5,
a harness must assert it did the work it reports. --self-test runs each
assertion against a fixture whose expected value is known and fails loudly;
make cost runs it before any reported number.
5. Metric feasibility check
Per InnerLoop §Step 4, the acceptance table is checked against the contracts in this same spec:
- AC-1's $93.15 is reachable only if CA-06 holds (both trees enumerated). Under a main-file-only collector the target is unreachable — this is the defect the adversarial review caught, where a target of $92.21 would have been hit only by a broken collector.
- AC-3's 33% is a property of CB-WP-0001's commit subjects, not of the collector. It is a regression pin on the fixture, not a quality target; improving tagging discipline will change it, and that is expected.
- CA-07 (pinning) makes AC-1 reproducible; without it the target drifts upward on every run and the test is meaningless.
6. Known limitations, stated with the numbers
- Per-task cost covers ~67% of spend. Structural: 19 of 33 commits carry no task tag. Reported per CA-10, never silently dropped.
- Work spanning a commit is assigned whole to the later task. Bounded by one interval: p50 6.9 min, p90 17.7 min, max 36.8 min.
- The price sheet cannot express a time-boxed rate. Sonnet's intro price is a TOML comment, so the sheet is silently wrong for sonnet-priced work ($0.17 at the pin, 0.19%). This becomes an error, not a rounding issue, on 2026-08-31 when the intro rate expires and the comment and the data disagree in the other direction. Tracked as a schema defect against the price sheet.
- The State Hub cannot store what this spec measures. Its token event
schema has
tokens_in/tokens_outand no cache fields, so the dashboard necessarily shows a lossy projection. This is a limitation of the sink, not of the metric;make costremains the authority.
7. The live cost budget
Replaces the 8k/10k per-task token budget, which was never referenced or enforced, which CB-WP-0001 T08 exceeded by orders of magnitude with no signal, and which InnerLoop v1.1 now marks dead. It implied a control that did not exist.
The design constraint that shapes this. Per-task cost needs the commit that closes the task (CA-08), so it is unavoidably retrospective — a per-task budget can only ever report a breach after the money is spent. What is observable mid-task is spend since the last commit, because the transcript is an append-live file (CA-07). That is the only number a budget can actually fire on, so that is what the budget binds.
CB-01. The budget is the open remainder: USD spent since the last commit, computable at any instant by
make cost-budget.
CB-02. Thresholds: soft $10.00, hard $22.00.
- soft — state progress as a percentage and decide: continue, or commit what works and decompose.
- hard — stop. Commit what works, or decompose. Uncommitted work is also unattributable work, so a hard breach costs twice: once in money, once in the 33% attribution gap.
Calibration, measured over the 32 non-empty commit intervals of
CB-WP-0001 (cb-cost --pin fc76445):
| spend per interval | |
|---|---|
| min | $0.07 |
| p50 | $1.40 |
| p90 | $9.36 |
| max | $10.80 (T07, the Cargo workspace scaffold) |
Soft is set at $10.00, just below the observed maximum and just above p90: on the pass this is calibrated from it would have fired exactly once, on the single most expensive interval. Hard is set at $22.00, roughly 2× the observed maximum — a value never reached in 32 intervals, so reaching it means the session is doing something the calibration data has no example of, which is precisely when stopping is right.
Why these are not retargeted-after-the-fact. T07's concern is a target moved to accommodate a measurement. These thresholds are set above every observed value, so they bind on future work rather than ratifying present work — no interval in the calibration set breaches hard, and one breaches soft. If a future pass routinely breaches, that is signal, not a reason to raise them.
Known limit: the budget cannot fire between turns without something
running the command. It is available to an agent at any point, in CI, and
in make all; it is not a daemon. An agent that never runs it gets no
signal — which is exactly the dead-policy failure this replaces, one level
up. The mitigation is that it costs nothing to run and is on the one
command surface.
9. The price sheet
benchmarks/baselines/model-prices.toml is the single source of rates.
Two defects were fixed here, both of the same shape — a schema that
could not hold the fact it needed, which is the criticism this
capability's survey levelled at the State Hub:
- A time-boxed rate lived in a comment. Sonnet's intro price
(2.00/10.00 through 2026-08-31) was written as
# intro …and was therefore invisible to the collector, which used the standard rate. Cost at the pin: $0.17, 0.19% of the pass. The real hazard was the date — on 2026-08-31 the promotion expires and a comment-only representation would have been wrong in the opposite direction, with nothing to notice. Now expressed aspromo_*fields and applied per CA-16. - The 90-day staleness rule had no check. It was prose in
MetricsAndScenarios.md§1a and every M-D2-CST verdict silently inherited it. Nowrecorded+max_age_daysin the sheet, enforced by CA-17 and exercised bymake cost-test.
Applying CA-16 moved AC-1 from $93.32 to $93.15. This is a legitimate retarget under the distinction CB-WP-0003 T07 draws: the instrument disproved the target, rather than the implementation missing it. The instrument's output is in the same commit.
8. Revisions to M-D2-CST
specs/MetricsAndScenarios.md §1a is superseded by this spec. M-D2-CST is
redefined from "tokens × pricepoint" — which named no instrument and was
never computed — to: USD per completed workplan task, per CA-08
attribution, produced by make cost. M-D2-TOK is retained but demoted:
tokens are the input to the cost model, not a comparable figure across
models or across cache states.