clay-borg/research/CB-RES-0002-cost-accounting.md
tegwick 2f086d26b6 T06: wire cost into the loop
- InnerLoop definition-of-done gains a cost row: M-D2-CST may no longer
  be recorded uncomputable, and composition must be reported, not only a
  total.
- InnerLoop Step 5 gains the --self-test contract: every tool that
  reports a number exposes one, and it runs before the number does.
  Rationale attached, because the case that motivated it is the one
  review cannot catch — survey and reviewer both verified the same large
  sample and both missed the small one.
- make cost / cost-test / cost-pin on the one command surface; cost-test
  in `make all` and in CI.
- Hub now holds the measured figure for CB-WP-0001: 80.6M in / 323.6k
  out against the 401,100 it previously estimated, low by ~200x. The
  event states plainly that the hub schema cannot represent the 88% of
  cost that is cache, and names `make cost-pin` as the authority.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 08:52:33 +02:00

333 lines
18 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# CB-RES-0002: agentic cost accounting
capability: meta.loop.cost-accounting
status: approved # adversarial review 2026-07-31: 16 findings, 15 conceded
tier: L (structural L, chaos d10=2 → no override)
runnable-baseline: invoked — every candidate below was exercised against the
CB-WP-0001 session on this machine, not cited
review-trail: history/260731-cost-accounting-{research,challenge,response}.md
Survey of instruments that can attribute the USD cost of agentic work to a
unit of work, so that M-D2-CST (`specs/MetricsAndScenarios.md` §1a) becomes
computable. CB-WP-0001 specified that metric completely and recorded it as
*uncomputable*; the premise of this workplan is that the data existed the
whole time.
That premise survives. The workplan's **numbers do not** — see §Correction.
---
## Correction to this workplan's own Purpose section
CB-WP-0002's Purpose reports the CB-WP-0001 session at **$248.46**, from
131,863,164 cache-read tokens priced at Fable 5. Both halves are wrong, and
in the same direction — too high. The survey found this by re-deriving the
number rather than adopting it.
**Error 1 — per-line summation double-counts.** A single API response is
written to the transcript as *several* JSONL lines, split by content block
(`thinking`, `text`, `tool_use`), and **every one of those lines repeats the
complete `usage` object**. Measured on the CB-WP-0001 transcript: 657
assistant lines carry only 346 distinct `requestId`s. Group sizes run 16:
| lines per requestId | 1 | 2 | 3 | 4 | 5 | 6 | total |
|---|---|---|---|---|---|---|---|
| groups | 140 | 120 | 74 | 6 | 5 | 1 | **346** |
| lines | 140 | 240 | 222 | 24 | 25 | 6 | **657** |
Both rows are checksummed against the file: groups sum to 346, lines to 657.
Positive control on the dedup: across all 206 multi-line groups the `usage`
object is byte-identical (206/206 identical, 0 differing), and no group
mixes models. Every assistant line carries a `requestId` — there is no null
key for `setdefault` to collapse. The duplication is a transcript-format
artifact, not repeated billing. Summing per line inflates by ≈1.9×.
**Error 2 — single-model pricing on a multi-model session.** The session ran
three models, not one. Both columns are shown because the gap between them
*is* error 1:
| model | JSONL lines | API responses (deduped) | pinned cost |
|---|---|---|---|
| claude-opus-5 | 382 | 213 | $36.19 (39%) |
| claude-fable-5 | 250 | 118 | **$55.50 (60%)** |
| claude-sonnet-5 | 24 | 14 | $0.52 (0.6%) |
| `<synthetic>` | 1 | 1 | $0.00 |
| **total** | **657** | **346** | |
§1a *already required* per-model pricing; the Purpose section did not apply
its own rule. Note the inversion, which matters more than the correction:
**Fable is 35% of the calls and 60% of the dollars; Opus is 60% of the calls
and 39% of the dollars.** A count-majority is not a cost-majority — the
error this document is about, one level down.
The `<synthetic>` entry is an error placeholder carrying a valid `requestId`
and a complete **all-zero** `usage` object, not a missing one. Its dollar
impact is exactly $0.00, but a collector guarding on `if not usage` and one
guarding on `if model not in prices` take different branches; T04 must name
which.
**Corrected totals** for the same transcript, all three methods run over
the identical unpinned line set so the methods are comparable:
| method | responses | output | cache read | cost |
|---|---|---|---|---|
| per-line, all-Fable (the Purpose method) | 657 | 700,690 | 161,408,840 | $289.12 |
| per-line, per-model | 657 | 700,690 | 161,408,840 | $210.05 |
| **deduped, per-model (correct)** | **346** | **318,230** | **81,100,498** | **$93.15** |
**Third error, found while verifying the second: the transcript is a live
file.** Re-running the deduped figure minutes later returned 356
responses and $94.04 — this survey's own session appends to the same
JSONL it is measuring. An unpinned total is not a repeatable number. The
acceptance target is therefore pinned by timestamp:
| CB-WP-0001, pinned ≤ `2026-07-31T02:17:59Z` (commit `fc76445`) | value |
|---|---|
| responses | 339 (206 opus-5, 118 fable-5, 14 sonnet-5, 1 synthetic) |
| output | 313,900 tok → $10.66 |
| cache read | 80,453,702 tok → $59.59 |
| cache write 1h | 1,672,854 tok → $21.95 |
| input | 676 tok → $0.00 |
| main transcript | **$92.21** — 88.4% cache, 256:1 cache-read:output |
| + subagent tree (7 responses, ran 23:1123:14Z, inside the pin) | $1.11 |
| **TRUE TOTAL** | **$93.32** |
*(The subagent figure was $0.66 when this survey was written. T04's
positive control found the cause: `output_tokens` is a running count in
the `subagents/` tree, so first-wins dedup under-counted it. See
`specs/CostAccounting.md` CA-02a.)*
The subagent line is not a footnote. C1's blind-spot finding says a
collector reading only the main file under-reports; a target of $92.21 would
have been hit only by a collector *with* that blind spot, and failed by a
correct one. The acceptance target is **$93.32, stated as its two
components**, so a collector reading one tree is diagnosed rather than
merely failed.
The reported figure was **~2.7× the real cost**. This is the fourth
instance of the harness-does-nothing error class from
`history/260731-inner-loop-retrospective.md`, wearing a new coat: not a
harness that measured nothing, but an arithmetic that measured the same
thing twice. Both produce a number that looks fine.
The qualitative headline survives the correction and gets stronger: cache
reads are **81.1M tokens against 318k of output**, ~255:1. Cost in an
agentic loop is context × turns.
---
## Candidates
### C1 — Session transcript JSONL
`~/.claude/projects/<slug>/<session>.jsonl`, one JSON object per line.
Assistant lines carry `message.usage` with exact billing counters:
`input_tokens`, `output_tokens`, `cache_read_input_tokens`, and
`cache_creation.{ephemeral_1h,ephemeral_5m}_input_tokens`, plus
`message.model`, `requestId`, and an ISO-8601 `timestamp`.
- **Granularity:** per API response, once deduplicated by `requestId`.
- **Accuracy:** exact — these are the counters the invoice is computed from.
There is no sampling or rounding.
- **Verified non-issue:** `usage.iterations[]` is a sub-breakdown, not an
additional charge. Every assistant line carries `usage`; the iteration
outputs sum exactly to the top-level `output_tokens` in every case, and
no message had more than one iteration. Summing `iterations` *instead of*
the top-level fields is safe; summing *both* would double-count.
- **Cache-write rates are per-TTL and must not be aggregated.**
`cache_creation_input_tokens` equals `ephemeral_5m + ephemeral_1h` in all
775 usage lines across both transcripts, and the two bill at different
multipliers (1.25 vs 2.0). Pricing the top-level aggregate at a single
rate is a silent error: the subagent's writes are **entirely 5m** (37,467
tokens), and pricing them at 1h inflates that transcript by **+43%**
($0.658 → $0.939). The main session happens to be all-1h, so the pinned
total is insensitive — but fan-out passes are exactly where 5m dominates.
- **Survives compaction:** yes. `/compact` writes a summary message into the
same file (`isCompactSummary`, `compactMetadata`) and the session
continues; no usage is lost. Compaction is visible as an event, so its
cost is itself measurable.
- **Attribution:** none built in — a transcript is a flat message stream
with timestamps. It must be joined against an external time index.
- **Blind spot 1 — subagents are a separate tree.** Subagent cost is **not**
in the main transcript. `isSidechain` is `false` on all 657 lines;
subagent work lives in `<session>/subagents/agent-*.jsonl` with an
`agent-*.meta.json` naming the agentType and model. CB-WP-0001 spawned one
(the adversarial review). A collector reading only the main file silently
under-reports.
- **Blind spot 2 — one repo, several transcripts, overlapping in time.**
The project directory holds more than one session. `f1eb1147` overlaps
`8cbd5701` for **4 h 13 m**, carrying ~$6 on each side — ~$12 that no
wall-clock join can separate, beginning 14 seconds after the pin. A
collector must therefore key attribution on **`sessionId`, not only time**,
and must enumerate every transcript for the repo rather than one file.
### C2 — Custodian State Hub token API
`record_token_event`, the `update_task_status` token tiers, and
`get_token_summary`. Exercised against CB-WP-0001's workplan
(`a1b434dc-…`), which returned:
```text
tokens_in 362,000 tokens_out 39,100 event_count 7 by model: claude-fable-5
```
- **Granularity:** per task — the best of any candidate, and the only one
that is natively *about* the unit of work.
- **Accuracy:** poor, and structurally so. Three independent defects:
1. **The schema has no cache fields.** `tokens_in`/`tokens_out` cannot
represent the finding this workplan exists to report. Cache read alone
is **64.6% of cost** and all cache is **88.4%**; the hub cannot express
either at any fidelity.
2. **The recorded numbers are estimates.** 7 events for 9 tasks, at
round figures — the skill's Tier-3 heuristic (1000/500) and Tier-1
eyeball estimates. Against a deduped transcript output of 318,230,
the hub's 39,100 is off by ~8×; against total input it is off by
~450×.
3. **Model attribution is wrong.** Everything is filed under
`claude-fable-5` on a session that was majority Opus 5.
- **Survives compaction:** yes — it is server-side and independent of the
client.
- **Verdict:** durable and task-shaped, but its numbers are unusable as a
cost source. Its role is as a **sink** for numbers computed elsewhere,
not a source. Even as a sink it can only carry a lossy projection until
the schema grows cache fields.
### C3 — Claude Code status bar
- **Granularity:** whole session, live.
- **Accuracy:** unknown and unauditable — it is rendered text.
- **Machine-readable:** no. Not configured here (`statusLine` is absent
from `~/.claude/settings.json`), and it is not reachable from inside a
tool call regardless.
- **Verdict:** eliminated. The ralph-workplan skill's "read tokens from the
status bar" instruction is the proximate cause of C2's bad numbers — it
asks an agent to report a figure it cannot read, and an agent that cannot
read it estimates instead. This should be raised against the skill.
### C4 — Anthropic usage / billing API
- **Granularity:** organization and API-key, by day.
- **Accuracy:** authoritative — it *is* the invoice.
- **Attribution:** none to a task, and none to a session. Cannot separate
clay-borg from the other twenty-plus projects on this machine.
- **Availability:** requires an admin key; none is configured here.
- **Verdict:** not usable for M-D2-CST, but valuable as an **external
reconciliation check** if an admin key is ever provisioned — it is the
only candidate that can catch a systematic error in C1's price model.
Left as a stated non-dependency.
### C5 — Git commit history (attribution index, not a cost source)
Not a cost instrument; the missing half of C1. The loop already commits per
task iteration with the task in the subject line, giving durable, timestamped
boundaries at exactly the granularity M-D2-CST wants:
```text
a09d76f 2026-07-31T02:14:34+02:00 T08 iter 1: scenario runner executes; …
b58a913 2026-07-31T02:19:34+02:00 T08 iter 2: Reveal, Resolve, End; …
```
- **Interval convention:** *ending-at-commit*, `(prev_commit, this_commit]`.
This is the only convention consistent with commit-at-end-of-task; under
the alternative every task's cost shifts one interval.
- Spacing over the 33 pinned commits: min 0.1, p50 **6.9**, p90 17.7, max
**36.8** minutes. Usually finer than a task; the 37-minute max is not.
- Durable, versioned, and free; requires no change to how work is done.
- **Coverage is the real limit, and it is sized:** only **14 of 33** commits
name a task (`T##`) in the subject. The other 19 — `chore(consistency)`,
`CI:`, `AM-4:`, workplan additions — hold **$30.32 of $92.21, or 33% of
cost**. Attribution to a *task* therefore covers two-thirds of spend at
best; the remainder is real work that must be reported as its own line,
not discarded.
- **Boundary conditions verified clean:** 0 responses before the first
commit, 0 after the last, exactly 1 empty interval (`Initial commit`).
- **Known hazards:** offsets are not uniform — `git log` shows `+02:00` ×34
and `+00:00` ×1, so a collector must parse `%cI` and convert, never
subtract a fixed offset. Commits made outside a session create empty
intervals.
---
## Baselines (benchmark-to-beat)
| Dimension | Baseline holder | Metric | Value | Provenance |
|---|---|---|---|---|
| D1 ease of specification | C2 hub | fields needed to record a task's cost | 4 (`task_id`, `tokens_in`, `tokens_out`, `model`) — but cannot express cache | measured (API schema) |
| D2 efficiency | C2 hub | cost of producing a number | ~0 (one API call) — number is an estimate, off by ~8× on output | measured |
| D2 efficiency | C1 transcript | cost of producing a number | one file read, 5.1 MB, ~1 s; exact | measured |
| D3 speed | C1 transcript | parse of a full session | 2,040 lines / 5.1 MB in <1 s in CPython | measured |
| D2 efficiency (**the deciding row**) | C1 transcript | **error against the billing counters** | **$0.00 the transcript *is* the counter set; C2's error on the same work is $92.21 $0.03 recorded 100%** | measured |
| D4 optionality | C4 billing API | authority | invoice-grade, zero attribution, admin key absent (asserted, not measured used only to keep C4 as an optional check) | availability measured |
**Benchmark-to-beat for the collector:** reproduce **$93.32** for repo
`clay-borg` pinned at `2026-07-31T02:17:59Z` as its two components,
**$92.21 main transcript + $1.11 subagent tree** from the committed price
sheet, with the unattributed remainder reported as its own line (expected
**33%**, $30.32) and reconciliation asserted rather than assumed.
---
## Verdict
**C1 (transcript) leads on accuracy and is the only exact candidate.**
**C5 (git commits) supplies the attribution index C1 lacks.** C2 is the
durable sink. C3 is eliminated. C4 is an optional external check.
The expected shape is therefore: enumerate **every** transcript for the repo
including the `subagents/` tree dedup by `requestId` price per message at
its own model's rate and **per cache TTL** from
`benchmarks/baselines/model-prices.toml` attribute to a task by joining
`(prev_commit, this_commit]` intervals **within a `sessionId`** emit
per-task cost, an unattributed-remainder line, and a composition breakdown
push a lossy summary to C2.
**What none of them do well — the surpass opportunity.** Every candidate
reports *totals*. None reports **composition**, and composition is where
the actionable finding lives: 81.1M cache-read tokens against 318k of
output means cost is driven by how much context is re-read per turn, which
no total can show. A metric that had reported only dollars would have been
correct and useless.
**Risks in the baselines themselves.**
1. **The $248.46 figure was wrong and was nearly adopted as this
workplan's acceptance target.** T05's reconciliation test must be
against a number this survey re-derived, not against the Purpose
section. The Purpose section needs correcting.
2. **Dedup is load-bearing.** If the transcript format ever splits one
response across two `requestId`s, dedup silently under-reports
the opposite error, and the more dangerous one. The collector must
assert its dedup assumption (identical usage within a group) at
runtime rather than trusting this survey's one-time check.
3. **Subagent transcripts are a separate tree.** Measured: CB-WP-0001's
one subagent (adversarial review, Fable 5, 7 responses, 158,096 cache
reads) cost **$1.11**, invisible to any collector reading only the
main file. Small here; not small for a pass that fans out.
4. **The price sheet has a 90-day staleness rule** 1a) and no automated
check. Every number this capability produces inherits that.
5. **The transcript is append-live.** It is written by the session that
reads it, so any total is a reading at an instant. Every committed
number from this capability states its pin (timestamp or commit), and
the collector takes a pin argument rather than defaulting to "all".
6. **Concurrent sessions are not hypothetical — they are in this data.**
`f1eb1147` and `8cbd5701` overlap for 4 h 13 m on the same repo, ~$12
inseparable by wall-clock. (The session that wrote the first draft of
this sentence *was* the overlap.) Attribution keys on `sessionId` first,
time second. Compaction inside a task is safe it stays in one file and
one session.
7. **A third of spend has no task, structurally.** 19 of 33 commits carry no
task tag; $30.32 of $92.21. Any per-task cost table is a view over ~two
thirds of the money, and must say so wherever it is reported the same
limit-with-the-number rule the coverage gate carries.
8. **The price sheet cannot express a time-boxed rate.** Sonnet's intro
price lives in a TOML *comment*, so a collector reading the sheet
silently uses the wrong number ($0.17 at the pin, 0.19%). Small now, and
the same class of defect this survey levels at C2: a schema that cannot
hold the fact it needs. Raised for T03.
9. **Hand-typed tables are the actual failure surface.** Every computed
figure in this survey reproduced to the cent under adversarial
re-derivation; three hand-written markdown tables did not (a group-size
row that failed its own checksum, a per-line count labelled as deduped,
a stale 87%). Numbers a tool can emit should be emitted by the tool.
Raised for T03 and T06.