clay-borg/history/260731-cost-accounting-research.md
tegwick 060eb8cf6a T01: cost-accounting survey — and a 2.7x correction to our own numbers
Tier L (structural L, chaos d10=2). Every candidate exercised locally
against the CB-WP-0001 transcript rather than cited.

Verdict: session transcript JSONL is the only exact instrument; git
commits supply the attribution index it lacks; the State Hub is a
durable sink whose schema cannot express cache tokens at all; the
status bar and billing API are eliminated on availability.

Three errors found in this workplan's own Purpose section, all
overstating cost:

  - one API response is written as up to six JSONL lines, each
    repeating the same usage object (657 lines, 346 responses)
  - a multi-model session was priced entirely at Fable 5 rates,
    violating a per-model rule MetricsAndScenarios §1a already had
  - the transcript is a live file, so an unpinned total is not
    repeatable (346 -> 356 responses between two runs minutes apart)

CB-WP-0001 pinned at fc76445: $92.21, not $248.46. The qualitative
finding sharpens rather than weakens — 88.4% of spend is cache, at
256:1 cache-read to output tokens.

Neither error was of the harness-does-nothing class the positive
control was written for; both sums ran over real data. What caught
them was re-deriving a number instead of quoting it. Raised for T07
and for CB-WP-0003 T01.

Also measured: the one subagent (adversarial review) cost $0.66 and is
invisible to a collector reading only the main transcript.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 04:27:52 +02:00

3.8 KiB
Raw Blame History

2026-07-31 — how the cost-accounting survey was conducted

Trail for research/CB-RES-0002-cost-accounting.md (CB-WP-0002 T01). Tier L, chaos d10 = 2 (no override).

Method

No web sources. Every candidate was exercised locally against the CB-WP-0001 session transcript (~/.claude/projects/-home-worsch-clay-borg/8cbd5701-….jsonl, 5.1 MB, 2,040 lines) and the clay-borg git history. The survey's own rule — measured beats cited — applied to itself, which is the only reason the errors below were found.

What was measured, in order

  1. Transcript shape. Counted line types and collected the union of usage keys. Found 654 usage-bearing assistant lines, and four distinct message.model values where one was assumed.
  2. usage.iterations. Suspected as a double-count source. Checked whether iteration outputs sum to the top-level output_tokens: they match in all 654 cases, and no message had >1 iteration. Cleared.
  3. requestId cardinality. 346 distinct ids across 657 lines. This was the finding that overturned the workplan's numbers.
  4. Dedup validation. For all 208 multi-line groups, compared the serialized usage object across the group: 208 identical, 0 differing, 0 mixed-model. Block-type patterns confirmed the split is thinking / text / tool_use, i.e. a transcript-writer artifact.
  5. Re-priced three ways (per-line all-Fable, per-line per-model, deduped per-model) to isolate how much of the error came from each mistake: $289.12 / $210.05 / $93.15.
  6. Hub API. Called get_token_summary on CB-WP-0001's workplan. 401,100 tokens over 7 events for 9 tasks, all filed under Fable 5. Compared against the transcript to quantify the gap (~8× on output).
  7. Status bar and billing API. Checked ~/.claude/settings.json for statusLine (absent) and for an admin key (absent). Both eliminated on availability before any further evaluation.
  8. Git boundaries. Read git log with ISO timestamps to confirm commit density and that subjects name tasks. Noted the UTC/+02:00 mismatch against transcript timestamps.
  9. Subagent tree. Found <session>/subagents/ while checking whether isSidechain was ever true (it never is). Priced the one subagent transcript: $0.66.

Dead ends

  • isSidechain as the subagent signal. It is present on every line and false on every line in this session; it does not mark subagent work in the main file because subagent work is not in the main file. Looking for a flag wasted a step that a directory listing answered.
  • tool-results/ sidecar directory (340 KB) was inspected as a possible cost source. It holds raw tool outputs for replay, carries no usage data, and is not billed separately. Not a candidate.

The finding worth carrying forward

The workplan that exists to make cost measurable opened with a mismeasured cost, off by ~2.7×, and its two errors were of different kinds: a format artifact (double-counted lines) and a rule that existed but was not applied (per-model pricing, already mandated by specs/MetricsAndScenarios.md §1a).

The second is the more uncomfortable one. InnerLoop v1.0's answer to bad numbers is the positive control, and a positive control would not have caught either of these — both sums were over real data that really existed. What caught them was re-deriving a number instead of quoting it. That belongs in the loop as its own rule, and is raised for T07: a number inherited from a previous pass is re-derived before it is used as a target, or it is cited as unverified.

Reported into CB-WP-0003 T01's rule-enforceability audit as well: this is a fifth error instance, and it is not of the harness-does-nothing class that T10 predicted. The class is "trusted arithmetic over real data".