clay-borg/history/260731-cost-accounting-retrospective.md
tegwick 42180fbc57
Some checks failed
ci / check (push) Failing after 3s
T07: InnerLoop v1.1 — the instrument must emit its own target
Answers the question the workplan posed: did v1.0's "every metric names
its instrument" rule stop a metric being written without a working
instrument? No. AC-1 named `cb-cost --pin` before that tool existed and
set a hand-computed target of $92.87; the tool returned $93.32. The rule
was satisfied completely and the metric was still wrong.

v1.1 adds:
  - the instrument must exist and the target must come out of it;
    targets are provisional until the tool emits them
  - a number inherited from earlier work is re-derived before use as a
    target, or cited as unverified
  - every reporting tool exposes --self-test, run before the number
  - cost is in the definition of done; M-D2-CST may not be uncomputable

The cost of CB-WP-0001 was stated four times before it was right --
$248.46, $92.21, $92.87, $93.32 -- and each correction came from a
different mechanism: re-derivation, adversarial review, and the positive
control. None found more than one. That is the case for keeping all
three.

CB-WP-0003 T10 predicted the next error would be harness-does-nothing.
It was not, twice. Trusted arithmetic over real data would pass a
positive control; and a property verified on 206/206 groups of the main
transcript is false in the 8-response subagent tree that neither the
survey nor the reviewer examined separately. Review structurally cannot
catch the second -- re-deriving on the same sample reproduces the same
blind spot.

Workplan status: done.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 08:53:52 +02:00

112 lines
5.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# 2026-07-31 — retrospective: what CB-WP-0002 taught the loop
Pass: CB-WP-0002, cost accounting. Produced `tools/cb-cost.py`,
`specs/CostAccounting.md`, ADR-0003, CB-EV-0002, and InnerLoop **v1.1**.
## The question the workplan asked
> CB-WP-0001 produced a fully specified metric that could not be computed,
> and v1.0's "every metric names its instrument" rule was written to stop
> that. Did it?
**No.** The rule was satisfied completely and the metric was still wrong.
`specs/CostAccounting.md` AC-1 named `cb-cost --pin fc76445` as its
instrument — a command that did not exist when the row was written — and
set the target to **$92.87**, computed by hand. When the tool was built one
task later it returned **$93.32**. The hand computation carried a dedup bug
that the tool's own positive control caught on first contact.
So the v1.0 rule stops "a metric with no instrument named" and does not stop
"a metric whose target the instrument never produced". The distinction did
not exist when the rule was written because the failure it was written for
was the coarser one.
**v1.1's answer:** the instrument must exist and the target must come out of
it. Where the tool is built later in the pass, the target is `provisional:`
until the tool emits it, and the spec is amended to whatever it returns —
not the reverse.
## The number, and its four corrections
The cost of CB-WP-0001 was stated four times before it was right:
| value | method | what was wrong |
|---|---|---|
| $248.46 | inherited, quoted | per-line summation (~1.9×) + single-model pricing |
| $92.21 | re-derived, deduped, per-model | omitted the subagent tree |
| $92.87 | + subagent at first-wins dedup | `output_tokens` is a running count |
| **$93.32** | emitted by `cb-cost` | — |
Each correction was found by a different mechanism, and this is the useful
part:
1. **$248.46 → $92.21** by *re-deriving instead of quoting*. Now a v1.1
rule.
2. **$92.21 → $92.87** by *adversarial review*, which noticed the target
contradicted the survey's own blind-spot finding. Cost: $1.11.
3. **$92.87 → $93.32** by the *positive control*, which fired on real data
and refused to print a number.
No single mechanism found more than one. That is the argument for keeping
all three rather than consolidating.
## The error class the last pass predicted, and what actually happened
CB-WP-0003 T10 predicted the next error would be the harness-does-nothing
class, since four instances had already been seen. It was not — twice.
- The two errors in $248.46 were **trusted arithmetic over real data**. Both
sums ran over data that genuinely existed. A positive control asserting
"did this harness do work?" would have answered yes, correctly, and
reported a wrong number.
- The error in $92.87 was **a property verified on the large sample and
assumed on the small one**. The dedup invariant was checked on the main
transcript (206/206 groups) by the survey, and independently re-checked by
the adversarial reviewer, who also used the main transcript. It is false
in the 8-response subagent tree that neither examined separately.
The second is the one worth carrying forward, because **review structurally
cannot catch it**: the reviewer's job is to re-derive the author's claims,
and re-deriving on the same sample reproduces the same blind spot. Only an
assertion running over *all* the data at execution time catches it. That is
now a v1.1 rule (`--self-test` on every reporting tool, run before the
number).
## What the pass bought, in its own units
The capability measures itself, which is the first time this project has
been able to say what a pass cost while the pass was running:
- **CB-WP-0001 (measured retroactively):** $93.32 pinned at `fc76445`.
88.0% cache, 11.9% output, 249:1 context re-read to text written.
- **Adversarial review, this pass:** $1.11 — ~1% of the pass it reviewed —
and it found three approval-blocking defects. Second consecutive pass
where that trade was decisive. Two data points now support CB-WP-0003 T03.
- **`/compact`, measured for the first time:** 542,991 → 19,974 tokens, a
27× context reduction. Cost per turn fell 3.1×, from $0.457 across the
136 turns before it to $0.149 across the 202 after.
That last number changes advice this project was about to give itself.
CB-WP-0003 T04 was going to prescribe one task per session on the theory
that long sessions are quadratic. They are **bounded**-quadratic: cost grows
with context between compactions and resets at each one. The failure mode is
a long *uncompacted* session, and whether a fresh session beats a compaction
is now a measurable question rather than a matter of taste. T04 should
measure it before prescribing.
## Raised, not resolved
- **The price sheet cannot express a time-boxed rate.** Sonnet's intro price
is a TOML comment. Costs $0.17 today (0.19%); becomes a real error on
**2026-08-31** when the intro rate expires and the comment and the data
disagree in the other direction.
- **32.5% of spend has no task**, because 19 of 33 commits carry no `T##`
tag. Reported as its own line rather than hidden, but the underlying fix
is commit hygiene, not tooling.
- **Outside this repo:** the `ralph-workplan` skill instructs agents to read
token counts from the Claude Code status bar. That is not readable from a
tool call, so an agent asked for it estimates instead — the proximate
cause of the hub holding 401,100 tokens for a workplan that actually
consumed 80.9M. Flagged for the maintainer; `make cost` is the authority
in the meantime.