The metric six evidence files reported as 0.0% was never measured. --shape-budget windows since the last commit, and it is read while writing the evidence file — right after a commit, when the window holds one or two responses. SH-1/SH-2 are location statistics and survive n=2. SH-3 is a rate: at n=2 its only possible values are 0%, 50%, 100%. cb-cost.py gains sh3_line(), which below a minimum sample prints "insufficient sample" and no verdict. The floor is derived: if the true rate were exactly the 20% target, P(zero batched in n) = 0.8^n, and at n=14 that is 4.4% — so "0 batched in 14" rules out a target-meeting rate at ~95%. Below that the tool has nothing to say and now says so. The window was kept rather than split. SH-3 could have been given a per-pass window, but the budget's purpose is the open remainder since the last commit, and giving one of three metrics a different window makes "the window" ambiguous in a tool three specs cite. Four controls, three mutations, each red for its stated reason — including the one the evidence files actually hit, where a refusal is printed as a measured zero. SessionShape.md §4 carries the correction with the real per-pass figures (1.1%-6.3%), beside the eighth trusted-arithmetic instance. This is the ninth, and the second in this same metric. It also shows what the frozen 0.0% hid: against the pinned 7.8%-8.6%, batching has got worse, and six passes reported a breach that was moving underneath them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
10 KiB
Session Shape
Status: v1.0 — 2026-07-31, from CB-WP-0003 T04.
Instrument: make cost (CostAccounting.md).
How to shape an agentic session so it does not cost more than the work is worth. Every claim here carries the measurement it rests on. This spec exists because the numbers were surprising and because the advice this task was created to give turned out to be wrong.
1. The cost model
cost ≈ Σ over turns ( context_size × cache_read_rate ) + output + cache writes
Measured on the CB-WP-0001 pass: 88.0% of spend is cache, 11.9% is output, at 249:1 cache-read tokens to output tokens. Cost is driven by how much context is re-read per turn, not by how much the model writes.
Marginal cost of one turn, at the Opus cache-read rate ($0.50/M):
| context | $/turn |
|---|---|
| 20,000 | $0.010 |
| 50,000 | $0.025 |
| 100,000 | $0.050 |
| 200,000 | $0.100 |
| 400,000 | $0.200 |
| 540,000 | $0.270 |
SS-01. A turn's cost is a function of context size alone, to within ~12%. Optimizing output length is not a cost lever; optimizing context is.
2. Compaction is the primary control
The original premise of this spec was wrong and is recorded rather than deleted. CB-WP-0003 T04 was written to prescribe one task per session, on the theory that mean context grows with turn count and long sessions are therefore quadratic. Measured:
| segment | turns | mean context | total | $/turn |
|---|---|---|---|---|
| start → compact 1 | 136 | 304,178 | $62.19 | $0.457 |
| compact 1 → compact 2 | 202 | 193,493 | $30.01 | $0.149 |
The 202 turns after the first compaction cost less than half the 136 before it. Sessions are bounded-quadratic: context grows between compactions and resets at each one.
Compaction events, read from compactMetadata in the transcript:
| pre | post | reduction | duration | |
|---|---|---|---|---|
| C1 | 542,991 | 19,974 | 27× | 151 s |
| C2 | 344,954 | 19,035 | 18× | 97 s |
SS-02. The failure mode is a long uncompacted session, not a long session. Turn count is not the variable to control; context is.
2.1 When compaction pays
A compaction costs its own summarization call plus one cache
re-establishment. The re-establishment is measurable — the first post-compact
turn wrote 36,149 cache tokens and cost $0.382. The summarization call
itself does not appear in the transcript's usage records, so it is
bounded rather than measured: at worst it is one full read of the
pre-compaction context (543k tokens ≈ $2.72 at the uncached Opus input
rate), at best a cache read of the same (≈ $0.27).
Savings are immediate and permanent: dropping from 543k to 20k context saves $0.262 per subsequent turn.
SS-03. Break-even is 2–11 turns depending on where the summarization call falls in that range. Any session expecting more than ~11 further turns should compact. In practice this means: compact whenever context exceeds ~300k and work remains.
Measured context distribution across the CB-WP-0001 session shows how much was spent above that line — p75 was 308,325 and p90 was 481,722, so roughly a quarter of all turns ran at more than 2× the recommended ceiling.
| percentile | context |
|---|---|
| p50 | 156,771 |
| p75 | 308,325 |
| p90 | 481,722 |
| p99 | 553,252 |
3. Fresh session vs compaction
A fresh session is not free, and this is the measurement that overturned the original advice. Cold-start context floors, measured over the first 20 turns of each session:
| session | start type | mean context, first 20 turns | cost |
|---|---|---|---|
8cbd5701 |
cold (system + CLAUDE.md + orientation) | 50,872 | $2.17 |
f1eb1147 |
seeded by a compact summary | 65,819 | $1.90 |
A cold session starts ~51k and must then re-read the committed artifacts to become productive — the survey, the spec, the ADR — which is exactly the context a compaction summary already contains. A compact-seeded session starts ~66k with that reading already done.
SS-04. Prefer compaction over a fresh session for continuing work. A fresh session's lower floor (51k vs 66k) is repaid within the first few turns of re-reading artifacts the summary already held.
Prefer a fresh session when the task changes, because then the prior context is not an asset — it is 66k of irrelevance re-read every turn.
4. Batching
SS-05. Independent tool calls in one turn cost nothing extra: a turn is billed for its context regardless of how many tool calls it carries. Two calls in one turn cost half of two calls in two turns.
Measured batching rate:
| session | responses w/ tool calls | tool calls | responses batching ≥2 | calls in batched turns |
|---|---|---|---|---|
8cbd5701 |
361 | 406 | 31 (8.6%) | 76 (18.7%) |
f1eb1147 |
146 | 158 | 12 (8.2%) | 24 (15.2%) |
pinned at fc76445, all transcripts (instrument) |
322 | 362 | 25 (7.8%) | 65 (18.0%) |
Correction to a prior claim. CB-WP-0001's retrospective and the first
draft of CB-WP-0003 T04 both stated "0 of 330 tool calls were batched".
That is wrong: 330 was the count of responses making exactly one call,
not the total number of calls. The real figure is 7.8–8.6% of responses and
18.0–18.7% of calls, depending on the window. The claim was never re-derived before being carried into a
workplan — the failure mode InnerLoop v1.1 rule 13 exists for, and the
eighth recorded error instance, in the trusted-arithmetic class that
history/260731-inner-loop-rule-audit.md identified as having no
executable defence.
Correction (CB-WP-0013 T01): SH-3 was never read as measured
Six evidence files quote SH-3 0.0%, and none of them measured it.
cb-cost --shape-budget reports the window since the last commit, and
it is read while writing the evidence file — immediately after a commit,
when that window holds one or two responses. SH-1 and SH-2 are location
statistics and survive a sample of two. SH-3 is a rate, and at n=2 its
only possible values are 0%, 50% and 100%. One window, three metrics, and
it is only wrong for one of them.
Measured over real per-pass windows instead:
| window since | responses | SH-3 |
|---|---|---|
0d2ab22 |
115 | 6.2% |
6fb0aea |
113 | 6.3% |
331e7e9 |
95 | 1.1% |
c5fa610 |
74 | 1.4% |
883b608 |
70 | 1.4% |
84d6886 |
24 | 4.3% |
The substantive claim survives — SH-3 is badly unmet — but the constant 0.0% hid something the real numbers show: against the 7.8–8.6% pinned above, batching has got worse, and a metric frozen at a constant could not have shown a trend at all. Six passes reported a breach and the breach was moving underneath them.
This is the ninth recorded instance of the trusted-arithmetic class, and
the second in this same metric — the eighth is the 0 of 330 tool calls
figure corrected immediately above. SH-3 is the number this project keeps
getting wrong, which is itself a reason to doubt the metric rather than the
behaviour (§SH-3 disposition).
The fix, cb-cost.py sh3_line(): below a minimum sample the tool
prints insufficient sample and no verdict. The floor is derived, not
round — if the true rate were exactly the 20% target, the chance of seeing
zero batched turns in n responses is 0.8^n; at n = 14 that is 4.4%,
so "0 batched in 14" rules out a target-meeting rate at ~95%. Below that
the tool has nothing to say and now says so.
At an 8.6% batching rate there is real headroom, but the saving is bounded: eliminating every avoidable single-call turn would remove at most a few percent of turns, worth roughly $2–4 on a $93 pass. Batch because it is free, not because it is the lever. The lever is §2.
5. What this spec does not support
- n = 2 sessions, one repo, one model mix. Every ratio here is from CB-WP-0001/0002 and should be re-measured before being treated as general.
- The compaction summarization call is bounded, not measured. SS-03's 2–11 turn break-even is a range because of it. Narrowing it requires a cost source that sees non-transcript calls — the billing API (CB-RES-0002 C4), which needs an admin key this machine does not have.
- Context size is read as
cache_read_input_tokens, which is the cached prefix, not the full prompt. On a cache miss the true context is larger and the turn costs 10× more. Cache misses were not isolated in this measurement. - No claim about quality. Compaction discards detail. Every number here says compaction is cheaper; none says the work is as good. A pass that compacts and then repeats work it forgot has not saved anything, and this spec cannot detect that.
6. Acceptance metrics
| ID | Metric | Target | Measured (pinned fc76445) |
Verdict | Instrument |
|---|---|---|---|---|---|
| SH-1 | mean context per turn | ≤ 200,000 | 232,982 | unmet | make cost-pin |
| SH-2 | p90 context per turn | ≤ 300,000 | 492,042 | unmet | make cost-pin |
| SH-3 | batching rate (responses with ≥2 tool calls) | ≥ 20% | 7.8% | unmet | make cost-pin |
cb-cost emits all three (session shape block), so these are ratified
under InnerLoop v1.1 rather than hand-derived. Supporting counts from the
same run: 362 tool calls in 322 responses, 65 of them in batched turns.
All three targets are unmet, and that is the finding, not a defect in the targets. Per the loop's definition of done, an unmet metric is an output:
- SH-1/SH-2 are unmet because CB-WP-0001 ran 136 turns before its first compaction. §2 is the remedy and it is now written down.
- SH-3 is set above the measured rate deliberately. Whether 20% is reachable or arbitrary is unknown until a pass tries; the first pass that reports it will say. It is the cheapest of the three to move, and the least valuable — §4 puts the ceiling at $2–4 on a $93 pass.
Targets are not being retargeted to match the measurement. That is the structure CB-WP-0003 T07 exists to prevent, and doing it here — in the commit that first measures them — would be the exact defect.