clay-borg/specs/SessionShape.md

227 lines
10 KiB
Markdown
Raw Permalink Normal View History

T04: specs/SessionShape.md — compaction is the lever, not session length The task's original premise was wrong and is recorded rather than deleted. It was written to prescribe one task per session; measurement says the variable is context, not turn count. Measured: - compaction cut context 27x (542,991 -> 19,974) and cost/turn 3.1x; the 202 turns after C1 cost less than half the 136 before it - a turn costs $0.010 at 20k context and $0.270 at 540k - break-even for a compaction is 2-11 turns, so: compact whenever context exceeds ~300k and work remains - a fresh session is NOT free -- cold start floors at ~51k and must then re-read the artifacts a compact summary already holds (~66k). Prefer compaction to continue work; prefer a fresh session when the task changes, because then prior context is pure overhead. cb-cost now emits SH-1..SH-3 so the targets come from the instrument rather than from analysis, per InnerLoop v1.1. All three are UNMET (mean context 232,982 vs 200,000; p90 492,042 vs 300,000; batching 7.8% vs 20%) and are reported unmet rather than retargeted -- retargeting in the commit that first measures is precisely what T07 exists to prevent. Eighth error instance found while writing this: CB-WP-0001's claim that "0 of 330 tool calls were batched" is wrong. 330 was the count of single-call responses, not the total; 31 responses batched, covering 76 calls. It was carried into this workplan unverified. Trusted-arithmetic class -- the one the T01 audit flagged as having no executable defence, confirming that finding within hours of making it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 09:19:32 +02:00
# Session Shape
Status: **v1.0** — 2026-07-31, from CB-WP-0003 T04.
Instrument: `make cost` ([CostAccounting.md](CostAccounting.md)).
How to shape an agentic session so it does not cost more than the work is
worth. **Every claim here carries the measurement it rests on.** This spec
exists because the numbers were surprising and because the advice this task
was created to give turned out to be wrong.
---
## 1. The cost model
```text
cost ≈ Σ over turns ( context_size × cache_read_rate ) + output + cache writes
```
Measured on the CB-WP-0001 pass: **88.0% of spend is cache, 11.9% is
output**, at 249:1 cache-read tokens to output tokens. Cost is driven by
**how much context is re-read per turn**, not by how much the model writes.
Marginal cost of one turn, at the Opus cache-read rate ($0.50/M):
| context | $/turn |
|---|---|
| 20,000 | $0.010 |
| 50,000 | $0.025 |
| 100,000 | $0.050 |
| 200,000 | $0.100 |
| 400,000 | $0.200 |
| 540,000 | $0.270 |
> **SS-01.** A turn's cost is a function of context size alone, to within
> ~12%. Optimizing output length is not a cost lever; optimizing context is.
## 2. Compaction is the primary control
**The original premise of this spec was wrong and is recorded rather than
deleted.** CB-WP-0003 T04 was written to prescribe *one task per session*,
on the theory that mean context grows with turn count and long sessions are
therefore quadratic. Measured:
| segment | turns | mean context | total | $/turn |
|---|---|---|---|---|
| start → compact 1 | 136 | 304,178 | $62.19 | **$0.457** |
| compact 1 → compact 2 | 202 | 193,493 | $30.01 | **$0.149** |
The 202 turns *after* the first compaction cost less than half the 136
before it. Sessions are **bounded-quadratic**: context grows between
compactions and resets at each one.
Compaction events, read from `compactMetadata` in the transcript:
| | pre | post | reduction | duration |
|---|---|---|---|---|
| C1 | 542,991 | 19,974 | **27×** | 151 s |
| C2 | 344,954 | 19,035 | **18×** | 97 s |
> **SS-02.** The failure mode is a long **uncompacted** session, not a long
> session. Turn count is not the variable to control; context is.
### 2.1 When compaction pays
A compaction costs its own summarization call plus one cache
re-establishment. The re-establishment is measurable — the first post-compact
turn wrote 36,149 cache tokens and cost **$0.382**. The summarization call
itself does **not** appear in the transcript's `usage` records, so it is
bounded rather than measured: at worst it is one full read of the
pre-compaction context (543k tokens ≈ $2.72 at the uncached Opus input
rate), at best a cache read of the same (≈ $0.27).
Savings are immediate and permanent: dropping from 543k to 20k context saves
**$0.262 per subsequent turn**.
> **SS-03.** Break-even is **211 turns** depending on where the
> summarization call falls in that range. Any session expecting more than
> ~11 further turns should compact. In practice this means: **compact
> whenever context exceeds ~300k and work remains.**
Measured context distribution across the CB-WP-0001 session shows how much
was spent above that line — p75 was 308,325 and p90 was 481,722, so roughly
a quarter of all turns ran at more than 2× the recommended ceiling.
| percentile | context |
|---|---|
| p50 | 156,771 |
| p75 | 308,325 |
| p90 | 481,722 |
| p99 | 553,252 |
## 3. Fresh session vs compaction
A fresh session is **not** free, and this is the measurement that overturned
the original advice. Cold-start context floors, measured over the first 20
turns of each session:
| session | start type | mean context, first 20 turns | cost |
|---|---|---|---|
| `8cbd5701` | cold (system + CLAUDE.md + orientation) | 50,872 | $2.17 |
| `f1eb1147` | seeded by a compact summary | 65,819 | $1.90 |
A cold session starts ~51k and must then **re-read the committed artifacts**
to become productive — the survey, the spec, the ADR — which is exactly the
context a compaction summary already contains. A compact-seeded session
starts ~66k with that reading already done.
> **SS-04.** Prefer compaction over a fresh session for continuing work.
> A fresh session's lower floor (51k vs 66k) is repaid within the first few
> turns of re-reading artifacts the summary already held.
>
> Prefer a fresh session when the *task changes*, because then the prior
> context is not an asset — it is 66k of irrelevance re-read every turn.
## 4. Batching
> **SS-05.** Independent tool calls in one turn cost nothing extra: a turn
> is billed for its context regardless of how many tool calls it carries.
> Two calls in one turn cost half of two calls in two turns.
Measured batching rate:
| session | responses w/ tool calls | tool calls | responses batching ≥2 | calls in batched turns |
|---|---|---|---|---|
| `8cbd5701` | 361 | 406 | 31 (8.6%) | 76 (18.7%) |
| `f1eb1147` | 146 | 158 | 12 (8.2%) | 24 (15.2%) |
| pinned at `fc76445`, all transcripts (instrument) | 322 | 362 | 25 (7.8%) | 65 (18.0%) |
**Correction to a prior claim.** CB-WP-0001's retrospective and the first
draft of CB-WP-0003 T04 both stated *"0 of 330 tool calls were batched"*.
That is wrong: 330 was the count of responses making exactly **one** call,
not the total number of calls. The real figure is 7.88.6% of responses and
18.018.7% of calls, depending on the window. The claim was never re-derived before being carried into a
workplan — the failure mode InnerLoop v1.1 rule 13 exists for, and the
**eighth** recorded error instance, in the trusted-arithmetic class that
`history/260731-inner-loop-rule-audit.md` identified as having no
executable defence.
CB-WP-0013-T01: SH-3 refuses a window that cannot carry a rate The metric six evidence files reported as 0.0% was never measured. --shape-budget windows since the last commit, and it is read while writing the evidence file — right after a commit, when the window holds one or two responses. SH-1/SH-2 are location statistics and survive n=2. SH-3 is a rate: at n=2 its only possible values are 0%, 50%, 100%. cb-cost.py gains sh3_line(), which below a minimum sample prints "insufficient sample" and no verdict. The floor is derived: if the true rate were exactly the 20% target, P(zero batched in n) = 0.8^n, and at n=14 that is 4.4% — so "0 batched in 14" rules out a target-meeting rate at ~95%. Below that the tool has nothing to say and now says so. The window was kept rather than split. SH-3 could have been given a per-pass window, but the budget's purpose is the open remainder since the last commit, and giving one of three metrics a different window makes "the window" ambiguous in a tool three specs cite. Four controls, three mutations, each red for its stated reason — including the one the evidence files actually hit, where a refusal is printed as a measured zero. SessionShape.md §4 carries the correction with the real per-pass figures (1.1%-6.3%), beside the eighth trusted-arithmetic instance. This is the ninth, and the second in this same metric. It also shows what the frozen 0.0% hid: against the pinned 7.8%-8.6%, batching has got worse, and six passes reported a breach that was moving underneath them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 07:23:04 +02:00
### Correction (CB-WP-0013 T01): SH-3 was never read as measured
**Six evidence files quote `SH-3 0.0%`, and none of them measured it.**
`cb-cost --shape-budget` reports the window **since the last commit**, and
it is read while writing the evidence file — immediately *after* a commit,
when that window holds one or two responses. SH-1 and SH-2 are location
statistics and survive a sample of two. SH-3 is a **rate**, and at n=2 its
only possible values are 0%, 50% and 100%. One window, three metrics, and
it is only wrong for one of them.
Measured over real per-pass windows instead:
| window since | responses | SH-3 |
|---|---:|---:|
| `0d2ab22` | 115 | 6.2% |
| `6fb0aea` | 113 | 6.3% |
| `331e7e9` | 95 | 1.1% |
| `c5fa610` | 74 | 1.4% |
| `883b608` | 70 | 1.4% |
| `84d6886` | 24 | 4.3% |
The substantive claim survives — SH-3 is badly unmet — but the constant
0.0% hid something the real numbers show: against the **7.88.6%** pinned
above, **batching has got worse**, and a metric frozen at a constant could
not have shown a trend at all. Six passes reported a breach and the breach
was moving underneath them.
**This is the ninth recorded instance of the trusted-arithmetic class, and
the second in this same metric** — the eighth is the `0 of 330 tool calls`
figure corrected immediately above. SH-3 is the number this project keeps
getting wrong, which is itself a reason to doubt the metric rather than the
behaviour (§SH-3 disposition).
**The fix**, `cb-cost.py` `sh3_line()`: below a minimum sample the tool
prints `insufficient sample` and no verdict. The floor is derived, not
round — if the true rate were exactly the 20% target, the chance of seeing
zero batched turns in `n` responses is `0.8^n`; at **n = 14** that is 4.4%,
so "0 batched in 14" rules out a target-meeting rate at ~95%. Below that
the tool has nothing to say and now says so.
T04: specs/SessionShape.md — compaction is the lever, not session length The task's original premise was wrong and is recorded rather than deleted. It was written to prescribe one task per session; measurement says the variable is context, not turn count. Measured: - compaction cut context 27x (542,991 -> 19,974) and cost/turn 3.1x; the 202 turns after C1 cost less than half the 136 before it - a turn costs $0.010 at 20k context and $0.270 at 540k - break-even for a compaction is 2-11 turns, so: compact whenever context exceeds ~300k and work remains - a fresh session is NOT free -- cold start floors at ~51k and must then re-read the artifacts a compact summary already holds (~66k). Prefer compaction to continue work; prefer a fresh session when the task changes, because then prior context is pure overhead. cb-cost now emits SH-1..SH-3 so the targets come from the instrument rather than from analysis, per InnerLoop v1.1. All three are UNMET (mean context 232,982 vs 200,000; p90 492,042 vs 300,000; batching 7.8% vs 20%) and are reported unmet rather than retargeted -- retargeting in the commit that first measures is precisely what T07 exists to prevent. Eighth error instance found while writing this: CB-WP-0001's claim that "0 of 330 tool calls were batched" is wrong. 330 was the count of single-call responses, not the total; 31 responses batched, covering 76 calls. It was carried into this workplan unverified. Trusted-arithmetic class -- the one the T01 audit flagged as having no executable defence, confirming that finding within hours of making it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 09:19:32 +02:00
At an 8.6% batching rate there is real headroom, but the saving is bounded:
eliminating every avoidable single-call turn would remove at most a few
percent of turns, worth roughly $24 on a $93 pass. **Batch because it is
free, not because it is the lever.** The lever is §2.
## 5. What this spec does not support
- **n = 2 sessions, one repo, one model mix.** Every ratio here is from
CB-WP-0001/0002 and should be re-measured before being treated as general.
- **The compaction summarization call is bounded, not measured.** SS-03's
211 turn break-even is a range because of it. Narrowing it requires a
cost source that sees non-transcript calls — the billing API (CB-RES-0002
C4), which needs an admin key this machine does not have.
- **Context size is read as `cache_read_input_tokens`**, which is the cached
prefix, not the full prompt. On a cache miss the true context is larger
and the turn costs 10× more. Cache misses were not isolated in this
measurement.
- **No claim about quality.** Compaction discards detail. Every number here
says compaction is cheaper; none says the work is as good. A pass that
compacts and then repeats work it forgot has not saved anything, and this
spec cannot detect that.
## 6. Acceptance metrics
| ID | Metric | Target | Measured (pinned `fc76445`) | Verdict | Instrument |
|---|---|---|---|---|---|
| SH-1 | mean context per turn | ≤ 200,000 | **232,982** | **unmet** | `make cost-pin` |
| SH-2 | p90 context per turn | ≤ 300,000 | **492,042** | **unmet** | `make cost-pin` |
| SH-3 | batching rate (responses with ≥2 tool calls) | ≥ 20% | **7.8%** | **unmet** | `make cost-pin` |
`cb-cost` emits all three (`session shape` block), so these are ratified
under InnerLoop v1.1 rather than hand-derived. Supporting counts from the
same run: **362 tool calls in 322 responses, 65 of them in batched turns.**
**All three targets are unmet, and that is the finding, not a defect in the
targets.** Per the loop's definition of done, an unmet metric is an output:
- SH-1/SH-2 are unmet because CB-WP-0001 ran 136 turns before its first
compaction. §2 is the remedy and it is now written down.
- SH-3 is set above the measured rate deliberately. Whether 20% is
reachable or arbitrary is unknown until a pass tries; the first pass that
reports it will say. It is the cheapest of the three to move, and the
least valuable — §4 puts the ceiling at $24 on a $93 pass.
Targets are **not** being retargeted to match the measurement. That is the
structure CB-WP-0003 T07 exists to prevent, and doing it here — in the
commit that first measures them — would be the exact defect.