# CB-RES-0003: which agent turns can become deterministic compute capability: meta.loop.mechanical-work status: draft tier: M (structural M — no new capability port; touches the one command surface and the cost instrument; chaos d4=2 → no override) instrument: `make cost-mix` (`tools/cb-cost.py`, tool-mix block) Review of where token-priced agent turns did work a deterministic tool could have done, so that capacity moves to judgment rather than mechanics. Commissioned 2026-07-31. **Method.** Every turn in both clay-borg session transcripts (573 responses, $134 total) was classified by the tool calls it made. The classifier is committed as `classify_tool()` in `tools/cb-cost.py` and its output is emitted by `make cost-mix` — the numbers below are reproducible, and the same command measures whether any fix worked. --- ## The measurement ```text mech environment setup 84 turns $ 15.33 mech ad-hoc text patching 75 turns $ 13.86 git 37 turns $ 13.85 mech hub task status 25 turns $ 7.46 mech orientation / inspect 49 turns $ 6.87 hub other 32 turns $ 6.22 mech ad-hoc transcript analysis 39 turns $ 4.56 mech workplan status edit 21 turns $ 4.06 make (gates) 3 turns $ 1.82 MECHANICAL (deduplicated) 290 turns $ 51.26 = 38% of pass ``` **38% of spend went through turns whose tool calls were mechanical.** A turn's whole cost is charged to every category it touched, so the rows overlap and the deduplicated line is the honest total. `git` and `make` are excluded from *mechanical* deliberately: a commit message is judgment, and running a gate is the point. ## Candidate 1 — environment setup: 84 turns, $15.33 The single largest category, and pure friction. `cd` (51) and `export PATH="$HOME/.cargo/bin:$PATH"` (10+) recur because the shell does not persist state between calls and `cargo` is not on the default path. Every `make` invocation that touches Rust had to be prefixed. - **Deterministic replacement:** put `~/.cargo/bin` on the path the agent starts with, and make every `make` target self-sufficient (the Makefile already knows where it is). `tools/dep-weight.py` already special-cases this with a helpful error — evidence the friction was noticed and patched at the leaf instead of the root. - **Expected effect:** these turns do not become cheaper, they **stop existing**. Predicted saving: **~$12–15 per two-session pass**, the cleanest win in this review. - **Risk:** near zero. Nothing depends on the current behaviour. ## Candidate 2 — ad-hoc text patching: 75 turns, $13.86 Inline `python3 - <<'PY'` heredocs that open a markdown file, string-replace, and write it back. Written fresh each time, unreviewed, and the mechanism behind the **duplicated-fact-drift** class: a number changes and every artifact quoting it must be found and patched by hand. The acceptance figure moved four times and each move cost a sweep across three or four files. - **Deterministic replacement:** a *fact registry*. Numbers that appear in more than one artifact are declared once (`facts.toml` or emitted by the tool that measures them) and injected into markdown by a generator, with `make facts-check` failing when a committed artifact disagrees with the registry. - **Expected effect:** removes the sweep, and gives the DFD class its first executable gate — which InnerLoop v1.2 currently states only as prose. Predicted saving: **~$6–9 per pass**, plus the error class. - **Risk:** moderate. A registry that is itself hand-maintained just moves the problem. It must be generated from instruments where possible, and the check must be the enforcement, not the generation. ## Candidate 3 — hub task status: 25 turns, $7.46 Every `update_task_status` call was hand-written, including `tokens_in`/`tokens_out` — **which were estimates I typed**, in a session whose entire subject was that estimated token counts are worthless. The hub holds numbers derived from the very habit CB-WP-0002 disproved. - **Deterministic replacement:** `make task-done T=T05` — flips the workplan file, reads the measured cost for that task from `cb-cost`, and pushes the hub event with real numbers. One command replaces an edit, a status call, and a fabricated figure. - **Expected effect:** **~$7 per pass**, and the hub stops holding fiction. Combines with candidate 5. - **Risk:** low. `cb-cost --by-task` already produces the figure. ## Candidate 4 — orientation / inspect: 49 turns, $6.87 `grep`/`ls`/`wc` to answer "what is the state of this repo" — which workplan is active, which tasks are open, which gates pass, what is uncommitted. - **Deterministic replacement:** `make status` printing the loop state in one shot: active workplan, task counts, gate results, open cost since last commit, provisional item ages. - **Expected effect:** turns ~10 orientation turns into 1 at the start of a session, and shrinks cold-start context (SS-04). Predicted saving: **~$4 per pass**. - **Risk:** low, but the saving is softer than it looks — some inspection is genuinely exploratory and will not disappear. ## Candidate 5 — workplan status edit: 21 turns, $4.06 Heredocs doing `s.replace("status: todo", "status: done")` on a workplan file. Purely mechanical, and error-prone: it silently does nothing if the task is already done or the ID is mistyped. - **Deterministic replacement:** folded into candidate 3's `make task-done`, which can *fail* on an unknown task ID instead of no-op'ing. - **Expected effect:** **~$4 per pass** and one class of silent no-op removed. ## Candidate 6 — ad-hoc transcript analysis: 39 turns, $4.56 Already partly solved: `cb-cost` subsumed most of this during CB-WP-0002, and this review's own classifier is now committed rather than ad-hoc. The residue is one-off questions (context percentiles, compaction boundaries, per-model splits) that were each written fresh. - **Deterministic replacement:** promote the recurring ones to flags. Most already exist (`--by-task`, `--composition`, session shape, tool mix). - **Expected effect:** **~$2 per pass**, diminishing. Listed for completeness, not priority. ## What must NOT be automated Stating this because a review that only finds savings is not a review. - **`git` (37 turns, $13.85)** is the second-most expensive category and is mostly *commit message authorship* — the highest-output-token turns in the corpus. That output is the project's reasoning record. Automating it would save money and destroy the thing that makes corrections cheap. - **`make` gates (3 turns, $1.82)** are already deterministic; the agent merely invokes them. Correctly cheap. - **Judgment work is invisible in this table** — writing a spec, choosing an attribution model, deciding a target is legitimate. That is where the remaining 62% went, and it is what the freed capacity should buy. ## Verdict and expected total | # | candidate | turns | measured | predicted saving/pass | confidence | |---|---|---|---|---|---| | 1 | environment setup | 84 | $15.33 | **$12–15** | high | | 2 | fact registry (text patching) | 75 | $13.86 | **$6–9** | medium | | 3+5 | `make task-done` | 46 | $11.52 | **$9–11** | high | | 4 | `make status` | 49 | $6.87 | **$4** | medium | | 6 | cb-cost flags | 39 | $4.56 | **$2** | low | | | **total** | | **$51.26** | **$33–41** | | Predicted recovery is **25–30% of a pass**, against a measured 38% mechanical share. The gap is deliberate: some inspection and some patching is genuinely exploratory and will not vanish. **The prediction is falsifiable and must be checked.** `make cost-mix` emits the same categories, so the next pass measures whether these turns disappeared or merely relocated — the failure mode being that an agent which no longer writes heredocs simply writes more prose instead. A control loop that does not test for relocation is not a control loop. ## Risks in this review itself - **A turn's cost is charged to every category it touched**, so per-row figures overstate. The deduplicated $51.26 is the defensible number; per-candidate savings are apportioned from it and are estimates. - **n = 2 sessions, one repo, one agent.** The mix is a property of how this project was built, not a general law. - **The classifier is a regex over shell commands.** It cannot see intent: a `cd` that precedes real work is charged the whole turn. This inflates candidate 1, which is why its predicted saving is below its measured cost despite being the most certain fix.