clay-borg/workplans/CB-WP-0004-mechanical-work.md
tegwick 3f1dbac164 CB-WP-0004 T01: fix environment friction at the root
CB-RES-0003 measured 84 turns and $15.33 — the largest mechanical
category — spent prefixing commands with `cd` and
`export PATH="$HOME/.cargo/bin:$PATH"`. Both causes are now fixed once
instead of per-leaf.

tools/repo.py resolves the repo root from __file__ and cargo from PATH
then the standard rustup locations. Every tool imports ROOT from it, so
the repo path is stated once rather than redefined in four files —
single source of fact, the rule DFD earned in InnerLoop v1.2.
rule-coverage and dep-weight now call enter_root(), which is why their
relative paths did not need rewriting one by one.

The Makefile derives REPO from MAKEFILE_LIST and resolves CARGO the same
way, so `make -C <repo> <target>` works from any directory with no
prefix.

make env-test is the positive control, and is in `make all`: every tool
runs from / with PATH=/usr/bin:/bin. Without it this fix could regress
silently and invalidate T05's measurement — the whole point of the
control loop.

dep-weight's "cargo not on PATH" error is kept rather than deleted. It
should now be unreachable, and --self-test asserts cargo_bin() resolves
unaided; a control that never fires is cheaper than a regression.

loop-lint failed on repo.py on its first run — a reporting tool with a
positive control but no --self-test entry point. Second time the gate
has caught work from its own pass within the hour.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 10:13:52 +02:00

231 lines
9.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
id: CB-WP-0004
title: "Move mechanical turns off the token budget, and prove it worked"
status: in_progress
state_hub_workstream_id: "6880ac78-d817-41b9-b267-f12ff9deea28"
---
# Purpose
`research/CB-RES-0003-agent-vs-deterministic.md` measured every turn in
both clay-borg sessions by the tool calls it made:
```text
MECHANICAL (deduplicated) 290 turns $ 51.26 = 38% of pass
```
**38% of spend went through turns doing work a deterministic tool could
do.** The largest single category is `cd` and `export PATH` — 84 turns and
$15.33 of pure environment friction. The second is inline heredocs
string-patching markdown, which is also the mechanism behind the
duplicated-fact-drift error class InnerLoop v1.2 named and could not gate.
This workplan converts the five worthwhile categories to classic compute
and **measures whether that actually recovered anything**. The predicted
recovery is $3341 per pass, 2530%.
**The control loop is the point, not a formality.** The named failure mode
is *relocation*: an agent that can no longer write a heredoc may simply
write more prose, and the pass costs the same. `make cost-mix` emits the
same categories that produced the baseline, so the claim is falsifiable by
the same instrument that made it. A saving that cannot be demonstrated in
that table did not happen.
Per InnerLoop v1.2, targets here are **provisional until the instrument
emits them**, and no target may be moved in the commit that measures it
unless the instrument disproved it (§Step 4, correction vs retarget).
## Phase A — The certain wins
## Task: Remove environment friction
```task
id: CB-WP-0004-T01
status: done
priority: high
state_hub_task_id: "3ddfd3e2-8596-4969-a069-09577933fcbc"
```
84 turns / $15.33, the largest category and the least interesting work in
the corpus. `cargo` is not on the default path, so every Rust-touching
command carried `export PATH="$HOME/.cargo/bin:$PATH"`, and the shell does
not persist `cd`.
Fix at the root, not the leaf: `tools/dep-weight.py` already special-cases
the missing `cargo` with a helpful error, which is evidence the friction
was noticed and patched in the wrong place.
Deliver: every `make` target runs from a clean shell with no prefix, from
any directory. Document the one-line environment requirement in
`README.md` if one remains. Remove the leaf workaround in `dep-weight.py`
only if it becomes unreachable — a positive control that never fires is
still cheaper than a regression.
**Predicted:** environment-setup turns → **< 10** (from 84),
**$1215** recovered. Highest confidence in the review.
**Delivered.** `tools/repo.py` resolves the repo root from `__file__` and
`cargo` from PATH-then-rustup-locations; every tool imports it, so `REPO`
is now stated once rather than four times. The Makefile derives `REPO`
from `MAKEFILE_LIST` and resolves `CARGO` the same way. `make env-test`
is the positive control it runs every tool from `/` with
`PATH=/usr/bin:/bin`, and is wired into `make all`, so this cannot
silently regress and invalidate T05's measurement.
The leaf workaround in `dep-weight.py` was kept, not deleted: it is now
unreachable on a rustup machine, and `--self-test` asserts `cargo_bin()`
resolves unaided. Per the task text, a control that never fires is
cheaper than a regression.
`loop-lint` failed on `repo.py` immediately a new tool with a positive
control but no `--self-test` entry point. Second time the gate has caught
its own pass's work within the hour.
## Task: `make task-done` — one command for a task close
```task
id: CB-WP-0004-T02
status: todo
priority: high
state_hub_task_id: "8a1b59c7-6318-47ba-8cbc-83bb44dd221f"
```
Merges two categories: workplan status edits (21 turns, $4.06) and hub
task-status calls (25 turns, $7.46).
`make task-done T=CB-WP-0004-T02` must:
1. flip `status: todo` `done` in the workplan file, **failing loudly on
an unknown or already-done task** the heredocs it replaces silently
no-op'd on a typo;
2. read that task's **measured** cost and tokens from `cb-cost --by-task`;
3. push the hub event with the real numbers.
The third point is the one that matters beyond cost. Every
`update_task_status` in this project so far carried **hand-typed token
estimates**, in a repo whose central finding is that estimated token counts
are worthless. The hub currently holds fiction produced by the exact habit
CB-WP-0002 disproved.
**Predicted:** those 46 turns **~6**, **$911** recovered, and the hub
stops holding estimates. Add `--self-test` per InnerLoop v1.1.
## Task: `make status` — one-shot orientation
```task
id: CB-WP-0004-T03
status: todo
priority: medium
state_hub_task_id: "5466a510-37a5-4491-b93e-509cd400cc23"
```
49 turns / $6.87 of `grep`/`ls`/`wc` answering "what is the state of this
repo". Replace with one command printing: active workplan and task counts,
gate results, open spend since the last commit (CB-01), provisional item
ages, and any `loop-lint` findings.
Also shrinks cold-start context, which `specs/SessionShape.md` §3 measures
at ~51k for a fresh session the artifacts a new session reads to orient
are exactly what this prints.
**Predicted:** ~10 orientation turns **1 per session**, **$4** recovered.
Confidence medium: some inspection is genuinely exploratory and will not
disappear, and the review says so.
## Phase B — The one that also closes an error class
## Task: Fact registry and `make facts-check`
```task
id: CB-WP-0004-T04
status: todo
priority: high
state_hub_task_id: "b5dfbf1e-619b-4e64-8463-49ff2ce38b48"
```
75 turns / $13.86 of heredocs opening a markdown file, string-replacing a
number, and writing it back the mechanism behind **duplicated-fact
drift**, the fourth error class, which InnerLoop v1.2 states as prose and
cannot currently gate.
Two instances on record: a price sheet inlined into a spec went stale
within an hour of the real sheet changing, and the acceptance figure
$92.21 $92.87 $93.32 $93.15 had to be chased across a survey, a
workplan, and an evidence file on every move.
Deliver a registry where a number appearing in more than one artifact is
declared once **generated by the instrument that measures it wherever
possible**, not hand-maintained plus `make facts-check` failing when a
committed artifact disagrees with it.
**The trap to avoid, stated up front:** a hand-maintained registry moves
the problem rather than solving it, and would itself become a copy that
drifts. If generation from instruments proves impractical, deliver only
the *check* (detect the same number stated differently in two artifacts)
and say so a gate with no generator still closes the class.
**Predicted:** **$69** recovered, plus DFD's first executable gate.
Confidence medium; this is the hardest task here and the most valuable.
## Phase C — Prove it, or withdraw the claim
## Task: Control loop — measure recovery and test for relocation
```task
id: CB-WP-0004-T05
status: todo
priority: high
state_hub_task_id: "922346ac-1e99-4e94-9dc6-35c8fa0eddc5"
```
The task this workplan exists for. Run `make cost-mix` over the sessions
that executed T01T04 and commit
`evidence/CB-EV-0003-mechanical-work.md` comparing against the committed
baseline:
| category | baseline turns | baseline $ | predicted | measured | verdict |
|---|---|---|---|---|---|
| environment setup | 84 | $15.33 | <10 turns | | |
| ad-hoc text patching | 75 | $13.86 | $69 saved | | |
| hub task status + workplan edit | 46 | $11.52 | ~6 turns | | |
| orientation / inspect | 49 | $6.87 | $4 saved | | |
| **mechanical total** | **290** | **$51.26** | **$3341 saved** | | |
**Three tests, all of which must be reported:**
1. **Did the mechanical turns disappear?** Per-category, against
prediction. An unmet prediction is reported unmet, not retargeted.
2. **Did they relocate?** Total pass cost and non-mechanical turn counts
must be compared too. If mechanical turns fell and prose turns rose by
as much, the saving is zero and this workplan failed that is the
result to publish.
3. **Did quality hold?** `make all` green, and the same class of findings
still surfacing. A cheaper pass that catches fewer errors is worse, and
the loop has no metric for this yet record the judgment explicitly
rather than implying the cost number settles it.
Normalize per unit of work, not per session: passes differ in size, so
report **mechanical share of pass cost** (baseline: 38%) alongside
absolute dollars.
## Task: Retrospective
```task
id: CB-WP-0004-T06
status: todo
priority: low
state_hub_task_id: "47892c23-ef09-4910-a9af-327494289e05"
```
The question to answer honestly: **does converting agent work to
deterministic tooling actually recover capacity, or does the work
reappear elsewhere?**
This is the first pass in the project to make a *quantitative prediction
before acting*. Whether the prediction held is more informative than the
saving itself a loop that can forecast its own economics can plan; one
that cannot is guessing with numbers attached.
Record the prediction error per candidate, and whether the review's stated
confidence levels (high/medium/low) tracked reality. If they did not, the
next review should stop stating confidence, or state it differently.