clay-borg/workplans/CB-WP-0004-mechanical-work.md
tegwick b79ea9690d CB-WP-0007 T01+T03: window the metric, budget it, cap meta at 25%
Scope cut first, on the maintainer's decision after a spend review: the
project is 38% product / 62% loop-meta, cost per response is 2.9x worse
than its best window, and INTENT stage 0 still lacks a CLI player and
bots. CB-WP-0005 and CB-WP-0006 cost ~$74 — 31% of all spend — for zero
measured efficiency gain. T02 and T04 are cancelled unstarted.

T01: SH-1/SH-2/SH-3 now report over the window since the last commit, and
the cumulative figure is retained but labelled "history, NOT the metric".
The prediction held decisively — window 655,744 mean context against
cumulative 255,307, a 2.6x gap against a 20% refutation threshold. A
cumulative mean over 1,094 responses cannot detect a worsening trend
because the history outvotes the present.

T03: `make shape-budget`, modelled on CB-01/CB-02. Soft thresholds are the
existing SessionShape targets; hard is 1.5x, set before the next
measurement per §Step 4. Deliberately not in `make all` — failing the
build on context would block committing, and committing is what closes the
attribution window and is the natural point to compact, so a gate that
blocks the remedy is a trap. It fires HARD on its first run: 656,574
against a 300,000 ceiling.

InnerLoop v1.5 establishes the soft 25% meta budget. Workplans declare
kind: product|meta|mixed and `make status` reports the share; mixed splits
50/50 and says so. Soft on purpose — a task already started may be
finished, because stopping mid-task to satisfy a ratio wastes the work.
What it forbids is opening new meta work above the line. A pass that
exceeds it must say so in its evidence and name the product work
displaced.

First reading: 68% OVER, of $74.22 attributed. Product reads $0.00 because
the only product workplan, CB-WP-0001, predates qualified task ids and its
bare T## labels collide across passes — stated in the output rather than
papered over.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 14:12:07 +02:00

16 KiB
Raw Blame History

id kind title status state_hub_workstream_id
CB-WP-0004 meta Move mechanical turns off the token budget, and prove it worked done 6880ac78-d817-41b9-b267-f12ff9deea28

Purpose

research/CB-RES-0003-agent-vs-deterministic.md measured every turn in both clay-borg sessions by the tool calls it made:

    MECHANICAL (deduplicated)          290 turns  $   51.26  = 38% of pass

38% of spend went through turns doing work a deterministic tool could do. The largest single category is cd and export PATH — 84 turns and $15.33 of pure environment friction. The second is inline heredocs string-patching markdown, which is also the mechanism behind the duplicated-fact-drift error class InnerLoop v1.2 named and could not gate.

This workplan converts the five worthwhile categories to classic compute and measures whether that actually recovered anything. The predicted recovery is $3341 per pass, 2530%.

The control loop is the point, not a formality. The named failure mode is relocation: an agent that can no longer write a heredoc may simply write more prose, and the pass costs the same. make cost-mix emits the same categories that produced the baseline, so the claim is falsifiable by the same instrument that made it. A saving that cannot be demonstrated in that table did not happen.

Per InnerLoop v1.2, targets here are provisional until the instrument emits them, and no target may be moved in the commit that measures it unless the instrument disproved it (§Step 4, correction vs retarget).

Phase A — The certain wins

Task: Remove environment friction

id: CB-WP-0004-T01
status: done
priority: high
state_hub_task_id: "3ddfd3e2-8596-4969-a069-09577933fcbc"

84 turns / $15.33, the largest category and the least interesting work in the corpus. cargo is not on the default path, so every Rust-touching command carried export PATH="$HOME/.cargo/bin:$PATH", and the shell does not persist cd.

Fix at the root, not the leaf: tools/dep-weight.py already special-cases the missing cargo with a helpful error, which is evidence the friction was noticed and patched in the wrong place.

Deliver: every make target runs from a clean shell with no prefix, from any directory. Document the one-line environment requirement in README.md if one remains. Remove the leaf workaround in dep-weight.py only if it becomes unreachable — a positive control that never fires is still cheaper than a regression.

Predicted: environment-setup turns → < 10 (from 84), $1215 recovered. Highest confidence in the review.

Delivered. tools/repo.py resolves the repo root from __file__ and cargo from PATH-then-rustup-locations; every tool imports it, so REPO is now stated once rather than four times. The Makefile derives REPO from MAKEFILE_LIST and resolves CARGO the same way. make env-test is the positive control — it runs every tool from / with PATH=/usr/bin:/bin, and is wired into make all, so this cannot silently regress and invalidate T05's measurement.

The leaf workaround in dep-weight.py was kept, not deleted: it is now unreachable on a rustup machine, and --self-test asserts cargo_bin() resolves unaided. Per the task text, a control that never fires is cheaper than a regression.

loop-lint failed on repo.py immediately — a new tool with a positive control but no --self-test entry point. Second time the gate has caught its own pass's work within the hour.

Task: make task-done — one command for a task close

id: CB-WP-0004-T02
status: done
priority: high
state_hub_task_id: "8a1b59c7-6318-47ba-8cbc-83bb44dd221f"

Merges two categories: workplan status edits (21 turns, $4.06) and hub task-status calls (25 turns, $7.46).

make task-done T=CB-WP-0004-T02 must:

  1. flip status: tododone in the workplan file, failing loudly on an unknown or already-done task — the heredocs it replaces silently no-op'd on a typo;
  2. read that task's measured cost and tokens from cb-cost --by-task;
  3. push the hub event with the real numbers.

The third point is the one that matters beyond cost. Every update_task_status in this project so far carried hand-typed token estimates, in a repo whose central finding is that estimated token counts are worthless. The hub currently holds fiction produced by the exact habit CB-WP-0002 disproved.

Predicted: those 46 turns → ~6, $911 recovered, and the hub stops holding estimates. Add --self-test per InnerLoop v1.1.

Delivered. tools/task-done.py + make task-done T=<id>. It refuses on an unknown id, a typo'd id, an already-done task, a task with no state_hub_task_id, and — the one that matters — a task with no measured spend, rather than reporting an estimate. cb-cost gained by_task_detail (cost, response count, model histogram, and token components per task), and task-done imports cb-cost rather than parsing its printed table, so the hub number is not a copy that can drift.

The positive control found a real defect before the tool was used once. Attribution keyed on a bare T\d\d from the commit subject, so CB-WP-0002 T01, CB-WP-0003 T01 and CB-WP-0004 T01 all landed in one bucket. The self-test reported $12.10 for "T01"; the qualified figure is $2.33 — a 5.2× overstatement that would have been pushed to the hub as a measured number, reproducing the fiction this task exists to end, in a new form. Fixed by task_label(): qualified subjects (CB-WP-0004 T01) key on the full id, unqualified ones stay bare and are never retro-assigned to a workplan. The pinned $93.15 benchmark is unchanged, confirming historical attribution was not disturbed.

That is the fourth instance of trusted arithmetic (TA) — a number believed because it was produced by a program rather than by hand.

Task: make status — one-shot orientation

id: CB-WP-0004-T03
status: done
priority: medium
state_hub_task_id: "5466a510-37a5-4491-b93e-509cd400cc23"

49 turns / $6.87 of grep/ls/wc answering "what is the state of this repo". Replace with one command printing: active workplan and task counts, gate results, open spend since the last commit (CB-01), provisional item ages, and any loop-lint findings.

Also shrinks cold-start context, which specs/SessionShape.md §3 measures at ~51k for a fresh session — the artifacts a new session reads to orient are exactly what this prints.

Predicted: ~10 orientation turns → 1 per session, $4 recovered. Confidence medium: some inspection is genuinely exploratory and will not disappear, and the review says so.

Delivered. tools/status.py + make status, 22 lines of output: commit and working-tree state, every workplan with its task counts, the next todo task with its heading, spend since the last commit against CB-01/CB-02, and the three fast gates.

Two deliberate constraints. It does not build — a status command that takes two minutes gets replaced by ls within a day — and it states its own limit in the output (slow gates not run here), so a green status cannot be misread as a green make all. The self-test asserts the output stays under 40 lines, because the failure mode for an orientation tool is becoming another thing to skim.

Its positive control refuses the confidently-empty report: a parser that found zero workplans and zero tasks would print a clean, wrong picture, which is worse than the greps it replaces. That control also caught a dot-all regex that returned twenty paragraphs of a previous task's prose as the "task heading".

Phase B — The one that also closes an error class

Task: Fact registry and make facts-check

id: CB-WP-0004-T04
status: done
priority: high
state_hub_task_id: "b5dfbf1e-619b-4e64-8463-49ff2ce38b48"

75 turns / $13.86 of heredocs opening a markdown file, string-replacing a number, and writing it back — the mechanism behind duplicated-fact drift, the fourth error class, which InnerLoop v1.2 states as prose and cannot currently gate.

Two instances on record: a price sheet inlined into a spec went stale within an hour of the real sheet changing, and the acceptance figure $92.21 → $92.87 → $93.32 → $93.15 had to be chased across a survey, a workplan, and an evidence file on every move.

Deliver a registry where a number appearing in more than one artifact is declared once — generated by the instrument that measures it wherever possible, not hand-maintained — plus make facts-check failing when a committed artifact disagrees with it.

The trap to avoid, stated up front: a hand-maintained registry moves the problem rather than solving it, and would itself become a copy that drifts. If generation from instruments proves impractical, deliver only the check (detect the same number stated differently in two artifacts) and say so — a gate with no generator still closes the class.

Predicted: $69 recovered, plus DFD's first executable gate. Confidence medium; this is the hardest task here and the most valuable.

Delivered — both halves, not just the check. facts.toml holds 15 facts and is generated by make facts-gen from cb-cost, dep-weight and rule-coverage; the file opens with # GENERATED — do not edit and the self-test asserts that line is still there. The trap named in this task — a hand-maintained registry that becomes another drifting copy — is closed by facts-check re-running the instruments and failing if the committed registry disagrees with them. A stale registry cannot certify stale artifacts.

An artifact quoting a fact tags it: **$93.15** <!-- fact:pinned_total -->. 17 occurrences across 5 artifacts are now under the gate.

Falsified before being believed. Changing specs/CostAccounting.md line 158 from $93.15 to $92.87 — the exact historical drift — produced exit 1 and specs/CostAccounting.md:158 claims fact:pinned_total but does not state $93.15. The gate was tested against the class it exists to catch, on a real artifact, not only in its self-test.

What it does not close, stated rather than implied. 22 untagged literal copies remain, across specs/InnerLoop.md, specs/GameKernel.md, research/CB-RES-0002 and the older workplans. They are reported, not failed: tagging is opt-in, a number can legitimately recur, and a gate that fires on coincidence gets routed around. Naming the uncovered surface is more useful than claiming the class is closed.

InnerLoop's single-source-of-fact rule moves from prose to executable — v1.3.

Phase C — Prove it, or withdraw the claim

Task: Control loop — measure recovery and test for relocation

id: CB-WP-0004-T05
status: done
priority: high
state_hub_task_id: "922346ac-1e99-4e94-9dc6-35c8fa0eddc5"

The task this workplan exists for. Run make cost-mix over the sessions that executed T01T04 and commit evidence/CB-EV-0003-mechanical-work.md comparing against the committed baseline:

category baseline turns baseline $ predicted measured verdict
environment setup 84 $15.33 <10 turns
ad-hoc text patching 75 $13.86 $69 saved
hub task status + workplan edit 46 $11.52 ~6 turns
orientation / inspect 49 $6.87 $4 saved
mechanical total 290 $51.26 $3341 saved

Three tests, all of which must be reported:

  1. Did the mechanical turns disappear? Per-category, against prediction. An unmet prediction is reported unmet, not retargeted.
  2. Did they relocate? Total pass cost and non-mechanical turn counts must be compared too. If mechanical turns fell and prose turns rose by as much, the saving is zero and this workplan failed — that is the result to publish.
  3. Did quality hold? make all green, and the same class of findings still surfacing. A cheaper pass that catches fewer errors is worse, and the loop has no metric for this yet — record the judgment explicitly rather than implying the cost number settles it.

Normalize per unit of work, not per session: passes differ in size, so report mechanical share of pass cost (baseline: 38%) alongside absolute dollars.

Measured — CB-EV-0003. cb-cost gained --since, so the baseline window (--pin 578dcbe, 662 responses, $135.60) and this pass (--since 578dcbe, 83 responses, $7.82) are two disjoint windows over the same transcripts.

Mechanical share fell 38.2% → 32.4%. Six points, against a predicted 2530. The aggregate prediction is reported unmet.

Two categories met their prediction and two did not:

  • environment setup 11.5% → 0.6% of pass (85 turns → 1) — met
  • hub + workplan closes 8.5% → 0.0% (46 turns → 0) — met
  • text patching 10.4% → 9.6% — not met
  • orientation 5.1% → 21.2% — not met, worse

The confound is stated in the evidence before any defence: this is the pass that built the tools, and the two categories that missed are precisely the two whose tools were under construction. Writing make status requires inspecting what it summarizes. The clean test is the next pass, which uses the tools without building them — that test is carried forward, not waived.

Relocation: not supported. Output tokens per response fell 896 → 681. The rise in output's share of cost is a shrinking denominator (cache-read fell as context shrank), not more prose. Cost per response halved, and the evidence explicitly refuses to claim it: that is compaction (mean context 232,982 → 117,822), not tooling.

Quality held. make all green with two new gates, and four findings surfaced this pass — three of them caught by controls written this pass, including a 5.2× attribution error that would have reached the hub as a measured number.

Task: Retrospective

id: CB-WP-0004-T06
status: done
priority: low
state_hub_task_id: "47892c23-ef09-4910-a9af-327494289e05"

The question to answer honestly: does converting agent work to deterministic tooling actually recover capacity, or does the work reappear elsewhere?

This is the first pass in the project to make a quantitative prediction before acting. Whether the prediction held is more informative than the saving itself — a loop that can forecast its own economics can plan; one that cannot is guessing with numbers attached.

Record the prediction error per candidate, and whether the review's stated confidence levels (high/medium/low) tracked reality. If they did not, the next review should stop stating confidence, or state it differently.

Recorded — 260731-mechanical-work-retrospective.md.

The answer to the question is neither of the two offered: tooling recovers capacity where it removes the possibility of doing the work by hand, and recovers nothing where it merely offers a better option. Sorting the five candidates by that test reproduces the results exactly — the two that removed the manual path met their predictions, the two that left grep and heredocs one keystroke away did not.

Confidence sorted the extremes and said nothing in the middle, which is where every judgement call lives: both high met, both medium missed, and the single low also met — miscalibrated at that end too. So the next review states the mechanism, not a confidence level: after this fix, is the manual path impossible, absurd, or merely inferior? Only the first two may carry a cost prediction.

This is the same principle ADR-0003 used to reject explicit session markers — nothing that depends on discipline gets to be the primary index. It governs tooling economics too, and this review did not apply it.

The most valuable output is not the $6 of 38 points recovered. It is that the loop made a falsifiable forecast, was wrong by 45×, and could say exactly which parts were wrong and why.