clay-borg/history/260731-mechanical-work-retrospective.md
tegwick 7b4d9e359f CB-WP-0004 T06: retrospective — what actually predicted the outcome
The question was whether converting agent work to deterministic tooling
recovers capacity or the work reappears elsewhere. Neither, cleanly:
tooling recovers capacity where it removes the possibility of doing the
work by hand, and recovers nothing where it merely offers a better
option.

Sorting the five candidates by that test reproduces the results exactly.
env-setup and task-done left no manual path — env-test fails the build if
the friction returns, and task-done refuses to emit an estimate, so the
old way means knowingly writing fiction. Both met. facts-check and
`make status` left grep and heredocs one keystroke away. Both missed.

This is the principle ADR-0003 already used to reject explicit session
markers — nothing that depends on discipline gets to be the primary
index. It governs tooling economics too, and CB-RES-0003 did not apply
it.

Stated confidence sorted the extremes and said nothing in the middle,
where every judgement call lives: both high met, both medium missed, and
the single low also met, so it was miscalibrated at that end too. Future
reviews state the mechanism instead — after this fix, is the manual path
impossible, absurd, or merely inferior? — and only the first two may
carry a cost prediction.

The confound is symmetric and does not explain the split: building
repo.py and task-done.py also required inspection and patching, and those
categories still went to zero. It inflates the size of the miss, not its
location.

The most valuable output is not the 6 points of 38 recovered. It is that
the loop made a falsifiable quantitative forecast, was wrong by 4-5x, and
could say exactly which parts were wrong and why.

CB-WP-0004 status -> done.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 10:28:35 +02:00

6.1 KiB
Raw Blame History

2026-07-31 — retrospective: does tooling recover capacity?

CB-WP-0004 T06. The pass built repo.py, task-done.py, status.py, facts.py, make env-test, make facts-check, cb-cost --since, and facts.toml; it measured itself in CB-EV-0003.

The question this task was given

Does converting agent work to deterministic tooling actually recover capacity, or does the work reappear elsewhere?

Neither, cleanly. It recovers capacity where the tool removes the possibility of doing the work by hand, and recovers nothing where the tool merely offers a better option. That distinction — not confidence, not category size — is what separated the hits from the misses, and it was not the distinction the review used.

Prediction error, per candidate

This is the first pass in the project to make a quantitative prediction before acting, so the error matters more than the saving.

# candidate stated confidence predicted measured (share of pass) verdict
1 environment setup high <10 turns, $1215 11.5% → 0.6%, 85 turns → 1 met
3+5 make task-done high ~6 turns, $911 8.5% → 0.0%, 46 turns → 0 met, exceeded
6 cb-cost flags low $2 3.4% → 1.0% met
2 fact registry medium $69 10.4% → 9.6% missed
4 make status medium $4, ~1 turn/session 5.1% → 21.2% missed, worse
aggregate 2530 points 6 points (38.2% → 32.4%) missed

The aggregate prediction was wrong by roughly 45×, in the same direction and for a related reason each time.

Did the stated confidence track reality?

Partly, and not usefully.

  • Both high-confidence candidates met their prediction. Both were correct.
  • Both medium-confidence candidates missed, one of them badly.
  • The single low-confidence candidate met its prediction, so confidence was miscalibrated at that end too — it under-predicted.

So confidence sorted the extremes correctly and said nothing useful in the middle, which is where every judgement call actually lives. The next review should stop stating a confidence level and state the mechanism instead — see below — because the mechanism was predictive and the confidence label was mostly a restatement of how sure the author felt.

The mechanism that actually predicted the outcome

Sorting the five candidates by whether they removed the work or offered an alternative reproduces the results exactly:

candidate after the fix, can the old way still happen? outcome
environment setup No. There is nothing to type; make env-test fails the build if the friction returns. met
make task-done No. The hand path is a heredoc plus a hub call; the tool refuses estimates, so using the old way means knowingly writing fiction. met
cb-cost flags Mostly no. The flags exist; writing a fresh analysis is strictly more work. met
fact registry Yes. facts-check gates drifted copies; it does not stop anyone patching markdown by hand. missed
make status Yes. grep still works, and is still one keystroke away. missed

Tooling recovers capacity when it makes the manual path impossible or absurd. It recovers nothing when it makes the manual path merely inferior — because at the moment of choosing, the manual path is specific to the question being asked and the tool is generic.

This is a design rule, not an observation about discipline. ADR-0003 already rejected explicit session markers on the grounds that nothing which depends on discipline gets to be the primary index. This pass found the same principle governs tooling economics, and the review did not apply it.

The confound, and why it does not dissolve the finding

This pass built the tools it measured, so the two misses are exactly the two whose tools were under construction: writing make status requires inspecting what it summarizes, and writing facts-check requires grepping every duplicated number in the repo. CB-EV-0003 states this before defending any number.

But the confound is symmetric and does not explain the split. Building repo.py and task-done.py also required inspection and patching, and those categories still went to zero — because after the fix there was nothing left to do by hand. The confound inflates the size of the miss; it does not explain why these two missed.

The clean test is still owed. A pass that uses these tools without building them is the only run that can confirm or refute the aggregate prediction, and it should carry the same three tests.

What this changes

  1. Reviews state the mechanism, not a confidence level. For each candidate: after this fix, is the manual path impossible, absurd, or merely inferior? Only the first two are predicted to recover anything; the third is proposed as a quality fix, with no cost claim.
  2. A "merely inferior" tool needs a gate to become a real saving. facts-check closed an error class without closing a cost category, and that is a legitimate result — but the workplan predicted dollars from it, and should not have.
  3. The relocation test earned its place and should stay. It came back negative (output per response fell 24%), and that negative is only credible because the test could have come back positive. It also forced the pass to disown the halved cost-per-response as a compaction effect, which was the most tempting wrong claim available.

Honest summary

The pass recovered 6 points of 38, against a predicted 2530, on a window it admits is confounded. Two of five categories are gone and gated; two are unchanged and now understood; one error class gained its first executable gate.

The most valuable output is not the saving. It is that the loop made a falsifiable quantitative forecast, was wrong by 45×, and could tell exactly which parts were wrong and why — which is the capability a loop needs before its cost projections are worth anything.