cb-cost gains --since, so the baseline (--pin578dcbe, 662 responses, $135.60) and this pass (--since578dcbe, 83 responses, $7.82) are two disjoint windows over the same transcripts. Every verdict is on share of pass, since the windows differ 17x in size. Test 1 — did mechanical turns disappear? Partly. Mechanical share fell 38.2% -> 32.4%. Six points against a predicted 25-30; reported unmet, and no target was moved in this commit. environment setup 11.5% -> 0.6% (85 turns -> 1) met, decisively hub + workplan 8.5% -> 0.0% (46 turns -> 0) met text patching 10.4% -> 9.6% not met orientation 5.1% -> 21.2% not met, worse The confound is stated before any defence of the numbers: this is the pass that built the tools, and the two categories that missed are exactly the two whose tools were under construction. The clean test is the next pass, and it is carried forward rather than waived. Test 2 — did the work relocate? Not into prose. Output tokens per response fell 896 -> 681. Output's rising share of cost is a shrinking denominator, not more writing. Cost per response halved and the evidence refuses to claim it: that is compaction (mean context 232,982 -> 117,822), and attributing it to tooling would repeat CB-WP-0002's original error in a new direction. Test 3 — did quality hold? Yes, recorded as explicit judgment. make all green with two gates that did not exist before, and four findings surfaced this pass, three caught by controls written this pass — one of them a 5.2x attribution error that would have reached the hub as a measured number. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
7.2 KiB
CB-EV-0003: did converting mechanical turns to tooling recover anything?
research: CB-RES-0003
workplan: CB-WP-0004
instrument: make cost-mix (tools/cb-cost.py --composition)
commit: measured on the tree at CB-WP-0004 T05
Two windows over the same transcripts, cut at 578dcbe — the commit
that published the baseline:
| window | command | responses | pass cost |
|---|---|---|---|
| baseline (CB-WP-0001…0003 + the review) | cb-cost --pin 578dcbe |
662 | $135.60 |
| this pass (CB-WP-0004 T01–T04) | cb-cost --since 578dcbe |
83 | $7.82 |
The windows differ 17× in size, so every verdict below is on share of pass, per the normalization this task required. Absolute dollars are shown but decide nothing.
Test 1 — did the mechanical turns disappear?
| category | baseline $ | baseline share | this pass $ | this share | predicted | verdict |
|---|---|---|---|---|---|---|
| environment setup | $15.58 (85 turns) | 11.5% | $0.05 (1 turn) | 0.6% | <10 turns | met, decisively |
| hub task status + workplan edit | $11.52 (46 turns) | 8.5% | $0.00 (0 turns) | 0.0% | ~6 turns | met |
| ad-hoc text patching | $14.11 (76 turns) | 10.4% | $0.75 (8 turns) | 9.6% | $6–9 saved | not met |
| orientation / inspect | $6.87 (49 turns) | 5.1% | $1.66 (16 turns) | 21.2% | ~1 turn/session | not met — worse |
| ad-hoc transcript analysis | $4.56 (39 turns) | 3.4% | $0.08 (1 turn) | 1.0% | $2 saved | met |
| mechanical total | $51.76 (292) | 38.2% | $2.53 (26) | 32.4% | 25–30% recovered | not met |
Mechanical share fell 38.2% → 32.4%. That is 6 points, not the 25–30 points the review predicted. The prediction is reported unmet. Per InnerLoop §Step 4 no target was moved in the commit that measured it.
The two that worked
Environment setup went from the single largest category to effectively
zero — 85 turns to 1, an 18× drop in share. make env-test runs in
make all, so it cannot regress silently.
Hub and workplan closes went to exactly zero hand-written turns. Four
tasks were closed this pass, each with one make task-done, and each hub
event carries measured tokens rather than a typed estimate — the first
time that has been true in this repo.
The two that did not
Text patching barely moved (10.4% → 9.6%). facts-check gates
drifted copies but nothing removed the act of patching markdown: the
Makefile edits and the fact-tagging in T04 were both done with the same
heredocs the task was meant to retire. The gate closed the error class;
it did not close the cost category. Those are different claims and the
review conflated them.
Orientation got worse — 5.1% → 21.2% of pass. This is the largest
single miss and it deserves the plain reading first: make status was
built this pass and then barely used. 16 turns of grep/sed/ls
still went to reading the repo.
The confound, stated before any defence of the numbers
This pass is the pass that built the tools. Writing make status
requires inspecting exactly the artifacts make status summarizes;
writing facts-check requires grepping for every duplicated number in
the repo. The two categories that missed are precisely the two whose
tools were under construction.
So this measurement understates the effect, and no part of it should be read as the settled answer. The clean test is the next pass, which uses the tools without building them. That test is not optional: a prediction that can only be confirmed by a differently-shaped future run is not yet confirmed, and CB-WP-0005 should carry it.
What this pass does establish is the half that is not confounded: environment setup and task closes are gone, they were 20% of baseline pass cost, and their tools have gates that keep them gone.
Test 2 — did the work relocate?
The named failure mode: an agent that can no longer write heredocs simply writes more prose, and the pass costs the same.
| signal | baseline | this pass | reading |
|---|---|---|---|
| output tokens per response | 896 | 681 | fell 24% — no prose inflation |
| output share of cost | 13.2% | 18.1% | rose, but see below |
| cost per response | $0.205 | $0.094 | fell 54% — not attributable to this work |
| non-mechanical share | 61.8% | 67.6% | rose 6 pts, mirroring the mechanical fall |
Relocation into prose is not supported. Output per response fell. The output share of cost rose only because cache-read share fell (64.6% → 62.5%) as context shrank — the denominator moved, not the numerator.
The halved cost per response is a compaction effect, not a T01–T04
effect, and claiming it would be the most tempting error available
here. specs/SessionShape.md measures mean context at 232,982 tokens in
the baseline window against 117,822 in this one, because this session
was compacted twice. Attributing that to tooling would repeat CB-WP-0002's
original mistake in a new direction.
The 6-point rise in non-mechanical share is arithmetic, not relocation: if mechanical work leaves and the same judgment work remains, judgment's share rises by construction.
Test 3 — did quality hold?
Recorded as explicit judgment, because the loop has no metric for this and a cost number does not settle it.
make allgreen, now including two gates that did not exist:env-testandfacts-check.- Findings kept surfacing at the same rate. This pass produced four,
each caught by a control rather than by review:
loop-lintfailed onrepo.py— a new tool with a positive control but no--self-testentry point.task-done's self-test reported $12.10 for "T01" where the qualified figure is $2.33 — a bareT\d\dattribution key collided across three workplans. That 5.2× overstatement would have been pushed to the hub as a measured number. Fourth instance of trusted arithmetic.status.py's heading regex returned twenty paragraphs of the preceding task's prose as a "task heading".facts-checkfailed on a live fact tag inside its own documentation example inspecs/InnerLoop.md.
- No scenario, benchmark or coverage number regressed. The pinned benchmark is $93.15, unchanged by the attribution-key fix, which is the evidence that historical attribution was not disturbed.
Verdict: quality held. Three of the four findings were caught by controls written this pass, which is the pattern the loop is trying to buy.
Verdict
| claim | status |
|---|---|
| environment setup eliminated | confirmed |
| task closes eliminated, hub on measured numbers | confirmed |
| DFD has an executable gate | confirmed (falsified against a real artifact) |
| text patching reduced | not met |
| orientation reduced | not met — worse this pass |
| 25–30% of pass recovered | not met — 6 points, on a confounded window |
| work did not relocate into prose | confirmed (output/response fell 24%) |
The workplan said a saving that cannot be demonstrated in this table did not happen. Two of five categories are demonstrated; the aggregate prediction is not. Both are recorded as measured.