CB-WP-0004 T02: make task-done — close a task on measured numbers

Replaces the three hand-done steps of a task close (46 turns, $11.52 per
CB-RES-0003): the heredoc flipping status in the workplan file, the
hand-written hub call, and the hand-typed token counts.

The third is the reason this task exists. Every update_task_status this
repo produced carried estimated tokens_in/tokens_out — in a project whose
central finding is that estimated token counts are worthless. task-done
reads the measured figure from the transcripts, or refuses; there is no
path through it that emits an estimate.

cb-cost gains by_task_detail: cost, response count, model histogram and
token components per task. task-done imports cb-cost rather than parsing
its printed table, so the hub figure is not a copy that can drift from
its source.

The positive control found a real defect before the tool ran once.
Attribution keyed on a bare T\d\d from the commit subject, so CB-WP-0002
T01, CB-WP-0003 T01 and CB-WP-0004 T01 shared a bucket: the self-test
reported $12.10 for "T01" where the qualified figure is $2.33. That 5.2x
overstatement would have been pushed to the hub as a *measured* number —
the same fiction in a new form. task_label() now keys qualified subjects
on the full id and leaves unqualified ones bare rather than
retro-assigning them to a workplan. The pinned $93.15 benchmark is
unchanged, so historical attribution was not disturbed.

Fourth instance of trusted arithmetic: a number believed because a
program produced it rather than a hand.

Refusals, all exercised by --self-test: unknown id, typo'd id,
already-done task, missing state_hub_task_id, no measured spend, and a
status flip that produced no change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
tegwick 2026-07-31 10:17:51 +02:00
parent 3f1dbac164
commit b52a9ec88a
5 changed files with 419 additions and 3 deletions

View file

@ -110,6 +110,28 @@ CB-WP-0002 disproved.
**Predicted:** those 46 turns → **~6**, **$911** recovered, and the hub
stops holding estimates. Add `--self-test` per InnerLoop v1.1.
**Delivered.** `tools/task-done.py` + `make task-done T=<id>`. It refuses
on an unknown id, a typo'd id, an already-done task, a task with no
`state_hub_task_id`, and — the one that matters — **a task with no
measured spend**, rather than reporting an estimate. `cb-cost` gained
`by_task_detail` (cost, response count, model histogram, and token
components per task), and `task-done` imports cb-cost rather than parsing
its printed table, so the hub number is not a copy that can drift.
**The positive control found a real defect before the tool was used
once.** Attribution keyed on a bare `T\d\d` from the commit subject, so
`CB-WP-0002 T01`, `CB-WP-0003 T01` and `CB-WP-0004 T01` all landed in one
bucket. The self-test reported **$12.10** for "T01"; the qualified figure
is **$2.33** — a 5.2× overstatement that would have been pushed to the
hub as a measured number, reproducing the fiction this task exists to
end, in a new form. Fixed by `task_label()`: qualified subjects
(`CB-WP-0004 T01`) key on the full id, unqualified ones stay bare and are
never retro-assigned to a workplan. The pinned $93.15 benchmark is
unchanged, confirming historical attribution was not disturbed.
That is the **fourth** instance of trusted arithmetic (TA) — a number
believed because it was produced by a program rather than by hand.
## Task: `make status` — one-shot orientation
```task