CB-WP-0004 T02: make task-done — close a task on measured numbers
Replaces the three hand-done steps of a task close (46 turns, $11.52 per CB-RES-0003): the heredoc flipping status in the workplan file, the hand-written hub call, and the hand-typed token counts. The third is the reason this task exists. Every update_task_status this repo produced carried estimated tokens_in/tokens_out — in a project whose central finding is that estimated token counts are worthless. task-done reads the measured figure from the transcripts, or refuses; there is no path through it that emits an estimate. cb-cost gains by_task_detail: cost, response count, model histogram and token components per task. task-done imports cb-cost rather than parsing its printed table, so the hub figure is not a copy that can drift from its source. The positive control found a real defect before the tool ran once. Attribution keyed on a bare T\d\d from the commit subject, so CB-WP-0002 T01, CB-WP-0003 T01 and CB-WP-0004 T01 shared a bucket: the self-test reported $12.10 for "T01" where the qualified figure is $2.33. That 5.2x overstatement would have been pushed to the hub as a *measured* number — the same fiction in a new form. task_label() now keys qualified subjects on the full id and leaves unqualified ones bare rather than retro-assigning them to a workplan. The pinned $93.15 benchmark is unchanged, so historical attribution was not disturbed. Fourth instance of trusted arithmetic: a number believed because a program produced it rather than a hand. Refusals, all exercised by --self-test: unknown id, typo'd id, already-done task, missing state_hub_task_id, no measured spend, and a status flip that produced no change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
3f1dbac164
commit
b52a9ec88a
5 changed files with 419 additions and 3 deletions
|
|
@ -110,6 +110,28 @@ CB-WP-0002 disproved.
|
|||
**Predicted:** those 46 turns → **~6**, **$9–11** recovered, and the hub
|
||||
stops holding estimates. Add `--self-test` per InnerLoop v1.1.
|
||||
|
||||
**Delivered.** `tools/task-done.py` + `make task-done T=<id>`. It refuses
|
||||
on an unknown id, a typo'd id, an already-done task, a task with no
|
||||
`state_hub_task_id`, and — the one that matters — **a task with no
|
||||
measured spend**, rather than reporting an estimate. `cb-cost` gained
|
||||
`by_task_detail` (cost, response count, model histogram, and token
|
||||
components per task), and `task-done` imports cb-cost rather than parsing
|
||||
its printed table, so the hub number is not a copy that can drift.
|
||||
|
||||
**The positive control found a real defect before the tool was used
|
||||
once.** Attribution keyed on a bare `T\d\d` from the commit subject, so
|
||||
`CB-WP-0002 T01`, `CB-WP-0003 T01` and `CB-WP-0004 T01` all landed in one
|
||||
bucket. The self-test reported **$12.10** for "T01"; the qualified figure
|
||||
is **$2.33** — a 5.2× overstatement that would have been pushed to the
|
||||
hub as a measured number, reproducing the fiction this task exists to
|
||||
end, in a new form. Fixed by `task_label()`: qualified subjects
|
||||
(`CB-WP-0004 T01`) key on the full id, unqualified ones stay bare and are
|
||||
never retro-assigned to a workplan. The pinned $93.15 benchmark is
|
||||
unchanged, confirming historical attribution was not disturbed.
|
||||
|
||||
That is the **fourth** instance of trusted arithmetic (TA) — a number
|
||||
believed because it was produced by a program rather than by hand.
|
||||
|
||||
## Task: `make status` — one-shot orientation
|
||||
|
||||
```task
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue