T03: point review at the harness, and state what review cannot catch
The review step targeted the survey. Every serious error in this project has been in measurement or build configuration, so the loop was adversarially reviewing the artifact cheapest to fix and leaving unreviewed the one where errors occur. Step 2 now routes by risk: when the claim rests on numbers, the reviewer gets the harness and the evidence file too, and must reproduce the number independently rather than read about it. The addition that matters more, because it was learned the hard way: a reviewer re-derives the author's claims and therefore inherits the author's SAMPLING. CB-WP-0002's dedup invariant was checked twice -- survey 206/206 groups, then the reviewer independently -- and both used the main transcript. It is false in the 8-response subagent tree neither looked at. Two independent verifications, one shared blind spot. Rule: the reviewer re-derives on a different sample than the author used, and where only one sample exists, says so rather than reporting a clean verify. Also recorded: what review demonstrably DOES do. $0.66 and $1.11 across two passes, ~1% of each, both finding approval-blocking defects. Cost is not a reason to skip it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
e21b9f4250
commit
4fd6322e17
2 changed files with 35 additions and 6 deletions
|
|
@ -124,7 +124,7 @@ dedup.
|
|||
|
||||
```task
|
||||
id: CB-WP-0003-T03
|
||||
status: todo
|
||||
status: done
|
||||
priority: medium
|
||||
state_hub_task_id: "1365b35d-4539-489a-beff-c072221e4702"
|
||||
```
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue