Revise CB-WP-0003 from CB-WP-0002's findings
Some checks failed
ci / check (push) Failing after 4s

The framing hypothesis -- a rule that cannot be executed is not a rule
-- survives but was too coarse. CB-WP-0002 corrected one number four
times and each correction came from a DIFFERENT mechanism, none finding
more than one. Two limits are now measured rather than assumed:

  - a positive control cannot catch trusted arithmetic (both errors in
    $248.46 were sums over data that really existed)
  - adversarial review cannot catch a same-sample blind spot: survey and
    reviewer both verified dedup on the main transcript, and it is false
    in the subagent tree neither examined

Task-level revisions:

  T01  also classify each rule by WHICH failure class it catches; seven
       instances across three classes are now on record
  T02  scope reduced -- the --self-test contract, cb-cost's five
       assertions, make cost-test in CI are delivered; remaining work is
       rule-coverage/dep-weight plus a check that new tools carry one
  T03  now backed by two measurements ($0.66, $1.11 at ~1% of a pass),
       and must also state what review CANNOT do
  T04  premise overturned -- sessions are bounded-quadratic, not
       quadratic. /compact cut context 27x and cost/turn 3.1x. Measure
       compaction against a fresh session before prescribing either
  T05  unblocked, with calibration data ($1.29-$21.02/task) and the
       constraint that a budget binds only on the 67.5% attribution
       reaches, only after the closing commit
  T07  sharpened -- CB-WP-0002 retargeted three times, correctly, so a
       blanket prohibition is wrong; separate instrument-disproved-target
       from implementation-missed-target
  T10  its prediction is already falsified; the successor question is
       whether class-by-class hardening converges at all
  T11  NEW -- the price sheet cannot express a dated rate and becomes
       wrong on 2026-08-31

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
tegwick 2026-07-31 09:11:08 +02:00
parent 85b86e3805
commit 06b3df6eca

View file

@ -12,22 +12,42 @@ found error modes the loop does not catch and policy that costs
something while delivering nothing. The CI fixes were taken immediately something while delivering nothing. The CI fixes were taken immediately
(commit `72c594e`); this workplan covers the rest. (commit `72c594e`); this workplan covers the rest.
The finding that frames all of it: The finding that framed it:
> **Writing a rule into `specs/InnerLoop.md` did not prevent the next > **Writing a rule into `specs/InnerLoop.md` did not prevent the next
> instance of the failure it was written for. Making it a CI step did.** > instance of the failure it was written for. Making it a CI step did.**
InnerLoop v1.0 added the positive-control rule after two measurement **Revised 2026-07-31 from CB-WP-0002**, which was run first and changed
errors. A third occurred immediately afterwards (`default-features` the premise. The original hypothesis — *a loop rule that cannot be
silently ignored, build green, feature flag gating nothing), and a executed is not a rule* — survives but is now too coarse. CB-WP-0002
fourth was found *only* when the rule became a CI step — the committed corrected one number four times ($248.46 → $92.21 → $92.87 → $93.32),
replay benchmark had never run to completion. Four instances of one and **each correction came from a different mechanism**:
error class; the prose rule prevented none; the executable rule caught
one on its first run.
**Working hypothesis for this workplan: a loop rule that cannot be | mechanism | caught | cost |
executed is not a rule.** Every item below is either made executable or |---|---|---|
deleted. | re-deriving instead of quoting | inherited double-count + single-model pricing | ~0 |
| adversarial review | target contradicted the survey's own finding | $1.11 |
| executable assertion (`--self-test`) | dedup invariant false in the subagent tree | ~0 |
**No mechanism found more than one.** So the goal is not to convert every
rule into a CI step; it is to know which class each mechanism catches and
to stop expecting one to cover another. Two specific limits are now
measured rather than assumed:
1. **A positive control cannot catch trusted arithmetic.** Both errors in
$248.46 were sums over data that genuinely existed. An assertion of the
form "did this harness do work?" answers yes, correctly, and the number
is still wrong.
2. **Adversarial review cannot catch a same-sample blind spot.** The dedup
invariant was verified on the main transcript by the survey *and*
independently by the reviewer; it is false in the 8-response subagent
tree neither examined. A reviewer's job is to re-derive the author's
claims, and re-deriving on the same sample reproduces the same gap.
CB-WP-0002 also falsified this workplan's own prediction (old T10): the
fifth error was **not** of the harness-does-nothing class. Two more
instances arrived, of two new classes. The retrospective task is rewritten
accordingly.
## Phase A — Make the rules executable ## Phase A — Make the rules executable
@ -40,7 +60,7 @@ priority: high
state_hub_task_id: "3d5d45fb-f931-407a-be71-d2d727279d5e" state_hub_task_id: "3d5d45fb-f931-407a-be71-d2d727279d5e"
``` ```
Go through `specs/InnerLoop.md` v1.0 rule by rule and classify each: Go through `specs/InnerLoop.md` **v1.1** rule by rule and classify each:
**executable** (a command fails when it is violated), **checkable** **executable** (a command fails when it is violated), **checkable**
(a human or agent can verify it in review, cheaply and objectively), or (a human or agent can verify it in review, cheaply and objectively), or
**decorative** (neither). Record the table in **decorative** (neither). Record the table in
@ -51,10 +71,19 @@ to guidance with that status stated, or delete it. The default is
delete — an unenforceable rule that reads like a requirement creates delete — an unenforceable rule that reads like a requirement creates
false assurance, which is worse than silence. false assurance, which is worse than silence.
**Revised:** classification is no longer enough. For each rule also record
**which failure class it catches**, drawn from the seven instances now on
record (four harness-does-nothing, two trusted-arithmetic, one
same-sample-blind-spot). A rule that catches a class no observed error
belongs to is a candidate for deletion even if it is perfectly executable.
Conversely, a class with no rule covering it is the gap worth the next
rule.
Known candidates to start from: whole-file loadability (~400 lines) is Known candidates to start from: whole-file loadability (~400 lines) is
mechanically checkable and unchecked; "evidence or it didn't happen" is mechanically checkable and unchecked — `specs/InnerLoop.md` is now near
executable via a committed-artifact check; the four implementation rules that limit and should be measured first; "evidence or it didn't happen" is
added in v1.0 are currently prose. executable via a committed-artifact check; the four v1.0 implementation
rules and the three v1.1 rules are currently prose.
## Task: Extend the positive-control gate beyond benchmarks ## Task: Extend the positive-control gate beyond benchmarks
@ -65,19 +94,28 @@ priority: high
state_hub_task_id: "f861e67f-ed09-489d-963d-06697da9c08e" state_hub_task_id: "f861e67f-ed09-489d-963d-06697da9c08e"
``` ```
`cargo bench -- --test` now covers benchmarks. The same error class **Partly delivered by CB-WP-0002.** Already done: the `--self-test`
lives anywhere a harness can succeed while doing nothing. Extend contract is stated in `specs/InnerLoop.md` v1.1 §Step 5; `cb-cost` exposes
coverage to at least: the scenario runner (done — cb-sim fails on an one with five assertions (AC-5..AC-9); `make cost` depends on
empty run), the coverage tool, `dep-weight`, and any future collector `make cost-test`; `cost-test` runs in CI and in `make all`. `cargo bench
including CB-WP-0002's. -- --test` and cb-sim's empty-run check were already in place.
Deliver a stated contract in `specs/InnerLoop.md`: **every tool that **Remaining scope:**
reports a number states what it asserts to prove it did the work**, and
a CI step that fails when a reporting tool has no such assertion. Prefer
a mechanical check (e.g. each tool exposes `--self-test`) over a
convention nobody can verify.
## Task: Point adversarial review at measurement, not only research - `tools/rule-coverage.py` and `tools/dep-weight.py` have positive-control
logic but no `--self-test` entry point, so nothing verifies the control
itself still works.
- No CI step fails when a *new* reporting tool ships without `--self-test`.
Prefer a mechanical check (enumerate `tools/*.py` + `tools/cb-*`, require
the flag) over a convention nobody can verify.
Carry forward from CB-WP-0002: a self-test that only exercises the happy
path is decorative. Each assertion must name a failure it detects, and the
strongest ones are regression pins on defects that actually occurred —
`cb-cost`'s AC-9 pins the exact `5, 5, 195` case that fooled first-wins
dedup.
## Task: Point adversarial review at measurement, and state its limits
```task ```task
id: CB-WP-0003-T03 id: CB-WP-0003-T03
@ -86,21 +124,32 @@ priority: medium
state_hub_task_id: "1365b35d-4539-489a-beff-c072221e4702" state_hub_task_id: "1365b35d-4539-489a-beff-c072221e4702"
``` ```
The review step is specified for tier-L *research*. All four real errors The review step is specified for tier-L *research*. The errors in this
in this project were in *measurement and build configuration*, which the project are in *measurement and build configuration*, which the review
review never touches. We adversarially review the artifact that is never formally targets.
cheapest to fix and leave unreviewed the one where errors actually
occur. **Now supported by two measurements, not an argument.** Reviews cost
$0.66 (CB-WP-0001) and $1.11 (CB-WP-0002) — ~1% of the pass each — and
each found approval-blocking defects. In CB-WP-0002 the review caught a
target that would have made the evidence file certify a broken collector.
Revise InnerLoop §Step 2 so the review target follows the risk: for a Revise InnerLoop §Step 2 so the review target follows the risk: for a
capability whose claim rests on numbers, the review reads the harness capability whose claim rests on numbers, the review reads the harness and
and the evidence file, not only the survey. Keep it one round; state the evidence file, not only the survey. Keep it one round; state
explicitly what the reviewer must attempt (reproduce the number, explicitly what the reviewer must attempt (reproduce the number, identify
identify what the harness would report if the work silently stopped). what the harness would report if the work silently stopped).
**Added:** state what review *cannot* do, with the case attached. A
reviewer re-derives the author's claims and therefore inherits the
author's sampling. CB-WP-0002's dedup invariant was checked by both
parties on the main transcript and is false in the subagent tree. The
instruction that follows: **a reviewer must re-derive on a different
sample than the author used**, and where only one sample exists, say so
rather than reporting a clean verify.
## Phase B — Session economics ## Phase B — Session economics
## Task: Measure and specify session shape ## Task: Measure session shape before prescribing it
```task ```task
id: CB-WP-0003-T04 id: CB-WP-0003-T04
@ -109,26 +158,41 @@ priority: high
state_hub_task_id: "382724e6-efd5-4ef8-a438-9d31a65dacdb" state_hub_task_id: "382724e6-efd5-4ef8-a438-9d31a65dacdb"
``` ```
Measured from the CB-WP-0001 session: **592 assistant turns, mean **Premise revised — the original conclusion was wrong.** This task was
context 245,321 tokens, $0.245 per turn in cache reads alone**, and of written to prescribe one task per session on the theory that cost ≈
330 tool calls **none** were batched into a multi-call turn. Cache reads turns × mean_context and mean_context grows with turns, making long
were 58% of the $248.46 total; output was 12%. sessions quadratic. CB-WP-0002 measured it:
Cost is therefore approximately `turns × mean_context`, and mean_context | segment | turns | mean context | total | $/turn |
grows with turns — a long session is quadratic. Running T01T09 in one |---|---|---|---|---|
context cost ~$145 in cache reads; nine task-scoped sessions at ~40k | start → compact 1 | 136 | 304,178 | $62.19 | **$0.457** |
context each would plausibly cost ~$25. | compact 1 → compact 2 | 202 | 193,493 | $30.01 | **$0.149** |
Write `specs/SessionShape.md`: one task per session as the default, `/compact` cut context 542,991 → 19,974 tokens (**27×**), and the 202
what a fresh session must be able to load from committed artifacts to turns after it cost less than half the 136 before. Sessions are
start cold, when to compact versus start clean, and batching of **bounded**-quadratic: cost grows with context between compactions and
independent tool calls. Every claim in it carries the measurement it resets at each one. The failure mode is a long *uncompacted* session, not
rests on — this spec exists because the numbers were surprising, not a long one.
because the advice sounds sensible.
Depends on CB-WP-0002 for per-task attribution to validate the estimate; So the deliverable is no longer a policy derived from a theory. **Measure
the session-level numbers above are already sufficient to write the the two remedies against each other first**, using `make cost`:
policy.
- cost of a compaction (the summarization call) plus the post-compact
ramp, versus
- cost of a fresh session, including the cold-start re-read of committed
artifacts that a fresh session must pay.
Then write `specs/SessionShape.md` with whichever wins, and the number.
Also carry over what is already measured and independent of that
question: 0 of 330 tool calls in CB-WP-0001 were batched into a
multi-call turn, and batching is free.
Every claim in the spec carries the measurement it rests on — this spec
exists because the numbers were surprising, and one of them has already
overturned the advice this task was created to give.
Unblocked: `tools/cb-cost.py` exists and `make cost` reports per-task
attribution.
## Task: Replace the dead token budget with a live cost budget ## Task: Replace the dead token budget with a live cost budget
@ -139,16 +203,25 @@ priority: medium
state_hub_task_id: "26c920ee-2c09-4b45-a908-4d343532db09" state_hub_task_id: "26c920ee-2c09-4b45-a908-4d343532db09"
``` ```
The global 8k soft / 10k hard per-task token budget was never The global 8k soft / 10k hard per-task token budget was never referenced
referenced or enforced during CB-WP-0001, and T08 exceeded it by orders or enforced during CB-WP-0001, and T08 exceeded it by orders of magnitude
of magnitude with no signal. It is dead policy: it implies a control with no signal. It is dead policy: it implies a control that does not
that does not exist. exist.
Replace it with a budget expressed in the unit that CB-WP-0002 makes Replace it with a budget in USD, the unit `make cost` now makes
measurable, with a defined action on breach and a way to observe the measurable, with a defined action on breach and a way to observe the
breach at the time it happens rather than in a retrospective. If no such breach **at the time it happens** rather than in a retrospective. If no
observation is possible, say so and delete the budget rather than such in-flight observation is possible, say so and delete the budget
restating it. rather than restating it.
**Calibration data now exists.** Per-task cost on CB-WP-0001 ranged
$1.29 (T09) to $21.02 (T08), with 32.5% of spend unattributed to any
task. Note the constraint this creates: a budget can only bind on the
67.5% that attribution reaches, and attribution is only computable
*after* the commit that closes a task. Any budget claiming to fire
mid-task must explain what it reads.
Unblocked: depends on CB-WP-0002, now complete.
## Phase C — Remove or fix the rest ## Phase C — Remove or fix the rest
@ -161,10 +234,10 @@ priority: low
state_hub_task_id: "8f3a6147-54e8-48f9-8543-c13b0e0b0278" state_hub_task_id: "8f3a6147-54e8-48f9-8543-c13b0e0b0278"
``` ```
The d10 tier roll never fired across the whole pass — roughly 0.2 The d10 tier roll has now been rolled across two workplans (CB-WP-0001:
expected firings across 23 tier decisions — so it is untested, and at 9; CB-WP-0002: 2) and has **never** triggered the override, as expected
1-in-10 it will stay untested for many more passes while adding a step at 1-in-10. It remains untested while adding a step to every tier
to every decision. decision.
Decide: raise the rate during a stated calibration period so the Decide: raise the rate during a stated calibration period so the
mechanism produces evidence, or delete it. Keeping an unevaluated mechanism produces evidence, or delete it. Keeping an unevaluated
@ -191,6 +264,19 @@ that motivated the change, and why the new target binds on future work
rather than merely passing present work. Apply retroactively to AM-4a rather than merely passing present work. Apply retroactively to AM-4a
and AM-4b — either ratify them by ADR or change them. and AM-4b — either ratify them by ADR or change them.
**Sharpened by CB-WP-0002, which did this three times in one pass.**
AC-1 moved $92.21 → $92.87 → $93.32, each time in the commit that
discovered the discrepancy. Those retargets were *correct* — the target
was wrong and the instrument was right — which is precisely why a blanket
prohibition is the wrong rule. Distinguish the two cases:
- **target corrected because the instrument disproved it** — legitimate,
requires the instrument's output in the commit; and
- **target moved because the implementation missed it** — requires an ADR.
The rule must separate these without relying on the implementer's
self-report of which one it was.
## Task: Give provisional items an expiry ## Task: Give provisional items an expiry
```task ```task
@ -227,7 +313,33 @@ Cheapest real strengthening to evaluate first: require every GR-id in a
`covers:` list to also appear in a doc comment in the aggregate, making `covers:` list to also appear in a doc comment in the aggregate, making
the spec→code→scenario chain mechanical rather than asserted. Consider the spec→code→scenario chain mechanical rather than asserted. Consider
mutation-style checking (does removing the rule's code break the mutation-style checking (does removing the rule's code break the
scenario that claims it?) and cost it before adopting. scenario that claims it?) and cost it before adopting — `make cost` can
now price the evaluation itself.
## Task: Fix the price sheet's time-boxed rate before 2026-08-31
```task
id: CB-WP-0003-T11
status: todo
priority: medium
state_hub_task_id: ""
```
**New, from CB-WP-0002.** `benchmarks/baselines/model-prices.toml`
encodes Sonnet at 3.00/15.00 with the intro rate 2.00/10.00 as a **TOML
comment**. A collector reading the sheet — which `cb-cost` does — silently
uses the wrong number. Today this costs $0.17 on a $93.32 pass (0.19%).
The deadline is real: **on 2026-08-31 the intro rate expires**, and the
comment and the data disagree in the opposite direction. Whoever reads the
sheet after that date gets a defensible number by accident rather than by
construction.
Give the schema a representation for a dated rate, and make the staleness
rule (§1a: refresh on price change or after 90 days) executable rather
than prose — it currently has no check at all, and every M-D2-CST verdict
inherits it. This is the same defect class the cost survey levelled at the
State Hub: a schema that cannot hold the fact it needs.
## Task: Retrospective ## Task: Retrospective
@ -238,9 +350,24 @@ priority: low
state_hub_task_id: "f43de208-92b1-4f0e-9236-a9d27f3ec451" state_hub_task_id: "f43de208-92b1-4f0e-9236-a9d27f3ec451"
``` ```
Revise `specs/InnerLoop.md` to v1.1 from this pass. The question to Revise `specs/InnerLoop.md` to v1.2 from this pass.
answer honestly: after making rules executable, did the *next* error
still slip through, and if so, what class was it? Four instances of the **The original question is already answered, and the answer was no.**
harness-does-nothing class were found before a gate caught one. Record This task asked whether, after making rules executable, the next error
whether the gate holds, and what the fifth error — of whatever class — would still slip through and of what class. CB-WP-0002 ran first and
turns out to be. supplied two data points: the next errors were **not** harness-does-nothing.
They were trusted arithmetic over real data (which a positive control
passes) and a property verified on the large sample and assumed on the
small one (which adversarial review reproduces rather than catches).
So the question for this retrospective is the harder successor:
**is the set of mechanisms now complete, or is each new pass still finding
a new class?** Seven error instances are on record across three classes.
If this pass produces an eighth in a fourth class, the honest conclusion
is that class-by-class hardening does not converge, and the loop should
optimize for *cheap detection and correction* rather than for prevention —
which is a different design.
Record the count either way. A retrospective that reports only what was
fixed, and not whether the fixing is converging, is the same shape of
false assurance this workplan exists to remove.