Scope cut first, on the maintainer's decision after a spend review: the project is 38% product / 62% loop-meta, cost per response is 2.9x worse than its best window, and INTENT stage 0 still lacks a CLI player and bots. CB-WP-0005 and CB-WP-0006 cost ~$74 — 31% of all spend — for zero measured efficiency gain. T02 and T04 are cancelled unstarted. T01: SH-1/SH-2/SH-3 now report over the window since the last commit, and the cumulative figure is retained but labelled "history, NOT the metric". The prediction held decisively — window 655,744 mean context against cumulative 255,307, a 2.6x gap against a 20% refutation threshold. A cumulative mean over 1,094 responses cannot detect a worsening trend because the history outvotes the present. T03: `make shape-budget`, modelled on CB-01/CB-02. Soft thresholds are the existing SessionShape targets; hard is 1.5x, set before the next measurement per §Step 4. Deliberately not in `make all` — failing the build on context would block committing, and committing is what closes the attribution window and is the natural point to compact, so a gate that blocks the remedy is a trap. It fires HARD on its first run: 656,574 against a 300,000 ceiling. InnerLoop v1.5 establishes the soft 25% meta budget. Workplans declare kind: product|meta|mixed and `make status` reports the share; mixed splits 50/50 and says so. Soft on purpose — a task already started may be finished, because stopping mid-task to satisfy a ratio wastes the work. What it forbids is opening new meta work above the line. A pass that exceeds it must say so in its evidence and name the product work displaced. First reading: 68% OVER, of $74.22 attributed. Product reads $0.00 because the only product workplan, CB-WP-0001, predates qualified task ids and its bare T## labels collide across passes — stated in the output rather than papered over. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
379 lines
15 KiB
Markdown
379 lines
15 KiB
Markdown
---
|
||
id: CB-WP-0003
|
||
kind: meta
|
||
title: "Harden the inner loop: executable rules, session economics, dead policy"
|
||
status: done
|
||
state_hub_workstream_id: "39d61dc0-870d-45c1-a595-bcf91f289dce"
|
||
---
|
||
|
||
# Purpose
|
||
|
||
A rigorous review of the loop after its first full pass (CB-WP-0001)
|
||
found error modes the loop does not catch and policy that costs
|
||
something while delivering nothing. The CI fixes were taken immediately
|
||
(commit `72c594e`); this workplan covers the rest.
|
||
|
||
The finding that framed it:
|
||
|
||
> **Writing a rule into `specs/InnerLoop.md` did not prevent the next
|
||
> instance of the failure it was written for. Making it a CI step did.**
|
||
|
||
**Revised 2026-07-31 from CB-WP-0002**, which was run first and changed
|
||
the premise. The original hypothesis — *a loop rule that cannot be
|
||
executed is not a rule* — survives but is now too coarse. CB-WP-0002
|
||
corrected one number four times ($248.46 → $92.21 → $92.87 → $93.32),
|
||
and **each correction came from a different mechanism**:
|
||
|
||
| mechanism | caught | cost |
|
||
|---|---|---|
|
||
| re-deriving instead of quoting | inherited double-count + single-model pricing | ~0 |
|
||
| adversarial review | target contradicted the survey's own finding | $1.11 |
|
||
| executable assertion (`--self-test`) | dedup invariant false in the subagent tree | ~0 |
|
||
|
||
**No mechanism found more than one.** So the goal is not to convert every
|
||
rule into a CI step; it is to know which class each mechanism catches and
|
||
to stop expecting one to cover another. Two specific limits are now
|
||
measured rather than assumed:
|
||
|
||
1. **A positive control cannot catch trusted arithmetic.** Both errors in
|
||
$248.46 were sums over data that genuinely existed. An assertion of the
|
||
form "did this harness do work?" answers yes, correctly, and the number
|
||
is still wrong.
|
||
2. **Adversarial review cannot catch a same-sample blind spot.** The dedup
|
||
invariant was verified on the main transcript by the survey *and*
|
||
independently by the reviewer; it is false in the 8-response subagent
|
||
tree neither examined. A reviewer's job is to re-derive the author's
|
||
claims, and re-deriving on the same sample reproduces the same gap.
|
||
|
||
CB-WP-0002 also falsified this workplan's own prediction (old T10): the
|
||
fifth error was **not** of the harness-does-nothing class. Two more
|
||
instances arrived, of two new classes. The retrospective task is rewritten
|
||
accordingly.
|
||
|
||
## Phase A — Make the rules executable
|
||
|
||
## Task: Audit every InnerLoop rule for enforceability
|
||
|
||
```task
|
||
id: CB-WP-0003-T01
|
||
status: done
|
||
priority: high
|
||
state_hub_task_id: "3d5d45fb-f931-407a-be71-d2d727279d5e"
|
||
```
|
||
|
||
Go through `specs/InnerLoop.md` **v1.1** rule by rule and classify each:
|
||
**executable** (a command fails when it is violated), **checkable**
|
||
(a human or agent can verify it in review, cheaply and objectively), or
|
||
**decorative** (neither). Record the table in
|
||
`history/YYMMDD-inner-loop-rule-audit.md`.
|
||
|
||
For each decorative rule, choose one of: make it executable, demote it
|
||
to guidance with that status stated, or delete it. The default is
|
||
delete — an unenforceable rule that reads like a requirement creates
|
||
false assurance, which is worse than silence.
|
||
|
||
**Revised:** classification is no longer enough. For each rule also record
|
||
**which failure class it catches**, drawn from the seven instances now on
|
||
record (four harness-does-nothing, two trusted-arithmetic, one
|
||
same-sample-blind-spot). A rule that catches a class no observed error
|
||
belongs to is a candidate for deletion even if it is perfectly executable.
|
||
Conversely, a class with no rule covering it is the gap worth the next
|
||
rule.
|
||
|
||
Known candidates to start from: whole-file loadability (~400 lines) is
|
||
mechanically checkable and unchecked — `specs/InnerLoop.md` is now near
|
||
that limit and should be measured first; "evidence or it didn't happen" is
|
||
executable via a committed-artifact check; the four v1.0 implementation
|
||
rules and the three v1.1 rules are currently prose.
|
||
|
||
## Task: Extend the positive-control gate beyond benchmarks
|
||
|
||
```task
|
||
id: CB-WP-0003-T02
|
||
status: done
|
||
priority: high
|
||
state_hub_task_id: "f861e67f-ed09-489d-963d-06697da9c08e"
|
||
```
|
||
|
||
**Partly delivered by CB-WP-0002.** Already done: the `--self-test`
|
||
contract is stated in `specs/InnerLoop.md` v1.1 §Step 5; `cb-cost` exposes
|
||
one with five assertions (AC-5..AC-9); `make cost` depends on
|
||
`make cost-test`; `cost-test` runs in CI and in `make all`. `cargo bench
|
||
-- --test` and cb-sim's empty-run check were already in place.
|
||
|
||
**Closed 2026-07-31 alongside T01** — both items below are done:
|
||
tools/rule-coverage.py and tools/dep-weight.py gained `--self-test`
|
||
(the former exposing a latent silent-pass on zero parsed rules), and
|
||
`make loop-lint` fails CI when any `tools/*.py` lacks the flag.
|
||
|
||
**Scope as written:**
|
||
|
||
- `tools/rule-coverage.py` and `tools/dep-weight.py` have positive-control
|
||
logic but no `--self-test` entry point, so nothing verifies the control
|
||
itself still works.
|
||
- No CI step fails when a *new* reporting tool ships without `--self-test`.
|
||
Prefer a mechanical check (enumerate `tools/*.py` + `tools/cb-*`, require
|
||
the flag) over a convention nobody can verify.
|
||
|
||
Carry forward from CB-WP-0002: a self-test that only exercises the happy
|
||
path is decorative. Each assertion must name a failure it detects, and the
|
||
strongest ones are regression pins on defects that actually occurred —
|
||
`cb-cost`'s AC-9 pins the exact `5, 5, 195` case that fooled first-wins
|
||
dedup.
|
||
|
||
## Task: Point adversarial review at measurement, and state its limits
|
||
|
||
```task
|
||
id: CB-WP-0003-T03
|
||
status: done
|
||
priority: medium
|
||
state_hub_task_id: "1365b35d-4539-489a-beff-c072221e4702"
|
||
```
|
||
|
||
The review step is specified for tier-L *research*. The errors in this
|
||
project are in *measurement and build configuration*, which the review
|
||
never formally targets.
|
||
|
||
**Now supported by two measurements, not an argument.** Reviews cost
|
||
$0.66 (CB-WP-0001) and $1.11 (CB-WP-0002) — ~1% of the pass each — and
|
||
each found approval-blocking defects. In CB-WP-0002 the review caught a
|
||
target that would have made the evidence file certify a broken collector.
|
||
|
||
Revise InnerLoop §Step 2 so the review target follows the risk: for a
|
||
capability whose claim rests on numbers, the review reads the harness and
|
||
the evidence file, not only the survey. Keep it one round; state
|
||
explicitly what the reviewer must attempt (reproduce the number, identify
|
||
what the harness would report if the work silently stopped).
|
||
|
||
**Added:** state what review *cannot* do, with the case attached. A
|
||
reviewer re-derives the author's claims and therefore inherits the
|
||
author's sampling. CB-WP-0002's dedup invariant was checked by both
|
||
parties on the main transcript and is false in the subagent tree. The
|
||
instruction that follows: **a reviewer must re-derive on a different
|
||
sample than the author used**, and where only one sample exists, say so
|
||
rather than reporting a clean verify.
|
||
|
||
## Phase B — Session economics
|
||
|
||
## Task: Measure session shape before prescribing it
|
||
|
||
```task
|
||
id: CB-WP-0003-T04
|
||
status: done
|
||
priority: high
|
||
state_hub_task_id: "382724e6-efd5-4ef8-a438-9d31a65dacdb"
|
||
```
|
||
|
||
**Premise revised — the original conclusion was wrong.** This task was
|
||
written to prescribe one task per session on the theory that cost ≈
|
||
turns × mean_context and mean_context grows with turns, making long
|
||
sessions quadratic. CB-WP-0002 measured it:
|
||
|
||
| segment | turns | mean context | total | $/turn |
|
||
|---|---|---|---|---|
|
||
| start → compact 1 | 136 | 304,178 | $62.19 | **$0.457** |
|
||
| compact 1 → compact 2 | 202 | 193,493 | $30.01 | **$0.149** |
|
||
|
||
`/compact` cut context 542,991 → 19,974 tokens (**27×**), and the 202
|
||
turns after it cost less than half the 136 before. Sessions are
|
||
**bounded**-quadratic: cost grows with context between compactions and
|
||
resets at each one. The failure mode is a long *uncompacted* session, not
|
||
a long one.
|
||
|
||
So the deliverable is no longer a policy derived from a theory. **Measure
|
||
the two remedies against each other first**, using `make cost`:
|
||
|
||
- cost of a compaction (the summarization call) plus the post-compact
|
||
ramp, versus
|
||
- cost of a fresh session, including the cold-start re-read of committed
|
||
artifacts that a fresh session must pay.
|
||
|
||
Then write `specs/SessionShape.md` with whichever wins, and the number.
|
||
Also carry over what is already measured and independent of that
|
||
question: 0 of 330 tool calls in CB-WP-0001 were batched into a
|
||
multi-call turn, and batching is free.
|
||
|
||
Every claim in the spec carries the measurement it rests on — this spec
|
||
exists because the numbers were surprising, and one of them has already
|
||
overturned the advice this task was created to give.
|
||
|
||
Unblocked: `tools/cb-cost.py` exists and `make cost` reports per-task
|
||
attribution.
|
||
|
||
## Task: Replace the dead token budget with a live cost budget
|
||
|
||
```task
|
||
id: CB-WP-0003-T05
|
||
status: done
|
||
priority: medium
|
||
state_hub_task_id: "26c920ee-2c09-4b45-a908-4d343532db09"
|
||
```
|
||
|
||
The global 8k soft / 10k hard per-task token budget was never referenced
|
||
or enforced during CB-WP-0001, and T08 exceeded it by orders of magnitude
|
||
with no signal. It is dead policy: it implies a control that does not
|
||
exist.
|
||
|
||
Replace it with a budget in USD, the unit `make cost` now makes
|
||
measurable, with a defined action on breach and a way to observe the
|
||
breach **at the time it happens** rather than in a retrospective. If no
|
||
such in-flight observation is possible, say so and delete the budget
|
||
rather than restating it.
|
||
|
||
**Calibration data now exists.** Per-task cost on CB-WP-0001 ranged
|
||
$1.29 (T09) to $21.02 (T08), with 32.5% of spend unattributed to any
|
||
task. Note the constraint this creates: a budget can only bind on the
|
||
67.5% that attribution reaches, and attribution is only computable
|
||
*after* the commit that closes a task. Any budget claiming to fire
|
||
mid-task must explain what it reads.
|
||
|
||
Unblocked: depends on CB-WP-0002, now complete.
|
||
|
||
## Phase C — Remove or fix the rest
|
||
|
||
## Task: Resolve the chaos roll
|
||
|
||
```task
|
||
id: CB-WP-0003-T06
|
||
status: done
|
||
priority: low
|
||
state_hub_task_id: "8f3a6147-54e8-48f9-8543-c13b0e0b0278"
|
||
```
|
||
|
||
The d10 tier roll has now been rolled across two workplans (CB-WP-0001:
|
||
9; CB-WP-0002: 2) and has **never** triggered the override, as expected
|
||
at 1-in-10. It remains untested while adding a step to every tier
|
||
decision.
|
||
|
||
Decide: raise the rate during a stated calibration period so the
|
||
mechanism produces evidence, or delete it. Keeping an unevaluated
|
||
mechanism at a rate that prevents its own evaluation is the one option
|
||
to reject.
|
||
|
||
## Task: Make retargeting a reviewed decision
|
||
|
||
```task
|
||
id: CB-WP-0003-T07
|
||
status: done
|
||
priority: medium
|
||
state_hub_task_id: "2f702821-7e64-4981-9885-82fa07f638aa"
|
||
```
|
||
|
||
AM-4's new targets were measured at 246,250 and set at 250,000 **in the
|
||
same commit, by the implementer, after seeing the number**. The
|
||
reasoning was recorded and is defensible, but the structure is exactly
|
||
what the loop exists to prevent.
|
||
|
||
Add to InnerLoop: a metric may not be retargeted in the commit that
|
||
measures it. A retarget is an ADR with the old target, the measurement
|
||
that motivated the change, and why the new target binds on future work
|
||
rather than merely passing present work. Apply retroactively to AM-4a
|
||
and AM-4b — either ratify them by ADR or change them.
|
||
|
||
**Sharpened by CB-WP-0002, which did this three times in one pass.**
|
||
AC-1 moved $92.21 → $92.87 → $93.32, each time in the commit that
|
||
discovered the discrepancy. Those retargets were *correct* — the target
|
||
was wrong and the instrument was right — which is precisely why a blanket
|
||
prohibition is the wrong rule. Distinguish the two cases:
|
||
|
||
- **target corrected because the instrument disproved it** — legitimate,
|
||
requires the instrument's output in the commit; and
|
||
- **target moved because the implementation missed it** — requires an ADR.
|
||
|
||
The rule must separate these without relying on the implementer's
|
||
self-report of which one it was.
|
||
|
||
## Task: Give provisional items an expiry
|
||
|
||
```task
|
||
id: CB-WP-0003-T08
|
||
status: done
|
||
priority: low
|
||
state_hub_task_id: "c2ee91ee-e551-48a7-b1a9-477c34c0690c"
|
||
```
|
||
|
||
Ten U-items plus GR-E02's "successes" are marked `provisional: true`
|
||
with no owner and no review date, so they can shape the kernel
|
||
indefinitely while looking handled.
|
||
|
||
Add an owner and a raised-on date to each provisional item, and make
|
||
`make coverage` report their age. Decide what happens when one goes
|
||
stale — the useful answer is probably that CI warns and the evidence
|
||
file must list them, not that the build breaks.
|
||
|
||
## Task: Strengthen the coverage gate beyond tag-counting
|
||
|
||
```task
|
||
id: CB-WP-0003-T09
|
||
status: done
|
||
priority: low
|
||
state_hub_task_id: "69ccc9ec-dc35-481c-8065-cef040f07f50"
|
||
```
|
||
|
||
`make coverage` compares rule IDs in the spec against `covers:` lists.
|
||
It proves no rule is unclaimed and no claimed rule is invented. It does
|
||
**not** prove a scenario exercises the rule it names, so 58/58 is weaker
|
||
evidence than it reads as.
|
||
|
||
Cheapest real strengthening to evaluate first: require every GR-id in a
|
||
`covers:` list to also appear in a doc comment in the aggregate, making
|
||
the spec→code→scenario chain mechanical rather than asserted. Consider
|
||
mutation-style checking (does removing the rule's code break the
|
||
scenario that claims it?) and cost it before adopting — `make cost` can
|
||
now price the evaluation itself.
|
||
|
||
## Task: Fix the price sheet's time-boxed rate before 2026-08-31
|
||
|
||
```task
|
||
id: CB-WP-0003-T11
|
||
status: done
|
||
priority: medium
|
||
state_hub_task_id: "809215d8-8a5c-4b10-9445-9bc6db9bda42"
|
||
```
|
||
|
||
**New, from CB-WP-0002.** `benchmarks/baselines/model-prices.toml`
|
||
encodes Sonnet at 3.00/15.00 with the intro rate 2.00/10.00 as a **TOML
|
||
comment**. A collector reading the sheet — which `cb-cost` does — silently
|
||
uses the wrong number. Today this costs $0.17 on a $93.32 pass (0.19%).
|
||
|
||
The deadline is real: **on 2026-08-31 the intro rate expires**, and the
|
||
comment and the data disagree in the opposite direction. Whoever reads the
|
||
sheet after that date gets a defensible number by accident rather than by
|
||
construction.
|
||
|
||
Give the schema a representation for a dated rate, and make the staleness
|
||
rule (§1a: refresh on price change or after 90 days) executable rather
|
||
than prose — it currently has no check at all, and every M-D2-CST verdict
|
||
inherits it. This is the same defect class the cost survey levelled at the
|
||
State Hub: a schema that cannot hold the fact it needs.
|
||
|
||
## Task: Retrospective
|
||
|
||
```task
|
||
id: CB-WP-0003-T10
|
||
status: done
|
||
priority: low
|
||
state_hub_task_id: "f43de208-92b1-4f0e-9236-a9d27f3ec451"
|
||
```
|
||
|
||
Revise `specs/InnerLoop.md` to v1.2 from this pass.
|
||
|
||
**The original question is already answered, and the answer was no.**
|
||
This task asked whether, after making rules executable, the next error
|
||
would still slip through and of what class. CB-WP-0002 ran first and
|
||
supplied two data points: the next errors were **not** harness-does-nothing.
|
||
They were trusted arithmetic over real data (which a positive control
|
||
passes) and a property verified on the large sample and assumed on the
|
||
small one (which adversarial review reproduces rather than catches).
|
||
|
||
So the question for this retrospective is the harder successor:
|
||
**is the set of mechanisms now complete, or is each new pass still finding
|
||
a new class?** Seven error instances are on record across three classes.
|
||
If this pass produces an eighth in a fourth class, the honest conclusion
|
||
is that class-by-class hardening does not converge, and the loop should
|
||
optimize for *cheap detection and correction* rather than for prevention —
|
||
which is a different design.
|
||
|
||
Record the count either way. A retrospective that reports only what was
|
||
fixed, and not whether the fixing is converging, is the same shape of
|
||
false assurance this workplan exists to remove.
|