CB-WP-0007 T01+T03: window the metric, budget it, cap meta at 25%
Scope cut first, on the maintainer's decision after a spend review: the project is 38% product / 62% loop-meta, cost per response is 2.9x worse than its best window, and INTENT stage 0 still lacks a CLI player and bots. CB-WP-0005 and CB-WP-0006 cost ~$74 — 31% of all spend — for zero measured efficiency gain. T02 and T04 are cancelled unstarted. T01: SH-1/SH-2/SH-3 now report over the window since the last commit, and the cumulative figure is retained but labelled "history, NOT the metric". The prediction held decisively — window 655,744 mean context against cumulative 255,307, a 2.6x gap against a 20% refutation threshold. A cumulative mean over 1,094 responses cannot detect a worsening trend because the history outvotes the present. T03: `make shape-budget`, modelled on CB-01/CB-02. Soft thresholds are the existing SessionShape targets; hard is 1.5x, set before the next measurement per §Step 4. Deliberately not in `make all` — failing the build on context would block committing, and committing is what closes the attribution window and is the natural point to compact, so a gate that blocks the remedy is a trap. It fires HARD on its first run: 656,574 against a 300,000 ceiling. InnerLoop v1.5 establishes the soft 25% meta budget. Workplans declare kind: product|meta|mixed and `make status` reports the share; mixed splits 50/50 and says so. Soft on purpose — a task already started may be finished, because stopping mid-task to satisfy a ratio wastes the work. What it forbids is opening new meta work above the line. A pass that exceeds it must say so in its evidence and name the product work displaced. First reading: 68% OVER, of $74.22 attributed. Product reads $0.00 because the only product workplan, CB-WP-0001, predates qualified task ids and its bare T## labels collide across passes — stated in the output rather than papered over. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
98e09c4394
commit
b79ea9690d
11 changed files with 229 additions and 26 deletions
|
|
@ -1,5 +1,6 @@
|
|||
---
|
||||
id: CB-WP-0001
|
||||
kind: product
|
||||
title: "Establish the assimilate-and-surpass inner loop via the GROUND game kernel"
|
||||
status: done
|
||||
state_hub_workstream_id: "a1b434dc-b1c6-46b5-bbd9-80a4e6b7620f"
|
||||
|
|
|
|||
|
|
@ -1,5 +1,6 @@
|
|||
---
|
||||
id: CB-WP-0002
|
||||
kind: meta
|
||||
title: "Make agentic cost measurable, so D2 claims are falsifiable"
|
||||
status: done
|
||||
state_hub_workstream_id: "b7c22f69-fbe9-48df-9619-007db79ae338"
|
||||
|
|
|
|||
|
|
@ -1,5 +1,6 @@
|
|||
---
|
||||
id: CB-WP-0003
|
||||
kind: meta
|
||||
title: "Harden the inner loop: executable rules, session economics, dead policy"
|
||||
status: done
|
||||
state_hub_workstream_id: "39d61dc0-870d-45c1-a595-bcf91f289dce"
|
||||
|
|
|
|||
|
|
@ -1,5 +1,6 @@
|
|||
---
|
||||
id: CB-WP-0004
|
||||
kind: meta
|
||||
title: "Move mechanical turns off the token budget, and prove it worked"
|
||||
status: done
|
||||
state_hub_workstream_id: "6880ac78-d817-41b9-b267-f12ff9deea28"
|
||||
|
|
|
|||
|
|
@ -1,5 +1,6 @@
|
|||
---
|
||||
id: CB-WP-0005
|
||||
kind: meta
|
||||
title: "Make the instruments count assertions, then fix what they expose"
|
||||
status: done
|
||||
state_hub_workstream_id: "0b95a1e3-7780-43d0-81e9-072ef7978734"
|
||||
|
|
|
|||
|
|
@ -1,5 +1,6 @@
|
|||
---
|
||||
id: CB-WP-0006
|
||||
kind: mixed
|
||||
title: "Instrument the acceptance table, then implement what it exposes"
|
||||
status: done
|
||||
state_hub_workstream_id: "8a6327cc-fd5c-4e2c-a29b-b437c27d1e71"
|
||||
|
|
|
|||
|
|
@ -1,7 +1,8 @@
|
|||
---
|
||||
id: CB-WP-0007
|
||||
kind: meta
|
||||
title: "Make session shape measurable in the window that matters, then enforce it"
|
||||
status: proposed
|
||||
status: in_progress
|
||||
state_hub_workstream_id: "bee19b76-fb60-4dcf-94d0-9242b43e0e42"
|
||||
---
|
||||
|
||||
|
|
@ -33,6 +34,24 @@ Per InnerLoop §Step 4, no target moves in the commit that measures it —
|
|||
and CB-RES-0005 D3 states that up front, because all three targets are
|
||||
currently unmet by wide margins and the temptation is obvious.
|
||||
|
||||
## Scope cut, 2026-08-01 (maintainer decision)
|
||||
|
||||
A spend review before starting found the project **38% product / 62%
|
||||
loop-meta**, with cost per response degraded 2.9× from its best window and
|
||||
INTENT stage 0 still missing a CLI player and bots. CB-WP-0005 and
|
||||
CB-WP-0006 cost **~$74 — 31% of all spend — for zero measured efficiency
|
||||
gain** (their return was correctness of claims, which is real but is not
|
||||
optimization).
|
||||
|
||||
So this workplan is **cut to T01 and T03**, the two tasks that attack the
|
||||
2.9× regression directly. **T02 and T04 are cancelled**, not deferred —
|
||||
SH-4 trend reporting and the batching trial are more instrument work, and
|
||||
the instrument-failure taxonomy is not converging. T06 becomes the
|
||||
project-level retrospective the maintainer asked for.
|
||||
|
||||
**A soft 25% meta budget is established here** (T03), because the review's
|
||||
finding needs a standing number, not a memory.
|
||||
|
||||
## Phase A — measure the right window
|
||||
|
||||
## Task: window the session-shape metrics
|
||||
|
|
@ -58,11 +77,16 @@ target compares against.
|
|||
**Refuted if** it lands within 20% of the cumulative number, in which case
|
||||
the aggregation was not the problem and this pass should re-plan.
|
||||
|
||||
## Task: SH-4 — a windowed trend the instrument can see
|
||||
## Task: SH-4 — a windowed trend the instrument can see (CANCELLED)
|
||||
|
||||
> **Cancelled unstarted 2026-08-01.** More instrument work, against a
|
||||
> review finding that instrument work has stopped paying. The ceiling T03
|
||||
> adds catches a bad pass; a trend line would catch a slow slide, which is
|
||||
> a real gap — recorded as open, not built.
|
||||
|
||||
```task
|
||||
id: CB-WP-0007-T02
|
||||
status: todo
|
||||
status: cancel
|
||||
priority: medium
|
||||
state_hub_task_id: "1d5a32ef-4086-4021-b874-24ed2160c4dd"
|
||||
```
|
||||
|
|
@ -119,11 +143,17 @@ Carries `--self-test`. Its positive control is the one this project keeps
|
|||
needing: a budget that reports `ok` because it measured nothing must
|
||||
abort instead.
|
||||
|
||||
## Task: batch deliberately, and report what the rate reaches
|
||||
## Task: batch deliberately, and report what the rate reaches (CANCELLED)
|
||||
|
||||
> **Cancelled unstarted 2026-08-01.** `SessionShape.md` §4 already puts
|
||||
> the ceiling at **$2–4 on a $93 pass** — the cheapest of the three
|
||||
> metrics to move and the least valuable. Spending a task on it while
|
||||
> stage 0 lacks a CLI player is the misallocation the review found.
|
||||
> SH-3 remains measured, unmet at 0.0%, and unfalsified.
|
||||
|
||||
```task
|
||||
id: CB-WP-0007-T04
|
||||
status: todo
|
||||
status: cancel
|
||||
priority: medium
|
||||
state_hub_task_id: "c284db6f-2e16-40fb-be8b-83ae5391c058"
|
||||
```
|
||||
|
|
@ -154,7 +184,10 @@ priority: high
|
|||
state_hub_task_id: "2f78b272-e1f9-4530-bc8c-3a9d82833546"
|
||||
```
|
||||
|
||||
Commit `evidence/CB-EV-0006-session-shape.md`. Four tests, all reported:
|
||||
Commit `evidence/CB-EV-0006-session-shape.md`. **Reduced with the scope
|
||||
cut** — tests 1, 2 and 4 remain; test 3 (SH-3) is cancelled with T04.
|
||||
|
||||
Four tests, all reported:
|
||||
|
||||
1. **Does the windowed metric differ from cumulative?** Against T01's
|
||||
prediction; refuted if within 20%.
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue