CB-WP-0007 T01+T03: window the metric, budget it, cap meta at 25%

Scope cut first, on the maintainer's decision after a spend review: the
project is 38% product / 62% loop-meta, cost per response is 2.9x worse
than its best window, and INTENT stage 0 still lacks a CLI player and
bots. CB-WP-0005 and CB-WP-0006 cost ~$74 — 31% of all spend — for zero
measured efficiency gain. T02 and T04 are cancelled unstarted.

T01: SH-1/SH-2/SH-3 now report over the window since the last commit, and
the cumulative figure is retained but labelled "history, NOT the metric".
The prediction held decisively — window 655,744 mean context against
cumulative 255,307, a 2.6x gap against a 20% refutation threshold. A
cumulative mean over 1,094 responses cannot detect a worsening trend
because the history outvotes the present.

T03: `make shape-budget`, modelled on CB-01/CB-02. Soft thresholds are the
existing SessionShape targets; hard is 1.5x, set before the next
measurement per §Step 4. Deliberately not in `make all` — failing the
build on context would block committing, and committing is what closes the
attribution window and is the natural point to compact, so a gate that
blocks the remedy is a trap. It fires HARD on its first run: 656,574
against a 300,000 ceiling.

InnerLoop v1.5 establishes the soft 25% meta budget. Workplans declare
kind: product|meta|mixed and `make status` reports the share; mixed splits
50/50 and says so. Soft on purpose — a task already started may be
finished, because stopping mid-task to satisfy a ratio wastes the work.
What it forbids is opening new meta work above the line. A pass that
exceeds it must say so in its evidence and name the product work
displaced.

First reading: 68% OVER, of $74.22 attributed. Product reads $0.00 because
the only product workplan, CB-WP-0001, predates qualified task ids and its
bare T## labels collide across passes — stated in the output rather than
papered over.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
tegwick 2026-08-01 14:12:07 +02:00
parent 98e09c4394
commit b79ea9690d
11 changed files with 229 additions and 26 deletions

View file

@ -24,7 +24,7 @@ TOOLS := $(REPO)/tools
# Every cargo recipe runs at the repo root; the shell does not persist cd. # Every cargo recipe runs at the repo root; the shell does not persist cd.
IN_REPO := cd $(REPO) && IN_REPO := cd $(REPO) &&
.PHONY: check test sim bench bench-test coverage dep-weight cost cost-test cost-pin cost-budget cost-mix loop-lint self-tests env-test task-done status facts-check facts-gen mutation-check size-metrics runtime-metrics build-time am6 replay-test loc all .PHONY: check test sim bench bench-test coverage dep-weight cost cost-test cost-pin cost-budget shape-budget cost-mix loop-lint self-tests env-test task-done status facts-check facts-gen mutation-check size-metrics runtime-metrics build-time am6 replay-test loc all
## fmt + clippy (deny warnings) + HashMap deny-lint ## fmt + clippy (deny warnings) + HashMap deny-lint
check: check:
@ -140,6 +140,13 @@ task-done:
@test -n "$(T)" || { echo "usage: make task-done T=CB-WP-0004-T02" >&2; exit 2; } @test -n "$(T)" || { echo "usage: make task-done T=CB-WP-0004-T02" >&2; exit 2; }
$(PY) $(TOOLS)/task-done.py $(T) $(ARGS) $(PY) $(TOOLS)/task-done.py $(T) $(ARGS)
# CB-WP-0007 T03 / InnerLoop v1.5: session shape for the window since the
# last commit. Deliberately NOT in `make all` — failing the build on
# context would block committing, and committing is the natural point to
# compact. A gate that blocks the remedy is a trap.
shape-budget: cost-test
$(PY) $(TOOLS)/cb-cost.py --shape-budget
# CB-01/CB-02: live spend since the last commit. # CB-01/CB-02: live spend since the last commit.
cost-budget: cost-test cost-budget: cost-test
$(PY) $(TOOLS)/cb-cost.py --budget $(PY) $(TOOLS)/cb-cost.py --budget

View file

@ -1,6 +1,31 @@
# The Inner Loop — Assimilate and Surpass # The Inner Loop — Assimilate and Surpass
Status: **v1.4** — corrected from CB-WP-0005 (assertion coverage) on Status: **v1.5** — corrected from CB-WP-0007 (session shape) on
2026-08-01. Change from v1.4: a **soft 25% meta budget** (below), after a
spend review found the project 38% product / 62% loop-meta with cost per
response degraded 2.9x from its best window.
> **Meta budget — soft, 25% of spend per pass.** Work on the loop's own
> instruments and process is capped at a quarter of a pass. Measured by
> `make status` from each workplan's `kind:` frontmatter
> (`product` | `meta` | `mixed`).
>
> **Soft on purpose.** A task already started may be finished — stopping
> mid-task to satisfy a ratio wastes the work and leaves the tree in a
> worse state than either finishing or never starting. What the budget
> forbids is *opening* new meta work above the line.
>
> A pass that exceeds it **says so in its evidence file and names the
> product work displaced**. That is the whole enforcement: this is a
> reporting budget, not a gate, for the same reason the session-shape
> budget is (CB-RES-0005 §4) — it constrains judgment, not artifacts.
>
> *(v1.5, from the CB-WP-0007 spend review: CB-WP-0005 and CB-WP-0006 cost
> ~$74, 31% of all spend, for zero measured efficiency gain. Their return
> was correctness of claims, which is real and is not optimization. The
> budget exists so that distinction has to be made out loud.)*
v1.4 — corrected from CB-WP-0005 (assertion coverage) on
2026-07-31. Change from v1.3: where a claim rests on numbers, the 2026-07-31. Change from v1.3: where a claim rests on numbers, the
adversarial reviewer must read the assertion behind each quoted number and adversarial reviewer must read the assertion behind each quoted number and
**mutate it** — re-running the command that prints a number is not **mutate it** — re-running the command that prints a number is not

View file

@ -396,7 +396,8 @@ def collect(slug, pin_ref=None, since_ref=None):
for r in responses: for r in responses:
r["cost"] = price_of(prices, r["model"], r["toks"], r["timestamp"]) r["cost"] = price_of(prices, r["model"], r["toks"], r["timestamp"])
attribute(responses, commit_index(pin)) commits = commit_index(pin)
attribute(responses, commits)
by_task = collections.defaultdict(float) by_task = collections.defaultdict(float)
# T02: per-task token detail, so a task close can report *measured* # T02: per-task token detail, so a task close can report *measured*
@ -454,9 +455,28 @@ def collect(slug, pin_ref=None, since_ref=None):
mech = [r for r in responses mech = [r for r in responses
if set(r.get("categories") or []) & MECHANICAL] if set(r.get("categories") or []) & MECHANICAL]
# The window the SH-* targets compare against: responses since the
# last commit when unpinned, or the whole pinned range when pinned
# (a pin is already a window someone chose deliberately).
if pin or since:
window_responses, window_label = responses, "the pinned/--since range"
else:
last = commits[-1][0] if commits else ""
window_responses = [r for r in responses if last and r["timestamp"] > last]
window_label = "since the last commit"
sub = sum(r["cost"] or 0 for r in responses if r["subagent"]) sub = sum(r["cost"] or 0 for r in responses if r["subagent"])
return { return {
"session_shape": session_shape(responses), "session_shape": session_shape(responses),
# CB-WP-0007 T01. SH-1/SH-2 as a cumulative mean over every
# response ever recorded cannot detect a worsening trend: the
# history outvotes the present. CB-WP-0006 ran at 503,464 mean
# context and the cumulative figure still printed 206,952.
# Both are reported; only the window is the metric.
"session_shape_window": session_shape(window_responses)
if window_responses else None,
"window_responses": len(window_responses),
"window_label": window_label,
"tool_mix": { "tool_mix": {
"turns": dict(mix_turns), "turns": dict(mix_turns),
"cost": dict(mix_cost), "cost": dict(mix_cost),
@ -550,16 +570,33 @@ def render(rep, by_task=False, composition=False):
f"{mix['mechanical_turns']:>5} turns ${mix['mechanical_cost']:>8,.2f}" f"{mix['mechanical_turns']:>5} turns ${mix['mechanical_cost']:>8,.2f}"
f" = {100*mix['mechanical_cost']/(rep['total'] or 1):.0f}% of pass") f" = {100*mix['mechanical_cost']/(rep['total'] or 1):.0f}% of pass")
def _shape_lines(sh, indent=" "):
print(f"{indent}SH-1 mean context {sh['SH-1_mean_context']:>12,.0f} tok "
f"[{'ok ' if sh['SH-1_mean_context']<=200_000 else 'FAIL'} target 200,000]")
print(f"{indent}SH-2 p90 context {sh['SH-2_p90_context']:>12,.0f} tok "
f"[{'ok ' if sh['SH-2_p90_context']<=300_000 else 'FAIL'} target 300,000]")
print(f"{indent}SH-3 batching rate {100*sh['SH-3_batching_rate']:>11.1f}% "
f"[{'ok ' if sh['SH-3_batching_rate']>=0.20 else 'FAIL'} target 20.0%]")
print(f"{indent} {sh['tool_calls']} tool calls in "
f"{sh['responses_with_tools']} responses; "
f"{sh['calls_in_batched_turns']} in batched turns")
win = rep.get("session_shape_window")
print(f"\n session shape (specs/SessionShape.md) — window: "
f"{rep.get('window_label', '?')}")
if win:
print(f" THE METRIC — {rep['window_responses']} response(s)")
_shape_lines(win)
else:
print(" THE METRIC — no responses in the window "
"(nothing since the last commit)")
sh = rep["session_shape"] sh = rep["session_shape"]
print("\n session shape (specs/SessionShape.md)") print(f"\n history, {rep['responses']} responses — context only, NOT the "
print(f" SH-1 mean context {sh['SH-1_mean_context']:>12,.0f} tok " f"metric.")
f"[{'ok ' if sh['SH-1_mean_context']<=200_000 else 'FAIL'} target 200,000]") print(" A cumulative mean cannot detect a worsening trend; the history")
print(f" SH-2 p90 context {sh['SH-2_p90_context']:>12,.0f} tok " print(" outvotes the present (CB-RES-0005 §1).")
f"[{'ok ' if sh['SH-2_p90_context']<=300_000 else 'FAIL'} target 300,000]") _shape_lines(sh)
print(f" SH-3 batching rate {100*sh['SH-3_batching_rate']:>11.1f}% "
f"[{'ok ' if sh['SH-3_batching_rate']>=0.20 else 'FAIL'} target 20.0%]")
print(f" {sh['tool_calls']} tool calls in {sh['responses_with_tools']} responses; "
f"{sh['calls_in_batched_turns']} in batched turns")
if rep["unpriced"]: if rep["unpriced"]:
print(f"\n UNPRICED ({len(rep['unpriced'])} responses, model not in sheet):") print(f"\n UNPRICED ({len(rep['unpriced'])} responses, model not in sheet):")
@ -690,6 +727,63 @@ def self_test():
return 0 if ok else 1 return 0 if ok else 1
# CB-WP-0007 T03 / CB-RES-0005 D2. Soft = the SessionShape target; hard is
# 1.5x, set here BEFORE the next measurement per InnerLoop §Step 4.
# Reported, never in `make all`: failing the build on context would block
# committing, and committing is what closes the attribution window and is
# the natural point to compact. A gate that blocks the remedy is a trap.
SHAPE_SOFT = {"SH-1": 200_000, "SH-2": 300_000}
SHAPE_HARD = {"SH-1": 300_000, "SH-2": 450_000}
def shape_budget(slug):
"""SH-* for the window since the last commit, with soft/hard verdicts."""
try:
rep = collect(slug, None)
except Abort as e:
print(f"ABORT — {e}", file=sys.stderr)
return 1
win = rep.get("session_shape_window")
head = subprocess.run(
["git", "-C", REPO, "log", "-1", "--format=%h %s"],
capture_output=True, text=True, check=True).stdout.strip()
print("session-shape budget — the window since the last commit")
print(f" last commit {head}")
# Positive control: a budget that reports ok because it measured
# nothing is the harness-does-nothing shape, in a reporting tool.
if not win or rep["window_responses"] == 0:
print(" no responses since the last commit — nothing to report",
file=sys.stderr)
return 1
print(f" window {rep['window_responses']} response(s)")
breach = 0
for key, label, val in (
("SH-1", "mean context", win["SH-1_mean_context"]),
("SH-2", "p90 context ", win["SH-2_p90_context"]),
):
soft, hard = SHAPE_SOFT[key], SHAPE_HARD[key]
mark = "ok " if val <= soft else ("SOFT" if val <= hard else "HARD")
print(f" {key} {label} {val:>10,.0f} tok [{mark}] "
f"soft {soft:,} / hard {hard:,}")
breach = max(breach, 0 if val <= soft else (1 if val <= hard else 2))
rate = win["SH-3_batching_rate"]
print(f" SH-3 batching {100*rate:>9.1f}% "
f"[{'ok ' if rate >= 0.20 else 'SOFT'}] floor 20.0%")
if breach >= 2:
print("\n HARD — compact before continuing. Context this size costs "
"~10x per turn\n against a compacted session "
"(specs/SessionShape.md §2).", file=sys.stderr)
return 1
if breach == 1:
print("\n SOFT — over target. Compaction is the remedy and it is free.")
else:
print("\n within budget")
return 0
def budget(slug, soft, hard): def budget(slug, soft, hard):
"""Live cost budget (specs/CostAccounting.md §7). """Live cost budget (specs/CostAccounting.md §7).
@ -734,6 +828,8 @@ def main():
ap.add_argument("--composition", action="store_true") ap.add_argument("--composition", action="store_true")
ap.add_argument("--session-shape", action="store_true", ap.add_argument("--session-shape", action="store_true",
help="SH-1..SH-3 (always shown in the default report)") help="SH-1..SH-3 (always shown in the default report)")
ap.add_argument("--shape-budget", action="store_true",
help="SH-1/SH-2/SH-3 for the window since the last commit")
ap.add_argument("--budget", action="store_true", ap.add_argument("--budget", action="store_true",
help="CB-01/CB-02: spend since the last commit, live") help="CB-01/CB-02: spend since the last commit, live")
ap.add_argument("--soft", type=float, default=10.00) ap.add_argument("--soft", type=float, default=10.00)
@ -745,6 +841,8 @@ def main():
if args.self_test: if args.self_test:
return self_test() return self_test()
if args.shape_budget:
return shape_budget(args.slug)
if args.budget: if args.budget:
return budget(args.slug, args.soft, args.hard) return budget(args.slug, args.soft, args.hard)

View file

@ -48,6 +48,7 @@ def parse_workplan(path):
m = re.search(rf'^{key}:\s*"?(.*?)"?\s*$', text, re.M) m = re.search(rf'^{key}:\s*"?(.*?)"?\s*$', text, re.M)
return m.group(1) if m else None return m.group(1) if m else None
kind = fm("kind") or "untagged"
tasks = [] tasks = []
for m in re.finditer(r"```task\n(.*?)```", text, re.S): for m in re.finditer(r"```task\n(.*?)```", text, re.S):
body = m.group(1) body = m.group(1)
@ -57,7 +58,7 @@ def parse_workplan(path):
return mm.group(1) if mm else None return mm.group(1) if mm else None
tasks.append((f("id"), f("status"), f("priority"))) tasks.append((f("id"), f("status"), f("priority")))
return fm("id"), fm("title"), fm("status"), tasks return fm("id"), fm("title"), fm("status"), tasks, kind
def workplans(): def workplans():
@ -69,7 +70,7 @@ def workplans():
def next_task(plans): def next_task(plans):
"""First todo task of the first non-done workplan, with its heading.""" """First todo task of the first non-done workplan, with its heading."""
for wid, _title, status, tasks in plans: for wid, _title, status, tasks, _kind in plans:
if status == "done": if status == "done":
continue continue
for tid, tstatus, prio in tasks: for tid, tstatus, prio in tasks:
@ -130,7 +131,7 @@ def report():
+ (f" ({len(dirty)} uncommitted)" if dirty else " (clean)")) + (f" ({len(dirty)} uncommitted)" if dirty else " (clean)"))
print("\n workplans") print("\n workplans")
for wid, title, status, tasks in plans: for wid, title, status, tasks, _kind in plans:
done = sum(1 for _, s, _ in tasks if s == "done") done = sum(1 for _, s, _ in tasks if s == "done")
if status == "done" and done == len(tasks): if status == "done" and done == len(tasks):
print(f" {wid} {status:<12} {done}/{len(tasks)}") print(f" {wid} {status:<12} {done}/{len(tasks)}")
@ -161,6 +162,34 @@ def report():
except Exception as e: except Exception as e:
print(f" UNAVAILABLE — {e}") print(f" UNAVAILABLE — {e}")
# InnerLoop v1.5 — soft 25% meta budget. Reported, never gated.
try:
_, _, mod = cost_lines()
rep = mod.collect("-home-worsch-clay-borg", None)
detail = rep["by_task_detail"]
by_kind = {"product": 0.0, "meta": 0.0, "mixed": 0.0}
attributed = 0.0
for wid, _t, _s, tasks, kind in plans:
for tid, _ts, _p in tasks:
d = detail.get(tid)
if d and kind in by_kind:
by_kind[kind] += d["cost"]
attributed += d["cost"]
if attributed > 0:
# `mixed` splits evenly; stated rather than hidden.
meta = by_kind["meta"] + by_kind["mixed"] / 2
share = 100 * meta / attributed
mark = "ok " if share <= 25 else "OVER"
print(f"\n meta budget [{mark}] {share:.0f}% of "
f"${attributed:,.2f} attributed (soft 25%, InnerLoop v1.5)")
print(f" product ${by_kind['product']:,.2f} "
f"meta ${by_kind['meta']:,.2f} "
f"mixed ${by_kind['mixed']:,.2f} (split 50/50)")
print(" NOTE: attributed tasks only — early workplans used bare")
print(" T## ids that collide across passes and are excluded.")
except Exception as e:
print(f"\n meta budget UNAVAILABLE — {e}")
print("\n fast gates") print("\n fast gates")
code, out = run_tool("loop-lint.py") code, out = run_tool("loop-lint.py")
last = out.splitlines()[-1] if out else "?" last = out.splitlines()[-1] if out else "?"
@ -209,18 +238,18 @@ def self_test():
fh.write(sample) fh.write(sample)
tmp = fh.name tmp = fh.name
try: try:
wid, title, status, tasks = parse_workplan(tmp) wid, title, status, tasks, _kind = parse_workplan(tmp)
check("workplan frontmatter parses", (wid, title, status) check("workplan frontmatter parses", (wid, title, status)
== ("CB-WP-0009", "Sample", "in_progress"), f"{wid} {status}") == ("CB-WP-0009", "Sample", "in_progress"), f"{wid} {status}")
check("both task blocks parse with status and priority", check("both task blocks parse with status and priority",
tasks == [("CB-WP-0009-T01", "done", "high"), tasks == [("CB-WP-0009-T01", "done", "high"),
("CB-WP-0009-T02", "todo", "medium")]) ("CB-WP-0009-T02", "todo", "medium")])
check("next task is the first todo, not the first task", check("next task is the first todo, not the first task",
next_task([(wid, title, status, tasks)]) next_task([(wid, title, status, tasks, "meta")])
== ("CB-WP-0009", "CB-WP-0009-T02", "medium")) == ("CB-WP-0009", "CB-WP-0009-T02", "medium"))
check("a fully-done workplan yields no next task", check("a fully-done workplan yields no next task",
next_task([("X", "t", "done", next_task([("X", "t", "done",
[("X-T01", "done", "high")])]) is None) [("X-T01", "done", "high")], "meta")]) is None)
finally: finally:
os.unlink(tmp) os.unlink(tmp)
@ -232,9 +261,14 @@ def self_test():
sum(len(t) for *_, t in plans) >= 20, sum(len(t) for *_, t in plans) >= 20,
f"{sum(len(t) for *_, t in plans)} task(s)") f"{sum(len(t) for *_, t in plans)} task(s)")
check("every real workplan has an id and a status", check("every real workplan has an id and a status",
all(w and s for w, _, s, _ in plans)) all(w and s for w, _, s, _, _ in plans))
# InnerLoop v1.5: the meta budget needs every workplan classified, or
# the ratio is computed over an unknown denominator.
check("every workplan declares kind: product|meta|mixed",
all(k in ("product", "meta", "mixed") for *_, k in plans),
", ".join(f"{w}={k}" for w, *_, k in plans))
check("every real task has an id and a status", check("every real task has an id and a status",
all(i and s for *_, ts in plans for i, s, _ in ts)) all(i and s for _, _, _, ts, _ in plans for i, s, _ in ts))
# A heading must be one line. The first version used a dot-all match # A heading must be one line. The first version used a dot-all match
# and returned the prose of the preceding task as the heading. # and returned the prose of the preceding task as the heading.

View file

@ -1,5 +1,6 @@
--- ---
id: CB-WP-0001 id: CB-WP-0001
kind: product
title: "Establish the assimilate-and-surpass inner loop via the GROUND game kernel" title: "Establish the assimilate-and-surpass inner loop via the GROUND game kernel"
status: done status: done
state_hub_workstream_id: "a1b434dc-b1c6-46b5-bbd9-80a4e6b7620f" state_hub_workstream_id: "a1b434dc-b1c6-46b5-bbd9-80a4e6b7620f"

View file

@ -1,5 +1,6 @@
--- ---
id: CB-WP-0002 id: CB-WP-0002
kind: meta
title: "Make agentic cost measurable, so D2 claims are falsifiable" title: "Make agentic cost measurable, so D2 claims are falsifiable"
status: done status: done
state_hub_workstream_id: "b7c22f69-fbe9-48df-9619-007db79ae338" state_hub_workstream_id: "b7c22f69-fbe9-48df-9619-007db79ae338"

View file

@ -1,5 +1,6 @@
--- ---
id: CB-WP-0003 id: CB-WP-0003
kind: meta
title: "Harden the inner loop: executable rules, session economics, dead policy" title: "Harden the inner loop: executable rules, session economics, dead policy"
status: done status: done
state_hub_workstream_id: "39d61dc0-870d-45c1-a595-bcf91f289dce" state_hub_workstream_id: "39d61dc0-870d-45c1-a595-bcf91f289dce"

View file

@ -1,5 +1,6 @@
--- ---
id: CB-WP-0004 id: CB-WP-0004
kind: meta
title: "Move mechanical turns off the token budget, and prove it worked" title: "Move mechanical turns off the token budget, and prove it worked"
status: done status: done
state_hub_workstream_id: "6880ac78-d817-41b9-b267-f12ff9deea28" state_hub_workstream_id: "6880ac78-d817-41b9-b267-f12ff9deea28"

View file

@ -1,5 +1,6 @@
--- ---
id: CB-WP-0005 id: CB-WP-0005
kind: meta
title: "Make the instruments count assertions, then fix what they expose" title: "Make the instruments count assertions, then fix what they expose"
status: done status: done
state_hub_workstream_id: "0b95a1e3-7780-43d0-81e9-072ef7978734" state_hub_workstream_id: "0b95a1e3-7780-43d0-81e9-072ef7978734"

View file

@ -1,5 +1,6 @@
--- ---
id: CB-WP-0006 id: CB-WP-0006
kind: mixed
title: "Instrument the acceptance table, then implement what it exposes" title: "Instrument the acceptance table, then implement what it exposes"
status: done status: done
state_hub_workstream_id: "8a6327cc-fd5c-4e2c-a29b-b437c27d1e71" state_hub_workstream_id: "8a6327cc-fd5c-4e2c-a29b-b437c27d1e71"

View file

@ -1,7 +1,8 @@
--- ---
id: CB-WP-0007 id: CB-WP-0007
kind: meta
title: "Make session shape measurable in the window that matters, then enforce it" title: "Make session shape measurable in the window that matters, then enforce it"
status: proposed status: in_progress
state_hub_workstream_id: "bee19b76-fb60-4dcf-94d0-9242b43e0e42" state_hub_workstream_id: "bee19b76-fb60-4dcf-94d0-9242b43e0e42"
--- ---
@ -33,6 +34,24 @@ Per InnerLoop §Step 4, no target moves in the commit that measures it —
and CB-RES-0005 D3 states that up front, because all three targets are and CB-RES-0005 D3 states that up front, because all three targets are
currently unmet by wide margins and the temptation is obvious. currently unmet by wide margins and the temptation is obvious.
## Scope cut, 2026-08-01 (maintainer decision)
A spend review before starting found the project **38% product / 62%
loop-meta**, with cost per response degraded 2.9× from its best window and
INTENT stage 0 still missing a CLI player and bots. CB-WP-0005 and
CB-WP-0006 cost **~$74 — 31% of all spend — for zero measured efficiency
gain** (their return was correctness of claims, which is real but is not
optimization).
So this workplan is **cut to T01 and T03**, the two tasks that attack the
2.9× regression directly. **T02 and T04 are cancelled**, not deferred —
SH-4 trend reporting and the batching trial are more instrument work, and
the instrument-failure taxonomy is not converging. T06 becomes the
project-level retrospective the maintainer asked for.
**A soft 25% meta budget is established here** (T03), because the review's
finding needs a standing number, not a memory.
## Phase A — measure the right window ## Phase A — measure the right window
## Task: window the session-shape metrics ## Task: window the session-shape metrics
@ -58,11 +77,16 @@ target compares against.
**Refuted if** it lands within 20% of the cumulative number, in which case **Refuted if** it lands within 20% of the cumulative number, in which case
the aggregation was not the problem and this pass should re-plan. the aggregation was not the problem and this pass should re-plan.
## Task: SH-4 — a windowed trend the instrument can see ## Task: SH-4 — a windowed trend the instrument can see (CANCELLED)
> **Cancelled unstarted 2026-08-01.** More instrument work, against a
> review finding that instrument work has stopped paying. The ceiling T03
> adds catches a bad pass; a trend line would catch a slow slide, which is
> a real gap — recorded as open, not built.
```task ```task
id: CB-WP-0007-T02 id: CB-WP-0007-T02
status: todo status: cancel
priority: medium priority: medium
state_hub_task_id: "1d5a32ef-4086-4021-b874-24ed2160c4dd" state_hub_task_id: "1d5a32ef-4086-4021-b874-24ed2160c4dd"
``` ```
@ -119,11 +143,17 @@ Carries `--self-test`. Its positive control is the one this project keeps
needing: a budget that reports `ok` because it measured nothing must needing: a budget that reports `ok` because it measured nothing must
abort instead. abort instead.
## Task: batch deliberately, and report what the rate reaches ## Task: batch deliberately, and report what the rate reaches (CANCELLED)
> **Cancelled unstarted 2026-08-01.** `SessionShape.md` §4 already puts
> the ceiling at **$24 on a $93 pass** — the cheapest of the three
> metrics to move and the least valuable. Spending a task on it while
> stage 0 lacks a CLI player is the misallocation the review found.
> SH-3 remains measured, unmet at 0.0%, and unfalsified.
```task ```task
id: CB-WP-0007-T04 id: CB-WP-0007-T04
status: todo status: cancel
priority: medium priority: medium
state_hub_task_id: "c284db6f-2e16-40fb-be8b-83ae5391c058" state_hub_task_id: "c284db6f-2e16-40fb-be8b-83ae5391c058"
``` ```
@ -154,7 +184,10 @@ priority: high
state_hub_task_id: "2f78b272-e1f9-4530-bc8c-3a9d82833546" state_hub_task_id: "2f78b272-e1f9-4530-bc8c-3a9d82833546"
``` ```
Commit `evidence/CB-EV-0006-session-shape.md`. Four tests, all reported: Commit `evidence/CB-EV-0006-session-shape.md`. **Reduced with the scope
cut** — tests 1, 2 and 4 remain; test 3 (SH-3) is cancelled with T04.
Four tests, all reported:
1. **Does the windowed metric differ from cumulative?** Against T01's 1. **Does the windowed metric differ from cumulative?** Against T01's
prediction; refuted if within 20%. prediction; refuted if within 20%.