CB-WP-0009-T04: CB-EV-0008 — the gate changes, measured
Some checks failed
ci / check (push) Failing after 4s

The windowed budget is confirmed for the failure it was written against
and not for the general claim: a single-pass window reads 0% for
CB-WP-0008 against 50% lifetime, but the trailing-3 window reads 45%
against 50%, inside the refutation band. The prediction was written
before the window size was chosen and did not say which comparison it
meant. Both readings are on record and whether 3 is the right window is
carried as open.

gate-review's first run: 9 gates, 0 due, 2 silent. The silent two are the
chaos roll and gate-review itself, both with dates. A registry where
everything looked productive would have been one written to look good.

D4 holds per pass, not per task: three of four tasks shipped a command,
and the two that did not are the spec change that makes the commands
normative and the evidence file that checks them.

Cost is the honest part. This pass cashed out three commands and ran at
$0.177/response — cheaper than every previous meta pass (0.228, 0.298,
0.362) and still 1.4x the product pass at 0.123. Partial support for D4,
not vindication. Context breached both shape targets because the pass ran
on an already-long session; reported, not gated.

Meta reads 45% of the trailing three against a soft 25%. Nothing was
displaced, but the number is over the line and the next pass should be
product.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
tegwick 2026-08-01 15:45:08 +02:00
parent 2b6bf29e8f
commit 06483bc587
2 changed files with 142 additions and 2 deletions

View file

@ -0,0 +1,134 @@
# CB-EV-0008 — did the gate changes fire?
CB-WP-0009 T04. Measured 2026-08-01 at `HEAD` after T03. This pass is
`kind: meta`, opened above the budget line on recorded maintainer
authorization (ADR-0006 D2's first use).
---
## 1. Does the windowed meta share differ from cumulative?
**Yes at one pass, no at three — and the prediction did not say which.**
| reading | value |
|---|---|
| CB-WP-0008 alone (product) | **0%** |
| trailing 3 passes, before this one | 36% |
| trailing 3 passes, now | **45%** |
| lifetime, all 9 passes | 50% |
T01 predicted the windowed figure would differ sharply from cumulative,
refuted within 20%. Against the **single-pass** window that is decisive:
0% versus 50%. Against the **trailing-3** window it is not: 45% versus
50% is a 10% relative gap, inside the refutation band.
Both are true, and the honest reading is that **the prediction was
underspecified** — it was written before the window size was chosen, and
T01 then chose three passes for a stated reason (one pass is a coin flip:
a meta pass reads 100%, a product pass 0%). So:
- D1 is **confirmed for the failure it was written against**: the v1.5
budget could not see a 100%-product pass, and now it can.
- D1 is **not** confirmed as "trailing-3 always differs from lifetime."
Right now it does not, because the recent mix happens to resemble the
lifetime mix. That is a fact about this project's history, not evidence
for or against the machinery.
**Carried as an open question**, not resolved by argument: whether three
is the right window. It should be revisited when a pass makes the two
diverge, which is the only cheap evidence available.
## 2. What does `gate-review` say on its first run?
```
9 gate(s) — 0 due, 2 silent, 0 drifted
```
The two silent entries are the uncomfortable ones, and they are supposed
to be:
- **CHAOS** — the chaos roll has never changed a tier. Its 12-declaration
window (4 used) exists precisely to find out whether it ever will.
- **GATE-REVIEW** — this registry itself, which has caught nothing
because it is one hour old, and which retires on 2026-12-31 if it has
not by then retired, tightened or forced the re-justification of
anything.
A registry where every gate looked current and productive would have been
a registry written to look good. Two of nine are on notice.
The `caught` lists are the load-bearing part, and they are honest in the
other direction too: **M-D1-MUT has six entries, LOOP-LINT two, and the
DFD gate one that landed this week** — `gr_scenarios` stale at 21 after
CB-WP-0008 T03 added three scenarios, caught by `make all` going red
before the push.
## 3. Did D4 hold for this pass?
**Three of four tasks shipped a command; the fourth is this file.**
| task | shipped |
|---|---|
| T01 | `cb-cost.pass_costs`, windowed `make status` |
| T02 | `gates.toml`, `make gate-review`, a `loop-lint` check |
| T03 | InnerLoop v1.6 — **prose** |
| T04 | this evidence file — **prose** |
D4 says a `meta` pass must ship *at least one* executable check, so the
rule holds. But the pass makes the boundary clear: **the spec change and
the evidence file are prose, and a rule requiring every task to cash out
would be wrong** — the spec edit is what makes the commands normative,
and the evidence file is what checks them.
So D4 stands as written (per pass, not per task), and this pass is on
record as its first exercise.
## 4. Cost
| pass | kind | responses | cost | $/response |
|---|---|---|---|---|
| CB-WP-0006 | meta | 158 | $57.22 | 0.362 |
| CB-WP-0007 | meta | 21 | $7.95 | 0.298* |
| CB-WP-0008 | product | 134 | $17.38 | **0.123** |
| **CB-WP-0009** | **meta** | **38** | **$6.73** | **0.177** |
\* the CB-WP-0007 figure is from the earlier hand-windowed measurement;
the table's per-pass windows are now computed by `pass_costs`.
**This is the interesting number.** D4's claim is that loop work that
cashes out into a command is cheap. This pass cashed out three commands
and ran at **0.177** — cheaper than every previous meta pass (0.228,
0.298, 0.362) and still **1.4× the product pass**. Partial support, not
vindication, and stated as such.
Session shape over the same window is worse than the product pass and
breaches both targets:
| metric | this pass | CB-WP-0008 | target |
|---|---|---|---|
| SH-1 mean context | **290,171** | 161,994 | ≤ 200,000 |
| SH-2 p90 context | **311,243** | 215,664 | ≤ 300,000 |
| SH-3 batching | 0.0% | 6.9% | ≥ 20% |
The mechanical cause is visible and boring: this pass ran on top of an
already-long session with no compaction, where CB-WP-0008 started
immediately after one. The shape budget reports it; nothing gates it; and
per CB-RES-0005 §4 nothing should.
**Meta share after this pass: 45% of the trailing three, against a soft
25%.** Reported under the rule this pass wrote. Nothing was displaced —
CB-WP-0008 shipped stage 0 first, and this work was maintainer-authorized
on top of it — but the number is over the line and the next pass should
be product to bring it back.
## 5. Open
- **Is 3 the right trailing window?** Unanswered; revisit when a pass
makes single-pass and trailing-3 diverge.
- **CHAOS and GATE-REVIEW have caught nothing.** Both have dates.
- **`mech environment setup` reads 17 turns / $2.90** this pass, which is
the cost classifier scoring inline `python3 - <<EOF` edits as
environment setup. The category is mislabelled rather than the work
being wrong — a small classifier defect, recorded, not chased.
- **The next pass should be `product`.** The budget says so, and this is
the first time it has said so about work that had not started yet.

View file

@ -2,7 +2,7 @@
id: CB-WP-0009
kind: meta
title: "Make control gates experimental: window them, review them, cash them out"
status: in_progress
status: done
authorized_above_budget: "maintainer, 2026-08-01 — 'improve our policies where it seems promising; we will experiment our way to sensible task- and context-specific control and review gates'. Meta read 59% at the time."
state_hub_workstream_id: "98c0af4a-415c-4cf5-92d0-a722eb70cb90"
---
@ -120,7 +120,7 @@ started at 399.
```task
id: CB-WP-0009-T04
status: todo
status: done
priority: high
state_hub_task_id: "e06f57fc-147f-4cba-b398-0db6b060381f"
```
@ -139,3 +139,9 @@ answered from output, not argument:
pass or needs the exception written down.
4. **Cost**, against the 0.1230.362 per-response range, and the meta
share this pass itself consumed — reported under the rule it wrote.
**Done 2026-08-01.** [CB-EV-0008](../evidence/CB-EV-0008-adaptive-gates.md).
The pass ran at **$0.177/response** — cheaper than every previous meta
pass and still 1.4× the product pass, which is partial support for D4 and
is reported as such. Meta share reads **45% over the trailing three**,
over the line; the next pass should be product.