Some checks failed
ci / check (push) Failing after 4s
CB-EV-0010. The pass's own verdict on the tier it ran at. Full-weight review withdrew the capability port the declaration was made to build. A tier-S pass has no step 2 and would have shipped it, and stage 2 would have found it unimplementable — which is what CommitWindow is already on record in this repo for doing. Two corrections of the survey's own numbers, compounding: AM-4a headroom 3,750 claimed -> 92,798 measured (25x) cheapest windowed 480,501 claimed -> 140,079 measured (3.4x) headline ratio 128x -> 1.5x (85x) The prediction from CB-EV-0009 §4 held: meta budget reads 0%, published in advance and unfalsified. A correction that is now a pattern: CB-EV-0009 reported CB-WP-0011 at 45 responses / $4.23 / 0.094; final is 71 / $7.02 / 0.099. Still the cheapest pass, so the conclusion stands. But that is the second consecutive evidence file to report its own pass's cost low — a pass cannot measure its own cost, and one quoting its own is quoting a floor. First priced tier comparison on a single subject: 0.123 $/response at L against 0.099 at S — 24% more, for a pass that found the two errors above. On one data point, step 2 is cheap. Not shipped, and said plainly: the emitted JavaScript has never been executed. The socket loop is tested end to end with synthetic HTTP and the page is asserted against as a parsed document, but no browser engine has run it. INTENT stage 1 therefore stays open even though all four of its named deliverables now exist. SH-3 reads 0.0% for a sixth consecutive pass and remains the oldest unargued number in the project. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
216 lines
9.9 KiB
Markdown
216 lines
9.9 KiB
Markdown
# CB-EV-0010 — tier L at full weight, and what it deleted
|
||
|
||
CB-WP-0012 T05. Measured 2026-08-02 at `84d6886`. Pass kind `product`.
|
||
|
||
The previous pass was the CHAOS gate's first fire, rolling this same
|
||
declaration from L down to S. This one rolled a 1 and ran at L. So for the
|
||
first time there are two passes on the same subject at two tiers, and the
|
||
comparison is the point.
|
||
|
||
---
|
||
|
||
## 1. What full-weight review actually did
|
||
|
||
**It deleted the thing the pass was declared to build.**
|
||
|
||
The declaration's structural trigger was *"creates a new capability port."*
|
||
The survey recommended `cb-render-api` + `cb-render-null` +
|
||
`cb-render-html`. The adversarial review (C3) found that recommendation
|
||
contradicted the survey's own citation of INTENT's second-use rule, and
|
||
the response conceded entirely. **What shipped has no port and no null
|
||
implementation** — a renderer against the existing `Project` trait, and a
|
||
deferral with a named trigger.
|
||
|
||
That is the strongest evidence yet that step 2 is not ceremony. A tier-S
|
||
pass has no step 2, would have shipped `cb-render-api`, and stage 2 would
|
||
have found it unimplementable and rewritten it — which is exactly what
|
||
`CommitWindow` is already on record in this repo for doing.
|
||
|
||
**Four of six challenges were conceded, not answered.** Two of them
|
||
invalidated the survey's main arguments:
|
||
|
||
| challenge | outcome |
|
||
|---|---|
|
||
| C1 the sub-100k region was never measured | conceded; **it is not empty** |
|
||
| C2 "cost zero" scored on an axis chosen to produce zero | conceded; one rule now covers browser, `sdl2` and `fltk` alike |
|
||
| C3 the port contradicts the second-use rule | conceded; **port withdrawn** |
|
||
| C6 the C1 measurements had no positive control | conceded; re-measured under it |
|
||
| C4 loopback server undersold | controls adopted into ADR-0007 |
|
||
| C5 the gate stops at the language boundary | controls adopted into ADR-0007 |
|
||
|
||
**Fidelity note, which caps all of the above.** The review ran in the same
|
||
session as the survey, not a separate one, because this environment's
|
||
standing instruction is not to spawn agents unasked. It therefore inherits
|
||
the author's sampling — the exact failure mode that produced CB-WP-0002's
|
||
dedup blind spot and CB-WP-0005's AM-7 defect. Treat the six challenges as
|
||
a lower bound.
|
||
|
||
## 2. The numbers the survey got wrong, and by how much
|
||
|
||
The survey's first draft argued the render port was unaffordable:
|
||
|
||
> *"The cheapest candidate that opens a window costs 128 times the entire
|
||
> remaining budget. This is not a near miss to be negotiated; it is two
|
||
> orders of magnitude."*
|
||
|
||
Both halves were wrong, in the same direction, for two independent
|
||
reasons — and they compound:
|
||
|
||
| | claimed | measured |
|
||
|---|---:|---:|
|
||
| AM-4a headroom | 3,750 | **92,798** (36.2% of the figure is proc-macro code that never ships) |
|
||
| cheapest windowed toolkit | 480,501 (`macroquad`) | **140,079** (`fltk`) |
|
||
| ratio | 128× | **1.5×** |
|
||
|
||
The instrument was wrong by 25×, the candidate sample by 3.4×, and the
|
||
headline claim by 85×. The 3,750 figure has been cited in every pass that
|
||
mentioned AM-4a, CB-WP-0011's reason for deferring this declaration
|
||
included.
|
||
|
||
**The recommendation survived both corrections**, and only because it was
|
||
re-derived rather than defended: the argument moved from *affordability*
|
||
to *allocation* — stage 2's `wgpu` bill is **1,741,979** marginal lines,
|
||
twelve times the stage-1 toolkit it would replace, so buying `fltk` means
|
||
spending 1.5× the remaining budget on something stage 2 discards. That
|
||
argument needs no particular value for AM-4a's target, which is why it is
|
||
worth more than the one it replaced.
|
||
|
||
## 3. AM-4a cannot survive stage 2 — raised, not decided
|
||
|
||
| | marginal lines |
|
||
|---|---:|
|
||
| AM-4a target, **total** | 250,000 |
|
||
| `wgpu` + `winit`, named by INTENT stage 2 | **1,741,979** |
|
||
|
||
**7× the entire target, 19× the corrected headroom.** No sequencing,
|
||
feature-gating or metric correction closes that. AM-4a as targeted is
|
||
incompatible with INTENT as written, and has been since both were written;
|
||
nothing in this pass caused it.
|
||
|
||
Reserved for the maintainer along with the acquisition rule of ADR-0007
|
||
Decision 3 — which is proposed by the pass that benefits from it, is
|
||
written to cost more than it saves, and is still an argument rather than a
|
||
ratification.
|
||
|
||
## 4. What shipped, and what has never been run
|
||
|
||
`cb-render-html`: HTML/SVG/JS emission, pointer-fact resolution, and a
|
||
loopback listener. `cb-play --serve PORT`. **Measured** marginal cost:
|
||
|
||
```
|
||
games-ground shipped: 23 third-party crates
|
||
cb-render-html: 23 third-party crates
|
||
new crates introduced: 0
|
||
```
|
||
|
||
AM-4a unmoved at 246,250. Own source 7,636 → 9,652.
|
||
|
||
Eight mutations run, each red for its stated reason. Two results worth
|
||
more than the six that behaved:
|
||
|
||
- **The coverage gate fired on its author again**, first run, before
|
||
commit: `ground_choices.*.choice`, `ground_choices.*.problem` and
|
||
`players.*.blame_from` were in neither list. The third is the one to
|
||
keep — an **empty vector is a leaf path of its own**, and a fixture
|
||
where every collection is populated would never have produced it. It
|
||
now renders as an explicit absence (`blamed by none`).
|
||
- **One mutation did not go red.** Removing the `Sec-Fetch-Site` arm alone
|
||
left the cross-site test green: the `Origin` check caught it
|
||
independently. Both had to go before the control bit. Defence in depth
|
||
is fine; a control that passes for a reason you did not intend has not
|
||
been demonstrated, and would have been reported as a clean result by
|
||
anyone who ran one mutation and stopped.
|
||
|
||
### What has never been executed
|
||
|
||
**The emitted JavaScript has never run.** Every test is on the Rust side:
|
||
the socket loop is driven by synthetic HTTP, and the page is asserted
|
||
against as a parsed document. No browser engine has executed `SCRIPT`.
|
||
|
||
So "hot-seat play" is delivered in the sense that the Rust half is tested
|
||
end-to-end over a real socket and the page is emitted correctly — and
|
||
**not** in the sense that anyone has played a game in a browser. Control 5
|
||
bounds the exposure (the JS is 20 lines and holds no game logic, by
|
||
assertion) but does not remove it.
|
||
|
||
**INTENT stage 1 is therefore left open.** Its four named deliverables now
|
||
all exist, which is the first time that has been true — but a stage marked
|
||
complete on the strength of code that has never been run once is the
|
||
failure mode stage 0 avoided by leaving the CLI player open for three
|
||
passes. Closing it is the maintainer's call, and it needs one person to
|
||
open the URL.
|
||
|
||
## 5. Cost, shape, and a prediction that held
|
||
|
||
| pass | kind | responses | cost | $/response |
|
||
|---|---|---|---|---|
|
||
| CB-WP-0008 | product | 134 | $17.38 | 0.123 |
|
||
| CB-WP-0009 | meta | 46 | $11.31 | 0.246 |
|
||
| CB-WP-0010 | product | 26 | $4.08 | 0.157 |
|
||
| CB-WP-0011 | product | 71 | $7.02 | **0.099** |
|
||
| **CB-WP-0012** | **product** | **72** | **$8.82** | **0.123** |
|
||
|
||
**Correcting CB-EV-0009 §4.** It reported CB-WP-0011 as `45 responses,
|
||
$4.23, 0.094 $/response` and called it the cheapest pass on record. The
|
||
final figures are **71 responses, $7.02, 0.099**. Still the cheapest, so
|
||
the conclusion holds — but the number does not, and this is the *second*
|
||
consecutive pass to report its own cost low mid-flight (CB-EV-0008 did the
|
||
same for CB-WP-0009). Twice is a pattern: **a pass cannot measure its own
|
||
cost, and every evidence file that quotes its own is quoting a floor.**
|
||
Future evidence files should quote the prior pass's final figure and mark
|
||
their own as provisional.
|
||
|
||
This tier-L pass cost 0.123 $/response against the rolled-down pass's
|
||
0.099 — **24% more per response**, for a pass that also deleted its own
|
||
deliverable and found two errors of a factor of 25 and 85. That is the
|
||
first priced comparison of the two tiers on the same subject, and on this
|
||
one data point step 2 is cheap.
|
||
|
||
**Meta budget: 0%**, exactly as CB-EV-0009 §4 predicted:
|
||
|
||
> *"if the next pass is `product`, the trailing-3 meta share drops to 0%,
|
||
> because CB-WP-0009 will be the pass that rolled off. If it does not, the
|
||
> windowing is wrong in a way neither CB-EV-0007 §3 nor CB-EV-0008 §1
|
||
> found."*
|
||
|
||
Falsifiable, published in advance, and it held. The windowing is doing
|
||
what ADR-0006 D1 says it does.
|
||
|
||
### Session shape
|
||
|
||
| metric | this pass | target |
|
||
|---|---|---|
|
||
| SH-1 mean context | 209,486 `[SOFT]` | ≤ 200,000 soft / 300,000 hard |
|
||
| SH-2 p90 context | 210,760 `[ok]` | ≤ 300,000 / 450,000 |
|
||
| SH-3 batching | 0.0% `[SOFT]` | ≥ 20% |
|
||
|
||
SH-1 drifted back over the soft line — a tier-L pass reads more than a
|
||
tier-S one, which is the mechanism, not an excuse. The standing prediction
|
||
from CB-EV-0009 (*"the next pass opened above the SH-1 hard line will cost
|
||
more per response than 0.123"*) is **untested**: this pass opened after a
|
||
compaction, below the line.
|
||
|
||
**SH-3 has now read 0.0% for six consecutive passes** against a 20% floor.
|
||
It is the oldest unargued number in the project: either the floor is wrong
|
||
or the behaviour is, and no pass has argued either. It is past time this
|
||
was a declaration of its own rather than a line in an evidence file.
|
||
|
||
## 6. Open
|
||
|
||
- **INTENT stage 1 stays open** — all four deliverables exist; none of the
|
||
browser half has been run. §4.
|
||
- **Two items reserved for the maintainer**: AM-4a vs stage 2, and
|
||
ADR-0007 Decision 3's acquisition rule. §3.
|
||
- **CHAOS gets its second entry**, and it is the more interesting one: the
|
||
first came from an override, this one from a *non*-override that deleted
|
||
its own structural trigger. Window open to 2026-09-30; 7 of 12
|
||
declarations used, 1 override.
|
||
- **GATE-REVIEW still has zero `caught`**, two passes older.
|
||
- **SH-3 at 0.0% for six passes.** §5.
|
||
- **`cb-play` is now three modes in one binary** — play, inspect, serve.
|
||
CB-EV-0009 §5 named the third mode as the second use at which the
|
||
single-binary shape should be reconsidered. That is now due, and this
|
||
pass did not do it.
|
||
- **The proc-macro correction to AM-4a is filed and unimplemented.**
|
||
ADR-0007 Decision 4. Until it lands, every AM-4a figure in this repo
|
||
overstates the load by 36%.
|