clay-borg/evidence/CB-EV-0010-render-port.md
tegwick c38ecd9da3
Some checks failed
ci / check (push) Failing after 4s
CB-WP-0012-T05: evidence — tier L deleted its own deliverable
CB-EV-0010. The pass's own verdict on the tier it ran at.

Full-weight review withdrew the capability port the declaration was made
to build. A tier-S pass has no step 2 and would have shipped it, and
stage 2 would have found it unimplementable — which is what CommitWindow
is already on record in this repo for doing.

Two corrections of the survey's own numbers, compounding:
  AM-4a headroom       3,750 claimed -> 92,798 measured   (25x)
  cheapest windowed  480,501 claimed -> 140,079 measured  (3.4x)
  headline ratio         128x        -> 1.5x              (85x)

The prediction from CB-EV-0009 §4 held: meta budget reads 0%, published
in advance and unfalsified.

A correction that is now a pattern: CB-EV-0009 reported CB-WP-0011 at
45 responses / $4.23 / 0.094; final is 71 / $7.02 / 0.099. Still the
cheapest pass, so the conclusion stands. But that is the second
consecutive evidence file to report its own pass's cost low — a pass
cannot measure its own cost, and one quoting its own is quoting a floor.

First priced tier comparison on a single subject: 0.123 $/response at L
against 0.099 at S — 24% more, for a pass that found the two errors
above. On one data point, step 2 is cheap.

Not shipped, and said plainly: the emitted JavaScript has never been
executed. The socket loop is tested end to end with synthetic HTTP and
the page is asserted against as a parsed document, but no browser engine
has run it. INTENT stage 1 therefore stays open even though all four of
its named deliverables now exist.

SH-3 reads 0.0% for a sixth consecutive pass and remains the oldest
unargued number in the project.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 04:30:10 +02:00

216 lines
9.9 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# CB-EV-0010 — tier L at full weight, and what it deleted
CB-WP-0012 T05. Measured 2026-08-02 at `84d6886`. Pass kind `product`.
The previous pass was the CHAOS gate's first fire, rolling this same
declaration from L down to S. This one rolled a 1 and ran at L. So for the
first time there are two passes on the same subject at two tiers, and the
comparison is the point.
---
## 1. What full-weight review actually did
**It deleted the thing the pass was declared to build.**
The declaration's structural trigger was *"creates a new capability port."*
The survey recommended `cb-render-api` + `cb-render-null` +
`cb-render-html`. The adversarial review (C3) found that recommendation
contradicted the survey's own citation of INTENT's second-use rule, and
the response conceded entirely. **What shipped has no port and no null
implementation** — a renderer against the existing `Project` trait, and a
deferral with a named trigger.
That is the strongest evidence yet that step 2 is not ceremony. A tier-S
pass has no step 2, would have shipped `cb-render-api`, and stage 2 would
have found it unimplementable and rewritten it — which is exactly what
`CommitWindow` is already on record in this repo for doing.
**Four of six challenges were conceded, not answered.** Two of them
invalidated the survey's main arguments:
| challenge | outcome |
|---|---|
| C1 the sub-100k region was never measured | conceded; **it is not empty** |
| C2 "cost zero" scored on an axis chosen to produce zero | conceded; one rule now covers browser, `sdl2` and `fltk` alike |
| C3 the port contradicts the second-use rule | conceded; **port withdrawn** |
| C6 the C1 measurements had no positive control | conceded; re-measured under it |
| C4 loopback server undersold | controls adopted into ADR-0007 |
| C5 the gate stops at the language boundary | controls adopted into ADR-0007 |
**Fidelity note, which caps all of the above.** The review ran in the same
session as the survey, not a separate one, because this environment's
standing instruction is not to spawn agents unasked. It therefore inherits
the author's sampling — the exact failure mode that produced CB-WP-0002's
dedup blind spot and CB-WP-0005's AM-7 defect. Treat the six challenges as
a lower bound.
## 2. The numbers the survey got wrong, and by how much
The survey's first draft argued the render port was unaffordable:
> *"The cheapest candidate that opens a window costs 128 times the entire
> remaining budget. This is not a near miss to be negotiated; it is two
> orders of magnitude."*
Both halves were wrong, in the same direction, for two independent
reasons — and they compound:
| | claimed | measured |
|---|---:|---:|
| AM-4a headroom | 3,750 | **92,798** (36.2% of the figure is proc-macro code that never ships) |
| cheapest windowed toolkit | 480,501 (`macroquad`) | **140,079** (`fltk`) |
| ratio | 128× | **1.5×** |
The instrument was wrong by 25×, the candidate sample by 3.4×, and the
headline claim by 85×. The 3,750 figure has been cited in every pass that
mentioned AM-4a, CB-WP-0011's reason for deferring this declaration
included.
**The recommendation survived both corrections**, and only because it was
re-derived rather than defended: the argument moved from *affordability*
to *allocation* — stage 2's `wgpu` bill is **1,741,979** marginal lines,
twelve times the stage-1 toolkit it would replace, so buying `fltk` means
spending 1.5× the remaining budget on something stage 2 discards. That
argument needs no particular value for AM-4a's target, which is why it is
worth more than the one it replaced.
## 3. AM-4a cannot survive stage 2 — raised, not decided
| | marginal lines |
|---|---:|
| AM-4a target, **total** | 250,000 |
| `wgpu` + `winit`, named by INTENT stage 2 | **1,741,979** |
**7× the entire target, 19× the corrected headroom.** No sequencing,
feature-gating or metric correction closes that. AM-4a as targeted is
incompatible with INTENT as written, and has been since both were written;
nothing in this pass caused it.
Reserved for the maintainer along with the acquisition rule of ADR-0007
Decision 3 — which is proposed by the pass that benefits from it, is
written to cost more than it saves, and is still an argument rather than a
ratification.
## 4. What shipped, and what has never been run
`cb-render-html`: HTML/SVG/JS emission, pointer-fact resolution, and a
loopback listener. `cb-play --serve PORT`. **Measured** marginal cost:
```
games-ground shipped: 23 third-party crates
cb-render-html: 23 third-party crates
new crates introduced: 0
```
AM-4a unmoved at 246,250. Own source 7,636 → 9,652.
Eight mutations run, each red for its stated reason. Two results worth
more than the six that behaved:
- **The coverage gate fired on its author again**, first run, before
commit: `ground_choices.*.choice`, `ground_choices.*.problem` and
`players.*.blame_from` were in neither list. The third is the one to
keep — an **empty vector is a leaf path of its own**, and a fixture
where every collection is populated would never have produced it. It
now renders as an explicit absence (`blamed by none`).
- **One mutation did not go red.** Removing the `Sec-Fetch-Site` arm alone
left the cross-site test green: the `Origin` check caught it
independently. Both had to go before the control bit. Defence in depth
is fine; a control that passes for a reason you did not intend has not
been demonstrated, and would have been reported as a clean result by
anyone who ran one mutation and stopped.
### What has never been executed
**The emitted JavaScript has never run.** Every test is on the Rust side:
the socket loop is driven by synthetic HTTP, and the page is asserted
against as a parsed document. No browser engine has executed `SCRIPT`.
So "hot-seat play" is delivered in the sense that the Rust half is tested
end-to-end over a real socket and the page is emitted correctly — and
**not** in the sense that anyone has played a game in a browser. Control 5
bounds the exposure (the JS is 20 lines and holds no game logic, by
assertion) but does not remove it.
**INTENT stage 1 is therefore left open.** Its four named deliverables now
all exist, which is the first time that has been true — but a stage marked
complete on the strength of code that has never been run once is the
failure mode stage 0 avoided by leaving the CLI player open for three
passes. Closing it is the maintainer's call, and it needs one person to
open the URL.
## 5. Cost, shape, and a prediction that held
| pass | kind | responses | cost | $/response |
|---|---|---|---|---|
| CB-WP-0008 | product | 134 | $17.38 | 0.123 |
| CB-WP-0009 | meta | 46 | $11.31 | 0.246 |
| CB-WP-0010 | product | 26 | $4.08 | 0.157 |
| CB-WP-0011 | product | 71 | $7.02 | **0.099** |
| **CB-WP-0012** | **product** | **72** | **$8.82** | **0.123** |
**Correcting CB-EV-0009 §4.** It reported CB-WP-0011 as `45 responses,
$4.23, 0.094 $/response` and called it the cheapest pass on record. The
final figures are **71 responses, $7.02, 0.099**. Still the cheapest, so
the conclusion holds — but the number does not, and this is the *second*
consecutive pass to report its own cost low mid-flight (CB-EV-0008 did the
same for CB-WP-0009). Twice is a pattern: **a pass cannot measure its own
cost, and every evidence file that quotes its own is quoting a floor.**
Future evidence files should quote the prior pass's final figure and mark
their own as provisional.
This tier-L pass cost 0.123 $/response against the rolled-down pass's
0.099 — **24% more per response**, for a pass that also deleted its own
deliverable and found two errors of a factor of 25 and 85. That is the
first priced comparison of the two tiers on the same subject, and on this
one data point step 2 is cheap.
**Meta budget: 0%**, exactly as CB-EV-0009 §4 predicted:
> *"if the next pass is `product`, the trailing-3 meta share drops to 0%,
> because CB-WP-0009 will be the pass that rolled off. If it does not, the
> windowing is wrong in a way neither CB-EV-0007 §3 nor CB-EV-0008 §1
> found."*
Falsifiable, published in advance, and it held. The windowing is doing
what ADR-0006 D1 says it does.
### Session shape
| metric | this pass | target |
|---|---|---|
| SH-1 mean context | 209,486 `[SOFT]` | ≤ 200,000 soft / 300,000 hard |
| SH-2 p90 context | 210,760 `[ok]` | ≤ 300,000 / 450,000 |
| SH-3 batching | 0.0% `[SOFT]` | ≥ 20% |
SH-1 drifted back over the soft line — a tier-L pass reads more than a
tier-S one, which is the mechanism, not an excuse. The standing prediction
from CB-EV-0009 (*"the next pass opened above the SH-1 hard line will cost
more per response than 0.123"*) is **untested**: this pass opened after a
compaction, below the line.
**SH-3 has now read 0.0% for six consecutive passes** against a 20% floor.
It is the oldest unargued number in the project: either the floor is wrong
or the behaviour is, and no pass has argued either. It is past time this
was a declaration of its own rather than a line in an evidence file.
## 6. Open
- **INTENT stage 1 stays open** — all four deliverables exist; none of the
browser half has been run. §4.
- **Two items reserved for the maintainer**: AM-4a vs stage 2, and
ADR-0007 Decision 3's acquisition rule. §3.
- **CHAOS gets its second entry**, and it is the more interesting one: the
first came from an override, this one from a *non*-override that deleted
its own structural trigger. Window open to 2026-09-30; 7 of 12
declarations used, 1 override.
- **GATE-REVIEW still has zero `caught`**, two passes older.
- **SH-3 at 0.0% for six passes.** §5.
- **`cb-play` is now three modes in one binary** — play, inspect, serve.
CB-EV-0009 §5 named the third mode as the second use at which the
single-binary shape should be reconsidered. That is now due, and this
pass did not do it.
- **The proc-macro correction to AM-4a is filed and unimplemented.**
ADR-0007 Decision 4. Until it lands, every AM-4a figure in this repo
overstates the load by 36%.