clay-borg/evidence/CB-EV-0012-execute-the-javascript.md
tegwick f768bc4a41
Some checks failed
ci / check (push) Failing after 4s
CB-WP-0014-T03: hot-seat evidenced; stage 1 open on one human check
CB-EV-0012. Stage 1, deliverable by deliverable:

  relationship-graph visualization   emitted and gated, NEVER SEEN
  drag-to-propose                    evidenced end to end
  debug inspector                    evidenced (CB-WP-0011)
  hot-seat play                      evidenced here

Hot-seat was the one closest to being claimed on the strength of the code
path existing. SeatPolicy hands every human seat a handle on one shared
Server, so turn-taking "obviously" worked — and nothing drove more than
one seat until now. The property that matters is not that two turns
happen but that the same tab, asked twice, shows two different hands.
Mutating the projection to serve P1's view to every seat turns it red.

The stage stays open on ONE named blocker rather than a vague
reservation: no browser is available to this loop, so the visualization
is evidenced only as correctly emitted. Everything testable from here has
been tested. What remains is `cb-play --serve 0`, open the URL, confirm
the table reads and a drag works. INTENT carries that note now.

The self-quoting rule from CB-EV-0011 §4 is ADOPTED: an evidence file
quotes the previous pass's final cost and never its own. CB-WP-0013
reported itself at $5.78/34 mid-flight; final is $8.26/47, under by 43%.
Four for four, always low.

Meta budget 29% [OVER] soft 25%, driven by CB-WP-0013 in a trailing three
with two cheap product passes; it was an instrument repair, which
ADR-0006 D2 exempts.

SH-1 at 347,720 [HARD] against a 300,000 ceiling. Compaction is the
remedy and this session cannot do it for itself. CB-EV-0009's standing
prediction is now live and testable for the first time in three passes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 07:58:59 +02:00

150 lines
7.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# CB-EV-0012 — the loop is closed, and the budget is blind
CB-WP-0014 T03. Measured 2026-08-02 at `55212d7`+. Pass kind `product`,
tier **M** (chaos d4=2, no override). Declaration 9 of 12.
**Applying CB-EV-0011 §4's proposed rule, and hereby adopting it:** the
cost table below quotes the **previous** pass's final figure and marks
this pass's own as provisional. See §5.
---
## 1. What was actually blocking stage 1, and what it cost to find out
CB-EV-0010 §4 said the emitted JavaScript had never run. That sentence was
carried for a pass on an assumption nobody checked. **`node` v24.11.1 was
on the machine the whole time.**
That is the fourth consecutive pass whose central finding was a claim
carried because nobody ran the command that settles it — after the 3,750
headroom figure, the "empty" sub-100k dependency region, and SH-3's 0.0%.
The pattern is stable enough now to name: **this project's characteristic
error is not faulty reasoning, it is unexecuted verification.** Every one
was cheap to check and none had been checked.
## 2. The loop, closed
The real server serves the real page → QuickJS runs *that page's own two
`<script>` blocks* → the gesture goes over a real socket → the seat's
`Choice` comes back.
Before this, every link was tested and the chain was not. The page was
asserted against as a parsed document; the socket was driven by synthetic
HTTP this test suite wrote itself. **A page whose JavaScript sent
something else entirely would have passed everything.**
Control 5 — *the page may not construct commands* — was held up by a test
that grepped the emitted script for game vocabulary. It is now asserted on
what the script actually puts on the wire, and the mutation that makes the
JS post `command=SelectAction&target=…` turns it red.
**A wrong assertion worth keeping.** The first draft required the body not
to contain `"attack"`. It legitimately does — the body is
`down=action-attack&up=seat-1`, and `action-attack` is *the id of an
element a finger landed on*. An element may name an action; that is not
the page deciding anything. What distinguishes reporting from deciding is
the **shape**: exactly two fields, `down` and `up`, carrying two ids and
nothing derived from them.
## 3. AM-4b is blind to more source than its own target
ADR-0009's first draft priced `quick-js` at *"35% of AM-4b's remaining
headroom."* After it landed, `make dep-weight` reported AM-4b
**unchanged**.
| | crates | lines |
|---|---:|---:|
| AM-4b as instrumented (`games-ground`, `--edges normal`) | 29 | 317,021 |
| the whole workspace, including dev edges | 57 | 725,258 |
| **uncounted** | **28** | **408,237** |
AM-4b measures one package and no dev edges. `criterion`, `clap`,
`ciborium` and now `quick-js` are invisible to it.
**This is the same defect the ADR refuses `node` for** — a real
acquisition scoring zero because the instrument does not look there. The
decision still stands, because `quick-js` is pinned and auditable and
`node` is neither, but it stands on the **acquisition rule alone**. The
affordability argument is withdrawn: there is none to be had until AM-4b
can see what it is buying.
Third defect in the AM-4 family, and the second found by a pass that was
*using* the instrument rather than auditing it:
| defect | status |
|---|---|
| shipped-runtime counted proc-macro crates (89,048 lines) | fixed, ADR-0008 D2 |
| AM-4b's own proc-macro share | unmeasured, owed |
| **AM-4b's scope: one package, no dev edges (408,237 lines)** | **found here, owed** |
## 4. INTENT stage 1, deliverable by deliverable
| deliverable | status |
|---|---|
| card/token/hand/**relationship-graph** visualization | **emitted and gated** — 41 of 42 leaf paths asserted present in the *parsed* document. **Never seen.** |
| drag-to-propose | **evidenced end to end** — JS gesture → socket → `Choice` |
| debug inspector | **evidenced** (CB-WP-0011) — walks bundles and scenarios, hash-verified |
| hot-seat play | **evidenced** — two seats, one listener, the projection follows the seat |
The last one was the closest to being claimed on the strength of the code
path existing. `SeatPolicy` hands every human seat a handle on one shared
`Server`, so turn-taking "obviously" worked — and nothing drove more than
one seat until this pass. The property that matters is not that two turns
happen but that **the same tab, asked twice, shows two different hands**;
mutating the projection to serve P1's view to every seat turns it red.
### The stage stays open, on one named blocker
Three of four deliverables are evidenced by executing code. The first is
evidenced only as *correctly emitted*: QuickJS has no layout engine, no
browser is available here, and no test in this repo can establish that an
SVG table is legible or that a drag feels like a drag.
**Everything testable from here has been tested.** What remains is one
human action: run `cb-play --serve 0`, open the printed URL, and confirm
the table reads and a drag works. That is the whole of the remaining
blocker, it is the maintainer's to do, and it is named here so the stage
is no longer carried on a vague reservation.
Closing it unilaterally would be the failure mode stage 0 avoided.
Carrying it *vaguely* was the failure mode this pass found.
## 5. Cost, shape, and the rule now adopted
| pass | kind | responses | cost | $/response |
|---|---|---|---|---|
| CB-WP-0012 | product | 97 | $14.52 | 0.150 |
| **CB-WP-0013** (final) | meta | 47 | **$8.26** | **0.176** |
| CB-WP-0014 | product | *provisional — see below* | | |
**The rule is adopted.** CB-EV-0011 proposed that an evidence file quote
the previous pass's final cost and never its own. CB-WP-0013 reported
itself at $5.78/34 responses mid-flight; final is **$8.26/47** — under by
43%. Four for four, always low. This file quotes CB-WP-0013 and declines
to quote itself, and future evidence files should do the same.
**Meta budget: 29% `[OVER]`** the soft 25%, driven by CB-WP-0013's $8.26
sitting in a trailing three with two cheap product passes. Reported under
the rule that requires reporting it; no product work was displaced, and
CB-WP-0013 was an instrument repair, which ADR-0006 D2 exempts anyway.
**SH-1 is at 347,720 `[HARD]`** against a 300,000 ceiling. The remedy is a
compaction and it is free, and it is not something this session can do for
itself. **CB-EV-0009's standing prediction is now live and testable**: if
the next pass opens above the SH-1 hard line, it should cost more per
response than 0.123. It has been unfalsified for three passes because
every one of them opened below the line.
## 6. Open
- **AM-4b's scope defect (408,237 uncounted lines)** and its unmeasured
proc-macro share. §3.
- **`python3` as a toolchain dependency was never argued.** ADR-0009
declined to lean on it as precedent and left it owed.
- **INTENT stage 1: one human verification.** §4.
- **AM-4a still cannot survive stage 2** — 1,741,979 against 161,000,
10.8×. Reserved for the maintainer since ADR-0007.
- **ADR-0007 D3's acquisition rule** remains unratified — and it has now
decided two dependency questions, which raises rather than lowers the
cost of leaving it unratified.
- **Chaos: 9 of 12 declarations, 1 override.**