clay-borg/workplans/CB-WP-0014-execute-the-javascript.md

208 lines
8.9 KiB
Markdown
Raw Normal View History

---
id: CB-WP-0014
kind: product
title: "Execute the JavaScript, and close or keep stage 1"
CB-WP-0014-T03: hot-seat evidenced; stage 1 open on one human check CB-EV-0012. Stage 1, deliverable by deliverable: relationship-graph visualization emitted and gated, NEVER SEEN drag-to-propose evidenced end to end debug inspector evidenced (CB-WP-0011) hot-seat play evidenced here Hot-seat was the one closest to being claimed on the strength of the code path existing. SeatPolicy hands every human seat a handle on one shared Server, so turn-taking "obviously" worked — and nothing drove more than one seat until now. The property that matters is not that two turns happen but that the same tab, asked twice, shows two different hands. Mutating the projection to serve P1's view to every seat turns it red. The stage stays open on ONE named blocker rather than a vague reservation: no browser is available to this loop, so the visualization is evidenced only as correctly emitted. Everything testable from here has been tested. What remains is `cb-play --serve 0`, open the URL, confirm the table reads and a drag works. INTENT carries that note now. The self-quoting rule from CB-EV-0011 §4 is ADOPTED: an evidence file quotes the previous pass's final cost and never its own. CB-WP-0013 reported itself at $5.78/34 mid-flight; final is $8.26/47, under by 43%. Four for four, always low. Meta budget 29% [OVER] soft 25%, driven by CB-WP-0013 in a trailing three with two cheap product passes; it was an instrument repair, which ADR-0006 D2 exempts. SH-1 at 347,720 [HARD] against a 300,000 ceiling. Compaction is the remedy and this session cannot do it for itself. CB-EV-0009's standing prediction is now live and testable for the first time in three passes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 07:58:59 +02:00
status: done
state_hub_workstream_id: "a8602f84-d073-4873-b845-afe9d2620a5d"
---
# Purpose
```
structural tier M (adds an external dependency to the toolchain —
InnerLoop v1.6 tier table)
chaos d4 = 2 → no override
declared tier M
```
Declaration 9 of 12. Tier M merges the survey and the ADR into one
document; the adversarial review is optional.
## The one thing blocking stage 1
CB-EV-0010 §4, stated plainly at the time:
> **The emitted JavaScript has never run.** Every test is on the Rust
> side: the socket loop is driven by synthetic HTTP, and the page is
> asserted against as a parsed document. No browser engine has executed
> `SCRIPT`.
All four of INTENT stage 1's named deliverables now exist. That sentence
is the only reason the stage is still open, and it has been carried
unexamined for one pass on the assumption that executing it was not
possible here.
**It is possible.** `node` v24.11.1 is on this machine. The assumption was
never checked, which makes it the same shape as every other finding in the
last three passes: a claim carried because nobody ran the one command that
would settle it.
## What executing it actually proves
Not that the table *looks* right — that needs a rendering engine and a
human eye, and neither is available here. What it proves is the contract
ADR-0007 control 5 rests on:
> the page reports **raw pointer facts and nothing else**, and Rust
> decides what command they mean.
That contract is currently asserted by a test that greps the emitted
script for game vocabulary. Grepping for the absence of words is a weak
proxy for "this code cannot construct a command". Running it, feeding it
synthetic pointer events, and observing exactly what it puts on the wire
is the real check.
## Task: decide whether a JS engine joins the toolchain
```task
id: CB-WP-0014-T01
CB-WP-0014-T01/T02: execute the JavaScript — and find AM-4b blind ADR-0009: embed quick-js; node is refused. Measured marginal cost against the dev-toolchain graph, under the positive control: boa_engine 896,410 rquickjs 69,985 quick-js 11,434 node 0 <- and that zero is the problem ADR-0007 D3's acquisition rule biting its author. CI runs on rust:1.97, which has no node, so the test would make our build fetch a JS runtime of tens of millions of unaudited lines while scoring zero on the only instrument that governs dependencies. A browser is exempt because a developer has one regardless of us; a CI-installed runtime is not. The loop is now closed: the real server serves the real page, QuickJS runs that page's own scripts, the gesture goes over a real socket, and the seat's Choice comes back. Before this, every link was tested and the chain was not — a page whose JavaScript sent something else entirely would have passed everything. Three controls, each red for its stated reason: the JS posting a command name instead of ids, the gesture not being delivered (EXPECT-VACUOUS), and the token stripped from the endpoint. A wrong assertion worth keeping: the first draft required the body not to contain "attack". It legitimately does — action-attack is the id of an element a finger landed on. An element may name an action; that is not the page deciding. The real test is the shape: exactly two fields, down and up, carrying two ids and nothing derived from them. AND the ADR's own cost argument was wrong. It claimed 35% of AM-4b's headroom; after landing AM-4b did not move at all. It measures games-ground --edges normal — one package, no dev edges. Measured, the workspace including dev edges is 725,258 lines against AM-4b's 317,021: 408,237 uncounted, MORE THAN THE TARGET ITSELF (criterion, clap, ciborium, quick-js). The decision stands on the acquisition rule; the affordability argument is withdrawn. Third defect in the AM-4 family. Also fixed structurally rather than by raising a limit: `make status` had grown past its 40-line readability gate as workplans accumulated. Closed workplans now collapse to one line, so the report is fixed-size. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 07:56:02 +02:00
status: done
priority: high
state_hub_task_id: "ad118b97-e59a-41ec-9054-e314aa9482f6"
```
Write `decisions/ADR-0009-*.md` (tier M: survey and decision in one).
**This is a real dependency decision, not a formality.** CI runs on the
`rust:1.97` image, which has no `node`. Adding a test that needs one means
either installing it in CI or letting the test skip — and **a test that
skips silently is the harness-does-nothing class this project has found
five times.** If it can skip, it must fail loudly instead.
Decide, and price it under ADR-0007 Decision 3's acquisition rule — the
rule this project adopted precisely so that "it's only a dev tool" is not
an automatic pass:
- `python3` is already a toolchain dependency, so this is the second, not
the first. Say whether that makes it cheaper or is merely a precedent
being leaned on.
- What the alternatives cost: no execution at all (status quo, and the
reason stage 1 is open), or a JS interpreter as a Rust crate (measure it
— do not estimate, and note that AM-4b now has an unmeasured proc-macro
share of its own).
- Whether the engine is required by `make all` or by a separate target.
**AM-4b is the budget this lands in**, and it is the instrument
CB-WP-0013 explicitly declined to correct. Do not use its uncorrected
headroom as an argument.
CB-WP-0014-T01/T02: execute the JavaScript — and find AM-4b blind ADR-0009: embed quick-js; node is refused. Measured marginal cost against the dev-toolchain graph, under the positive control: boa_engine 896,410 rquickjs 69,985 quick-js 11,434 node 0 <- and that zero is the problem ADR-0007 D3's acquisition rule biting its author. CI runs on rust:1.97, which has no node, so the test would make our build fetch a JS runtime of tens of millions of unaudited lines while scoring zero on the only instrument that governs dependencies. A browser is exempt because a developer has one regardless of us; a CI-installed runtime is not. The loop is now closed: the real server serves the real page, QuickJS runs that page's own scripts, the gesture goes over a real socket, and the seat's Choice comes back. Before this, every link was tested and the chain was not — a page whose JavaScript sent something else entirely would have passed everything. Three controls, each red for its stated reason: the JS posting a command name instead of ids, the gesture not being delivered (EXPECT-VACUOUS), and the token stripped from the endpoint. A wrong assertion worth keeping: the first draft required the body not to contain "attack". It legitimately does — action-attack is the id of an element a finger landed on. An element may name an action; that is not the page deciding. The real test is the shape: exactly two fields, down and up, carrying two ids and nothing derived from them. AND the ADR's own cost argument was wrong. It claimed 35% of AM-4b's headroom; after landing AM-4b did not move at all. It measures games-ground --edges normal — one package, no dev edges. Measured, the workspace including dev edges is 725,258 lines against AM-4b's 317,021: 408,237 uncounted, MORE THAN THE TARGET ITSELF (criterion, clap, ciborium, quick-js). The decision stands on the acquisition rule; the affordability argument is withdrawn. Third defect in the AM-4 family. Also fixed structurally rather than by raising a limit: `make status` had grown past its 40-line readability gate as workplans accumulated. Closed workplans now collapse to one line, so the report is fixed-size. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 07:56:02 +02:00
**Done 2026-08-02.**
[ADR-0009](../decisions/ADR-0009-embed-the-js-engine.md) — **embed
`quick-js`; `node` is refused.** Measured, marginal, under the control:
`boa_engine` 896,410 · `rquickjs` 69,985 · **`quick-js` 11,434** · `node`
**0**, and that zero is the problem.
This is ADR-0007 D3's acquisition rule biting its author. CI runs on
`rust:1.97`, which has no `node` — so the test would make *our build*
fetch a JS runtime of tens of millions of unaudited lines, scoring zero on
the only instrument that governs dependencies. A browser is exempt because
a developer has one regardless of us; a CI-installed runtime is not.
**And the cost argument in the first draft was wrong.** It claimed 35% of
AM-4b's headroom. After landing, AM-4b did not move at all — it measures
`games-ground --edges normal`, one package, no dev edges. Measured: the
workspace including dev edges is **725,258** lines against AM-4b's
**317,021**, so **408,237 lines are uncounted — more than the target
itself**. The decision stands on the acquisition rule; the affordability
argument is withdrawn, because there is none to be had until the
instrument can see what it is buying. Filed as the third AM-4 defect.
## Task: run it, end to end, through a real socket
```task
id: CB-WP-0014-T02
CB-WP-0014-T01/T02: execute the JavaScript — and find AM-4b blind ADR-0009: embed quick-js; node is refused. Measured marginal cost against the dev-toolchain graph, under the positive control: boa_engine 896,410 rquickjs 69,985 quick-js 11,434 node 0 <- and that zero is the problem ADR-0007 D3's acquisition rule biting its author. CI runs on rust:1.97, which has no node, so the test would make our build fetch a JS runtime of tens of millions of unaudited lines while scoring zero on the only instrument that governs dependencies. A browser is exempt because a developer has one regardless of us; a CI-installed runtime is not. The loop is now closed: the real server serves the real page, QuickJS runs that page's own scripts, the gesture goes over a real socket, and the seat's Choice comes back. Before this, every link was tested and the chain was not — a page whose JavaScript sent something else entirely would have passed everything. Three controls, each red for its stated reason: the JS posting a command name instead of ids, the gesture not being delivered (EXPECT-VACUOUS), and the token stripped from the endpoint. A wrong assertion worth keeping: the first draft required the body not to contain "attack". It legitimately does — action-attack is the id of an element a finger landed on. An element may name an action; that is not the page deciding. The real test is the shape: exactly two fields, down and up, carrying two ids and nothing derived from them. AND the ADR's own cost argument was wrong. It claimed 35% of AM-4b's headroom; after landing AM-4b did not move at all. It measures games-ground --edges normal — one package, no dev edges. Measured, the workspace including dev edges is 725,258 lines against AM-4b's 317,021: 408,237 uncounted, MORE THAN THE TARGET ITSELF (criterion, clap, ciborium, quick-js). The decision stands on the acquisition rule; the affordability argument is withdrawn. Third defect in the AM-4 family. Also fixed structurally rather than by raising a limit: `make status` had grown past its 40-line readability gate as workplans accumulated. Closed workplans now collapse to one line, so the report is fixed-size. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 07:56:02 +02:00
status: done
priority: high
state_hub_task_id: "56f09f33-cdb9-4d58-80c9-73e5ee1cf589"
```
Execute `cb_render_html::doc::SCRIPT` in a real JS engine against a
minimal DOM, synthesize a pointer-down and pointer-up on two real element
ids from a real emitted document, and let its `fetch` reach the **actual
`Server`** — the same `next_choice` loop `cb-play --serve` runs.
Acceptance: the seat's `Choice` comes back from a pointer gesture that
originated in JavaScript. That is the first time anything in this project
has closed that loop.
**Controls, and they are the point.** The test must go red when:
- the JS is mutated to POST a command name instead of element ids — this
is **control 5 asserted rather than grepped**, and if it does not fire,
the grep test was the only thing holding the contract up;
- the pointer events are not delivered at all — the EXPECT-VACUOUS case, a
harness that runs a script and asserts nothing about what it did;
- the token is stripped from the endpoint the page was given.
CB-WP-0014-T01/T02: execute the JavaScript — and find AM-4b blind ADR-0009: embed quick-js; node is refused. Measured marginal cost against the dev-toolchain graph, under the positive control: boa_engine 896,410 rquickjs 69,985 quick-js 11,434 node 0 <- and that zero is the problem ADR-0007 D3's acquisition rule biting its author. CI runs on rust:1.97, which has no node, so the test would make our build fetch a JS runtime of tens of millions of unaudited lines while scoring zero on the only instrument that governs dependencies. A browser is exempt because a developer has one regardless of us; a CI-installed runtime is not. The loop is now closed: the real server serves the real page, QuickJS runs that page's own scripts, the gesture goes over a real socket, and the seat's Choice comes back. Before this, every link was tested and the chain was not — a page whose JavaScript sent something else entirely would have passed everything. Three controls, each red for its stated reason: the JS posting a command name instead of ids, the gesture not being delivered (EXPECT-VACUOUS), and the token stripped from the endpoint. A wrong assertion worth keeping: the first draft required the body not to contain "attack". It legitimately does — action-attack is the id of an element a finger landed on. An element may name an action; that is not the page deciding. The real test is the shape: exactly two fields, down and up, carrying two ids and nothing derived from them. AND the ADR's own cost argument was wrong. It claimed 35% of AM-4b's headroom; after landing AM-4b did not move at all. It measures games-ground --edges normal — one package, no dev edges. Measured, the workspace including dev edges is 725,258 lines against AM-4b's 317,021: 408,237 uncounted, MORE THAN THE TARGET ITSELF (criterion, clap, ciborium, quick-js). The decision stands on the acquisition rule; the affordability argument is withdrawn. Third defect in the AM-4 family. Also fixed structurally rather than by raising a limit: `make status` had grown past its 40-line readability gate as workplans accumulated. Closed workplans now collapse to one line, so the report is fixed-size. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 07:56:02 +02:00
**Done 2026-08-02.** `crates/cb-render-html/src/jsrun.rs` and
`hotseat::tests::a_gesture_in_javascript_becomes_a_move_in_the_game`.
**The loop is closed.** The real server serves the real page; QuickJS runs
*that page's own scripts*; the gesture it produces goes over a real socket;
the seat's `Choice` comes back. Both `<script>` blocks are lifted from the
served document, so the endpoint under test is the one the page actually
carried — token included.
Three controls, each red for its stated reason:
| mutation | result |
|---|---|
| the JS posts `command=SelectAction&target=…` | body assertion red — **control 5 asserted, not grepped** |
| the gesture is not delivered (EXPECT-VACUOUS) | `expected exactly one fetch, got []` |
| the token dropped from the endpoint | the end-to-end token assertion red |
**A wrong assertion, worth keeping.** The first draft asserted the body
must not contain `"attack"`. It legitimately does: the body is
`down=action-attack&up=seat-1`, and `action-attack` is *the id of an
element a finger landed on*. An element may name an action; that is not
the page deciding anything. The real test of reporting-versus-deciding is
the **shape** — exactly two fields, `down` and `up`, carrying two ids and
nothing derived from them. That is what it asserts now.
## Task: close stage 1, or say what still blocks it
```task
id: CB-WP-0014-T03
CB-WP-0014-T03: hot-seat evidenced; stage 1 open on one human check CB-EV-0012. Stage 1, deliverable by deliverable: relationship-graph visualization emitted and gated, NEVER SEEN drag-to-propose evidenced end to end debug inspector evidenced (CB-WP-0011) hot-seat play evidenced here Hot-seat was the one closest to being claimed on the strength of the code path existing. SeatPolicy hands every human seat a handle on one shared Server, so turn-taking "obviously" worked — and nothing drove more than one seat until now. The property that matters is not that two turns happen but that the same tab, asked twice, shows two different hands. Mutating the projection to serve P1's view to every seat turns it red. The stage stays open on ONE named blocker rather than a vague reservation: no browser is available to this loop, so the visualization is evidenced only as correctly emitted. Everything testable from here has been tested. What remains is `cb-play --serve 0`, open the URL, confirm the table reads and a drag works. INTENT carries that note now. The self-quoting rule from CB-EV-0011 §4 is ADOPTED: an evidence file quotes the previous pass's final cost and never its own. CB-WP-0013 reported itself at $5.78/34 mid-flight; final is $8.26/47, under by 43%. Four for four, always low. Meta budget 29% [OVER] soft 25%, driven by CB-WP-0013 in a trailing three with two cheap product passes; it was an instrument repair, which ADR-0006 D2 exempts. SH-1 at 347,720 [HARD] against a 300,000 ceiling. Compaction is the remedy and this session cannot do it for itself. CB-EV-0009's standing prediction is now live and testable for the first time in three passes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 07:58:59 +02:00
status: done
priority: high
state_hub_task_id: "108c9f6d-8324-44b3-ad7e-114f9787fe38"
```
`evidence/CB-EV-0012-*.md`, and then a decision on INTENT.
**Be precise about what is and is not established.** A passing JS test
proves the input contract and the loop. It does **not** prove that an SVG
table renders legibly, that a drag feels like a drag, or that a human can
play a game — none of which any test here can reach.
State which of stage 1's four deliverables are evidenced and which rest on
inspection, then close the stage or name the remaining blocker. Carrying
it open by default is as unexamined as closing it by default was going to
be.
**Adopt or reject the self-quoting rule** proposed in CB-EV-0011 §4: an
evidence file quotes the previous pass's final cost and never its own.
This pass is the first that can apply it — CB-WP-0013's figure is now
final. Decide it rather than inheriting it.
CB-WP-0014-T03: hot-seat evidenced; stage 1 open on one human check CB-EV-0012. Stage 1, deliverable by deliverable: relationship-graph visualization emitted and gated, NEVER SEEN drag-to-propose evidenced end to end debug inspector evidenced (CB-WP-0011) hot-seat play evidenced here Hot-seat was the one closest to being claimed on the strength of the code path existing. SeatPolicy hands every human seat a handle on one shared Server, so turn-taking "obviously" worked — and nothing drove more than one seat until now. The property that matters is not that two turns happen but that the same tab, asked twice, shows two different hands. Mutating the projection to serve P1's view to every seat turns it red. The stage stays open on ONE named blocker rather than a vague reservation: no browser is available to this loop, so the visualization is evidenced only as correctly emitted. Everything testable from here has been tested. What remains is `cb-play --serve 0`, open the URL, confirm the table reads and a drag works. INTENT carries that note now. The self-quoting rule from CB-EV-0011 §4 is ADOPTED: an evidence file quotes the previous pass's final cost and never its own. CB-WP-0013 reported itself at $5.78/34 mid-flight; final is $8.26/47, under by 43%. Four for four, always low. Meta budget 29% [OVER] soft 25%, driven by CB-WP-0013 in a trailing three with two cheap product passes; it was an instrument repair, which ADR-0006 D2 exempts. SH-1 at 347,720 [HARD] against a 300,000 ceiling. Compaction is the remedy and this session cannot do it for itself. CB-EV-0009's standing prediction is now live and testable for the first time in three passes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 07:58:59 +02:00
**Done 2026-08-02.**
[CB-EV-0012](../evidence/CB-EV-0012-execute-the-javascript.md).
- **Stage 1 stays open on one named blocker**, not a vague reservation.
Three of four deliverables are evidenced by executing code; the
visualization is evidenced only as *correctly emitted*, because QuickJS
has no layout engine and no browser is available here. What remains is
one human action: `cb-play --serve 0`, open the URL, confirm the table
reads and a drag works.
- **Hot-seat was the deliverable closest to being claimed on the strength
of the code path existing.** Nothing drove more than one seat until this
pass. Now: two seats, one listener, and the projection follows the seat —
mutating it to serve P1's view to everyone turns it red.
- **The self-quoting rule is ADOPTED.** CB-WP-0013 reported itself at
$5.78/34 mid-flight; final $8.26/47, under by 43%. Four for four.
- **AM-4b is blind to 408,237 lines** — more than its own target. Found by
a pass that was using it, not auditing it.
- **SH-1 at 347,720 `[HARD]`.** CB-EV-0009's standing prediction is now
live: the next pass opened above that line should cost more than 0.123
per response.