T07: agentic realization over a stdlib browser surface

Two decisions taken with the operator: stdlib HTML driver instead of
Playwright (F-0004), and a deterministic discovery runtime instead of a live
model. Both sit behind interfaces so the alternatives drop in later.

- html.py: stdlib DOM parse and query
- agentic.py: DiscoveryRuntime (agentic arm, ignores data-td by construction)
  and RecordedSelectorRuntime (control arm, uses the strongest identifier the
  page offers)
- browser.py: per-actor sessions over real HTTP, constructed per call so no
  actor inherits another's connection state
- cost/nondeterminism metrics recorded from the first run

F-0005 (CONCEPT_DRIFT): the H-001 result is a narrowing. Where test ids are
preserved, discovery 9/9 and recorded selectors 9/9 - the semantic action buys
nothing. Where they are dropped, discovery 2/3 and recorded 0/3. The concept
model presents semantic actions as generally superior; the evidence says
conditionally superior.

M21 and M22 added mid-task: the deciding side of the axis was N=1. M22 (field
names renamed) defeats the heuristic and is the first concrete evidence that a
live model would add capability, not just cost.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 1629012@bnt-lap001
Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39
This commit is contained in:
tegwick 2026-08-22 23:50:29 +02:00
parent 925ff2dd91
commit 44faf3de8e
23 changed files with 1008 additions and 9 deletions

View file

@ -307,7 +307,7 @@ fail*. All four `td://self/...` identifiers are covered. 72 tests pass overall.
```task
id: TD-WP-0002-T07
status: todo
status: done
priority: high
state_hub_task_id: "b58cf7ce-dc18-5218-8eac-179291626467"
```
@ -324,6 +324,36 @@ run than simply asking an agent to rewrite the broken test, the crystallization
argument is an aesthetic preference rather than a value proposition. This data is
free to collect from run one and impossible to backfill.
**Done 2026-08-22.** `html.py` (stdlib DOM), `agentic.py` (two runtimes),
`browser.py` (per-actor sessions over real HTTP), `scenarios/browser_grant.py`.
107 tests pass. Two decisions taken with the operator: **stdlib HTML driver
instead of Playwright** (F-0004) and **a deterministic discovery runtime instead
of a live model**, both behind interfaces that let the alternatives drop in later.
The headline result is a narrowing, not a confirmation:
| | Discovery (agentic) | Recorded selectors (control) |
|---|---|---|
| test ids preserved (9 mutations) | 9/9 | **9/9** |
| test ids dropped (3 mutations) | 2/3 | 0/3 |
- **F-0005 — semantic actions earn their keep more narrowly than claimed.** Where
an application keeps stable identifiers, the conventional approach matches the
agentic one exactly and is cheaper, faster and deterministic. The semantic
action wins only where identifiers are absent or not carried forward. Filed as
`CONCEPT_DRIFT`: the concept model overstates this and should be revised to
match the evidence. Three mutations on the deciding side is directionally clear
and statistically nothing — recorded rather than rounded up.
- **M22 marks where a scripted runtime stops.** Renaming form fields defeats the
heuristic, but the page still carries a "Person" label a model could read. This
is the first concrete evidence that a live model would add *capability* rather
than only cost — worth more than the general argument that it might.
- Two mutations (M21, M22) were added mid-task because the deciding side of the
test-id axis was N=1 after the first run. Extending the instrument when the
evidence shows it is too thin is the intended behaviour.
- Recovery happened with **no claim or invariant diff in any run**, and the one
failure failed *loudly*`RealizationFailed` in evidence, not a silent pass.
## Adaptation detection and the defect-vs-adaptation classifier
```task