T07: agentic realization over a stdlib browser surface
Two decisions taken with the operator: stdlib HTML driver instead of Playwright (F-0004), and a deterministic discovery runtime instead of a live model. Both sit behind interfaces so the alternatives drop in later. - html.py: stdlib DOM parse and query - agentic.py: DiscoveryRuntime (agentic arm, ignores data-td by construction) and RecordedSelectorRuntime (control arm, uses the strongest identifier the page offers) - browser.py: per-actor sessions over real HTTP, constructed per call so no actor inherits another's connection state - cost/nondeterminism metrics recorded from the first run F-0005 (CONCEPT_DRIFT): the H-001 result is a narrowing. Where test ids are preserved, discovery 9/9 and recorded selectors 9/9 - the semantic action buys nothing. Where they are dropped, discovery 2/3 and recorded 0/3. The concept model presents semantic actions as generally superior; the evidence says conditionally superior. M21 and M22 added mid-task: the deciding side of the axis was N=1. M22 (field names renamed) defeats the heuristic and is the first concrete evidence that a live model would add capability, not just cost. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Assistant: claude-code Assistant-Model: opus Assistant-Process: 1629012@bnt-lap001 Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39
This commit is contained in:
parent
925ff2dd91
commit
44faf3de8e
23 changed files with 1008 additions and 9 deletions
|
|
@ -307,7 +307,7 @@ fail*. All four `td://self/...` identifiers are covered. 72 tests pass overall.
|
|||
|
||||
```task
|
||||
id: TD-WP-0002-T07
|
||||
status: todo
|
||||
status: done
|
||||
priority: high
|
||||
state_hub_task_id: "b58cf7ce-dc18-5218-8eac-179291626467"
|
||||
```
|
||||
|
|
@ -324,6 +324,36 @@ run than simply asking an agent to rewrite the broken test, the crystallization
|
|||
argument is an aesthetic preference rather than a value proposition. This data is
|
||||
free to collect from run one and impossible to backfill.
|
||||
|
||||
**Done 2026-08-22.** `html.py` (stdlib DOM), `agentic.py` (two runtimes),
|
||||
`browser.py` (per-actor sessions over real HTTP), `scenarios/browser_grant.py`.
|
||||
107 tests pass. Two decisions taken with the operator: **stdlib HTML driver
|
||||
instead of Playwright** (F-0004) and **a deterministic discovery runtime instead
|
||||
of a live model**, both behind interfaces that let the alternatives drop in later.
|
||||
|
||||
The headline result is a narrowing, not a confirmation:
|
||||
|
||||
| | Discovery (agentic) | Recorded selectors (control) |
|
||||
|---|---|---|
|
||||
| test ids preserved (9 mutations) | 9/9 | **9/9** |
|
||||
| test ids dropped (3 mutations) | 2/3 | 0/3 |
|
||||
|
||||
- **F-0005 — semantic actions earn their keep more narrowly than claimed.** Where
|
||||
an application keeps stable identifiers, the conventional approach matches the
|
||||
agentic one exactly and is cheaper, faster and deterministic. The semantic
|
||||
action wins only where identifiers are absent or not carried forward. Filed as
|
||||
`CONCEPT_DRIFT`: the concept model overstates this and should be revised to
|
||||
match the evidence. Three mutations on the deciding side is directionally clear
|
||||
and statistically nothing — recorded rather than rounded up.
|
||||
- **M22 marks where a scripted runtime stops.** Renaming form fields defeats the
|
||||
heuristic, but the page still carries a "Person" label a model could read. This
|
||||
is the first concrete evidence that a live model would add *capability* rather
|
||||
than only cost — worth more than the general argument that it might.
|
||||
- Two mutations (M21, M22) were added mid-task because the deciding side of the
|
||||
test-id axis was N=1 after the first run. Extending the instrument when the
|
||||
evidence shows it is too thin is the intended behaviour.
|
||||
- Recovery happened with **no claim or invariant diff in any run**, and the one
|
||||
failure failed *loudly* — `RealizationFailed` in evidence, not a silent pass.
|
||||
|
||||
## Adaptation detection and the defect-vs-adaptation classifier
|
||||
|
||||
```task
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue