T07: agentic realization over a stdlib browser surface

Two decisions taken with the operator: stdlib HTML driver instead of
Playwright (F-0004), and a deterministic discovery runtime instead of a live
model. Both sit behind interfaces so the alternatives drop in later.

- html.py: stdlib DOM parse and query
- agentic.py: DiscoveryRuntime (agentic arm, ignores data-td by construction)
  and RecordedSelectorRuntime (control arm, uses the strongest identifier the
  page offers)
- browser.py: per-actor sessions over real HTTP, constructed per call so no
  actor inherits another's connection state
- cost/nondeterminism metrics recorded from the first run

F-0005 (CONCEPT_DRIFT): the H-001 result is a narrowing. Where test ids are
preserved, discovery 9/9 and recorded selectors 9/9 - the semantic action buys
nothing. Where they are dropped, discovery 2/3 and recorded 0/3. The concept
model presents semantic actions as generally superior; the evidence says
conditionally superior.

M21 and M22 added mid-task: the deciding side of the axis was N=1. M22 (field
names renamed) defeats the heuristic and is the first concrete evidence that a
live model would add capability, not just cost.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 1629012@bnt-lap001
Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39
This commit is contained in:
tegwick 2026-08-22 23:50:29 +02:00
parent 925ff2dd91
commit 44faf3de8e
23 changed files with 1008 additions and 9 deletions

View file

@ -1,7 +1,7 @@
---
id: H-001
title: Semantic Action Stability
status: PROPOSED
status: EXPERIMENTING
created: "2026-08-22"
experiments: [E-001]
concepts: [C-semantic-action]
@ -36,6 +36,28 @@ The comparison is unfair if arm B is built naively — a brittle straw man makes
H-001 trivially true and worthless. Arm B uses the most robust selector strategy
reasonably available (roles, labels, test ids where the lab provides them).
## Result so far (TD-WP-0002-T07)
Twelve mechanical mutations, both arms, same semantic action:
| | Discovery (agentic) | Recorded selectors (control) |
|---|---|---|
| test ids preserved (9) | 9/9 | **9/9** |
| test ids dropped (3) | 2/3 | 0/3 |
**Not falsified, but substantially narrowed.** Where stable identifiers survive,
the control arm matches the agentic arm exactly — the semantic action buys
nothing. It earns its keep only where identifiers are absent or not carried
forward.
Three mutations on the deciding side is directionally clear and statistically
nothing. See `research/findings/F-0005-...` — the concept model overstates this
and should be revised to match the evidence.
Scope caveat (F-0004): the driver tests **structural** durability only. Visual
relayout is untested.
## Status log
- 2026-08-22 `PROPOSED`. No evidence.
- 2026-08-22 `EXPERIMENTING`. Partial evidence from T07; awaiting E-001 in full.

View file

@ -1,7 +1,7 @@
---
id: H-002
title: Mechanical Adaptation
status: PROPOSED
status: EXPERIMENTING
created: "2026-08-22"
experiments: [E-001]
concepts: [C-adaptation]
@ -33,8 +33,20 @@ at all.
(`docs/TestDriverClassificationDesign.md` D-02). Any non-empty diff is both a
falsification signal **and** a framework defect, since no write path should exist.
## Result so far (TD-WP-0002-T07)
Discovery recovered from 11 of 12 mechanical mutations with **no claim or
invariant diff in any run**, as D-02 requires structurally. The single failure
(M22, field names renamed) failed *loudly* — a `RealizationFailed` recorded in
evidence, not a silent pass. That distinction is the one that matters: the
framework reported that it could not act, rather than reporting that nothing was
wrong.
Full classification of recovery vs defect is T08.
## Status log
- 2026-08-22 `PROPOSED`. Design decision D-02 makes the second falsification
- 2026-08-22 `PROPOSED`.
- 2026-08-22 `EXPERIMENTING`. Recovery demonstrated; classification pending T08. Design decision D-02 makes the second falsification
branch structurally unreachable; the measurement is retained anyway, as an
assertion that the architecture is what we believe it is.