T07: agentic realization over a stdlib browser surface

Two decisions taken with the operator: stdlib HTML driver instead of
Playwright (F-0004), and a deterministic discovery runtime instead of a live
model. Both sit behind interfaces so the alternatives drop in later.

- html.py: stdlib DOM parse and query
- agentic.py: DiscoveryRuntime (agentic arm, ignores data-td by construction)
  and RecordedSelectorRuntime (control arm, uses the strongest identifier the
  page offers)
- browser.py: per-actor sessions over real HTTP, constructed per call so no
  actor inherits another's connection state
- cost/nondeterminism metrics recorded from the first run

F-0005 (CONCEPT_DRIFT): the H-001 result is a narrowing. Where test ids are
preserved, discovery 9/9 and recorded selectors 9/9 - the semantic action buys
nothing. Where they are dropped, discovery 2/3 and recorded 0/3. The concept
model presents semantic actions as generally superior; the evidence says
conditionally superior.

M21 and M22 added mid-task: the deciding side of the axis was N=1. M22 (field
names renamed) defeats the heuristic and is the first concrete evidence that a
live model would add capability, not just cost.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 1629012@bnt-lap001
Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39
This commit is contained in:
tegwick 2026-08-22 23:50:29 +02:00
parent 925ff2dd91
commit 44faf3de8e
23 changed files with 1008 additions and 9 deletions

View file

@ -1,6 +1,6 @@
# Concept ↔ Implementation Fitness Map
**Updated:** 2026-08-22 (TD-WP-0002-T06)
**Updated:** 2026-08-22 (TD-WP-0002-T07)
Traces each important concept to the implementation, experiment and evidence that
support it. **Unsupported entries are the point of this map** — a concept with no
@ -13,8 +13,12 @@ Support levels follow `TestDriverImprovementLoop.md` §13:
## Current state
`C-semantic-action` is the first concept to reach `C2`: it has an experiment
behind it (the T07 two-arm comparison), and that experiment narrowed the claim
rather than confirming it. Everything else still rests on unit tests.
The deterministic kernel exists (T04) and its guarantees are covered by unit
tests. **Levels have not moved.** A passing unit test is not an experiment: it
tests. **Levels do not move for those.** A passing unit test is not an experiment: it
shows the code does what its author intended, not that the concept holds under
the mutations it claims to survive. Levels rise when E-001/E-002/E-003 produce
evidence, not before. The implementation column below moves; the level column
@ -26,7 +30,7 @@ were aspirational, not evidenced.
|---|---|---|---|---|---|
| `C-use-case` | C1 | `intent.py` | — | — | Is a use case expressible without leaking mechanics? |
| `C-actor-isolation` | C1 | `world.py` | E-001 | `td://self/actor-isolation` | **F-0003** — only observable when the scenario plants canaries. |
| `C-semantic-action` | C1 | `actions.py` | E-001 | — | Does identity survive restructuring better than a recorded sequence? (H-001) |
| `C-semantic-action` | **C2** | `actions.py`, `agentic.py` | E-001 (partial) | T07 arm comparison | **F-0005** — supported only where stable identifiers are absent. Narrower than the concept model claims. |
| `C-oracle-independence` | C1 | `runner.py`, `oracles.py` | E-001, E-003 | — | Independence of components ≠ independence of belief. (H-004) |
| `C-evidence-pack` | C1 | `evidence.py` | — | `td://self/evidence-reproducibility` | Verdicts are reproducible from S3 alone, on passing and failing runs. |
| `C-observation-channel` | C1 | `lab/app.py` | — | — | **D-07** — required of every system under test. Adoption cost unknown. |

View file

@ -0,0 +1,60 @@
---
id: F-0004
type: framework-finding
class: FRAMEWORK_LIMITATION
status: open
discovered: "2026-08-22"
discovered_by: TD-WP-0002-T07
workplan: TD-WP-0002
task: TD-WP-0002-T07
carried_to: TD-WP-0002-T10
---
# F-0004 — The browser surface is driven without a browser
## Decision
`TD-WP-0002-T07` names Playwright as the browser driver. It was not used.
Instead, `src/testdriver/html.py` parses the lab's server-rendered HTML with
`html.parser`, and `src/testdriver/browser.py` drives it over real HTTP.
Reasons, in order of weight:
1. **The lab's surface has no JavaScript.** It is server-rendered forms. A
browser engine would not change what H-001 or H-002 can be measured against —
the mutations that matter are DOM restructuring, rewording and field renaming,
all of which a parser faces in full.
2. **Establishing Playwright is a real cost to the operator** — installation plus
a licence review, and the usual snag is the Chromium binaries the installer
pulls rather than the library licence itself. Blocking the vertical spike on
that was not worth it for a surface that gains nothing from it.
3. **The stack is deliberately boring.** Zero dependencies remains true.
`Driver` is a Protocol, so a Playwright implementation sits alongside this one
without touching the kernel, the oracles or the evidence format.
## What this costs — stated, not buried
The driver cannot:
- execute JavaScript, so **no single-page-application surface can be driven**;
- take screenshots, so that evidence type listed in `INTENT.md` is unavailable;
- read an accessibility tree, or observe visual layout at all.
The third is the most consequential for the thesis. A real mechanical change
often moves a control *visually* while leaving the DOM largely intact, and this
driver is blind to that entire class. H-001's mutation set is therefore narrower
than the hypothesis's wording implies: it tests **structural** durability, not
**visual** durability.
That should be said plainly whenever the H-001 result is quoted.
## Carried to T10
Two questions, neither answerable now:
1. Does a Playwright driver actually change the H-001 result, or merely widen the
surface it can reach? Worth one experiment, not a rewrite.
2. Is the screenshot evidence type in `INTENT.md` load-bearing, or was it listed
because screenshots are conventional in this space? If nothing has needed one
by T10, that is a candidate for compression.

View file

@ -0,0 +1,80 @@
---
id: F-0005
type: framework-finding
class: CONCEPT_DRIFT
status: open
discovered: "2026-08-22"
discovered_by: TD-WP-0002-T07
workplan: TD-WP-0002
task: TD-WP-0002-T07
hypotheses: [H-001]
carried_to: TD-WP-0002-T10
---
# F-0005 — Semantic actions earn their keep more narrowly than claimed
## The claim as written
H-001, from `TestDriverImprovementLoop.md` § 3:
> A semantic action survives implementation restructuring better than a recorded
> UI interaction sequence.
`INTENT.md` treats this as foundational — semantic actions are "the bridge
between agentic exploration and deterministic crystallization".
## The measurement
Twelve mechanical mutations, two arms, same semantic action
(`grant_access(Bob, R, READ)`), same surface, same oracles.
| | Discovery (agentic) | Recorded selectors (control) |
|---|---|---|
| test ids **preserved** (9 mutations) | 9/9 | **9/9** |
| test ids **dropped** (3 mutations) | 2/3 | 0/3 |
The control arm is not a straw man: it uses stable `data-td` attributes, which is
what a well-instrumented application provides and what good practice recommends.
`test_the_control_arm_is_not_a_straw_man` asserts it keeps winning where those
attributes survive.
## What this actually says
**Where an application is well instrumented and keeps its identifiers, the
semantic action buys nothing.** Nine mutations, two arms, identical results. The
conventional approach is not merely adequate there — it is cheaper, faster and
deterministic.
The semantic action earns its keep in exactly one circumstance: **when stable
identifiers are absent or are not carried forward through a change.** That is a
real and common circumstance — a rewrite rarely preserves test ids, and a large
share of applications never had them — but it is much narrower than "survives
implementation restructuring better", which reads as a general claim.
## The second boundary: M22
Discovery fails on M22, where form field names change (`subject_id`
`recipient`) with test ids dropped. The heuristic runtime scores candidates
partly on field names, so renaming them removes a signal it depends on.
This is an honest limit rather than a bug. It marks where a scripted runtime
stops and where a model plausibly starts: the page still carries the label
"Person" next to the field, which a model could read and a keyword heuristic
cannot. **M22 is the first concrete piece of evidence that a live model would add
capability rather than merely cost** — worth more than a general argument that it
might.
## Consequences
1. **H-001 must never be quoted as a single rate.** Split by the test-id axis or
it is misleading. `research/hypotheses/H-001-semantic-action-stability.md` now
records the split.
2. **The concept model overstates this.** `INTENT.md` and the Concept Model
present semantic actions as generally superior. The evidence says
conditionally superior. Classified `CONCEPT_DRIFT` — the concept should be
revised to match the evidence (§ 5 path 2), not the other way round.
3. **Three mutations is still thin.** 2/3 and 0/3 are directionally clear and
statistically nothing. Any stronger statement needs more mutations on the
dropped-identifier side. Recorded rather than rounded up.
4. See also **F-0004**: this driver tests structural durability only. Visual
relayout, the other major mechanical change class, is untested entirely.

View file

@ -1,7 +1,7 @@
---
id: H-001
title: Semantic Action Stability
status: PROPOSED
status: EXPERIMENTING
created: "2026-08-22"
experiments: [E-001]
concepts: [C-semantic-action]
@ -36,6 +36,28 @@ The comparison is unfair if arm B is built naively — a brittle straw man makes
H-001 trivially true and worthless. Arm B uses the most robust selector strategy
reasonably available (roles, labels, test ids where the lab provides them).
## Result so far (TD-WP-0002-T07)
Twelve mechanical mutations, both arms, same semantic action:
| | Discovery (agentic) | Recorded selectors (control) |
|---|---|---|
| test ids preserved (9) | 9/9 | **9/9** |
| test ids dropped (3) | 2/3 | 0/3 |
**Not falsified, but substantially narrowed.** Where stable identifiers survive,
the control arm matches the agentic arm exactly — the semantic action buys
nothing. It earns its keep only where identifiers are absent or not carried
forward.
Three mutations on the deciding side is directionally clear and statistically
nothing. See `research/findings/F-0005-...` — the concept model overstates this
and should be revised to match the evidence.
Scope caveat (F-0004): the driver tests **structural** durability only. Visual
relayout is untested.
## Status log
- 2026-08-22 `PROPOSED`. No evidence.
- 2026-08-22 `EXPERIMENTING`. Partial evidence from T07; awaiting E-001 in full.

View file

@ -1,7 +1,7 @@
---
id: H-002
title: Mechanical Adaptation
status: PROPOSED
status: EXPERIMENTING
created: "2026-08-22"
experiments: [E-001]
concepts: [C-adaptation]
@ -33,8 +33,20 @@ at all.
(`docs/TestDriverClassificationDesign.md` D-02). Any non-empty diff is both a
falsification signal **and** a framework defect, since no write path should exist.
## Result so far (TD-WP-0002-T07)
Discovery recovered from 11 of 12 mechanical mutations with **no claim or
invariant diff in any run**, as D-02 requires structurally. The single failure
(M22, field names renamed) failed *loudly* — a `RealizationFailed` recorded in
evidence, not a silent pass. That distinction is the one that matters: the
framework reported that it could not act, rather than reporting that nothing was
wrong.
Full classification of recovery vs defect is T08.
## Status log
- 2026-08-22 `PROPOSED`. Design decision D-02 makes the second falsification
- 2026-08-22 `PROPOSED`.
- 2026-08-22 `EXPERIMENTING`. Recovery demonstrated; classification pending T08. Design decision D-02 makes the second falsification
branch structurally unreachable; the measurement is retained anyway, as an
assertion that the architecture is what we believe it is.