Two decisions taken with the operator: stdlib HTML driver instead of Playwright (F-0004), and a deterministic discovery runtime instead of a live model. Both sit behind interfaces so the alternatives drop in later. - html.py: stdlib DOM parse and query - agentic.py: DiscoveryRuntime (agentic arm, ignores data-td by construction) and RecordedSelectorRuntime (control arm, uses the strongest identifier the page offers) - browser.py: per-actor sessions over real HTTP, constructed per call so no actor inherits another's connection state - cost/nondeterminism metrics recorded from the first run F-0005 (CONCEPT_DRIFT): the H-001 result is a narrowing. Where test ids are preserved, discovery 9/9 and recorded selectors 9/9 - the semantic action buys nothing. Where they are dropped, discovery 2/3 and recorded 0/3. The concept model presents semantic actions as generally superior; the evidence says conditionally superior. M21 and M22 added mid-task: the deciding side of the axis was N=1. M22 (field names renamed) defeats the heuristic and is the first concrete evidence that a live model would add capability, not just cost. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Assistant: claude-code Assistant-Model: opus Assistant-Process: 1629012@bnt-lap001 Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39
3.3 KiB
| id | type | class | status | discovered | discovered_by | workplan | task | hypotheses | carried_to | |
|---|---|---|---|---|---|---|---|---|---|---|
| F-0005 | framework-finding | CONCEPT_DRIFT | open | 2026-08-22 | TD-WP-0002-T07 | TD-WP-0002 | TD-WP-0002-T07 |
|
TD-WP-0002-T10 |
F-0005 — Semantic actions earn their keep more narrowly than claimed
The claim as written
H-001, from TestDriverImprovementLoop.md § 3:
A semantic action survives implementation restructuring better than a recorded UI interaction sequence.
INTENT.md treats this as foundational — semantic actions are "the bridge
between agentic exploration and deterministic crystallization".
The measurement
Twelve mechanical mutations, two arms, same semantic action
(grant_access(Bob, R, READ)), same surface, same oracles.
| Discovery (agentic) | Recorded selectors (control) | |
|---|---|---|
| test ids preserved (9 mutations) | 9/9 | 9/9 |
| test ids dropped (3 mutations) | 2/3 | 0/3 |
The control arm is not a straw man: it uses stable data-td attributes, which is
what a well-instrumented application provides and what good practice recommends.
test_the_control_arm_is_not_a_straw_man asserts it keeps winning where those
attributes survive.
What this actually says
Where an application is well instrumented and keeps its identifiers, the semantic action buys nothing. Nine mutations, two arms, identical results. The conventional approach is not merely adequate there — it is cheaper, faster and deterministic.
The semantic action earns its keep in exactly one circumstance: when stable identifiers are absent or are not carried forward through a change. That is a real and common circumstance — a rewrite rarely preserves test ids, and a large share of applications never had them — but it is much narrower than "survives implementation restructuring better", which reads as a general claim.
The second boundary: M22
Discovery fails on M22, where form field names change (subject_id →
recipient) with test ids dropped. The heuristic runtime scores candidates
partly on field names, so renaming them removes a signal it depends on.
This is an honest limit rather than a bug. It marks where a scripted runtime stops and where a model plausibly starts: the page still carries the label "Person" next to the field, which a model could read and a keyword heuristic cannot. M22 is the first concrete piece of evidence that a live model would add capability rather than merely cost — worth more than a general argument that it might.
Consequences
- H-001 must never be quoted as a single rate. Split by the test-id axis or
it is misleading.
research/hypotheses/H-001-semantic-action-stability.mdnow records the split. - The concept model overstates this.
INTENT.mdand the Concept Model present semantic actions as generally superior. The evidence says conditionally superior. ClassifiedCONCEPT_DRIFT— the concept should be revised to match the evidence (§ 5 path 2), not the other way round. - Three mutations is still thin. 2/3 and 0/3 are directionally clear and statistically nothing. Any stronger statement needs more mutations on the dropped-identifier side. Recorded rather than rounded up.
- See also F-0004: this driver tests structural durability only. Visual relayout, the other major mechanical change class, is untested entirely.