--- id: F-0005 type: framework-finding class: CONCEPT_DRIFT status: open discovered: "2026-08-22" discovered_by: TD-WP-0002-T07 workplan: TD-WP-0002 task: TD-WP-0002-T07 hypotheses: [H-001] carried_to: TD-WP-0002-T10 --- # F-0005 — Semantic actions earn their keep more narrowly than claimed ## The claim as written H-001, from `TestDriverImprovementLoop.md` § 3: > A semantic action survives implementation restructuring better than a recorded > UI interaction sequence. `INTENT.md` treats this as foundational — semantic actions are "the bridge between agentic exploration and deterministic crystallization". ## The measurement Twelve mechanical mutations, two arms, same semantic action (`grant_access(Bob, R, READ)`), same surface, same oracles. | | Discovery (agentic) | Recorded selectors (control) | |---|---|---| | test ids **preserved** (9 mutations) | 9/9 | **9/9** | | test ids **dropped** (3 mutations) | 2/3 | 0/3 | The control arm is not a straw man: it uses stable `data-td` attributes, which is what a well-instrumented application provides and what good practice recommends. `test_the_control_arm_is_not_a_straw_man` asserts it keeps winning where those attributes survive. ## What this actually says **Where an application is well instrumented and keeps its identifiers, the semantic action buys nothing.** Nine mutations, two arms, identical results. The conventional approach is not merely adequate there — it is cheaper, faster and deterministic. The semantic action earns its keep in exactly one circumstance: **when stable identifiers are absent or are not carried forward through a change.** That is a real and common circumstance — a rewrite rarely preserves test ids, and a large share of applications never had them — but it is much narrower than "survives implementation restructuring better", which reads as a general claim. ## The second boundary: M22 Discovery fails on M22, where form field names change (`subject_id` → `recipient`) with test ids dropped. The heuristic runtime scores candidates partly on field names, so renaming them removes a signal it depends on. This is an honest limit rather than a bug. It marks where a scripted runtime stops and where a model plausibly starts: the page still carries the label "Person" next to the field, which a model could read and a keyword heuristic cannot. **M22 is the first concrete piece of evidence that a live model would add capability rather than merely cost** — worth more than a general argument that it might. ## Consequences 1. **H-001 must never be quoted as a single rate.** Split by the test-id axis or it is misleading. `research/hypotheses/H-001-semantic-action-stability.md` now records the split. 2. **The concept model overstates this.** `INTENT.md` and the Concept Model present semantic actions as generally superior. The evidence says conditionally superior. Classified `CONCEPT_DRIFT` — the concept should be revised to match the evidence (§ 5 path 2), not the other way round. 3. **Three mutations is still thin.** 2/3 and 0/3 are directionally clear and statistically nothing. Any stronger statement needs more mutations on the dropped-identifier side. Recorded rather than rounded up. 4. See also **F-0004**: this driver tests structural durability only. Visual relayout, the other major mechanical change class, is untested entirely.