test-driver/research/hypotheses/H-001-semantic-action-stability.md

79 lines
3 KiB
Markdown
Raw Normal View History

---
id: H-001
title: Semantic Action Stability
status: EXPERIMENTING
created: "2026-08-22"
experiments: [E-001]
concepts: [C-semantic-action]
---
# H-001 — Semantic Action Stability
## Claim
A semantic action survives implementation restructuring better than a recorded UI
interaction sequence.
## Falsification condition
Across the labelled mechanical mutations in the lab (T05), a recorded interaction
sequence survives **at least as many** mutations as the semantic action does.
If mechanics-free identity buys no measurable durability, the central abstraction
is decorative and `SemanticAction` should be reduced to a naming convention.
## Measurement
Mechanical Recovery Rate for each of two arms over the same mutation set:
- **arm A** — semantic action realized by an agentic driver;
- **arm B** — a recorded selector-based sequence captured against the baseline.
Arm B is a genuine control and must be run, not assumed to fail.
## Threats to validity
The comparison is unfair if arm B is built naively — a brittle straw man makes
H-001 trivially true and worthless. Arm B uses the most robust selector strategy
reasonably available (roles, labels, test ids where the lab provides them).
## Result so far (TD-WP-0002-T07)
Twelve mechanical mutations, both arms, same semantic action:
| | Discovery (agentic) | Recorded selectors (control) |
|---|---|---|
| test ids preserved (9) | 9/9 | **9/9** |
| test ids dropped (3) | 2/3 | 0/3 |
**Not falsified, but substantially narrowed.** Where stable identifiers survive,
the control arm matches the agentic arm exactly — the semantic action buys
nothing. It earns its keep only where identifiers are absent or not carried
forward.
Three mutations on the deciding side is directionally clear and statistically
nothing. See `research/findings/F-0005-...` — the concept model overstates this
and should be revised to match the evidence.
Scope caveat (F-0004): the driver tests **structural** durability only. Visual
relayout is untested.
## Status log
- 2026-08-22 `PROPOSED`. No evidence.
- 2026-08-22 `EXPERIMENTING`. Partial evidence from T07; awaiting E-001 in full.
## Expanded measurement — 2026-09-28 (TD-WP-0003-T03)
38 mutations, including ten that drop test ids and seven new paired structural
controls. Discovery and recorded selectors both pass 17/17 preserved-id cases;
on dropped ids discovery passes 9/10 and recorded selectors 0/10. M22 remains
unrealizable. The full journey accepts 26/27 mechanical cases; FAR remains 0/7
on the unchanged defect set. Claim/provenance indices remain unchanged.
Reproduce: `PYTHONPATH=src:. python3 tools/measure_e001.py /tmp/e001.json`.
[Raw result](../evidence/2026-09-28-e001.json).
These selected and correlated synthetic HTML cases do not establish population
reliability, visual/browser coverage, model capability or economics. See
[review](../../docs/TestDriverGeneralisationReview.md) for the mechanism list and
limits. H-001 remains supported only in the narrowed identifier-loss setting.