Two decisions taken with the operator: stdlib HTML driver instead of Playwright (F-0004), and a deterministic discovery runtime instead of a live model. Both sit behind interfaces so the alternatives drop in later. - html.py: stdlib DOM parse and query - agentic.py: DiscoveryRuntime (agentic arm, ignores data-td by construction) and RecordedSelectorRuntime (control arm, uses the strongest identifier the page offers) - browser.py: per-actor sessions over real HTTP, constructed per call so no actor inherits another's connection state - cost/nondeterminism metrics recorded from the first run F-0005 (CONCEPT_DRIFT): the H-001 result is a narrowing. Where test ids are preserved, discovery 9/9 and recorded selectors 9/9 - the semantic action buys nothing. Where they are dropped, discovery 2/3 and recorded 0/3. The concept model presents semantic actions as generally superior; the evidence says conditionally superior. M21 and M22 added mid-task: the deciding side of the axis was N=1. M22 (field names renamed) defeats the heuristic and is the first concrete evidence that a live model would add capability, not just cost. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Assistant: claude-code Assistant-Model: opus Assistant-Process: 1629012@bnt-lap001 Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39 |
||
|---|---|---|
| .. | ||
| concepts | ||
| decisions | ||
| experiments | ||
| findings | ||
| hypotheses | ||
| README.md | ||
Research Control Plane
Deliberately small. This directory exists so that claims about test-driver can be falsified rather than accumulated. It is plain files — no CLI, no schema, no tooling — until there are enough readings to justify tooling.
research/
├── hypotheses/ H-NNN — a claim with a falsification condition
├── experiments/ E-NNN — a planned or executed test of a hypothesis
├── findings/ F-NNNN — findings about test-driver itself
├── concepts/ the Concept ↔ Implementation Fitness Map
└── decisions/ pointers to decisions recorded in State Hub
Identifier convention
| Prefix | Scope | Example |
|---|---|---|
H-NNN |
Hypothesis | H-001 |
E-NNN |
Experiment | E-001 |
F-NNNN |
Framework Finding | F-0001 |
C-<slug> |
Concept in the fitness map | C-actor-isolation |
D-NN |
Decision, scoped to its design note | D-07 |
TD-WP-NNNN-TNN |
Workplan task (State Hub) | TD-WP-0002-T04 |
Identifiers are stable and never reused. A rejected hypothesis keeps its number.
Hypothesis lifecycle
PROPOSED → EXPERIMENTING → SUPPORTED → PRACTICALLY_VALIDATED → ARCHITECTURAL
└──────→ REJECTED
A hypothesis may be reopened if later evidence contradicts it. Reopening is recorded in the file, not by creating a new identifier.
Rules
- Every hypothesis states what would falsify it, in terms of an observable outcome, before any experiment runs. A hypothesis with no falsification condition is an opinion.
- Concepts with no supporting evidence are marked as such, not quietly retained. The fitness map is expected to contain unsupported entries; hiding them defeats its purpose.
- Subtraction counts as progress. A rejected hypothesis or a removed abstraction is a result, not a setback.