Agentic framework for integration, end2end, multiuserinteraction, security testing based on usecases.
Find a file
tegwick 44faf3de8e T07: agentic realization over a stdlib browser surface
Two decisions taken with the operator: stdlib HTML driver instead of
Playwright (F-0004), and a deterministic discovery runtime instead of a live
model. Both sit behind interfaces so the alternatives drop in later.

- html.py: stdlib DOM parse and query
- agentic.py: DiscoveryRuntime (agentic arm, ignores data-td by construction)
  and RecordedSelectorRuntime (control arm, uses the strongest identifier the
  page offers)
- browser.py: per-actor sessions over real HTTP, constructed per call so no
  actor inherits another's connection state
- cost/nondeterminism metrics recorded from the first run

F-0005 (CONCEPT_DRIFT): the H-001 result is a narrowing. Where test ids are
preserved, discovery 9/9 and recorded selectors 9/9 - the semantic action buys
nothing. Where they are dropped, discovery 2/3 and recorded 0/3. The concept
model presents semantic actions as generally superior; the evidence says
conditionally superior.

M21 and M22 added mid-task: the deciding side of the axis was N=1. M22 (field
names renamed) defeats the heuristic and is the first concrete evidence that a
live model would add capability, not just cost.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 1629012@bnt-lap001
Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39
2026-08-22 23:50:29 +02:00
docs T02: adaptation classification and intent provenance design 2026-08-22 23:09:32 +02:00
history Register with State Hub, persist concept assessment, seed first workplan 2026-08-22 22:40:39 +02:00
lab T07: agentic realization over a stdlib browser surface 2026-08-22 23:50:29 +02:00
research T07: agentic realization over a stdlib browser surface 2026-08-22 23:50:29 +02:00
scenarios T07: agentic realization over a stdlib browser surface 2026-08-22 23:50:29 +02:00
src/testdriver T07: agentic realization over a stdlib browser surface 2026-08-22 23:50:29 +02:00
tests T07: agentic realization over a stdlib browser surface 2026-08-22 23:50:29 +02:00
workplans T07: agentic realization over a stdlib browser surface 2026-08-22 23:50:29 +02:00
.custodian-brief.md chore(consistency): sync task status from DB [auto] 2026-08-22 23:38:48 +02:00
.gitignore Register with State Hub, persist concept assessment, seed first workplan 2026-08-22 22:40:39 +02:00
.repo-classification.yaml Register with State Hub, persist concept assessment, seed first workplan 2026-08-22 22:40:39 +02:00
AGENTS.md T04: deterministic semantic kernel 2026-08-22 23:21:07 +02:00
INTENT.md T01: establish canonical milestone sequence, record F-0001 2026-08-22 23:07:31 +02:00
pyproject.toml T04: deterministic semantic kernel 2026-08-22 23:21:07 +02:00
README.md T04: deterministic semantic kernel 2026-08-22 23:21:07 +02:00
SCOPE.md Register with State Hub, persist concept assessment, seed first workplan 2026-08-22 22:40:39 +02:00
WORK-RECORDS.md T07: agentic realization over a stdlib browser surface 2026-08-22 23:50:29 +02:00

test-driver

Agentic framework for integration, end-to-end, multi-user interaction and security testing, driven by use cases.

Tests mature alongside the software they protect: fluid and agentic while behaviour is changing, deterministic once it settles. See INTENT.md for the thesis and SCOPE.md for boundaries.

Status: research prototype. The deterministic kernel runs; agentic realization, adaptation classification and crystallization are not built yet. Current work: workplans/TD-WP-0002-vertical-spike-crystallization.md.

Run

python3 -m pytest -q          # the whole suite
python3 -m pytest -q tests/test_reference_scenario.py

No third-party dependencies. Python ≥ 3.11, pytest for the suite.

Layout

src/testdriver/   the kernel — intent, world, actions, drivers,
                  observers, oracles, evidence, runner
lab/              the system under test
scenarios/        reference scenarios
research/         hypotheses, experiments, findings, fitness map
docs/             concept model, improvement loop, milestones, design notes
history/          assessments and completed-work write-ups

Reading order

  1. INTENT.md — the thesis
  2. docs/TestDriverConceptModel.md — canonical concept set
  3. docs/TestDriverClassificationDesign.md — why adaptation cannot normalize a defect, and where model judgment is and is not permitted
  4. docs/TestDriverInitialMilestones.md — canonical milestones M0M10