test-driver/docs/TestDriverImprovementLoop.md
tegwick 7249c6a403 Register with State Hub, persist concept assessment, seed first workplan
- history/2026-08-22-concept-assessment-swot.md: SWOT assessment of the
  concept corpus with recommendations for the first workplan
- statehub register: infotech domain, TD-WP prefix, generated AGENTS.md,
  .custodian-brief.md and TD-WP-0001 bootstrap workplan
- .repo-classification.yaml: category research, domain infotech
- SCOPE.md rewritten with real repo boundaries
- TD-WP-0002: vertical spike reordering M0-M10 into one end-to-end thread
  that can falsify the crystallization thesis early
- commit previously untracked INTENT.md and docs/

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 1629012@bnt-lap001
Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39
2026-08-22 22:40:39 +02:00

14 KiB
Executable file
Raw Permalink Blame History

TestDriver Improvement Loop

Status: Concept v0.1
Purpose: Establish a self-improvement loop that keeps the implementation of test-driver aligned with its conceptual model while generating evidence about which concepts actually work.


1. Intent

test-driver is itself an evolving software system. Its implementation must therefore be subject to the same principles it applies to systems under test:

  • intended behavior must remain distinguishable from implementation details;
  • change must generate evidence rather than silently redefine correctness;
  • exploratory mechanisms should harden into deterministic mechanisms as understanding grows;
  • failures should improve future verification;
  • obsolete complexity should be allowed to disappear.

The improvement loop exists to continuously reconcile:

  1. Concept — what test-driver claims should be true;
  2. Implementation — what the framework actually does;
  3. Evidence — what experiments and executions demonstrate.

The loop should make conceptual drift visible and turn important framework failures into durable self-verification assets.


2. Core Principle

Every meaningful implementation increment should generate evidence about whether test-driver is becoming a better realization of its own concept.

The framework therefore has two simultaneous feedback loops:

             Concept / Intent / Hypotheses
                        |
                        v
                 Select Experiment
                        |
                        v
                    Implement
                        |
                        v
                 Exercise on SUT
                        |
                        v
                     Evidence
                        |
             +----------+----------+
             |                     |
             v                     v
       Product Finding       Framework Finding
             |                     |
             v                     v
      Improve target         Improve test-driver
                                   |
                                   v
                         Reconcile with Concept
                                   |
                                   +-----> next cycle

A run may therefore produce findings about the system under test and findings about the verification framework itself.


3. Development as Experimental Work

Implementation work should increasingly be framed as hypotheses rather than feature requests.

Examples:

H-001 — Semantic Action Stability

A semantic action survives implementation restructuring better than a recorded UI interaction sequence.

H-002 — Mechanical Adaptation

An agentic driver can recover from a mechanical implementation change without modifying the semantics of the protected use case.

H-003 — Crystallization

A sufficiently stable agentic execution can be converted into deterministic test code without losing relevant oracle coverage.

H-004 — Independent Judgment

Separating actor execution from deterministic oracles reduces false-positive adaptation to defective behavior.

H-005 — Verification Energy

Historical evidence about defects caught, adaptations required, false positives and duplication can identify verification assets whose continued execution is more valuable than others.

Each hypothesis should have:

  • an identifier;
  • a claim;
  • a falsification condition;
  • one or more experiments;
  • evidence;
  • a status;
  • resulting implementation or concept changes.

Suggested lifecycle:

PROPOSED
   |
   v
EXPERIMENTING
   |
   +------> REJECTED
   |
   v
SUPPORTED
   |
   v
PRACTICALLY_VALIDATED
   |
   v
ARCHITECTURAL

A hypothesis may also be reopened if later evidence contradicts it.


4. ConceptImplementation Fitness Map

Every important concept should become traceable to the implementation and evidence that support it.

Conceptual relationship:

Concept
   |
   +-- implementation
   +-- experiment
   +-- evidence
   +-- self-verification
   +-- unresolved questions

Example:

concept: actor-isolation
claim: >
  One actor must not obtain private state, credentials, observations,
  or memory belonging to another actor except through modeled
  communication channels.

implementation:
  - testdriver/runtime/actor_context.py

experiments:
  - H-006

self_verifications:
  - td://self/actor-isolation

status: supported

The map should expose two forms of drift.

Conceptual Orphaning

A concept is claimed but has no implementation or verification evidence.

Implementation Orphaning

A subsystem or abstraction exists without a concept, requirement, or experiment that explains why it is needed.

Both should be visible during review.


5. Concept Drift

A Concept Drift Finding occurs when implementation behavior diverges from an established concept without an explicit decision to revise that concept.

Example:

Concept:
  Agents discover paths; independent oracles judge outcomes.

Implementation:
  Browser agent declares the scenario successful.

Finding:
  CONCEPT_DRIFT

A concept-drift finding must resolve in one of three ways:

  1. implementation changes to match the concept;
  2. the concept is deliberately revised;
  3. an experiment demonstrates that the distinction is no longer useful.

Implementation reality must not silently redefine the conceptual model.


6. Self-Verification

test-driver should become a system under test for test-driver.

Self-verification use cases use the namespace:

td://self/...

Initial candidates:

td://self/actor-isolation
td://self/oracle-independence
td://self/evidence-reproducibility
td://self/mechanical-adaptation
td://self/semantic-change-detection
td://self/crystallization
td://self/test-retirement

These scenarios should verify framework-level promises rather than implementation internals whenever possible.

Example:

td://self/mechanical-adaptation

  1. execute a stable use case against lab version A;
  2. change the UI mechanically without changing semantics;
  3. rerun the use case;
  4. allow agentic navigation to recover;
  5. verify that the same semantic action and oracles remain valid;
  6. record the adaptation and evidence.

Expected result:

mechanical implementation change
        |
        v
adaptation detected
        |
        v
alternative realization discovered
        |
        v
semantic action preserved
        |
        v
original oracle still passes

7. Test-Driver Lab

A purpose-built mutable application should provide controlled evolutionary pressure for the framework.

Suggested repository or module name:

test-driver-lab

The lab should be intentionally small but support:

  • multiple users;
  • organizations or tenants;
  • authentication;
  • resources;
  • sharing;
  • permissions;
  • simple workflows;
  • audit events;
  • API interaction;
  • browser interaction.

The lab should also support deliberate implementation mutations.

Examples:

M01  move or rename a UI control
M02  replace the DOM structure
M03  change a compatible API representation
M04  add a legitimate workflow step
M05  introduce an authorization defect
M06  introduce eventual-consistency delay
M07  introduce intermittent dependency failure
M08  remove or deprecate a capability
M09  create a concurrency race
M10  change the intended business requirement

The purpose is not to build a representative product. It is to create a controlled environment in which claims about test-driver can be falsified.


8. Dual Mutation

Mutation should operate in two directions.

Use-Case Mutation

Mutate actor, resource, order, timing, state, or permissions to test the robustness of the system under test.

UseCase Mutation
      |
      v
tests robustness of application

Implementation Mutation

Mutate the target implementation while holding the use-case semantics constant to test the robustness of test-driver.

Implementation Mutation
      |
      v
tests robustness of test-driver

This duality allows the framework to test both the application and its own verification strategy.


9. Framework Findings

The framework should maintain finding classes distinct from ordinary product defects.

Initial set:

PRODUCT_DEFECT

The system under test violates unchanged intent.

TEST_DEFECT

The verification asset or oracle is incorrect.

MECHANICAL_ADAPTATION

Implementation mechanics changed while protected semantics remain equivalent.

SEMANTIC_CHANGE

The intended product behavior has changed.

CONCEPT_DRIFT

The implementation of test-driver no longer matches an established framework concept.

FRAMEWORK_LIMITATION

A valid scenario cannot be expressed, executed, observed, or judged adequately.

EVIDENCE_FAILURE

A finding cannot be reproduced or supported from the retained evidence.

UNNECESSARY_COMPLEXITY

An abstraction or subsystem adds material maintenance cost without sufficient conceptual or experimental justification.

These findings feed the improvement loop.


10. Improvement Cycle

The canonical loop is:

OBSERVE
   |
   v
CLASSIFY
   |
   v
EXPLAIN
   |
   v
PROPOSE
   |
   v
EXPERIMENT
   |
   v
MEASURE
   |
   v
ACCEPT / REJECT
   |
   v
CRYSTALLIZE

Observe

Collect evidence from test-driver runs, self-tests, implementation work and lab experiments.

Classify

Determine whether the observation represents a product defect, test defect, adaptation, concept drift, framework limitation or other finding class.

Explain

Produce the smallest useful causal explanation supported by evidence.

Propose

Generate one or more candidate improvements.

Experiment

Change one relevant variable where practical and attempt to falsify the proposed improvement.

Measure

Evaluate the result against explicit success criteria.

Accept / Reject

Retain improvements that produce sufficient evidence. Reject or revise those that do not.

Crystallize

Convert learned behavior into a more deterministic, cheaper and more maintainable form whenever possible.


11. Agentic Roles

Self-improvement should not rely on one omnipotent self-modifying agent.

Distinct roles create productive tension.

Builder

Implements the current hypothesis or improvement proposal.

Experimenter

Designs experiments intended to falsify claims.

Critic

Looks for false success, hidden assumptions and semantic drift.

Auditor

Checks concept-to-implementation traceability.

Maintainer

Looks for unnecessary abstractions, duplication and maintenance burden.

These are conceptual roles. Initially they may simply correspond to separate prompts or workflow phases.


12. Fitness Scorecard

The framework should be measured against its thesis rather than implementation volume.

Initial dimensions:

Dimension Example Measure
Adaptability Mechanical changes recovered automatically
Semantic integrity False semantic adaptations
Detection Seeded defects correctly discovered
Reproducibility Findings replayable from retained evidence
Crystallization Agentic assets converted to deterministic execution
Efficiency Cost per verified use case
Autonomy Human interventions per 100 runs
Robustness Success across controlled implementation mutations
Traceability Concepts connected to implementation and evidence
Simplicity Complexity required per supported capability

A particularly important safety metric is:

False Adaptation Rate: the frequency with which test-driver treats an actual product defect as a legitimate adaptation.

This should be aggressively minimized.


13. Concept Maturity

Concepts should mature based on evidence rather than attractive terminology.

Suggested levels:

C0  Idea
C1  Hypothesis
C2  Experimentally Supported
C3  Practically Validated
C4  Architectural Invariant

Examples at the beginning of the research prototype may be approximately:

Actor Isolation          C2
Independent Oracles      C2
Semantic Actions         C1
Crystallization          C1
Verification Energy      C0-C1

These classifications are provisional and should change with evidence.


14. Compression

Self-improvement must include deletion.

Learning does not necessarily imply adding features or abstractions.

At regular intervals, perform a compression review:

  • Which concepts can be merged?
  • Which abstractions lack evidence?
  • Which agentic mechanisms can crystallize into deterministic code?
  • Which metadata has never informed a decision?
  • Which subsystem can be removed?
  • Which verification assets have become redundant?

The desired outcome is not maximal capability count.

It is:

the smallest framework that reliably realizes the validated test-driver concepts.


15. Improvement Evidence

Every accepted framework improvement should retain:

Improvement ID
Triggering finding(s)
Affected concept(s)
Hypothesis
Experiment
Before state
After state
Evidence
Measured outcome
Decision
Resulting self-verification
Resulting deterministic regression, if applicable

This creates a lineage from conceptual claim through evidence to implementation.


16. Initial Control Loop

The first working version does not require autonomous self-modification.

A minimal loop is sufficient:

Framework run
     |
     v
Framework finding
     |
     v
Human/agent classification
     |
     v
Improvement hypothesis
     |
     v
Controlled lab experiment
     |
     v
Evidence
     |
     v
Accept / reject
     |
     v
Self-verification added

Only after this loop reliably produces good improvements should more of the process become agentic.


17. Success Condition

The improvement loop is successful when test-driver can repeatedly demonstrate that:

  1. conceptual claims are traceable to implementation and evidence;
  2. implementation changes that violate those claims are detected;
  3. controlled experiments can distinguish defects, adaptations and semantic changes;
  4. important framework failures become durable self-verifications;
  5. agentic mechanisms harden into deterministic mechanisms where possible;
  6. the framework becomes simpler or more effective as evidence accumulates;
  7. framework evolution does not silently redefine correctness.

The long-term objective is a self-hosting verification system whose own evolution is governed by the evidence-driven principles it applies to other software.