test-driver/docs/TestDriverImprovementLoop.md
tegwick 7249c6a403 Register with State Hub, persist concept assessment, seed first workplan
- history/2026-08-22-concept-assessment-swot.md: SWOT assessment of the
  concept corpus with recommendations for the first workplan
- statehub register: infotech domain, TD-WP prefix, generated AGENTS.md,
  .custodian-brief.md and TD-WP-0001 bootstrap workplan
- .repo-classification.yaml: category research, domain infotech
- SCOPE.md rewritten with real repo boundaries
- TD-WP-0002: vertical spike reordering M0-M10 into one end-to-end thread
  that can falsify the crystallization thesis early
- commit previously untracked INTENT.md and docs/

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 1629012@bnt-lap001
Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39
2026-08-22 22:40:39 +02:00

609 lines
14 KiB
Markdown
Executable file
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# TestDriver Improvement Loop
**Status:** Concept v0.1
**Purpose:** Establish a self-improvement loop that keeps the implementation of `test-driver` aligned with its conceptual model while generating evidence about which concepts actually work.
---
## 1. Intent
`test-driver` is itself an evolving software system. Its implementation must therefore be subject to the same principles it applies to systems under test:
- intended behavior must remain distinguishable from implementation details;
- change must generate evidence rather than silently redefine correctness;
- exploratory mechanisms should harden into deterministic mechanisms as understanding grows;
- failures should improve future verification;
- obsolete complexity should be allowed to disappear.
The improvement loop exists to continuously reconcile:
1. **Concept** — what test-driver claims should be true;
2. **Implementation** — what the framework actually does;
3. **Evidence** — what experiments and executions demonstrate.
The loop should make conceptual drift visible and turn important framework failures into durable self-verification assets.
---
## 2. Core Principle
> Every meaningful implementation increment should generate evidence about whether test-driver is becoming a better realization of its own concept.
The framework therefore has two simultaneous feedback loops:
```text
Concept / Intent / Hypotheses
|
v
Select Experiment
|
v
Implement
|
v
Exercise on SUT
|
v
Evidence
|
+----------+----------+
| |
v v
Product Finding Framework Finding
| |
v v
Improve target Improve test-driver
|
v
Reconcile with Concept
|
+-----> next cycle
```
A run may therefore produce findings about the system under test and findings about the verification framework itself.
---
## 3. Development as Experimental Work
Implementation work should increasingly be framed as hypotheses rather than feature requests.
Examples:
### H-001 — Semantic Action Stability
> A semantic action survives implementation restructuring better than a recorded UI interaction sequence.
### H-002 — Mechanical Adaptation
> An agentic driver can recover from a mechanical implementation change without modifying the semantics of the protected use case.
### H-003 — Crystallization
> A sufficiently stable agentic execution can be converted into deterministic test code without losing relevant oracle coverage.
### H-004 — Independent Judgment
> Separating actor execution from deterministic oracles reduces false-positive adaptation to defective behavior.
### H-005 — Verification Energy
> Historical evidence about defects caught, adaptations required, false positives and duplication can identify verification assets whose continued execution is more valuable than others.
Each hypothesis should have:
- an identifier;
- a claim;
- a falsification condition;
- one or more experiments;
- evidence;
- a status;
- resulting implementation or concept changes.
Suggested lifecycle:
```text
PROPOSED
|
v
EXPERIMENTING
|
+------> REJECTED
|
v
SUPPORTED
|
v
PRACTICALLY_VALIDATED
|
v
ARCHITECTURAL
```
A hypothesis may also be reopened if later evidence contradicts it.
---
## 4. ConceptImplementation Fitness Map
Every important concept should become traceable to the implementation and evidence that support it.
Conceptual relationship:
```text
Concept
|
+-- implementation
+-- experiment
+-- evidence
+-- self-verification
+-- unresolved questions
```
Example:
```yaml
concept: actor-isolation
claim: >
One actor must not obtain private state, credentials, observations,
or memory belonging to another actor except through modeled
communication channels.
implementation:
- testdriver/runtime/actor_context.py
experiments:
- H-006
self_verifications:
- td://self/actor-isolation
status: supported
```
The map should expose two forms of drift.
### Conceptual Orphaning
A concept is claimed but has no implementation or verification evidence.
### Implementation Orphaning
A subsystem or abstraction exists without a concept, requirement, or experiment that explains why it is needed.
Both should be visible during review.
---
## 5. Concept Drift
A **Concept Drift Finding** occurs when implementation behavior diverges from an established concept without an explicit decision to revise that concept.
Example:
```text
Concept:
Agents discover paths; independent oracles judge outcomes.
Implementation:
Browser agent declares the scenario successful.
Finding:
CONCEPT_DRIFT
```
A concept-drift finding must resolve in one of three ways:
1. implementation changes to match the concept;
2. the concept is deliberately revised;
3. an experiment demonstrates that the distinction is no longer useful.
Implementation reality must not silently redefine the conceptual model.
---
## 6. Self-Verification
`test-driver` should become a system under test for `test-driver`.
Self-verification use cases use the namespace:
```text
td://self/...
```
Initial candidates:
```text
td://self/actor-isolation
td://self/oracle-independence
td://self/evidence-reproducibility
td://self/mechanical-adaptation
td://self/semantic-change-detection
td://self/crystallization
td://self/test-retirement
```
These scenarios should verify framework-level promises rather than implementation internals whenever possible.
Example:
### `td://self/mechanical-adaptation`
1. execute a stable use case against lab version A;
2. change the UI mechanically without changing semantics;
3. rerun the use case;
4. allow agentic navigation to recover;
5. verify that the same semantic action and oracles remain valid;
6. record the adaptation and evidence.
Expected result:
```text
mechanical implementation change
|
v
adaptation detected
|
v
alternative realization discovered
|
v
semantic action preserved
|
v
original oracle still passes
```
---
## 7. Test-Driver Lab
A purpose-built mutable application should provide controlled evolutionary pressure for the framework.
Suggested repository or module name:
```text
test-driver-lab
```
The lab should be intentionally small but support:
- multiple users;
- organizations or tenants;
- authentication;
- resources;
- sharing;
- permissions;
- simple workflows;
- audit events;
- API interaction;
- browser interaction.
The lab should also support deliberate implementation mutations.
Examples:
```text
M01 move or rename a UI control
M02 replace the DOM structure
M03 change a compatible API representation
M04 add a legitimate workflow step
M05 introduce an authorization defect
M06 introduce eventual-consistency delay
M07 introduce intermittent dependency failure
M08 remove or deprecate a capability
M09 create a concurrency race
M10 change the intended business requirement
```
The purpose is not to build a representative product. It is to create a controlled environment in which claims about test-driver can be falsified.
---
## 8. Dual Mutation
Mutation should operate in two directions.
### Use-Case Mutation
Mutate actor, resource, order, timing, state, or permissions to test the robustness of the system under test.
```text
UseCase Mutation
|
v
tests robustness of application
```
### Implementation Mutation
Mutate the target implementation while holding the use-case semantics constant to test the robustness of test-driver.
```text
Implementation Mutation
|
v
tests robustness of test-driver
```
This duality allows the framework to test both the application and its own verification strategy.
---
## 9. Framework Findings
The framework should maintain finding classes distinct from ordinary product defects.
Initial set:
### PRODUCT_DEFECT
The system under test violates unchanged intent.
### TEST_DEFECT
The verification asset or oracle is incorrect.
### MECHANICAL_ADAPTATION
Implementation mechanics changed while protected semantics remain equivalent.
### SEMANTIC_CHANGE
The intended product behavior has changed.
### CONCEPT_DRIFT
The implementation of test-driver no longer matches an established framework concept.
### FRAMEWORK_LIMITATION
A valid scenario cannot be expressed, executed, observed, or judged adequately.
### EVIDENCE_FAILURE
A finding cannot be reproduced or supported from the retained evidence.
### UNNECESSARY_COMPLEXITY
An abstraction or subsystem adds material maintenance cost without sufficient conceptual or experimental justification.
These findings feed the improvement loop.
---
## 10. Improvement Cycle
The canonical loop is:
```text
OBSERVE
|
v
CLASSIFY
|
v
EXPLAIN
|
v
PROPOSE
|
v
EXPERIMENT
|
v
MEASURE
|
v
ACCEPT / REJECT
|
v
CRYSTALLIZE
```
### Observe
Collect evidence from test-driver runs, self-tests, implementation work and lab experiments.
### Classify
Determine whether the observation represents a product defect, test defect, adaptation, concept drift, framework limitation or other finding class.
### Explain
Produce the smallest useful causal explanation supported by evidence.
### Propose
Generate one or more candidate improvements.
### Experiment
Change one relevant variable where practical and attempt to falsify the proposed improvement.
### Measure
Evaluate the result against explicit success criteria.
### Accept / Reject
Retain improvements that produce sufficient evidence. Reject or revise those that do not.
### Crystallize
Convert learned behavior into a more deterministic, cheaper and more maintainable form whenever possible.
---
## 11. Agentic Roles
Self-improvement should not rely on one omnipotent self-modifying agent.
Distinct roles create productive tension.
### Builder
Implements the current hypothesis or improvement proposal.
### Experimenter
Designs experiments intended to falsify claims.
### Critic
Looks for false success, hidden assumptions and semantic drift.
### Auditor
Checks concept-to-implementation traceability.
### Maintainer
Looks for unnecessary abstractions, duplication and maintenance burden.
These are conceptual roles. Initially they may simply correspond to separate prompts or workflow phases.
---
## 12. Fitness Scorecard
The framework should be measured against its thesis rather than implementation volume.
Initial dimensions:
| Dimension | Example Measure |
|---|---|
| Adaptability | Mechanical changes recovered automatically |
| Semantic integrity | False semantic adaptations |
| Detection | Seeded defects correctly discovered |
| Reproducibility | Findings replayable from retained evidence |
| Crystallization | Agentic assets converted to deterministic execution |
| Efficiency | Cost per verified use case |
| Autonomy | Human interventions per 100 runs |
| Robustness | Success across controlled implementation mutations |
| Traceability | Concepts connected to implementation and evidence |
| Simplicity | Complexity required per supported capability |
A particularly important safety metric is:
> **False Adaptation Rate:** the frequency with which test-driver treats an actual product defect as a legitimate adaptation.
This should be aggressively minimized.
---
## 13. Concept Maturity
Concepts should mature based on evidence rather than attractive terminology.
Suggested levels:
```text
C0 Idea
C1 Hypothesis
C2 Experimentally Supported
C3 Practically Validated
C4 Architectural Invariant
```
Examples at the beginning of the research prototype may be approximately:
```text
Actor Isolation C2
Independent Oracles C2
Semantic Actions C1
Crystallization C1
Verification Energy C0-C1
```
These classifications are provisional and should change with evidence.
---
## 14. Compression
Self-improvement must include deletion.
Learning does not necessarily imply adding features or abstractions.
At regular intervals, perform a compression review:
- Which concepts can be merged?
- Which abstractions lack evidence?
- Which agentic mechanisms can crystallize into deterministic code?
- Which metadata has never informed a decision?
- Which subsystem can be removed?
- Which verification assets have become redundant?
The desired outcome is not maximal capability count.
It is:
> **the smallest framework that reliably realizes the validated test-driver concepts.**
---
## 15. Improvement Evidence
Every accepted framework improvement should retain:
```text
Improvement ID
Triggering finding(s)
Affected concept(s)
Hypothesis
Experiment
Before state
After state
Evidence
Measured outcome
Decision
Resulting self-verification
Resulting deterministic regression, if applicable
```
This creates a lineage from conceptual claim through evidence to implementation.
---
## 16. Initial Control Loop
The first working version does not require autonomous self-modification.
A minimal loop is sufficient:
```text
Framework run
|
v
Framework finding
|
v
Human/agent classification
|
v
Improvement hypothesis
|
v
Controlled lab experiment
|
v
Evidence
|
v
Accept / reject
|
v
Self-verification added
```
Only after this loop reliably produces good improvements should more of the process become agentic.
---
## 17. Success Condition
The improvement loop is successful when test-driver can repeatedly demonstrate that:
1. conceptual claims are traceable to implementation and evidence;
2. implementation changes that violate those claims are detected;
3. controlled experiments can distinguish defects, adaptations and semantic changes;
4. important framework failures become durable self-verifications;
5. agentic mechanisms harden into deterministic mechanisms where possible;
6. the framework becomes simpler or more effective as evidence accumulates;
7. framework evolution does not silently redefine correctness.
The long-term objective is a self-hosting verification system whose own evolution is governed by the evidence-driven principles it applies to other software.