**Purpose:** Establish a self-improvement loop that keeps the implementation of `test-driver` aligned with its conceptual model while generating evidence about which concepts actually work.
---
## 1. Intent
`test-driver` is itself an evolving software system. Its implementation must therefore be subject to the same principles it applies to systems under test:
- intended behavior must remain distinguishable from implementation details;
- change must generate evidence rather than silently redefine correctness;
- exploratory mechanisms should harden into deterministic mechanisms as understanding grows;
- failures should improve future verification;
- obsolete complexity should be allowed to disappear.
The improvement loop exists to continuously reconcile:
1.**Concept** — what test-driver claims should be true;
2.**Implementation** — what the framework actually does;
3.**Evidence** — what experiments and executions demonstrate.
The loop should make conceptual drift visible and turn important framework failures into durable self-verification assets.
---
## 2. Core Principle
> Every meaningful implementation increment should generate evidence about whether test-driver is becoming a better realization of its own concept.
The framework therefore has two simultaneous feedback loops:
```text
Concept / Intent / Hypotheses
|
v
Select Experiment
|
v
Implement
|
v
Exercise on SUT
|
v
Evidence
|
+----------+----------+
| |
v v
Product Finding Framework Finding
| |
v v
Improve target Improve test-driver
|
v
Reconcile with Concept
|
+-----> next cycle
```
A run may therefore produce findings about the system under test and findings about the verification framework itself.
---
## 3. Development as Experimental Work
Implementation work should increasingly be framed as hypotheses rather than feature requests.
Examples:
### H-001 — Semantic Action Stability
> A semantic action survives implementation restructuring better than a recorded UI interaction sequence.
### H-002 — Mechanical Adaptation
> An agentic driver can recover from a mechanical implementation change without modifying the semantics of the protected use case.
### H-003 — Crystallization
> A sufficiently stable agentic execution can be converted into deterministic test code without losing relevant oracle coverage.
### H-004 — Independent Judgment
> Separating actor execution from deterministic oracles reduces false-positive adaptation to defective behavior.
### H-005 — Verification Energy
> Historical evidence about defects caught, adaptations required, false positives and duplication can identify verification assets whose continued execution is more valuable than others.
Each hypothesis should have:
- an identifier;
- a claim;
- a falsification condition;
- one or more experiments;
- evidence;
- a status;
- resulting implementation or concept changes.
Suggested lifecycle:
```text
PROPOSED
|
v
EXPERIMENTING
|
+------> REJECTED
|
v
SUPPORTED
|
v
PRACTICALLY_VALIDATED
|
v
ARCHITECTURAL
```
A hypothesis may also be reopened if later evidence contradicts it.
---
## 4. Concept–Implementation Fitness Map
Every important concept should become traceable to the implementation and evidence that support it.
Conceptual relationship:
```text
Concept
|
+-- implementation
+-- experiment
+-- evidence
+-- self-verification
+-- unresolved questions
```
Example:
```yaml
concept: actor-isolation
claim: >
One actor must not obtain private state, credentials, observations,
or memory belonging to another actor except through modeled
communication channels.
implementation:
- testdriver/runtime/actor_context.py
experiments:
- H-006
self_verifications:
- td://self/actor-isolation
status: supported
```
The map should expose two forms of drift.
### Conceptual Orphaning
A concept is claimed but has no implementation or verification evidence.
### Implementation Orphaning
A subsystem or abstraction exists without a concept, requirement, or experiment that explains why it is needed.
Both should be visible during review.
---
## 5. Concept Drift
A **Concept Drift Finding** occurs when implementation behavior diverges from an established concept without an explicit decision to revise that concept.
| Reproducibility | Findings replayable from retained evidence |
| Crystallization | Agentic assets converted to deterministic execution |
| Efficiency | Cost per verified use case |
| Autonomy | Human interventions per 100 runs |
| Robustness | Success across controlled implementation mutations |
| Traceability | Concepts connected to implementation and evidence |
| Simplicity | Complexity required per supported capability |
A particularly important safety metric is:
> **False Adaptation Rate:** the frequency with which test-driver treats an actual product defect as a legitimate adaptation.
This should be aggressively minimized.
---
## 13. Concept Maturity
Concepts should mature based on evidence rather than attractive terminology.
Suggested levels:
```text
C0 Idea
C1 Hypothesis
C2 Experimentally Supported
C3 Practically Validated
C4 Architectural Invariant
```
Examples at the beginning of the research prototype may be approximately:
```text
Actor Isolation C2
Independent Oracles C2
Semantic Actions C1
Crystallization C1
Verification Energy C0-C1
```
These classifications are provisional and should change with evidence.
---
## 14. Compression
Self-improvement must include deletion.
Learning does not necessarily imply adding features or abstractions.
At regular intervals, perform a compression review:
- Which concepts can be merged?
- Which abstractions lack evidence?
- Which agentic mechanisms can crystallize into deterministic code?
- Which metadata has never informed a decision?
- Which subsystem can be removed?
- Which verification assets have become redundant?
The desired outcome is not maximal capability count.
It is:
> **the smallest framework that reliably realizes the validated test-driver concepts.**
---
## 15. Improvement Evidence
Every accepted framework improvement should retain:
```text
Improvement ID
Triggering finding(s)
Affected concept(s)
Hypothesis
Experiment
Before state
After state
Evidence
Measured outcome
Decision
Resulting self-verification
Resulting deterministic regression, if applicable
```
This creates a lineage from conceptual claim through evidence to implementation.
---
## 16. Initial Control Loop
The first working version does not require autonomous self-modification.
A minimal loop is sufficient:
```text
Framework run
|
v
Framework finding
|
v
Human/agent classification
|
v
Improvement hypothesis
|
v
Controlled lab experiment
|
v
Evidence
|
v
Accept / reject
|
v
Self-verification added
```
Only after this loop reliably produces good improvements should more of the process become agentic.
---
## 17. Success Condition
The improvement loop is successful when test-driver can repeatedly demonstrate that:
1. conceptual claims are traceable to implementation and evidence;
2. implementation changes that violate those claims are detected;
3. controlled experiments can distinguish defects, adaptations and semantic changes;
4. important framework failures become durable self-verifications;
5. agentic mechanisms harden into deterministic mechanisms where possible;
6. the framework becomes simpler or more effective as evidence accumulates;
7. framework evolution does not silently redefine correctness.
The long-term objective is a self-hosting verification system whose own evolution is governed by the evidence-driven principles it applies to other software.