- history/2026-08-22-concept-assessment-swot.md: SWOT assessment of the concept corpus with recommendations for the first workplan - statehub register: infotech domain, TD-WP prefix, generated AGENTS.md, .custodian-brief.md and TD-WP-0001 bootstrap workplan - .repo-classification.yaml: category research, domain infotech - SCOPE.md rewritten with real repo boundaries - TD-WP-0002: vertical spike reordering M0-M10 into one end-to-end thread that can falsify the crystallization thesis early - commit previously untracked INTENT.md and docs/ Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Assistant: claude-code Assistant-Model: opus Assistant-Process: 1629012@bnt-lap001 Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39
609 lines
14 KiB
Markdown
Executable file
609 lines
14 KiB
Markdown
Executable file
# TestDriver Improvement Loop
|
||
|
||
**Status:** Concept v0.1
|
||
**Purpose:** Establish a self-improvement loop that keeps the implementation of `test-driver` aligned with its conceptual model while generating evidence about which concepts actually work.
|
||
|
||
---
|
||
|
||
## 1. Intent
|
||
|
||
`test-driver` is itself an evolving software system. Its implementation must therefore be subject to the same principles it applies to systems under test:
|
||
|
||
- intended behavior must remain distinguishable from implementation details;
|
||
- change must generate evidence rather than silently redefine correctness;
|
||
- exploratory mechanisms should harden into deterministic mechanisms as understanding grows;
|
||
- failures should improve future verification;
|
||
- obsolete complexity should be allowed to disappear.
|
||
|
||
The improvement loop exists to continuously reconcile:
|
||
|
||
1. **Concept** — what test-driver claims should be true;
|
||
2. **Implementation** — what the framework actually does;
|
||
3. **Evidence** — what experiments and executions demonstrate.
|
||
|
||
The loop should make conceptual drift visible and turn important framework failures into durable self-verification assets.
|
||
|
||
---
|
||
|
||
## 2. Core Principle
|
||
|
||
> Every meaningful implementation increment should generate evidence about whether test-driver is becoming a better realization of its own concept.
|
||
|
||
The framework therefore has two simultaneous feedback loops:
|
||
|
||
```text
|
||
Concept / Intent / Hypotheses
|
||
|
|
||
v
|
||
Select Experiment
|
||
|
|
||
v
|
||
Implement
|
||
|
|
||
v
|
||
Exercise on SUT
|
||
|
|
||
v
|
||
Evidence
|
||
|
|
||
+----------+----------+
|
||
| |
|
||
v v
|
||
Product Finding Framework Finding
|
||
| |
|
||
v v
|
||
Improve target Improve test-driver
|
||
|
|
||
v
|
||
Reconcile with Concept
|
||
|
|
||
+-----> next cycle
|
||
```
|
||
|
||
A run may therefore produce findings about the system under test and findings about the verification framework itself.
|
||
|
||
---
|
||
|
||
## 3. Development as Experimental Work
|
||
|
||
Implementation work should increasingly be framed as hypotheses rather than feature requests.
|
||
|
||
Examples:
|
||
|
||
### H-001 — Semantic Action Stability
|
||
|
||
> A semantic action survives implementation restructuring better than a recorded UI interaction sequence.
|
||
|
||
### H-002 — Mechanical Adaptation
|
||
|
||
> An agentic driver can recover from a mechanical implementation change without modifying the semantics of the protected use case.
|
||
|
||
### H-003 — Crystallization
|
||
|
||
> A sufficiently stable agentic execution can be converted into deterministic test code without losing relevant oracle coverage.
|
||
|
||
### H-004 — Independent Judgment
|
||
|
||
> Separating actor execution from deterministic oracles reduces false-positive adaptation to defective behavior.
|
||
|
||
### H-005 — Verification Energy
|
||
|
||
> Historical evidence about defects caught, adaptations required, false positives and duplication can identify verification assets whose continued execution is more valuable than others.
|
||
|
||
Each hypothesis should have:
|
||
|
||
- an identifier;
|
||
- a claim;
|
||
- a falsification condition;
|
||
- one or more experiments;
|
||
- evidence;
|
||
- a status;
|
||
- resulting implementation or concept changes.
|
||
|
||
Suggested lifecycle:
|
||
|
||
```text
|
||
PROPOSED
|
||
|
|
||
v
|
||
EXPERIMENTING
|
||
|
|
||
+------> REJECTED
|
||
|
|
||
v
|
||
SUPPORTED
|
||
|
|
||
v
|
||
PRACTICALLY_VALIDATED
|
||
|
|
||
v
|
||
ARCHITECTURAL
|
||
```
|
||
|
||
A hypothesis may also be reopened if later evidence contradicts it.
|
||
|
||
---
|
||
|
||
## 4. Concept–Implementation Fitness Map
|
||
|
||
Every important concept should become traceable to the implementation and evidence that support it.
|
||
|
||
Conceptual relationship:
|
||
|
||
```text
|
||
Concept
|
||
|
|
||
+-- implementation
|
||
+-- experiment
|
||
+-- evidence
|
||
+-- self-verification
|
||
+-- unresolved questions
|
||
```
|
||
|
||
Example:
|
||
|
||
```yaml
|
||
concept: actor-isolation
|
||
claim: >
|
||
One actor must not obtain private state, credentials, observations,
|
||
or memory belonging to another actor except through modeled
|
||
communication channels.
|
||
|
||
implementation:
|
||
- testdriver/runtime/actor_context.py
|
||
|
||
experiments:
|
||
- H-006
|
||
|
||
self_verifications:
|
||
- td://self/actor-isolation
|
||
|
||
status: supported
|
||
```
|
||
|
||
The map should expose two forms of drift.
|
||
|
||
### Conceptual Orphaning
|
||
|
||
A concept is claimed but has no implementation or verification evidence.
|
||
|
||
### Implementation Orphaning
|
||
|
||
A subsystem or abstraction exists without a concept, requirement, or experiment that explains why it is needed.
|
||
|
||
Both should be visible during review.
|
||
|
||
---
|
||
|
||
## 5. Concept Drift
|
||
|
||
A **Concept Drift Finding** occurs when implementation behavior diverges from an established concept without an explicit decision to revise that concept.
|
||
|
||
Example:
|
||
|
||
```text
|
||
Concept:
|
||
Agents discover paths; independent oracles judge outcomes.
|
||
|
||
Implementation:
|
||
Browser agent declares the scenario successful.
|
||
|
||
Finding:
|
||
CONCEPT_DRIFT
|
||
```
|
||
|
||
A concept-drift finding must resolve in one of three ways:
|
||
|
||
1. implementation changes to match the concept;
|
||
2. the concept is deliberately revised;
|
||
3. an experiment demonstrates that the distinction is no longer useful.
|
||
|
||
Implementation reality must not silently redefine the conceptual model.
|
||
|
||
---
|
||
|
||
## 6. Self-Verification
|
||
|
||
`test-driver` should become a system under test for `test-driver`.
|
||
|
||
Self-verification use cases use the namespace:
|
||
|
||
```text
|
||
td://self/...
|
||
```
|
||
|
||
Initial candidates:
|
||
|
||
```text
|
||
td://self/actor-isolation
|
||
td://self/oracle-independence
|
||
td://self/evidence-reproducibility
|
||
td://self/mechanical-adaptation
|
||
td://self/semantic-change-detection
|
||
td://self/crystallization
|
||
td://self/test-retirement
|
||
```
|
||
|
||
These scenarios should verify framework-level promises rather than implementation internals whenever possible.
|
||
|
||
Example:
|
||
|
||
### `td://self/mechanical-adaptation`
|
||
|
||
1. execute a stable use case against lab version A;
|
||
2. change the UI mechanically without changing semantics;
|
||
3. rerun the use case;
|
||
4. allow agentic navigation to recover;
|
||
5. verify that the same semantic action and oracles remain valid;
|
||
6. record the adaptation and evidence.
|
||
|
||
Expected result:
|
||
|
||
```text
|
||
mechanical implementation change
|
||
|
|
||
v
|
||
adaptation detected
|
||
|
|
||
v
|
||
alternative realization discovered
|
||
|
|
||
v
|
||
semantic action preserved
|
||
|
|
||
v
|
||
original oracle still passes
|
||
```
|
||
|
||
---
|
||
|
||
## 7. Test-Driver Lab
|
||
|
||
A purpose-built mutable application should provide controlled evolutionary pressure for the framework.
|
||
|
||
Suggested repository or module name:
|
||
|
||
```text
|
||
test-driver-lab
|
||
```
|
||
|
||
The lab should be intentionally small but support:
|
||
|
||
- multiple users;
|
||
- organizations or tenants;
|
||
- authentication;
|
||
- resources;
|
||
- sharing;
|
||
- permissions;
|
||
- simple workflows;
|
||
- audit events;
|
||
- API interaction;
|
||
- browser interaction.
|
||
|
||
The lab should also support deliberate implementation mutations.
|
||
|
||
Examples:
|
||
|
||
```text
|
||
M01 move or rename a UI control
|
||
M02 replace the DOM structure
|
||
M03 change a compatible API representation
|
||
M04 add a legitimate workflow step
|
||
M05 introduce an authorization defect
|
||
M06 introduce eventual-consistency delay
|
||
M07 introduce intermittent dependency failure
|
||
M08 remove or deprecate a capability
|
||
M09 create a concurrency race
|
||
M10 change the intended business requirement
|
||
```
|
||
|
||
The purpose is not to build a representative product. It is to create a controlled environment in which claims about test-driver can be falsified.
|
||
|
||
---
|
||
|
||
## 8. Dual Mutation
|
||
|
||
Mutation should operate in two directions.
|
||
|
||
### Use-Case Mutation
|
||
|
||
Mutate actor, resource, order, timing, state, or permissions to test the robustness of the system under test.
|
||
|
||
```text
|
||
UseCase Mutation
|
||
|
|
||
v
|
||
tests robustness of application
|
||
```
|
||
|
||
### Implementation Mutation
|
||
|
||
Mutate the target implementation while holding the use-case semantics constant to test the robustness of test-driver.
|
||
|
||
```text
|
||
Implementation Mutation
|
||
|
|
||
v
|
||
tests robustness of test-driver
|
||
```
|
||
|
||
This duality allows the framework to test both the application and its own verification strategy.
|
||
|
||
---
|
||
|
||
## 9. Framework Findings
|
||
|
||
The framework should maintain finding classes distinct from ordinary product defects.
|
||
|
||
Initial set:
|
||
|
||
### PRODUCT_DEFECT
|
||
The system under test violates unchanged intent.
|
||
|
||
### TEST_DEFECT
|
||
The verification asset or oracle is incorrect.
|
||
|
||
### MECHANICAL_ADAPTATION
|
||
Implementation mechanics changed while protected semantics remain equivalent.
|
||
|
||
### SEMANTIC_CHANGE
|
||
The intended product behavior has changed.
|
||
|
||
### CONCEPT_DRIFT
|
||
The implementation of test-driver no longer matches an established framework concept.
|
||
|
||
### FRAMEWORK_LIMITATION
|
||
A valid scenario cannot be expressed, executed, observed, or judged adequately.
|
||
|
||
### EVIDENCE_FAILURE
|
||
A finding cannot be reproduced or supported from the retained evidence.
|
||
|
||
### UNNECESSARY_COMPLEXITY
|
||
An abstraction or subsystem adds material maintenance cost without sufficient conceptual or experimental justification.
|
||
|
||
These findings feed the improvement loop.
|
||
|
||
---
|
||
|
||
## 10. Improvement Cycle
|
||
|
||
The canonical loop is:
|
||
|
||
```text
|
||
OBSERVE
|
||
|
|
||
v
|
||
CLASSIFY
|
||
|
|
||
v
|
||
EXPLAIN
|
||
|
|
||
v
|
||
PROPOSE
|
||
|
|
||
v
|
||
EXPERIMENT
|
||
|
|
||
v
|
||
MEASURE
|
||
|
|
||
v
|
||
ACCEPT / REJECT
|
||
|
|
||
v
|
||
CRYSTALLIZE
|
||
```
|
||
|
||
### Observe
|
||
|
||
Collect evidence from test-driver runs, self-tests, implementation work and lab experiments.
|
||
|
||
### Classify
|
||
|
||
Determine whether the observation represents a product defect, test defect, adaptation, concept drift, framework limitation or other finding class.
|
||
|
||
### Explain
|
||
|
||
Produce the smallest useful causal explanation supported by evidence.
|
||
|
||
### Propose
|
||
|
||
Generate one or more candidate improvements.
|
||
|
||
### Experiment
|
||
|
||
Change one relevant variable where practical and attempt to falsify the proposed improvement.
|
||
|
||
### Measure
|
||
|
||
Evaluate the result against explicit success criteria.
|
||
|
||
### Accept / Reject
|
||
|
||
Retain improvements that produce sufficient evidence. Reject or revise those that do not.
|
||
|
||
### Crystallize
|
||
|
||
Convert learned behavior into a more deterministic, cheaper and more maintainable form whenever possible.
|
||
|
||
---
|
||
|
||
## 11. Agentic Roles
|
||
|
||
Self-improvement should not rely on one omnipotent self-modifying agent.
|
||
|
||
Distinct roles create productive tension.
|
||
|
||
### Builder
|
||
|
||
Implements the current hypothesis or improvement proposal.
|
||
|
||
### Experimenter
|
||
|
||
Designs experiments intended to falsify claims.
|
||
|
||
### Critic
|
||
|
||
Looks for false success, hidden assumptions and semantic drift.
|
||
|
||
### Auditor
|
||
|
||
Checks concept-to-implementation traceability.
|
||
|
||
### Maintainer
|
||
|
||
Looks for unnecessary abstractions, duplication and maintenance burden.
|
||
|
||
These are conceptual roles. Initially they may simply correspond to separate prompts or workflow phases.
|
||
|
||
---
|
||
|
||
## 12. Fitness Scorecard
|
||
|
||
The framework should be measured against its thesis rather than implementation volume.
|
||
|
||
Initial dimensions:
|
||
|
||
| Dimension | Example Measure |
|
||
|---|---|
|
||
| Adaptability | Mechanical changes recovered automatically |
|
||
| Semantic integrity | False semantic adaptations |
|
||
| Detection | Seeded defects correctly discovered |
|
||
| Reproducibility | Findings replayable from retained evidence |
|
||
| Crystallization | Agentic assets converted to deterministic execution |
|
||
| Efficiency | Cost per verified use case |
|
||
| Autonomy | Human interventions per 100 runs |
|
||
| Robustness | Success across controlled implementation mutations |
|
||
| Traceability | Concepts connected to implementation and evidence |
|
||
| Simplicity | Complexity required per supported capability |
|
||
|
||
A particularly important safety metric is:
|
||
|
||
> **False Adaptation Rate:** the frequency with which test-driver treats an actual product defect as a legitimate adaptation.
|
||
|
||
This should be aggressively minimized.
|
||
|
||
---
|
||
|
||
## 13. Concept Maturity
|
||
|
||
Concepts should mature based on evidence rather than attractive terminology.
|
||
|
||
Suggested levels:
|
||
|
||
```text
|
||
C0 Idea
|
||
C1 Hypothesis
|
||
C2 Experimentally Supported
|
||
C3 Practically Validated
|
||
C4 Architectural Invariant
|
||
```
|
||
|
||
Examples at the beginning of the research prototype may be approximately:
|
||
|
||
```text
|
||
Actor Isolation C2
|
||
Independent Oracles C2
|
||
Semantic Actions C1
|
||
Crystallization C1
|
||
Verification Energy C0-C1
|
||
```
|
||
|
||
These classifications are provisional and should change with evidence.
|
||
|
||
---
|
||
|
||
## 14. Compression
|
||
|
||
Self-improvement must include deletion.
|
||
|
||
Learning does not necessarily imply adding features or abstractions.
|
||
|
||
At regular intervals, perform a compression review:
|
||
|
||
- Which concepts can be merged?
|
||
- Which abstractions lack evidence?
|
||
- Which agentic mechanisms can crystallize into deterministic code?
|
||
- Which metadata has never informed a decision?
|
||
- Which subsystem can be removed?
|
||
- Which verification assets have become redundant?
|
||
|
||
The desired outcome is not maximal capability count.
|
||
|
||
It is:
|
||
|
||
> **the smallest framework that reliably realizes the validated test-driver concepts.**
|
||
|
||
---
|
||
|
||
## 15. Improvement Evidence
|
||
|
||
Every accepted framework improvement should retain:
|
||
|
||
```text
|
||
Improvement ID
|
||
Triggering finding(s)
|
||
Affected concept(s)
|
||
Hypothesis
|
||
Experiment
|
||
Before state
|
||
After state
|
||
Evidence
|
||
Measured outcome
|
||
Decision
|
||
Resulting self-verification
|
||
Resulting deterministic regression, if applicable
|
||
```
|
||
|
||
This creates a lineage from conceptual claim through evidence to implementation.
|
||
|
||
---
|
||
|
||
## 16. Initial Control Loop
|
||
|
||
The first working version does not require autonomous self-modification.
|
||
|
||
A minimal loop is sufficient:
|
||
|
||
```text
|
||
Framework run
|
||
|
|
||
v
|
||
Framework finding
|
||
|
|
||
v
|
||
Human/agent classification
|
||
|
|
||
v
|
||
Improvement hypothesis
|
||
|
|
||
v
|
||
Controlled lab experiment
|
||
|
|
||
v
|
||
Evidence
|
||
|
|
||
v
|
||
Accept / reject
|
||
|
|
||
v
|
||
Self-verification added
|
||
```
|
||
|
||
Only after this loop reliably produces good improvements should more of the process become agentic.
|
||
|
||
---
|
||
|
||
## 17. Success Condition
|
||
|
||
The improvement loop is successful when test-driver can repeatedly demonstrate that:
|
||
|
||
1. conceptual claims are traceable to implementation and evidence;
|
||
2. implementation changes that violate those claims are detected;
|
||
3. controlled experiments can distinguish defects, adaptations and semantic changes;
|
||
4. important framework failures become durable self-verifications;
|
||
5. agentic mechanisms harden into deterministic mechanisms where possible;
|
||
6. the framework becomes simpler or more effective as evidence accumulates;
|
||
7. framework evolution does not silently redefine correctness.
|
||
|
||
The long-term objective is a self-hosting verification system whose own evolution is governed by the evidence-driven principles it applies to other software.
|