Register with State Hub, persist concept assessment, seed first workplan
- history/2026-08-22-concept-assessment-swot.md: SWOT assessment of the concept corpus with recommendations for the first workplan - statehub register: infotech domain, TD-WP prefix, generated AGENTS.md, .custodian-brief.md and TD-WP-0001 bootstrap workplan - .repo-classification.yaml: category research, domain infotech - SCOPE.md rewritten with real repo boundaries - TD-WP-0002: vertical spike reordering M0-M10 into one end-to-end thread that can falsify the crystallization thesis early - commit previously untracked INTENT.md and docs/ Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Assistant: claude-code Assistant-Model: opus Assistant-Process: 1629012@bnt-lap001 Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39
This commit is contained in:
parent
b471bed707
commit
7249c6a403
13 changed files with 3900 additions and 0 deletions
609
docs/TestDriverImprovementLoop.md
Executable file
609
docs/TestDriverImprovementLoop.md
Executable file
|
|
@ -0,0 +1,609 @@
|
|||
# TestDriver Improvement Loop
|
||||
|
||||
**Status:** Concept v0.1
|
||||
**Purpose:** Establish a self-improvement loop that keeps the implementation of `test-driver` aligned with its conceptual model while generating evidence about which concepts actually work.
|
||||
|
||||
---
|
||||
|
||||
## 1. Intent
|
||||
|
||||
`test-driver` is itself an evolving software system. Its implementation must therefore be subject to the same principles it applies to systems under test:
|
||||
|
||||
- intended behavior must remain distinguishable from implementation details;
|
||||
- change must generate evidence rather than silently redefine correctness;
|
||||
- exploratory mechanisms should harden into deterministic mechanisms as understanding grows;
|
||||
- failures should improve future verification;
|
||||
- obsolete complexity should be allowed to disappear.
|
||||
|
||||
The improvement loop exists to continuously reconcile:
|
||||
|
||||
1. **Concept** — what test-driver claims should be true;
|
||||
2. **Implementation** — what the framework actually does;
|
||||
3. **Evidence** — what experiments and executions demonstrate.
|
||||
|
||||
The loop should make conceptual drift visible and turn important framework failures into durable self-verification assets.
|
||||
|
||||
---
|
||||
|
||||
## 2. Core Principle
|
||||
|
||||
> Every meaningful implementation increment should generate evidence about whether test-driver is becoming a better realization of its own concept.
|
||||
|
||||
The framework therefore has two simultaneous feedback loops:
|
||||
|
||||
```text
|
||||
Concept / Intent / Hypotheses
|
||||
|
|
||||
v
|
||||
Select Experiment
|
||||
|
|
||||
v
|
||||
Implement
|
||||
|
|
||||
v
|
||||
Exercise on SUT
|
||||
|
|
||||
v
|
||||
Evidence
|
||||
|
|
||||
+----------+----------+
|
||||
| |
|
||||
v v
|
||||
Product Finding Framework Finding
|
||||
| |
|
||||
v v
|
||||
Improve target Improve test-driver
|
||||
|
|
||||
v
|
||||
Reconcile with Concept
|
||||
|
|
||||
+-----> next cycle
|
||||
```
|
||||
|
||||
A run may therefore produce findings about the system under test and findings about the verification framework itself.
|
||||
|
||||
---
|
||||
|
||||
## 3. Development as Experimental Work
|
||||
|
||||
Implementation work should increasingly be framed as hypotheses rather than feature requests.
|
||||
|
||||
Examples:
|
||||
|
||||
### H-001 — Semantic Action Stability
|
||||
|
||||
> A semantic action survives implementation restructuring better than a recorded UI interaction sequence.
|
||||
|
||||
### H-002 — Mechanical Adaptation
|
||||
|
||||
> An agentic driver can recover from a mechanical implementation change without modifying the semantics of the protected use case.
|
||||
|
||||
### H-003 — Crystallization
|
||||
|
||||
> A sufficiently stable agentic execution can be converted into deterministic test code without losing relevant oracle coverage.
|
||||
|
||||
### H-004 — Independent Judgment
|
||||
|
||||
> Separating actor execution from deterministic oracles reduces false-positive adaptation to defective behavior.
|
||||
|
||||
### H-005 — Verification Energy
|
||||
|
||||
> Historical evidence about defects caught, adaptations required, false positives and duplication can identify verification assets whose continued execution is more valuable than others.
|
||||
|
||||
Each hypothesis should have:
|
||||
|
||||
- an identifier;
|
||||
- a claim;
|
||||
- a falsification condition;
|
||||
- one or more experiments;
|
||||
- evidence;
|
||||
- a status;
|
||||
- resulting implementation or concept changes.
|
||||
|
||||
Suggested lifecycle:
|
||||
|
||||
```text
|
||||
PROPOSED
|
||||
|
|
||||
v
|
||||
EXPERIMENTING
|
||||
|
|
||||
+------> REJECTED
|
||||
|
|
||||
v
|
||||
SUPPORTED
|
||||
|
|
||||
v
|
||||
PRACTICALLY_VALIDATED
|
||||
|
|
||||
v
|
||||
ARCHITECTURAL
|
||||
```
|
||||
|
||||
A hypothesis may also be reopened if later evidence contradicts it.
|
||||
|
||||
---
|
||||
|
||||
## 4. Concept–Implementation Fitness Map
|
||||
|
||||
Every important concept should become traceable to the implementation and evidence that support it.
|
||||
|
||||
Conceptual relationship:
|
||||
|
||||
```text
|
||||
Concept
|
||||
|
|
||||
+-- implementation
|
||||
+-- experiment
|
||||
+-- evidence
|
||||
+-- self-verification
|
||||
+-- unresolved questions
|
||||
```
|
||||
|
||||
Example:
|
||||
|
||||
```yaml
|
||||
concept: actor-isolation
|
||||
claim: >
|
||||
One actor must not obtain private state, credentials, observations,
|
||||
or memory belonging to another actor except through modeled
|
||||
communication channels.
|
||||
|
||||
implementation:
|
||||
- testdriver/runtime/actor_context.py
|
||||
|
||||
experiments:
|
||||
- H-006
|
||||
|
||||
self_verifications:
|
||||
- td://self/actor-isolation
|
||||
|
||||
status: supported
|
||||
```
|
||||
|
||||
The map should expose two forms of drift.
|
||||
|
||||
### Conceptual Orphaning
|
||||
|
||||
A concept is claimed but has no implementation or verification evidence.
|
||||
|
||||
### Implementation Orphaning
|
||||
|
||||
A subsystem or abstraction exists without a concept, requirement, or experiment that explains why it is needed.
|
||||
|
||||
Both should be visible during review.
|
||||
|
||||
---
|
||||
|
||||
## 5. Concept Drift
|
||||
|
||||
A **Concept Drift Finding** occurs when implementation behavior diverges from an established concept without an explicit decision to revise that concept.
|
||||
|
||||
Example:
|
||||
|
||||
```text
|
||||
Concept:
|
||||
Agents discover paths; independent oracles judge outcomes.
|
||||
|
||||
Implementation:
|
||||
Browser agent declares the scenario successful.
|
||||
|
||||
Finding:
|
||||
CONCEPT_DRIFT
|
||||
```
|
||||
|
||||
A concept-drift finding must resolve in one of three ways:
|
||||
|
||||
1. implementation changes to match the concept;
|
||||
2. the concept is deliberately revised;
|
||||
3. an experiment demonstrates that the distinction is no longer useful.
|
||||
|
||||
Implementation reality must not silently redefine the conceptual model.
|
||||
|
||||
---
|
||||
|
||||
## 6. Self-Verification
|
||||
|
||||
`test-driver` should become a system under test for `test-driver`.
|
||||
|
||||
Self-verification use cases use the namespace:
|
||||
|
||||
```text
|
||||
td://self/...
|
||||
```
|
||||
|
||||
Initial candidates:
|
||||
|
||||
```text
|
||||
td://self/actor-isolation
|
||||
td://self/oracle-independence
|
||||
td://self/evidence-reproducibility
|
||||
td://self/mechanical-adaptation
|
||||
td://self/semantic-change-detection
|
||||
td://self/crystallization
|
||||
td://self/test-retirement
|
||||
```
|
||||
|
||||
These scenarios should verify framework-level promises rather than implementation internals whenever possible.
|
||||
|
||||
Example:
|
||||
|
||||
### `td://self/mechanical-adaptation`
|
||||
|
||||
1. execute a stable use case against lab version A;
|
||||
2. change the UI mechanically without changing semantics;
|
||||
3. rerun the use case;
|
||||
4. allow agentic navigation to recover;
|
||||
5. verify that the same semantic action and oracles remain valid;
|
||||
6. record the adaptation and evidence.
|
||||
|
||||
Expected result:
|
||||
|
||||
```text
|
||||
mechanical implementation change
|
||||
|
|
||||
v
|
||||
adaptation detected
|
||||
|
|
||||
v
|
||||
alternative realization discovered
|
||||
|
|
||||
v
|
||||
semantic action preserved
|
||||
|
|
||||
v
|
||||
original oracle still passes
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 7. Test-Driver Lab
|
||||
|
||||
A purpose-built mutable application should provide controlled evolutionary pressure for the framework.
|
||||
|
||||
Suggested repository or module name:
|
||||
|
||||
```text
|
||||
test-driver-lab
|
||||
```
|
||||
|
||||
The lab should be intentionally small but support:
|
||||
|
||||
- multiple users;
|
||||
- organizations or tenants;
|
||||
- authentication;
|
||||
- resources;
|
||||
- sharing;
|
||||
- permissions;
|
||||
- simple workflows;
|
||||
- audit events;
|
||||
- API interaction;
|
||||
- browser interaction.
|
||||
|
||||
The lab should also support deliberate implementation mutations.
|
||||
|
||||
Examples:
|
||||
|
||||
```text
|
||||
M01 move or rename a UI control
|
||||
M02 replace the DOM structure
|
||||
M03 change a compatible API representation
|
||||
M04 add a legitimate workflow step
|
||||
M05 introduce an authorization defect
|
||||
M06 introduce eventual-consistency delay
|
||||
M07 introduce intermittent dependency failure
|
||||
M08 remove or deprecate a capability
|
||||
M09 create a concurrency race
|
||||
M10 change the intended business requirement
|
||||
```
|
||||
|
||||
The purpose is not to build a representative product. It is to create a controlled environment in which claims about test-driver can be falsified.
|
||||
|
||||
---
|
||||
|
||||
## 8. Dual Mutation
|
||||
|
||||
Mutation should operate in two directions.
|
||||
|
||||
### Use-Case Mutation
|
||||
|
||||
Mutate actor, resource, order, timing, state, or permissions to test the robustness of the system under test.
|
||||
|
||||
```text
|
||||
UseCase Mutation
|
||||
|
|
||||
v
|
||||
tests robustness of application
|
||||
```
|
||||
|
||||
### Implementation Mutation
|
||||
|
||||
Mutate the target implementation while holding the use-case semantics constant to test the robustness of test-driver.
|
||||
|
||||
```text
|
||||
Implementation Mutation
|
||||
|
|
||||
v
|
||||
tests robustness of test-driver
|
||||
```
|
||||
|
||||
This duality allows the framework to test both the application and its own verification strategy.
|
||||
|
||||
---
|
||||
|
||||
## 9. Framework Findings
|
||||
|
||||
The framework should maintain finding classes distinct from ordinary product defects.
|
||||
|
||||
Initial set:
|
||||
|
||||
### PRODUCT_DEFECT
|
||||
The system under test violates unchanged intent.
|
||||
|
||||
### TEST_DEFECT
|
||||
The verification asset or oracle is incorrect.
|
||||
|
||||
### MECHANICAL_ADAPTATION
|
||||
Implementation mechanics changed while protected semantics remain equivalent.
|
||||
|
||||
### SEMANTIC_CHANGE
|
||||
The intended product behavior has changed.
|
||||
|
||||
### CONCEPT_DRIFT
|
||||
The implementation of test-driver no longer matches an established framework concept.
|
||||
|
||||
### FRAMEWORK_LIMITATION
|
||||
A valid scenario cannot be expressed, executed, observed, or judged adequately.
|
||||
|
||||
### EVIDENCE_FAILURE
|
||||
A finding cannot be reproduced or supported from the retained evidence.
|
||||
|
||||
### UNNECESSARY_COMPLEXITY
|
||||
An abstraction or subsystem adds material maintenance cost without sufficient conceptual or experimental justification.
|
||||
|
||||
These findings feed the improvement loop.
|
||||
|
||||
---
|
||||
|
||||
## 10. Improvement Cycle
|
||||
|
||||
The canonical loop is:
|
||||
|
||||
```text
|
||||
OBSERVE
|
||||
|
|
||||
v
|
||||
CLASSIFY
|
||||
|
|
||||
v
|
||||
EXPLAIN
|
||||
|
|
||||
v
|
||||
PROPOSE
|
||||
|
|
||||
v
|
||||
EXPERIMENT
|
||||
|
|
||||
v
|
||||
MEASURE
|
||||
|
|
||||
v
|
||||
ACCEPT / REJECT
|
||||
|
|
||||
v
|
||||
CRYSTALLIZE
|
||||
```
|
||||
|
||||
### Observe
|
||||
|
||||
Collect evidence from test-driver runs, self-tests, implementation work and lab experiments.
|
||||
|
||||
### Classify
|
||||
|
||||
Determine whether the observation represents a product defect, test defect, adaptation, concept drift, framework limitation or other finding class.
|
||||
|
||||
### Explain
|
||||
|
||||
Produce the smallest useful causal explanation supported by evidence.
|
||||
|
||||
### Propose
|
||||
|
||||
Generate one or more candidate improvements.
|
||||
|
||||
### Experiment
|
||||
|
||||
Change one relevant variable where practical and attempt to falsify the proposed improvement.
|
||||
|
||||
### Measure
|
||||
|
||||
Evaluate the result against explicit success criteria.
|
||||
|
||||
### Accept / Reject
|
||||
|
||||
Retain improvements that produce sufficient evidence. Reject or revise those that do not.
|
||||
|
||||
### Crystallize
|
||||
|
||||
Convert learned behavior into a more deterministic, cheaper and more maintainable form whenever possible.
|
||||
|
||||
---
|
||||
|
||||
## 11. Agentic Roles
|
||||
|
||||
Self-improvement should not rely on one omnipotent self-modifying agent.
|
||||
|
||||
Distinct roles create productive tension.
|
||||
|
||||
### Builder
|
||||
|
||||
Implements the current hypothesis or improvement proposal.
|
||||
|
||||
### Experimenter
|
||||
|
||||
Designs experiments intended to falsify claims.
|
||||
|
||||
### Critic
|
||||
|
||||
Looks for false success, hidden assumptions and semantic drift.
|
||||
|
||||
### Auditor
|
||||
|
||||
Checks concept-to-implementation traceability.
|
||||
|
||||
### Maintainer
|
||||
|
||||
Looks for unnecessary abstractions, duplication and maintenance burden.
|
||||
|
||||
These are conceptual roles. Initially they may simply correspond to separate prompts or workflow phases.
|
||||
|
||||
---
|
||||
|
||||
## 12. Fitness Scorecard
|
||||
|
||||
The framework should be measured against its thesis rather than implementation volume.
|
||||
|
||||
Initial dimensions:
|
||||
|
||||
| Dimension | Example Measure |
|
||||
|---|---|
|
||||
| Adaptability | Mechanical changes recovered automatically |
|
||||
| Semantic integrity | False semantic adaptations |
|
||||
| Detection | Seeded defects correctly discovered |
|
||||
| Reproducibility | Findings replayable from retained evidence |
|
||||
| Crystallization | Agentic assets converted to deterministic execution |
|
||||
| Efficiency | Cost per verified use case |
|
||||
| Autonomy | Human interventions per 100 runs |
|
||||
| Robustness | Success across controlled implementation mutations |
|
||||
| Traceability | Concepts connected to implementation and evidence |
|
||||
| Simplicity | Complexity required per supported capability |
|
||||
|
||||
A particularly important safety metric is:
|
||||
|
||||
> **False Adaptation Rate:** the frequency with which test-driver treats an actual product defect as a legitimate adaptation.
|
||||
|
||||
This should be aggressively minimized.
|
||||
|
||||
---
|
||||
|
||||
## 13. Concept Maturity
|
||||
|
||||
Concepts should mature based on evidence rather than attractive terminology.
|
||||
|
||||
Suggested levels:
|
||||
|
||||
```text
|
||||
C0 Idea
|
||||
C1 Hypothesis
|
||||
C2 Experimentally Supported
|
||||
C3 Practically Validated
|
||||
C4 Architectural Invariant
|
||||
```
|
||||
|
||||
Examples at the beginning of the research prototype may be approximately:
|
||||
|
||||
```text
|
||||
Actor Isolation C2
|
||||
Independent Oracles C2
|
||||
Semantic Actions C1
|
||||
Crystallization C1
|
||||
Verification Energy C0-C1
|
||||
```
|
||||
|
||||
These classifications are provisional and should change with evidence.
|
||||
|
||||
---
|
||||
|
||||
## 14. Compression
|
||||
|
||||
Self-improvement must include deletion.
|
||||
|
||||
Learning does not necessarily imply adding features or abstractions.
|
||||
|
||||
At regular intervals, perform a compression review:
|
||||
|
||||
- Which concepts can be merged?
|
||||
- Which abstractions lack evidence?
|
||||
- Which agentic mechanisms can crystallize into deterministic code?
|
||||
- Which metadata has never informed a decision?
|
||||
- Which subsystem can be removed?
|
||||
- Which verification assets have become redundant?
|
||||
|
||||
The desired outcome is not maximal capability count.
|
||||
|
||||
It is:
|
||||
|
||||
> **the smallest framework that reliably realizes the validated test-driver concepts.**
|
||||
|
||||
---
|
||||
|
||||
## 15. Improvement Evidence
|
||||
|
||||
Every accepted framework improvement should retain:
|
||||
|
||||
```text
|
||||
Improvement ID
|
||||
Triggering finding(s)
|
||||
Affected concept(s)
|
||||
Hypothesis
|
||||
Experiment
|
||||
Before state
|
||||
After state
|
||||
Evidence
|
||||
Measured outcome
|
||||
Decision
|
||||
Resulting self-verification
|
||||
Resulting deterministic regression, if applicable
|
||||
```
|
||||
|
||||
This creates a lineage from conceptual claim through evidence to implementation.
|
||||
|
||||
---
|
||||
|
||||
## 16. Initial Control Loop
|
||||
|
||||
The first working version does not require autonomous self-modification.
|
||||
|
||||
A minimal loop is sufficient:
|
||||
|
||||
```text
|
||||
Framework run
|
||||
|
|
||||
v
|
||||
Framework finding
|
||||
|
|
||||
v
|
||||
Human/agent classification
|
||||
|
|
||||
v
|
||||
Improvement hypothesis
|
||||
|
|
||||
v
|
||||
Controlled lab experiment
|
||||
|
|
||||
v
|
||||
Evidence
|
||||
|
|
||||
v
|
||||
Accept / reject
|
||||
|
|
||||
v
|
||||
Self-verification added
|
||||
```
|
||||
|
||||
Only after this loop reliably produces good improvements should more of the process become agentic.
|
||||
|
||||
---
|
||||
|
||||
## 17. Success Condition
|
||||
|
||||
The improvement loop is successful when test-driver can repeatedly demonstrate that:
|
||||
|
||||
1. conceptual claims are traceable to implementation and evidence;
|
||||
2. implementation changes that violate those claims are detected;
|
||||
3. controlled experiments can distinguish defects, adaptations and semantic changes;
|
||||
4. important framework failures become durable self-verifications;
|
||||
5. agentic mechanisms harden into deterministic mechanisms where possible;
|
||||
6. the framework becomes simpler or more effective as evidence accumulates;
|
||||
7. framework evolution does not silently redefine correctness.
|
||||
|
||||
The long-term objective is a self-hosting verification system whose own evolution is governed by the evidence-driven principles it applies to other software.
|
||||
Loading…
Add table
Add a link
Reference in a new issue