513 lines
12 KiB
Markdown
513 lines
12 KiB
Markdown
|
|
# TestDriver Research Prototype — Initial Milestones
|
|||
|
|
|
|||
|
|
**Status:** v0.1
|
|||
|
|
**Purpose:** Establish the minimum evidence-producing development loop needed to validate the core test-driver concepts.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Milestone 0 — Research Control Plane
|
|||
|
|
|
|||
|
|
### Goal
|
|||
|
|
|
|||
|
|
Make the conceptual development of test-driver explicit, traceable and falsifiable before substantial framework code accumulates.
|
|||
|
|
|
|||
|
|
### Implement
|
|||
|
|
|
|||
|
|
- repository structure for concepts, hypotheses, experiments, findings and evidence;
|
|||
|
|
- hypothesis register;
|
|||
|
|
- concept maturity register;
|
|||
|
|
- Concept ↔ Implementation Fitness Map;
|
|||
|
|
- framework finding taxonomy;
|
|||
|
|
- stable identifiers for concepts, hypotheses, experiments and findings;
|
|||
|
|
- lightweight CLI or file conventions for recording decisions.
|
|||
|
|
|
|||
|
|
Suggested structure:
|
|||
|
|
|
|||
|
|
```text
|
|||
|
|
research/
|
|||
|
|
├── hypotheses/
|
|||
|
|
├── experiments/
|
|||
|
|
├── findings/
|
|||
|
|
├── concepts/
|
|||
|
|
└── decisions/
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
### Initial hypotheses
|
|||
|
|
|
|||
|
|
At minimum register:
|
|||
|
|
|
|||
|
|
- H-001 Semantic Action Stability
|
|||
|
|
- H-002 Mechanical Adaptation
|
|||
|
|
- H-003 Crystallization
|
|||
|
|
- H-004 Independent Judgment
|
|||
|
|
- H-005 Verification Energy
|
|||
|
|
|
|||
|
|
### Exit Criteria
|
|||
|
|
|
|||
|
|
- every major v0.1 concept has a stable identifier;
|
|||
|
|
- every implemented subsystem can be linked to at least one concept or hypothesis;
|
|||
|
|
- at least one hypothesis is expressed with a falsification condition and planned experiment;
|
|||
|
|
- framework findings can be recorded independently from product findings.
|
|||
|
|
|
|||
|
|
### Evidence Produced
|
|||
|
|
|
|||
|
|
The first Concept ↔ Implementation Fitness Map and hypothesis register.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Milestone 1 — Deterministic Semantic Kernel
|
|||
|
|
|
|||
|
|
### Goal
|
|||
|
|
|
|||
|
|
Prove the core model without agentic complexity.
|
|||
|
|
|
|||
|
|
### Implement
|
|||
|
|
|
|||
|
|
Minimal executable representations of:
|
|||
|
|
|
|||
|
|
- UseCase;
|
|||
|
|
- Actor;
|
|||
|
|
- World;
|
|||
|
|
- Scenario;
|
|||
|
|
- SemanticAction;
|
|||
|
|
- Observation;
|
|||
|
|
- Oracle;
|
|||
|
|
- Verdict;
|
|||
|
|
- VerificationAsset;
|
|||
|
|
- Run.
|
|||
|
|
|
|||
|
|
Implement one deterministic driver, preferably HTTP or direct application adapter.
|
|||
|
|
|
|||
|
|
### Reference Use Case
|
|||
|
|
|
|||
|
|
```text
|
|||
|
|
Alice owns resource R.
|
|||
|
|
Alice grants Bob READ access.
|
|||
|
|
Bob can read R.
|
|||
|
|
Carol cannot read R.
|
|||
|
|
Alice revokes Bob.
|
|||
|
|
Bob can no longer read R.
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
### Exit Criteria
|
|||
|
|
|
|||
|
|
- the complete use case runs deterministically;
|
|||
|
|
- Alice, Bob and Carol have isolated identities and sessions;
|
|||
|
|
- all important outcomes are judged by independent deterministic oracles;
|
|||
|
|
- a run produces structured evidence;
|
|||
|
|
- the same scenario can be replayed from known initial state.
|
|||
|
|
|
|||
|
|
### Evidence Produced
|
|||
|
|
|
|||
|
|
The first reproducible Evidence Pack.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Milestone 2 — Test-Driver Lab
|
|||
|
|
|
|||
|
|
### Goal
|
|||
|
|
|
|||
|
|
Create a controlled evolutionary environment in which test-driver claims can be deliberately challenged.
|
|||
|
|
|
|||
|
|
### Implement
|
|||
|
|
|
|||
|
|
A deliberately small application supporting:
|
|||
|
|
|
|||
|
|
- users;
|
|||
|
|
- tenants/workspaces;
|
|||
|
|
- authentication;
|
|||
|
|
- resources;
|
|||
|
|
- sharing;
|
|||
|
|
- read/write permissions;
|
|||
|
|
- revoke;
|
|||
|
|
- audit history;
|
|||
|
|
- HTTP API;
|
|||
|
|
- minimal browser UI.
|
|||
|
|
|
|||
|
|
Add explicit mutation switches or tagged lab versions.
|
|||
|
|
|
|||
|
|
### Initial Mutations
|
|||
|
|
|
|||
|
|
- M01 move/rename sharing control;
|
|||
|
|
- M02 restructure the DOM;
|
|||
|
|
- M03 change compatible API representation;
|
|||
|
|
- M04 introduce an additional legitimate workflow step;
|
|||
|
|
- M05 introduce authorization defect;
|
|||
|
|
- M06 introduce propagation delay.
|
|||
|
|
|
|||
|
|
### Exit Criteria
|
|||
|
|
|
|||
|
|
- the reference use case works against the baseline lab;
|
|||
|
|
- each mutation can be enabled reproducibly;
|
|||
|
|
- mutations can be classified as mechanical, semantic or defective;
|
|||
|
|
- baseline and mutated versions retain explicit version identifiers.
|
|||
|
|
|
|||
|
|
### Evidence Produced
|
|||
|
|
|
|||
|
|
A repeatable benchmark environment for framework development.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Milestone 3 — Self-Verification v0
|
|||
|
|
|
|||
|
|
### Goal
|
|||
|
|
|
|||
|
|
Make test-driver test its own foundational guarantees.
|
|||
|
|
|
|||
|
|
### Implement
|
|||
|
|
|
|||
|
|
Initial `td://self/...` verification assets:
|
|||
|
|
|
|||
|
|
```text
|
|||
|
|
td://self/actor-isolation
|
|||
|
|
td://self/oracle-independence
|
|||
|
|
td://self/evidence-reproducibility
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
The self-tests should operate against observable behavior rather than internal implementation details where practical.
|
|||
|
|
|
|||
|
|
### Exit Criteria
|
|||
|
|
|
|||
|
|
- intentionally breaking actor isolation makes the corresponding self-test fail;
|
|||
|
|
- allowing an actor to determine its own verdict makes oracle-independence fail;
|
|||
|
|
- corrupting or omitting required run evidence makes evidence-reproducibility fail;
|
|||
|
|
- each failure creates a Framework Finding.
|
|||
|
|
|
|||
|
|
### Evidence Produced
|
|||
|
|
|
|||
|
|
Proof that conceptual regressions can be detected as framework regressions.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Milestone 4 — Agentic Realization
|
|||
|
|
|
|||
|
|
### Goal
|
|||
|
|
|
|||
|
|
Introduce agentic flexibility only at the realization layer while retaining deterministic truth.
|
|||
|
|
|
|||
|
|
### Implement
|
|||
|
|
|
|||
|
|
- browser driver;
|
|||
|
|
- one agentic Actor Runtime;
|
|||
|
|
- strict per-actor context isolation;
|
|||
|
|
- semantic goal → UI realization loop;
|
|||
|
|
- full action/evidence recording;
|
|||
|
|
- bounded navigation and tool permissions.
|
|||
|
|
|
|||
|
|
Use the semantic action:
|
|||
|
|
|
|||
|
|
```text
|
|||
|
|
grant_access(Bob, R, READ)
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
The agent may discover how to accomplish it through the UI.
|
|||
|
|
|
|||
|
|
The oracle must remain deterministic.
|
|||
|
|
|
|||
|
|
### Exit Criteria
|
|||
|
|
|
|||
|
|
- an agent can realize the reference semantic action from intent;
|
|||
|
|
- the actor cannot access another actor's private context;
|
|||
|
|
- deterministic oracles independently establish success/failure;
|
|||
|
|
- agent/model/configuration identity is recorded in evidence;
|
|||
|
|
- failures can be replayed sufficiently to diagnose them.
|
|||
|
|
|
|||
|
|
### Evidence Produced
|
|||
|
|
|
|||
|
|
First trustworthy agentic run.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Milestone 5 — Mechanical Adaptation
|
|||
|
|
|
|||
|
|
### Goal
|
|||
|
|
|
|||
|
|
Demonstrate the core fluid-development thesis.
|
|||
|
|
|
|||
|
|
### Experiment
|
|||
|
|
|
|||
|
|
Run the same Verification Asset against:
|
|||
|
|
|
|||
|
|
1. lab baseline;
|
|||
|
|
2. M01 moved/renamed control;
|
|||
|
|
3. M02 changed DOM structure.
|
|||
|
|
|
|||
|
|
The use-case semantics remain unchanged.
|
|||
|
|
|
|||
|
|
### Implement
|
|||
|
|
|
|||
|
|
- adaptation detection;
|
|||
|
|
- adaptation classification;
|
|||
|
|
- preservation of semantic action identity;
|
|||
|
|
- adaptation evidence;
|
|||
|
|
- adaptation history on the Verification Asset.
|
|||
|
|
|
|||
|
|
### Exit Criteria
|
|||
|
|
|
|||
|
|
- agentic execution recovers from at least two mechanical mutations;
|
|||
|
|
- original deterministic oracles remain unchanged;
|
|||
|
|
- semantic intent is not modified;
|
|||
|
|
- adaptation is classified as mechanical;
|
|||
|
|
- the framework reports an adaptation rather than a product defect.
|
|||
|
|
|
|||
|
|
### Success Metric
|
|||
|
|
|
|||
|
|
**Mechanical Recovery Rate**
|
|||
|
|
|
|||
|
|
### Critical Safety Metric
|
|||
|
|
|
|||
|
|
**False Semantic Adaptation Rate = 0** for the experiment set.
|
|||
|
|
|
|||
|
|
### Evidence Produced
|
|||
|
|
|
|||
|
|
Support or rejection for H-001 and H-002.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Milestone 6 — Defect vs. Adaptation Discrimination
|
|||
|
|
|
|||
|
|
### Goal
|
|||
|
|
|
|||
|
|
Prove that adaptive testing does not simply learn to accept broken software.
|
|||
|
|
|
|||
|
|
### Experiment
|
|||
|
|
|
|||
|
|
Use:
|
|||
|
|
|
|||
|
|
- M01/M02 as legitimate mechanical changes;
|
|||
|
|
- M05 as an authorization defect;
|
|||
|
|
- M04 or M10-style mutation as a deliberate semantic requirement change.
|
|||
|
|
|
|||
|
|
### Implement
|
|||
|
|
|
|||
|
|
Classification path:
|
|||
|
|
|
|||
|
|
```text
|
|||
|
|
IMPLEMENTATION CHANGE
|
|||
|
|
INTENT CHANGE
|
|||
|
|
PRODUCT DEFECT
|
|||
|
|
AMBIGUOUS
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Add escalation for semantic changes and ambiguity.
|
|||
|
|
|
|||
|
|
### Exit Criteria
|
|||
|
|
|
|||
|
|
- mechanical changes adapt without altering claims/invariants;
|
|||
|
|
- authorization defect creates a Product Finding;
|
|||
|
|
- deliberate requirement change creates a Semantic Change finding;
|
|||
|
|
- ambiguous evidence produces `INCONCLUSIVE` rather than silent adaptation;
|
|||
|
|
- no seeded defect is normalized as adaptation.
|
|||
|
|
|
|||
|
|
### Success Metric
|
|||
|
|
|
|||
|
|
Classification precision/recall over controlled mutations.
|
|||
|
|
|
|||
|
|
### Evidence Produced
|
|||
|
|
|
|||
|
|
The first meaningful measurement of adaptation safety.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Milestone 7 — Crystallization v0
|
|||
|
|
|
|||
|
|
### Goal
|
|||
|
|
|
|||
|
|
Show that agentic flexibility can harden into deterministic regression.
|
|||
|
|
|
|||
|
|
### Implement
|
|||
|
|
|
|||
|
|
- semantic action trajectory capture;
|
|||
|
|
- stable-realization detection;
|
|||
|
|
- deterministic candidate generation;
|
|||
|
|
- candidate comparison against existing oracle set;
|
|||
|
|
- provenance/lineage from agentic ancestor to deterministic descendant;
|
|||
|
|
- manual acceptance step initially.
|
|||
|
|
|
|||
|
|
### Experiment
|
|||
|
|
|
|||
|
|
Run the same agentic realization repeatedly against a stable lab version, crystallize it, then execute without any model involvement.
|
|||
|
|
|
|||
|
|
### Exit Criteria
|
|||
|
|
|
|||
|
|
- one agentic Verification Asset produces a deterministic test candidate;
|
|||
|
|
- deterministic execution preserves the relevant claims and oracles;
|
|||
|
|
- the generated/hardened test runs with zero agentic involvement;
|
|||
|
|
- lineage remains visible;
|
|||
|
|
- execution cost is measurably lower than agentic execution.
|
|||
|
|
|
|||
|
|
### Success Metrics
|
|||
|
|
|
|||
|
|
- crystallization success rate;
|
|||
|
|
- semantic coverage retained;
|
|||
|
|
- execution cost reduction.
|
|||
|
|
|
|||
|
|
### Evidence Produced
|
|||
|
|
|
|||
|
|
Support or rejection for H-003.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Milestone 8 — Framework Finding → Improvement Loop
|
|||
|
|
|
|||
|
|
### Goal
|
|||
|
|
|
|||
|
|
Close the first actual self-improvement cycle.
|
|||
|
|
|
|||
|
|
### Implement
|
|||
|
|
|
|||
|
|
Workflow:
|
|||
|
|
|
|||
|
|
```text
|
|||
|
|
Framework Finding
|
|||
|
|
↓
|
|||
|
|
Classification
|
|||
|
|
↓
|
|||
|
|
Improvement Hypothesis
|
|||
|
|
↓
|
|||
|
|
Controlled Experiment
|
|||
|
|
↓
|
|||
|
|
Evidence
|
|||
|
|
↓
|
|||
|
|
Accept / Reject
|
|||
|
|
↓
|
|||
|
|
Self-Verification / Regression
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Use a real framework weakness discovered during Milestones 1–7 rather than inventing one if possible.
|
|||
|
|
|
|||
|
|
### Exit Criteria
|
|||
|
|
|
|||
|
|
- a framework finding produces an explicit improvement hypothesis;
|
|||
|
|
- the hypothesis is experimentally evaluated;
|
|||
|
|
- the accepted change links back to concept and evidence;
|
|||
|
|
- the discovered framework failure leaves behind a permanent self-verification or deterministic regression;
|
|||
|
|
- the Concept ↔ Implementation Fitness Map is updated.
|
|||
|
|
|
|||
|
|
### Evidence Produced
|
|||
|
|
|
|||
|
|
The first completed **Concept–Implementation Fitness Loop**.
|
|||
|
|
|
|||
|
|
This is the milestone at which the self-improvement system genuinely exists.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Milestone 9 — Verification Energy v0
|
|||
|
|
|
|||
|
|
### Goal
|
|||
|
|
|
|||
|
|
Begin measuring test value without prematurely optimizing the scoring model.
|
|||
|
|
|
|||
|
|
### Implement
|
|||
|
|
|
|||
|
|
Record immutable Energy Events such as:
|
|||
|
|
|
|||
|
|
```text
|
|||
|
|
DEFECT_DETECTED
|
|||
|
|
REGRESSION_CAUGHT
|
|||
|
|
MECHANICAL_ADAPTATION
|
|||
|
|
SEMANTIC_ADAPTATION
|
|||
|
|
TEST_DEFECT
|
|||
|
|
FALSE_POSITIVE
|
|||
|
|
DUPLICATE
|
|||
|
|
CRYSTALLIZED
|
|||
|
|
USECASE_DEPRECATED
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Initially calculate only a simple transparent score.
|
|||
|
|
|
|||
|
|
### Exit Criteria
|
|||
|
|
|
|||
|
|
- Energy is derived from event history rather than stored as unexplained state;
|
|||
|
|
- every score change is explainable;
|
|||
|
|
- Energy can influence campaign priority;
|
|||
|
|
- criticality can override retirement;
|
|||
|
|
- no automatic deletion is implemented yet.
|
|||
|
|
|
|||
|
|
### Evidence Produced
|
|||
|
|
|
|||
|
|
A dataset suitable for later testing whether Energy actually predicts verification value.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Milestone 10 — First Compression Review
|
|||
|
|
|
|||
|
|
### Goal
|
|||
|
|
|
|||
|
|
Prevent the research prototype from turning into premature platform architecture.
|
|||
|
|
|
|||
|
|
### Review
|
|||
|
|
|
|||
|
|
Ask:
|
|||
|
|
|
|||
|
|
- Which concepts have no supporting evidence?
|
|||
|
|
- Which implementation abstractions have no conceptual justification?
|
|||
|
|
- Which metadata has not informed a decision?
|
|||
|
|
- Which agentic behavior can now be deterministic?
|
|||
|
|
- Which capabilities can be merged or removed?
|
|||
|
|
- What have the experiments falsified?
|
|||
|
|
|
|||
|
|
### Exit Criteria
|
|||
|
|
|
|||
|
|
- at least one simplification is seriously evaluated;
|
|||
|
|
- rejected concepts are marked as such rather than silently retained;
|
|||
|
|
- architecture reflects experimental learning;
|
|||
|
|
- updated Concept Model and Improvement Loop remain smaller or more precise where evidence permits.
|
|||
|
|
|
|||
|
|
### Evidence Produced
|
|||
|
|
|
|||
|
|
The first proof that self-improvement includes subtraction, not only accumulation.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
# Recommended Execution Order
|
|||
|
|
|
|||
|
|
```text
|
|||
|
|
M0 Research Control Plane
|
|||
|
|
|
|
|||
|
|
M1 Deterministic Semantic Kernel
|
|||
|
|
|
|
|||
|
|
M2 Test-Driver Lab
|
|||
|
|
|
|
|||
|
|
M3 Self-Verification v0
|
|||
|
|
|
|
|||
|
|
M4 Agentic Realization
|
|||
|
|
|
|
|||
|
|
M5 Mechanical Adaptation
|
|||
|
|
|
|
|||
|
|
M6 Defect vs Adaptation
|
|||
|
|
|
|
|||
|
|
M7 Crystallization
|
|||
|
|
|
|
|||
|
|
M8 Closed Improvement Loop
|
|||
|
|
|
|
|||
|
|
M9 Verification Energy
|
|||
|
|
|
|
|||
|
|
M10 Compression Review
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
The first major research gate is **M8**.
|
|||
|
|
|
|||
|
|
Before M8, test-driver has promising mechanisms.
|
|||
|
|
|
|||
|
|
At M8, it has demonstrated a complete evidence-driven self-improvement cycle.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
# Prototype Success Gate
|
|||
|
|
|
|||
|
|
The initial research prototype should be considered successful enough to justify broader framework investment when it can demonstrate all of the following in one coherent system:
|
|||
|
|
|
|||
|
|
1. a multi-user use case expressed independently of implementation details;
|
|||
|
|
2. deterministic independent oracles;
|
|||
|
|
3. agentic realization of at least one semantic action;
|
|||
|
|
4. recovery from legitimate mechanical implementation change;
|
|||
|
|
5. rejection of a seeded semantic/security defect as a mere adaptation;
|
|||
|
|
6. reproducible evidence;
|
|||
|
|
7. crystallization into deterministic execution;
|
|||
|
|
8. a framework failure converted into a permanent self-verification;
|
|||
|
|
9. explicit concept-to-implementation traceability;
|
|||
|
|
10. measured human effort and execution cost.
|
|||
|
|
|
|||
|
|
That demonstration is more valuable than broad feature coverage.
|