test-driver/docs/TestDriverInitialMilestones.md

523 lines
12 KiB
Markdown
Raw Permalink Normal View History

# TestDriver Research Prototype — Initial Milestones
**Status:** v0.1 — **canonical milestone sequence**
**Purpose:** Establish the minimum evidence-producing development loop needed to validate the core test-driver concepts.
> This document is the single canonical source for milestone identifiers `M0``M10`.
> The competing sequences formerly in `INTENT.md` and `TestDriverConceptModel.md` § 15
> are superseded — see framework finding
> `research/findings/F-0001-milestone-sequence-drift.md`.
>
> **Canonical order is not execution order.** The milestones define *what must be
> true*; `workplans/TD-WP-0002-vertical-spike-crystallization.md` defines *the order
> in which we get there*, driving one thin thread through M1M7 rather than
> completing each milestone in turn.
---
## Milestone 0 — Research Control Plane
### Goal
Make the conceptual development of test-driver explicit, traceable and falsifiable before substantial framework code accumulates.
### Implement
- repository structure for concepts, hypotheses, experiments, findings and evidence;
- hypothesis register;
- concept maturity register;
- Concept ↔ Implementation Fitness Map;
- framework finding taxonomy;
- stable identifiers for concepts, hypotheses, experiments and findings;
- lightweight CLI or file conventions for recording decisions.
Suggested structure:
```text
research/
├── hypotheses/
├── experiments/
├── findings/
├── concepts/
└── decisions/
```
### Initial hypotheses
At minimum register:
- H-001 Semantic Action Stability
- H-002 Mechanical Adaptation
- H-003 Crystallization
- H-004 Independent Judgment
- H-005 Verification Energy
### Exit Criteria
- every major v0.1 concept has a stable identifier;
- every implemented subsystem can be linked to at least one concept or hypothesis;
- at least one hypothesis is expressed with a falsification condition and planned experiment;
- framework findings can be recorded independently from product findings.
### Evidence Produced
The first Concept ↔ Implementation Fitness Map and hypothesis register.
---
## Milestone 1 — Deterministic Semantic Kernel
### Goal
Prove the core model without agentic complexity.
### Implement
Minimal executable representations of:
- UseCase;
- Actor;
- World;
- Scenario;
- SemanticAction;
- Observation;
- Oracle;
- Verdict;
- VerificationAsset;
- Run.
Implement one deterministic driver, preferably HTTP or direct application adapter.
### Reference Use Case
```text
Alice owns resource R.
Alice grants Bob READ access.
Bob can read R.
Carol cannot read R.
Alice revokes Bob.
Bob can no longer read R.
```
### Exit Criteria
- the complete use case runs deterministically;
- Alice, Bob and Carol have isolated identities and sessions;
- all important outcomes are judged by independent deterministic oracles;
- a run produces structured evidence;
- the same scenario can be replayed from known initial state.
### Evidence Produced
The first reproducible Evidence Pack.
---
## Milestone 2 — Test-Driver Lab
### Goal
Create a controlled evolutionary environment in which test-driver claims can be deliberately challenged.
### Implement
A deliberately small application supporting:
- users;
- tenants/workspaces;
- authentication;
- resources;
- sharing;
- read/write permissions;
- revoke;
- audit history;
- HTTP API;
- minimal browser UI.
Add explicit mutation switches or tagged lab versions.
### Initial Mutations
- M01 move/rename sharing control;
- M02 restructure the DOM;
- M03 change compatible API representation;
- M04 introduce an additional legitimate workflow step;
- M05 introduce authorization defect;
- M06 introduce propagation delay.
### Exit Criteria
- the reference use case works against the baseline lab;
- each mutation can be enabled reproducibly;
- mutations can be classified as mechanical, semantic or defective;
- baseline and mutated versions retain explicit version identifiers.
### Evidence Produced
A repeatable benchmark environment for framework development.
---
## Milestone 3 — Self-Verification v0
### Goal
Make test-driver test its own foundational guarantees.
### Implement
Initial `td://self/...` verification assets:
```text
td://self/actor-isolation
td://self/oracle-independence
td://self/evidence-reproducibility
```
The self-tests should operate against observable behavior rather than internal implementation details where practical.
### Exit Criteria
- intentionally breaking actor isolation makes the corresponding self-test fail;
- allowing an actor to determine its own verdict makes oracle-independence fail;
- corrupting or omitting required run evidence makes evidence-reproducibility fail;
- each failure creates a Framework Finding.
### Evidence Produced
Proof that conceptual regressions can be detected as framework regressions.
---
## Milestone 4 — Agentic Realization
### Goal
Introduce agentic flexibility only at the realization layer while retaining deterministic truth.
### Implement
- browser driver;
- one agentic Actor Runtime;
- strict per-actor context isolation;
- semantic goal → UI realization loop;
- full action/evidence recording;
- bounded navigation and tool permissions.
Use the semantic action:
```text
grant_access(Bob, R, READ)
```
The agent may discover how to accomplish it through the UI.
The oracle must remain deterministic.
### Exit Criteria
- an agent can realize the reference semantic action from intent;
- the actor cannot access another actor's private context;
- deterministic oracles independently establish success/failure;
- agent/model/configuration identity is recorded in evidence;
- failures can be replayed sufficiently to diagnose them.
### Evidence Produced
First trustworthy agentic run.
---
## Milestone 5 — Mechanical Adaptation
### Goal
Demonstrate the core fluid-development thesis.
### Experiment
Run the same Verification Asset against:
1. lab baseline;
2. M01 moved/renamed control;
3. M02 changed DOM structure.
The use-case semantics remain unchanged.
### Implement
- adaptation detection;
- adaptation classification;
- preservation of semantic action identity;
- adaptation evidence;
- adaptation history on the Verification Asset.
### Exit Criteria
- agentic execution recovers from at least two mechanical mutations;
- original deterministic oracles remain unchanged;
- semantic intent is not modified;
- adaptation is classified as mechanical;
- the framework reports an adaptation rather than a product defect.
### Success Metric
**Mechanical Recovery Rate**
### Critical Safety Metric
**False Semantic Adaptation Rate = 0** for the experiment set.
### Evidence Produced
Support or rejection for H-001 and H-002.
---
## Milestone 6 — Defect vs. Adaptation Discrimination
### Goal
Prove that adaptive testing does not simply learn to accept broken software.
### Experiment
Use:
- M01/M02 as legitimate mechanical changes;
- M05 as an authorization defect;
- M04 or M10-style mutation as a deliberate semantic requirement change.
### Implement
Classification path:
```text
IMPLEMENTATION CHANGE
INTENT CHANGE
PRODUCT DEFECT
AMBIGUOUS
```
Add escalation for semantic changes and ambiguity.
### Exit Criteria
- mechanical changes adapt without altering claims/invariants;
- authorization defect creates a Product Finding;
- deliberate requirement change creates a Semantic Change finding;
- ambiguous evidence produces `INCONCLUSIVE` rather than silent adaptation;
- no seeded defect is normalized as adaptation.
### Success Metric
Classification precision/recall over controlled mutations.
### Evidence Produced
The first meaningful measurement of adaptation safety.
---
## Milestone 7 — Crystallization v0
### Goal
Show that agentic flexibility can harden into deterministic regression.
### Implement
- semantic action trajectory capture;
- stable-realization detection;
- deterministic candidate generation;
- candidate comparison against existing oracle set;
- provenance/lineage from agentic ancestor to deterministic descendant;
- manual acceptance step initially.
### Experiment
Run the same agentic realization repeatedly against a stable lab version, crystallize it, then execute without any model involvement.
### Exit Criteria
- one agentic Verification Asset produces a deterministic test candidate;
- deterministic execution preserves the relevant claims and oracles;
- the generated/hardened test runs with zero agentic involvement;
- lineage remains visible;
- execution cost is measurably lower than agentic execution.
### Success Metrics
- crystallization success rate;
- semantic coverage retained;
- execution cost reduction.
### Evidence Produced
Support or rejection for H-003.
---
## Milestone 8 — Framework Finding → Improvement Loop
### Goal
Close the first actual self-improvement cycle.
### Implement
Workflow:
```text
Framework Finding
Classification
Improvement Hypothesis
Controlled Experiment
Evidence
Accept / Reject
Self-Verification / Regression
```
Use a real framework weakness discovered during Milestones 17 rather than inventing one if possible.
### Exit Criteria
- a framework finding produces an explicit improvement hypothesis;
- the hypothesis is experimentally evaluated;
- the accepted change links back to concept and evidence;
- the discovered framework failure leaves behind a permanent self-verification or deterministic regression;
- the Concept ↔ Implementation Fitness Map is updated.
### Evidence Produced
The first completed **ConceptImplementation Fitness Loop**.
This is the milestone at which the self-improvement system genuinely exists.
---
## Milestone 9 — Verification Energy v0
### Goal
Begin measuring test value without prematurely optimizing the scoring model.
### Implement
Record immutable Energy Events such as:
```text
DEFECT_DETECTED
REGRESSION_CAUGHT
MECHANICAL_ADAPTATION
SEMANTIC_ADAPTATION
TEST_DEFECT
FALSE_POSITIVE
DUPLICATE
CRYSTALLIZED
USECASE_DEPRECATED
```
Initially calculate only a simple transparent score.
### Exit Criteria
- Energy is derived from event history rather than stored as unexplained state;
- every score change is explainable;
- Energy can influence campaign priority;
- criticality can override retirement;
- no automatic deletion is implemented yet.
### Evidence Produced
A dataset suitable for later testing whether Energy actually predicts verification value.
---
## Milestone 10 — First Compression Review
### Goal
Prevent the research prototype from turning into premature platform architecture.
### Review
Ask:
- Which concepts have no supporting evidence?
- Which implementation abstractions have no conceptual justification?
- Which metadata has not informed a decision?
- Which agentic behavior can now be deterministic?
- Which capabilities can be merged or removed?
- What have the experiments falsified?
### Exit Criteria
- at least one simplification is seriously evaluated;
- rejected concepts are marked as such rather than silently retained;
- architecture reflects experimental learning;
- updated Concept Model and Improvement Loop remain smaller or more precise where evidence permits.
### Evidence Produced
The first proof that self-improvement includes subtraction, not only accumulation.
---
# Recommended Execution Order
```text
M0 Research Control Plane
|
M1 Deterministic Semantic Kernel
|
M2 Test-Driver Lab
|
M3 Self-Verification v0
|
M4 Agentic Realization
|
M5 Mechanical Adaptation
|
M6 Defect vs Adaptation
|
M7 Crystallization
|
M8 Closed Improvement Loop
|
M9 Verification Energy
|
M10 Compression Review
```
The first major research gate is **M8**.
Before M8, test-driver has promising mechanisms.
At M8, it has demonstrated a complete evidence-driven self-improvement cycle.
---
# Prototype Success Gate
The initial research prototype should be considered successful enough to justify broader framework investment when it can demonstrate all of the following in one coherent system:
1. a multi-user use case expressed independently of implementation details;
2. deterministic independent oracles;
3. agentic realization of at least one semantic action;
4. recovery from legitimate mechanical implementation change;
5. rejection of a seeded semantic/security defect as a mere adaptation;
6. reproducible evidence;
7. crystallization into deterministic execution;
8. a framework failure converted into a permanent self-verification;
9. explicit concept-to-implementation traceability;
10. measured human effort and execution cost.
That demonstration is more valuable than broad feature coverage.