Complete local generalisation tasks and document remaining experiment blockers

Assistant: codex
Assistant-Model: gpt-6-astra
Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
This commit is contained in:
tegwick 2026-09-28 12:06:24 +02:00
parent d54af628df
commit cf682281d7
36 changed files with 5394 additions and 279 deletions

View file

@ -1,5 +1,8 @@
# Test Driver Concept Model
> Current scope, 2026-09-28: TD-WP-0003 removes speculative lifecycle scoring
> and declared Temperature. See [settlement and compatibility](TestDriverGeneralisationReview.md).
**Status:** v0.1 Draft
**Framework:** `test-driver`
@ -58,17 +61,10 @@ Agentic tests should not remain agentic merely because they started that way.
As behavior and implementation become stable, agentic realizations are progressively hardened and finally crystallized into deterministic tests.
### 2.5 Useful tests accumulate energy
### 2.5–2.6 Evidence, not lifecycle scores
Verification assets gain energy when they demonstrate value, for example by catching genuine failures or protecting important behavior.
They lose energy when they repeatedly become invalid, require unnecessary adaptation, become flaky, duplicate stronger verification or protect behavior that is no longer relevant.
### 2.6 Verification assets may die
Tests are not immortal repository artifacts.
When a verification asset loses relevance and reaches sufficiently low energy, it may be removed from active campaigns, archived or retired.
Retain observations and verdicts. Asset selection and retirement remain explicit
maintainer choices; there is no energy score or automated retirement policy.
### 2.7 Implementation is not the truth
@ -514,10 +510,6 @@ protects:
- least-privilege
maturity: adaptive
temperature: warm
energy:
value: 72
execution:
mode: agentic
@ -585,29 +577,10 @@ U4 Contractual
A stable use case may temporarily require agentic tests when a new implementation is introduced.
## 9.3 Temperature
## 9.3 Measured stability
**Temperature** represents implementation or capability fluidity.
Initial model:
```text
HOT actively being invented
WARM frequently changing
COOL stabilizing
COLD mature or contractual
```
Temperature influences the preferred execution strategy:
| Temperature | Preferred verification mode |
|---|---|
| HOT | exploratory and agentic |
| WARM | adaptive with emerging deterministic coverage |
| COOL | hardened regression with occasional exploration |
| COLD | predominantly deterministic |
A capability may cool as it stabilizes and heat up again during major redesign or migration.
Declared Temperature was removed on 2026-09-28 (F-0008). Crystallization uses
observed trajectory stability through `assess_stability`, not a temperature label.
## 9.4 Adaptation
@ -665,53 +638,11 @@ Explore
Crystallization is not necessarily permanent. A major implementation change may temporarily move an asset back toward adaptive or agentic execution.
## 9.7 Energy
## 9.7–9.8 Removed lifecycle scores
**Energy** expresses the current value of retaining, maintaining and executing a verification asset.
Energy should be derived from observable events rather than being an unexplained score.
Illustrative initial event model:
| Event | Energy change |
|---|---:|
| Detects confirmed product defect | +25 |
| Detects confirmed security defect | +40 |
| Prevents regression after previous defect | +20 |
| Exercises recently changed relevant capability | +2 |
| Requires harmless mechanical adaptation | -2 |
| Requires workflow adaptation | -10 |
| Test itself was wrong | -20 |
| False positive or flaky failure | -15 |
| Duplicates stronger existing verification | -20 |
| Protected use case is deprecated | -100 |
Energy history should be retained:
```yaml
energy:
value: 83
history:
- event: defect-detected
delta: 25
finding: TD-143
- event: navigation-adaptation
delta: -2
```
Energy is not synonymous with correctness.
It is better interpreted as:
> the current value of spending verification attention and resources on this asset.
## 9.8 Confidence
**Confidence** represents how strongly the framework currently believes that a verification asset faithfully represents the intended behavior it claims to protect.
Confidence should remain conceptually separate from energy.
A high-energy test may still have low confidence if its behavior is poorly specified.
Energy and Confidence are removed from the current concept set. Neither was
used in a decision across the three reference use cases. H-005 is
dormant-indefinite; there is no capture-only energy subsystem.
## 9.9 Lineage
@ -735,18 +666,10 @@ Lineage should make it possible to answer:
> Why does this test exist?
## 9.10 Retirement
## 9.10 Asset removal
A verification asset may move through:
```text
active
-> low-frequency
-> archival
-> retired
```
Energy reaching zero may trigger retirement consideration, but critical security, regulatory or contractual invariants may define a retirement floor or prohibition.
Maintainers may remove obsolete tests with an explicit review of the protected
claims. No computed Retirement concept or policy is implemented or scheduled.
---
@ -770,51 +693,11 @@ When one side changes, the framework evaluates whether the other relationships r
---
# 11. Test Metabolism
# 11–12. Removed planning abstractions
Energy, temperature, maturity and execution cost together create a **Test Metabolism**.
The framework should preferentially spend verification resources where they are most valuable.
A future campaign planner may consider:
```text
Priority =
Test Energy
x Changed-System Proximity
x Use-Case Criticality
x Risk
x Time Since Last Execution
```
The precise formula is intentionally deferred.
The conceptual requirement is that test selection should be dynamic rather than assuming every historical test has equal present value.
---
# 12. Campaign
A **Campaign** is a strategy for selecting and executing verification assets and scenario variants.
Initial campaign examples:
```text
PR smoke
regression
release qualification
authorization
tenant isolation
concurrency
resilience
exploratory
```
Campaigns manage combinatorial explosion by selecting relevant slices of:
```text
actors x roles x states x data x surfaces x schedules x mutations
```
Metabolism and Campaign were removed from the current concept set on 2026-09-28.
The experiments choose explicit scenario and mutation lists. No evidence justifies
a planner or score-driven scheduling; these are not deferred implementation promises.
---
@ -898,9 +781,6 @@ Observers --> Evidence --> Oracle Engine --> Verdict
Metadata Store:
maturity
temperature
energy
confidence
lineage
```
@ -955,17 +835,11 @@ Finding
Investigator
VerificationAsset
Maturity
Temperature
Adaptation
Hardening
Crystallization
Energy
Confidence
Lineage
Retirement
Campaign
TriadicVerification
TestMetabolism
```
---
@ -1003,15 +877,9 @@ The defining model of test-driver is:
v
Verdict
|
Finding / Value
|
v
Energy
Finding
HOT ------> WARM ------> COOL ------> COLD
| | | |
agentic adaptive hardened deterministic
+-------------- crystallization ------>
Repeated stable trajectories --> crystallization --> deterministic regression
```
The fundamental promise is: