Complete local generalisation tasks and document remaining experiment blockers
Assistant: codex Assistant-Model: gpt-6-astra Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
This commit is contained in:
parent
d54af628df
commit
cf682281d7
36 changed files with 5394 additions and 279 deletions
|
|
@ -1,5 +1,8 @@
|
|||
# Test Driver Concept Model
|
||||
|
||||
> Current scope, 2026-09-28: TD-WP-0003 removes speculative lifecycle scoring
|
||||
> and declared Temperature. See [settlement and compatibility](TestDriverGeneralisationReview.md).
|
||||
|
||||
**Status:** v0.1 Draft
|
||||
**Framework:** `test-driver`
|
||||
|
||||
|
|
@ -58,17 +61,10 @@ Agentic tests should not remain agentic merely because they started that way.
|
|||
|
||||
As behavior and implementation become stable, agentic realizations are progressively hardened and finally crystallized into deterministic tests.
|
||||
|
||||
### 2.5 Useful tests accumulate energy
|
||||
### 2.5–2.6 Evidence, not lifecycle scores
|
||||
|
||||
Verification assets gain energy when they demonstrate value, for example by catching genuine failures or protecting important behavior.
|
||||
|
||||
They lose energy when they repeatedly become invalid, require unnecessary adaptation, become flaky, duplicate stronger verification or protect behavior that is no longer relevant.
|
||||
|
||||
### 2.6 Verification assets may die
|
||||
|
||||
Tests are not immortal repository artifacts.
|
||||
|
||||
When a verification asset loses relevance and reaches sufficiently low energy, it may be removed from active campaigns, archived or retired.
|
||||
Retain observations and verdicts. Asset selection and retirement remain explicit
|
||||
maintainer choices; there is no energy score or automated retirement policy.
|
||||
|
||||
### 2.7 Implementation is not the truth
|
||||
|
||||
|
|
@ -514,10 +510,6 @@ protects:
|
|||
- least-privilege
|
||||
|
||||
maturity: adaptive
|
||||
temperature: warm
|
||||
|
||||
energy:
|
||||
value: 72
|
||||
|
||||
execution:
|
||||
mode: agentic
|
||||
|
|
@ -585,29 +577,10 @@ U4 Contractual
|
|||
|
||||
A stable use case may temporarily require agentic tests when a new implementation is introduced.
|
||||
|
||||
## 9.3 Temperature
|
||||
## 9.3 Measured stability
|
||||
|
||||
**Temperature** represents implementation or capability fluidity.
|
||||
|
||||
Initial model:
|
||||
|
||||
```text
|
||||
HOT actively being invented
|
||||
WARM frequently changing
|
||||
COOL stabilizing
|
||||
COLD mature or contractual
|
||||
```
|
||||
|
||||
Temperature influences the preferred execution strategy:
|
||||
|
||||
| Temperature | Preferred verification mode |
|
||||
|---|---|
|
||||
| HOT | exploratory and agentic |
|
||||
| WARM | adaptive with emerging deterministic coverage |
|
||||
| COOL | hardened regression with occasional exploration |
|
||||
| COLD | predominantly deterministic |
|
||||
|
||||
A capability may cool as it stabilizes and heat up again during major redesign or migration.
|
||||
Declared Temperature was removed on 2026-09-28 (F-0008). Crystallization uses
|
||||
observed trajectory stability through `assess_stability`, not a temperature label.
|
||||
|
||||
## 9.4 Adaptation
|
||||
|
||||
|
|
@ -665,53 +638,11 @@ Explore
|
|||
|
||||
Crystallization is not necessarily permanent. A major implementation change may temporarily move an asset back toward adaptive or agentic execution.
|
||||
|
||||
## 9.7 Energy
|
||||
## 9.7–9.8 Removed lifecycle scores
|
||||
|
||||
**Energy** expresses the current value of retaining, maintaining and executing a verification asset.
|
||||
|
||||
Energy should be derived from observable events rather than being an unexplained score.
|
||||
|
||||
Illustrative initial event model:
|
||||
|
||||
| Event | Energy change |
|
||||
|---|---:|
|
||||
| Detects confirmed product defect | +25 |
|
||||
| Detects confirmed security defect | +40 |
|
||||
| Prevents regression after previous defect | +20 |
|
||||
| Exercises recently changed relevant capability | +2 |
|
||||
| Requires harmless mechanical adaptation | -2 |
|
||||
| Requires workflow adaptation | -10 |
|
||||
| Test itself was wrong | -20 |
|
||||
| False positive or flaky failure | -15 |
|
||||
| Duplicates stronger existing verification | -20 |
|
||||
| Protected use case is deprecated | -100 |
|
||||
|
||||
Energy history should be retained:
|
||||
|
||||
```yaml
|
||||
energy:
|
||||
value: 83
|
||||
history:
|
||||
- event: defect-detected
|
||||
delta: 25
|
||||
finding: TD-143
|
||||
- event: navigation-adaptation
|
||||
delta: -2
|
||||
```
|
||||
|
||||
Energy is not synonymous with correctness.
|
||||
|
||||
It is better interpreted as:
|
||||
|
||||
> the current value of spending verification attention and resources on this asset.
|
||||
|
||||
## 9.8 Confidence
|
||||
|
||||
**Confidence** represents how strongly the framework currently believes that a verification asset faithfully represents the intended behavior it claims to protect.
|
||||
|
||||
Confidence should remain conceptually separate from energy.
|
||||
|
||||
A high-energy test may still have low confidence if its behavior is poorly specified.
|
||||
Energy and Confidence are removed from the current concept set. Neither was
|
||||
used in a decision across the three reference use cases. H-005 is
|
||||
dormant-indefinite; there is no capture-only energy subsystem.
|
||||
|
||||
## 9.9 Lineage
|
||||
|
||||
|
|
@ -735,18 +666,10 @@ Lineage should make it possible to answer:
|
|||
|
||||
> Why does this test exist?
|
||||
|
||||
## 9.10 Retirement
|
||||
## 9.10 Asset removal
|
||||
|
||||
A verification asset may move through:
|
||||
|
||||
```text
|
||||
active
|
||||
-> low-frequency
|
||||
-> archival
|
||||
-> retired
|
||||
```
|
||||
|
||||
Energy reaching zero may trigger retirement consideration, but critical security, regulatory or contractual invariants may define a retirement floor or prohibition.
|
||||
Maintainers may remove obsolete tests with an explicit review of the protected
|
||||
claims. No computed Retirement concept or policy is implemented or scheduled.
|
||||
|
||||
---
|
||||
|
||||
|
|
@ -770,51 +693,11 @@ When one side changes, the framework evaluates whether the other relationships r
|
|||
|
||||
---
|
||||
|
||||
# 11. Test Metabolism
|
||||
# 11–12. Removed planning abstractions
|
||||
|
||||
Energy, temperature, maturity and execution cost together create a **Test Metabolism**.
|
||||
|
||||
The framework should preferentially spend verification resources where they are most valuable.
|
||||
|
||||
A future campaign planner may consider:
|
||||
|
||||
```text
|
||||
Priority =
|
||||
Test Energy
|
||||
x Changed-System Proximity
|
||||
x Use-Case Criticality
|
||||
x Risk
|
||||
x Time Since Last Execution
|
||||
```
|
||||
|
||||
The precise formula is intentionally deferred.
|
||||
|
||||
The conceptual requirement is that test selection should be dynamic rather than assuming every historical test has equal present value.
|
||||
|
||||
---
|
||||
|
||||
# 12. Campaign
|
||||
|
||||
A **Campaign** is a strategy for selecting and executing verification assets and scenario variants.
|
||||
|
||||
Initial campaign examples:
|
||||
|
||||
```text
|
||||
PR smoke
|
||||
regression
|
||||
release qualification
|
||||
authorization
|
||||
tenant isolation
|
||||
concurrency
|
||||
resilience
|
||||
exploratory
|
||||
```
|
||||
|
||||
Campaigns manage combinatorial explosion by selecting relevant slices of:
|
||||
|
||||
```text
|
||||
actors x roles x states x data x surfaces x schedules x mutations
|
||||
```
|
||||
Metabolism and Campaign were removed from the current concept set on 2026-09-28.
|
||||
The experiments choose explicit scenario and mutation lists. No evidence justifies
|
||||
a planner or score-driven scheduling; these are not deferred implementation promises.
|
||||
|
||||
---
|
||||
|
||||
|
|
@ -898,9 +781,6 @@ Observers --> Evidence --> Oracle Engine --> Verdict
|
|||
|
||||
Metadata Store:
|
||||
maturity
|
||||
temperature
|
||||
energy
|
||||
confidence
|
||||
lineage
|
||||
```
|
||||
|
||||
|
|
@ -955,17 +835,11 @@ Finding
|
|||
Investigator
|
||||
VerificationAsset
|
||||
Maturity
|
||||
Temperature
|
||||
Adaptation
|
||||
Hardening
|
||||
Crystallization
|
||||
Energy
|
||||
Confidence
|
||||
Lineage
|
||||
Retirement
|
||||
Campaign
|
||||
TriadicVerification
|
||||
TestMetabolism
|
||||
```
|
||||
|
||||
---
|
||||
|
|
@ -1003,15 +877,9 @@ The defining model of test-driver is:
|
|||
v
|
||||
Verdict
|
||||
|
|
||||
Finding / Value
|
||||
|
|
||||
v
|
||||
Energy
|
||||
Finding
|
||||
|
||||
HOT ------> WARM ------> COOL ------> COLD
|
||||
| | | |
|
||||
agentic adaptive hardened deterministic
|
||||
+-------------- crystallization ------>
|
||||
Repeated stable trajectories --> crystallization --> deterministic regression
|
||||
```
|
||||
|
||||
The fundamental promise is:
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue