Complete local generalisation tasks and document remaining experiment blockers

Assistant: codex
Assistant-Model: gpt-6-astra
Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
This commit is contained in:
tegwick 2026-09-28 12:06:24 +02:00
parent d54af628df
commit cf682281d7
36 changed files with 5394 additions and 279 deletions

View file

@ -1,5 +1,8 @@
# Test Driver Concept Model
> Current scope, 2026-09-28: TD-WP-0003 removes speculative lifecycle scoring
> and declared Temperature. See [settlement and compatibility](TestDriverGeneralisationReview.md).
**Status:** v0.1 Draft
**Framework:** `test-driver`
@ -58,17 +61,10 @@ Agentic tests should not remain agentic merely because they started that way.
As behavior and implementation become stable, agentic realizations are progressively hardened and finally crystallized into deterministic tests.
### 2.5 Useful tests accumulate energy
### 2.5–2.6 Evidence, not lifecycle scores
Verification assets gain energy when they demonstrate value, for example by catching genuine failures or protecting important behavior.
They lose energy when they repeatedly become invalid, require unnecessary adaptation, become flaky, duplicate stronger verification or protect behavior that is no longer relevant.
### 2.6 Verification assets may die
Tests are not immortal repository artifacts.
When a verification asset loses relevance and reaches sufficiently low energy, it may be removed from active campaigns, archived or retired.
Retain observations and verdicts. Asset selection and retirement remain explicit
maintainer choices; there is no energy score or automated retirement policy.
### 2.7 Implementation is not the truth
@ -514,10 +510,6 @@ protects:
- least-privilege
maturity: adaptive
temperature: warm
energy:
value: 72
execution:
mode: agentic
@ -585,29 +577,10 @@ U4 Contractual
A stable use case may temporarily require agentic tests when a new implementation is introduced.
## 9.3 Temperature
## 9.3 Measured stability
**Temperature** represents implementation or capability fluidity.
Initial model:
```text
HOT actively being invented
WARM frequently changing
COOL stabilizing
COLD mature or contractual
```
Temperature influences the preferred execution strategy:
| Temperature | Preferred verification mode |
|---|---|
| HOT | exploratory and agentic |
| WARM | adaptive with emerging deterministic coverage |
| COOL | hardened regression with occasional exploration |
| COLD | predominantly deterministic |
A capability may cool as it stabilizes and heat up again during major redesign or migration.
Declared Temperature was removed on 2026-09-28 (F-0008). Crystallization uses
observed trajectory stability through `assess_stability`, not a temperature label.
## 9.4 Adaptation
@ -665,53 +638,11 @@ Explore
Crystallization is not necessarily permanent. A major implementation change may temporarily move an asset back toward adaptive or agentic execution.
## 9.7 Energy
## 9.7–9.8 Removed lifecycle scores
**Energy** expresses the current value of retaining, maintaining and executing a verification asset.
Energy should be derived from observable events rather than being an unexplained score.
Illustrative initial event model:
| Event | Energy change |
|---|---:|
| Detects confirmed product defect | +25 |
| Detects confirmed security defect | +40 |
| Prevents regression after previous defect | +20 |
| Exercises recently changed relevant capability | +2 |
| Requires harmless mechanical adaptation | -2 |
| Requires workflow adaptation | -10 |
| Test itself was wrong | -20 |
| False positive or flaky failure | -15 |
| Duplicates stronger existing verification | -20 |
| Protected use case is deprecated | -100 |
Energy history should be retained:
```yaml
energy:
value: 83
history:
- event: defect-detected
delta: 25
finding: TD-143
- event: navigation-adaptation
delta: -2
```
Energy is not synonymous with correctness.
It is better interpreted as:
> the current value of spending verification attention and resources on this asset.
## 9.8 Confidence
**Confidence** represents how strongly the framework currently believes that a verification asset faithfully represents the intended behavior it claims to protect.
Confidence should remain conceptually separate from energy.
A high-energy test may still have low confidence if its behavior is poorly specified.
Energy and Confidence are removed from the current concept set. Neither was
used in a decision across the three reference use cases. H-005 is
dormant-indefinite; there is no capture-only energy subsystem.
## 9.9 Lineage
@ -735,18 +666,10 @@ Lineage should make it possible to answer:
> Why does this test exist?
## 9.10 Retirement
## 9.10 Asset removal
A verification asset may move through:
```text
active
-> low-frequency
-> archival
-> retired
```
Energy reaching zero may trigger retirement consideration, but critical security, regulatory or contractual invariants may define a retirement floor or prohibition.
Maintainers may remove obsolete tests with an explicit review of the protected
claims. No computed Retirement concept or policy is implemented or scheduled.
---
@ -770,51 +693,11 @@ When one side changes, the framework evaluates whether the other relationships r
---
# 11. Test Metabolism
# 11–12. Removed planning abstractions
Energy, temperature, maturity and execution cost together create a **Test Metabolism**.
The framework should preferentially spend verification resources where they are most valuable.
A future campaign planner may consider:
```text
Priority =
Test Energy
x Changed-System Proximity
x Use-Case Criticality
x Risk
x Time Since Last Execution
```
The precise formula is intentionally deferred.
The conceptual requirement is that test selection should be dynamic rather than assuming every historical test has equal present value.
---
# 12. Campaign
A **Campaign** is a strategy for selecting and executing verification assets and scenario variants.
Initial campaign examples:
```text
PR smoke
regression
release qualification
authorization
tenant isolation
concurrency
resilience
exploratory
```
Campaigns manage combinatorial explosion by selecting relevant slices of:
```text
actors x roles x states x data x surfaces x schedules x mutations
```
Metabolism and Campaign were removed from the current concept set on 2026-09-28.
The experiments choose explicit scenario and mutation lists. No evidence justifies
a planner or score-driven scheduling; these are not deferred implementation promises.
---
@ -898,9 +781,6 @@ Observers --> Evidence --> Oracle Engine --> Verdict
Metadata Store:
maturity
temperature
energy
confidence
lineage
```
@ -955,17 +835,11 @@ Finding
Investigator
VerificationAsset
Maturity
Temperature
Adaptation
Hardening
Crystallization
Energy
Confidence
Lineage
Retirement
Campaign
TriadicVerification
TestMetabolism
```
---
@ -1003,15 +877,9 @@ The defining model of test-driver is:
v
Verdict
|
Finding / Value
|
v
Energy
Finding
HOT ------> WARM ------> COOL ------> COLD
| | | |
agentic adaptive hardened deterministic
+-------------- crystallization ------>
Repeated stable trajectories --> crystallization --> deterministic regression
```
The fundamental promise is:

View file

@ -0,0 +1,146 @@
# TD-WP-0003 generalisation and settlement — 2026-09-28
The kernel now expresses sharing/revocation, delegated approval and tenant
lifecycle without additional framework concepts. This is evidence of local
expressiveness, not readiness for autonomous verification of a real system.
## New use cases (T02)
The agent wrote [synthetic contracts](../usecases/generalisation-spec.md) before
the implementations. Claims are `agent-from-spec`, not human-authored. The same
agent wrote the synthetic requirements and labs; this does not establish
independence from a real product team or requirements quality.
- `scenarios/delegated_approval.py`: premature execution, submission, delegation,
rejected self-approval, rejected former-reviewer approval, delegate approval,
requester execution. Three independently enabled defects must fail.
- `scenarios/tenant_lifecycle.py`: two administrators use the same local id;
foreign delete attempts fail; deletion and recreation preserve the other
tenant. Four independently enabled defects must fail.
`tests/test_generalisation.py` verifies baselines, each defect, missing-evidence
INCONCLUSIVE, deterministic replay and verdict reproduction from serialized S3.
Separate observers read state and probe enforcement; drivers only emit S1.
No claims or invariants are learned from the lab's responses.
Ordered `Step`s already express the required Schedule. `Scenario.variant`
identifies each lab fault. A separate scheduler, Lens or Variant object was not
needed. There is no concurrent scheduling or time-window evidence here.
The resistance was in the adapters, not new concepts: `StateObserver`'s standard
snapshot is specific to resource grants, so each domain needs its own collector
implementing `snapshot()` and `name`. Expected refusal must be a completed
protocol response with an independently observed receipt; setting
`Realization.raised` would suppress dependent claims as INCONCLUSIVE. Unexpected
exceptions still abort these fixture drivers; they are not silently swallowed.
## Structural experiment (T03)
[Machine-readable results](../research/evidence/2026-09-28-e001.json) retain each
arm's surface, judgments, metrics and the journey classification/signals.
Reproduce with:
```bash
PYTHONPATH=src:. python3 tools/measure_e001.py /tmp/e001.json
```
| Identifier axis | Discovery | Recorded selectors |
|---|---:|---:|
| Preserved | 17/17 | 17/17 |
| Dropped | 9/10 | 0/10 |
M25–M31 add reordered fields, dialog relocation, anchor affordances, fieldset
grouping, textarea input, an unrelated preceding form, and sidebar relocation.
M32–M38 pair those mechanisms with preserved identifiers. Existing M02/M21/M22
supply rewritten DOM, reworded controls and renamed fields. New labels were fixed
from the synthetic contract before running and are agent-authored, not claimed
as newly human-reviewed ground truth.
The 38-entry catalogue contains 27 mechanical, four semantic and seven defect
mutations. Mechanical acceptance is 26/27. M22 still defeats the heuristic.
False adaptation is 0/7 on the existing defect set; adding mechanical variants
does not enlarge that safety denominator. No claim/invariant index changed.
M25 and M32 are UNCHANGED because the coarse surface fingerprint ignores field
order and test ids; this is not a claim that their HTML is identical.
These are selected mechanisms, not random independent samples. Paired variants
are correlated; 9/10 is not a population reliability estimate. The control is a
stable-test-id replay, not every possible conventional selector strategy. Both
arms remain token-free and test structural HTML only. H-001 stays narrowed to
identifier-loss cases, and F-0004/F-0005/F-0007's external validity limits remain.
## Concept settlement and compatibility (T04/T05)
**Keep claims as Python predicates.** Across all three runnable use cases the
repeated shapes are equality, membership and comparisons between enforcement
and state. The audit-core consumer additionally compares sanitized absence
surfaces, bounded execution, cleanup and receipt binding. A small vocabulary
would either duplicate Python or need frequent escape hatches; serialization
would create a second language and a second statement of intent.
This publishes the interface decision requested by T05: the existing
`Claim(id, text, provenance, predicate, after_step, source_ref=None)` and
`Invariant(id, text, provenance, predicate, source_ref=None)` constructors stay
compatible. Predicates receive independent observation mappings and return
booleans. Required evidence must use explicit lookup, not a pass-producing
fallback. The oracle converts missing keys or predicate exceptions to
INCONCLUSIVE. Predicates must not reach into actors or the live SUT.
Generated regression artifacts continue to import their original scenario
claims; consumers must ship those modules with the artifact. Fully standalone
claim serialization is deliberately rejected. The audit-core consumer needs no
migration; its tests remain in the default suite.
**Remove unused lifecycle concepts now.** Temperature, Energy, Confidence,
Campaign, Metabolism and automated Retirement have informed no decision in
these three use cases. They leave the current model, not another deferred gate.
Trajectory stability continues to drive crystallization. H-005 becomes
`DORMANT-INDEFINITE`; F-0008 is resolved.
This is a research-prototype compatibility change: `testdriver.energy`, public
`EnergyEvent`/`EnergyEventType` exports and the `EvidencePack.energy_events`
field are removed. Consumers must stop importing or requiring them. Historical
JSON may retain that field; the current classifier already ignores it.
Execution identity, timestamps, judgments, S1 realization costs and lineage
remain; none of these imply a computed lifecycle score. There is no dependency
addition. Historical concept proposals carry a supersession notice.
## Authoring measurement (T07 remains waiting)
[Timing receipt](../research/evidence/2026-09-28-authoring.json) preserves the
actual measurements and line counts. The timer was placed inside the shell
write operation, after code composition. The approval interval also included
composition of the next case, and the tenant interval largely measured file
writes and pytest. These intervals are **not valid authoring-cost measurements**
and must not be compared as productivity numbers. Human maintenance effort and
model latency/cost were not captured. The first attempt cannot be reconstructed.
The observable concepts consulted and friction are recorded above. T07 retains
the remaining work: an independent author must implement these contracts afresh
with a timer started before reading/design, stopping after first passing defect
checks, separately recording tooling wait and review. Do not count a rerun or
copy of these implementations as a new authoring sample. No replacement task
or workplan is created.
## Real-system readiness (T08)
**Verdict: not ready for autonomous real-system application.** Local research and
attended review of prepared claims can continue. Completing the review does not
satisfy the workplan's live-model success gate or authorize a production run.
| Required criterion | Current evidence | Assessment |
|---|---|---|
| Independent observation channel can read records and probe enforcement without trusting actors | Three labs; audit-core declares an independent observer but has no runnable driver/collector here | Not met outside synthetic labs; adapter construction, permissions and operating cost remain unmeasured |
| Claims trace to independently approved intent predating the tested behavior | Local provenance guards; audit-core cites approved engagement and workplan, while its precedent receipt is calibration only | Structurally supported, operational chain must be checked for the exact next engagement; a ticket written after code is insufficient without human promotion |
| Legitimate behavioral ambiguity is escalated without rewriting intent | M12/M19 are behaviorally indistinguishable; AMBIGUOUS and BEHAVIOUR_CHANGED require review | Demonstrated in lab, real backlog dispute resolution is still attended; missing evidence stays INCONCLUSIVE |
| False-adaptation consequences are identified for the engagement | Audit-core tenant disclosure, forged foreign append or leftover identity could be concealed by a false pass | Zero tolerance for automatic acceptance of regression; 0/7 synthetic defects is no guarantee against unseen attacks |
| Execution, custody, timing, cleanup and target revision are enforced | Audit-core's PhaseContract is durable intent, explicitly not a runnable Scenario; tests use supplied observations | Not met: no causal/time-window runtime, custody or cluster driver, independent cleanup collector or fresh admitted window here |
| Capability, economics and maintenance costs justify adoption | Token-free structural discovery, deterministic crystallization and invalid first authoring timer | Not met; T01/T06/T07 retain the unanswered measurements |
The audit-core material referenced here is the checked-in
`usecases/audit_core_e2_tenant_boundary.py` contract and its fixture-based tests,
not a fresh inspection of any deployment or a revalidation of the 2026-08-22
engagement. A real run needs the exact engagement's target-owner acceptance,
independent observation/cleanup receipts and appropriate approved access. A
successful historical receipt cannot substitute for those. The existing T01,
T06 and T07 tasks retain the experiment blockers; this review creates no
production implementation commitment or new work record.

View file

@ -1,5 +1,9 @@
# TestDriver Improvement Loop
> Scope update (2026-09-28): lifecycle Energy, Temperature, Confidence, Campaign,
> Metabolism and automated Retirement proposals below are historical, superseded
> by the [TD-WP-0003 settlement](TestDriverGeneralisationReview.md).
**Status:** Concept v0.1
**Purpose:** Establish a self-improvement loop that keeps the implementation of `test-driver` aligned with its conceptual model while generating evidence about which concepts actually work.

View file

@ -1,5 +1,9 @@
# TestDriver Research Prototype — Initial Milestones
> Scope update (2026-09-28): lifecycle Energy, Temperature, Confidence, Campaign,
> Metabolism and automated Retirement proposals below are historical, superseded
> by the [TD-WP-0003 settlement](TestDriverGeneralisationReview.md).
**Status:** v0.1 — **canonical milestone sequence**
**Purpose:** Establish the minimum evidence-producing development loop needed to validate the core test-driver concepts.

View file

@ -1,5 +1,9 @@
# Stage 1 Test Driver Validation
> Scope update (2026-09-28): lifecycle Energy, Temperature, Confidence, Campaign,
> Metabolism and automated Retirement proposals below are historical, superseded
> by the [TD-WP-0003 settlement](TestDriverGeneralisationReview.md).
The biggest risk is not technical feasibility. It is that **test-driver becomes conceptually elegant but too broad to prove itself quickly**.
I would improve the odds of success by treating the next phase as an experiment in whether three specific ideas actually work: