Complete local generalisation tasks and document remaining experiment blockers
Assistant: codex Assistant-Model: gpt-6-astra Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
This commit is contained in:
parent
d54af628df
commit
cf682281d7
36 changed files with 5394 additions and 279 deletions
|
|
@ -1,5 +1,8 @@
|
|||
# Test Driver Concept Model
|
||||
|
||||
> Current scope, 2026-09-28: TD-WP-0003 removes speculative lifecycle scoring
|
||||
> and declared Temperature. See [settlement and compatibility](TestDriverGeneralisationReview.md).
|
||||
|
||||
**Status:** v0.1 Draft
|
||||
**Framework:** `test-driver`
|
||||
|
||||
|
|
@ -58,17 +61,10 @@ Agentic tests should not remain agentic merely because they started that way.
|
|||
|
||||
As behavior and implementation become stable, agentic realizations are progressively hardened and finally crystallized into deterministic tests.
|
||||
|
||||
### 2.5 Useful tests accumulate energy
|
||||
### 2.5–2.6 Evidence, not lifecycle scores
|
||||
|
||||
Verification assets gain energy when they demonstrate value, for example by catching genuine failures or protecting important behavior.
|
||||
|
||||
They lose energy when they repeatedly become invalid, require unnecessary adaptation, become flaky, duplicate stronger verification or protect behavior that is no longer relevant.
|
||||
|
||||
### 2.6 Verification assets may die
|
||||
|
||||
Tests are not immortal repository artifacts.
|
||||
|
||||
When a verification asset loses relevance and reaches sufficiently low energy, it may be removed from active campaigns, archived or retired.
|
||||
Retain observations and verdicts. Asset selection and retirement remain explicit
|
||||
maintainer choices; there is no energy score or automated retirement policy.
|
||||
|
||||
### 2.7 Implementation is not the truth
|
||||
|
||||
|
|
@ -514,10 +510,6 @@ protects:
|
|||
- least-privilege
|
||||
|
||||
maturity: adaptive
|
||||
temperature: warm
|
||||
|
||||
energy:
|
||||
value: 72
|
||||
|
||||
execution:
|
||||
mode: agentic
|
||||
|
|
@ -585,29 +577,10 @@ U4 Contractual
|
|||
|
||||
A stable use case may temporarily require agentic tests when a new implementation is introduced.
|
||||
|
||||
## 9.3 Temperature
|
||||
## 9.3 Measured stability
|
||||
|
||||
**Temperature** represents implementation or capability fluidity.
|
||||
|
||||
Initial model:
|
||||
|
||||
```text
|
||||
HOT actively being invented
|
||||
WARM frequently changing
|
||||
COOL stabilizing
|
||||
COLD mature or contractual
|
||||
```
|
||||
|
||||
Temperature influences the preferred execution strategy:
|
||||
|
||||
| Temperature | Preferred verification mode |
|
||||
|---|---|
|
||||
| HOT | exploratory and agentic |
|
||||
| WARM | adaptive with emerging deterministic coverage |
|
||||
| COOL | hardened regression with occasional exploration |
|
||||
| COLD | predominantly deterministic |
|
||||
|
||||
A capability may cool as it stabilizes and heat up again during major redesign or migration.
|
||||
Declared Temperature was removed on 2026-09-28 (F-0008). Crystallization uses
|
||||
observed trajectory stability through `assess_stability`, not a temperature label.
|
||||
|
||||
## 9.4 Adaptation
|
||||
|
||||
|
|
@ -665,53 +638,11 @@ Explore
|
|||
|
||||
Crystallization is not necessarily permanent. A major implementation change may temporarily move an asset back toward adaptive or agentic execution.
|
||||
|
||||
## 9.7 Energy
|
||||
## 9.7–9.8 Removed lifecycle scores
|
||||
|
||||
**Energy** expresses the current value of retaining, maintaining and executing a verification asset.
|
||||
|
||||
Energy should be derived from observable events rather than being an unexplained score.
|
||||
|
||||
Illustrative initial event model:
|
||||
|
||||
| Event | Energy change |
|
||||
|---|---:|
|
||||
| Detects confirmed product defect | +25 |
|
||||
| Detects confirmed security defect | +40 |
|
||||
| Prevents regression after previous defect | +20 |
|
||||
| Exercises recently changed relevant capability | +2 |
|
||||
| Requires harmless mechanical adaptation | -2 |
|
||||
| Requires workflow adaptation | -10 |
|
||||
| Test itself was wrong | -20 |
|
||||
| False positive or flaky failure | -15 |
|
||||
| Duplicates stronger existing verification | -20 |
|
||||
| Protected use case is deprecated | -100 |
|
||||
|
||||
Energy history should be retained:
|
||||
|
||||
```yaml
|
||||
energy:
|
||||
value: 83
|
||||
history:
|
||||
- event: defect-detected
|
||||
delta: 25
|
||||
finding: TD-143
|
||||
- event: navigation-adaptation
|
||||
delta: -2
|
||||
```
|
||||
|
||||
Energy is not synonymous with correctness.
|
||||
|
||||
It is better interpreted as:
|
||||
|
||||
> the current value of spending verification attention and resources on this asset.
|
||||
|
||||
## 9.8 Confidence
|
||||
|
||||
**Confidence** represents how strongly the framework currently believes that a verification asset faithfully represents the intended behavior it claims to protect.
|
||||
|
||||
Confidence should remain conceptually separate from energy.
|
||||
|
||||
A high-energy test may still have low confidence if its behavior is poorly specified.
|
||||
Energy and Confidence are removed from the current concept set. Neither was
|
||||
used in a decision across the three reference use cases. H-005 is
|
||||
dormant-indefinite; there is no capture-only energy subsystem.
|
||||
|
||||
## 9.9 Lineage
|
||||
|
||||
|
|
@ -735,18 +666,10 @@ Lineage should make it possible to answer:
|
|||
|
||||
> Why does this test exist?
|
||||
|
||||
## 9.10 Retirement
|
||||
## 9.10 Asset removal
|
||||
|
||||
A verification asset may move through:
|
||||
|
||||
```text
|
||||
active
|
||||
-> low-frequency
|
||||
-> archival
|
||||
-> retired
|
||||
```
|
||||
|
||||
Energy reaching zero may trigger retirement consideration, but critical security, regulatory or contractual invariants may define a retirement floor or prohibition.
|
||||
Maintainers may remove obsolete tests with an explicit review of the protected
|
||||
claims. No computed Retirement concept or policy is implemented or scheduled.
|
||||
|
||||
---
|
||||
|
||||
|
|
@ -770,51 +693,11 @@ When one side changes, the framework evaluates whether the other relationships r
|
|||
|
||||
---
|
||||
|
||||
# 11. Test Metabolism
|
||||
# 11–12. Removed planning abstractions
|
||||
|
||||
Energy, temperature, maturity and execution cost together create a **Test Metabolism**.
|
||||
|
||||
The framework should preferentially spend verification resources where they are most valuable.
|
||||
|
||||
A future campaign planner may consider:
|
||||
|
||||
```text
|
||||
Priority =
|
||||
Test Energy
|
||||
x Changed-System Proximity
|
||||
x Use-Case Criticality
|
||||
x Risk
|
||||
x Time Since Last Execution
|
||||
```
|
||||
|
||||
The precise formula is intentionally deferred.
|
||||
|
||||
The conceptual requirement is that test selection should be dynamic rather than assuming every historical test has equal present value.
|
||||
|
||||
---
|
||||
|
||||
# 12. Campaign
|
||||
|
||||
A **Campaign** is a strategy for selecting and executing verification assets and scenario variants.
|
||||
|
||||
Initial campaign examples:
|
||||
|
||||
```text
|
||||
PR smoke
|
||||
regression
|
||||
release qualification
|
||||
authorization
|
||||
tenant isolation
|
||||
concurrency
|
||||
resilience
|
||||
exploratory
|
||||
```
|
||||
|
||||
Campaigns manage combinatorial explosion by selecting relevant slices of:
|
||||
|
||||
```text
|
||||
actors x roles x states x data x surfaces x schedules x mutations
|
||||
```
|
||||
Metabolism and Campaign were removed from the current concept set on 2026-09-28.
|
||||
The experiments choose explicit scenario and mutation lists. No evidence justifies
|
||||
a planner or score-driven scheduling; these are not deferred implementation promises.
|
||||
|
||||
---
|
||||
|
||||
|
|
@ -898,9 +781,6 @@ Observers --> Evidence --> Oracle Engine --> Verdict
|
|||
|
||||
Metadata Store:
|
||||
maturity
|
||||
temperature
|
||||
energy
|
||||
confidence
|
||||
lineage
|
||||
```
|
||||
|
||||
|
|
@ -955,17 +835,11 @@ Finding
|
|||
Investigator
|
||||
VerificationAsset
|
||||
Maturity
|
||||
Temperature
|
||||
Adaptation
|
||||
Hardening
|
||||
Crystallization
|
||||
Energy
|
||||
Confidence
|
||||
Lineage
|
||||
Retirement
|
||||
Campaign
|
||||
TriadicVerification
|
||||
TestMetabolism
|
||||
```
|
||||
|
||||
---
|
||||
|
|
@ -1003,15 +877,9 @@ The defining model of test-driver is:
|
|||
v
|
||||
Verdict
|
||||
|
|
||||
Finding / Value
|
||||
|
|
||||
v
|
||||
Energy
|
||||
Finding
|
||||
|
||||
HOT ------> WARM ------> COOL ------> COLD
|
||||
| | | |
|
||||
agentic adaptive hardened deterministic
|
||||
+-------------- crystallization ------>
|
||||
Repeated stable trajectories --> crystallization --> deterministic regression
|
||||
```
|
||||
|
||||
The fundamental promise is:
|
||||
|
|
|
|||
146
docs/TestDriverGeneralisationReview.md
Normal file
146
docs/TestDriverGeneralisationReview.md
Normal file
|
|
@ -0,0 +1,146 @@
|
|||
# TD-WP-0003 generalisation and settlement — 2026-09-28
|
||||
|
||||
The kernel now expresses sharing/revocation, delegated approval and tenant
|
||||
lifecycle without additional framework concepts. This is evidence of local
|
||||
expressiveness, not readiness for autonomous verification of a real system.
|
||||
|
||||
## New use cases (T02)
|
||||
|
||||
The agent wrote [synthetic contracts](../usecases/generalisation-spec.md) before
|
||||
the implementations. Claims are `agent-from-spec`, not human-authored. The same
|
||||
agent wrote the synthetic requirements and labs; this does not establish
|
||||
independence from a real product team or requirements quality.
|
||||
|
||||
- `scenarios/delegated_approval.py`: premature execution, submission, delegation,
|
||||
rejected self-approval, rejected former-reviewer approval, delegate approval,
|
||||
requester execution. Three independently enabled defects must fail.
|
||||
- `scenarios/tenant_lifecycle.py`: two administrators use the same local id;
|
||||
foreign delete attempts fail; deletion and recreation preserve the other
|
||||
tenant. Four independently enabled defects must fail.
|
||||
|
||||
`tests/test_generalisation.py` verifies baselines, each defect, missing-evidence
|
||||
INCONCLUSIVE, deterministic replay and verdict reproduction from serialized S3.
|
||||
Separate observers read state and probe enforcement; drivers only emit S1.
|
||||
No claims or invariants are learned from the lab's responses.
|
||||
|
||||
Ordered `Step`s already express the required Schedule. `Scenario.variant`
|
||||
identifies each lab fault. A separate scheduler, Lens or Variant object was not
|
||||
needed. There is no concurrent scheduling or time-window evidence here.
|
||||
|
||||
The resistance was in the adapters, not new concepts: `StateObserver`'s standard
|
||||
snapshot is specific to resource grants, so each domain needs its own collector
|
||||
implementing `snapshot()` and `name`. Expected refusal must be a completed
|
||||
protocol response with an independently observed receipt; setting
|
||||
`Realization.raised` would suppress dependent claims as INCONCLUSIVE. Unexpected
|
||||
exceptions still abort these fixture drivers; they are not silently swallowed.
|
||||
|
||||
## Structural experiment (T03)
|
||||
|
||||
[Machine-readable results](../research/evidence/2026-09-28-e001.json) retain each
|
||||
arm's surface, judgments, metrics and the journey classification/signals.
|
||||
Reproduce with:
|
||||
|
||||
```bash
|
||||
PYTHONPATH=src:. python3 tools/measure_e001.py /tmp/e001.json
|
||||
```
|
||||
|
||||
| Identifier axis | Discovery | Recorded selectors |
|
||||
|---|---:|---:|
|
||||
| Preserved | 17/17 | 17/17 |
|
||||
| Dropped | 9/10 | 0/10 |
|
||||
|
||||
M25–M31 add reordered fields, dialog relocation, anchor affordances, fieldset
|
||||
grouping, textarea input, an unrelated preceding form, and sidebar relocation.
|
||||
M32–M38 pair those mechanisms with preserved identifiers. Existing M02/M21/M22
|
||||
supply rewritten DOM, reworded controls and renamed fields. New labels were fixed
|
||||
from the synthetic contract before running and are agent-authored, not claimed
|
||||
as newly human-reviewed ground truth.
|
||||
|
||||
The 38-entry catalogue contains 27 mechanical, four semantic and seven defect
|
||||
mutations. Mechanical acceptance is 26/27. M22 still defeats the heuristic.
|
||||
False adaptation is 0/7 on the existing defect set; adding mechanical variants
|
||||
does not enlarge that safety denominator. No claim/invariant index changed.
|
||||
M25 and M32 are UNCHANGED because the coarse surface fingerprint ignores field
|
||||
order and test ids; this is not a claim that their HTML is identical.
|
||||
|
||||
These are selected mechanisms, not random independent samples. Paired variants
|
||||
are correlated; 9/10 is not a population reliability estimate. The control is a
|
||||
stable-test-id replay, not every possible conventional selector strategy. Both
|
||||
arms remain token-free and test structural HTML only. H-001 stays narrowed to
|
||||
identifier-loss cases, and F-0004/F-0005/F-0007's external validity limits remain.
|
||||
|
||||
## Concept settlement and compatibility (T04/T05)
|
||||
|
||||
**Keep claims as Python predicates.** Across all three runnable use cases the
|
||||
repeated shapes are equality, membership and comparisons between enforcement
|
||||
and state. The audit-core consumer additionally compares sanitized absence
|
||||
surfaces, bounded execution, cleanup and receipt binding. A small vocabulary
|
||||
would either duplicate Python or need frequent escape hatches; serialization
|
||||
would create a second language and a second statement of intent.
|
||||
|
||||
This publishes the interface decision requested by T05: the existing
|
||||
`Claim(id, text, provenance, predicate, after_step, source_ref=None)` and
|
||||
`Invariant(id, text, provenance, predicate, source_ref=None)` constructors stay
|
||||
compatible. Predicates receive independent observation mappings and return
|
||||
booleans. Required evidence must use explicit lookup, not a pass-producing
|
||||
fallback. The oracle converts missing keys or predicate exceptions to
|
||||
INCONCLUSIVE. Predicates must not reach into actors or the live SUT.
|
||||
Generated regression artifacts continue to import their original scenario
|
||||
claims; consumers must ship those modules with the artifact. Fully standalone
|
||||
claim serialization is deliberately rejected. The audit-core consumer needs no
|
||||
migration; its tests remain in the default suite.
|
||||
|
||||
**Remove unused lifecycle concepts now.** Temperature, Energy, Confidence,
|
||||
Campaign, Metabolism and automated Retirement have informed no decision in
|
||||
these three use cases. They leave the current model, not another deferred gate.
|
||||
Trajectory stability continues to drive crystallization. H-005 becomes
|
||||
`DORMANT-INDEFINITE`; F-0008 is resolved.
|
||||
|
||||
This is a research-prototype compatibility change: `testdriver.energy`, public
|
||||
`EnergyEvent`/`EnergyEventType` exports and the `EvidencePack.energy_events`
|
||||
field are removed. Consumers must stop importing or requiring them. Historical
|
||||
JSON may retain that field; the current classifier already ignores it.
|
||||
Execution identity, timestamps, judgments, S1 realization costs and lineage
|
||||
remain; none of these imply a computed lifecycle score. There is no dependency
|
||||
addition. Historical concept proposals carry a supersession notice.
|
||||
|
||||
## Authoring measurement (T07 remains waiting)
|
||||
|
||||
[Timing receipt](../research/evidence/2026-09-28-authoring.json) preserves the
|
||||
actual measurements and line counts. The timer was placed inside the shell
|
||||
write operation, after code composition. The approval interval also included
|
||||
composition of the next case, and the tenant interval largely measured file
|
||||
writes and pytest. These intervals are **not valid authoring-cost measurements**
|
||||
and must not be compared as productivity numbers. Human maintenance effort and
|
||||
model latency/cost were not captured. The first attempt cannot be reconstructed.
|
||||
|
||||
The observable concepts consulted and friction are recorded above. T07 retains
|
||||
the remaining work: an independent author must implement these contracts afresh
|
||||
with a timer started before reading/design, stopping after first passing defect
|
||||
checks, separately recording tooling wait and review. Do not count a rerun or
|
||||
copy of these implementations as a new authoring sample. No replacement task
|
||||
or workplan is created.
|
||||
|
||||
## Real-system readiness (T08)
|
||||
|
||||
**Verdict: not ready for autonomous real-system application.** Local research and
|
||||
attended review of prepared claims can continue. Completing the review does not
|
||||
satisfy the workplan's live-model success gate or authorize a production run.
|
||||
|
||||
| Required criterion | Current evidence | Assessment |
|
||||
|---|---|---|
|
||||
| Independent observation channel can read records and probe enforcement without trusting actors | Three labs; audit-core declares an independent observer but has no runnable driver/collector here | Not met outside synthetic labs; adapter construction, permissions and operating cost remain unmeasured |
|
||||
| Claims trace to independently approved intent predating the tested behavior | Local provenance guards; audit-core cites approved engagement and workplan, while its precedent receipt is calibration only | Structurally supported, operational chain must be checked for the exact next engagement; a ticket written after code is insufficient without human promotion |
|
||||
| Legitimate behavioral ambiguity is escalated without rewriting intent | M12/M19 are behaviorally indistinguishable; AMBIGUOUS and BEHAVIOUR_CHANGED require review | Demonstrated in lab, real backlog dispute resolution is still attended; missing evidence stays INCONCLUSIVE |
|
||||
| False-adaptation consequences are identified for the engagement | Audit-core tenant disclosure, forged foreign append or leftover identity could be concealed by a false pass | Zero tolerance for automatic acceptance of regression; 0/7 synthetic defects is no guarantee against unseen attacks |
|
||||
| Execution, custody, timing, cleanup and target revision are enforced | Audit-core's PhaseContract is durable intent, explicitly not a runnable Scenario; tests use supplied observations | Not met: no causal/time-window runtime, custody or cluster driver, independent cleanup collector or fresh admitted window here |
|
||||
| Capability, economics and maintenance costs justify adoption | Token-free structural discovery, deterministic crystallization and invalid first authoring timer | Not met; T01/T06/T07 retain the unanswered measurements |
|
||||
|
||||
The audit-core material referenced here is the checked-in
|
||||
`usecases/audit_core_e2_tenant_boundary.py` contract and its fixture-based tests,
|
||||
not a fresh inspection of any deployment or a revalidation of the 2026-08-22
|
||||
engagement. A real run needs the exact engagement's target-owner acceptance,
|
||||
independent observation/cleanup receipts and appropriate approved access. A
|
||||
successful historical receipt cannot substitute for those. The existing T01,
|
||||
T06 and T07 tasks retain the experiment blockers; this review creates no
|
||||
production implementation commitment or new work record.
|
||||
|
|
@ -1,5 +1,9 @@
|
|||
# TestDriver Improvement Loop
|
||||
|
||||
> Scope update (2026-09-28): lifecycle Energy, Temperature, Confidence, Campaign,
|
||||
> Metabolism and automated Retirement proposals below are historical, superseded
|
||||
> by the [TD-WP-0003 settlement](TestDriverGeneralisationReview.md).
|
||||
|
||||
**Status:** Concept v0.1
|
||||
**Purpose:** Establish a self-improvement loop that keeps the implementation of `test-driver` aligned with its conceptual model while generating evidence about which concepts actually work.
|
||||
|
||||
|
|
|
|||
|
|
@ -1,5 +1,9 @@
|
|||
# TestDriver Research Prototype — Initial Milestones
|
||||
|
||||
> Scope update (2026-09-28): lifecycle Energy, Temperature, Confidence, Campaign,
|
||||
> Metabolism and automated Retirement proposals below are historical, superseded
|
||||
> by the [TD-WP-0003 settlement](TestDriverGeneralisationReview.md).
|
||||
|
||||
**Status:** v0.1 — **canonical milestone sequence**
|
||||
**Purpose:** Establish the minimum evidence-producing development loop needed to validate the core test-driver concepts.
|
||||
|
||||
|
|
|
|||
|
|
@ -1,5 +1,9 @@
|
|||
# Stage 1 Test Driver Validation
|
||||
|
||||
> Scope update (2026-09-28): lifecycle Energy, Temperature, Confidence, Campaign,
|
||||
> Metabolism and automated Retirement proposals below are historical, superseded
|
||||
> by the [TD-WP-0003 settlement](TestDriverGeneralisationReview.md).
|
||||
|
||||
The biggest risk is not technical feasibility. It is that **test-driver becomes conceptually elegant but too broad to prove itself quickly**.
|
||||
|
||||
I would improve the odds of success by treating the next phase as an experiment in whether three specific ideas actually work:
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue