Complete local generalisation tasks and document remaining experiment blockers
Assistant: codex Assistant-Model: gpt-6-astra Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
This commit is contained in:
parent
d54af628df
commit
cf682281d7
36 changed files with 5394 additions and 279 deletions
|
|
@ -1,5 +1,9 @@
|
||||||
# test-driver
|
# test-driver
|
||||||
|
|
||||||
|
> Scope update (2026-09-28): lifecycle Energy, Temperature, Confidence, Campaign,
|
||||||
|
> Metabolism and automated Retirement proposals below are historical, superseded
|
||||||
|
> by the [TD-WP-0003 settlement](docs/TestDriverGeneralisationReview.md).
|
||||||
|
|
||||||
## Intent
|
## Intent
|
||||||
|
|
||||||
`test-driver` is a use-case-driven verification framework for integration, end-to-end, multi-user interaction, authorization, security and resilience testing in software systems that evolve through fast and increasingly agentic development cycles.
|
`test-driver` is a use-case-driven verification framework for integration, end-to-end, multi-user interaction, authorization, security and resilience testing in software systems that evolve through fast and increasingly agentic development cycles.
|
||||||
|
|
|
||||||
|
|
@ -7,9 +7,12 @@ Tests mature alongside the software they protect: fluid and agentic while
|
||||||
behaviour is changing, deterministic once it settles. See `INTENT.md` for the
|
behaviour is changing, deterministic once it settles. See `INTENT.md` for the
|
||||||
thesis and `SCOPE.md` for boundaries.
|
thesis and `SCOPE.md` for boundaries.
|
||||||
|
|
||||||
**Status:** research prototype. The deterministic kernel runs; agentic
|
**Status:** research prototype. Deterministic claims, heuristic discovery,
|
||||||
realization, adaptation classification and crystallization are not built yet.
|
adaptation classification and crystallization run locally. Three synthetic
|
||||||
Current work: `workplans/TD-WP-0002-vertical-spike-crystallization.md`.
|
multi-user use cases and a 38-mutation catalogue exercise the model. Live-model
|
||||||
|
economics, browser-engine coverage and independent authoring-cost measurement
|
||||||
|
remain blocked in `workplans/TD-WP-0003-generalise-and-settle.md`.
|
||||||
|
See [readiness and interface decisions](docs/TestDriverGeneralisationReview.md).
|
||||||
|
|
||||||
## Run
|
## Run
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -10,7 +10,7 @@
|
||||||
| --- | --- | --- | --- | --- |
|
| --- | --- | --- | --- | --- |
|
||||||
| workplan | TD-WP-0001 | finished | — | workplans/TD-WP-0001-statehub-bootstrap.md |
|
| workplan | TD-WP-0001 | finished | — | workplans/TD-WP-0001-statehub-bootstrap.md |
|
||||||
| workplan | TD-WP-0002 | finished | — | workplans/TD-WP-0002-vertical-spike-crystallization.md |
|
| workplan | TD-WP-0002 | finished | — | workplans/TD-WP-0002-vertical-spike-crystallization.md |
|
||||||
| workplan | TD-WP-0003 | proposed | — | workplans/TD-WP-0003-generalise-and-settle.md |
|
| workplan | TD-WP-0003 | blocked | — | workplans/TD-WP-0003-generalise-and-settle.md |
|
||||||
| task | TD-WP-0001-T01 | done | — | workplans/TD-WP-0001-statehub-bootstrap.md |
|
| task | TD-WP-0001-T01 | done | — | workplans/TD-WP-0001-statehub-bootstrap.md |
|
||||||
| task | TD-WP-0001-T02 | done | — | workplans/TD-WP-0001-statehub-bootstrap.md |
|
| task | TD-WP-0001-T02 | done | — | workplans/TD-WP-0001-statehub-bootstrap.md |
|
||||||
| task | TD-WP-0001-T03 | done | — | workplans/TD-WP-0001-statehub-bootstrap.md |
|
| task | TD-WP-0001-T03 | done | — | workplans/TD-WP-0001-statehub-bootstrap.md |
|
||||||
|
|
@ -24,11 +24,11 @@
|
||||||
| task | TD-WP-0002-T08 | done | — | workplans/TD-WP-0002-vertical-spike-crystallization.md |
|
| task | TD-WP-0002-T08 | done | — | workplans/TD-WP-0002-vertical-spike-crystallization.md |
|
||||||
| task | TD-WP-0002-T09 | done | — | workplans/TD-WP-0002-vertical-spike-crystallization.md |
|
| task | TD-WP-0002-T09 | done | — | workplans/TD-WP-0002-vertical-spike-crystallization.md |
|
||||||
| task | TD-WP-0002-T10 | done | — | workplans/TD-WP-0002-vertical-spike-crystallization.md |
|
| task | TD-WP-0002-T10 | done | — | workplans/TD-WP-0002-vertical-spike-crystallization.md |
|
||||||
| task | TD-WP-0003-T01 | todo | — | workplans/TD-WP-0003-generalise-and-settle.md |
|
| task | TD-WP-0003-T01 | wait | — | workplans/TD-WP-0003-generalise-and-settle.md |
|
||||||
| task | TD-WP-0003-T02 | todo | — | workplans/TD-WP-0003-generalise-and-settle.md |
|
| task | TD-WP-0003-T02 | done | — | workplans/TD-WP-0003-generalise-and-settle.md |
|
||||||
| task | TD-WP-0003-T03 | todo | — | workplans/TD-WP-0003-generalise-and-settle.md |
|
| task | TD-WP-0003-T03 | done | — | workplans/TD-WP-0003-generalise-and-settle.md |
|
||||||
| task | TD-WP-0003-T04 | todo | — | workplans/TD-WP-0003-generalise-and-settle.md |
|
| task | TD-WP-0003-T04 | done | — | workplans/TD-WP-0003-generalise-and-settle.md |
|
||||||
| task | TD-WP-0003-T05 | todo | — | workplans/TD-WP-0003-generalise-and-settle.md |
|
| task | TD-WP-0003-T05 | done | — | workplans/TD-WP-0003-generalise-and-settle.md |
|
||||||
| task | TD-WP-0003-T06 | wait | — | workplans/TD-WP-0003-generalise-and-settle.md |
|
| task | TD-WP-0003-T06 | wait | — | workplans/TD-WP-0003-generalise-and-settle.md |
|
||||||
| task | TD-WP-0003-T07 | todo | — | workplans/TD-WP-0003-generalise-and-settle.md |
|
| task | TD-WP-0003-T07 | wait | — | workplans/TD-WP-0003-generalise-and-settle.md |
|
||||||
| task | TD-WP-0003-T08 | todo | — | workplans/TD-WP-0003-generalise-and-settle.md |
|
| task | TD-WP-0003-T08 | done | — | workplans/TD-WP-0003-generalise-and-settle.md |
|
||||||
|
|
|
||||||
|
|
@ -1,5 +1,8 @@
|
||||||
# Test Driver Concept Model
|
# Test Driver Concept Model
|
||||||
|
|
||||||
|
> Current scope, 2026-09-28: TD-WP-0003 removes speculative lifecycle scoring
|
||||||
|
> and declared Temperature. See [settlement and compatibility](TestDriverGeneralisationReview.md).
|
||||||
|
|
||||||
**Status:** v0.1 Draft
|
**Status:** v0.1 Draft
|
||||||
**Framework:** `test-driver`
|
**Framework:** `test-driver`
|
||||||
|
|
||||||
|
|
@ -58,17 +61,10 @@ Agentic tests should not remain agentic merely because they started that way.
|
||||||
|
|
||||||
As behavior and implementation become stable, agentic realizations are progressively hardened and finally crystallized into deterministic tests.
|
As behavior and implementation become stable, agentic realizations are progressively hardened and finally crystallized into deterministic tests.
|
||||||
|
|
||||||
### 2.5 Useful tests accumulate energy
|
### 2.5–2.6 Evidence, not lifecycle scores
|
||||||
|
|
||||||
Verification assets gain energy when they demonstrate value, for example by catching genuine failures or protecting important behavior.
|
Retain observations and verdicts. Asset selection and retirement remain explicit
|
||||||
|
maintainer choices; there is no energy score or automated retirement policy.
|
||||||
They lose energy when they repeatedly become invalid, require unnecessary adaptation, become flaky, duplicate stronger verification or protect behavior that is no longer relevant.
|
|
||||||
|
|
||||||
### 2.6 Verification assets may die
|
|
||||||
|
|
||||||
Tests are not immortal repository artifacts.
|
|
||||||
|
|
||||||
When a verification asset loses relevance and reaches sufficiently low energy, it may be removed from active campaigns, archived or retired.
|
|
||||||
|
|
||||||
### 2.7 Implementation is not the truth
|
### 2.7 Implementation is not the truth
|
||||||
|
|
||||||
|
|
@ -514,10 +510,6 @@ protects:
|
||||||
- least-privilege
|
- least-privilege
|
||||||
|
|
||||||
maturity: adaptive
|
maturity: adaptive
|
||||||
temperature: warm
|
|
||||||
|
|
||||||
energy:
|
|
||||||
value: 72
|
|
||||||
|
|
||||||
execution:
|
execution:
|
||||||
mode: agentic
|
mode: agentic
|
||||||
|
|
@ -585,29 +577,10 @@ U4 Contractual
|
||||||
|
|
||||||
A stable use case may temporarily require agentic tests when a new implementation is introduced.
|
A stable use case may temporarily require agentic tests when a new implementation is introduced.
|
||||||
|
|
||||||
## 9.3 Temperature
|
## 9.3 Measured stability
|
||||||
|
|
||||||
**Temperature** represents implementation or capability fluidity.
|
Declared Temperature was removed on 2026-09-28 (F-0008). Crystallization uses
|
||||||
|
observed trajectory stability through `assess_stability`, not a temperature label.
|
||||||
Initial model:
|
|
||||||
|
|
||||||
```text
|
|
||||||
HOT actively being invented
|
|
||||||
WARM frequently changing
|
|
||||||
COOL stabilizing
|
|
||||||
COLD mature or contractual
|
|
||||||
```
|
|
||||||
|
|
||||||
Temperature influences the preferred execution strategy:
|
|
||||||
|
|
||||||
| Temperature | Preferred verification mode |
|
|
||||||
|---|---|
|
|
||||||
| HOT | exploratory and agentic |
|
|
||||||
| WARM | adaptive with emerging deterministic coverage |
|
|
||||||
| COOL | hardened regression with occasional exploration |
|
|
||||||
| COLD | predominantly deterministic |
|
|
||||||
|
|
||||||
A capability may cool as it stabilizes and heat up again during major redesign or migration.
|
|
||||||
|
|
||||||
## 9.4 Adaptation
|
## 9.4 Adaptation
|
||||||
|
|
||||||
|
|
@ -665,53 +638,11 @@ Explore
|
||||||
|
|
||||||
Crystallization is not necessarily permanent. A major implementation change may temporarily move an asset back toward adaptive or agentic execution.
|
Crystallization is not necessarily permanent. A major implementation change may temporarily move an asset back toward adaptive or agentic execution.
|
||||||
|
|
||||||
## 9.7 Energy
|
## 9.7–9.8 Removed lifecycle scores
|
||||||
|
|
||||||
**Energy** expresses the current value of retaining, maintaining and executing a verification asset.
|
Energy and Confidence are removed from the current concept set. Neither was
|
||||||
|
used in a decision across the three reference use cases. H-005 is
|
||||||
Energy should be derived from observable events rather than being an unexplained score.
|
dormant-indefinite; there is no capture-only energy subsystem.
|
||||||
|
|
||||||
Illustrative initial event model:
|
|
||||||
|
|
||||||
| Event | Energy change |
|
|
||||||
|---|---:|
|
|
||||||
| Detects confirmed product defect | +25 |
|
|
||||||
| Detects confirmed security defect | +40 |
|
|
||||||
| Prevents regression after previous defect | +20 |
|
|
||||||
| Exercises recently changed relevant capability | +2 |
|
|
||||||
| Requires harmless mechanical adaptation | -2 |
|
|
||||||
| Requires workflow adaptation | -10 |
|
|
||||||
| Test itself was wrong | -20 |
|
|
||||||
| False positive or flaky failure | -15 |
|
|
||||||
| Duplicates stronger existing verification | -20 |
|
|
||||||
| Protected use case is deprecated | -100 |
|
|
||||||
|
|
||||||
Energy history should be retained:
|
|
||||||
|
|
||||||
```yaml
|
|
||||||
energy:
|
|
||||||
value: 83
|
|
||||||
history:
|
|
||||||
- event: defect-detected
|
|
||||||
delta: 25
|
|
||||||
finding: TD-143
|
|
||||||
- event: navigation-adaptation
|
|
||||||
delta: -2
|
|
||||||
```
|
|
||||||
|
|
||||||
Energy is not synonymous with correctness.
|
|
||||||
|
|
||||||
It is better interpreted as:
|
|
||||||
|
|
||||||
> the current value of spending verification attention and resources on this asset.
|
|
||||||
|
|
||||||
## 9.8 Confidence
|
|
||||||
|
|
||||||
**Confidence** represents how strongly the framework currently believes that a verification asset faithfully represents the intended behavior it claims to protect.
|
|
||||||
|
|
||||||
Confidence should remain conceptually separate from energy.
|
|
||||||
|
|
||||||
A high-energy test may still have low confidence if its behavior is poorly specified.
|
|
||||||
|
|
||||||
## 9.9 Lineage
|
## 9.9 Lineage
|
||||||
|
|
||||||
|
|
@ -735,18 +666,10 @@ Lineage should make it possible to answer:
|
||||||
|
|
||||||
> Why does this test exist?
|
> Why does this test exist?
|
||||||
|
|
||||||
## 9.10 Retirement
|
## 9.10 Asset removal
|
||||||
|
|
||||||
A verification asset may move through:
|
Maintainers may remove obsolete tests with an explicit review of the protected
|
||||||
|
claims. No computed Retirement concept or policy is implemented or scheduled.
|
||||||
```text
|
|
||||||
active
|
|
||||||
-> low-frequency
|
|
||||||
-> archival
|
|
||||||
-> retired
|
|
||||||
```
|
|
||||||
|
|
||||||
Energy reaching zero may trigger retirement consideration, but critical security, regulatory or contractual invariants may define a retirement floor or prohibition.
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
|
@ -770,51 +693,11 @@ When one side changes, the framework evaluates whether the other relationships r
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
# 11. Test Metabolism
|
# 11–12. Removed planning abstractions
|
||||||
|
|
||||||
Energy, temperature, maturity and execution cost together create a **Test Metabolism**.
|
Metabolism and Campaign were removed from the current concept set on 2026-09-28.
|
||||||
|
The experiments choose explicit scenario and mutation lists. No evidence justifies
|
||||||
The framework should preferentially spend verification resources where they are most valuable.
|
a planner or score-driven scheduling; these are not deferred implementation promises.
|
||||||
|
|
||||||
A future campaign planner may consider:
|
|
||||||
|
|
||||||
```text
|
|
||||||
Priority =
|
|
||||||
Test Energy
|
|
||||||
x Changed-System Proximity
|
|
||||||
x Use-Case Criticality
|
|
||||||
x Risk
|
|
||||||
x Time Since Last Execution
|
|
||||||
```
|
|
||||||
|
|
||||||
The precise formula is intentionally deferred.
|
|
||||||
|
|
||||||
The conceptual requirement is that test selection should be dynamic rather than assuming every historical test has equal present value.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
# 12. Campaign
|
|
||||||
|
|
||||||
A **Campaign** is a strategy for selecting and executing verification assets and scenario variants.
|
|
||||||
|
|
||||||
Initial campaign examples:
|
|
||||||
|
|
||||||
```text
|
|
||||||
PR smoke
|
|
||||||
regression
|
|
||||||
release qualification
|
|
||||||
authorization
|
|
||||||
tenant isolation
|
|
||||||
concurrency
|
|
||||||
resilience
|
|
||||||
exploratory
|
|
||||||
```
|
|
||||||
|
|
||||||
Campaigns manage combinatorial explosion by selecting relevant slices of:
|
|
||||||
|
|
||||||
```text
|
|
||||||
actors x roles x states x data x surfaces x schedules x mutations
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
|
@ -898,9 +781,6 @@ Observers --> Evidence --> Oracle Engine --> Verdict
|
||||||
|
|
||||||
Metadata Store:
|
Metadata Store:
|
||||||
maturity
|
maturity
|
||||||
temperature
|
|
||||||
energy
|
|
||||||
confidence
|
|
||||||
lineage
|
lineage
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|
@ -955,17 +835,11 @@ Finding
|
||||||
Investigator
|
Investigator
|
||||||
VerificationAsset
|
VerificationAsset
|
||||||
Maturity
|
Maturity
|
||||||
Temperature
|
|
||||||
Adaptation
|
Adaptation
|
||||||
Hardening
|
Hardening
|
||||||
Crystallization
|
Crystallization
|
||||||
Energy
|
|
||||||
Confidence
|
|
||||||
Lineage
|
Lineage
|
||||||
Retirement
|
|
||||||
Campaign
|
|
||||||
TriadicVerification
|
TriadicVerification
|
||||||
TestMetabolism
|
|
||||||
```
|
```
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
@ -1003,15 +877,9 @@ The defining model of test-driver is:
|
||||||
v
|
v
|
||||||
Verdict
|
Verdict
|
||||||
|
|
|
|
||||||
Finding / Value
|
Finding
|
||||||
|
|
|
||||||
v
|
|
||||||
Energy
|
|
||||||
|
|
||||||
HOT ------> WARM ------> COOL ------> COLD
|
Repeated stable trajectories --> crystallization --> deterministic regression
|
||||||
| | | |
|
|
||||||
agentic adaptive hardened deterministic
|
|
||||||
+-------------- crystallization ------>
|
|
||||||
```
|
```
|
||||||
|
|
||||||
The fundamental promise is:
|
The fundamental promise is:
|
||||||
|
|
|
||||||
146
docs/TestDriverGeneralisationReview.md
Normal file
146
docs/TestDriverGeneralisationReview.md
Normal file
|
|
@ -0,0 +1,146 @@
|
||||||
|
# TD-WP-0003 generalisation and settlement — 2026-09-28
|
||||||
|
|
||||||
|
The kernel now expresses sharing/revocation, delegated approval and tenant
|
||||||
|
lifecycle without additional framework concepts. This is evidence of local
|
||||||
|
expressiveness, not readiness for autonomous verification of a real system.
|
||||||
|
|
||||||
|
## New use cases (T02)
|
||||||
|
|
||||||
|
The agent wrote [synthetic contracts](../usecases/generalisation-spec.md) before
|
||||||
|
the implementations. Claims are `agent-from-spec`, not human-authored. The same
|
||||||
|
agent wrote the synthetic requirements and labs; this does not establish
|
||||||
|
independence from a real product team or requirements quality.
|
||||||
|
|
||||||
|
- `scenarios/delegated_approval.py`: premature execution, submission, delegation,
|
||||||
|
rejected self-approval, rejected former-reviewer approval, delegate approval,
|
||||||
|
requester execution. Three independently enabled defects must fail.
|
||||||
|
- `scenarios/tenant_lifecycle.py`: two administrators use the same local id;
|
||||||
|
foreign delete attempts fail; deletion and recreation preserve the other
|
||||||
|
tenant. Four independently enabled defects must fail.
|
||||||
|
|
||||||
|
`tests/test_generalisation.py` verifies baselines, each defect, missing-evidence
|
||||||
|
INCONCLUSIVE, deterministic replay and verdict reproduction from serialized S3.
|
||||||
|
Separate observers read state and probe enforcement; drivers only emit S1.
|
||||||
|
No claims or invariants are learned from the lab's responses.
|
||||||
|
|
||||||
|
Ordered `Step`s already express the required Schedule. `Scenario.variant`
|
||||||
|
identifies each lab fault. A separate scheduler, Lens or Variant object was not
|
||||||
|
needed. There is no concurrent scheduling or time-window evidence here.
|
||||||
|
|
||||||
|
The resistance was in the adapters, not new concepts: `StateObserver`'s standard
|
||||||
|
snapshot is specific to resource grants, so each domain needs its own collector
|
||||||
|
implementing `snapshot()` and `name`. Expected refusal must be a completed
|
||||||
|
protocol response with an independently observed receipt; setting
|
||||||
|
`Realization.raised` would suppress dependent claims as INCONCLUSIVE. Unexpected
|
||||||
|
exceptions still abort these fixture drivers; they are not silently swallowed.
|
||||||
|
|
||||||
|
## Structural experiment (T03)
|
||||||
|
|
||||||
|
[Machine-readable results](../research/evidence/2026-09-28-e001.json) retain each
|
||||||
|
arm's surface, judgments, metrics and the journey classification/signals.
|
||||||
|
Reproduce with:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
PYTHONPATH=src:. python3 tools/measure_e001.py /tmp/e001.json
|
||||||
|
```
|
||||||
|
|
||||||
|
| Identifier axis | Discovery | Recorded selectors |
|
||||||
|
|---|---:|---:|
|
||||||
|
| Preserved | 17/17 | 17/17 |
|
||||||
|
| Dropped | 9/10 | 0/10 |
|
||||||
|
|
||||||
|
M25–M31 add reordered fields, dialog relocation, anchor affordances, fieldset
|
||||||
|
grouping, textarea input, an unrelated preceding form, and sidebar relocation.
|
||||||
|
M32–M38 pair those mechanisms with preserved identifiers. Existing M02/M21/M22
|
||||||
|
supply rewritten DOM, reworded controls and renamed fields. New labels were fixed
|
||||||
|
from the synthetic contract before running and are agent-authored, not claimed
|
||||||
|
as newly human-reviewed ground truth.
|
||||||
|
|
||||||
|
The 38-entry catalogue contains 27 mechanical, four semantic and seven defect
|
||||||
|
mutations. Mechanical acceptance is 26/27. M22 still defeats the heuristic.
|
||||||
|
False adaptation is 0/7 on the existing defect set; adding mechanical variants
|
||||||
|
does not enlarge that safety denominator. No claim/invariant index changed.
|
||||||
|
M25 and M32 are UNCHANGED because the coarse surface fingerprint ignores field
|
||||||
|
order and test ids; this is not a claim that their HTML is identical.
|
||||||
|
|
||||||
|
These are selected mechanisms, not random independent samples. Paired variants
|
||||||
|
are correlated; 9/10 is not a population reliability estimate. The control is a
|
||||||
|
stable-test-id replay, not every possible conventional selector strategy. Both
|
||||||
|
arms remain token-free and test structural HTML only. H-001 stays narrowed to
|
||||||
|
identifier-loss cases, and F-0004/F-0005/F-0007's external validity limits remain.
|
||||||
|
|
||||||
|
## Concept settlement and compatibility (T04/T05)
|
||||||
|
|
||||||
|
**Keep claims as Python predicates.** Across all three runnable use cases the
|
||||||
|
repeated shapes are equality, membership and comparisons between enforcement
|
||||||
|
and state. The audit-core consumer additionally compares sanitized absence
|
||||||
|
surfaces, bounded execution, cleanup and receipt binding. A small vocabulary
|
||||||
|
would either duplicate Python or need frequent escape hatches; serialization
|
||||||
|
would create a second language and a second statement of intent.
|
||||||
|
|
||||||
|
This publishes the interface decision requested by T05: the existing
|
||||||
|
`Claim(id, text, provenance, predicate, after_step, source_ref=None)` and
|
||||||
|
`Invariant(id, text, provenance, predicate, source_ref=None)` constructors stay
|
||||||
|
compatible. Predicates receive independent observation mappings and return
|
||||||
|
booleans. Required evidence must use explicit lookup, not a pass-producing
|
||||||
|
fallback. The oracle converts missing keys or predicate exceptions to
|
||||||
|
INCONCLUSIVE. Predicates must not reach into actors or the live SUT.
|
||||||
|
Generated regression artifacts continue to import their original scenario
|
||||||
|
claims; consumers must ship those modules with the artifact. Fully standalone
|
||||||
|
claim serialization is deliberately rejected. The audit-core consumer needs no
|
||||||
|
migration; its tests remain in the default suite.
|
||||||
|
|
||||||
|
**Remove unused lifecycle concepts now.** Temperature, Energy, Confidence,
|
||||||
|
Campaign, Metabolism and automated Retirement have informed no decision in
|
||||||
|
these three use cases. They leave the current model, not another deferred gate.
|
||||||
|
Trajectory stability continues to drive crystallization. H-005 becomes
|
||||||
|
`DORMANT-INDEFINITE`; F-0008 is resolved.
|
||||||
|
|
||||||
|
This is a research-prototype compatibility change: `testdriver.energy`, public
|
||||||
|
`EnergyEvent`/`EnergyEventType` exports and the `EvidencePack.energy_events`
|
||||||
|
field are removed. Consumers must stop importing or requiring them. Historical
|
||||||
|
JSON may retain that field; the current classifier already ignores it.
|
||||||
|
Execution identity, timestamps, judgments, S1 realization costs and lineage
|
||||||
|
remain; none of these imply a computed lifecycle score. There is no dependency
|
||||||
|
addition. Historical concept proposals carry a supersession notice.
|
||||||
|
|
||||||
|
## Authoring measurement (T07 remains waiting)
|
||||||
|
|
||||||
|
[Timing receipt](../research/evidence/2026-09-28-authoring.json) preserves the
|
||||||
|
actual measurements and line counts. The timer was placed inside the shell
|
||||||
|
write operation, after code composition. The approval interval also included
|
||||||
|
composition of the next case, and the tenant interval largely measured file
|
||||||
|
writes and pytest. These intervals are **not valid authoring-cost measurements**
|
||||||
|
and must not be compared as productivity numbers. Human maintenance effort and
|
||||||
|
model latency/cost were not captured. The first attempt cannot be reconstructed.
|
||||||
|
|
||||||
|
The observable concepts consulted and friction are recorded above. T07 retains
|
||||||
|
the remaining work: an independent author must implement these contracts afresh
|
||||||
|
with a timer started before reading/design, stopping after first passing defect
|
||||||
|
checks, separately recording tooling wait and review. Do not count a rerun or
|
||||||
|
copy of these implementations as a new authoring sample. No replacement task
|
||||||
|
or workplan is created.
|
||||||
|
|
||||||
|
## Real-system readiness (T08)
|
||||||
|
|
||||||
|
**Verdict: not ready for autonomous real-system application.** Local research and
|
||||||
|
attended review of prepared claims can continue. Completing the review does not
|
||||||
|
satisfy the workplan's live-model success gate or authorize a production run.
|
||||||
|
|
||||||
|
| Required criterion | Current evidence | Assessment |
|
||||||
|
|---|---|---|
|
||||||
|
| Independent observation channel can read records and probe enforcement without trusting actors | Three labs; audit-core declares an independent observer but has no runnable driver/collector here | Not met outside synthetic labs; adapter construction, permissions and operating cost remain unmeasured |
|
||||||
|
| Claims trace to independently approved intent predating the tested behavior | Local provenance guards; audit-core cites approved engagement and workplan, while its precedent receipt is calibration only | Structurally supported, operational chain must be checked for the exact next engagement; a ticket written after code is insufficient without human promotion |
|
||||||
|
| Legitimate behavioral ambiguity is escalated without rewriting intent | M12/M19 are behaviorally indistinguishable; AMBIGUOUS and BEHAVIOUR_CHANGED require review | Demonstrated in lab, real backlog dispute resolution is still attended; missing evidence stays INCONCLUSIVE |
|
||||||
|
| False-adaptation consequences are identified for the engagement | Audit-core tenant disclosure, forged foreign append or leftover identity could be concealed by a false pass | Zero tolerance for automatic acceptance of regression; 0/7 synthetic defects is no guarantee against unseen attacks |
|
||||||
|
| Execution, custody, timing, cleanup and target revision are enforced | Audit-core's PhaseContract is durable intent, explicitly not a runnable Scenario; tests use supplied observations | Not met: no causal/time-window runtime, custody or cluster driver, independent cleanup collector or fresh admitted window here |
|
||||||
|
| Capability, economics and maintenance costs justify adoption | Token-free structural discovery, deterministic crystallization and invalid first authoring timer | Not met; T01/T06/T07 retain the unanswered measurements |
|
||||||
|
|
||||||
|
The audit-core material referenced here is the checked-in
|
||||||
|
`usecases/audit_core_e2_tenant_boundary.py` contract and its fixture-based tests,
|
||||||
|
not a fresh inspection of any deployment or a revalidation of the 2026-08-22
|
||||||
|
engagement. A real run needs the exact engagement's target-owner acceptance,
|
||||||
|
independent observation/cleanup receipts and appropriate approved access. A
|
||||||
|
successful historical receipt cannot substitute for those. The existing T01,
|
||||||
|
T06 and T07 tasks retain the experiment blockers; this review creates no
|
||||||
|
production implementation commitment or new work record.
|
||||||
|
|
@ -1,5 +1,9 @@
|
||||||
# TestDriver Improvement Loop
|
# TestDriver Improvement Loop
|
||||||
|
|
||||||
|
> Scope update (2026-09-28): lifecycle Energy, Temperature, Confidence, Campaign,
|
||||||
|
> Metabolism and automated Retirement proposals below are historical, superseded
|
||||||
|
> by the [TD-WP-0003 settlement](TestDriverGeneralisationReview.md).
|
||||||
|
|
||||||
**Status:** Concept v0.1
|
**Status:** Concept v0.1
|
||||||
**Purpose:** Establish a self-improvement loop that keeps the implementation of `test-driver` aligned with its conceptual model while generating evidence about which concepts actually work.
|
**Purpose:** Establish a self-improvement loop that keeps the implementation of `test-driver` aligned with its conceptual model while generating evidence about which concepts actually work.
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -1,5 +1,9 @@
|
||||||
# TestDriver Research Prototype — Initial Milestones
|
# TestDriver Research Prototype — Initial Milestones
|
||||||
|
|
||||||
|
> Scope update (2026-09-28): lifecycle Energy, Temperature, Confidence, Campaign,
|
||||||
|
> Metabolism and automated Retirement proposals below are historical, superseded
|
||||||
|
> by the [TD-WP-0003 settlement](TestDriverGeneralisationReview.md).
|
||||||
|
|
||||||
**Status:** v0.1 — **canonical milestone sequence**
|
**Status:** v0.1 — **canonical milestone sequence**
|
||||||
**Purpose:** Establish the minimum evidence-producing development loop needed to validate the core test-driver concepts.
|
**Purpose:** Establish the minimum evidence-producing development loop needed to validate the core test-driver concepts.
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -1,5 +1,9 @@
|
||||||
# Stage 1 Test Driver Validation
|
# Stage 1 Test Driver Validation
|
||||||
|
|
||||||
|
> Scope update (2026-09-28): lifecycle Energy, Temperature, Confidence, Campaign,
|
||||||
|
> Metabolism and automated Retirement proposals below are historical, superseded
|
||||||
|
> by the [TD-WP-0003 settlement](TestDriverGeneralisationReview.md).
|
||||||
|
|
||||||
The biggest risk is not technical feasibility. It is that **test-driver becomes conceptually elegant but too broad to prove itself quickly**.
|
The biggest risk is not technical feasibility. It is that **test-driver becomes conceptually elegant but too broad to prove itself quickly**.
|
||||||
|
|
||||||
I would improve the odds of success by treating the next phase as an experiment in whether three specific ideas actually work:
|
I would improve the odds of success by treating the next phase as an experiment in whether three specific ideas actually work:
|
||||||
|
|
|
||||||
|
|
@ -83,3 +83,13 @@ Measured in `tests/test_classification.py`. **False Adaptation Rate = 0/7.**
|
||||||
|
|
||||||
The matrix is asserted in `tests/test_lab_ground_truth.py`. A moved cell fails
|
The matrix is asserted in `tests/test_lab_ground_truth.py`. A moved cell fails
|
||||||
the suite: changing the measuring instrument must be a deliberate, reviewed act.
|
the suite: changing the measuring instrument must be a deliberate, reviewed act.
|
||||||
|
|
||||||
|
## T03 extension — 2026-09-28
|
||||||
|
|
||||||
|
M25–M31 drop identifiers while changing field order, dialog placement, anchor
|
||||||
|
affordances, fieldset grouping, textarea input, unrelated-form competition and
|
||||||
|
sidebar placement. M32–M38 pair the same mechanisms with identifiers preserved.
|
||||||
|
All fourteen are mechanical by the pre-run synthetic contract: grant behavior
|
||||||
|
and the endpoint are unchanged. These additions are agent-authored instrument
|
||||||
|
labels, not a claim of new human review. Both arms and the full classifier are
|
||||||
|
measured in `research/evidence/2026-09-28-e001.json`.
|
||||||
|
|
|
||||||
|
|
@ -98,6 +98,7 @@ class LabApp:
|
||||||
ui_test_ids: str = "stable"
|
ui_test_ids: str = "stable"
|
||||||
ui_field_names: str = "canonical"
|
ui_field_names: str = "canonical"
|
||||||
ui_confirm_revoke: bool = False
|
ui_confirm_revoke: bool = False
|
||||||
|
ui_form_layout: str = "plain"
|
||||||
api_path_style: str = "long"
|
api_path_style: str = "long"
|
||||||
api_grant_path: str = "grant"
|
api_grant_path: str = "grant"
|
||||||
tenant_sharing_announced: bool = False
|
tenant_sharing_announced: bool = False
|
||||||
|
|
|
||||||
|
|
@ -66,6 +66,14 @@ def _render(app: LabApp, user_id: str, resource_id: str) -> str:
|
||||||
else subject_field + permission_field
|
else subject_field + permission_field
|
||||||
)
|
)
|
||||||
|
|
||||||
|
if app.ui_form_layout == "fieldset":
|
||||||
|
fields = f'<fieldset><legend>Access</legend>{fields}</fieldset>'
|
||||||
|
elif app.ui_form_layout == "textarea":
|
||||||
|
fields = fields.replace(
|
||||||
|
f'<input id="subject" name="{subject_name}" data-td="subject">',
|
||||||
|
f'<textarea id="subject" name="{subject_name}" data-td="subject"></textarea>',
|
||||||
|
)
|
||||||
|
|
||||||
share_form = (
|
share_form = (
|
||||||
f'<form id="share" data-td="share-form" method="post" '
|
f'<form id="share" data-td="share-form" method="post" '
|
||||||
f'action="{_resource_path(app, resource_id)}/{app.api_grant_path}">'
|
f'action="{_resource_path(app, resource_id)}/{app.api_grant_path}">'
|
||||||
|
|
@ -73,6 +81,11 @@ def _render(app: LabApp, user_id: str, resource_id: str) -> str:
|
||||||
f'<{tag}{href}{role} id="share-submit" data-td="share-submit">'
|
f'<{tag}{href}{role} id="share-submit" data-td="share-submit">'
|
||||||
f"{share_label}</{tag}></form>"
|
f"{share_label}</{tag}></form>"
|
||||||
)
|
)
|
||||||
|
if app.ui_form_layout == "decoy":
|
||||||
|
share_form = (
|
||||||
|
'<form method="post" action="/search"><input name="query">'
|
||||||
|
'<button>Search resources</button></form>' + share_form
|
||||||
|
)
|
||||||
if app.ui_share_control == "removed":
|
if app.ui_share_control == "removed":
|
||||||
# The control is gone from the UI. The API endpoint is still open — the
|
# The control is gone from the UI. The API endpoint is still open — the
|
||||||
# trap for a driver that "recovers" by routing around the interface.
|
# trap for a driver that "recovers" by routing around the interface.
|
||||||
|
|
@ -107,6 +120,11 @@ def _render(app: LabApp, user_id: str, resource_id: str) -> str:
|
||||||
f'<h1 data-td="resource-title">Resource {resource_id}</h1>'
|
f'<h1 data-td="resource-title">Resource {resource_id}</h1>'
|
||||||
f'<ul data-td="grant-list">{grants}</ul>{share_form}{revoke_form}'
|
f'<ul data-td="grant-list">{grants}</ul>{share_form}{revoke_form}'
|
||||||
)
|
)
|
||||||
|
if app.ui_form_layout == "sidebar":
|
||||||
|
body = (
|
||||||
|
f'<main><h1>Resource {resource_id}</h1><ul>{grants}</ul>{revoke_form}</main>'
|
||||||
|
f'<aside aria-label="Access management">{share_form}</aside>'
|
||||||
|
)
|
||||||
if app.ui_dom_style == "nested":
|
if app.ui_dom_style == "nested":
|
||||||
body = (
|
body = (
|
||||||
'<div class="shell"><section class="panel"><div class="panel-inner">'
|
'<div class="shell"><section class="panel"><div class="panel-inner">'
|
||||||
|
|
|
||||||
|
|
@ -1,12 +1,13 @@
|
||||||
"""The labelled mutation catalogue — the project's measuring instrument.
|
"""The labelled mutation catalogue — the project's measuring instrument.
|
||||||
|
|
||||||
Every claim test-driver makes is measured against this catalogue, so its quality
|
The structural adaptation claims are measured against this catalogue, so its quality
|
||||||
caps the credibility of every downstream result. Six mutations, as the milestones
|
caps the credibility of every downstream result. Six mutations, as the milestones
|
||||||
document originally sketched, cannot support any statement about precision or
|
document originally sketched, cannot support any statement about precision or
|
||||||
recall; there are twenty here.
|
recall; there are 38 here after TD-WP-0003-T03.
|
||||||
|
|
||||||
Each mutation carries a **ground-truth label**, decided by a human from the use
|
Each mutation carries a **ground-truth label** recorded before its measurement.
|
||||||
case and recorded before any run:
|
The original labels came with the use case; M25–M38 are explicitly agent-authored
|
||||||
|
synthetic extensions (not newly human-reviewed):
|
||||||
|
|
||||||
MECHANICAL the surface changed; protected semantics are identical.
|
MECHANICAL the surface changed; protected semantics are identical.
|
||||||
test-driver should recover and report an adaptation.
|
test-driver should recover and report an adaptation.
|
||||||
|
|
@ -216,11 +217,37 @@ CATALOGUE: tuple[Mutation, ...] = (
|
||||||
"A user of another tenant reads the resource.", _m20),
|
"A user of another tenant reads the resource.", _m20),
|
||||||
)
|
)
|
||||||
|
|
||||||
|
# T03 extension: labels fixed from the synthetic contract before execution.
|
||||||
|
# These are agent-authored instrument variants, not newly human-reviewed labels.
|
||||||
|
# Paired variants isolate identifier loss from the structural mechanism.
|
||||||
|
EXTENDED_LAYOUTS = (
|
||||||
|
("M25", "M32", "Reordered fields", {"ui_field_order": "reversed"}),
|
||||||
|
("M26", "M33", "Relocated into an open dialog", {"ui_share_control": "modal"}),
|
||||||
|
("M27", "M34", "Anchor affordances", {"ui_button_element": "anchor"}),
|
||||||
|
("M28", "M35", "Fields grouped in a fieldset", {"ui_form_layout": "fieldset"}),
|
||||||
|
("M29", "M36", "Subject input becomes a textarea", {"ui_form_layout": "textarea"}),
|
||||||
|
("M30", "M37", "Unrelated form precedes the control", {"ui_form_layout": "decoy"}),
|
||||||
|
("M31", "M38", "Control moved after revoke into a sidebar", {"ui_form_layout": "sidebar"}),
|
||||||
|
)
|
||||||
|
for dropped_id, preserved_id, title, flags in EXTENDED_LAYOUTS:
|
||||||
|
for mutation_id, preserved in ((dropped_id, False), (preserved_id, True)):
|
||||||
|
CATALOGUE += (Mutation(
|
||||||
|
mutation_id, title + ("; ids preserved" if preserved else "; ids dropped"),
|
||||||
|
"MECHANICAL", "ui",
|
||||||
|
"Agent-authored synthetic variant: the same explicit grant form and "
|
||||||
|
"endpoint remain available; stored permissions and enforcement are unchanged.",
|
||||||
|
lambda app, flags=flags, preserved=preserved: _m(
|
||||||
|
app, **flags, ui_test_ids="stable" if preserved else "dropped"
|
||||||
|
),
|
||||||
|
preserves_test_ids=preserved,
|
||||||
|
),)
|
||||||
|
|
||||||
|
|
||||||
BY_ID: dict[str, Mutation] = {m.id: m for m in CATALOGUE}
|
BY_ID: dict[str, Mutation] = {m.id: m for m in CATALOGUE}
|
||||||
|
|
||||||
|
|
||||||
def expected_classification(mutation_id: str) -> Label:
|
def expected_classification(mutation_id: str) -> Label:
|
||||||
"""Ground truth. Recorded by a human before any run — never inferred."""
|
"""Pre-recorded instrument labels, never inferred from run outcomes."""
|
||||||
return BY_ID[mutation_id].label
|
return BY_ID[mutation_id].label
|
||||||
|
|
||||||
|
|
||||||
|
|
|
||||||
134
lab/workflows.py
Normal file
134
lab/workflows.py
Normal file
|
|
@ -0,0 +1,134 @@
|
||||||
|
"""Synthetic workflow lab, implemented from usecases/generalisation-spec.md."""
|
||||||
|
from copy import deepcopy
|
||||||
|
from dataclasses import dataclass
|
||||||
|
|
||||||
|
from testdriver.actions import Surface
|
||||||
|
from testdriver.drivers import Realization
|
||||||
|
|
||||||
|
|
||||||
|
class WorkflowDriver:
|
||||||
|
"""A completed denial is a protocol response, not a failed realization."""
|
||||||
|
|
||||||
|
surface = Surface("workflow", "direct")
|
||||||
|
|
||||||
|
def __init__(self, lab):
|
||||||
|
self.lab = lab
|
||||||
|
|
||||||
|
def realize(self, actor, action):
|
||||||
|
action.check_surface(self.surface.id)
|
||||||
|
response = self.lab.request(actor.credentials["token"], action.name, **action.args)
|
||||||
|
return Realization(self.surface.id, {"operation": action.name, "response": response})
|
||||||
|
|
||||||
|
|
||||||
|
class ApprovalLab:
|
||||||
|
tokens = {"requester": "fixture-requester", "reviewer": "fixture-reviewer",
|
||||||
|
"delegate": "fixture-delegate"}
|
||||||
|
|
||||||
|
def __init__(self, defect=None):
|
||||||
|
self.defect = defect
|
||||||
|
self.state = "draft"
|
||||||
|
self.reviewer = "reviewer"
|
||||||
|
self.approved_by = None
|
||||||
|
self.receipts = []
|
||||||
|
|
||||||
|
def may_execute(self, actor):
|
||||||
|
return actor == "requester" and (
|
||||||
|
self.state == "approved" or self.defect == "premature-execution"
|
||||||
|
)
|
||||||
|
|
||||||
|
def request(self, token, operation, **args):
|
||||||
|
actor = next((a for a, t in self.tokens.items() if t == token), None)
|
||||||
|
allowed = False
|
||||||
|
if operation == "submit":
|
||||||
|
allowed = actor == "requester" and self.state == "draft"
|
||||||
|
if allowed:
|
||||||
|
self.state = "pending"
|
||||||
|
elif operation == "delegate":
|
||||||
|
allowed = (actor == self.reviewer and self.state == "pending"
|
||||||
|
and args["to"] == "delegate")
|
||||||
|
if allowed:
|
||||||
|
self.reviewer = args["to"]
|
||||||
|
elif operation == "approve":
|
||||||
|
allowed = self.state == "pending" and (
|
||||||
|
actor == self.reviewer
|
||||||
|
or (actor == "requester" and self.defect == "self-approval")
|
||||||
|
or (actor == "reviewer" and self.defect == "stale-delegation")
|
||||||
|
)
|
||||||
|
if allowed:
|
||||||
|
self.state, self.approved_by = "approved", actor
|
||||||
|
elif operation == "execute":
|
||||||
|
allowed = self.may_execute(actor)
|
||||||
|
if allowed:
|
||||||
|
self.state = "executed"
|
||||||
|
else:
|
||||||
|
raise ValueError(operation)
|
||||||
|
receipt = {"actor": actor, "operation": operation, "allowed": allowed}
|
||||||
|
self.receipts.append(receipt)
|
||||||
|
return deepcopy(receipt)
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class ApprovalObserver:
|
||||||
|
lab: ApprovalLab
|
||||||
|
name: str = "approval-observer"
|
||||||
|
|
||||||
|
def snapshot(self):
|
||||||
|
return deepcopy({
|
||||||
|
"state": self.lab.state,
|
||||||
|
"reviewer": self.lab.reviewer,
|
||||||
|
"approved_by": self.lab.approved_by,
|
||||||
|
"receipts": self.lab.receipts,
|
||||||
|
"execution_gate": {actor: self.lab.may_execute(actor) for actor in self.lab.tokens},
|
||||||
|
})
|
||||||
|
|
||||||
|
|
||||||
|
class TenantLab:
|
||||||
|
tokens = {"admin-a": "fixture-admin-a", "admin-b": "fixture-admin-b"}
|
||||||
|
tenants = {"admin-a": "A", "admin-b": "B"}
|
||||||
|
|
||||||
|
def __init__(self, defect=None):
|
||||||
|
self.defect = defect
|
||||||
|
self.records = {}
|
||||||
|
self.deleted = {}
|
||||||
|
self.receipts = []
|
||||||
|
|
||||||
|
def read(self, actor, tenant, resource):
|
||||||
|
if self.tenants.get(actor) != tenant and self.defect != "foreign-read":
|
||||||
|
return {"allowed": False, "content": None}
|
||||||
|
return {"allowed": (tenant, resource) in self.records,
|
||||||
|
"content": self.records.get((tenant, resource))}
|
||||||
|
|
||||||
|
def request(self, token, operation, tenant, resource="R", content=None):
|
||||||
|
actor = next((a for a, t in self.tokens.items() if t == token), None)
|
||||||
|
allowed = self.tenants.get(actor) == tenant
|
||||||
|
key = (tenant, resource)
|
||||||
|
if operation == "create":
|
||||||
|
allowed = allowed and key not in self.records
|
||||||
|
if allowed:
|
||||||
|
self.records[key] = (self.deleted.get(key, content)
|
||||||
|
if self.defect == "resurrect-content" else content)
|
||||||
|
elif operation == "delete":
|
||||||
|
allowed = (allowed or self.defect == "foreign-delete") and key in self.records
|
||||||
|
if allowed:
|
||||||
|
self.deleted[key] = self.records.pop(key)
|
||||||
|
if self.defect == "unscoped-delete":
|
||||||
|
self.records = {k: v for k, v in self.records.items() if k[1] != resource}
|
||||||
|
else:
|
||||||
|
raise ValueError(operation)
|
||||||
|
receipt = {"actor": actor, "operation": operation, "tenant": tenant, "allowed": allowed}
|
||||||
|
self.receipts.append(receipt)
|
||||||
|
return deepcopy(receipt)
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class TenantObserver:
|
||||||
|
lab: TenantLab
|
||||||
|
name: str = "tenant-observer"
|
||||||
|
|
||||||
|
def snapshot(self):
|
||||||
|
return deepcopy({
|
||||||
|
"records": {tenant: self.lab.records.get((tenant, "R")) for tenant in ("A", "B")},
|
||||||
|
"reads": {f"{actor}:{tenant}": self.lab.read(actor, tenant, "R")
|
||||||
|
for actor in self.lab.tokens for tenant in ("A", "B")},
|
||||||
|
"receipts": self.lab.receipts,
|
||||||
|
})
|
||||||
|
|
@ -1,6 +1,6 @@
|
||||||
# Concept ↔ Implementation Fitness Map
|
# Concept ↔ Implementation Fitness Map
|
||||||
|
|
||||||
**Updated:** 2026-08-23 (TD-WP-0002-T10)
|
**Updated:** 2026-09-28 (TD-WP-0003-T02/T03/T04)
|
||||||
|
|
||||||
Traces each important concept to the implementation, experiment and evidence that
|
Traces each important concept to the implementation, experiment and evidence that
|
||||||
support it. **Unsupported entries are the point of this map** — a concept with no
|
support it. **Unsupported entries are the point of this map** — a concept with no
|
||||||
|
|
@ -28,43 +28,39 @@ were aspirational, not evidenced.
|
||||||
|
|
||||||
| Concept | Level | Implementation | Experiment | Evidence | Open question |
|
| Concept | Level | Implementation | Experiment | Evidence | Open question |
|
||||||
|---|---|---|---|---|---|
|
|---|---|---|---|---|---|
|
||||||
| `C-use-case` | C1 | `intent.py` | — | — | Is a use case expressible without leaking mechanics? |
|
| `C-use-case` | C1 | `intent.py`, three reference scenarios | T02 | `tests/test_generalisation.py` | Three synthetic use cases fit; external authoring still unmeasured. |
|
||||||
| `C-actor-isolation` | **C2** | `world.py`, `runner.py` | E-001 | `td://self/actor-isolation`, examined every run | **F-0003 resolved** — automatic canaries; observable on every scenario. |
|
| `C-actor-isolation` | **C2** | `world.py`, `runner.py` | E-001 | `td://self/actor-isolation`, examined every run | **F-0003 resolved** — automatic canaries; observable on every scenario. |
|
||||||
| `C-semantic-action` | **C2** | `actions.py`, `agentic.py` | E-001 (partial) | T07 arm comparison | **F-0005** — supported only where stable identifiers are absent. Narrower than the concept model claims. |
|
| `C-semantic-action` | **C2** | `actions.py`, `agentic.py` | E-001 | 2026-09-28 two-arm receipt | **F-0005** — supported only where stable identifiers are absent. Narrower than the concept model claims. |
|
||||||
| `C-oracle-independence` | **C2** | `runner.py`, `oracles.py` | E-001, E-003 | — | Independence of components ≠ independence of belief. (H-004) |
|
| `C-oracle-independence` | **C2** | `runner.py`, `oracles.py` | E-001, E-003 | — | Independence of components ≠ independence of belief. (H-004) |
|
||||||
| `C-evidence-pack` | C1 | `evidence.py` | — | `td://self/evidence-reproducibility` | Verdicts are reproducible from S3 alone, on passing and failing runs. |
|
| `C-evidence-pack` | C1 | `evidence.py` | — | `td://self/evidence-reproducibility` | Verdicts are reproducible from S3 alone, on passing and failing runs. |
|
||||||
| `C-observation-channel` | C1 | `lab/app.py` | — | — | **D-07** — required of every system under test. Adoption cost unknown. |
|
| `C-observation-channel` | C1 | `lab/app.py` | — | — | **D-07** — required of every system under test. Adoption cost unknown. |
|
||||||
| `C-adaptation` | **C2** | `classification.py`, `agentic.py` | E-001 | T08 matrix | 11/12 mechanical absorbed without a human. |
|
| `C-adaptation` | **C2** | `classification.py`, `agentic.py` | E-001 | T08 matrix | 26/27 mechanical absorbed; ten dropped-id cases now included. |
|
||||||
| `C-classification` | **C2** | `classification.py` | E-001, E-003 | T08 matrix, FAR 0/7 | **F-0006** — cannot infer SEMANTIC vs DEFECT; collapses to one escalating outcome. |
|
| `C-classification` | **C2** | `classification.py` | E-001, E-003 | T08 matrix, FAR 0/7 | **F-0006** — cannot infer SEMANTIC vs DEFECT; collapses to one escalating outcome. |
|
||||||
| `C-crystallization` | **C2** | `crystallization.py` | E-002 | T09 agreement matrix | Fidelity supported; **F-0007** — the economic case is unmeasurable with a token-free runtime. |
|
| `C-crystallization` | **C2** | `crystallization.py` | E-002 | T09 agreement matrix | Fidelity supported; **F-0007** — the economic case is unmeasurable with a token-free runtime. |
|
||||||
| `C-intent-provenance` | C1 | `provenance.py` | E-003 | `td://self/intent-independence` | Constrains provenance, not quality. Accepted residual. |
|
| `C-intent-provenance` | C1 | `provenance.py` | E-003 | `td://self/intent-independence` | Constrains provenance, not quality. Accepted residual. |
|
||||||
| `C-lineage` | C1 | `scenario.py`, generated headers | E-002 | generated module docstring | Parent pointer plus provenance in the artefact; no graph. |
|
| `C-lineage` | C1 | `scenario.py`, generated headers | E-002 | generated module docstring | Parent pointer plus provenance in the artefact; no graph. |
|
||||||
| `C-energy` | C0 | `energy.py`, capture only | — | none | Dormant. **Gated**: if the next workplan ends with no decision having used the history, remove. |
|
| `C-energy` | Removed | — | T04 | Generalisation review | No demonstrated decision; not an implementation promise. |
|
||||||
| `C-temperature` | **C0 — under review** | — | — | none | **F-0008** — crystallization was built without it; measured stability did the work. Gated for removal. |
|
| `C-temperature` | Removed | — | T04 | Generalisation review | No demonstrated decision; not an implementation promise. |
|
||||||
| `C-confidence` | C0 | — | — | none | Declared, unimplemented, never consulted. Same position as `C-temperature` without a built alternative. |
|
| `C-confidence` | Removed | — | T04 | Generalisation review | No demonstrated decision; not an implementation promise. |
|
||||||
| `C-campaign` | C0 | — | — | — | Deferred. |
|
| `C-campaign` | Removed | — | T04 | Generalisation review | No demonstrated decision; not an implementation promise. |
|
||||||
| `C-metabolism` | C0 | — | — | — | Deferred. Depends on C-energy. |
|
| `C-metabolism` | Removed | — | T04 | Generalisation review | No demonstrated decision; not an implementation promise. |
|
||||||
| `C-retirement` | C0 | — | — | — | Deferred. Depends on C-energy. |
|
| `C-retirement` | Removed | — | T04 | Generalisation review | No demonstrated decision; not an implementation promise. |
|
||||||
| `C-security-mutation` | C1 | `lab/mutations.py` | E-003 | `lab/GROUND-TRUTH.md` | Catalogue is hand-written; no derivation mechanism from use cases yet. |
|
| `C-security-mutation` | C1 | `lab/mutations.py` | E-003 | `lab/GROUND-TRUTH.md` | Catalogue is hand-written; no derivation mechanism from use cases yet. |
|
||||||
|
|
||||||
## Orphan check
|
## Generalisation and compression review — 2026-09-28
|
||||||
|
|
||||||
**Conceptual orphans** — concepts with no planned implementation in `TD-WP-0002`:
|
No new kernel concept was required for delegation/sequencing or tenant lifecycle.
|
||||||
`C-temperature`, `C-confidence`, `C-campaign`, `C-metabolism`, `C-retirement`.
|
Ordered `Step`s supply the Schedule; scenario `variant` identifies seeded defects.
|
||||||
|
No independent scheduler or Variant class was needed. Integration/security lenses
|
||||||
|
remain descriptive groupings; no Lens runtime object was consulted or validated.
|
||||||
|
The generic observer interface works, but each new domain needs a custom snapshot
|
||||||
|
collector: observation-channel adoption cost remains a real concern.
|
||||||
|
|
||||||
All five are deferred *by explicit decision*, not oversight. They are the group
|
Temperature, Energy, Confidence, Campaign, Metabolism and Retirement are removed,
|
||||||
most at risk of being built because they are easy and satisfying, and never
|
not deferred to another gate. `energy.py`, its exports and evidence field are
|
||||||
validated. They are revisited at T10, where the question is not "when do we build
|
removed; H-005 is dormant-indefinite and F-0008 is resolved. Ordinary evidence,
|
||||||
these" but "does the evidence justify keeping them in the model at all".
|
judgments, realization metrics and lineage remain. See
|
||||||
|
[review](../../docs/TestDriverGeneralisationReview.md) for compatibility and limits.
|
||||||
|
|
||||||
**Implementation orphans** — none. Every module in `src/testdriver/` traces to a
|
Previous compression (T10): `Verdict.SUSPICIOUS`, `Step.expect_refusal`,
|
||||||
concept above.
|
`ActorIsolationError`, `World.seed`, `EvidencePack.latest()`, `Trajectory.method`.
|
||||||
|
|
||||||
**Removed at T10** (compression pass — see the gate review § 3):
|
|
||||||
`Verdict.SUSPICIOUS`, `Step.expect_refusal`, `ActorIsolationError`,
|
|
||||||
`World.seed`, `EvidencePack.latest()`, `Trajectory.method`. Each was declared and
|
|
||||||
never used; `SUSPICIOUS` was additionally a verdict no oracle could emit.
|
|
||||||
|
|
||||||
**Gated for removal**: `C-temperature` (F-0008) and `energy.py`. Both survive the
|
|
||||||
current review on the strength of being cheap, not of being used. If the next
|
|
||||||
workplan closes without a decision consulting either, they go.
|
|
||||||
|
|
|
||||||
|
|
@ -7,3 +7,4 @@ file is a pointer table, not a second source of truth.
|
||||||
| Ref | Title | Hub ID | Source |
|
| Ref | Title | Hub ID | Source |
|
||||||
|---|---|---|---|
|
|---|---|---|---|
|
||||||
| D-01 … D-07 | Stratified evidence and claim provenance as the basis for adaptation safety | `fef5213f-ce9b-44c2-b327-a0b0ba4b6270` | `docs/TestDriverClassificationDesign.md` |
|
| D-01 … D-07 | Stratified evidence and claim provenance as the basis for adaptation safety | `fef5213f-ce9b-44c2-b327-a0b0ba4b6270` | `docs/TestDriverClassificationDesign.md` |
|
||||||
|
| TD-WP-0003 T04/T05/T08 | Python claims, concept removal and not-ready assessment | `d80734d1-01b8-4a98-8e72-c86f1806d572` | `docs/TestDriverGeneralisationReview.md` |
|
||||||
|
|
|
||||||
19
research/evidence/2026-09-28-authoring.json
Normal file
19
research/evidence/2026-09-28-authoring.json
Normal file
|
|
@ -0,0 +1,19 @@
|
||||||
|
{
|
||||||
|
"approval": {
|
||||||
|
"started": "2026-09-28T09:56:40.697848+00:00",
|
||||||
|
"first_passing_tests": "2026-09-28T09:57:56.524430+00:00",
|
||||||
|
"scenario_lines": 73,
|
||||||
|
"scenario": "scenarios/delegated_approval.py",
|
||||||
|
"limitation": "Timer started after code composition; interval is not authoring time."
|
||||||
|
},
|
||||||
|
"tenant": {
|
||||||
|
"started": "2026-09-28T09:57:56.524430+00:00",
|
||||||
|
"first_passing_tests": "2026-09-28T09:57:57.123098+00:00",
|
||||||
|
"scenario_lines": 70,
|
||||||
|
"scenario": "scenarios/tenant_lifecycle.py",
|
||||||
|
"limitation": "Timer started after code composition; interval is not authoring time."
|
||||||
|
},
|
||||||
|
"valid_authoring_measurement": false,
|
||||||
|
"shared_adapter_and_lab_lines": 134,
|
||||||
|
"shared_test_lines": 80
|
||||||
|
}
|
||||||
4524
research/evidence/2026-09-28-e001.json
Normal file
4524
research/evidence/2026-09-28-e001.json
Normal file
File diff suppressed because it is too large
Load diff
|
|
@ -1,7 +1,7 @@
|
||||||
---
|
---
|
||||||
id: E-001
|
id: E-001
|
||||||
title: Mechanical recovery and defect discrimination over the labelled mutation set
|
title: Mechanical recovery and defect discrimination over the labelled mutation set
|
||||||
status: PLANNED
|
status: EXECUTED
|
||||||
hypotheses: [H-001, H-002, H-004]
|
hypotheses: [H-001, H-002, H-004]
|
||||||
task: TD-WP-0002-T08
|
task: TD-WP-0002-T08
|
||||||
created: "2026-08-22"
|
created: "2026-08-22"
|
||||||
|
|
@ -37,3 +37,18 @@ tokens · wall time · retries.
|
||||||
|
|
||||||
`EXECUTED` 2026-08-22 (T08). Arm A run over 23 mutations; arm B over
|
`EXECUTED` 2026-08-22 (T08). Arm A run over 23 mutations; arm B over
|
||||||
the 12 mechanical ones. Results in H-001 and H-004. FAR 0/7.
|
the 12 mechanical ones. Results in H-001 and H-004. FAR 0/7.
|
||||||
|
|
||||||
|
## Expanded measurement — 2026-09-28 (TD-WP-0003-T03)
|
||||||
|
|
||||||
|
38 mutations, including ten that drop test ids and seven new paired structural
|
||||||
|
controls. Discovery and recorded selectors both pass 17/17 preserved-id cases;
|
||||||
|
on dropped ids discovery passes 9/10 and recorded selectors 0/10. M22 remains
|
||||||
|
unrealizable. The full journey accepts 26/27 mechanical cases; FAR remains 0/7
|
||||||
|
on the unchanged defect set. Claim/provenance indices remain unchanged.
|
||||||
|
|
||||||
|
Reproduce: `PYTHONPATH=src:. python3 tools/measure_e001.py /tmp/e001.json`.
|
||||||
|
[Raw result](../evidence/2026-09-28-e001.json).
|
||||||
|
These selected and correlated synthetic HTML cases do not establish population
|
||||||
|
reliability, visual/browser coverage, model capability or economics. See
|
||||||
|
[review](../../docs/TestDriverGeneralisationReview.md) for the mechanism list and
|
||||||
|
limits. H-001 remains supported only in the narrowed identifier-loss setting.
|
||||||
|
|
|
||||||
|
|
@ -85,3 +85,10 @@ Fully standalone generation would need claims expressible in a serializable form
|
||||||
rather than as Python predicates. That is a real design question — it is the same
|
rather than as Python predicates. That is a real design question — it is the same
|
||||||
question as "should scenarios be YAML", deferred at T04 — and both should be
|
question as "should scenarios be YAML", deferred at T04 — and both should be
|
||||||
answered together at T10, with evidence about which predicates actually recur.
|
answered together at T10, with evidence about which predicates actually recur.
|
||||||
|
|
||||||
|
## Interface settlement — 2026-09-28
|
||||||
|
|
||||||
|
TD-WP-0003-T05 deliberately retains Python predicates and imported original
|
||||||
|
claims. Standalone serialization is not a promised deliverable. See
|
||||||
|
`docs/TestDriverGeneralisationReview.md` for the published contract. The economic
|
||||||
|
finding stays open under TD-WP-0003-T01; no live-model measurement was performed.
|
||||||
|
|
|
||||||
|
|
@ -2,12 +2,12 @@
|
||||||
id: F-0008
|
id: F-0008
|
||||||
type: framework-finding
|
type: framework-finding
|
||||||
class: UNNECESSARY_COMPLEXITY
|
class: UNNECESSARY_COMPLEXITY
|
||||||
status: open
|
status: resolved
|
||||||
discovered: "2026-08-23"
|
discovered: "2026-08-23"
|
||||||
discovered_by: TD-WP-0002-T10
|
discovered_by: TD-WP-0002-T10
|
||||||
workplan: TD-WP-0002
|
workplan: TD-WP-0002
|
||||||
task: TD-WP-0002-T10
|
task: TD-WP-0002-T10
|
||||||
carried_to: next workplan
|
carried_to: TD-WP-0003-T04
|
||||||
---
|
---
|
||||||
|
|
||||||
# F-0008 — Temperature may be redundant; measured stability did the work
|
# F-0008 — Temperature may be redundant; measured stability did the work
|
||||||
|
|
@ -68,3 +68,9 @@ is currently the clearest instance in the corpus.
|
||||||
`Confidence`, `Campaign`, `Metabolism` and `Retirement` are in the same position —
|
`Confidence`, `Campaign`, `Metabolism` and `Retirement` are in the same position —
|
||||||
declared, unimplemented, never consulted — but Temperature is the one with a
|
declared, unimplemented, never consulted — but Temperature is the one with a
|
||||||
built alternative already doing its job, which makes it the decidable case.
|
built alternative already doing its job, which makes it the decidable case.
|
||||||
|
|
||||||
|
## Resolution — 2026-09-28
|
||||||
|
|
||||||
|
T04 removes Temperature from the current concept set. Three synthetic use cases
|
||||||
|
and the expanded structural experiment require no temperature decision; observed
|
||||||
|
trajectory stability continues to govern crystallization. No replacement gate.
|
||||||
|
|
|
||||||
|
|
@ -61,3 +61,18 @@ relayout is untested.
|
||||||
|
|
||||||
- 2026-08-22 `PROPOSED`. No evidence.
|
- 2026-08-22 `PROPOSED`. No evidence.
|
||||||
- 2026-08-22 `EXPERIMENTING`. Partial evidence from T07; awaiting E-001 in full.
|
- 2026-08-22 `EXPERIMENTING`. Partial evidence from T07; awaiting E-001 in full.
|
||||||
|
|
||||||
|
## Expanded measurement — 2026-09-28 (TD-WP-0003-T03)
|
||||||
|
|
||||||
|
38 mutations, including ten that drop test ids and seven new paired structural
|
||||||
|
controls. Discovery and recorded selectors both pass 17/17 preserved-id cases;
|
||||||
|
on dropped ids discovery passes 9/10 and recorded selectors 0/10. M22 remains
|
||||||
|
unrealizable. The full journey accepts 26/27 mechanical cases; FAR remains 0/7
|
||||||
|
on the unchanged defect set. Claim/provenance indices remain unchanged.
|
||||||
|
|
||||||
|
Reproduce: `PYTHONPATH=src:. python3 tools/measure_e001.py /tmp/e001.json`.
|
||||||
|
[Raw result](../evidence/2026-09-28-e001.json).
|
||||||
|
These selected and correlated synthetic HTML cases do not establish population
|
||||||
|
reliability, visual/browser coverage, model capability or economics. See
|
||||||
|
[review](../../docs/TestDriverGeneralisationReview.md) for the mechanism list and
|
||||||
|
limits. H-001 remains supported only in the narrowed identifier-loss setting.
|
||||||
|
|
|
||||||
|
|
@ -1,7 +1,7 @@
|
||||||
---
|
---
|
||||||
id: H-005
|
id: H-005
|
||||||
title: Verification Energy
|
title: Verification Energy
|
||||||
status: PROPOSED
|
status: DORMANT-INDEFINITE
|
||||||
created: "2026-08-22"
|
created: "2026-08-22"
|
||||||
experiments: []
|
experiments: []
|
||||||
concepts: [C-energy]
|
concepts: [C-energy]
|
||||||
|
|
@ -28,9 +28,9 @@ dormant.** Validating it requires event history across many assets over months
|
||||||
history the spike will not accumulate. Implementing a scoring function now would
|
history the spike will not accumulate. Implementing a scoring function now would
|
||||||
produce a number that cannot be checked, which is worse than no number.
|
produce a number that cannot be checked, which is worse than no number.
|
||||||
|
|
||||||
`TD-WP-0002` therefore records **raw immutable `EnergyEvent`s from the first run**
|
`TD-WP-0002` recorded raw events, but no decision used them. TD-WP-0003-T04
|
||||||
and implements no scoring, decay, or selection logic. Events cannot be
|
removed capture and scoring concepts on 2026-09-28. Ordinary run evidence and
|
||||||
reconstructed later; scores can always be computed later.
|
verdicts remain; no energy history is promised for future reconstruction.
|
||||||
|
|
||||||
This is the hypothesis most likely to be **cheaply built and never validated**,
|
This is the hypothesis most likely to be **cheaply built and never validated**,
|
||||||
which is precisely why it is fenced off.
|
which is precisely why it is fenced off.
|
||||||
|
|
@ -38,3 +38,5 @@ which is precisely why it is fenced off.
|
||||||
## Status log
|
## Status log
|
||||||
|
|
||||||
- 2026-08-22 `PROPOSED`, dormant by decision. Event capture only.
|
- 2026-08-22 `PROPOSED`, dormant by decision. Event capture only.
|
||||||
|
|
||||||
|
- 2026-09-28 `DORMANT-INDEFINITE`. Capture removed; no new implementation gate.
|
||||||
|
|
|
||||||
73
scenarios/delegated_approval.py
Normal file
73
scenarios/delegated_approval.py
Normal file
|
|
@ -0,0 +1,73 @@
|
||||||
|
"""Sequenced delegation contract; no new kernel concepts."""
|
||||||
|
from testdriver import (
|
||||||
|
Actor, Cast, Claim, Invariant, Oracle, Provenance, Scenario, SemanticAction,
|
||||||
|
Step, UseCase, VerificationAsset, World,
|
||||||
|
)
|
||||||
|
from lab.workflows import ApprovalLab, ApprovalObserver, WorkflowDriver
|
||||||
|
|
||||||
|
SOURCE = "usecases/generalisation-spec.md#delegated-approval"
|
||||||
|
PROVENANCE = Provenance.AGENT_FROM_SPEC
|
||||||
|
|
||||||
|
|
||||||
|
def receipt(obs, operation, actor, allowed):
|
||||||
|
return obs["receipts"][-1] == {
|
||||||
|
"operation": operation, "actor": actor, "allowed": allowed,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def execution_matches_record(obs):
|
||||||
|
return obs["execution_gate"] == {
|
||||||
|
"requester": obs["state"] == "approved", "reviewer": False, "delegate": False,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
USE_CASE = UseCase(
|
||||||
|
"uc-delegated-approval", "Delegate, approve, execute",
|
||||||
|
"A separate reviewer delegates approval before the requester executes.",
|
||||||
|
PROVENANCE, source_ref=SOURCE,
|
||||||
|
claims=(
|
||||||
|
Claim("approval-premature", "Execution before approval is denied", PROVENANCE,
|
||||||
|
lambda o: receipt(o, "execute", "requester", False) and o["state"] == "draft",
|
||||||
|
"premature", SOURCE),
|
||||||
|
Claim("approval-submit", "Submission enters pending review", PROVENANCE,
|
||||||
|
lambda o: o["state"] == "pending", "submit", SOURCE),
|
||||||
|
Claim("approval-delegate", "Authority passes to the delegate", PROVENANCE,
|
||||||
|
lambda o: o["reviewer"] == "delegate", "delegate", SOURCE),
|
||||||
|
Claim("approval-self", "Requester cannot self-approve", PROVENANCE,
|
||||||
|
lambda o: receipt(o, "approve", "requester", False) and o["state"] == "pending",
|
||||||
|
"self-approve", SOURCE),
|
||||||
|
Claim("approval-former", "Former reviewer cannot approve", PROVENANCE,
|
||||||
|
lambda o: receipt(o, "approve", "reviewer", False) and o["state"] == "pending",
|
||||||
|
"former-approve", SOURCE),
|
||||||
|
Claim("approval-approved", "Delegate approves", PROVENANCE,
|
||||||
|
lambda o: o["state"] == "approved" and o["approved_by"] == "delegate",
|
||||||
|
"approve", SOURCE),
|
||||||
|
Claim("approval-executed", "Requester executes approved work", PROVENANCE,
|
||||||
|
lambda o: o["state"] == "executed", "execute", SOURCE),
|
||||||
|
),
|
||||||
|
invariants=(Invariant("approval-gate", "Execution gate matches stored workflow",
|
||||||
|
PROVENANCE, execution_matches_record, SOURCE),),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def build(defect=None):
|
||||||
|
lab = ApprovalLab(defect)
|
||||||
|
cast = Cast()
|
||||||
|
for actor, token in lab.tokens.items():
|
||||||
|
cast.add(Actor(actor, actor.title(), credentials={"token": token}))
|
||||||
|
schedule = (
|
||||||
|
("premature", "requester", "execute", {}),
|
||||||
|
("submit", "requester", "submit", {}),
|
||||||
|
("delegate", "reviewer", "delegate", {"to": "delegate"}),
|
||||||
|
("self-approve", "requester", "approve", {}),
|
||||||
|
("former-approve", "reviewer", "approve", {}),
|
||||||
|
("approve", "delegate", "approve", {}),
|
||||||
|
("execute", "requester", "execute", {}),
|
||||||
|
)
|
||||||
|
scenario = Scenario("sc-delegated-approval", USE_CASE, tuple(
|
||||||
|
Step(id, actor, SemanticAction(operation, args, frozenset({"workflow"})))
|
||||||
|
for id, actor, operation, args in schedule
|
||||||
|
), variant=defect or "baseline")
|
||||||
|
return (World("w-approval", lab, f"approval-1/{defect or 'baseline'}", cast),
|
||||||
|
WorkflowDriver(lab), ApprovalObserver(lab),
|
||||||
|
VerificationAsset("va-approval", scenario), Oracle())
|
||||||
70
scenarios/tenant_lifecycle.py
Normal file
70
scenarios/tenant_lifecycle.py
Normal file
|
|
@ -0,0 +1,70 @@
|
||||||
|
"""Same-name resources, explicit foreign access and delete/recreate isolation."""
|
||||||
|
from testdriver import (
|
||||||
|
Actor, Cast, Claim, Invariant, Oracle, Provenance, Scenario, SemanticAction,
|
||||||
|
Step, UseCase, VerificationAsset, World,
|
||||||
|
)
|
||||||
|
from lab.workflows import TenantLab, TenantObserver, WorkflowDriver
|
||||||
|
|
||||||
|
SOURCE = "usecases/generalisation-spec.md#tenant-lifecycle"
|
||||||
|
PROVENANCE = Provenance.AGENT_FROM_SPEC
|
||||||
|
|
||||||
|
|
||||||
|
def isolated_reads(obs):
|
||||||
|
for actor, own, foreign in (("admin-a", "A", "B"), ("admin-b", "B", "A")):
|
||||||
|
if obs["reads"][f"{actor}:{foreign}"] != {"allowed": False, "content": None}:
|
||||||
|
return False
|
||||||
|
record = obs["records"][own]
|
||||||
|
if obs["reads"][f"{actor}:{own}"] != {"allowed": record is not None, "content": record}:
|
||||||
|
return False
|
||||||
|
return True
|
||||||
|
|
||||||
|
|
||||||
|
def foreign_delete_refused(obs):
|
||||||
|
return (obs["receipts"][-1]["allowed"] is False
|
||||||
|
and obs["records"] == {"A": "alpha", "B": "beta"})
|
||||||
|
|
||||||
|
|
||||||
|
USE_CASE = UseCase(
|
||||||
|
"uc-tenant-lifecycle", "Tenant-scoped deletion and recreation",
|
||||||
|
"Two administrators use R independently; deleting and recreating A never affects B.",
|
||||||
|
PROVENANCE, source_ref=SOURCE,
|
||||||
|
claims=(
|
||||||
|
Claim("tenant-create-a", "A holds alpha", PROVENANCE,
|
||||||
|
lambda o: o["records"] == {"A": "alpha", "B": None}, "create-a", SOURCE),
|
||||||
|
Claim("tenant-create-b", "Same local name preserves both records", PROVENANCE,
|
||||||
|
lambda o: o["records"] == {"A": "alpha", "B": "beta"}, "create-b", SOURCE),
|
||||||
|
Claim("tenant-deny-a", "A cannot delete B", PROVENANCE,
|
||||||
|
foreign_delete_refused, "foreign-a", SOURCE),
|
||||||
|
Claim("tenant-deny-b", "B cannot delete A", PROVENANCE,
|
||||||
|
foreign_delete_refused, "foreign-b", SOURCE),
|
||||||
|
Claim("tenant-delete", "Deleting A preserves B", PROVENANCE,
|
||||||
|
lambda o: o["records"] == {"A": None, "B": "beta"}, "delete-a", SOURCE),
|
||||||
|
Claim("tenant-recreate", "Recreate does not resurrect or overwrite", PROVENANCE,
|
||||||
|
lambda o: o["records"] == {"A": "new-alpha", "B": "beta"}, "recreate-a", SOURCE),
|
||||||
|
),
|
||||||
|
invariants=(Invariant("tenant-reads", "Read enforcement matches tenant records",
|
||||||
|
PROVENANCE, isolated_reads, SOURCE),),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def build(defect=None):
|
||||||
|
lab = TenantLab(defect)
|
||||||
|
cast = Cast()
|
||||||
|
for actor, token in lab.tokens.items():
|
||||||
|
cast.add(Actor(actor, actor.title(), credentials={"token": token}))
|
||||||
|
schedule = (
|
||||||
|
("create-a", "admin-a", "create", "A", "alpha"),
|
||||||
|
("create-b", "admin-b", "create", "B", "beta"),
|
||||||
|
("foreign-a", "admin-a", "delete", "B", None),
|
||||||
|
("foreign-b", "admin-b", "delete", "A", None),
|
||||||
|
("delete-a", "admin-a", "delete", "A", None),
|
||||||
|
("recreate-a", "admin-a", "create", "A", "new-alpha"),
|
||||||
|
)
|
||||||
|
scenario = Scenario("sc-tenant-lifecycle", USE_CASE, tuple(
|
||||||
|
Step(id, actor, SemanticAction(operation, {"tenant": tenant, "content": content},
|
||||||
|
frozenset({"workflow"})))
|
||||||
|
for id, actor, operation, tenant, content in schedule
|
||||||
|
), variant=defect or "baseline")
|
||||||
|
return (World("w-tenant", lab, f"tenant-1/{defect or 'baseline'}", cast),
|
||||||
|
WorkflowDriver(lab), TenantObserver(lab),
|
||||||
|
VerificationAsset("va-tenant", scenario), Oracle())
|
||||||
|
|
@ -2,7 +2,6 @@
|
||||||
|
|
||||||
from .actions import SemanticAction, Surface, SurfaceNotPermitted
|
from .actions import SemanticAction, Surface, SurfaceNotPermitted
|
||||||
from .drivers import DirectDriver, Realization
|
from .drivers import DirectDriver, Realization
|
||||||
from .energy import EnergyEvent, EnergyEventType
|
|
||||||
from .evidence import EvidencePack, Observation, Stratum
|
from .evidence import EvidencePack, Observation, Stratum
|
||||||
from .intent import Claim, Invariant, UseCase
|
from .intent import Claim, Invariant, UseCase
|
||||||
from .observers import StateObserver, Watch
|
from .observers import StateObserver, Watch
|
||||||
|
|
@ -14,7 +13,7 @@ from .world import Actor, Cast, World
|
||||||
|
|
||||||
__all__ = [
|
__all__ = [
|
||||||
"Actor", "Cast", "Claim", "CollectorIndependenceError",
|
"Actor", "Cast", "Claim", "CollectorIndependenceError",
|
||||||
"DirectDriver", "EnergyEvent", "EnergyEventType", "EvidencePack",
|
"DirectDriver", "EvidencePack",
|
||||||
"InadmissibleProvenance", "Invariant", "Judgment", "Observation", "Oracle",
|
"InadmissibleProvenance", "Invariant", "Judgment", "Observation", "Oracle",
|
||||||
"Provenance", "Realization", "RunResult", "Runner", "Scenario",
|
"Provenance", "Realization", "RunResult", "Runner", "Scenario",
|
||||||
"SemanticAction", "StateObserver", "Step", "Stratum", "Surface",
|
"SemanticAction", "StateObserver", "Step", "Stratum", "Surface",
|
||||||
|
|
|
||||||
|
|
@ -1,45 +0,0 @@
|
||||||
"""Energy events — capture only.
|
|
||||||
|
|
||||||
H-005 is dormant by decision: verification energy is not testable at the current
|
|
||||||
scale, and a scoring function producing a number nobody can check is worse than
|
|
||||||
no number. Events are recorded from the first run because history cannot be
|
|
||||||
reconstructed later; scores always can.
|
|
||||||
|
|
||||||
There is deliberately no score() function in this module.
|
|
||||||
"""
|
|
||||||
|
|
||||||
from __future__ import annotations
|
|
||||||
|
|
||||||
from dataclasses import asdict, dataclass, field
|
|
||||||
from datetime import datetime, timezone
|
|
||||||
from enum import Enum
|
|
||||||
from typing import Any
|
|
||||||
|
|
||||||
|
|
||||||
class EnergyEventType(str, Enum):
|
|
||||||
DEFECT_DETECTED = "DEFECT_DETECTED"
|
|
||||||
REGRESSION_CAUGHT = "REGRESSION_CAUGHT"
|
|
||||||
MECHANICAL_ADAPTATION = "MECHANICAL_ADAPTATION"
|
|
||||||
SEMANTIC_ADAPTATION = "SEMANTIC_ADAPTATION"
|
|
||||||
TEST_DEFECT = "TEST_DEFECT"
|
|
||||||
FALSE_POSITIVE = "FALSE_POSITIVE"
|
|
||||||
DUPLICATE = "DUPLICATE"
|
|
||||||
CRYSTALLIZED = "CRYSTALLIZED"
|
|
||||||
USECASE_DEPRECATED = "USECASE_DEPRECATED"
|
|
||||||
EXECUTED = "EXECUTED"
|
|
||||||
|
|
||||||
|
|
||||||
@dataclass(frozen=True, slots=True)
|
|
||||||
class EnergyEvent:
|
|
||||||
"""Immutable. Energy is derived from event history, never stored as state."""
|
|
||||||
|
|
||||||
asset_id: str
|
|
||||||
run_id: str
|
|
||||||
event_type: EnergyEventType
|
|
||||||
detail: dict[str, Any] = field(default_factory=dict)
|
|
||||||
at: str = field(
|
|
||||||
default_factory=lambda: datetime.now(timezone.utc).isoformat()
|
|
||||||
)
|
|
||||||
|
|
||||||
def as_dict(self) -> dict[str, Any]:
|
|
||||||
return {**asdict(self), "event_type": self.event_type.value}
|
|
||||||
|
|
@ -63,7 +63,6 @@ class EvidencePack:
|
||||||
finished_at: str | None = None
|
finished_at: str | None = None
|
||||||
observations: list[Observation] = field(default_factory=list)
|
observations: list[Observation] = field(default_factory=list)
|
||||||
verdicts: list[dict[str, Any]] = field(default_factory=list)
|
verdicts: list[dict[str, Any]] = field(default_factory=list)
|
||||||
energy_events: list[dict[str, Any]] = field(default_factory=list)
|
|
||||||
provenance_index: dict[str, str] = field(default_factory=dict)
|
provenance_index: dict[str, str] = field(default_factory=dict)
|
||||||
|
|
||||||
def record(self, observation: Observation) -> None:
|
def record(self, observation: Observation) -> None:
|
||||||
|
|
|
||||||
|
|
@ -20,7 +20,12 @@ Predicate = Callable[[Mapping[str, object]], bool]
|
||||||
|
|
||||||
@dataclass(frozen=True, slots=True)
|
@dataclass(frozen=True, slots=True)
|
||||||
class Claim:
|
class Claim:
|
||||||
"""A statement that must hold at a specific point in a scenario."""
|
"""A statement that must hold at a specific point in a scenario.
|
||||||
|
|
||||||
|
Public contract (TD-WP-0003-T05): predicates remain Python callables over
|
||||||
|
independent snapshots. Generated tests import these original claims; no
|
||||||
|
serialized claim language is promised. See TestDriverGeneralisationReview.
|
||||||
|
"""
|
||||||
|
|
||||||
id: str
|
id: str
|
||||||
text: str
|
text: str
|
||||||
|
|
|
||||||
|
|
@ -15,7 +15,6 @@ from typing import Any
|
||||||
|
|
||||||
from .actions import SurfaceNotPermitted
|
from .actions import SurfaceNotPermitted
|
||||||
from .drivers import Driver
|
from .drivers import Driver
|
||||||
from .energy import EnergyEvent, EnergyEventType
|
|
||||||
from .evidence import EvidencePack, Observation, Stratum
|
from .evidence import EvidencePack, Observation, Stratum
|
||||||
from .observers import StateObserver
|
from .observers import StateObserver
|
||||||
from .oracles import Judgment, Oracle, Verdict, overall
|
from .oracles import Judgment, Oracle, Verdict, overall
|
||||||
|
|
@ -125,9 +124,6 @@ class Runner:
|
||||||
**{c.id: c.provenance.value for c in scenario.use_case.claims},
|
**{c.id: c.provenance.value for c in scenario.use_case.claims},
|
||||||
**{i.id: i.provenance.value for i in scenario.use_case.invariants},
|
**{i.id: i.provenance.value for i in scenario.use_case.invariants},
|
||||||
}
|
}
|
||||||
pack.energy_events.append(
|
|
||||||
EnergyEvent(asset.id, run_id, EnergyEventType.EXECUTED).as_dict()
|
|
||||||
)
|
|
||||||
|
|
||||||
isolation = self._isolation_violations()
|
isolation = self._isolation_violations()
|
||||||
self._record(
|
self._record(
|
||||||
|
|
@ -229,8 +225,4 @@ class Runner:
|
||||||
pack.finished_at = datetime.now(timezone.utc).isoformat()
|
pack.finished_at = datetime.now(timezone.utc).isoformat()
|
||||||
|
|
||||||
result_verdict = overall(judgments)
|
result_verdict = overall(judgments)
|
||||||
if result_verdict is Verdict.FAIL:
|
|
||||||
pack.energy_events.append(
|
|
||||||
EnergyEvent(asset.id, run_id, EnergyEventType.DEFECT_DETECTED).as_dict()
|
|
||||||
)
|
|
||||||
return RunResult(run_id, result_verdict, judgments, pack)
|
return RunResult(run_id, result_verdict, judgments, pack)
|
||||||
|
|
|
||||||
|
|
@ -173,3 +173,19 @@ def test_the_control_arm_is_not_a_straw_man():
|
||||||
preserved = [m for m in MECHANICAL if m.preserves_test_ids]
|
preserved = [m for m in MECHANICAL if m.preserves_test_ids]
|
||||||
assert all(realized_ok(m.id, runtime=runtime) for m in preserved)
|
assert all(realized_ok(m.id, runtime=runtime) for m in preserved)
|
||||||
assert len(preserved) >= 9
|
assert len(preserved) >= 9
|
||||||
|
|
||||||
|
|
||||||
|
def test_deciding_axis_has_ten_mutations_and_paired_structural_controls():
|
||||||
|
from lab.mutations import EXTENDED_LAYOUTS
|
||||||
|
from lab.http_api import render_resource_page
|
||||||
|
from lab.mutations import build_lab
|
||||||
|
assert sum(not m.preserves_test_ids for m in MECHANICAL) >= 10
|
||||||
|
for dropped, preserved, _, _ in EXTENDED_LAYOUTS:
|
||||||
|
dropped_app, _ = build_lab(dropped)
|
||||||
|
preserved_app, _ = build_lab(preserved)
|
||||||
|
# Version metadata differs; control markup must differ only in test ids.
|
||||||
|
import re
|
||||||
|
def shape(app):
|
||||||
|
html = render_resource_page(app, "alice", "R")
|
||||||
|
return re.sub(r'\s*data-td(?:-version)?="[^"]*"', "", html)
|
||||||
|
assert shape(dropped_app) == shape(preserved_app)
|
||||||
|
|
|
||||||
|
|
@ -65,6 +65,12 @@ EXPECTED = {
|
||||||
"M24": Classification.MECHANICAL_ADAPTATION,
|
"M24": Classification.MECHANICAL_ADAPTATION,
|
||||||
}
|
}
|
||||||
|
|
||||||
|
EXPECTED.update({
|
||||||
|
f"M{number}": (Classification.UNCHANGED if number in (25, 32)
|
||||||
|
else Classification.MECHANICAL_ADAPTATION)
|
||||||
|
for number in range(25, 39)
|
||||||
|
})
|
||||||
|
|
||||||
|
|
||||||
@pytest.mark.parametrize("mutation_id", sorted(EXPECTED))
|
@pytest.mark.parametrize("mutation_id", sorted(EXPECTED))
|
||||||
def test_response_matrix(baseline, mutation_id):
|
def test_response_matrix(baseline, mutation_id):
|
||||||
|
|
|
||||||
81
tests/test_generalisation.py
Normal file
81
tests/test_generalisation.py
Normal file
|
|
@ -0,0 +1,81 @@
|
||||||
|
"""Independent checks that new use cases discriminate defects and evidence loss."""
|
||||||
|
import json
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
from testdriver import Oracle, Runner, Stratum, Verdict
|
||||||
|
from scenarios.delegated_approval import build as approval
|
||||||
|
from scenarios.tenant_lifecycle import build as tenant
|
||||||
|
|
||||||
|
|
||||||
|
def run(builder, defect=None):
|
||||||
|
world, driver, observer, asset, oracle = builder(defect)
|
||||||
|
return Runner(world, driver, observer, oracle).run(asset)
|
||||||
|
|
||||||
|
|
||||||
|
def test_approval_baseline():
|
||||||
|
assert run(approval).verdict is Verdict.PASS
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("defect,claim", [
|
||||||
|
("premature-execution", "approval-premature"),
|
||||||
|
("self-approval", "approval-self"),
|
||||||
|
("stale-delegation", "approval-former"),
|
||||||
|
])
|
||||||
|
def test_approval_defects(defect, claim):
|
||||||
|
assert run(approval, defect).judgment(claim).verdict is Verdict.FAIL
|
||||||
|
|
||||||
|
|
||||||
|
def test_approval_denials_are_judged_not_skipped():
|
||||||
|
result = run(approval)
|
||||||
|
assert len(result.judgments) == 14
|
||||||
|
assert all(j.verdict is Verdict.PASS for j in result.judgments)
|
||||||
|
|
||||||
|
|
||||||
|
def test_tenant_baseline():
|
||||||
|
assert run(tenant).verdict is Verdict.PASS
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("defect,claim", [
|
||||||
|
("foreign-read", "tenant-reads"),
|
||||||
|
("foreign-delete", "tenant-deny-a"),
|
||||||
|
("unscoped-delete", "tenant-delete"),
|
||||||
|
("resurrect-content", "tenant-recreate"),
|
||||||
|
])
|
||||||
|
def test_tenant_defects(defect, claim):
|
||||||
|
result = run(tenant, defect)
|
||||||
|
assert any(j.assertion_id == claim and j.verdict is Verdict.FAIL for j in result.judgments)
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("builder,defect", [
|
||||||
|
(approval, None), (approval, "self-approval"),
|
||||||
|
(tenant, None), (tenant, "foreign-read"),
|
||||||
|
])
|
||||||
|
def test_verdicts_reproduce_from_serialized_s3_without_actors(builder, defect):
|
||||||
|
result = run(builder, defect)
|
||||||
|
pack = json.loads(result.evidence.to_json())
|
||||||
|
assertions = builder()[3].scenario.use_case
|
||||||
|
snapshots = {o["step_id"]: o["data"] for o in pack["observations"]
|
||||||
|
if o["kind"] == "state_snapshot"}
|
||||||
|
by_id = {a.id: a for a in (*assertions.claims, *assertions.invariants)}
|
||||||
|
for j in pack["verdicts"]:
|
||||||
|
replayed = Oracle().judge(by_id[j["assertion_id"]], snapshots[j["step_id"]], j["step_id"])
|
||||||
|
assert replayed.verdict.value == j["verdict"]
|
||||||
|
actors = set(builder()[0].cast.actors)
|
||||||
|
assert not actors.intersection(o.collector for o in result.evidence.observations
|
||||||
|
if o.stratum is not Stratum.SURFACE)
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("builder", [approval, tenant])
|
||||||
|
def test_missing_evidence_is_inconclusive_for_every_assertion(builder):
|
||||||
|
case = builder()[3].scenario.use_case
|
||||||
|
for assertion in (*case.claims, *case.invariants):
|
||||||
|
assert Oracle().judge(assertion, {"unrelated": True}, None).verdict is Verdict.INCONCLUSIVE
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("builder", [approval, tenant])
|
||||||
|
def test_replay_keeps_judgments_and_claims(builder):
|
||||||
|
case = builder()[3].scenario.use_case
|
||||||
|
before = (case.claims, case.invariants)
|
||||||
|
assert run(builder).judgments == run(builder).judgments
|
||||||
|
assert (case.claims, case.invariants) == before
|
||||||
|
|
@ -130,10 +130,3 @@ def test_seeded_authorization_defect_fails_the_run():
|
||||||
# The claim set is untouched by the failure — there is no path to adapt it.
|
# The claim set is untouched by the failure — there is no path to adapt it.
|
||||||
by_id = {c.id: c for c in USE_CASE.claims}
|
by_id = {c.id: c for c in USE_CASE.claims}
|
||||||
assert by_id["c-bob-revoked"].text == "Bob cannot read R after revocation"
|
assert by_id["c-bob-revoked"].text == "Bob cannot read R after revocation"
|
||||||
|
|
||||||
|
|
||||||
def test_defect_run_emits_an_energy_event():
|
|
||||||
"""Energy events are captured; no score is computed (H-005 is dormant)."""
|
|
||||||
import testdriver.energy as energy
|
|
||||||
|
|
||||||
assert not hasattr(energy, "score")
|
|
||||||
|
|
|
||||||
67
tools/measure_e001.py
Normal file
67
tools/measure_e001.py
Normal file
|
|
@ -0,0 +1,67 @@
|
||||||
|
"""Reproduce E-001: PYTHONPATH=src:. python3 tools/measure_e001.py OUTPUT.json."""
|
||||||
|
import json
|
||||||
|
import sys
|
||||||
|
from collections import defaultdict
|
||||||
|
from datetime import datetime, timezone
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
from lab.mutations import CATALOGUE
|
||||||
|
from scenarios.browser_grant import baseline_recordings, build_agentic, lab_server
|
||||||
|
from scenarios.full_journey import build_journey, journey_lab_server
|
||||||
|
from testdriver import Runner, Stratum
|
||||||
|
from testdriver.agentic import DiscoveryRuntime, RecordedSelectorRuntime
|
||||||
|
from testdriver.classification import classify
|
||||||
|
|
||||||
|
|
||||||
|
def journey(*mutations):
|
||||||
|
with journey_lab_server(*mutations) as (app, tokens, url):
|
||||||
|
world, driver, observer, asset, oracle = build_journey(app, tokens, url)
|
||||||
|
return json.loads(Runner(world, driver, observer, oracle).run(asset).evidence.to_json())
|
||||||
|
|
||||||
|
|
||||||
|
def measure():
|
||||||
|
baseline = journey()
|
||||||
|
selectors = baseline_recordings()
|
||||||
|
rows = []
|
||||||
|
for mutation in CATALOGUE:
|
||||||
|
pack = journey(mutation.id)
|
||||||
|
outcome = classify(baseline, pack)
|
||||||
|
row = {"mutation": mutation.id, "label": mutation.label,
|
||||||
|
"preserves_test_ids": mutation.preserves_test_ids,
|
||||||
|
"classification": outcome.as_dict(), "arms": {}}
|
||||||
|
if mutation.label == "MECHANICAL":
|
||||||
|
for name, runtime in (("discovery", DiscoveryRuntime()),
|
||||||
|
("recorded", RecordedSelectorRuntime(selectors))):
|
||||||
|
with lab_server(mutation.id) as (app, tokens, url):
|
||||||
|
world, driver, observer, asset, oracle = build_agentic(app, tokens, url, runtime)
|
||||||
|
result = Runner(world, driver, observer, oracle).run(asset)
|
||||||
|
surface = result.evidence.of_stratum(Stratum.SURFACE)[0].data
|
||||||
|
row["arms"][name] = {
|
||||||
|
"realized": surface["raised"] is None,
|
||||||
|
"verdict": result.verdict.value,
|
||||||
|
"surface": surface,
|
||||||
|
"judgments": [j.as_dict() for j in result.judgments],
|
||||||
|
}
|
||||||
|
rows.append(row)
|
||||||
|
summary = defaultdict(lambda: {"total": 0, "discovery": 0, "recorded": 0})
|
||||||
|
for row in rows:
|
||||||
|
if not row["arms"]:
|
||||||
|
continue
|
||||||
|
group = summary["preserved" if row["preserves_test_ids"] else "dropped"]
|
||||||
|
group["total"] += 1
|
||||||
|
for arm in ("discovery", "recorded"):
|
||||||
|
group[arm] += int(row["arms"][arm]["realized"]
|
||||||
|
and row["arms"][arm]["verdict"] == "PASS")
|
||||||
|
false_adaptations = [row["mutation"] for row in rows if row["label"] == "DEFECT"
|
||||||
|
and row["classification"]["classification"]
|
||||||
|
in ("UNCHANGED", "MECHANICAL_ADAPTATION")]
|
||||||
|
return {"measured_at": datetime.now(timezone.utc).isoformat(),
|
||||||
|
"scope": "Synthetic server-rendered HTML, deterministic runtimes; no browser engine or live model",
|
||||||
|
"summary": dict(summary), "false_adaptations": false_adaptations,
|
||||||
|
"defect_count": sum(row["label"] == "DEFECT" for row in rows), "rows": rows}
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
result = measure()
|
||||||
|
Path(sys.argv[1]).write_text(json.dumps(result, indent=2) + "\n")
|
||||||
|
print(json.dumps({k: v for k, v in result.items() if k != "rows"}, indent=2))
|
||||||
27
usecases/generalisation-spec.md
Normal file
27
usecases/generalisation-spec.md
Normal file
|
|
@ -0,0 +1,27 @@
|
||||||
|
# Synthetic generalisation contracts — 2026-09-28
|
||||||
|
|
||||||
|
These contracts are written before the corresponding lab implementations for
|
||||||
|
TD-WP-0003-T02. They are agent-authored research specifications, not externally
|
||||||
|
validated product requirements. Scenario claims use `agent-from-spec` and refer
|
||||||
|
here. Same-session authorship limits epistemic independence; passing these labs
|
||||||
|
must not be presented as independent validation of a real product.
|
||||||
|
|
||||||
|
## Delegated approval
|
||||||
|
|
||||||
|
A requester submits a request for review. Only the assigned reviewer can delegate
|
||||||
|
it to another reviewer. The requester cannot approve their own request; the
|
||||||
|
former reviewer loses approval authority after delegation. The delegate may
|
||||||
|
approve only after submission, and only the requester may then execute it.
|
||||||
|
Premature execution and unauthorized approval are refused without state change.
|
||||||
|
The persisted workflow and the independently probed execution gate must agree.
|
||||||
|
A complete run includes premature execution, submission, delegation, attempted
|
||||||
|
self-approval, former-reviewer approval, delegate approval and execution.
|
||||||
|
|
||||||
|
## Tenant lifecycle
|
||||||
|
|
||||||
|
Two tenant administrators independently use the same local resource name `R`.
|
||||||
|
Each can create, read and delete their own tenant's instance. Neither may read or
|
||||||
|
delete the other tenant's instance, including by supplying an explicit foreign
|
||||||
|
tenant id. Deleting tenant A's instance preserves tenant B's content. Recreating
|
||||||
|
A's instance must not resurrect its old content or affect B. At every step,
|
||||||
|
foreign reads remain denied and successful reads match the stored tenant record.
|
||||||
|
|
@ -4,17 +4,25 @@ type: workplan
|
||||||
title: "Generalise the model and settle the open questions"
|
title: "Generalise the model and settle the open questions"
|
||||||
domain: infotech
|
domain: infotech
|
||||||
repo: test-driver
|
repo: test-driver
|
||||||
status: proposed
|
status: blocked
|
||||||
flavor: planning
|
flavor: implementation
|
||||||
owner: codex
|
owner: codex
|
||||||
topic_slug: custodian
|
topic_slug: custodian
|
||||||
created: "2026-08-23"
|
created: "2026-08-23"
|
||||||
updated: "2026-08-23"
|
updated: "2026-09-28"
|
||||||
state_hub_workstream_id: "36a082df-1288-567d-9a0b-f2e0f292786c"
|
state_hub_workstream_id: "36a082df-1288-567d-9a0b-f2e0f292786c"
|
||||||
---
|
---
|
||||||
|
|
||||||
# Generalise the model and settle the open questions
|
# Generalise the model and settle the open questions
|
||||||
|
|
||||||
|
## Closeout status — 2026-09-28
|
||||||
|
|
||||||
|
T02/T03/T04/T05/T08 are done. T01/T06/T07 remain waiting for external experiment
|
||||||
|
choices or a valid independent authoring sample, so the workplan is **blocked**.
|
||||||
|
No new task or workplan was created. The live-model success gate is still unmet;
|
||||||
|
this is a partial closeout, not a declaration that the workplan succeeded.
|
||||||
|
Detailed evidence and interface changes: `docs/TestDriverGeneralisationReview.md`.
|
||||||
|
|
||||||
## Why this workplan exists
|
## Why this workplan exists
|
||||||
|
|
||||||
`TD-WP-0002` passed its gate: False Adaptation Rate 0/7, one asset crystallized
|
`TD-WP-0002` passed its gate: False Adaptation Rate 0/7, one asset crystallized
|
||||||
|
|
@ -76,7 +84,7 @@ consumes the kernel as a library. Two consequences for this workplan:
|
||||||
|
|
||||||
```task
|
```task
|
||||||
id: TD-WP-0003-T01
|
id: TD-WP-0003-T01
|
||||||
status: todo
|
status: wait
|
||||||
priority: high
|
priority: high
|
||||||
state_hub_task_id: "0090ce51-0086-5079-a094-fb63b3b53415"
|
state_hub_task_id: "0090ce51-0086-5079-a094-fb63b3b53415"
|
||||||
```
|
```
|
||||||
|
|
@ -107,11 +115,13 @@ reliably, and the difference will not show in an average.
|
||||||
the live runtime ever runs in the default test suite (recommendation: no — keep
|
the live runtime ever runs in the default test suite (recommendation: no — keep
|
||||||
the suite deterministic and free, run the live arm on demand).
|
the suite deterministic and free, run the live arm on demand).
|
||||||
|
|
||||||
|
**2026-09-28 closeout — wait.** Await an explicit model, fixed run count, cost ceiling and suite policy from Bernd Worsch before the first live call. No live calls or costs were incurred; local heuristic evidence cannot close capability/economics.
|
||||||
|
|
||||||
## Two further use cases
|
## Two further use cases
|
||||||
|
|
||||||
```task
|
```task
|
||||||
id: TD-WP-0003-T02
|
id: TD-WP-0003-T02
|
||||||
status: todo
|
status: done
|
||||||
priority: high
|
priority: high
|
||||||
state_hub_task_id: "eedf5894-e540-5fa7-9dc3-7e2a5c892bd6"
|
state_hub_task_id: "eedf5894-e540-5fa7-9dc3-7e2a5c892bd6"
|
||||||
```
|
```
|
||||||
|
|
@ -133,11 +143,13 @@ express them**, and every such growth is evidence:
|
||||||
|
|
||||||
Record the answer in the fitness map either way.
|
Record the answer in the fitness map either way.
|
||||||
|
|
||||||
|
**2026-09-28 closeout — done.** Added delegated approval and tenant lifecycle from prewritten synthetic specifications, using existing kernel concepts and independent snapshot collectors. Baseline, seven seeded defects, missing evidence and S3 replay are covered in tests/test_generalisation.py. Same-agent synthetic authorship is explicitly limited, not human provenance.
|
||||||
|
|
||||||
## Strengthen the instrument on the deciding axis
|
## Strengthen the instrument on the deciding axis
|
||||||
|
|
||||||
```task
|
```task
|
||||||
id: TD-WP-0003-T03
|
id: TD-WP-0003-T03
|
||||||
status: todo
|
status: done
|
||||||
priority: high
|
priority: high
|
||||||
state_hub_task_id: "95406a5d-bcb6-5bcb-abed-834ad16aeb69"
|
state_hub_task_id: "95406a5d-bcb6-5bcb-abed-834ad16aeb69"
|
||||||
```
|
```
|
||||||
|
|
@ -153,11 +165,13 @@ side proportionate so the comparison stays fair.
|
||||||
Then re-run E-001 and restate H-001 with a number that means something. Be
|
Then re-run E-001 and restate H-001 with a number that means something. Be
|
||||||
prepared for the honest outcome that the narrowed claim narrows further.
|
prepared for the honest outcome that the narrowed claim narrows further.
|
||||||
|
|
||||||
|
**2026-09-28 closeout — done.** Expanded the catalogue to 38 mutations, with ten dropped-identifier mechanisms and seven new paired preserved-id controls. Reproduced E-001: preserved 17/17 in both arms; dropped 9/10 discovery versus 0/10 recorded; full-journey mechanical acceptance 26/27, FAR 0/7. Raw receipt: research/evidence/2026-09-28-e001.json.
|
||||||
|
|
||||||
## Settle every gated concept
|
## Settle every gated concept
|
||||||
|
|
||||||
```task
|
```task
|
||||||
id: TD-WP-0003-T04
|
id: TD-WP-0003-T04
|
||||||
status: todo
|
status: done
|
||||||
priority: medium
|
priority: medium
|
||||||
state_hub_task_id: "458dc9aa-3d0c-5e67-9f89-3ae8a5a5d297"
|
state_hub_task_id: "458dc9aa-3d0c-5e67-9f89-3ae8a5a5d297"
|
||||||
```
|
```
|
||||||
|
|
@ -176,11 +190,13 @@ evidence to say whether they are deferred or dead.
|
||||||
|
|
||||||
Removing a concept is a result. Extending a gate is not.
|
Removing a concept is a result. Extending a gate is not.
|
||||||
|
|
||||||
|
**2026-09-28 closeout — done.** Removed Temperature, Energy, Confidence, Campaign, Metabolism and automated Retirement from the current model; removed energy.py, exports and capture field. H-005 is dormant-indefinite; F-0008 resolved. Compatibility changes and superseded historical proposals are explicit.
|
||||||
|
|
||||||
## Decide how claims are expressed
|
## Decide how claims are expressed
|
||||||
|
|
||||||
```task
|
```task
|
||||||
id: TD-WP-0003-T05
|
id: TD-WP-0003-T05
|
||||||
status: todo
|
status: done
|
||||||
priority: medium
|
priority: medium
|
||||||
state_hub_task_id: "34cc60cd-b966-54f8-a92d-cfe64587822f"
|
state_hub_task_id: "34cc60cd-b966-54f8-a92d-cfe64587822f"
|
||||||
```
|
```
|
||||||
|
|
@ -202,6 +218,8 @@ Three candidate answers, and the middle one should not win by default:
|
||||||
Whatever is chosen, publish it as an interface change — the audit-core session
|
Whatever is chosen, publish it as an interface change — the audit-core session
|
||||||
consumes `Claim` directly.
|
consumes `Claim` directly.
|
||||||
|
|
||||||
|
**2026-09-28 closeout — done.** Published the decision to retain Python predicates and imported original claims, after comparing all three runnable use cases and the audit-core consumer. Claim and Invariant signatures remain compatible; standalone serialization is deliberately rejected. See docs/TestDriverGeneralisationReview.md.
|
||||||
|
|
||||||
## A browser-engine driver, if it is warranted
|
## A browser-engine driver, if it is warranted
|
||||||
|
|
||||||
```task
|
```task
|
||||||
|
|
@ -226,11 +244,13 @@ as the default.
|
||||||
Do not start this before T01–T03. It widens the surface; those three settle what
|
Do not start this before T01–T03. It widens the surface; those three settle what
|
||||||
the surface is for.
|
the surface is for.
|
||||||
|
|
||||||
|
**2026-09-28 closeout — wait.** Still requires the existing Playwright setup/licence decision and T01 live-model prerequisite. T02/T03 are now complete. Structural HTML results do not satisfy the browser-engine/visual experiment.
|
||||||
|
|
||||||
## Measure the cost of expressing a use case
|
## Measure the cost of expressing a use case
|
||||||
|
|
||||||
```task
|
```task
|
||||||
id: TD-WP-0003-T07
|
id: TD-WP-0003-T07
|
||||||
status: todo
|
status: wait
|
||||||
priority: medium
|
priority: medium
|
||||||
state_hub_task_id: "62f08e27-f307-560f-9583-64c08835e29e"
|
state_hub_task_id: "62f08e27-f307-560f-9583-64c08835e29e"
|
||||||
```
|
```
|
||||||
|
|
@ -248,11 +268,13 @@ This feeds the metric the project ultimately cares about — *verified behaviour
|
||||||
unit of human maintenance effort* — and it is the number an adopter will ask for
|
unit of human maintenance effort* — and it is the number an adopter will ask for
|
||||||
first. It cannot be reconstructed later.
|
first. It cannot be reconstructed later.
|
||||||
|
|
||||||
|
**2026-09-28 closeout — wait.** Partial evidence only: concepts, adapter friction, line counts and raw timer receipts recorded in docs/TestDriverGeneralisationReview.md and research/evidence/2026-09-28-authoring.json. Timer placement missed code composition, so first-authoring elapsed costs are invalid and cannot be reconstructed. Await an independently timed fresh implementation by an author who does not copy these fixtures; retain this task, no replacement record.
|
||||||
|
|
||||||
## Readiness review for real-system application
|
## Readiness review for real-system application
|
||||||
|
|
||||||
```task
|
```task
|
||||||
id: TD-WP-0003-T08
|
id: TD-WP-0003-T08
|
||||||
status: todo
|
status: done
|
||||||
priority: high
|
priority: high
|
||||||
state_hub_task_id: "9e040e94-7245-5b4b-b177-cc337f904ceb"
|
state_hub_task_id: "9e040e94-7245-5b4b-b177-cc337f904ceb"
|
||||||
```
|
```
|
||||||
|
|
@ -275,3 +297,5 @@ At minimum the criteria must cover:
|
||||||
|
|
||||||
Then state a verdict, including the verdict "not yet, and here is what is missing".
|
Then state a verdict, including the verdict "not yet, and here is what is missing".
|
||||||
A readiness review that cannot conclude *not ready* is not a review.
|
A readiness review that cannot conclude *not ready* is not a review.
|
||||||
|
|
||||||
|
**2026-09-28 closeout — done.** Published explicit real-system gates and a NOT READY verdict for autonomous application, including D-07 cost, provenance from real backlogs, ambiguity, engagement-specific false-adaptation harm, custody/timing/cleanup and economics. The audit-core contract remains intent plus fixture calibration, not a runnable production engagement.
|
||||||
|
|
|
||||||
Loading…
Add table
Add a link
Reference in a new issue