Complete local generalisation tasks and document remaining experiment blockers

Assistant: codex
Assistant-Model: gpt-6-astra
Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
This commit is contained in:
tegwick 2026-09-28 12:06:24 +02:00
parent d54af628df
commit cf682281d7
36 changed files with 5394 additions and 279 deletions

View file

@ -1,5 +1,9 @@
# test-driver # test-driver
> Scope update (2026-09-28): lifecycle Energy, Temperature, Confidence, Campaign,
> Metabolism and automated Retirement proposals below are historical, superseded
> by the [TD-WP-0003 settlement](docs/TestDriverGeneralisationReview.md).
## Intent ## Intent
`test-driver` is a use-case-driven verification framework for integration, end-to-end, multi-user interaction, authorization, security and resilience testing in software systems that evolve through fast and increasingly agentic development cycles. `test-driver` is a use-case-driven verification framework for integration, end-to-end, multi-user interaction, authorization, security and resilience testing in software systems that evolve through fast and increasingly agentic development cycles.

View file

@ -7,9 +7,12 @@ Tests mature alongside the software they protect: fluid and agentic while
behaviour is changing, deterministic once it settles. See `INTENT.md` for the behaviour is changing, deterministic once it settles. See `INTENT.md` for the
thesis and `SCOPE.md` for boundaries. thesis and `SCOPE.md` for boundaries.
**Status:** research prototype. The deterministic kernel runs; agentic **Status:** research prototype. Deterministic claims, heuristic discovery,
realization, adaptation classification and crystallization are not built yet. adaptation classification and crystallization run locally. Three synthetic
Current work: `workplans/TD-WP-0002-vertical-spike-crystallization.md`. multi-user use cases and a 38-mutation catalogue exercise the model. Live-model
economics, browser-engine coverage and independent authoring-cost measurement
remain blocked in `workplans/TD-WP-0003-generalise-and-settle.md`.
See [readiness and interface decisions](docs/TestDriverGeneralisationReview.md).
## Run ## Run

View file

@ -10,7 +10,7 @@
| --- | --- | --- | --- | --- | | --- | --- | --- | --- | --- |
| workplan | TD-WP-0001 | finished | — | workplans/TD-WP-0001-statehub-bootstrap.md | | workplan | TD-WP-0001 | finished | — | workplans/TD-WP-0001-statehub-bootstrap.md |
| workplan | TD-WP-0002 | finished | — | workplans/TD-WP-0002-vertical-spike-crystallization.md | | workplan | TD-WP-0002 | finished | — | workplans/TD-WP-0002-vertical-spike-crystallization.md |
| workplan | TD-WP-0003 | proposed | — | workplans/TD-WP-0003-generalise-and-settle.md | | workplan | TD-WP-0003 | blocked | — | workplans/TD-WP-0003-generalise-and-settle.md |
| task | TD-WP-0001-T01 | done | — | workplans/TD-WP-0001-statehub-bootstrap.md | | task | TD-WP-0001-T01 | done | — | workplans/TD-WP-0001-statehub-bootstrap.md |
| task | TD-WP-0001-T02 | done | — | workplans/TD-WP-0001-statehub-bootstrap.md | | task | TD-WP-0001-T02 | done | — | workplans/TD-WP-0001-statehub-bootstrap.md |
| task | TD-WP-0001-T03 | done | — | workplans/TD-WP-0001-statehub-bootstrap.md | | task | TD-WP-0001-T03 | done | — | workplans/TD-WP-0001-statehub-bootstrap.md |
@ -24,11 +24,11 @@
| task | TD-WP-0002-T08 | done | — | workplans/TD-WP-0002-vertical-spike-crystallization.md | | task | TD-WP-0002-T08 | done | — | workplans/TD-WP-0002-vertical-spike-crystallization.md |
| task | TD-WP-0002-T09 | done | — | workplans/TD-WP-0002-vertical-spike-crystallization.md | | task | TD-WP-0002-T09 | done | — | workplans/TD-WP-0002-vertical-spike-crystallization.md |
| task | TD-WP-0002-T10 | done | — | workplans/TD-WP-0002-vertical-spike-crystallization.md | | task | TD-WP-0002-T10 | done | — | workplans/TD-WP-0002-vertical-spike-crystallization.md |
| task | TD-WP-0003-T01 | todo | — | workplans/TD-WP-0003-generalise-and-settle.md | | task | TD-WP-0003-T01 | wait | — | workplans/TD-WP-0003-generalise-and-settle.md |
| task | TD-WP-0003-T02 | todo | — | workplans/TD-WP-0003-generalise-and-settle.md | | task | TD-WP-0003-T02 | done | — | workplans/TD-WP-0003-generalise-and-settle.md |
| task | TD-WP-0003-T03 | todo | — | workplans/TD-WP-0003-generalise-and-settle.md | | task | TD-WP-0003-T03 | done | — | workplans/TD-WP-0003-generalise-and-settle.md |
| task | TD-WP-0003-T04 | todo | — | workplans/TD-WP-0003-generalise-and-settle.md | | task | TD-WP-0003-T04 | done | — | workplans/TD-WP-0003-generalise-and-settle.md |
| task | TD-WP-0003-T05 | todo | — | workplans/TD-WP-0003-generalise-and-settle.md | | task | TD-WP-0003-T05 | done | — | workplans/TD-WP-0003-generalise-and-settle.md |
| task | TD-WP-0003-T06 | wait | — | workplans/TD-WP-0003-generalise-and-settle.md | | task | TD-WP-0003-T06 | wait | — | workplans/TD-WP-0003-generalise-and-settle.md |
| task | TD-WP-0003-T07 | todo | — | workplans/TD-WP-0003-generalise-and-settle.md | | task | TD-WP-0003-T07 | wait | — | workplans/TD-WP-0003-generalise-and-settle.md |
| task | TD-WP-0003-T08 | todo | — | workplans/TD-WP-0003-generalise-and-settle.md | | task | TD-WP-0003-T08 | done | — | workplans/TD-WP-0003-generalise-and-settle.md |

View file

@ -1,5 +1,8 @@
# Test Driver Concept Model # Test Driver Concept Model
> Current scope, 2026-09-28: TD-WP-0003 removes speculative lifecycle scoring
> and declared Temperature. See [settlement and compatibility](TestDriverGeneralisationReview.md).
**Status:** v0.1 Draft **Status:** v0.1 Draft
**Framework:** `test-driver` **Framework:** `test-driver`
@ -58,17 +61,10 @@ Agentic tests should not remain agentic merely because they started that way.
As behavior and implementation become stable, agentic realizations are progressively hardened and finally crystallized into deterministic tests. As behavior and implementation become stable, agentic realizations are progressively hardened and finally crystallized into deterministic tests.
### 2.5 Useful tests accumulate energy ### 2.5–2.6 Evidence, not lifecycle scores
Verification assets gain energy when they demonstrate value, for example by catching genuine failures or protecting important behavior. Retain observations and verdicts. Asset selection and retirement remain explicit
maintainer choices; there is no energy score or automated retirement policy.
They lose energy when they repeatedly become invalid, require unnecessary adaptation, become flaky, duplicate stronger verification or protect behavior that is no longer relevant.
### 2.6 Verification assets may die
Tests are not immortal repository artifacts.
When a verification asset loses relevance and reaches sufficiently low energy, it may be removed from active campaigns, archived or retired.
### 2.7 Implementation is not the truth ### 2.7 Implementation is not the truth
@ -514,10 +510,6 @@ protects:
- least-privilege - least-privilege
maturity: adaptive maturity: adaptive
temperature: warm
energy:
value: 72
execution: execution:
mode: agentic mode: agentic
@ -585,29 +577,10 @@ U4 Contractual
A stable use case may temporarily require agentic tests when a new implementation is introduced. A stable use case may temporarily require agentic tests when a new implementation is introduced.
## 9.3 Temperature ## 9.3 Measured stability
**Temperature** represents implementation or capability fluidity. Declared Temperature was removed on 2026-09-28 (F-0008). Crystallization uses
observed trajectory stability through `assess_stability`, not a temperature label.
Initial model:
```text
HOT actively being invented
WARM frequently changing
COOL stabilizing
COLD mature or contractual
```
Temperature influences the preferred execution strategy:
| Temperature | Preferred verification mode |
|---|---|
| HOT | exploratory and agentic |
| WARM | adaptive with emerging deterministic coverage |
| COOL | hardened regression with occasional exploration |
| COLD | predominantly deterministic |
A capability may cool as it stabilizes and heat up again during major redesign or migration.
## 9.4 Adaptation ## 9.4 Adaptation
@ -665,53 +638,11 @@ Explore
Crystallization is not necessarily permanent. A major implementation change may temporarily move an asset back toward adaptive or agentic execution. Crystallization is not necessarily permanent. A major implementation change may temporarily move an asset back toward adaptive or agentic execution.
## 9.7 Energy ## 9.7–9.8 Removed lifecycle scores
**Energy** expresses the current value of retaining, maintaining and executing a verification asset. Energy and Confidence are removed from the current concept set. Neither was
used in a decision across the three reference use cases. H-005 is
Energy should be derived from observable events rather than being an unexplained score. dormant-indefinite; there is no capture-only energy subsystem.
Illustrative initial event model:
| Event | Energy change |
|---|---:|
| Detects confirmed product defect | +25 |
| Detects confirmed security defect | +40 |
| Prevents regression after previous defect | +20 |
| Exercises recently changed relevant capability | +2 |
| Requires harmless mechanical adaptation | -2 |
| Requires workflow adaptation | -10 |
| Test itself was wrong | -20 |
| False positive or flaky failure | -15 |
| Duplicates stronger existing verification | -20 |
| Protected use case is deprecated | -100 |
Energy history should be retained:
```yaml
energy:
value: 83
history:
- event: defect-detected
delta: 25
finding: TD-143
- event: navigation-adaptation
delta: -2
```
Energy is not synonymous with correctness.
It is better interpreted as:
> the current value of spending verification attention and resources on this asset.
## 9.8 Confidence
**Confidence** represents how strongly the framework currently believes that a verification asset faithfully represents the intended behavior it claims to protect.
Confidence should remain conceptually separate from energy.
A high-energy test may still have low confidence if its behavior is poorly specified.
## 9.9 Lineage ## 9.9 Lineage
@ -735,18 +666,10 @@ Lineage should make it possible to answer:
> Why does this test exist? > Why does this test exist?
## 9.10 Retirement ## 9.10 Asset removal
A verification asset may move through: Maintainers may remove obsolete tests with an explicit review of the protected
claims. No computed Retirement concept or policy is implemented or scheduled.
```text
active
-> low-frequency
-> archival
-> retired
```
Energy reaching zero may trigger retirement consideration, but critical security, regulatory or contractual invariants may define a retirement floor or prohibition.
--- ---
@ -770,51 +693,11 @@ When one side changes, the framework evaluates whether the other relationships r
--- ---
# 11. Test Metabolism # 11–12. Removed planning abstractions
Energy, temperature, maturity and execution cost together create a **Test Metabolism**. Metabolism and Campaign were removed from the current concept set on 2026-09-28.
The experiments choose explicit scenario and mutation lists. No evidence justifies
The framework should preferentially spend verification resources where they are most valuable. a planner or score-driven scheduling; these are not deferred implementation promises.
A future campaign planner may consider:
```text
Priority =
Test Energy
x Changed-System Proximity
x Use-Case Criticality
x Risk
x Time Since Last Execution
```
The precise formula is intentionally deferred.
The conceptual requirement is that test selection should be dynamic rather than assuming every historical test has equal present value.
---
# 12. Campaign
A **Campaign** is a strategy for selecting and executing verification assets and scenario variants.
Initial campaign examples:
```text
PR smoke
regression
release qualification
authorization
tenant isolation
concurrency
resilience
exploratory
```
Campaigns manage combinatorial explosion by selecting relevant slices of:
```text
actors x roles x states x data x surfaces x schedules x mutations
```
--- ---
@ -898,9 +781,6 @@ Observers --> Evidence --> Oracle Engine --> Verdict
Metadata Store: Metadata Store:
maturity maturity
temperature
energy
confidence
lineage lineage
``` ```
@ -955,17 +835,11 @@ Finding
Investigator Investigator
VerificationAsset VerificationAsset
Maturity Maturity
Temperature
Adaptation Adaptation
Hardening Hardening
Crystallization Crystallization
Energy
Confidence
Lineage Lineage
Retirement
Campaign
TriadicVerification TriadicVerification
TestMetabolism
``` ```
--- ---
@ -1003,15 +877,9 @@ The defining model of test-driver is:
v v
Verdict Verdict
| |
Finding / Value Finding
|
v
Energy
HOT ------> WARM ------> COOL ------> COLD Repeated stable trajectories --> crystallization --> deterministic regression
| | | |
agentic adaptive hardened deterministic
+-------------- crystallization ------>
``` ```
The fundamental promise is: The fundamental promise is:

View file

@ -0,0 +1,146 @@
# TD-WP-0003 generalisation and settlement — 2026-09-28
The kernel now expresses sharing/revocation, delegated approval and tenant
lifecycle without additional framework concepts. This is evidence of local
expressiveness, not readiness for autonomous verification of a real system.
## New use cases (T02)
The agent wrote [synthetic contracts](../usecases/generalisation-spec.md) before
the implementations. Claims are `agent-from-spec`, not human-authored. The same
agent wrote the synthetic requirements and labs; this does not establish
independence from a real product team or requirements quality.
- `scenarios/delegated_approval.py`: premature execution, submission, delegation,
rejected self-approval, rejected former-reviewer approval, delegate approval,
requester execution. Three independently enabled defects must fail.
- `scenarios/tenant_lifecycle.py`: two administrators use the same local id;
foreign delete attempts fail; deletion and recreation preserve the other
tenant. Four independently enabled defects must fail.
`tests/test_generalisation.py` verifies baselines, each defect, missing-evidence
INCONCLUSIVE, deterministic replay and verdict reproduction from serialized S3.
Separate observers read state and probe enforcement; drivers only emit S1.
No claims or invariants are learned from the lab's responses.
Ordered `Step`s already express the required Schedule. `Scenario.variant`
identifies each lab fault. A separate scheduler, Lens or Variant object was not
needed. There is no concurrent scheduling or time-window evidence here.
The resistance was in the adapters, not new concepts: `StateObserver`'s standard
snapshot is specific to resource grants, so each domain needs its own collector
implementing `snapshot()` and `name`. Expected refusal must be a completed
protocol response with an independently observed receipt; setting
`Realization.raised` would suppress dependent claims as INCONCLUSIVE. Unexpected
exceptions still abort these fixture drivers; they are not silently swallowed.
## Structural experiment (T03)
[Machine-readable results](../research/evidence/2026-09-28-e001.json) retain each
arm's surface, judgments, metrics and the journey classification/signals.
Reproduce with:
```bash
PYTHONPATH=src:. python3 tools/measure_e001.py /tmp/e001.json
```
| Identifier axis | Discovery | Recorded selectors |
|---|---:|---:|
| Preserved | 17/17 | 17/17 |
| Dropped | 9/10 | 0/10 |
M25–M31 add reordered fields, dialog relocation, anchor affordances, fieldset
grouping, textarea input, an unrelated preceding form, and sidebar relocation.
M32–M38 pair those mechanisms with preserved identifiers. Existing M02/M21/M22
supply rewritten DOM, reworded controls and renamed fields. New labels were fixed
from the synthetic contract before running and are agent-authored, not claimed
as newly human-reviewed ground truth.
The 38-entry catalogue contains 27 mechanical, four semantic and seven defect
mutations. Mechanical acceptance is 26/27. M22 still defeats the heuristic.
False adaptation is 0/7 on the existing defect set; adding mechanical variants
does not enlarge that safety denominator. No claim/invariant index changed.
M25 and M32 are UNCHANGED because the coarse surface fingerprint ignores field
order and test ids; this is not a claim that their HTML is identical.
These are selected mechanisms, not random independent samples. Paired variants
are correlated; 9/10 is not a population reliability estimate. The control is a
stable-test-id replay, not every possible conventional selector strategy. Both
arms remain token-free and test structural HTML only. H-001 stays narrowed to
identifier-loss cases, and F-0004/F-0005/F-0007's external validity limits remain.
## Concept settlement and compatibility (T04/T05)
**Keep claims as Python predicates.** Across all three runnable use cases the
repeated shapes are equality, membership and comparisons between enforcement
and state. The audit-core consumer additionally compares sanitized absence
surfaces, bounded execution, cleanup and receipt binding. A small vocabulary
would either duplicate Python or need frequent escape hatches; serialization
would create a second language and a second statement of intent.
This publishes the interface decision requested by T05: the existing
`Claim(id, text, provenance, predicate, after_step, source_ref=None)` and
`Invariant(id, text, provenance, predicate, source_ref=None)` constructors stay
compatible. Predicates receive independent observation mappings and return
booleans. Required evidence must use explicit lookup, not a pass-producing
fallback. The oracle converts missing keys or predicate exceptions to
INCONCLUSIVE. Predicates must not reach into actors or the live SUT.
Generated regression artifacts continue to import their original scenario
claims; consumers must ship those modules with the artifact. Fully standalone
claim serialization is deliberately rejected. The audit-core consumer needs no
migration; its tests remain in the default suite.
**Remove unused lifecycle concepts now.** Temperature, Energy, Confidence,
Campaign, Metabolism and automated Retirement have informed no decision in
these three use cases. They leave the current model, not another deferred gate.
Trajectory stability continues to drive crystallization. H-005 becomes
`DORMANT-INDEFINITE`; F-0008 is resolved.
This is a research-prototype compatibility change: `testdriver.energy`, public
`EnergyEvent`/`EnergyEventType` exports and the `EvidencePack.energy_events`
field are removed. Consumers must stop importing or requiring them. Historical
JSON may retain that field; the current classifier already ignores it.
Execution identity, timestamps, judgments, S1 realization costs and lineage
remain; none of these imply a computed lifecycle score. There is no dependency
addition. Historical concept proposals carry a supersession notice.
## Authoring measurement (T07 remains waiting)
[Timing receipt](../research/evidence/2026-09-28-authoring.json) preserves the
actual measurements and line counts. The timer was placed inside the shell
write operation, after code composition. The approval interval also included
composition of the next case, and the tenant interval largely measured file
writes and pytest. These intervals are **not valid authoring-cost measurements**
and must not be compared as productivity numbers. Human maintenance effort and
model latency/cost were not captured. The first attempt cannot be reconstructed.
The observable concepts consulted and friction are recorded above. T07 retains
the remaining work: an independent author must implement these contracts afresh
with a timer started before reading/design, stopping after first passing defect
checks, separately recording tooling wait and review. Do not count a rerun or
copy of these implementations as a new authoring sample. No replacement task
or workplan is created.
## Real-system readiness (T08)
**Verdict: not ready for autonomous real-system application.** Local research and
attended review of prepared claims can continue. Completing the review does not
satisfy the workplan's live-model success gate or authorize a production run.
| Required criterion | Current evidence | Assessment |
|---|---|---|
| Independent observation channel can read records and probe enforcement without trusting actors | Three labs; audit-core declares an independent observer but has no runnable driver/collector here | Not met outside synthetic labs; adapter construction, permissions and operating cost remain unmeasured |
| Claims trace to independently approved intent predating the tested behavior | Local provenance guards; audit-core cites approved engagement and workplan, while its precedent receipt is calibration only | Structurally supported, operational chain must be checked for the exact next engagement; a ticket written after code is insufficient without human promotion |
| Legitimate behavioral ambiguity is escalated without rewriting intent | M12/M19 are behaviorally indistinguishable; AMBIGUOUS and BEHAVIOUR_CHANGED require review | Demonstrated in lab, real backlog dispute resolution is still attended; missing evidence stays INCONCLUSIVE |
| False-adaptation consequences are identified for the engagement | Audit-core tenant disclosure, forged foreign append or leftover identity could be concealed by a false pass | Zero tolerance for automatic acceptance of regression; 0/7 synthetic defects is no guarantee against unseen attacks |
| Execution, custody, timing, cleanup and target revision are enforced | Audit-core's PhaseContract is durable intent, explicitly not a runnable Scenario; tests use supplied observations | Not met: no causal/time-window runtime, custody or cluster driver, independent cleanup collector or fresh admitted window here |
| Capability, economics and maintenance costs justify adoption | Token-free structural discovery, deterministic crystallization and invalid first authoring timer | Not met; T01/T06/T07 retain the unanswered measurements |
The audit-core material referenced here is the checked-in
`usecases/audit_core_e2_tenant_boundary.py` contract and its fixture-based tests,
not a fresh inspection of any deployment or a revalidation of the 2026-08-22
engagement. A real run needs the exact engagement's target-owner acceptance,
independent observation/cleanup receipts and appropriate approved access. A
successful historical receipt cannot substitute for those. The existing T01,
T06 and T07 tasks retain the experiment blockers; this review creates no
production implementation commitment or new work record.

View file

@ -1,5 +1,9 @@
# TestDriver Improvement Loop # TestDriver Improvement Loop
> Scope update (2026-09-28): lifecycle Energy, Temperature, Confidence, Campaign,
> Metabolism and automated Retirement proposals below are historical, superseded
> by the [TD-WP-0003 settlement](TestDriverGeneralisationReview.md).
**Status:** Concept v0.1 **Status:** Concept v0.1
**Purpose:** Establish a self-improvement loop that keeps the implementation of `test-driver` aligned with its conceptual model while generating evidence about which concepts actually work. **Purpose:** Establish a self-improvement loop that keeps the implementation of `test-driver` aligned with its conceptual model while generating evidence about which concepts actually work.

View file

@ -1,5 +1,9 @@
# TestDriver Research Prototype — Initial Milestones # TestDriver Research Prototype — Initial Milestones
> Scope update (2026-09-28): lifecycle Energy, Temperature, Confidence, Campaign,
> Metabolism and automated Retirement proposals below are historical, superseded
> by the [TD-WP-0003 settlement](TestDriverGeneralisationReview.md).
**Status:** v0.1 — **canonical milestone sequence** **Status:** v0.1 — **canonical milestone sequence**
**Purpose:** Establish the minimum evidence-producing development loop needed to validate the core test-driver concepts. **Purpose:** Establish the minimum evidence-producing development loop needed to validate the core test-driver concepts.

View file

@ -1,5 +1,9 @@
# Stage 1 Test Driver Validation # Stage 1 Test Driver Validation
> Scope update (2026-09-28): lifecycle Energy, Temperature, Confidence, Campaign,
> Metabolism and automated Retirement proposals below are historical, superseded
> by the [TD-WP-0003 settlement](TestDriverGeneralisationReview.md).
The biggest risk is not technical feasibility. It is that **test-driver becomes conceptually elegant but too broad to prove itself quickly**. The biggest risk is not technical feasibility. It is that **test-driver becomes conceptually elegant but too broad to prove itself quickly**.
I would improve the odds of success by treating the next phase as an experiment in whether three specific ideas actually work: I would improve the odds of success by treating the next phase as an experiment in whether three specific ideas actually work:

View file

@ -83,3 +83,13 @@ Measured in `tests/test_classification.py`. **False Adaptation Rate = 0/7.**
The matrix is asserted in `tests/test_lab_ground_truth.py`. A moved cell fails The matrix is asserted in `tests/test_lab_ground_truth.py`. A moved cell fails
the suite: changing the measuring instrument must be a deliberate, reviewed act. the suite: changing the measuring instrument must be a deliberate, reviewed act.
## T03 extension — 2026-09-28
M25–M31 drop identifiers while changing field order, dialog placement, anchor
affordances, fieldset grouping, textarea input, unrelated-form competition and
sidebar placement. M32–M38 pair the same mechanisms with identifiers preserved.
All fourteen are mechanical by the pre-run synthetic contract: grant behavior
and the endpoint are unchanged. These additions are agent-authored instrument
labels, not a claim of new human review. Both arms and the full classifier are
measured in `research/evidence/2026-09-28-e001.json`.

View file

@ -98,6 +98,7 @@ class LabApp:
ui_test_ids: str = "stable" ui_test_ids: str = "stable"
ui_field_names: str = "canonical" ui_field_names: str = "canonical"
ui_confirm_revoke: bool = False ui_confirm_revoke: bool = False
ui_form_layout: str = "plain"
api_path_style: str = "long" api_path_style: str = "long"
api_grant_path: str = "grant" api_grant_path: str = "grant"
tenant_sharing_announced: bool = False tenant_sharing_announced: bool = False

View file

@ -66,6 +66,14 @@ def _render(app: LabApp, user_id: str, resource_id: str) -> str:
else subject_field + permission_field else subject_field + permission_field
) )
if app.ui_form_layout == "fieldset":
fields = f'<fieldset><legend>Access</legend>{fields}</fieldset>'
elif app.ui_form_layout == "textarea":
fields = fields.replace(
f'<input id="subject" name="{subject_name}" data-td="subject">',
f'<textarea id="subject" name="{subject_name}" data-td="subject"></textarea>',
)
share_form = ( share_form = (
f'<form id="share" data-td="share-form" method="post" ' f'<form id="share" data-td="share-form" method="post" '
f'action="{_resource_path(app, resource_id)}/{app.api_grant_path}">' f'action="{_resource_path(app, resource_id)}/{app.api_grant_path}">'
@ -73,6 +81,11 @@ def _render(app: LabApp, user_id: str, resource_id: str) -> str:
f'<{tag}{href}{role} id="share-submit" data-td="share-submit">' f'<{tag}{href}{role} id="share-submit" data-td="share-submit">'
f"{share_label}</{tag}></form>" f"{share_label}</{tag}></form>"
) )
if app.ui_form_layout == "decoy":
share_form = (
'<form method="post" action="/search"><input name="query">'
'<button>Search resources</button></form>' + share_form
)
if app.ui_share_control == "removed": if app.ui_share_control == "removed":
# The control is gone from the UI. The API endpoint is still open — the # The control is gone from the UI. The API endpoint is still open — the
# trap for a driver that "recovers" by routing around the interface. # trap for a driver that "recovers" by routing around the interface.
@ -107,6 +120,11 @@ def _render(app: LabApp, user_id: str, resource_id: str) -> str:
f'<h1 data-td="resource-title">Resource {resource_id}</h1>' f'<h1 data-td="resource-title">Resource {resource_id}</h1>'
f'<ul data-td="grant-list">{grants}</ul>{share_form}{revoke_form}' f'<ul data-td="grant-list">{grants}</ul>{share_form}{revoke_form}'
) )
if app.ui_form_layout == "sidebar":
body = (
f'<main><h1>Resource {resource_id}</h1><ul>{grants}</ul>{revoke_form}</main>'
f'<aside aria-label="Access management">{share_form}</aside>'
)
if app.ui_dom_style == "nested": if app.ui_dom_style == "nested":
body = ( body = (
'<div class="shell"><section class="panel"><div class="panel-inner">' '<div class="shell"><section class="panel"><div class="panel-inner">'

View file

@ -1,12 +1,13 @@
"""The labelled mutation catalogue — the project's measuring instrument. """The labelled mutation catalogue — the project's measuring instrument.
Every claim test-driver makes is measured against this catalogue, so its quality The structural adaptation claims are measured against this catalogue, so its quality
caps the credibility of every downstream result. Six mutations, as the milestones caps the credibility of every downstream result. Six mutations, as the milestones
document originally sketched, cannot support any statement about precision or document originally sketched, cannot support any statement about precision or
recall; there are twenty here. recall; there are 38 here after TD-WP-0003-T03.
Each mutation carries a **ground-truth label**, decided by a human from the use Each mutation carries a **ground-truth label** recorded before its measurement.
case and recorded before any run: The original labels came with the use case; M25–M38 are explicitly agent-authored
synthetic extensions (not newly human-reviewed):
MECHANICAL the surface changed; protected semantics are identical. MECHANICAL the surface changed; protected semantics are identical.
test-driver should recover and report an adaptation. test-driver should recover and report an adaptation.
@ -216,11 +217,37 @@ CATALOGUE: tuple[Mutation, ...] = (
"A user of another tenant reads the resource.", _m20), "A user of another tenant reads the resource.", _m20),
) )
# T03 extension: labels fixed from the synthetic contract before execution.
# These are agent-authored instrument variants, not newly human-reviewed labels.
# Paired variants isolate identifier loss from the structural mechanism.
EXTENDED_LAYOUTS = (
("M25", "M32", "Reordered fields", {"ui_field_order": "reversed"}),
("M26", "M33", "Relocated into an open dialog", {"ui_share_control": "modal"}),
("M27", "M34", "Anchor affordances", {"ui_button_element": "anchor"}),
("M28", "M35", "Fields grouped in a fieldset", {"ui_form_layout": "fieldset"}),
("M29", "M36", "Subject input becomes a textarea", {"ui_form_layout": "textarea"}),
("M30", "M37", "Unrelated form precedes the control", {"ui_form_layout": "decoy"}),
("M31", "M38", "Control moved after revoke into a sidebar", {"ui_form_layout": "sidebar"}),
)
for dropped_id, preserved_id, title, flags in EXTENDED_LAYOUTS:
for mutation_id, preserved in ((dropped_id, False), (preserved_id, True)):
CATALOGUE += (Mutation(
mutation_id, title + ("; ids preserved" if preserved else "; ids dropped"),
"MECHANICAL", "ui",
"Agent-authored synthetic variant: the same explicit grant form and "
"endpoint remain available; stored permissions and enforcement are unchanged.",
lambda app, flags=flags, preserved=preserved: _m(
app, **flags, ui_test_ids="stable" if preserved else "dropped"
),
preserves_test_ids=preserved,
),)
BY_ID: dict[str, Mutation] = {m.id: m for m in CATALOGUE} BY_ID: dict[str, Mutation] = {m.id: m for m in CATALOGUE}
def expected_classification(mutation_id: str) -> Label: def expected_classification(mutation_id: str) -> Label:
"""Ground truth. Recorded by a human before any run — never inferred.""" """Pre-recorded instrument labels, never inferred from run outcomes."""
return BY_ID[mutation_id].label return BY_ID[mutation_id].label

134
lab/workflows.py Normal file
View file

@ -0,0 +1,134 @@
"""Synthetic workflow lab, implemented from usecases/generalisation-spec.md."""
from copy import deepcopy
from dataclasses import dataclass
from testdriver.actions import Surface
from testdriver.drivers import Realization
class WorkflowDriver:
"""A completed denial is a protocol response, not a failed realization."""
surface = Surface("workflow", "direct")
def __init__(self, lab):
self.lab = lab
def realize(self, actor, action):
action.check_surface(self.surface.id)
response = self.lab.request(actor.credentials["token"], action.name, **action.args)
return Realization(self.surface.id, {"operation": action.name, "response": response})
class ApprovalLab:
tokens = {"requester": "fixture-requester", "reviewer": "fixture-reviewer",
"delegate": "fixture-delegate"}
def __init__(self, defect=None):
self.defect = defect
self.state = "draft"
self.reviewer = "reviewer"
self.approved_by = None
self.receipts = []
def may_execute(self, actor):
return actor == "requester" and (
self.state == "approved" or self.defect == "premature-execution"
)
def request(self, token, operation, **args):
actor = next((a for a, t in self.tokens.items() if t == token), None)
allowed = False
if operation == "submit":
allowed = actor == "requester" and self.state == "draft"
if allowed:
self.state = "pending"
elif operation == "delegate":
allowed = (actor == self.reviewer and self.state == "pending"
and args["to"] == "delegate")
if allowed:
self.reviewer = args["to"]
elif operation == "approve":
allowed = self.state == "pending" and (
actor == self.reviewer
or (actor == "requester" and self.defect == "self-approval")
or (actor == "reviewer" and self.defect == "stale-delegation")
)
if allowed:
self.state, self.approved_by = "approved", actor
elif operation == "execute":
allowed = self.may_execute(actor)
if allowed:
self.state = "executed"
else:
raise ValueError(operation)
receipt = {"actor": actor, "operation": operation, "allowed": allowed}
self.receipts.append(receipt)
return deepcopy(receipt)
@dataclass
class ApprovalObserver:
lab: ApprovalLab
name: str = "approval-observer"
def snapshot(self):
return deepcopy({
"state": self.lab.state,
"reviewer": self.lab.reviewer,
"approved_by": self.lab.approved_by,
"receipts": self.lab.receipts,
"execution_gate": {actor: self.lab.may_execute(actor) for actor in self.lab.tokens},
})
class TenantLab:
tokens = {"admin-a": "fixture-admin-a", "admin-b": "fixture-admin-b"}
tenants = {"admin-a": "A", "admin-b": "B"}
def __init__(self, defect=None):
self.defect = defect
self.records = {}
self.deleted = {}
self.receipts = []
def read(self, actor, tenant, resource):
if self.tenants.get(actor) != tenant and self.defect != "foreign-read":
return {"allowed": False, "content": None}
return {"allowed": (tenant, resource) in self.records,
"content": self.records.get((tenant, resource))}
def request(self, token, operation, tenant, resource="R", content=None):
actor = next((a for a, t in self.tokens.items() if t == token), None)
allowed = self.tenants.get(actor) == tenant
key = (tenant, resource)
if operation == "create":
allowed = allowed and key not in self.records
if allowed:
self.records[key] = (self.deleted.get(key, content)
if self.defect == "resurrect-content" else content)
elif operation == "delete":
allowed = (allowed or self.defect == "foreign-delete") and key in self.records
if allowed:
self.deleted[key] = self.records.pop(key)
if self.defect == "unscoped-delete":
self.records = {k: v for k, v in self.records.items() if k[1] != resource}
else:
raise ValueError(operation)
receipt = {"actor": actor, "operation": operation, "tenant": tenant, "allowed": allowed}
self.receipts.append(receipt)
return deepcopy(receipt)
@dataclass
class TenantObserver:
lab: TenantLab
name: str = "tenant-observer"
def snapshot(self):
return deepcopy({
"records": {tenant: self.lab.records.get((tenant, "R")) for tenant in ("A", "B")},
"reads": {f"{actor}:{tenant}": self.lab.read(actor, tenant, "R")
for actor in self.lab.tokens for tenant in ("A", "B")},
"receipts": self.lab.receipts,
})

View file

@ -1,6 +1,6 @@
# Concept ↔ Implementation Fitness Map # Concept ↔ Implementation Fitness Map
**Updated:** 2026-08-23 (TD-WP-0002-T10) **Updated:** 2026-09-28 (TD-WP-0003-T02/T03/T04)
Traces each important concept to the implementation, experiment and evidence that Traces each important concept to the implementation, experiment and evidence that
support it. **Unsupported entries are the point of this map** — a concept with no support it. **Unsupported entries are the point of this map** — a concept with no
@ -28,43 +28,39 @@ were aspirational, not evidenced.
| Concept | Level | Implementation | Experiment | Evidence | Open question | | Concept | Level | Implementation | Experiment | Evidence | Open question |
|---|---|---|---|---|---| |---|---|---|---|---|---|
| `C-use-case` | C1 | `intent.py` | — | — | Is a use case expressible without leaking mechanics? | | `C-use-case` | C1 | `intent.py`, three reference scenarios | T02 | `tests/test_generalisation.py` | Three synthetic use cases fit; external authoring still unmeasured. |
| `C-actor-isolation` | **C2** | `world.py`, `runner.py` | E-001 | `td://self/actor-isolation`, examined every run | **F-0003 resolved** — automatic canaries; observable on every scenario. | | `C-actor-isolation` | **C2** | `world.py`, `runner.py` | E-001 | `td://self/actor-isolation`, examined every run | **F-0003 resolved** — automatic canaries; observable on every scenario. |
| `C-semantic-action` | **C2** | `actions.py`, `agentic.py` | E-001 (partial) | T07 arm comparison | **F-0005** — supported only where stable identifiers are absent. Narrower than the concept model claims. | | `C-semantic-action` | **C2** | `actions.py`, `agentic.py` | E-001 | 2026-09-28 two-arm receipt | **F-0005** — supported only where stable identifiers are absent. Narrower than the concept model claims. |
| `C-oracle-independence` | **C2** | `runner.py`, `oracles.py` | E-001, E-003 | — | Independence of components ≠ independence of belief. (H-004) | | `C-oracle-independence` | **C2** | `runner.py`, `oracles.py` | E-001, E-003 | — | Independence of components ≠ independence of belief. (H-004) |
| `C-evidence-pack` | C1 | `evidence.py` | — | `td://self/evidence-reproducibility` | Verdicts are reproducible from S3 alone, on passing and failing runs. | | `C-evidence-pack` | C1 | `evidence.py` | — | `td://self/evidence-reproducibility` | Verdicts are reproducible from S3 alone, on passing and failing runs. |
| `C-observation-channel` | C1 | `lab/app.py` | — | — | **D-07** — required of every system under test. Adoption cost unknown. | | `C-observation-channel` | C1 | `lab/app.py` | — | — | **D-07** — required of every system under test. Adoption cost unknown. |
| `C-adaptation` | **C2** | `classification.py`, `agentic.py` | E-001 | T08 matrix | 11/12 mechanical absorbed without a human. | | `C-adaptation` | **C2** | `classification.py`, `agentic.py` | E-001 | T08 matrix | 26/27 mechanical absorbed; ten dropped-id cases now included. |
| `C-classification` | **C2** | `classification.py` | E-001, E-003 | T08 matrix, FAR 0/7 | **F-0006** — cannot infer SEMANTIC vs DEFECT; collapses to one escalating outcome. | | `C-classification` | **C2** | `classification.py` | E-001, E-003 | T08 matrix, FAR 0/7 | **F-0006** — cannot infer SEMANTIC vs DEFECT; collapses to one escalating outcome. |
| `C-crystallization` | **C2** | `crystallization.py` | E-002 | T09 agreement matrix | Fidelity supported; **F-0007** — the economic case is unmeasurable with a token-free runtime. | | `C-crystallization` | **C2** | `crystallization.py` | E-002 | T09 agreement matrix | Fidelity supported; **F-0007** — the economic case is unmeasurable with a token-free runtime. |
| `C-intent-provenance` | C1 | `provenance.py` | E-003 | `td://self/intent-independence` | Constrains provenance, not quality. Accepted residual. | | `C-intent-provenance` | C1 | `provenance.py` | E-003 | `td://self/intent-independence` | Constrains provenance, not quality. Accepted residual. |
| `C-lineage` | C1 | `scenario.py`, generated headers | E-002 | generated module docstring | Parent pointer plus provenance in the artefact; no graph. | | `C-lineage` | C1 | `scenario.py`, generated headers | E-002 | generated module docstring | Parent pointer plus provenance in the artefact; no graph. |
| `C-energy` | C0 | `energy.py`, capture only | — | none | Dormant. **Gated**: if the next workplan ends with no decision having used the history, remove. | | `C-energy` | Removed | — | T04 | Generalisation review | No demonstrated decision; not an implementation promise. |
| `C-temperature` | **C0 — under review** | — | — | none | **F-0008** — crystallization was built without it; measured stability did the work. Gated for removal. | | `C-temperature` | Removed | — | T04 | Generalisation review | No demonstrated decision; not an implementation promise. |
| `C-confidence` | C0 | — | — | none | Declared, unimplemented, never consulted. Same position as `C-temperature` without a built alternative. | | `C-confidence` | Removed | — | T04 | Generalisation review | No demonstrated decision; not an implementation promise. |
| `C-campaign` | C0 | — | — | — | Deferred. | | `C-campaign` | Removed | — | T04 | Generalisation review | No demonstrated decision; not an implementation promise. |
| `C-metabolism` | C0 | — | — | — | Deferred. Depends on C-energy. | | `C-metabolism` | Removed | — | T04 | Generalisation review | No demonstrated decision; not an implementation promise. |
| `C-retirement` | C0 | — | — | — | Deferred. Depends on C-energy. | | `C-retirement` | Removed | — | T04 | Generalisation review | No demonstrated decision; not an implementation promise. |
| `C-security-mutation` | C1 | `lab/mutations.py` | E-003 | `lab/GROUND-TRUTH.md` | Catalogue is hand-written; no derivation mechanism from use cases yet. | | `C-security-mutation` | C1 | `lab/mutations.py` | E-003 | `lab/GROUND-TRUTH.md` | Catalogue is hand-written; no derivation mechanism from use cases yet. |
## Orphan check ## Generalisation and compression review — 2026-09-28
**Conceptual orphans** — concepts with no planned implementation in `TD-WP-0002`: No new kernel concept was required for delegation/sequencing or tenant lifecycle.
`C-temperature`, `C-confidence`, `C-campaign`, `C-metabolism`, `C-retirement`. Ordered `Step`s supply the Schedule; scenario `variant` identifies seeded defects.
No independent scheduler or Variant class was needed. Integration/security lenses
remain descriptive groupings; no Lens runtime object was consulted or validated.
The generic observer interface works, but each new domain needs a custom snapshot
collector: observation-channel adoption cost remains a real concern.
All five are deferred *by explicit decision*, not oversight. They are the group Temperature, Energy, Confidence, Campaign, Metabolism and Retirement are removed,
most at risk of being built because they are easy and satisfying, and never not deferred to another gate. `energy.py`, its exports and evidence field are
validated. They are revisited at T10, where the question is not "when do we build removed; H-005 is dormant-indefinite and F-0008 is resolved. Ordinary evidence,
these" but "does the evidence justify keeping them in the model at all". judgments, realization metrics and lineage remain. See
[review](../../docs/TestDriverGeneralisationReview.md) for compatibility and limits.
**Implementation orphans** — none. Every module in `src/testdriver/` traces to a Previous compression (T10): `Verdict.SUSPICIOUS`, `Step.expect_refusal`,
concept above. `ActorIsolationError`, `World.seed`, `EvidencePack.latest()`, `Trajectory.method`.
**Removed at T10** (compression pass — see the gate review § 3):
`Verdict.SUSPICIOUS`, `Step.expect_refusal`, `ActorIsolationError`,
`World.seed`, `EvidencePack.latest()`, `Trajectory.method`. Each was declared and
never used; `SUSPICIOUS` was additionally a verdict no oracle could emit.
**Gated for removal**: `C-temperature` (F-0008) and `energy.py`. Both survive the
current review on the strength of being cheap, not of being used. If the next
workplan closes without a decision consulting either, they go.

View file

@ -7,3 +7,4 @@ file is a pointer table, not a second source of truth.
| Ref | Title | Hub ID | Source | | Ref | Title | Hub ID | Source |
|---|---|---|---| |---|---|---|---|
| D-01 … D-07 | Stratified evidence and claim provenance as the basis for adaptation safety | `fef5213f-ce9b-44c2-b327-a0b0ba4b6270` | `docs/TestDriverClassificationDesign.md` | | D-01 … D-07 | Stratified evidence and claim provenance as the basis for adaptation safety | `fef5213f-ce9b-44c2-b327-a0b0ba4b6270` | `docs/TestDriverClassificationDesign.md` |
| TD-WP-0003 T04/T05/T08 | Python claims, concept removal and not-ready assessment | `d80734d1-01b8-4a98-8e72-c86f1806d572` | `docs/TestDriverGeneralisationReview.md` |

View file

@ -0,0 +1,19 @@
{
"approval": {
"started": "2026-09-28T09:56:40.697848+00:00",
"first_passing_tests": "2026-09-28T09:57:56.524430+00:00",
"scenario_lines": 73,
"scenario": "scenarios/delegated_approval.py",
"limitation": "Timer started after code composition; interval is not authoring time."
},
"tenant": {
"started": "2026-09-28T09:57:56.524430+00:00",
"first_passing_tests": "2026-09-28T09:57:57.123098+00:00",
"scenario_lines": 70,
"scenario": "scenarios/tenant_lifecycle.py",
"limitation": "Timer started after code composition; interval is not authoring time."
},
"valid_authoring_measurement": false,
"shared_adapter_and_lab_lines": 134,
"shared_test_lines": 80
}

File diff suppressed because it is too large Load diff

View file

@ -1,7 +1,7 @@
--- ---
id: E-001 id: E-001
title: Mechanical recovery and defect discrimination over the labelled mutation set title: Mechanical recovery and defect discrimination over the labelled mutation set
status: PLANNED status: EXECUTED
hypotheses: [H-001, H-002, H-004] hypotheses: [H-001, H-002, H-004]
task: TD-WP-0002-T08 task: TD-WP-0002-T08
created: "2026-08-22" created: "2026-08-22"
@ -37,3 +37,18 @@ tokens · wall time · retries.
`EXECUTED` 2026-08-22 (T08). Arm A run over 23 mutations; arm B over `EXECUTED` 2026-08-22 (T08). Arm A run over 23 mutations; arm B over
the 12 mechanical ones. Results in H-001 and H-004. FAR 0/7. the 12 mechanical ones. Results in H-001 and H-004. FAR 0/7.
## Expanded measurement — 2026-09-28 (TD-WP-0003-T03)
38 mutations, including ten that drop test ids and seven new paired structural
controls. Discovery and recorded selectors both pass 17/17 preserved-id cases;
on dropped ids discovery passes 9/10 and recorded selectors 0/10. M22 remains
unrealizable. The full journey accepts 26/27 mechanical cases; FAR remains 0/7
on the unchanged defect set. Claim/provenance indices remain unchanged.
Reproduce: `PYTHONPATH=src:. python3 tools/measure_e001.py /tmp/e001.json`.
[Raw result](../evidence/2026-09-28-e001.json).
These selected and correlated synthetic HTML cases do not establish population
reliability, visual/browser coverage, model capability or economics. See
[review](../../docs/TestDriverGeneralisationReview.md) for the mechanism list and
limits. H-001 remains supported only in the narrowed identifier-loss setting.

View file

@ -85,3 +85,10 @@ Fully standalone generation would need claims expressible in a serializable form
rather than as Python predicates. That is a real design question — it is the same rather than as Python predicates. That is a real design question — it is the same
question as "should scenarios be YAML", deferred at T04 — and both should be question as "should scenarios be YAML", deferred at T04 — and both should be
answered together at T10, with evidence about which predicates actually recur. answered together at T10, with evidence about which predicates actually recur.
## Interface settlement — 2026-09-28
TD-WP-0003-T05 deliberately retains Python predicates and imported original
claims. Standalone serialization is not a promised deliverable. See
`docs/TestDriverGeneralisationReview.md` for the published contract. The economic
finding stays open under TD-WP-0003-T01; no live-model measurement was performed.

View file

@ -2,12 +2,12 @@
id: F-0008 id: F-0008
type: framework-finding type: framework-finding
class: UNNECESSARY_COMPLEXITY class: UNNECESSARY_COMPLEXITY
status: open status: resolved
discovered: "2026-08-23" discovered: "2026-08-23"
discovered_by: TD-WP-0002-T10 discovered_by: TD-WP-0002-T10
workplan: TD-WP-0002 workplan: TD-WP-0002
task: TD-WP-0002-T10 task: TD-WP-0002-T10
carried_to: next workplan carried_to: TD-WP-0003-T04
--- ---
# F-0008 — Temperature may be redundant; measured stability did the work # F-0008 — Temperature may be redundant; measured stability did the work
@ -68,3 +68,9 @@ is currently the clearest instance in the corpus.
`Confidence`, `Campaign`, `Metabolism` and `Retirement` are in the same position — `Confidence`, `Campaign`, `Metabolism` and `Retirement` are in the same position —
declared, unimplemented, never consulted — but Temperature is the one with a declared, unimplemented, never consulted — but Temperature is the one with a
built alternative already doing its job, which makes it the decidable case. built alternative already doing its job, which makes it the decidable case.
## Resolution — 2026-09-28
T04 removes Temperature from the current concept set. Three synthetic use cases
and the expanded structural experiment require no temperature decision; observed
trajectory stability continues to govern crystallization. No replacement gate.

View file

@ -61,3 +61,18 @@ relayout is untested.
- 2026-08-22 `PROPOSED`. No evidence. - 2026-08-22 `PROPOSED`. No evidence.
- 2026-08-22 `EXPERIMENTING`. Partial evidence from T07; awaiting E-001 in full. - 2026-08-22 `EXPERIMENTING`. Partial evidence from T07; awaiting E-001 in full.
## Expanded measurement — 2026-09-28 (TD-WP-0003-T03)
38 mutations, including ten that drop test ids and seven new paired structural
controls. Discovery and recorded selectors both pass 17/17 preserved-id cases;
on dropped ids discovery passes 9/10 and recorded selectors 0/10. M22 remains
unrealizable. The full journey accepts 26/27 mechanical cases; FAR remains 0/7
on the unchanged defect set. Claim/provenance indices remain unchanged.
Reproduce: `PYTHONPATH=src:. python3 tools/measure_e001.py /tmp/e001.json`.
[Raw result](../evidence/2026-09-28-e001.json).
These selected and correlated synthetic HTML cases do not establish population
reliability, visual/browser coverage, model capability or economics. See
[review](../../docs/TestDriverGeneralisationReview.md) for the mechanism list and
limits. H-001 remains supported only in the narrowed identifier-loss setting.

View file

@ -1,7 +1,7 @@
--- ---
id: H-005 id: H-005
title: Verification Energy title: Verification Energy
status: PROPOSED status: DORMANT-INDEFINITE
created: "2026-08-22" created: "2026-08-22"
experiments: [] experiments: []
concepts: [C-energy] concepts: [C-energy]
@ -28,9 +28,9 @@ dormant.** Validating it requires event history across many assets over months
history the spike will not accumulate. Implementing a scoring function now would history the spike will not accumulate. Implementing a scoring function now would
produce a number that cannot be checked, which is worse than no number. produce a number that cannot be checked, which is worse than no number.
`TD-WP-0002` therefore records **raw immutable `EnergyEvent`s from the first run** `TD-WP-0002` recorded raw events, but no decision used them. TD-WP-0003-T04
and implements no scoring, decay, or selection logic. Events cannot be removed capture and scoring concepts on 2026-09-28. Ordinary run evidence and
reconstructed later; scores can always be computed later. verdicts remain; no energy history is promised for future reconstruction.
This is the hypothesis most likely to be **cheaply built and never validated**, This is the hypothesis most likely to be **cheaply built and never validated**,
which is precisely why it is fenced off. which is precisely why it is fenced off.
@ -38,3 +38,5 @@ which is precisely why it is fenced off.
## Status log ## Status log
- 2026-08-22 `PROPOSED`, dormant by decision. Event capture only. - 2026-08-22 `PROPOSED`, dormant by decision. Event capture only.
- 2026-09-28 `DORMANT-INDEFINITE`. Capture removed; no new implementation gate.

View file

@ -0,0 +1,73 @@
"""Sequenced delegation contract; no new kernel concepts."""
from testdriver import (
Actor, Cast, Claim, Invariant, Oracle, Provenance, Scenario, SemanticAction,
Step, UseCase, VerificationAsset, World,
)
from lab.workflows import ApprovalLab, ApprovalObserver, WorkflowDriver
SOURCE = "usecases/generalisation-spec.md#delegated-approval"
PROVENANCE = Provenance.AGENT_FROM_SPEC
def receipt(obs, operation, actor, allowed):
return obs["receipts"][-1] == {
"operation": operation, "actor": actor, "allowed": allowed,
}
def execution_matches_record(obs):
return obs["execution_gate"] == {
"requester": obs["state"] == "approved", "reviewer": False, "delegate": False,
}
USE_CASE = UseCase(
"uc-delegated-approval", "Delegate, approve, execute",
"A separate reviewer delegates approval before the requester executes.",
PROVENANCE, source_ref=SOURCE,
claims=(
Claim("approval-premature", "Execution before approval is denied", PROVENANCE,
lambda o: receipt(o, "execute", "requester", False) and o["state"] == "draft",
"premature", SOURCE),
Claim("approval-submit", "Submission enters pending review", PROVENANCE,
lambda o: o["state"] == "pending", "submit", SOURCE),
Claim("approval-delegate", "Authority passes to the delegate", PROVENANCE,
lambda o: o["reviewer"] == "delegate", "delegate", SOURCE),
Claim("approval-self", "Requester cannot self-approve", PROVENANCE,
lambda o: receipt(o, "approve", "requester", False) and o["state"] == "pending",
"self-approve", SOURCE),
Claim("approval-former", "Former reviewer cannot approve", PROVENANCE,
lambda o: receipt(o, "approve", "reviewer", False) and o["state"] == "pending",
"former-approve", SOURCE),
Claim("approval-approved", "Delegate approves", PROVENANCE,
lambda o: o["state"] == "approved" and o["approved_by"] == "delegate",
"approve", SOURCE),
Claim("approval-executed", "Requester executes approved work", PROVENANCE,
lambda o: o["state"] == "executed", "execute", SOURCE),
),
invariants=(Invariant("approval-gate", "Execution gate matches stored workflow",
PROVENANCE, execution_matches_record, SOURCE),),
)
def build(defect=None):
lab = ApprovalLab(defect)
cast = Cast()
for actor, token in lab.tokens.items():
cast.add(Actor(actor, actor.title(), credentials={"token": token}))
schedule = (
("premature", "requester", "execute", {}),
("submit", "requester", "submit", {}),
("delegate", "reviewer", "delegate", {"to": "delegate"}),
("self-approve", "requester", "approve", {}),
("former-approve", "reviewer", "approve", {}),
("approve", "delegate", "approve", {}),
("execute", "requester", "execute", {}),
)
scenario = Scenario("sc-delegated-approval", USE_CASE, tuple(
Step(id, actor, SemanticAction(operation, args, frozenset({"workflow"})))
for id, actor, operation, args in schedule
), variant=defect or "baseline")
return (World("w-approval", lab, f"approval-1/{defect or 'baseline'}", cast),
WorkflowDriver(lab), ApprovalObserver(lab),
VerificationAsset("va-approval", scenario), Oracle())

View file

@ -0,0 +1,70 @@
"""Same-name resources, explicit foreign access and delete/recreate isolation."""
from testdriver import (
Actor, Cast, Claim, Invariant, Oracle, Provenance, Scenario, SemanticAction,
Step, UseCase, VerificationAsset, World,
)
from lab.workflows import TenantLab, TenantObserver, WorkflowDriver
SOURCE = "usecases/generalisation-spec.md#tenant-lifecycle"
PROVENANCE = Provenance.AGENT_FROM_SPEC
def isolated_reads(obs):
for actor, own, foreign in (("admin-a", "A", "B"), ("admin-b", "B", "A")):
if obs["reads"][f"{actor}:{foreign}"] != {"allowed": False, "content": None}:
return False
record = obs["records"][own]
if obs["reads"][f"{actor}:{own}"] != {"allowed": record is not None, "content": record}:
return False
return True
def foreign_delete_refused(obs):
return (obs["receipts"][-1]["allowed"] is False
and obs["records"] == {"A": "alpha", "B": "beta"})
USE_CASE = UseCase(
"uc-tenant-lifecycle", "Tenant-scoped deletion and recreation",
"Two administrators use R independently; deleting and recreating A never affects B.",
PROVENANCE, source_ref=SOURCE,
claims=(
Claim("tenant-create-a", "A holds alpha", PROVENANCE,
lambda o: o["records"] == {"A": "alpha", "B": None}, "create-a", SOURCE),
Claim("tenant-create-b", "Same local name preserves both records", PROVENANCE,
lambda o: o["records"] == {"A": "alpha", "B": "beta"}, "create-b", SOURCE),
Claim("tenant-deny-a", "A cannot delete B", PROVENANCE,
foreign_delete_refused, "foreign-a", SOURCE),
Claim("tenant-deny-b", "B cannot delete A", PROVENANCE,
foreign_delete_refused, "foreign-b", SOURCE),
Claim("tenant-delete", "Deleting A preserves B", PROVENANCE,
lambda o: o["records"] == {"A": None, "B": "beta"}, "delete-a", SOURCE),
Claim("tenant-recreate", "Recreate does not resurrect or overwrite", PROVENANCE,
lambda o: o["records"] == {"A": "new-alpha", "B": "beta"}, "recreate-a", SOURCE),
),
invariants=(Invariant("tenant-reads", "Read enforcement matches tenant records",
PROVENANCE, isolated_reads, SOURCE),),
)
def build(defect=None):
lab = TenantLab(defect)
cast = Cast()
for actor, token in lab.tokens.items():
cast.add(Actor(actor, actor.title(), credentials={"token": token}))
schedule = (
("create-a", "admin-a", "create", "A", "alpha"),
("create-b", "admin-b", "create", "B", "beta"),
("foreign-a", "admin-a", "delete", "B", None),
("foreign-b", "admin-b", "delete", "A", None),
("delete-a", "admin-a", "delete", "A", None),
("recreate-a", "admin-a", "create", "A", "new-alpha"),
)
scenario = Scenario("sc-tenant-lifecycle", USE_CASE, tuple(
Step(id, actor, SemanticAction(operation, {"tenant": tenant, "content": content},
frozenset({"workflow"})))
for id, actor, operation, tenant, content in schedule
), variant=defect or "baseline")
return (World("w-tenant", lab, f"tenant-1/{defect or 'baseline'}", cast),
WorkflowDriver(lab), TenantObserver(lab),
VerificationAsset("va-tenant", scenario), Oracle())

View file

@ -2,7 +2,6 @@
from .actions import SemanticAction, Surface, SurfaceNotPermitted from .actions import SemanticAction, Surface, SurfaceNotPermitted
from .drivers import DirectDriver, Realization from .drivers import DirectDriver, Realization
from .energy import EnergyEvent, EnergyEventType
from .evidence import EvidencePack, Observation, Stratum from .evidence import EvidencePack, Observation, Stratum
from .intent import Claim, Invariant, UseCase from .intent import Claim, Invariant, UseCase
from .observers import StateObserver, Watch from .observers import StateObserver, Watch
@ -14,7 +13,7 @@ from .world import Actor, Cast, World
__all__ = [ __all__ = [
"Actor", "Cast", "Claim", "CollectorIndependenceError", "Actor", "Cast", "Claim", "CollectorIndependenceError",
"DirectDriver", "EnergyEvent", "EnergyEventType", "EvidencePack", "DirectDriver", "EvidencePack",
"InadmissibleProvenance", "Invariant", "Judgment", "Observation", "Oracle", "InadmissibleProvenance", "Invariant", "Judgment", "Observation", "Oracle",
"Provenance", "Realization", "RunResult", "Runner", "Scenario", "Provenance", "Realization", "RunResult", "Runner", "Scenario",
"SemanticAction", "StateObserver", "Step", "Stratum", "Surface", "SemanticAction", "StateObserver", "Step", "Stratum", "Surface",

View file

@ -1,45 +0,0 @@
"""Energy events — capture only.
H-005 is dormant by decision: verification energy is not testable at the current
scale, and a scoring function producing a number nobody can check is worse than
no number. Events are recorded from the first run because history cannot be
reconstructed later; scores always can.
There is deliberately no score() function in this module.
"""
from __future__ import annotations
from dataclasses import asdict, dataclass, field
from datetime import datetime, timezone
from enum import Enum
from typing import Any
class EnergyEventType(str, Enum):
DEFECT_DETECTED = "DEFECT_DETECTED"
REGRESSION_CAUGHT = "REGRESSION_CAUGHT"
MECHANICAL_ADAPTATION = "MECHANICAL_ADAPTATION"
SEMANTIC_ADAPTATION = "SEMANTIC_ADAPTATION"
TEST_DEFECT = "TEST_DEFECT"
FALSE_POSITIVE = "FALSE_POSITIVE"
DUPLICATE = "DUPLICATE"
CRYSTALLIZED = "CRYSTALLIZED"
USECASE_DEPRECATED = "USECASE_DEPRECATED"
EXECUTED = "EXECUTED"
@dataclass(frozen=True, slots=True)
class EnergyEvent:
"""Immutable. Energy is derived from event history, never stored as state."""
asset_id: str
run_id: str
event_type: EnergyEventType
detail: dict[str, Any] = field(default_factory=dict)
at: str = field(
default_factory=lambda: datetime.now(timezone.utc).isoformat()
)
def as_dict(self) -> dict[str, Any]:
return {**asdict(self), "event_type": self.event_type.value}

View file

@ -63,7 +63,6 @@ class EvidencePack:
finished_at: str | None = None finished_at: str | None = None
observations: list[Observation] = field(default_factory=list) observations: list[Observation] = field(default_factory=list)
verdicts: list[dict[str, Any]] = field(default_factory=list) verdicts: list[dict[str, Any]] = field(default_factory=list)
energy_events: list[dict[str, Any]] = field(default_factory=list)
provenance_index: dict[str, str] = field(default_factory=dict) provenance_index: dict[str, str] = field(default_factory=dict)
def record(self, observation: Observation) -> None: def record(self, observation: Observation) -> None:

View file

@ -20,7 +20,12 @@ Predicate = Callable[[Mapping[str, object]], bool]
@dataclass(frozen=True, slots=True) @dataclass(frozen=True, slots=True)
class Claim: class Claim:
"""A statement that must hold at a specific point in a scenario.""" """A statement that must hold at a specific point in a scenario.
Public contract (TD-WP-0003-T05): predicates remain Python callables over
independent snapshots. Generated tests import these original claims; no
serialized claim language is promised. See TestDriverGeneralisationReview.
"""
id: str id: str
text: str text: str

View file

@ -15,7 +15,6 @@ from typing import Any
from .actions import SurfaceNotPermitted from .actions import SurfaceNotPermitted
from .drivers import Driver from .drivers import Driver
from .energy import EnergyEvent, EnergyEventType
from .evidence import EvidencePack, Observation, Stratum from .evidence import EvidencePack, Observation, Stratum
from .observers import StateObserver from .observers import StateObserver
from .oracles import Judgment, Oracle, Verdict, overall from .oracles import Judgment, Oracle, Verdict, overall
@ -125,9 +124,6 @@ class Runner:
**{c.id: c.provenance.value for c in scenario.use_case.claims}, **{c.id: c.provenance.value for c in scenario.use_case.claims},
**{i.id: i.provenance.value for i in scenario.use_case.invariants}, **{i.id: i.provenance.value for i in scenario.use_case.invariants},
} }
pack.energy_events.append(
EnergyEvent(asset.id, run_id, EnergyEventType.EXECUTED).as_dict()
)
isolation = self._isolation_violations() isolation = self._isolation_violations()
self._record( self._record(
@ -229,8 +225,4 @@ class Runner:
pack.finished_at = datetime.now(timezone.utc).isoformat() pack.finished_at = datetime.now(timezone.utc).isoformat()
result_verdict = overall(judgments) result_verdict = overall(judgments)
if result_verdict is Verdict.FAIL:
pack.energy_events.append(
EnergyEvent(asset.id, run_id, EnergyEventType.DEFECT_DETECTED).as_dict()
)
return RunResult(run_id, result_verdict, judgments, pack) return RunResult(run_id, result_verdict, judgments, pack)

View file

@ -173,3 +173,19 @@ def test_the_control_arm_is_not_a_straw_man():
preserved = [m for m in MECHANICAL if m.preserves_test_ids] preserved = [m for m in MECHANICAL if m.preserves_test_ids]
assert all(realized_ok(m.id, runtime=runtime) for m in preserved) assert all(realized_ok(m.id, runtime=runtime) for m in preserved)
assert len(preserved) >= 9 assert len(preserved) >= 9
def test_deciding_axis_has_ten_mutations_and_paired_structural_controls():
from lab.mutations import EXTENDED_LAYOUTS
from lab.http_api import render_resource_page
from lab.mutations import build_lab
assert sum(not m.preserves_test_ids for m in MECHANICAL) >= 10
for dropped, preserved, _, _ in EXTENDED_LAYOUTS:
dropped_app, _ = build_lab(dropped)
preserved_app, _ = build_lab(preserved)
# Version metadata differs; control markup must differ only in test ids.
import re
def shape(app):
html = render_resource_page(app, "alice", "R")
return re.sub(r'\s*data-td(?:-version)?="[^"]*"', "", html)
assert shape(dropped_app) == shape(preserved_app)

View file

@ -65,6 +65,12 @@ EXPECTED = {
"M24": Classification.MECHANICAL_ADAPTATION, "M24": Classification.MECHANICAL_ADAPTATION,
} }
EXPECTED.update({
f"M{number}": (Classification.UNCHANGED if number in (25, 32)
else Classification.MECHANICAL_ADAPTATION)
for number in range(25, 39)
})
@pytest.mark.parametrize("mutation_id", sorted(EXPECTED)) @pytest.mark.parametrize("mutation_id", sorted(EXPECTED))
def test_response_matrix(baseline, mutation_id): def test_response_matrix(baseline, mutation_id):

View file

@ -0,0 +1,81 @@
"""Independent checks that new use cases discriminate defects and evidence loss."""
import json
import pytest
from testdriver import Oracle, Runner, Stratum, Verdict
from scenarios.delegated_approval import build as approval
from scenarios.tenant_lifecycle import build as tenant
def run(builder, defect=None):
world, driver, observer, asset, oracle = builder(defect)
return Runner(world, driver, observer, oracle).run(asset)
def test_approval_baseline():
assert run(approval).verdict is Verdict.PASS
@pytest.mark.parametrize("defect,claim", [
("premature-execution", "approval-premature"),
("self-approval", "approval-self"),
("stale-delegation", "approval-former"),
])
def test_approval_defects(defect, claim):
assert run(approval, defect).judgment(claim).verdict is Verdict.FAIL
def test_approval_denials_are_judged_not_skipped():
result = run(approval)
assert len(result.judgments) == 14
assert all(j.verdict is Verdict.PASS for j in result.judgments)
def test_tenant_baseline():
assert run(tenant).verdict is Verdict.PASS
@pytest.mark.parametrize("defect,claim", [
("foreign-read", "tenant-reads"),
("foreign-delete", "tenant-deny-a"),
("unscoped-delete", "tenant-delete"),
("resurrect-content", "tenant-recreate"),
])
def test_tenant_defects(defect, claim):
result = run(tenant, defect)
assert any(j.assertion_id == claim and j.verdict is Verdict.FAIL for j in result.judgments)
@pytest.mark.parametrize("builder,defect", [
(approval, None), (approval, "self-approval"),
(tenant, None), (tenant, "foreign-read"),
])
def test_verdicts_reproduce_from_serialized_s3_without_actors(builder, defect):
result = run(builder, defect)
pack = json.loads(result.evidence.to_json())
assertions = builder()[3].scenario.use_case
snapshots = {o["step_id"]: o["data"] for o in pack["observations"]
if o["kind"] == "state_snapshot"}
by_id = {a.id: a for a in (*assertions.claims, *assertions.invariants)}
for j in pack["verdicts"]:
replayed = Oracle().judge(by_id[j["assertion_id"]], snapshots[j["step_id"]], j["step_id"])
assert replayed.verdict.value == j["verdict"]
actors = set(builder()[0].cast.actors)
assert not actors.intersection(o.collector for o in result.evidence.observations
if o.stratum is not Stratum.SURFACE)
@pytest.mark.parametrize("builder", [approval, tenant])
def test_missing_evidence_is_inconclusive_for_every_assertion(builder):
case = builder()[3].scenario.use_case
for assertion in (*case.claims, *case.invariants):
assert Oracle().judge(assertion, {"unrelated": True}, None).verdict is Verdict.INCONCLUSIVE
@pytest.mark.parametrize("builder", [approval, tenant])
def test_replay_keeps_judgments_and_claims(builder):
case = builder()[3].scenario.use_case
before = (case.claims, case.invariants)
assert run(builder).judgments == run(builder).judgments
assert (case.claims, case.invariants) == before

View file

@ -130,10 +130,3 @@ def test_seeded_authorization_defect_fails_the_run():
# The claim set is untouched by the failure — there is no path to adapt it. # The claim set is untouched by the failure — there is no path to adapt it.
by_id = {c.id: c for c in USE_CASE.claims} by_id = {c.id: c for c in USE_CASE.claims}
assert by_id["c-bob-revoked"].text == "Bob cannot read R after revocation" assert by_id["c-bob-revoked"].text == "Bob cannot read R after revocation"
def test_defect_run_emits_an_energy_event():
"""Energy events are captured; no score is computed (H-005 is dormant)."""
import testdriver.energy as energy
assert not hasattr(energy, "score")

67
tools/measure_e001.py Normal file
View file

@ -0,0 +1,67 @@
"""Reproduce E-001: PYTHONPATH=src:. python3 tools/measure_e001.py OUTPUT.json."""
import json
import sys
from collections import defaultdict
from datetime import datetime, timezone
from pathlib import Path
from lab.mutations import CATALOGUE
from scenarios.browser_grant import baseline_recordings, build_agentic, lab_server
from scenarios.full_journey import build_journey, journey_lab_server
from testdriver import Runner, Stratum
from testdriver.agentic import DiscoveryRuntime, RecordedSelectorRuntime
from testdriver.classification import classify
def journey(*mutations):
with journey_lab_server(*mutations) as (app, tokens, url):
world, driver, observer, asset, oracle = build_journey(app, tokens, url)
return json.loads(Runner(world, driver, observer, oracle).run(asset).evidence.to_json())
def measure():
baseline = journey()
selectors = baseline_recordings()
rows = []
for mutation in CATALOGUE:
pack = journey(mutation.id)
outcome = classify(baseline, pack)
row = {"mutation": mutation.id, "label": mutation.label,
"preserves_test_ids": mutation.preserves_test_ids,
"classification": outcome.as_dict(), "arms": {}}
if mutation.label == "MECHANICAL":
for name, runtime in (("discovery", DiscoveryRuntime()),
("recorded", RecordedSelectorRuntime(selectors))):
with lab_server(mutation.id) as (app, tokens, url):
world, driver, observer, asset, oracle = build_agentic(app, tokens, url, runtime)
result = Runner(world, driver, observer, oracle).run(asset)
surface = result.evidence.of_stratum(Stratum.SURFACE)[0].data
row["arms"][name] = {
"realized": surface["raised"] is None,
"verdict": result.verdict.value,
"surface": surface,
"judgments": [j.as_dict() for j in result.judgments],
}
rows.append(row)
summary = defaultdict(lambda: {"total": 0, "discovery": 0, "recorded": 0})
for row in rows:
if not row["arms"]:
continue
group = summary["preserved" if row["preserves_test_ids"] else "dropped"]
group["total"] += 1
for arm in ("discovery", "recorded"):
group[arm] += int(row["arms"][arm]["realized"]
and row["arms"][arm]["verdict"] == "PASS")
false_adaptations = [row["mutation"] for row in rows if row["label"] == "DEFECT"
and row["classification"]["classification"]
in ("UNCHANGED", "MECHANICAL_ADAPTATION")]
return {"measured_at": datetime.now(timezone.utc).isoformat(),
"scope": "Synthetic server-rendered HTML, deterministic runtimes; no browser engine or live model",
"summary": dict(summary), "false_adaptations": false_adaptations,
"defect_count": sum(row["label"] == "DEFECT" for row in rows), "rows": rows}
if __name__ == "__main__":
result = measure()
Path(sys.argv[1]).write_text(json.dumps(result, indent=2) + "\n")
print(json.dumps({k: v for k, v in result.items() if k != "rows"}, indent=2))

View file

@ -0,0 +1,27 @@
# Synthetic generalisation contracts — 2026-09-28
These contracts are written before the corresponding lab implementations for
TD-WP-0003-T02. They are agent-authored research specifications, not externally
validated product requirements. Scenario claims use `agent-from-spec` and refer
here. Same-session authorship limits epistemic independence; passing these labs
must not be presented as independent validation of a real product.
## Delegated approval
A requester submits a request for review. Only the assigned reviewer can delegate
it to another reviewer. The requester cannot approve their own request; the
former reviewer loses approval authority after delegation. The delegate may
approve only after submission, and only the requester may then execute it.
Premature execution and unauthorized approval are refused without state change.
The persisted workflow and the independently probed execution gate must agree.
A complete run includes premature execution, submission, delegation, attempted
self-approval, former-reviewer approval, delegate approval and execution.
## Tenant lifecycle
Two tenant administrators independently use the same local resource name `R`.
Each can create, read and delete their own tenant's instance. Neither may read or
delete the other tenant's instance, including by supplying an explicit foreign
tenant id. Deleting tenant A's instance preserves tenant B's content. Recreating
A's instance must not resurrect its old content or affect B. At every step,
foreign reads remain denied and successful reads match the stored tenant record.

View file

@ -4,17 +4,25 @@ type: workplan
title: "Generalise the model and settle the open questions" title: "Generalise the model and settle the open questions"
domain: infotech domain: infotech
repo: test-driver repo: test-driver
status: proposed status: blocked
flavor: planning flavor: implementation
owner: codex owner: codex
topic_slug: custodian topic_slug: custodian
created: "2026-08-23" created: "2026-08-23"
updated: "2026-08-23" updated: "2026-09-28"
state_hub_workstream_id: "36a082df-1288-567d-9a0b-f2e0f292786c" state_hub_workstream_id: "36a082df-1288-567d-9a0b-f2e0f292786c"
--- ---
# Generalise the model and settle the open questions # Generalise the model and settle the open questions
## Closeout status — 2026-09-28
T02/T03/T04/T05/T08 are done. T01/T06/T07 remain waiting for external experiment
choices or a valid independent authoring sample, so the workplan is **blocked**.
No new task or workplan was created. The live-model success gate is still unmet;
this is a partial closeout, not a declaration that the workplan succeeded.
Detailed evidence and interface changes: `docs/TestDriverGeneralisationReview.md`.
## Why this workplan exists ## Why this workplan exists
`TD-WP-0002` passed its gate: False Adaptation Rate 0/7, one asset crystallized `TD-WP-0002` passed its gate: False Adaptation Rate 0/7, one asset crystallized
@ -76,7 +84,7 @@ consumes the kernel as a library. Two consequences for this workplan:
```task ```task
id: TD-WP-0003-T01 id: TD-WP-0003-T01
status: todo status: wait
priority: high priority: high
state_hub_task_id: "0090ce51-0086-5079-a094-fb63b3b53415" state_hub_task_id: "0090ce51-0086-5079-a094-fb63b3b53415"
``` ```
@ -107,11 +115,13 @@ reliably, and the difference will not show in an average.
the live runtime ever runs in the default test suite (recommendation: no — keep the live runtime ever runs in the default test suite (recommendation: no — keep
the suite deterministic and free, run the live arm on demand). the suite deterministic and free, run the live arm on demand).
**2026-09-28 closeout — wait.** Await an explicit model, fixed run count, cost ceiling and suite policy from Bernd Worsch before the first live call. No live calls or costs were incurred; local heuristic evidence cannot close capability/economics.
## Two further use cases ## Two further use cases
```task ```task
id: TD-WP-0003-T02 id: TD-WP-0003-T02
status: todo status: done
priority: high priority: high
state_hub_task_id: "eedf5894-e540-5fa7-9dc3-7e2a5c892bd6" state_hub_task_id: "eedf5894-e540-5fa7-9dc3-7e2a5c892bd6"
``` ```
@ -133,11 +143,13 @@ express them**, and every such growth is evidence:
Record the answer in the fitness map either way. Record the answer in the fitness map either way.
**2026-09-28 closeout — done.** Added delegated approval and tenant lifecycle from prewritten synthetic specifications, using existing kernel concepts and independent snapshot collectors. Baseline, seven seeded defects, missing evidence and S3 replay are covered in tests/test_generalisation.py. Same-agent synthetic authorship is explicitly limited, not human provenance.
## Strengthen the instrument on the deciding axis ## Strengthen the instrument on the deciding axis
```task ```task
id: TD-WP-0003-T03 id: TD-WP-0003-T03
status: todo status: done
priority: high priority: high
state_hub_task_id: "95406a5d-bcb6-5bcb-abed-834ad16aeb69" state_hub_task_id: "95406a5d-bcb6-5bcb-abed-834ad16aeb69"
``` ```
@ -153,11 +165,13 @@ side proportionate so the comparison stays fair.
Then re-run E-001 and restate H-001 with a number that means something. Be Then re-run E-001 and restate H-001 with a number that means something. Be
prepared for the honest outcome that the narrowed claim narrows further. prepared for the honest outcome that the narrowed claim narrows further.
**2026-09-28 closeout — done.** Expanded the catalogue to 38 mutations, with ten dropped-identifier mechanisms and seven new paired preserved-id controls. Reproduced E-001: preserved 17/17 in both arms; dropped 9/10 discovery versus 0/10 recorded; full-journey mechanical acceptance 26/27, FAR 0/7. Raw receipt: research/evidence/2026-09-28-e001.json.
## Settle every gated concept ## Settle every gated concept
```task ```task
id: TD-WP-0003-T04 id: TD-WP-0003-T04
status: todo status: done
priority: medium priority: medium
state_hub_task_id: "458dc9aa-3d0c-5e67-9f89-3ae8a5a5d297" state_hub_task_id: "458dc9aa-3d0c-5e67-9f89-3ae8a5a5d297"
``` ```
@ -176,11 +190,13 @@ evidence to say whether they are deferred or dead.
Removing a concept is a result. Extending a gate is not. Removing a concept is a result. Extending a gate is not.
**2026-09-28 closeout — done.** Removed Temperature, Energy, Confidence, Campaign, Metabolism and automated Retirement from the current model; removed energy.py, exports and capture field. H-005 is dormant-indefinite; F-0008 resolved. Compatibility changes and superseded historical proposals are explicit.
## Decide how claims are expressed ## Decide how claims are expressed
```task ```task
id: TD-WP-0003-T05 id: TD-WP-0003-T05
status: todo status: done
priority: medium priority: medium
state_hub_task_id: "34cc60cd-b966-54f8-a92d-cfe64587822f" state_hub_task_id: "34cc60cd-b966-54f8-a92d-cfe64587822f"
``` ```
@ -202,6 +218,8 @@ Three candidate answers, and the middle one should not win by default:
Whatever is chosen, publish it as an interface change — the audit-core session Whatever is chosen, publish it as an interface change — the audit-core session
consumes `Claim` directly. consumes `Claim` directly.
**2026-09-28 closeout — done.** Published the decision to retain Python predicates and imported original claims, after comparing all three runnable use cases and the audit-core consumer. Claim and Invariant signatures remain compatible; standalone serialization is deliberately rejected. See docs/TestDriverGeneralisationReview.md.
## A browser-engine driver, if it is warranted ## A browser-engine driver, if it is warranted
```task ```task
@ -226,11 +244,13 @@ as the default.
Do not start this before T01–T03. It widens the surface; those three settle what Do not start this before T01–T03. It widens the surface; those three settle what
the surface is for. the surface is for.
**2026-09-28 closeout — wait.** Still requires the existing Playwright setup/licence decision and T01 live-model prerequisite. T02/T03 are now complete. Structural HTML results do not satisfy the browser-engine/visual experiment.
## Measure the cost of expressing a use case ## Measure the cost of expressing a use case
```task ```task
id: TD-WP-0003-T07 id: TD-WP-0003-T07
status: todo status: wait
priority: medium priority: medium
state_hub_task_id: "62f08e27-f307-560f-9583-64c08835e29e" state_hub_task_id: "62f08e27-f307-560f-9583-64c08835e29e"
``` ```
@ -248,11 +268,13 @@ This feeds the metric the project ultimately cares about — *verified behaviour
unit of human maintenance effort* — and it is the number an adopter will ask for unit of human maintenance effort* — and it is the number an adopter will ask for
first. It cannot be reconstructed later. first. It cannot be reconstructed later.
**2026-09-28 closeout — wait.** Partial evidence only: concepts, adapter friction, line counts and raw timer receipts recorded in docs/TestDriverGeneralisationReview.md and research/evidence/2026-09-28-authoring.json. Timer placement missed code composition, so first-authoring elapsed costs are invalid and cannot be reconstructed. Await an independently timed fresh implementation by an author who does not copy these fixtures; retain this task, no replacement record.
## Readiness review for real-system application ## Readiness review for real-system application
```task ```task
id: TD-WP-0003-T08 id: TD-WP-0003-T08
status: todo status: done
priority: high priority: high
state_hub_task_id: "9e040e94-7245-5b4b-b177-cc337f904ceb" state_hub_task_id: "9e040e94-7245-5b4b-b177-cc337f904ceb"
``` ```
@ -275,3 +297,5 @@ At minimum the criteria must cover:
Then state a verdict, including the verdict "not yet, and here is what is missing". Then state a verdict, including the verdict "not yet, and here is what is missing".
A readiness review that cannot conclude *not ready* is not a review. A readiness review that cannot conclude *not ready* is not a review.
**2026-09-28 closeout — done.** Published explicit real-system gates and a NOT READY verdict for autonomous application, including D-07 cost, provenance from real backlogs, ambiguity, engagement-specific false-adaptation harm, custody/timing/cleanup and economics. The audit-core contract remains intent plus fixture calibration, not a runnable production engagement.