Align scope with intent and add durable evidence and bounded variants

Assistant: codex
Assistant-Model: gpt-6-astra
Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
This commit is contained in:
tegwick 2026-09-28 14:31:22 +02:00
parent 8822480b6e
commit eaf5d348e4
14 changed files with 767 additions and 92 deletions

View file

@ -13,6 +13,10 @@ multi-user use cases and a 38-mutation catalogue exercise the model. Live-model
economics, browser-engine coverage and independent authoring-cost measurement economics, browser-engine coverage and independent authoring-cost measurement
remain blocked in `workplans/TD-WP-0003-generalise-and-settle.md`. remain blocked in `workplans/TD-WP-0003-generalise-and-settle.md`.
See [readiness and interface decisions](docs/TestDriverGeneralisationReview.md). See [readiness and interface decisions](docs/TestDriverGeneralisationReview.md).
Opt-in local evidence receipts and explicit actor/argument variants are available;
see [usage](docs/TestDriverEvidenceAndVariants.md) and the
[scope/intent assessment](history/2026-09-28-121933-scope-intent-assessment.md).
An independently owned real-system pilot remains blocked in TD-WP-0004.
## Run ## Run

149
SCOPE.md
View file

@ -1,103 +1,80 @@
# SCOPE # SCOPE
> This file helps you quickly understand what this repository is about, Assessed 2026-09-28. `test-driver` is an executable **local research prototype**
> when it is relevant, and when it is not. for use-case-driven verification, adaptation classification and deterministic
replay. It is not yet a validated autonomous test platform for real systems.
--- ## What it can do
## One-liner | Capability | Implemented boundary and evidence |
|---|---|
| Deterministic verification | Python UseCase/Claim/Invariant inputs, ordered semantic steps, separate drivers and observers, explicit PASS/FAIL/INCONCLUSIVE. Non-boolean or missing predicate evidence is inconclusive. `src/testdriver/runner.py`, `oracles.py`; `tests/test_kernel_guarantees.py`. |
| Multi-user/domain examples | Sharing and revocation, delegated approval, and tenant deletion/recreation run against three synthetic lab domains. Actor stores/canaries and S2/S3 collector attribution are checked. These are in-process diagnostics, not process isolation. `scenarios/`, `tests/test_generalisation.py`. |
| Mechanical discovery | A deterministic heuristic finds forms in server-rendered HTML; a recorded-selector control provides comparison. This is not live-model inference or a JavaScript browser engine. `agentic.py`, `browser.py`. |
| Evidence-based classification | Compare complete passing reference runs with candidate evidence, explicit intent revisions and postconditions. Distinguish mechanical changes, observed behavior changes, changed intent, realization failures and ambiguity. Cannot infer whether a product behavior change was authorized. `classification.py`, `revisions.py`. |
| Crystallization | Check repeated eligible, distinct runs for stable paths; replay frozen paths by step; generate a single-action pytest module importing original predicates. One checked-in browser descendant runs without model involvement. No automatic full T0–T5 lifecycle or general multi-step code generator. `crystallization.py`, `crystallized/`. |
| Retained evidence and lineage | Optional local `EvidenceStore` writes versioned, checksummed JSON receipts atomically without replacing an existing run. Receipts contain asset/parent/maturity/variant, intent revisions, scheduled coverage, observations and final verdict. Reloaded packs can be classified. See [usage](docs/TestDriverEvidenceAndVariants.md). |
| Bounded adversarial variants | `variants.substitute` derives one actor or argument substitution, retaining the exact independent use case, assertions, postconditions and surface permissions. Resource, tenant and privilege values use the same operation. Scenario revisions prevent changed action intent from qualifying as mere mechanical adaptation. No automatic claim authoring. |
| Local research calibration | A 38-mutation catalogue, independent fixture judgments, mechanical/defect experiments and tests that deliberately break framework guarantees. Lab false-adaptation results are bounded experiment results, not a universal safety proof. `lab/GROUND-TRUTH.md`, `research/`, `tests/selfverification/`. |
| HTTP boundary checks | Browser and frozen/standalone HTTP bind bearer credentials to a configured origin, including redirects. Runner also validates reported surfaces. Drivers still need pre-action checks; a post-call check cannot undo effects or attest an untruthful report. |
`test-driver` is a use-case-driven verification framework whose tests mature ## Where it is useful
alongside the software they protect — fluid and agentic while behaviour is hot,
deterministic once it cools.
--- Use it to investigate whether independent Python claims survive changed
mechanics, to build small sequential multi-user domain fixtures, to calibrate
mutation detection, or to replay a demonstrated stable HTML action. A new domain
needs an explicitly authored scenario, driver and independent observation adapter.
## Core Idea The audit-core E2 module is an approved-intent contract and fixture calibration,
**not** a runnable production engagement. It has no deployed custody, cluster,
time-window or independent cleanup adapters in this repository.
A test is treated not as code but as a **verification asset** with identity, ## Limits and readiness
intent, evidence, lineage, maturity, temperature and energy. Use cases are the
primary behavioural source; integration, journey, multi-user, security and
resilience tests are projections of the same use case rather than separate
suites.
Verification assets progress along a maturity continuum - No natural-language use-case parser or autonomous scenario planner. Intent is
(`T0 Exploratory → T5 Deterministic`) called **Crystallization**. Agents may explicitly authored Python; source/provenance labels are not external attestation.
explore and realise semantic actions against unstable interfaces; oracles remain - No live-model experiment or validated agentic cost benefit. The runtime protocol
deterministic and independent from actors, so the framework can distinguish a exists; the demonstrated runtime is heuristic and token-free.
legitimate mechanical change from a product defect rather than adapting to - No Playwright/browser engine, JavaScript execution, screenshots or DOM timing.
whatever the implementation happens to do. - Sequential execution only: no actual grant/revoke races, causal deadlines,
concurrency scheduler or general dependency-fault/resilience engine.
- No generic CLI, Kubernetes, message-bus or multi-service production adapters.
- Evidence storage is opt-in and local. It is not encryption, signing, automatic
redaction, retention management or safe storage of arbitrary secrets. Unexpected
driver/observer exceptions still propagate; no finalized receipt is promised for
an interrupted/crashed process. Strict JSON rejects unsupported evidence values.
- No automatic finding-to-regression pipeline, asset registry, maturity promotion,
thawing or retirement. Parent ids and mutation history provide limited lineage.
**Current state:** research prototype. Concept corpus is complete Temperature, Energy, Confidence, Campaign, Metabolism and automated Retirement
(`INTENT.md`, `docs/`); implementation has not started. See were **deliberately removed**, not postponed implementation requirements. The
`history/2026-08-22-concept-assessment-swot.md` for the standing assessment and superseding note in INTENT.md and [settlement](docs/TestDriverGeneralisationReview.md)
the reasoning behind the current workplan sequence. take precedence over its historical lifecycle sections.
--- Outstanding work is tracked in [TD-WP-0003](workplans/TD-WP-0003-generalise-and-settle.md)
and [TD-WP-0004](workplans/TD-WP-0004-scope-evidence-and-variants.md), with priorities
and evidence in the [scope/intent assessment](history/2026-09-28-121933-scope-intent-assessment.md).
## In Scope ## Non-goals
- The conceptual model: UseCase, Actor, Scenario, SemanticAction, Observation, Replacing pytest, browser engines or CI; a distributed test cloud; a vulnerability
Oracle, Verdict, VerificationAsset, Finding, Adaptation, Crystallization. scanner, load-testing or observability platform; automatic rewriting of semantic
- A deterministic semantic scenario kernel and its evidence format. requirements; model-based verdicts; exhaustive scenario permutations. These remain
- The **test-driver lab** — a small mutable application under test carrying outside the intended initial project.
labelled mechanical, semantic and defect mutations as ground truth.
- Adaptation detection and the defect-vs-adaptation classifier.
- Crystallization of agentic realisations into deterministic regression tests.
- Security testing expressed as mutation of ordinary use cases.
- Self-verification of the framework's own foundational guarantees.
- The research control plane: hypotheses, experiments, findings, fitness map.
--- ## Actual stack and entry points
## Out of Scope Python ≥3.11, dataclasses and the standard library; pytest for testing. Labs use
in-memory domain state and a local stdlib HTTP server. No runtime third-party
dependencies, database, SQLite, Pydantic, YAML parser or browser engine is used.
- Replacing unit-test frameworks, browser automation engines, or CI systems. ```bash
- Building a load-testing, fuzzing, vulnerability-scanning or observability python3 -m pytest -q
platform. python3 -m pytest -q tests/test_reference_scenario.py tests/test_scope_delivery.py
- Test-management SaaS, distributed test clouds, or multi-tenant hosting. ```
- Making all tests agentic, or using model judgment where a deterministic oracle
is available.
- Rewriting semantic requirements to match implementation behaviour.
- Scale and performance verification before the conceptual model is proven.
--- Read `INTENT.md` for the thesis, this file for current capability,
`docs/TestDriverEvidenceAndVariants.md` for retained-run/variant examples,
## Relevant When `docs/TestDriverGeneralisationReview.md` for readiness constraints, and `workplans/`
for execution status. The concept corpus contains historical proposals and is not
- You need the test-driver conceptual vocabulary or its canonical concept set. an implementation inventory.
- You are working on the crystallization, adaptation-classification, or
semantic-action binding mechanisms.
- You are extending the lab or its mutation catalogue.
- You need the framework's hypotheses, fitness scorecard, or evidence format.
---
## Not Relevant When
- You need ordinary unit or component tests for another repo — use that repo's
own test stack.
- You are looking for fleet coordination or cross-repo memory — that is State
Hub.
---
## Getting Oriented
1. `INTENT.md` — purpose, thesis, design heuristics, non-goals.
2. `docs/TestDriverConceptModel.md` — the canonical concept set v0.1.
3. `docs/TestDriverImprovementLoop.md` — hypotheses, findings taxonomy,
self-improvement cycle.
4. `docs/TestDriverInitialMilestones.md` — M0–M10 and the prototype success gate.
5. `history/2026-08-22-concept-assessment-swot.md` — assessment and the reasons
the first workplan reorders those milestones into a vertical spike.
6. `workplans/` — current work. Agent instructions: `AGENTS.md`.
---
## Stack
Deliberately boring, per `docs/TestDriverResearchPrototype.md`: Python, pytest,
Playwright, Pydantic/dataclasses, YAML, SQLite. One process, one database, one
browser engine, one application under test. Novelty belongs in the verification
model, not the infrastructure.

View file

@ -11,6 +11,7 @@
| workplan | TD-WP-0001 | finished | — | workplans/TD-WP-0001-statehub-bootstrap.md | | workplan | TD-WP-0001 | finished | — | workplans/TD-WP-0001-statehub-bootstrap.md |
| workplan | TD-WP-0002 | finished | — | workplans/TD-WP-0002-vertical-spike-crystallization.md | | workplan | TD-WP-0002 | finished | — | workplans/TD-WP-0002-vertical-spike-crystallization.md |
| workplan | TD-WP-0003 | blocked | — | workplans/TD-WP-0003-generalise-and-settle.md | | workplan | TD-WP-0003 | blocked | — | workplans/TD-WP-0003-generalise-and-settle.md |
| workplan | TD-WP-0004 | blocked | — | workplans/TD-WP-0004-scope-evidence-and-variants.md |
| task | TD-WP-0001-T01 | done | — | workplans/TD-WP-0001-statehub-bootstrap.md | | task | TD-WP-0001-T01 | done | — | workplans/TD-WP-0001-statehub-bootstrap.md |
| task | TD-WP-0001-T02 | done | — | workplans/TD-WP-0001-statehub-bootstrap.md | | task | TD-WP-0001-T02 | done | — | workplans/TD-WP-0001-statehub-bootstrap.md |
| task | TD-WP-0001-T03 | done | — | workplans/TD-WP-0001-statehub-bootstrap.md | | task | TD-WP-0001-T03 | done | — | workplans/TD-WP-0001-statehub-bootstrap.md |
@ -32,3 +33,8 @@
| task | TD-WP-0003-T06 | wait | — | workplans/TD-WP-0003-generalise-and-settle.md | | task | TD-WP-0003-T06 | wait | — | workplans/TD-WP-0003-generalise-and-settle.md |
| task | TD-WP-0003-T07 | wait | — | workplans/TD-WP-0003-generalise-and-settle.md | | task | TD-WP-0003-T07 | wait | — | workplans/TD-WP-0003-generalise-and-settle.md |
| task | TD-WP-0003-T08 | done | — | workplans/TD-WP-0003-generalise-and-settle.md | | task | TD-WP-0003-T08 | done | — | workplans/TD-WP-0003-generalise-and-settle.md |
| task | TD-WP-0004-T01 | done | — | workplans/TD-WP-0004-scope-evidence-and-variants.md |
| task | TD-WP-0004-T02 | done | — | workplans/TD-WP-0004-scope-evidence-and-variants.md |
| task | TD-WP-0004-T03 | done | — | workplans/TD-WP-0004-scope-evidence-and-variants.md |
| task | TD-WP-0004-T04 | done | — | workplans/TD-WP-0004-scope-evidence-and-variants.md |
| task | TD-WP-0004-T05 | wait | — | workplans/TD-WP-0004-scope-evidence-and-variants.md |

View file

@ -0,0 +1,85 @@
# Retained evidence and explicit adversarial variants
TD-WP-0004 adds two standard-library APIs. Use Python ≥3.11 from the repository;
pytest remains the only dependency of the test suite.
## Retain, load and compare runs
```bash
PYTHONPATH=src:. python3 - <<'PY'
from scenarios.alice_bob_carol import build
from testdriver import Runner
from testdriver.storage import EvidenceStore
from testdriver.classification import classify
store = EvidenceStore('/tmp/test-driver-evidence')
receipts = []
for _ in range(2):
world, driver, observer, asset, oracle = build()
result = Runner(world, driver, observer, oracle).run(asset, evidence_store=store)
receipts.append(store.load(result.run_id))
print(result.run_id, result.verdict.value)
print(classify(*receipts).classification.value) # UNCHANGED
PY
```
`EvidenceStore.write(pack)` returns its path. `load(run_id)` returns the evidence
mapping expected by classification/crystallization. The version-1 envelope wraps
evidence with a SHA-256 checksum. Files publish atomically without overwrite,
with mode 0600; newly created store directories request mode 0700. Concurrent
publication of the same run yields one winner and FileExistsError for others.
Directory fsync uses the local POSIX filesystem API. Storage failures propagate;
a caller must not report retained evidence if a write failed.
Serialization rejects unsupported values, non-string mapping keys and non-finite
numbers. Observations detach nested data when recorded, so later domain mutations
do not rewrite earlier snapshots. Receipts include the final run verdict and
asset id, parent, maturity, variant and adaptation history. Independent claim
revisions and scenario revisions bind evidence to the original run inputs.
Retention is opt-in. It covers finalized returned runs, including FAIL and guard
aborts; unexpected driver/observer exceptions or process termination do not promise
a complete receipt. The checksum detects damage, not malicious replacement with
a recomputed checksum. The store is not encryption, credential redaction, access
custody or a retention service. Select observations safe to retain before enabling
it; existing driver mechanics can contain action arguments.
## Derive a security question from the same intent
```bash
PYTHONPATH=src:. python3 - <<'PY'
from scenarios.alice_bob_carol import build
from testdriver import Runner
from testdriver.variants import substitute
world, driver, observer, parent, oracle = build()
variant = substitute(parent, variant_id='grant-to-carol', step_id='s2-grant',
arguments={'subject_id': 'carol'})
assert variant.scenario.use_case is parent.scenario.use_case
result = Runner(world, driver, observer, oracle).run(variant)
print(variant.parent_id, result.verdict.value) # FAIL: original independent claims still apply
PY
```
`substitute(..., actor_id='other')` changes the scheduled actor; that actor must
exist in the execution world's cast. `arguments={...}` changes only existing
argument keys and covers resource, tenant or privilege substitution without new
framework concepts. Mutable argument data is copied. Unknown/duplicate step ids,
unknown keys, empty actor ids and invalid variant ids are rejected. Callers choose
unique variant ids for distinct definitions; the API is not an asset registry.
The exact UseCase, claim/invariant predicates, postconditions, watches, step order
and permitted surfaces remain unchanged. Parent identity and mutation metadata
are retained without copying changed values into mutation history. A derived
variant does not become an approved new requirement, and a denial is not a default
PASS: it needs the independently authored assertions appropriate to the question.
Acceptance fingerprints include scheduled actor/action identities, argument values,
order, surface permissions and postcondition definitions. Changed definitions
require intent review, even when product verdicts happen to pass. Unsupported
runtime dependencies or old packs without this revision cannot be accepted or
crystallized automatically. Mechanical selector/endpoint choices remain outside
scenario intent and can still qualify as mechanical adaptation.
This API deliberately does not generate new assertions, skip/reorder steps,
schedule races or infer required evidence from a successful implementation run.

View file

@ -0,0 +1,107 @@
# Scope against intent — 2026-09-28 12:19:33 UTC
## Assessment basis
Reviewed the repository at `e419bfe` and the delivery in TD-WP-0004. This report
separates the baseline gaps from work implemented during this assessment. Sources:
[INTENT](../INTENT.md), [scope](../SCOPE.md), [readiness settlement](../docs/TestDriverGeneralisationReview.md),
[source](../src/testdriver/), [tests](../tests/),
[38-mutation ground truth](../lab/GROUND-TRUTH.md),
[E-001 results](../research/evidence/2026-09-28-e001.json),
and the [audit-core contract](../usecases/audit_core_e2_tenant_boundary.py).
No new production or paid-model experiment was performed.
The old SCOPE.md materially understated implementation ("implementation has not
started") while overstating the stack and lifecycle model. The repo already had
398 passing tests, three synthetic multi-user domains, qualified crystallization
and extensive acceptance guards. It had no browser engine or database, and the
Temperature/Energy/Campaign family had already been removed. SCOPE.md now describes
implemented behavior and its limits rather than restating the thesis.
## Intent coverage
| INTENT area | Demonstrated before this work | Gap / disposition |
|---|---|---|
| Independent intent, claims and oracles | Frozen Python assertions; provenance admission; S1/S2/S3 separation; missing/invalid results inconclusive; definition revisions | Strong local coverage, but provenance labels and matching collector names are not external attestation. Domain adapters and independent authors remain necessary. |
| Fluid-to-deterministic verification | Heuristic HTML discovery, stable-trajectory checks, frozen replay and one generated single-action test | Narrow demonstration, no live model, no general multi-step generator, no automatic full lifecycle. Keep claims bounded. |
| Multi-user isolation | Separate actor objects/sessions; store/canary guards; tenant and delegation fixtures | No process security boundary, actual concurrent scheduler or causal/time-window enforcement. |
| Security via use-case mutation | 38 seeded SUT mutations and hand-authored adversarial cases | No reusable scenario-side derivation at baseline. T03 delivers explicit actor/argument substitutions; skip/reorder/replay/concurrency and inferred claims are not added. |
| Durable evidence and lineage | In-memory JSON-serializable evidence, experimental JSON files, asset and parent fields | Generic runs lacked a retained receipt API and asset lineage in the pack. T02 delivers opt-in local receipts and reload, preserving existing judgment behavior. |
| Stable intent during adaptation | Claim/use-case revisions and recorded step coverage | Actor/action argument/permission/postcondition changes lacked a common intent revision. T03 adds scenario revisions and rejects changed intent for automatic acceptance/freezing. |
| Real integration / end-to-end | Local stdlib HTTP journey, direct domain adapters; audit-core contract with fixture tests | No independently validated real-system adapter, custody/cleanup execution or operational observation cost evidence. T05 explicitly waits for an approved pilot. |
| Resilience and security concurrency | Sequential mutations, expected refusals and missing-evidence checks | Not evidence of races, time-window handling or fault-tolerant external execution. Follow-on design depends on a real pilot's needs. |
| Lineage and finding improvement loop | Research findings/decisions, parent ids, generated lineage comments | No automated finding-to-asset ingestion or asset registry. T02/T03 retain basic lineage, not an orchestration service. |
| Energy/Temperature/Campaign/Retirement | Removed after generalisation found no decision-making use | Intent's superseding note governs. Deliberate de-scope, not a backlog to resurrect. |
| Scale / distributed execution | None | Explicit initial non-goals, not missing milestone work. |
## Most relevant gaps, ranked
1. **Independent real-system validation.** This most limits confidence in the
central thesis: synthetic labs do not measure the cost of independent
observation, real requirements provenance, ambiguity or false-adaptation harm.
Owner: **TD-WP-0004-T05**, waiting for an operator/target owner to select and
approve a bounded target, observation path, fixtures and cleanup/custody plan.
Implementing the existing audit-core contract without those inputs would not
produce valid evidence. This task remains live and the workplan stays blocked.
2. **Durable, reviewable run evidence and lineage.** High immediate value, small
infrastructure cost, useful to every future adapter. **Delivered in T02**:
versioned local receipts, checksum verification, no-overwrite atomic publish,
strict JSON and recorded asset ancestry/variant/final verdict. Opt-in retention
avoids silently persisting arbitrary adapter data. No signing/redaction claim.
3. **Reusable adversarial variants with unchanged semantic expectations.** This
is a direct part of INTENT's differentiation. **Delivered in T03**: explicit
actor/argument substitution in two synthetic domains, preserved claim identity,
descendant lineage and scenario-intent admission binding. It produces a test
variant, never a new inferred oracle; an expected denial still needs independently
authored expectations. A changed scenario requires review even if it passes.
4. **Measured model/browser/authoring value.** Existing records remain canonical:
**TD-WP-0003-T01** needs model, run count, spend ceiling and suite policy;
**T06** needs browser setup and T01; **T07** needs a prospectively timed fresh
independent author. This request does not supply those missing inputs. No
duplicate records or fabricated economic/authoring measurements were created.
5. **Broader scheduling and crystallization.** Actual concurrency/resilience and
general multi-step code generation could extend the thesis, but should follow
a pilot that demonstrates the requirement. They are explicitly outside this
workplan's implementation commitment, not silently declared complete. T05's
readiness review must decide which capability to fund before claiming the
broader INTENT success criterion.
## Delivered behavior and assurance limits
[TD-WP-0004](../workplans/TD-WP-0004-scope-evidence-and-variants.md) was registered
through `statehub fix-consistency`. [Evidence/variant usage](../docs/TestDriverEvidenceAndVariants.md)
provides reproducible examples. Tests verify storage roundtrips, same-run collisions
and concurrent publication, corruption/truncation, permissions, failing/aborted
runs, unsupported JSON, stable retained nested observations, preserved assertion
objects, two-domain adversarial outcomes and changed-scenario acceptance refusal.
Evidence is immutable by the store API, not tamper-proof: a party able to rewrite
both receipt and checksum can forge it. No automatic credential redaction is
promised; adapters remain responsible for retention-safe observations. Complete
receipts cover normal and guard-aborted returns, not unexpected exceptions or
process crashes. Historical packs lacking scenario revisions need rerunning for
automatic acceptance. Existing partial scenarios remain supported; revisions do
not manufacture missing assertions or evaluate unseen state.
## Readiness verdict
**Useful local research framework; not ready for autonomous real-system assurance.**
The central safety model and local maturation example exist. Durable evidence
and bounded variants improve reuse, but do not establish live-model economics,
independent production provenance, browser coverage or concurrent correctness.
The priority after local delivery is the approved independent pilot, together
with the already recorded experiment choices—not rebuilding retired concepts.
## Validation and work status
The full suite passed **428 tests in 169.03 seconds**, including 30 new delivery
regressions. An earlier existing-boundary subset passed 161 tests. Both documented
examples executed successfully; local relative links and `git diff --check` pass.
These counts describe regression coverage, not independent production trials.
TD-WP-0004-T01–T04 are complete. T05 is waiting and flagged for human input, so
TD-WP-0004 remains **blocked**. TD-WP-0003-T01/T06/T07 remain the existing external
experiment records. The scope update and local implementation do not close those
validation gaps.
Decision: `f3b35dce-6047-4a95-8083-4d7032d4e99f`.

View file

@ -13,3 +13,4 @@ file is a pointer table, not a second source of truth.
| TD-WP-0003 T08 isolation/replay follow-up | Isolation acceptance gate and step-bound frozen replay | `e8fabf6e-3e5e-424e-9a60-152bf3dec641` | `docs/TestDriverGeneralisationReview.md` | | TD-WP-0003 T08 isolation/replay follow-up | Isolation acceptance gate and step-bound frozen replay | `e8fabf6e-3e5e-424e-9a60-152bf3dec641` | `docs/TestDriverGeneralisationReview.md` |
| TD-WP-0003 T08 guard-integrity follow-up | Strict boolean judgments and intact isolation guards | `2ecac1c5-5254-4ce9-b092-b47c3e1283a9` | `docs/TestDriverGeneralisationReview.md` | | TD-WP-0003 T08 guard-integrity follow-up | Strict boolean judgments and intact isolation guards | `2ecac1c5-5254-4ce9-b092-b47c3e1283a9` | `docs/TestDriverGeneralisationReview.md` |
| TD-WP-0003 T08 HTTP/surface follow-up | Origin-bound credentials and runner surface validation | `88490ee8-839d-40dc-affa-b00a1c68e3c5` | `docs/TestDriverGeneralisationReview.md` | | TD-WP-0003 T08 HTTP/surface follow-up | Origin-bound credentials and runner surface validation | `88490ee8-839d-40dc-affa-b00a1c68e3c5` | `docs/TestDriverGeneralisationReview.md` |
| TD-WP-0004 | Evidence-backed scope, local receipts and intent-preserving variants | `f3b35dce-6047-4a95-8083-4d7032d4e99f` | `history/2026-09-28-121933-scope-intent-assessment.md` |

View file

@ -133,7 +133,8 @@ def _surface_fingerprint(pack: Mapping[str, Any]) -> str:
def _claim_fingerprint(pack: Mapping[str, Any]) -> str: def _claim_fingerprint(pack: Mapping[str, Any]) -> str:
return json.dumps({"provenance": pack["provenance_index"], return json.dumps({"provenance": pack["provenance_index"],
"revisions": pack["intent_revisions"]}, sort_keys=True) "revisions": pack["intent_revisions"],
"scenario_revision": pack["scenario_revision"]}, sort_keys=True)
def _has_complete_evidence(pack: Mapping[str, Any]) -> bool: def _has_complete_evidence(pack: Mapping[str, Any]) -> bool:
@ -189,6 +190,10 @@ def _has_complete_evidence(pack: Mapping[str, Any]) -> bool:
return False return False
if any(o["kind"] == "surface_violation" for o in observations): if any(o["kind"] == "surface_violation" for o in observations):
return False return False
revision = pack.get("scenario_revision")
if (not isinstance(revision, str) or len(revision) != 64
or any(c not in "0123456789abcdef" for c in revision)):
return False
revisions = pack["intent_revisions"] revisions = pack["intent_revisions"]
required_ids = {assertion for assertion, _ in expected_keys} | {pack["use_case_id"]} required_ids = {assertion for assertion, _ in expected_keys} | {pack["use_case_id"]}
return ( return (

View file

@ -13,7 +13,8 @@ See docs/TestDriverClassificationDesign.md, Part A.
from __future__ import annotations from __future__ import annotations
import json import json
from dataclasses import dataclass, field, asdict from copy import deepcopy
from dataclasses import dataclass, field, asdict, replace
from datetime import datetime, timezone from datetime import datetime, timezone
from enum import Enum from enum import Enum
from typing import Any from typing import Any
@ -69,8 +70,12 @@ class EvidencePack:
expected_judgments: list[dict[str, str]] = field(default_factory=list) expected_judgments: list[dict[str, str]] = field(default_factory=list)
intent_revisions: dict[str, str | None] = field(default_factory=dict) intent_revisions: dict[str, str | None] = field(default_factory=dict)
scenario_revision: str | None = None
asset: dict[str, Any] = field(default_factory=dict)
run_verdict: str | None = None
def record(self, observation: Observation) -> None: def record(self, observation: Observation) -> None:
self.observations.append(observation) self.observations.append(replace(observation, data=deepcopy(observation.data)))
def of_stratum(self, stratum: Stratum) -> list[Observation]: def of_stratum(self, stratum: Stratum) -> list[Observation]:
return [o for o in self.observations if o.stratum is stratum] return [o for o in self.observations if o.stratum is stratum]
@ -80,4 +85,15 @@ class EvidencePack:
payload["observations"] = [ payload["observations"] = [
{**asdict(o), "stratum": o.stratum.value} for o in self.observations {**asdict(o), "stratum": o.stratum.value} for o in self.observations
] ]
return json.dumps(payload, indent=2, sort_keys=True, default=str) def check_keys(value):
if isinstance(value, dict):
if any(not isinstance(key, str) for key in value):
raise TypeError("evidence objects require string keys")
for item in value.values():
check_keys(item)
elif isinstance(value, (list, tuple)):
for item in value:
check_keys(item)
check_keys(payload)
return json.dumps(payload, indent=2, sort_keys=True, allow_nan=False)

View file

@ -111,3 +111,17 @@ def intent_revisions(case: UseCase) -> dict[str, str | None]:
assertion.source_ref, getattr(assertion, "after_step", None), predicate, assertion.source_ref, getattr(assertion, "after_step", None), predicate,
]) ])
return revisions return revisions
def scenario_revision(scenario) -> str | None:
"""Bind action intent and execution order, excluding mechanical driver choices."""
try:
steps = [[step.id, step.actor_id, step.action.name,
_describe(dict(step.action.args), set()),
sorted(step.action.permitted_surfaces),
_describe(step.action.postcondition, set())]
for step in scenario.steps]
return _digest(["python-scenario-v1", sys.implementation.name,
list(sys.version_info[:3]), scenario.use_case.id, steps])
except (KeyError, TypeError, ValueError, RecursionError):
return None

View file

@ -9,6 +9,7 @@ independence claim structural rather than procedural.
from __future__ import annotations from __future__ import annotations
import uuid import uuid
from copy import deepcopy
from dataclasses import dataclass from dataclasses import dataclass
from datetime import datetime, timezone from datetime import datetime, timezone
from typing import Any from typing import Any
@ -18,7 +19,7 @@ from .drivers import Driver, realize_step
from .evidence import EvidencePack, Observation, Stratum from .evidence import EvidencePack, Observation, Stratum
from .observers import StateObserver from .observers import StateObserver
from .oracles import Judgment, Oracle, Verdict, overall from .oracles import Judgment, Oracle, Verdict, overall
from .revisions import intent_revisions from .revisions import intent_revisions, scenario_revision
from .scenario import Scenario, VerificationAsset from .scenario import Scenario, VerificationAsset
from .world import World from .world import World
@ -131,7 +132,7 @@ class Runner:
# -- execution -------------------------------------------------------- # -- execution --------------------------------------------------------
def run(self, asset: VerificationAsset) -> RunResult: def run(self, asset: VerificationAsset, *, evidence_store=None) -> RunResult:
scenario: Scenario = asset.scenario scenario: Scenario = asset.scenario
run_id = f"run-{uuid.uuid4().hex[:12]}" run_id = f"run-{uuid.uuid4().hex[:12]}"
pack = EvidencePack( pack = EvidencePack(
@ -140,6 +141,10 @@ class Runner:
use_case_id=scenario.use_case.id, use_case_id=scenario.use_case.id,
sut_version=self._world.sut_version, sut_version=self._world.sut_version,
intent_revisions=intent_revisions(scenario.use_case), intent_revisions=intent_revisions(scenario.use_case),
scenario_revision=scenario_revision(scenario),
asset={"id": asset.id, "parent_id": asset.parent_id,
"maturity": asset.maturity, "variant": scenario.variant,
"adaptation_history": deepcopy(asset.adaptation_history)},
scheduled_steps=[step.id for step in scenario.steps], scheduled_steps=[step.id for step in scenario.steps],
expected_judgments=[ expected_judgments=[
{"assertion_id": assertion.id, "step_id": step.id} {"assertion_id": assertion.id, "step_id": step.id}
@ -301,4 +306,7 @@ class Runner:
# Preserve observed failures, but a passing prefix cannot certify an abort. # Preserve observed failures, but a passing prefix cannot certify an abort.
if aborted and result_verdict is Verdict.PASS: if aborted and result_verdict is Verdict.PASS:
result_verdict = Verdict.INCONCLUSIVE result_verdict = Verdict.INCONCLUSIVE
pack.run_verdict = result_verdict.value
if evidence_store is not None:
evidence_store.write(pack)
return RunResult(run_id, result_verdict, judgments, pack) return RunResult(run_id, result_verdict, judgments, pack)

72
src/testdriver/storage.py Normal file
View file

@ -0,0 +1,72 @@
"""Local, versioned evidence receipts. Checksums detect damage, not forgery."""
import hashlib
import json
import os
from pathlib import Path
import re
import tempfile
from .evidence import EvidencePack
def _canonical(payload):
return json.dumps(payload, sort_keys=True, separators=(',', ':'), allow_nan=False).encode()
class EvidenceStore:
"""Atomically publish one immutable-by-API receipt per run, without overwrites.
Evidence must already be suitable for retention. This is not a secret scrubber,
authenticated signature, remote store or retention policy.
"""
def __init__(self, directory):
self.directory = Path(directory)
def _path(self, run_id):
if not isinstance(run_id, str) or not re.fullmatch(r'[A-Za-z0-9][A-Za-z0-9_-]{0,127}', run_id):
raise ValueError('invalid evidence run id')
return self.directory / f'{run_id}.json'
def write(self, pack: EvidencePack) -> Path:
path = self._path(pack.run_id)
if not pack.finished_at or pack.run_verdict not in ('PASS', 'FAIL', 'INCONCLUSIVE'):
raise ValueError('only finalized run evidence can be stored')
payload = json.loads(pack.to_json()) # strict serialization before any write
envelope = {'schema_version': 1, 'evidence': payload,
'sha256': hashlib.sha256(_canonical(payload)).hexdigest()}
encoded = _canonical(envelope) + b'\n'
self.directory.mkdir(mode=0o700, parents=True, exist_ok=True)
temporary = None
try:
with tempfile.NamedTemporaryFile(dir=self.directory, prefix='.evidence-', delete=False) as stream:
temporary = Path(stream.name)
stream.write(encoded)
stream.flush()
os.fsync(stream.fileno())
# link publishes atomically and fails if the destination already exists.
os.link(temporary, path)
directory_fd = os.open(self.directory, os.O_RDONLY | os.O_DIRECTORY)
try:
os.fsync(directory_fd)
finally:
os.close(directory_fd)
finally:
if temporary is not None:
temporary.unlink(missing_ok=True)
return path
def load(self, run_id) -> dict:
path = self._path(run_id)
try:
envelope = json.loads(path.read_text())
if type(envelope['schema_version']) is not int or envelope['schema_version'] != 1:
raise ValueError('unsupported evidence schema version')
payload = envelope['evidence']
if (hashlib.sha256(_canonical(payload)).hexdigest() != envelope['sha256']
or payload['run_id'] != run_id or not payload['finished_at']
or payload['run_verdict'] not in ('PASS', 'FAIL', 'INCONCLUSIVE')):
raise ValueError('invalid evidence receipt or checksum')
return payload
except (KeyError, TypeError, json.JSONDecodeError) as exc:
raise ValueError('malformed evidence receipt') from exc

View file

@ -0,0 +1,47 @@
"""Explicit adversarial substitutions preserve the independent assertion set."""
from copy import deepcopy
from dataclasses import replace
import re
from .scenario import VerificationAsset
def substitute(asset: VerificationAsset, *, variant_id: str, step_id: str,
actor_id: str | None = None, arguments: dict | None = None) -> VerificationAsset:
"""Derive one actor/resource/tenant/privilege substitution, never its verdict.
New actor identities must exist in the execution world's cast. Argument keys
must already exist on the action; domain values remain the author's choice.
Expectations, surfaces and postconditions stay exactly as independently authored.
"""
if not isinstance(variant_id, str) or not re.fullmatch(r'[A-Za-z0-9][A-Za-z0-9_-]{0,63}', variant_id):
raise ValueError('invalid variant id')
if actor_id is not None and (not isinstance(actor_id, str) or not actor_id):
raise ValueError('invalid actor id')
if arguments is not None and not isinstance(arguments, dict):
raise ValueError('argument substitutions must be a dictionary')
matching = [step for step in asset.scenario.steps if step.id == step_id]
if len(matching) != 1:
raise ValueError('substitution requires exactly one matching step')
target = matching[0]
if set(arguments or {}) - target.action.args.keys():
raise ValueError('substitution names an unknown argument')
if actor_id is None and not arguments:
raise ValueError('substitution requires an actor or argument change')
steps = []
for step in asset.scenario.steps:
args = deepcopy(dict(step.action.args))
if step.id == step_id:
args.update(deepcopy(arguments or {}))
steps.append(replace(step, actor_id=actor_id if step.id == step_id and actor_id is not None else step.actor_id,
action=replace(step.action, args=args)))
scenario = replace(asset.scenario, id=f'{asset.scenario.id}--{variant_id}',
variant=variant_id, steps=tuple(steps))
# Keep values out of lineage: mechanics/evidence may have their own retention policy.
history = deepcopy(asset.adaptation_history) + [{
'kind': 'adversarial-substitution', 'parent_id': asset.id,
'variant': variant_id, 'step_id': step_id, 'actor_changed': actor_id is not None,
'argument_keys': sorted(arguments or {}),
}]
return VerificationAsset(f'{asset.id}--{variant_id}', scenario, maturity=asset.maturity,
parent_id=asset.id, adaptation_history=history)

View file

@ -0,0 +1,199 @@
"""Intent preservation, durable receipts and reusable adversarial variants."""
from concurrent.futures import ThreadPoolExecutor
from dataclasses import replace
import json
import stat
import pytest
from scenarios.alice_bob_carol import build
from scenarios.tenant_lifecycle import build as tenant
from testdriver import Runner, Verdict
from testdriver.classification import classify, Classification
from testdriver.crystallization import assess_stability
from testdriver.evidence import Observation, Stratum
from testdriver.storage import EvidenceStore
from testdriver.variants import substitute
def run(builder=build, store=None):
world, driver, observer, asset, oracle = builder()
return Runner(world, driver, observer, oracle).run(asset, evidence_store=store)
def test_store_roundtrip_retains_lineage_and_admission(tmp_path):
store = EvidenceStore(tmp_path / 'evidence')
results = [run(store=store) for _ in range(3)]
packs = [store.load(result.run_id) for result in results]
assert packs[0] == json.loads(results[0].evidence.to_json())
assert packs[0]['asset']['id'] == build()[3].id
assert packs[0]['run_verdict'] == 'PASS'
assert classify(packs[0], packs[1]).safe_to_accept
assert assess_stability(packs).stable
assert stat.S_IMODE((store.directory / f'{results[0].run_id}.json').stat().st_mode) == 0o600
with pytest.raises(FileExistsError):
store.write(results[0].evidence)
assert not list(store.directory.glob('.evidence-*'))
def test_failed_and_guard_aborted_runs_are_retained(tmp_path):
store = EvidenceStore(tmp_path)
failing = run(lambda: build('M17'), store)
assert store.load(failing.run_id)['run_verdict'] == 'FAIL'
w, d, o, a, oracle = build()
w.cast['bob']._memory = w.cast['alice']._memory
aborted = Runner(w, d, o, oracle).run(a, evidence_store=store)
assert store.load(aborted.run_id)['run_verdict'] == 'INCONCLUSIVE'
@pytest.mark.parametrize('damage', ['checksum', 'schema', 'truncated', 'wrong-id'])
def test_corrupt_receipts_are_rejected(tmp_path, damage):
store = EvidenceStore(tmp_path)
result = run(store=store)
path = tmp_path / f'{result.run_id}.json'
envelope = json.loads(path.read_text())
if damage == 'checksum':
envelope['evidence']['run_verdict'] = 'FAIL'
elif damage == 'schema':
envelope['schema_version'] = 999
elif damage == 'wrong-id':
other = tmp_path / 'other.json'
other.write_text(path.read_text())
with pytest.raises(ValueError):
store.load('other')
return
path.write_text('{' if damage == 'truncated' else json.dumps(envelope))
with pytest.raises(ValueError):
store.load(result.run_id)
@pytest.mark.parametrize('run_id', ['../escape', '/absolute', '.', '', 'a/b'])
def test_store_rejects_unsafe_identifiers_without_writes(tmp_path, run_id):
pack = run().evidence
pack.run_id = run_id
with pytest.raises(ValueError):
EvidenceStore(tmp_path).write(pack)
assert not list(tmp_path.iterdir())
@pytest.mark.parametrize('bad', [object(), float('nan'), {1: 'coerced-key'}])
def test_unserializable_evidence_is_not_silently_stringified(tmp_path, bad):
pack = run().evidence
pack.asset['unsupported'] = bad
with pytest.raises((TypeError, ValueError)):
EvidenceStore(tmp_path).write(pack)
assert not list(tmp_path.iterdir())
def test_concurrent_writers_publish_exactly_one_complete_receipt(tmp_path):
store = EvidenceStore(tmp_path)
pack = run().evidence
def write():
try:
store.write(pack)
return True
except FileExistsError:
return False
with ThreadPoolExecutor(max_workers=4) as workers:
assert sum(workers.map(lambda _: write(), range(4))) == 1
assert store.load(pack.run_id)['run_id'] == pack.run_id
assert not list(tmp_path.glob('.evidence-*'))
def test_runner_propagates_storage_failure(tmp_path):
path = tmp_path / 'file'
path.write_text('not a directory')
with pytest.raises(FileExistsError):
run(store=EvidenceStore(path))
def test_recorded_observations_do_not_change_with_live_nested_objects():
pack = run().evidence
snapshot = {'nested': {'values': [1]}}
pack.record(Observation('custom', Stratum.JUDGMENT, 'observer', None, 'state', snapshot))
snapshot['nested']['values'].append(2)
assert pack.observations[-1].data == {'nested': {'values': [1]}}
def test_variant_preserves_claims_surfaces_and_parent_and_copies_arguments(tmp_path):
world, driver, observer, parent, oracle = build()
step = parent.scenario.steps[1]
variant = substitute(parent, variant_id='carol-read', step_id=step.id,
arguments={'subject_id': 'carol'})
assert variant.parent_id == parent.id
assert variant.scenario.use_case is parent.scenario.use_case
for original, changed in zip(parent.scenario.steps, variant.scenario.steps):
assert original.action.postcondition is changed.action.postcondition
assert original.action.permitted_surfaces == changed.action.permitted_surfaces
assert original.action.args is not changed.action.args
assert step.action.args['subject_id'] == 'bob'
result = Runner(world, driver, observer, oracle).run(variant, evidence_store=EvidenceStore(tmp_path))
assert result.verdict is Verdict.FAIL
assert result.evidence.asset['parent_id'] == parent.id
assert result.judgment('c-carol-denied').verdict is Verdict.FAIL
def test_actor_substitution_applies_to_another_domain_without_new_claims():
world, driver, observer, parent, oracle = tenant()
variant = substitute(parent, variant_id='foreign-creator', step_id='create-a', actor_id='admin-b')
result = Runner(world, driver, observer, oracle).run(variant)
assert result.judgment('tenant-create-a').verdict is Verdict.FAIL
assert variant.scenario.use_case is parent.scenario.use_case
@pytest.mark.parametrize('changes', [
{'step_id': 'absent', 'actor_id': 'bob'},
{'step_id': 's2-grant', 'arguments': {'absent': 1}},
{'step_id': 's2-grant', 'actor_id': ''},
{'step_id': 's2-grant'},
])
def test_invalid_substitutions_do_not_modify_parent(changes):
parent = build()[3]
before = repr(parent)
with pytest.raises(ValueError):
substitute(parent, variant_id='invalid', **changes)
assert repr(parent) == before
@pytest.mark.parametrize('change', ['actor', 'argument', 'surface', 'postcondition', 'order'])
def test_scenario_intent_changes_require_review_even_with_passing_verdicts(change):
baseline = json.loads(run().evidence.to_json())
world, driver, observer, asset, oracle = build()
steps = list(asset.scenario.steps)
step = steps[0]
if change == 'actor':
# Identical authorized operation via another identity, without changing claim definitions.
driver._tokens['carol'] = driver._tokens['alice']
steps[0] = replace(step, actor_id='carol')
elif change == 'argument':
steps[0] = replace(step, action=replace(step.action, args={**step.action.args, 'content': 'changed'}))
elif change == 'surface':
steps[0] = replace(step, action=replace(step.action, permitted_surfaces=frozenset({'api', 'browser'})))
elif change == 'postcondition':
steps[0] = replace(step, action=replace(step.action, postcondition=lambda obs: True))
else:
# A schedule-only assertion-free no-op pair can be reordered without changing outcomes.
steps.extend([replace(step, id='extra-a'), replace(step, id='extra-b')])
steps[-2:] = reversed(steps[-2:])
asset.scenario = replace(asset.scenario, steps=tuple(steps))
result = Runner(world, driver, observer, oracle).run(asset)
pack = json.loads(result.evidence.to_json())
outcome = classify(baseline, pack)
if change != 'order':
assert result.verdict is Verdict.PASS
assert outcome.classification is Classification.INTENT_CHANGED
assert not outcome.safe_to_accept
assert outcome.classification in (Classification.INTENT_CHANGED, Classification.AMBIGUOUS)
def test_legacy_packs_require_fresh_scenario_intent_evidence():
pack = json.loads(run().evidence.to_json())
del pack['scenario_revision']
assert not classify(pack, pack).safe_to_accept
def test_action_order_is_part_of_scenario_revision():
from testdriver.revisions import scenario_revision
scenario = build()[3].scenario
assert scenario_revision(scenario) != scenario_revision(
replace(scenario, steps=tuple(reversed(scenario.steps))))

View file

@ -0,0 +1,134 @@
---
id: TD-WP-0004
type: workplan
title: "Align scope with evidence and close durable-run and variant gaps"
domain: infotech
repo: test-driver
status: blocked
owner: codex
topic_slug: custodian
created: "2026-09-28"
updated: "2026-09-28"
state_hub_workstream_id: "efd2200b-4856-5655-a010-0c3b3a858391"
---
# Scope, retained evidence and adversarial variants
User-authorized assessment and implementation following the scope review at
`e419bfe`. No new dependencies. Keep the Python intent API, deterministic oracles,
sequential execution and standard-library HTTP. Do not revive the lifecycle
concepts retired by TD-WP-0003. Evidence is local and must not imply production
readiness or proof of arbitrary actor isolation.
## Assess actual scope against intent
```task
id: TD-WP-0004-T01
status: done
priority: high
state_hub_task_id: "1c75b826-b6de-5aed-bd8b-40ea4217cdcd"
```
Replace stale SCOPE.md claims with implementation-backed capabilities, limitations
and actual stack. Write `history/2026-09-28-121933-scope-intent-assessment.md` with
an intent/capability matrix, ranked gaps, evidence links and live work ownership.
Distinguish missing capabilities from deliberately removed concepts.
## Retain complete run evidence and asset lineage
```task
id: TD-WP-0004-T02
status: done
priority: high
state_hub_task_id: "e6c957db-e53d-5bcf-957b-ff0fac4c2e77"
```
Add an opt-in local evidence store with strict JSON, schema version, safe run-id
filenames, atomic no-overwrite publication, restrictive file permissions, reload
and tamper/corruption detection. Record asset id, parent, maturity and variant in
run evidence. Runner can persist completed/guard-aborted runs on request; storage
failure must be explicit. Do not serialize arbitrary Python objects as strings,
or claim storage is encryption, redaction, signing or a replacement for custody.
Validate roundtrip, collisions, corruption, unsupported data and failing runs.
## Derive bounded adversarial variants without rewriting claims
```task
id: TD-WP-0004-T03
status: done
priority: high
state_hub_task_id: "7ce7ca92-7fff-5d2b-8482-db64049c1429"
```
Provide reusable actor and argument substitution (including resource, tenant and
privilege values) for one scheduled step. Retain exact UseCase, predicates,
postconditions, schedule and permitted surfaces; copy mutable argument data and
record descendant identity/parent/mutation history. Reject unknown steps/keys and
invalid identities. No automatic discovery, concurrent scheduler, skip/reorder
or inferred security assertions.
Bind acceptance to a conservative revision of scheduled action intent (actor,
order, args, permitted surfaces and postcondition). Changed/unsupported scenario
intent must not be accepted/frozen as a mechanical change. Verify controls and
adversarial behavior against at least two synthetic domains, without modifying
lab implementations to manufacture the result.
## Publish reproducible usage and validate the implemented scope
```task
id: TD-WP-0004-T04
status: done
priority: medium
state_hub_task_id: "4d0be7bb-5c76-5d68-972c-97cda53e0621"
```
Document executable evidence-store and variant examples, retained-data limits and
new admission requirements. Run targeted regressions and the full suite. Update
the assessment with delivered capability, remaining gaps and validation counts.
Register/sync records and commit/push implementation and documents.
## Validate an independently owned real-system pilot
```task
id: TD-WP-0004-T05
status: wait
priority: high
state_hub_task_id: "aeacea6b-20db-5dd6-9585-4d7fefe412a2"
```
Blocked pending an operator/target-owner-selected system and approved bounded
engagement: independent requirements, observation adapter and cost, test fixture
and cleanup authority, custody/expiry routing and a recorded tolerance for false
adaptation. No production target or credential access is authorized by this
repo-local request. Audit-core E2 is a contract/calibration candidate, not an
executable pilot. When prerequisites exist, implement and execute its bounded
adapter and retain independent results; until then this task/workplan stays
blocked after local completion. Owner of selection/approval: Bernd Worsch and
the target owner.
Existing external work remains in TD-WP-0003-T01 (model/run/budget choices), T06
(browser setup and T01) and T07 (fresh independent timed authoring). Do not
duplicate those tasks. Longer-term concurrency, resilience fault scheduling and
multi-step artifact generation are follow-on decisions informed by the pilot,
not implementation commitments in this bounded workplan.
## Local implementation closeout — 2026-09-28
T01–T04 are done. SCOPE.md now reflects executable capability and actual stack;
`history/2026-09-28-121933-scope-intent-assessment.md` ranks the intent gaps and
records the delivered changes. EvidenceStore, detached/strict evidence with
lineage, reusable substitutions and scenario-definition admission binding are
implemented. Usage examples ran successfully; relative document links resolve.
Validation: 30 new delivery regressions passed; the existing acceptance,
classification, generalisation and completeness subset passed 161 tests. The
final full suite passed **428 tests in 169.03 seconds**. `git diff --check` is
clean. Decision: `f3b35dce-6047-4a95-8083-4d7032d4e99f`.
T05 remains **wait**, flagged for human input with target/engagement prerequisites;
therefore this workplan is **blocked**, not finished. No residual is hidden in
scope prose: the pilot remains live here, and model/browser/authoring experiments
remain TD-WP-0003-T01/T06/T07. Broader scheduling/lifecycle/adapter features remain
explicit scope limits to prioritize from actual pilot needs. No real-system,
paid-model or browser-engine execution was performed.