--- id: TD-WP-0004 type: workplan title: "Align scope with evidence and close durable-run and variant gaps" domain: infotech repo: test-driver status: blocked owner: codex topic_slug: custodian created: "2026-09-28" updated: "2026-09-28" state_hub_workstream_id: "efd2200b-4856-5655-a010-0c3b3a858391" --- # Scope, retained evidence and adversarial variants User-authorized assessment and implementation following the scope review at `e419bfe`. No new dependencies. Keep the Python intent API, deterministic oracles, sequential execution and standard-library HTTP. Do not revive the lifecycle concepts retired by TD-WP-0003. Evidence is local and must not imply production readiness or proof of arbitrary actor isolation. ## Assess actual scope against intent ```task id: TD-WP-0004-T01 status: done priority: high state_hub_task_id: "1c75b826-b6de-5aed-bd8b-40ea4217cdcd" ``` Replace stale SCOPE.md claims with implementation-backed capabilities, limitations and actual stack. Write `history/2026-09-28-121933-scope-intent-assessment.md` with an intent/capability matrix, ranked gaps, evidence links and live work ownership. Distinguish missing capabilities from deliberately removed concepts. ## Retain complete run evidence and asset lineage ```task id: TD-WP-0004-T02 status: done priority: high state_hub_task_id: "e6c957db-e53d-5bcf-957b-ff0fac4c2e77" ``` Add an opt-in local evidence store with strict JSON, schema version, safe run-id filenames, atomic no-overwrite publication, restrictive file permissions, reload and tamper/corruption detection. Record asset id, parent, maturity and variant in run evidence. Runner can persist completed/guard-aborted runs on request; storage failure must be explicit. Do not serialize arbitrary Python objects as strings, or claim storage is encryption, redaction, signing or a replacement for custody. Validate roundtrip, collisions, corruption, unsupported data and failing runs. ## Derive bounded adversarial variants without rewriting claims ```task id: TD-WP-0004-T03 status: done priority: high state_hub_task_id: "7ce7ca92-7fff-5d2b-8482-db64049c1429" ``` Provide reusable actor and argument substitution (including resource, tenant and privilege values) for one scheduled step. Retain exact UseCase, predicates, postconditions, schedule and permitted surfaces; copy mutable argument data and record descendant identity/parent/mutation history. Reject unknown steps/keys and invalid identities. No automatic discovery, concurrent scheduler, skip/reorder or inferred security assertions. Bind acceptance to a conservative revision of scheduled action intent (actor, order, args, permitted surfaces and postcondition). Changed/unsupported scenario intent must not be accepted/frozen as a mechanical change. Verify controls and adversarial behavior against at least two synthetic domains, without modifying lab implementations to manufacture the result. ## Publish reproducible usage and validate the implemented scope ```task id: TD-WP-0004-T04 status: done priority: medium state_hub_task_id: "4d0be7bb-5c76-5d68-972c-97cda53e0621" ``` Document executable evidence-store and variant examples, retained-data limits and new admission requirements. Run targeted regressions and the full suite. Update the assessment with delivered capability, remaining gaps and validation counts. Register/sync records and commit/push implementation and documents. ## Validate an independently owned real-system pilot ```task id: TD-WP-0004-T05 status: wait priority: high state_hub_task_id: "aeacea6b-20db-5dd6-9585-4d7fefe412a2" ``` Blocked pending an operator/target-owner-selected system and approved bounded engagement: independent requirements, observation adapter and cost, test fixture and cleanup authority, custody/expiry routing and a recorded tolerance for false adaptation. No production target or credential access is authorized by this repo-local request. Audit-core E2 is a contract/calibration candidate, not an executable pilot. When prerequisites exist, implement and execute its bounded adapter and retain independent results; until then this task/workplan stays blocked after local completion. Owner of selection/approval: Bernd Worsch and the target owner. Existing external work remains in TD-WP-0003-T01 (model/run/budget choices), T06 (browser setup and T01) and T07 (fresh independent timed authoring). Do not duplicate those tasks. Longer-term concurrency, resilience fault scheduling and multi-step artifact generation are follow-on decisions informed by the pilot, not implementation commitments in this bounded workplan. ## Local implementation closeout — 2026-09-28 T01–T04 are done. SCOPE.md now reflects executable capability and actual stack; `history/2026-09-28-121933-scope-intent-assessment.md` ranks the intent gaps and records the delivered changes. EvidenceStore, detached/strict evidence with lineage, reusable substitutions and scenario-definition admission binding are implemented. Usage examples ran successfully; relative document links resolve. Validation: 30 new delivery regressions passed; the existing acceptance, classification, generalisation and completeness subset passed 161 tests. The final full suite passed **428 tests in 169.03 seconds**. `git diff --check` is clean. Decision: `f3b35dce-6047-4a95-8083-4d7032d4e99f`. T05 remains **wait**, flagged for human input with target/engagement prerequisites; therefore this workplan is **blocked**, not finished. No residual is hidden in scope prose: the pilot remains live here, and model/browser/authoring experiments remain TD-WP-0003-T01/T06/T07. Broader scheduling/lifecycle/adapter features remain explicit scope limits to prioritize from actual pilot needs. No real-system, paid-model or browser-engine execution was performed. **2026-09-28 evidence/preflight follow-up — done (T02/T03).** Runner validates JSON-native observations before evaluation/retention. Lossy types (including tuples), invalid keys, non-finite numbers and cyclic data cannot silently change meaning in JSON; a bad snapshot stops execution with an evidence-failure record and pending assertions INCONCLUSIVE. Serialization uses the same strict value contract without dataclass or custom-value coercion. Runner preflights every scheduled actor, unique nonempty step id and cast key/id agreement before any SUT action. Intentionally partial scenarios remain valid. Validation: 12 new regressions, of which 10 failed before the fixes; **131 focused tests passed**, and the full suite passed **440 tests in 169.07 seconds**. `git diff --check` is clean. Decision: `aa6c3e31-614c-4030-990c-945acef08515`. No new tasks/workplans; T05 remains waiting and this workplan remains blocked. **2026-09-28 generated-judgment follow-up — done (T04).** Generated regression modules now embed the actual predicate evaluator used by Oracle, retaining strict boolean and missing/exception/invalid-result INCONCLUSIVE semantics. The adapter evaluates all protected assertions so FAIL dominates any INCONCLUSIVE result; otherwise an inconclusive result becomes a stdlib SkipTest explicitly labeled INCONCLUSIVE in pytest. Callers must inspect skipped outcomes, not treat exit zero as complete verification. The checked-in descendant was regenerated with its original predicate set, frozen route and lineage. No runtime test-driver import or additional dependency is introduced into generated artifacts. Validation: the initial 30 parity cases reproduced 24 failures and six controls; all 31 final parity cases pass, including native pytest reporting in a subprocess. The related subset passed 93 tests. Full suite: **471 passed in 167.82 seconds**. `git diff --check` is clean. Decision: `c06c8752-80db-4458-9eaf-6321a0ca710c`. No new task/workplan; T05 remains waiting and this workplan remains blocked.