test-driver/workplans/TD-WP-0004-scope-evidence-and-variants.md
tegwick 10077edb8c Reject lossy observations and validate schedules before execution
Assistant: codex
Assistant-Model: gpt-6-astra
Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
2026-09-28 14:50:27 +02:00

149 lines
6.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
id: TD-WP-0004
type: workplan
title: "Align scope with evidence and close durable-run and variant gaps"
domain: infotech
repo: test-driver
status: blocked
owner: codex
topic_slug: custodian
created: "2026-09-28"
updated: "2026-09-28"
state_hub_workstream_id: "efd2200b-4856-5655-a010-0c3b3a858391"
---
# Scope, retained evidence and adversarial variants
User-authorized assessment and implementation following the scope review at
`e419bfe`. No new dependencies. Keep the Python intent API, deterministic oracles,
sequential execution and standard-library HTTP. Do not revive the lifecycle
concepts retired by TD-WP-0003. Evidence is local and must not imply production
readiness or proof of arbitrary actor isolation.
## Assess actual scope against intent
```task
id: TD-WP-0004-T01
status: done
priority: high
state_hub_task_id: "1c75b826-b6de-5aed-bd8b-40ea4217cdcd"
```
Replace stale SCOPE.md claims with implementation-backed capabilities, limitations
and actual stack. Write `history/2026-09-28-121933-scope-intent-assessment.md` with
an intent/capability matrix, ranked gaps, evidence links and live work ownership.
Distinguish missing capabilities from deliberately removed concepts.
## Retain complete run evidence and asset lineage
```task
id: TD-WP-0004-T02
status: done
priority: high
state_hub_task_id: "e6c957db-e53d-5bcf-957b-ff0fac4c2e77"
```
Add an opt-in local evidence store with strict JSON, schema version, safe run-id
filenames, atomic no-overwrite publication, restrictive file permissions, reload
and tamper/corruption detection. Record asset id, parent, maturity and variant in
run evidence. Runner can persist completed/guard-aborted runs on request; storage
failure must be explicit. Do not serialize arbitrary Python objects as strings,
or claim storage is encryption, redaction, signing or a replacement for custody.
Validate roundtrip, collisions, corruption, unsupported data and failing runs.
## Derive bounded adversarial variants without rewriting claims
```task
id: TD-WP-0004-T03
status: done
priority: high
state_hub_task_id: "7ce7ca92-7fff-5d2b-8482-db64049c1429"
```
Provide reusable actor and argument substitution (including resource, tenant and
privilege values) for one scheduled step. Retain exact UseCase, predicates,
postconditions, schedule and permitted surfaces; copy mutable argument data and
record descendant identity/parent/mutation history. Reject unknown steps/keys and
invalid identities. No automatic discovery, concurrent scheduler, skip/reorder
or inferred security assertions.
Bind acceptance to a conservative revision of scheduled action intent (actor,
order, args, permitted surfaces and postcondition). Changed/unsupported scenario
intent must not be accepted/frozen as a mechanical change. Verify controls and
adversarial behavior against at least two synthetic domains, without modifying
lab implementations to manufacture the result.
## Publish reproducible usage and validate the implemented scope
```task
id: TD-WP-0004-T04
status: done
priority: medium
state_hub_task_id: "4d0be7bb-5c76-5d68-972c-97cda53e0621"
```
Document executable evidence-store and variant examples, retained-data limits and
new admission requirements. Run targeted regressions and the full suite. Update
the assessment with delivered capability, remaining gaps and validation counts.
Register/sync records and commit/push implementation and documents.
## Validate an independently owned real-system pilot
```task
id: TD-WP-0004-T05
status: wait
priority: high
state_hub_task_id: "aeacea6b-20db-5dd6-9585-4d7fefe412a2"
```
Blocked pending an operator/target-owner-selected system and approved bounded
engagement: independent requirements, observation adapter and cost, test fixture
and cleanup authority, custody/expiry routing and a recorded tolerance for false
adaptation. No production target or credential access is authorized by this
repo-local request. Audit-core E2 is a contract/calibration candidate, not an
executable pilot. When prerequisites exist, implement and execute its bounded
adapter and retain independent results; until then this task/workplan stays
blocked after local completion. Owner of selection/approval: Bernd Worsch and
the target owner.
Existing external work remains in TD-WP-0003-T01 (model/run/budget choices), T06
(browser setup and T01) and T07 (fresh independent timed authoring). Do not
duplicate those tasks. Longer-term concurrency, resilience fault scheduling and
multi-step artifact generation are follow-on decisions informed by the pilot,
not implementation commitments in this bounded workplan.
## Local implementation closeout — 2026-09-28
T01–T04 are done. SCOPE.md now reflects executable capability and actual stack;
`history/2026-09-28-121933-scope-intent-assessment.md` ranks the intent gaps and
records the delivered changes. EvidenceStore, detached/strict evidence with
lineage, reusable substitutions and scenario-definition admission binding are
implemented. Usage examples ran successfully; relative document links resolve.
Validation: 30 new delivery regressions passed; the existing acceptance,
classification, generalisation and completeness subset passed 161 tests. The
final full suite passed **428 tests in 169.03 seconds**. `git diff --check` is
clean. Decision: `f3b35dce-6047-4a95-8083-4d7032d4e99f`.
T05 remains **wait**, flagged for human input with target/engagement prerequisites;
therefore this workplan is **blocked**, not finished. No residual is hidden in
scope prose: the pilot remains live here, and model/browser/authoring experiments
remain TD-WP-0003-T01/T06/T07. Broader scheduling/lifecycle/adapter features remain
explicit scope limits to prioritize from actual pilot needs. No real-system,
paid-model or browser-engine execution was performed.
**2026-09-28 evidence/preflight follow-up — done (T02/T03).** Runner validates
JSON-native observations before evaluation/retention. Lossy types (including
tuples), invalid keys, non-finite numbers and cyclic data cannot silently change
meaning in JSON; a bad snapshot stops execution with an evidence-failure record
and pending assertions INCONCLUSIVE. Serialization uses the same strict value
contract without dataclass or custom-value coercion. Runner preflights every
scheduled actor, unique nonempty step id and cast key/id agreement before any
SUT action. Intentionally partial scenarios remain valid.
Validation: 12 new regressions, of which 10 failed before the fixes; **131 focused
tests passed**, and the full suite passed **440 tests in 169.07 seconds**.
`git diff --check` is clean. Decision: `aa6c3e31-614c-4030-990c-945acef08515`.
No new tasks/workplans; T05 remains waiting and this workplan remains blocked.