Assistant: codex Assistant-Model: gpt-6-astra Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
149 lines
6.6 KiB
Markdown
149 lines
6.6 KiB
Markdown
---
|
||
id: TD-WP-0004
|
||
type: workplan
|
||
title: "Align scope with evidence and close durable-run and variant gaps"
|
||
domain: infotech
|
||
repo: test-driver
|
||
status: blocked
|
||
owner: codex
|
||
topic_slug: custodian
|
||
created: "2026-09-28"
|
||
updated: "2026-09-28"
|
||
state_hub_workstream_id: "efd2200b-4856-5655-a010-0c3b3a858391"
|
||
---
|
||
|
||
# Scope, retained evidence and adversarial variants
|
||
|
||
User-authorized assessment and implementation following the scope review at
|
||
`e419bfe`. No new dependencies. Keep the Python intent API, deterministic oracles,
|
||
sequential execution and standard-library HTTP. Do not revive the lifecycle
|
||
concepts retired by TD-WP-0003. Evidence is local and must not imply production
|
||
readiness or proof of arbitrary actor isolation.
|
||
|
||
## Assess actual scope against intent
|
||
|
||
```task
|
||
id: TD-WP-0004-T01
|
||
status: done
|
||
priority: high
|
||
state_hub_task_id: "1c75b826-b6de-5aed-bd8b-40ea4217cdcd"
|
||
```
|
||
|
||
Replace stale SCOPE.md claims with implementation-backed capabilities, limitations
|
||
and actual stack. Write `history/2026-09-28-121933-scope-intent-assessment.md` with
|
||
an intent/capability matrix, ranked gaps, evidence links and live work ownership.
|
||
Distinguish missing capabilities from deliberately removed concepts.
|
||
|
||
## Retain complete run evidence and asset lineage
|
||
|
||
```task
|
||
id: TD-WP-0004-T02
|
||
status: done
|
||
priority: high
|
||
state_hub_task_id: "e6c957db-e53d-5bcf-957b-ff0fac4c2e77"
|
||
```
|
||
|
||
Add an opt-in local evidence store with strict JSON, schema version, safe run-id
|
||
filenames, atomic no-overwrite publication, restrictive file permissions, reload
|
||
and tamper/corruption detection. Record asset id, parent, maturity and variant in
|
||
run evidence. Runner can persist completed/guard-aborted runs on request; storage
|
||
failure must be explicit. Do not serialize arbitrary Python objects as strings,
|
||
or claim storage is encryption, redaction, signing or a replacement for custody.
|
||
Validate roundtrip, collisions, corruption, unsupported data and failing runs.
|
||
|
||
## Derive bounded adversarial variants without rewriting claims
|
||
|
||
```task
|
||
id: TD-WP-0004-T03
|
||
status: done
|
||
priority: high
|
||
state_hub_task_id: "7ce7ca92-7fff-5d2b-8482-db64049c1429"
|
||
```
|
||
|
||
Provide reusable actor and argument substitution (including resource, tenant and
|
||
privilege values) for one scheduled step. Retain exact UseCase, predicates,
|
||
postconditions, schedule and permitted surfaces; copy mutable argument data and
|
||
record descendant identity/parent/mutation history. Reject unknown steps/keys and
|
||
invalid identities. No automatic discovery, concurrent scheduler, skip/reorder
|
||
or inferred security assertions.
|
||
|
||
Bind acceptance to a conservative revision of scheduled action intent (actor,
|
||
order, args, permitted surfaces and postcondition). Changed/unsupported scenario
|
||
intent must not be accepted/frozen as a mechanical change. Verify controls and
|
||
adversarial behavior against at least two synthetic domains, without modifying
|
||
lab implementations to manufacture the result.
|
||
|
||
## Publish reproducible usage and validate the implemented scope
|
||
|
||
```task
|
||
id: TD-WP-0004-T04
|
||
status: done
|
||
priority: medium
|
||
state_hub_task_id: "4d0be7bb-5c76-5d68-972c-97cda53e0621"
|
||
```
|
||
|
||
Document executable evidence-store and variant examples, retained-data limits and
|
||
new admission requirements. Run targeted regressions and the full suite. Update
|
||
the assessment with delivered capability, remaining gaps and validation counts.
|
||
Register/sync records and commit/push implementation and documents.
|
||
|
||
## Validate an independently owned real-system pilot
|
||
|
||
```task
|
||
id: TD-WP-0004-T05
|
||
status: wait
|
||
priority: high
|
||
state_hub_task_id: "aeacea6b-20db-5dd6-9585-4d7fefe412a2"
|
||
```
|
||
|
||
Blocked pending an operator/target-owner-selected system and approved bounded
|
||
engagement: independent requirements, observation adapter and cost, test fixture
|
||
and cleanup authority, custody/expiry routing and a recorded tolerance for false
|
||
adaptation. No production target or credential access is authorized by this
|
||
repo-local request. Audit-core E2 is a contract/calibration candidate, not an
|
||
executable pilot. When prerequisites exist, implement and execute its bounded
|
||
adapter and retain independent results; until then this task/workplan stays
|
||
blocked after local completion. Owner of selection/approval: Bernd Worsch and
|
||
the target owner.
|
||
|
||
Existing external work remains in TD-WP-0003-T01 (model/run/budget choices), T06
|
||
(browser setup and T01) and T07 (fresh independent timed authoring). Do not
|
||
duplicate those tasks. Longer-term concurrency, resilience fault scheduling and
|
||
multi-step artifact generation are follow-on decisions informed by the pilot,
|
||
not implementation commitments in this bounded workplan.
|
||
|
||
|
||
## Local implementation closeout — 2026-09-28
|
||
|
||
T01–T04 are done. SCOPE.md now reflects executable capability and actual stack;
|
||
`history/2026-09-28-121933-scope-intent-assessment.md` ranks the intent gaps and
|
||
records the delivered changes. EvidenceStore, detached/strict evidence with
|
||
lineage, reusable substitutions and scenario-definition admission binding are
|
||
implemented. Usage examples ran successfully; relative document links resolve.
|
||
|
||
Validation: 30 new delivery regressions passed; the existing acceptance,
|
||
classification, generalisation and completeness subset passed 161 tests. The
|
||
final full suite passed **428 tests in 169.03 seconds**. `git diff --check` is
|
||
clean. Decision: `f3b35dce-6047-4a95-8083-4d7032d4e99f`.
|
||
|
||
T05 remains **wait**, flagged for human input with target/engagement prerequisites;
|
||
therefore this workplan is **blocked**, not finished. No residual is hidden in
|
||
scope prose: the pilot remains live here, and model/browser/authoring experiments
|
||
remain TD-WP-0003-T01/T06/T07. Broader scheduling/lifecycle/adapter features remain
|
||
explicit scope limits to prioritize from actual pilot needs. No real-system,
|
||
paid-model or browser-engine execution was performed.
|
||
|
||
|
||
**2026-09-28 evidence/preflight follow-up — done (T02/T03).** Runner validates
|
||
JSON-native observations before evaluation/retention. Lossy types (including
|
||
tuples), invalid keys, non-finite numbers and cyclic data cannot silently change
|
||
meaning in JSON; a bad snapshot stops execution with an evidence-failure record
|
||
and pending assertions INCONCLUSIVE. Serialization uses the same strict value
|
||
contract without dataclass or custom-value coercion. Runner preflights every
|
||
scheduled actor, unique nonempty step id and cast key/id agreement before any
|
||
SUT action. Intentionally partial scenarios remain valid.
|
||
|
||
Validation: 12 new regressions, of which 10 failed before the fixes; **131 focused
|
||
tests passed**, and the full suite passed **440 tests in 169.07 seconds**.
|
||
`git diff --check` is clean. Decision: `aa6c3e31-614c-4030-990c-945acef08515`.
|
||
No new tasks/workplans; T05 remains waiting and this workplan remains blocked.
|