Assistant: codex Assistant-Model: gpt-6-astra Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
4.3 KiB
Retained evidence and explicit adversarial variants
TD-WP-0004 adds two standard-library APIs. Use Python ≥3.11 from the repository; pytest remains the only dependency of the test suite.
Retain, load and compare runs
PYTHONPATH=src:. python3 - <<'PY'
from scenarios.alice_bob_carol import build
from testdriver import Runner
from testdriver.storage import EvidenceStore
from testdriver.classification import classify
store = EvidenceStore('/tmp/test-driver-evidence')
receipts = []
for _ in range(2):
world, driver, observer, asset, oracle = build()
result = Runner(world, driver, observer, oracle).run(asset, evidence_store=store)
receipts.append(store.load(result.run_id))
print(result.run_id, result.verdict.value)
print(classify(*receipts).classification.value) # UNCHANGED
PY
EvidenceStore.write(pack) returns its path. load(run_id) returns the evidence
mapping expected by classification/crystallization. The version-1 envelope wraps
evidence with a SHA-256 checksum. Files publish atomically without overwrite,
with mode 0600; newly created store directories request mode 0700. Concurrent
publication of the same run yields one winner and FileExistsError for others.
Directory fsync uses the local POSIX filesystem API. Storage failures propagate;
a caller must not report retained evidence if a write failed.
Serialization rejects unsupported values, non-string mapping keys and non-finite numbers. Observations detach nested data when recorded, so later domain mutations do not rewrite earlier snapshots. Receipts include the final run verdict and asset id, parent, maturity, variant and adaptation history. Independent claim revisions and scenario revisions bind evidence to the original run inputs.
Retention is opt-in. It covers finalized returned runs, including FAIL and guard aborts; unexpected driver/observer exceptions or process termination do not promise a complete receipt. The checksum detects damage, not malicious replacement with a recomputed checksum. The store is not encryption, credential redaction, access custody or a retention service. Select observations safe to retain before enabling it; existing driver mechanics can contain action arguments.
Derive a security question from the same intent
PYTHONPATH=src:. python3 - <<'PY'
from scenarios.alice_bob_carol import build
from testdriver import Runner
from testdriver.variants import substitute
world, driver, observer, parent, oracle = build()
variant = substitute(parent, variant_id='grant-to-carol', step_id='s2-grant',
arguments={'subject_id': 'carol'})
assert variant.scenario.use_case is parent.scenario.use_case
result = Runner(world, driver, observer, oracle).run(variant)
print(variant.parent_id, result.verdict.value) # FAIL: original independent claims still apply
PY
substitute(..., actor_id='other') changes the scheduled actor; that actor must
exist in the execution world's cast. arguments={...} changes only existing
argument keys and covers resource, tenant or privilege substitution without new
framework concepts. Mutable argument data is copied. Unknown/duplicate step ids,
unknown keys, empty actor ids and invalid variant ids are rejected. Callers choose
unique variant ids for distinct definitions; the API is not an asset registry.
The exact UseCase, claim/invariant predicates, postconditions, watches, step order and permitted surfaces remain unchanged. Parent identity and mutation metadata are retained without copying changed values into mutation history. A derived variant does not become an approved new requirement, and a denial is not a default PASS: it needs the independently authored assertions appropriate to the question.
Acceptance fingerprints include scheduled actor/action identities, argument values, order, surface permissions and postcondition definitions. Changed definitions require intent review, even when product verdicts happen to pass. Unsupported runtime dependencies or old packs without this revision cannot be accepted or crystallized automatically. Mechanical selector/endpoint choices remain outside scenario intent and can still qualify as mechanical adaptation.
This API deliberately does not generate new assertions, skip/reorder steps, schedule races or infer required evidence from a successful implementation run.