Assistant: codex Assistant-Model: gpt-6-astra Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
113 lines
6 KiB
Markdown
113 lines
6 KiB
Markdown
# Retained evidence and explicit adversarial variants
|
|
|
|
TD-WP-0004 adds two standard-library APIs. Use Python ≥3.11 from the repository;
|
|
pytest remains the only dependency of the test suite.
|
|
|
|
## Retain, load and compare runs
|
|
|
|
```bash
|
|
PYTHONPATH=src:. python3 - <<'PY'
|
|
from scenarios.alice_bob_carol import build
|
|
from testdriver import Runner
|
|
from testdriver.storage import EvidenceStore
|
|
from testdriver.classification import classify
|
|
|
|
store = EvidenceStore('/tmp/test-driver-evidence')
|
|
receipts = []
|
|
for _ in range(2):
|
|
world, driver, observer, asset, oracle = build()
|
|
result = Runner(world, driver, observer, oracle).run(asset, evidence_store=store)
|
|
receipts.append(store.load(result.run_id))
|
|
print(result.run_id, result.verdict.value)
|
|
print(classify(*receipts).classification.value) # UNCHANGED
|
|
PY
|
|
```
|
|
|
|
`EvidenceStore.write(pack)` returns its path. `load(run_id)` returns the evidence
|
|
mapping expected by classification/crystallization. The version-1 envelope wraps
|
|
evidence with a SHA-256 checksum. Files publish atomically without overwrite,
|
|
with mode 0600; newly created store directories request mode 0700. Concurrent
|
|
publication of the same run yields one winner and FileExistsError for others.
|
|
Directory fsync uses the local POSIX filesystem API. Storage failures propagate;
|
|
a caller must not report retained evidence if a write failed.
|
|
|
|
Observations must use JSON-native dictionaries with string keys, lists, strings,
|
|
integers, finite floats, booleans and null. Tuples and custom value/container types
|
|
are rejected rather than coerced. Runner checks and detaches the snapshot before
|
|
postconditions or oracles evaluate it; invalid evidence stops further actions,
|
|
records an evidence failure and makes pending judgments INCONCLUSIVE. Already
|
|
observed failures remain FAIL. Serialization also rejects unsupported values,
|
|
non-string mapping keys, cycles and non-finite numbers. Observations detach nested data when recorded, so later domain mutations
|
|
do not rewrite earlier snapshots. Receipts include the final run verdict and
|
|
asset id, parent, maturity, variant and adaptation history. Independent claim
|
|
revisions and scenario revisions bind evidence to the original run inputs.
|
|
|
|
Retention is opt-in. It covers finalized returned runs, including FAIL and guard
|
|
aborts; unexpected driver/observer exceptions or process termination do not promise
|
|
a complete receipt. The checksum detects damage, not malicious replacement with
|
|
a recomputed checksum. The store is not encryption, credential redaction, access
|
|
custody or a retention service. Select observations safe to retain before enabling
|
|
it; existing driver mechanics can contain action arguments.
|
|
|
|
## Derive a security question from the same intent
|
|
|
|
```bash
|
|
PYTHONPATH=src:. python3 - <<'PY'
|
|
from scenarios.alice_bob_carol import build
|
|
from testdriver import Runner
|
|
from testdriver.variants import substitute
|
|
|
|
world, driver, observer, parent, oracle = build()
|
|
variant = substitute(parent, variant_id='grant-to-carol', step_id='s2-grant',
|
|
arguments={'subject_id': 'carol'})
|
|
assert variant.scenario.use_case is parent.scenario.use_case
|
|
result = Runner(world, driver, observer, oracle).run(variant)
|
|
print(variant.parent_id, result.verdict.value) # FAIL: original independent claims still apply
|
|
PY
|
|
```
|
|
|
|
`substitute(..., actor_id='other')` changes the scheduled actor; that actor must
|
|
exist in the execution world's cast. Runner validates all scheduled actor
|
|
references, cast key/id agreement and nonempty unique step ids before any action;
|
|
invalid input raises ValueError with no execution or finalized receipt. `arguments={...}` changes only existing
|
|
argument keys and covers resource, tenant or privilege substitution without new
|
|
framework concepts. Mutable argument data is copied. Unknown/duplicate step ids,
|
|
unknown keys, empty actor ids and invalid variant ids are rejected. Callers choose
|
|
unique variant ids for distinct definitions; the API is not an asset registry.
|
|
|
|
The exact UseCase, claim/invariant predicates, postconditions, watches, step order
|
|
and permitted surfaces remain unchanged. Parent identity and mutation metadata
|
|
are retained without copying changed values into mutation history. A derived
|
|
variant does not become an approved new requirement, and a denial is not a default
|
|
PASS: it needs the independently authored assertions appropriate to the question.
|
|
|
|
Acceptance fingerprints include scheduled actor/action identities, argument values,
|
|
order, surface permissions and postcondition definitions. Changed definitions
|
|
require intent review, even when product verdicts happen to pass. Unsupported
|
|
runtime dependencies or old packs without this revision cannot be accepted or
|
|
crystallized automatically. Mechanical selector/endpoint choices remain outside
|
|
scenario intent and can still qualify as mechanical adaptation.
|
|
|
|
This API deliberately does not generate new assertions, skip/reorder steps,
|
|
schedule races or infer required evidence from a successful implementation run.
|
|
|
|
Evidence/preflight decision: `aa6c3e31-614c-4030-990c-945acef08515`.
|
|
|
|
## Generated regression outcomes
|
|
|
|
Generated pytest modules embed the same predicate evaluator used by Oracle and
|
|
continue to import the original independent predicates. They require no runtime
|
|
test-driver machinery; the outcome adapter uses stdlib unittest.SkipTest.
|
|
|
|
True means PASS and False means FAIL. Empty snapshots, missing observation keys,
|
|
predicate exceptions and non-boolean results mean INCONCLUSIVE. The module
|
|
evaluates all supplied assertions: any FAIL takes precedence over INCONCLUSIVE,
|
|
regardless of assertion order. With no failure but an inconclusive assertion,
|
|
pytest reports a skip labeled INCONCLUSIVE, not a pass or product failure.
|
|
Use `pytest -rs` to display those reasons. A zero pytest exit status alone is not
|
|
proof that every generated assertion was verified; inspect skipped outcomes.
|
|
|
|
The checked-in descendant is regenerated with its original protected predicates
|
|
and frozen path. Execution-based tests compare generated and checked-in judgments
|
|
with Oracle, and a subprocess test verifies actual pytest pass/fail/skip reporting.
|
|
Decision: `c06c8752-80db-4458-9eaf-6321a0ca710c`.
|