diff --git a/SCOPE.md b/SCOPE.md index 83d06fa..2e8d44a 100644 --- a/SCOPE.md +++ b/SCOPE.md @@ -42,7 +42,8 @@ time-window or independent cleanup adapters in this repository. - Evidence storage is opt-in and local. It is not encryption, signing, automatic redaction, retention management or safe storage of arbitrary secrets. Unexpected driver/observer exceptions still propagate; no finalized receipt is promised for - an interrupted/crashed process. Strict JSON rejects unsupported evidence values. + an interrupted/crashed process. JSON-native snapshots are validated before judgment; lossy values stop the run + with pending judgments INCONCLUSIVE. Full schedule/cast validation precedes actions. - No automatic finding-to-regression pipeline, asset registry, maturity promotion, thawing or retirement. Parent ids and mutation history provide limited lineage. diff --git a/docs/TestDriverEvidenceAndVariants.md b/docs/TestDriverEvidenceAndVariants.md index d7a6510..cea90a5 100644 --- a/docs/TestDriverEvidenceAndVariants.md +++ b/docs/TestDriverEvidenceAndVariants.md @@ -31,8 +31,13 @@ publication of the same run yields one winner and FileExistsError for others. Directory fsync uses the local POSIX filesystem API. Storage failures propagate; a caller must not report retained evidence if a write failed. -Serialization rejects unsupported values, non-string mapping keys and non-finite -numbers. Observations detach nested data when recorded, so later domain mutations +Observations must use JSON-native dictionaries with string keys, lists, strings, +integers, finite floats, booleans and null. Tuples and custom value/container types +are rejected rather than coerced. Runner checks and detaches the snapshot before +postconditions or oracles evaluate it; invalid evidence stops further actions, +records an evidence failure and makes pending judgments INCONCLUSIVE. Already +observed failures remain FAIL. Serialization also rejects unsupported values, +non-string mapping keys, cycles and non-finite numbers. Observations detach nested data when recorded, so later domain mutations do not rewrite earlier snapshots. Receipts include the final run verdict and asset id, parent, maturity, variant and adaptation history. Independent claim revisions and scenario revisions bind evidence to the original run inputs. @@ -62,7 +67,9 @@ PY ``` `substitute(..., actor_id='other')` changes the scheduled actor; that actor must -exist in the execution world's cast. `arguments={...}` changes only existing +exist in the execution world's cast. Runner validates all scheduled actor +references, cast key/id agreement and nonempty unique step ids before any action; +invalid input raises ValueError with no execution or finalized receipt. `arguments={...}` changes only existing argument keys and covers resource, tenant or privilege substitution without new framework concepts. Mutable argument data is copied. Unknown/duplicate step ids, unknown keys, empty actor ids and invalid variant ids are rejected. Callers choose @@ -83,3 +90,5 @@ scenario intent and can still qualify as mechanical adaptation. This API deliberately does not generate new assertions, skip/reorder steps, schedule races or infer required evidence from a successful implementation run. + +Evidence/preflight decision: `aa6c3e31-614c-4030-990c-945acef08515`. diff --git a/history/2026-09-28-121933-scope-intent-assessment.md b/history/2026-09-28-121933-scope-intent-assessment.md index 5670403..6144b3b 100644 --- a/history/2026-09-28-121933-scope-intent-assessment.md +++ b/history/2026-09-28-121933-scope-intent-assessment.md @@ -105,3 +105,16 @@ experiment records. The scope update and local implementation do not close those validation gaps. Decision: `f3b35dce-6047-4a95-8083-4d7032d4e99f`. + + +## Evidence/preflight correction — 2026-09-28 + +A subsequent review reproduced tuple-to-list conversion changing an oracle result +on receipt replay, and a late unknown actor causing failure after earlier SUT +mutations. The follow-up under TD-WP-0004-T02/T03 rejects non-JSON-native snapshots +before judgment, retains a safe evidence-failure receipt with pending assertions +INCONCLUSIVE, and validates the whole schedule/cast before execution. This closes +the reproduced cases without changing claim definitions or expanding pilot scope. + +Follow-up validation: **440 tests passed** in 169.07 seconds, including 12 new +regressions (10 reproduced failures before implementation). diff --git a/research/decisions/README.md b/research/decisions/README.md index 239f1da..73a572a 100644 --- a/research/decisions/README.md +++ b/research/decisions/README.md @@ -14,3 +14,4 @@ file is a pointer table, not a second source of truth. | TD-WP-0003 T08 guard-integrity follow-up | Strict boolean judgments and intact isolation guards | `2ecac1c5-5254-4ce9-b092-b47c3e1283a9` | `docs/TestDriverGeneralisationReview.md` | | TD-WP-0003 T08 HTTP/surface follow-up | Origin-bound credentials and runner surface validation | `88490ee8-839d-40dc-affa-b00a1c68e3c5` | `docs/TestDriverGeneralisationReview.md` | | TD-WP-0004 | Evidence-backed scope, local receipts and intent-preserving variants | `f3b35dce-6047-4a95-8083-4d7032d4e99f` | `history/2026-09-28-121933-scope-intent-assessment.md` | +| TD-WP-0004 evidence/preflight follow-up | Lossless observation types and schedule preflight | `aa6c3e31-614c-4030-990c-945acef08515` | `docs/TestDriverEvidenceAndVariants.md` | diff --git a/src/testdriver/evidence.py b/src/testdriver/evidence.py index c3d5880..71c9e0b 100644 --- a/src/testdriver/evidence.py +++ b/src/testdriver/evidence.py @@ -13,13 +13,35 @@ See docs/TestDriverClassificationDesign.md, Part A. from __future__ import annotations import json +import math from copy import deepcopy -from dataclasses import dataclass, field, asdict, replace +from dataclasses import dataclass, field, fields, replace from datetime import datetime, timezone from enum import Enum from typing import Any +class InvalidEvidence(TypeError): + """Evidence cannot be represented losslessly as JSON-native values.""" + + +def json_value(value, active=None): + """Validate and detach without coercing tuples, keys or custom value types.""" + if value is None or type(value) in (bool, int, str): + return value + if type(value) is float and math.isfinite(value): + return value + active = set() if active is None else active + if id(value) in active or len(active) >= 100: + raise InvalidEvidence("cyclic or excessively deep evidence") + nested = active | {id(value)} + if type(value) is list: + return [json_value(item, nested) for item in value] + if type(value) is dict and all(type(key) is str for key in value): + return {key: json_value(item, nested) for key, item in value.items()} + raise InvalidEvidence("evidence requires JSON-native values and string keys") + + class Stratum(str, Enum): SURFACE = "S1" REALIZATION = "S2" @@ -81,19 +103,10 @@ class EvidencePack: return [o for o in self.observations if o.stratum is stratum] def to_json(self) -> str: - payload = asdict(self) + payload = {item.name: getattr(self, item.name) for item in fields(self) + if item.name != "observations"} payload["observations"] = [ - {**asdict(o), "stratum": o.stratum.value} for o in self.observations + {**{item.name: getattr(o, item.name) for item in fields(o)}, + "stratum": o.stratum.value} for o in self.observations ] - def check_keys(value): - if isinstance(value, dict): - if any(not isinstance(key, str) for key in value): - raise TypeError("evidence objects require string keys") - for item in value.values(): - check_keys(item) - elif isinstance(value, (list, tuple)): - for item in value: - check_keys(item) - - check_keys(payload) - return json.dumps(payload, indent=2, sort_keys=True, allow_nan=False) + return json.dumps(json_value(payload), indent=2, sort_keys=True, allow_nan=False) diff --git a/src/testdriver/runner.py b/src/testdriver/runner.py index 8da6eba..d2d30e2 100644 --- a/src/testdriver/runner.py +++ b/src/testdriver/runner.py @@ -16,7 +16,7 @@ from typing import Any from .actions import SurfaceNotPermitted from .drivers import Driver, realize_step -from .evidence import EvidencePack, Observation, Stratum +from .evidence import EvidencePack, Observation, Stratum, InvalidEvidence, json_value from .observers import StateObserver from .oracles import Judgment, Oracle, Verdict, overall from .revisions import intent_revisions, scenario_revision @@ -134,6 +134,17 @@ class Runner: def run(self, asset: VerificationAsset, *, evidence_store=None) -> RunResult: scenario: Scenario = asset.scenario + # Validate the whole schedule before touching any actor or SUT. + step_ids = set() + for key, actor in self._world.cast.actors.items(): + if not isinstance(key, str) or not key or actor.id != key: + raise ValueError("cast keys must match nonempty actor identities") + for step in scenario.steps: + if not isinstance(step.id, str) or not step.id or step.id in step_ids: + raise ValueError("scenario step ids must be nonempty and unique") + step_ids.add(step.id) + if not isinstance(step.actor_id, str) or step.actor_id not in self._world.cast.actors: + raise ValueError("scenario step names an unknown actor") run_id = f"run-{uuid.uuid4().hex[:12]}" pack = EvidencePack( run_id=run_id, @@ -242,7 +253,19 @@ class Runner: break # --- S3: what is now true ------------------------------------ - snapshot = self._observer.snapshot() + raw_snapshot = self._observer.snapshot() + try: + snapshot = json_value(raw_snapshot) + if type(snapshot) is not dict: + raise InvalidEvidence("observation snapshot must be a JSON object") + except InvalidEvidence as exc: + self._record( + pack, Stratum.JUDGMENT, self._observer.name, "evidence_failure", + {"reason": str(exc)}, step.id, + ) + aborted = True + skip_remaining(step_index, "observation cannot be retained losslessly") + break self._record( pack, Stratum.JUDGMENT, self._observer.name, "state_snapshot", dict(snapshot), step.id, diff --git a/tests/test_run_preflight_and_json.py b/tests/test_run_preflight_and_json.py new file mode 100644 index 0000000..4a1e5cc --- /dev/null +++ b/tests/test_run_preflight_and_json.py @@ -0,0 +1,88 @@ +"""Reject unsafe run inputs before effects and lossy observation types before judgment.""" +from dataclasses import replace +import json + +import pytest + +from scenarios.alice_bob_carol import build +from testdriver import Runner, Verdict +from testdriver.classification import classify +from testdriver.storage import EvidenceStore +from testdriver.variants import substitute + + +@pytest.mark.parametrize('bad', [('a', 'b'), {'nested': [('a', 'b')]}, {1: 'value'}, + float('nan'), float('inf')]) +def test_lossy_observations_never_produce_a_verdict_that_cannot_be_replayed(tmp_path, bad): + w, d, o, a, oracle = build() + calls = [] + case = a.scenario.use_case + claim = replace(case.claims[0], predicate=lambda snapshot: snapshot['pair'] == ('a', 'b')) + a.scenario = replace(a.scenario, use_case=replace(case, claims=(claim, *case.claims[1:]))) + + class Observer: + name = o.name + def snapshot(self): + return {**o.snapshot(), 'pair': bad} + + class Driver: + def realize(self, actor, action): + calls.append(action.name) + return d.realize(actor, action) + + store = EvidenceStore(tmp_path) + result = Runner(w, Driver(), Observer(), oracle).run(a, evidence_store=store) + assert result.verdict is Verdict.INCONCLUSIVE + assert len(calls) == 1 + assert all(j.verdict is Verdict.INCONCLUSIVE for j in result.judgments) + retained = store.load(result.run_id) + assert retained['run_verdict'] == 'INCONCLUSIVE' + assert any(o['kind'] == 'evidence_failure' for o in retained['observations']) + assert not classify(retained, retained).safe_to_accept + + +def test_json_native_observations_keep_live_and_replayed_verdicts(tmp_path): + w, d, o, a, oracle = build() + result = Runner(w, d, o, oracle).run(a, evidence_store=EvidenceStore(tmp_path)) + retained = EvidenceStore(tmp_path).load(result.run_id) + snapshots = {o['step_id']: o['data'] for o in retained['observations'] + if o['kind'] == 'state_snapshot'} + assertions = {c.id: c for c in (*a.scenario.use_case.claims, *a.scenario.use_case.invariants)} + for verdict in retained['verdicts']: + replayed = oracle.judge(assertions[verdict['assertion_id']], snapshots[verdict['step_id']], verdict['step_id']) + assert replayed.verdict.value == verdict['verdict'] + assert result.verdict is Verdict.PASS + + +@pytest.mark.parametrize('bad', [('a', 'b'), {1: 'value'}]) +def test_serializer_does_not_silently_convert_unsupported_types(bad): + w, d, o, a, oracle = build() + pack = Runner(w, d, o, oracle).run(a).evidence + pack.asset['invalid'] = bad + with pytest.raises(TypeError): + pack.to_json() + + +@pytest.mark.parametrize('damage', ['unknown-actor', 'duplicate-step', 'empty-step', 'cast-mismatch']) +def test_entire_schedule_is_validated_before_any_side_effect(tmp_path, damage): + w, d, o, a, oracle = build() + calls = [] + if damage == 'unknown-actor': + a = substitute(a, variant_id='typo', step_id=a.scenario.steps[-1].id, actor_id='unknown') + elif damage == 'cast-mismatch': + w.cast['alice'].id = 'different' + else: + steps = list(a.scenario.steps) + steps[-1] = replace(steps[-1], id='' if damage == 'empty-step' else steps[0].id) + a.scenario = replace(a.scenario, steps=tuple(steps)) + + class Driver: + def realize(self, actor, action): + calls.append(action.name) + return d.realize(actor, action) + + with pytest.raises(ValueError, match='scenario|actor|step|cast'): + Runner(w, Driver(), o, oracle).run(a, evidence_store=EvidenceStore(tmp_path)) + assert calls == [] + assert not w.sut.resources + assert not list(tmp_path.iterdir()) diff --git a/workplans/TD-WP-0004-scope-evidence-and-variants.md b/workplans/TD-WP-0004-scope-evidence-and-variants.md index 8dbaedc..6d25a26 100644 --- a/workplans/TD-WP-0004-scope-evidence-and-variants.md +++ b/workplans/TD-WP-0004-scope-evidence-and-variants.md @@ -132,3 +132,18 @@ scope prose: the pilot remains live here, and model/browser/authoring experiment remain TD-WP-0003-T01/T06/T07. Broader scheduling/lifecycle/adapter features remain explicit scope limits to prioritize from actual pilot needs. No real-system, paid-model or browser-engine execution was performed. + + +**2026-09-28 evidence/preflight follow-up — done (T02/T03).** Runner validates +JSON-native observations before evaluation/retention. Lossy types (including +tuples), invalid keys, non-finite numbers and cyclic data cannot silently change +meaning in JSON; a bad snapshot stops execution with an evidence-failure record +and pending assertions INCONCLUSIVE. Serialization uses the same strict value +contract without dataclass or custom-value coercion. Runner preflights every +scheduled actor, unique nonempty step id and cast key/id agreement before any +SUT action. Intentionally partial scenarios remain valid. + +Validation: 12 new regressions, of which 10 failed before the fixes; **131 focused +tests passed**, and the full suite passed **440 tests in 169.07 seconds**. +`git diff --check` is clean. Decision: `aa6c3e31-614c-4030-990c-945acef08515`. +No new tasks/workplans; T05 remains waiting and this workplan remains blocked.