Reject lossy observations and validate schedules before execution
Assistant: codex Assistant-Model: gpt-6-astra Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
This commit is contained in:
parent
eaf5d348e4
commit
10077edb8c
8 changed files with 184 additions and 21 deletions
3
SCOPE.md
3
SCOPE.md
|
|
@ -42,7 +42,8 @@ time-window or independent cleanup adapters in this repository.
|
|||
- Evidence storage is opt-in and local. It is not encryption, signing, automatic
|
||||
redaction, retention management or safe storage of arbitrary secrets. Unexpected
|
||||
driver/observer exceptions still propagate; no finalized receipt is promised for
|
||||
an interrupted/crashed process. Strict JSON rejects unsupported evidence values.
|
||||
an interrupted/crashed process. JSON-native snapshots are validated before judgment; lossy values stop the run
|
||||
with pending judgments INCONCLUSIVE. Full schedule/cast validation precedes actions.
|
||||
- No automatic finding-to-regression pipeline, asset registry, maturity promotion,
|
||||
thawing or retirement. Parent ids and mutation history provide limited lineage.
|
||||
|
||||
|
|
|
|||
|
|
@ -31,8 +31,13 @@ publication of the same run yields one winner and FileExistsError for others.
|
|||
Directory fsync uses the local POSIX filesystem API. Storage failures propagate;
|
||||
a caller must not report retained evidence if a write failed.
|
||||
|
||||
Serialization rejects unsupported values, non-string mapping keys and non-finite
|
||||
numbers. Observations detach nested data when recorded, so later domain mutations
|
||||
Observations must use JSON-native dictionaries with string keys, lists, strings,
|
||||
integers, finite floats, booleans and null. Tuples and custom value/container types
|
||||
are rejected rather than coerced. Runner checks and detaches the snapshot before
|
||||
postconditions or oracles evaluate it; invalid evidence stops further actions,
|
||||
records an evidence failure and makes pending judgments INCONCLUSIVE. Already
|
||||
observed failures remain FAIL. Serialization also rejects unsupported values,
|
||||
non-string mapping keys, cycles and non-finite numbers. Observations detach nested data when recorded, so later domain mutations
|
||||
do not rewrite earlier snapshots. Receipts include the final run verdict and
|
||||
asset id, parent, maturity, variant and adaptation history. Independent claim
|
||||
revisions and scenario revisions bind evidence to the original run inputs.
|
||||
|
|
@ -62,7 +67,9 @@ PY
|
|||
```
|
||||
|
||||
`substitute(..., actor_id='other')` changes the scheduled actor; that actor must
|
||||
exist in the execution world's cast. `arguments={...}` changes only existing
|
||||
exist in the execution world's cast. Runner validates all scheduled actor
|
||||
references, cast key/id agreement and nonempty unique step ids before any action;
|
||||
invalid input raises ValueError with no execution or finalized receipt. `arguments={...}` changes only existing
|
||||
argument keys and covers resource, tenant or privilege substitution without new
|
||||
framework concepts. Mutable argument data is copied. Unknown/duplicate step ids,
|
||||
unknown keys, empty actor ids and invalid variant ids are rejected. Callers choose
|
||||
|
|
@ -83,3 +90,5 @@ scenario intent and can still qualify as mechanical adaptation.
|
|||
|
||||
This API deliberately does not generate new assertions, skip/reorder steps,
|
||||
schedule races or infer required evidence from a successful implementation run.
|
||||
|
||||
Evidence/preflight decision: `aa6c3e31-614c-4030-990c-945acef08515`.
|
||||
|
|
|
|||
|
|
@ -105,3 +105,16 @@ experiment records. The scope update and local implementation do not close those
|
|||
validation gaps.
|
||||
|
||||
Decision: `f3b35dce-6047-4a95-8083-4d7032d4e99f`.
|
||||
|
||||
|
||||
## Evidence/preflight correction — 2026-09-28
|
||||
|
||||
A subsequent review reproduced tuple-to-list conversion changing an oracle result
|
||||
on receipt replay, and a late unknown actor causing failure after earlier SUT
|
||||
mutations. The follow-up under TD-WP-0004-T02/T03 rejects non-JSON-native snapshots
|
||||
before judgment, retains a safe evidence-failure receipt with pending assertions
|
||||
INCONCLUSIVE, and validates the whole schedule/cast before execution. This closes
|
||||
the reproduced cases without changing claim definitions or expanding pilot scope.
|
||||
|
||||
Follow-up validation: **440 tests passed** in 169.07 seconds, including 12 new
|
||||
regressions (10 reproduced failures before implementation).
|
||||
|
|
|
|||
|
|
@ -14,3 +14,4 @@ file is a pointer table, not a second source of truth.
|
|||
| TD-WP-0003 T08 guard-integrity follow-up | Strict boolean judgments and intact isolation guards | `2ecac1c5-5254-4ce9-b092-b47c3e1283a9` | `docs/TestDriverGeneralisationReview.md` |
|
||||
| TD-WP-0003 T08 HTTP/surface follow-up | Origin-bound credentials and runner surface validation | `88490ee8-839d-40dc-affa-b00a1c68e3c5` | `docs/TestDriverGeneralisationReview.md` |
|
||||
| TD-WP-0004 | Evidence-backed scope, local receipts and intent-preserving variants | `f3b35dce-6047-4a95-8083-4d7032d4e99f` | `history/2026-09-28-121933-scope-intent-assessment.md` |
|
||||
| TD-WP-0004 evidence/preflight follow-up | Lossless observation types and schedule preflight | `aa6c3e31-614c-4030-990c-945acef08515` | `docs/TestDriverEvidenceAndVariants.md` |
|
||||
|
|
|
|||
|
|
@ -13,13 +13,35 @@ See docs/TestDriverClassificationDesign.md, Part A.
|
|||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import math
|
||||
from copy import deepcopy
|
||||
from dataclasses import dataclass, field, asdict, replace
|
||||
from dataclasses import dataclass, field, fields, replace
|
||||
from datetime import datetime, timezone
|
||||
from enum import Enum
|
||||
from typing import Any
|
||||
|
||||
|
||||
class InvalidEvidence(TypeError):
|
||||
"""Evidence cannot be represented losslessly as JSON-native values."""
|
||||
|
||||
|
||||
def json_value(value, active=None):
|
||||
"""Validate and detach without coercing tuples, keys or custom value types."""
|
||||
if value is None or type(value) in (bool, int, str):
|
||||
return value
|
||||
if type(value) is float and math.isfinite(value):
|
||||
return value
|
||||
active = set() if active is None else active
|
||||
if id(value) in active or len(active) >= 100:
|
||||
raise InvalidEvidence("cyclic or excessively deep evidence")
|
||||
nested = active | {id(value)}
|
||||
if type(value) is list:
|
||||
return [json_value(item, nested) for item in value]
|
||||
if type(value) is dict and all(type(key) is str for key in value):
|
||||
return {key: json_value(item, nested) for key, item in value.items()}
|
||||
raise InvalidEvidence("evidence requires JSON-native values and string keys")
|
||||
|
||||
|
||||
class Stratum(str, Enum):
|
||||
SURFACE = "S1"
|
||||
REALIZATION = "S2"
|
||||
|
|
@ -81,19 +103,10 @@ class EvidencePack:
|
|||
return [o for o in self.observations if o.stratum is stratum]
|
||||
|
||||
def to_json(self) -> str:
|
||||
payload = asdict(self)
|
||||
payload = {item.name: getattr(self, item.name) for item in fields(self)
|
||||
if item.name != "observations"}
|
||||
payload["observations"] = [
|
||||
{**asdict(o), "stratum": o.stratum.value} for o in self.observations
|
||||
{**{item.name: getattr(o, item.name) for item in fields(o)},
|
||||
"stratum": o.stratum.value} for o in self.observations
|
||||
]
|
||||
def check_keys(value):
|
||||
if isinstance(value, dict):
|
||||
if any(not isinstance(key, str) for key in value):
|
||||
raise TypeError("evidence objects require string keys")
|
||||
for item in value.values():
|
||||
check_keys(item)
|
||||
elif isinstance(value, (list, tuple)):
|
||||
for item in value:
|
||||
check_keys(item)
|
||||
|
||||
check_keys(payload)
|
||||
return json.dumps(payload, indent=2, sort_keys=True, allow_nan=False)
|
||||
return json.dumps(json_value(payload), indent=2, sort_keys=True, allow_nan=False)
|
||||
|
|
|
|||
|
|
@ -16,7 +16,7 @@ from typing import Any
|
|||
|
||||
from .actions import SurfaceNotPermitted
|
||||
from .drivers import Driver, realize_step
|
||||
from .evidence import EvidencePack, Observation, Stratum
|
||||
from .evidence import EvidencePack, Observation, Stratum, InvalidEvidence, json_value
|
||||
from .observers import StateObserver
|
||||
from .oracles import Judgment, Oracle, Verdict, overall
|
||||
from .revisions import intent_revisions, scenario_revision
|
||||
|
|
@ -134,6 +134,17 @@ class Runner:
|
|||
|
||||
def run(self, asset: VerificationAsset, *, evidence_store=None) -> RunResult:
|
||||
scenario: Scenario = asset.scenario
|
||||
# Validate the whole schedule before touching any actor or SUT.
|
||||
step_ids = set()
|
||||
for key, actor in self._world.cast.actors.items():
|
||||
if not isinstance(key, str) or not key or actor.id != key:
|
||||
raise ValueError("cast keys must match nonempty actor identities")
|
||||
for step in scenario.steps:
|
||||
if not isinstance(step.id, str) or not step.id or step.id in step_ids:
|
||||
raise ValueError("scenario step ids must be nonempty and unique")
|
||||
step_ids.add(step.id)
|
||||
if not isinstance(step.actor_id, str) or step.actor_id not in self._world.cast.actors:
|
||||
raise ValueError("scenario step names an unknown actor")
|
||||
run_id = f"run-{uuid.uuid4().hex[:12]}"
|
||||
pack = EvidencePack(
|
||||
run_id=run_id,
|
||||
|
|
@ -242,7 +253,19 @@ class Runner:
|
|||
break
|
||||
|
||||
# --- S3: what is now true ------------------------------------
|
||||
snapshot = self._observer.snapshot()
|
||||
raw_snapshot = self._observer.snapshot()
|
||||
try:
|
||||
snapshot = json_value(raw_snapshot)
|
||||
if type(snapshot) is not dict:
|
||||
raise InvalidEvidence("observation snapshot must be a JSON object")
|
||||
except InvalidEvidence as exc:
|
||||
self._record(
|
||||
pack, Stratum.JUDGMENT, self._observer.name, "evidence_failure",
|
||||
{"reason": str(exc)}, step.id,
|
||||
)
|
||||
aborted = True
|
||||
skip_remaining(step_index, "observation cannot be retained losslessly")
|
||||
break
|
||||
self._record(
|
||||
pack, Stratum.JUDGMENT, self._observer.name,
|
||||
"state_snapshot", dict(snapshot), step.id,
|
||||
|
|
|
|||
88
tests/test_run_preflight_and_json.py
Normal file
88
tests/test_run_preflight_and_json.py
Normal file
|
|
@ -0,0 +1,88 @@
|
|||
"""Reject unsafe run inputs before effects and lossy observation types before judgment."""
|
||||
from dataclasses import replace
|
||||
import json
|
||||
|
||||
import pytest
|
||||
|
||||
from scenarios.alice_bob_carol import build
|
||||
from testdriver import Runner, Verdict
|
||||
from testdriver.classification import classify
|
||||
from testdriver.storage import EvidenceStore
|
||||
from testdriver.variants import substitute
|
||||
|
||||
|
||||
@pytest.mark.parametrize('bad', [('a', 'b'), {'nested': [('a', 'b')]}, {1: 'value'},
|
||||
float('nan'), float('inf')])
|
||||
def test_lossy_observations_never_produce_a_verdict_that_cannot_be_replayed(tmp_path, bad):
|
||||
w, d, o, a, oracle = build()
|
||||
calls = []
|
||||
case = a.scenario.use_case
|
||||
claim = replace(case.claims[0], predicate=lambda snapshot: snapshot['pair'] == ('a', 'b'))
|
||||
a.scenario = replace(a.scenario, use_case=replace(case, claims=(claim, *case.claims[1:])))
|
||||
|
||||
class Observer:
|
||||
name = o.name
|
||||
def snapshot(self):
|
||||
return {**o.snapshot(), 'pair': bad}
|
||||
|
||||
class Driver:
|
||||
def realize(self, actor, action):
|
||||
calls.append(action.name)
|
||||
return d.realize(actor, action)
|
||||
|
||||
store = EvidenceStore(tmp_path)
|
||||
result = Runner(w, Driver(), Observer(), oracle).run(a, evidence_store=store)
|
||||
assert result.verdict is Verdict.INCONCLUSIVE
|
||||
assert len(calls) == 1
|
||||
assert all(j.verdict is Verdict.INCONCLUSIVE for j in result.judgments)
|
||||
retained = store.load(result.run_id)
|
||||
assert retained['run_verdict'] == 'INCONCLUSIVE'
|
||||
assert any(o['kind'] == 'evidence_failure' for o in retained['observations'])
|
||||
assert not classify(retained, retained).safe_to_accept
|
||||
|
||||
|
||||
def test_json_native_observations_keep_live_and_replayed_verdicts(tmp_path):
|
||||
w, d, o, a, oracle = build()
|
||||
result = Runner(w, d, o, oracle).run(a, evidence_store=EvidenceStore(tmp_path))
|
||||
retained = EvidenceStore(tmp_path).load(result.run_id)
|
||||
snapshots = {o['step_id']: o['data'] for o in retained['observations']
|
||||
if o['kind'] == 'state_snapshot'}
|
||||
assertions = {c.id: c for c in (*a.scenario.use_case.claims, *a.scenario.use_case.invariants)}
|
||||
for verdict in retained['verdicts']:
|
||||
replayed = oracle.judge(assertions[verdict['assertion_id']], snapshots[verdict['step_id']], verdict['step_id'])
|
||||
assert replayed.verdict.value == verdict['verdict']
|
||||
assert result.verdict is Verdict.PASS
|
||||
|
||||
|
||||
@pytest.mark.parametrize('bad', [('a', 'b'), {1: 'value'}])
|
||||
def test_serializer_does_not_silently_convert_unsupported_types(bad):
|
||||
w, d, o, a, oracle = build()
|
||||
pack = Runner(w, d, o, oracle).run(a).evidence
|
||||
pack.asset['invalid'] = bad
|
||||
with pytest.raises(TypeError):
|
||||
pack.to_json()
|
||||
|
||||
|
||||
@pytest.mark.parametrize('damage', ['unknown-actor', 'duplicate-step', 'empty-step', 'cast-mismatch'])
|
||||
def test_entire_schedule_is_validated_before_any_side_effect(tmp_path, damage):
|
||||
w, d, o, a, oracle = build()
|
||||
calls = []
|
||||
if damage == 'unknown-actor':
|
||||
a = substitute(a, variant_id='typo', step_id=a.scenario.steps[-1].id, actor_id='unknown')
|
||||
elif damage == 'cast-mismatch':
|
||||
w.cast['alice'].id = 'different'
|
||||
else:
|
||||
steps = list(a.scenario.steps)
|
||||
steps[-1] = replace(steps[-1], id='' if damage == 'empty-step' else steps[0].id)
|
||||
a.scenario = replace(a.scenario, steps=tuple(steps))
|
||||
|
||||
class Driver:
|
||||
def realize(self, actor, action):
|
||||
calls.append(action.name)
|
||||
return d.realize(actor, action)
|
||||
|
||||
with pytest.raises(ValueError, match='scenario|actor|step|cast'):
|
||||
Runner(w, Driver(), o, oracle).run(a, evidence_store=EvidenceStore(tmp_path))
|
||||
assert calls == []
|
||||
assert not w.sut.resources
|
||||
assert not list(tmp_path.iterdir())
|
||||
|
|
@ -132,3 +132,18 @@ scope prose: the pilot remains live here, and model/browser/authoring experiment
|
|||
remain TD-WP-0003-T01/T06/T07. Broader scheduling/lifecycle/adapter features remain
|
||||
explicit scope limits to prioritize from actual pilot needs. No real-system,
|
||||
paid-model or browser-engine execution was performed.
|
||||
|
||||
|
||||
**2026-09-28 evidence/preflight follow-up — done (T02/T03).** Runner validates
|
||||
JSON-native observations before evaluation/retention. Lossy types (including
|
||||
tuples), invalid keys, non-finite numbers and cyclic data cannot silently change
|
||||
meaning in JSON; a bad snapshot stops execution with an evidence-failure record
|
||||
and pending assertions INCONCLUSIVE. Serialization uses the same strict value
|
||||
contract without dataclass or custom-value coercion. Runner preflights every
|
||||
scheduled actor, unique nonempty step id and cast key/id agreement before any
|
||||
SUT action. Intentionally partial scenarios remain valid.
|
||||
|
||||
Validation: 12 new regressions, of which 10 failed before the fixes; **131 focused
|
||||
tests passed**, and the full suite passed **440 tests in 169.07 seconds**.
|
||||
`git diff --check` is clean. Decision: `aa6c3e31-614c-4030-990c-945acef08515`.
|
||||
No new tasks/workplans; T05 remains waiting and this workplan remains blocked.
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue