Reject lossy observations and validate schedules before execution

Assistant: codex
Assistant-Model: gpt-6-astra
Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
This commit is contained in:
tegwick 2026-09-28 14:50:27 +02:00
parent eaf5d348e4
commit 10077edb8c
8 changed files with 184 additions and 21 deletions

View file

@ -42,7 +42,8 @@ time-window or independent cleanup adapters in this repository.
- Evidence storage is opt-in and local. It is not encryption, signing, automatic
redaction, retention management or safe storage of arbitrary secrets. Unexpected
driver/observer exceptions still propagate; no finalized receipt is promised for
an interrupted/crashed process. Strict JSON rejects unsupported evidence values.
an interrupted/crashed process. JSON-native snapshots are validated before judgment; lossy values stop the run
with pending judgments INCONCLUSIVE. Full schedule/cast validation precedes actions.
- No automatic finding-to-regression pipeline, asset registry, maturity promotion,
thawing or retirement. Parent ids and mutation history provide limited lineage.

View file

@ -31,8 +31,13 @@ publication of the same run yields one winner and FileExistsError for others.
Directory fsync uses the local POSIX filesystem API. Storage failures propagate;
a caller must not report retained evidence if a write failed.
Serialization rejects unsupported values, non-string mapping keys and non-finite
numbers. Observations detach nested data when recorded, so later domain mutations
Observations must use JSON-native dictionaries with string keys, lists, strings,
integers, finite floats, booleans and null. Tuples and custom value/container types
are rejected rather than coerced. Runner checks and detaches the snapshot before
postconditions or oracles evaluate it; invalid evidence stops further actions,
records an evidence failure and makes pending judgments INCONCLUSIVE. Already
observed failures remain FAIL. Serialization also rejects unsupported values,
non-string mapping keys, cycles and non-finite numbers. Observations detach nested data when recorded, so later domain mutations
do not rewrite earlier snapshots. Receipts include the final run verdict and
asset id, parent, maturity, variant and adaptation history. Independent claim
revisions and scenario revisions bind evidence to the original run inputs.
@ -62,7 +67,9 @@ PY
```
`substitute(..., actor_id='other')` changes the scheduled actor; that actor must
exist in the execution world's cast. `arguments={...}` changes only existing
exist in the execution world's cast. Runner validates all scheduled actor
references, cast key/id agreement and nonempty unique step ids before any action;
invalid input raises ValueError with no execution or finalized receipt. `arguments={...}` changes only existing
argument keys and covers resource, tenant or privilege substitution without new
framework concepts. Mutable argument data is copied. Unknown/duplicate step ids,
unknown keys, empty actor ids and invalid variant ids are rejected. Callers choose
@ -83,3 +90,5 @@ scenario intent and can still qualify as mechanical adaptation.
This API deliberately does not generate new assertions, skip/reorder steps,
schedule races or infer required evidence from a successful implementation run.
Evidence/preflight decision: `aa6c3e31-614c-4030-990c-945acef08515`.

View file

@ -105,3 +105,16 @@ experiment records. The scope update and local implementation do not close those
validation gaps.
Decision: `f3b35dce-6047-4a95-8083-4d7032d4e99f`.
## Evidence/preflight correction — 2026-09-28
A subsequent review reproduced tuple-to-list conversion changing an oracle result
on receipt replay, and a late unknown actor causing failure after earlier SUT
mutations. The follow-up under TD-WP-0004-T02/T03 rejects non-JSON-native snapshots
before judgment, retains a safe evidence-failure receipt with pending assertions
INCONCLUSIVE, and validates the whole schedule/cast before execution. This closes
the reproduced cases without changing claim definitions or expanding pilot scope.
Follow-up validation: **440 tests passed** in 169.07 seconds, including 12 new
regressions (10 reproduced failures before implementation).

View file

@ -14,3 +14,4 @@ file is a pointer table, not a second source of truth.
| TD-WP-0003 T08 guard-integrity follow-up | Strict boolean judgments and intact isolation guards | `2ecac1c5-5254-4ce9-b092-b47c3e1283a9` | `docs/TestDriverGeneralisationReview.md` |
| TD-WP-0003 T08 HTTP/surface follow-up | Origin-bound credentials and runner surface validation | `88490ee8-839d-40dc-affa-b00a1c68e3c5` | `docs/TestDriverGeneralisationReview.md` |
| TD-WP-0004 | Evidence-backed scope, local receipts and intent-preserving variants | `f3b35dce-6047-4a95-8083-4d7032d4e99f` | `history/2026-09-28-121933-scope-intent-assessment.md` |
| TD-WP-0004 evidence/preflight follow-up | Lossless observation types and schedule preflight | `aa6c3e31-614c-4030-990c-945acef08515` | `docs/TestDriverEvidenceAndVariants.md` |

View file

@ -13,13 +13,35 @@ See docs/TestDriverClassificationDesign.md, Part A.
from __future__ import annotations
import json
import math
from copy import deepcopy
from dataclasses import dataclass, field, asdict, replace
from dataclasses import dataclass, field, fields, replace
from datetime import datetime, timezone
from enum import Enum
from typing import Any
class InvalidEvidence(TypeError):
"""Evidence cannot be represented losslessly as JSON-native values."""
def json_value(value, active=None):
"""Validate and detach without coercing tuples, keys or custom value types."""
if value is None or type(value) in (bool, int, str):
return value
if type(value) is float and math.isfinite(value):
return value
active = set() if active is None else active
if id(value) in active or len(active) >= 100:
raise InvalidEvidence("cyclic or excessively deep evidence")
nested = active | {id(value)}
if type(value) is list:
return [json_value(item, nested) for item in value]
if type(value) is dict and all(type(key) is str for key in value):
return {key: json_value(item, nested) for key, item in value.items()}
raise InvalidEvidence("evidence requires JSON-native values and string keys")
class Stratum(str, Enum):
SURFACE = "S1"
REALIZATION = "S2"
@ -81,19 +103,10 @@ class EvidencePack:
return [o for o in self.observations if o.stratum is stratum]
def to_json(self) -> str:
payload = asdict(self)
payload = {item.name: getattr(self, item.name) for item in fields(self)
if item.name != "observations"}
payload["observations"] = [
{**asdict(o), "stratum": o.stratum.value} for o in self.observations
{**{item.name: getattr(o, item.name) for item in fields(o)},
"stratum": o.stratum.value} for o in self.observations
]
def check_keys(value):
if isinstance(value, dict):
if any(not isinstance(key, str) for key in value):
raise TypeError("evidence objects require string keys")
for item in value.values():
check_keys(item)
elif isinstance(value, (list, tuple)):
for item in value:
check_keys(item)
check_keys(payload)
return json.dumps(payload, indent=2, sort_keys=True, allow_nan=False)
return json.dumps(json_value(payload), indent=2, sort_keys=True, allow_nan=False)

View file

@ -16,7 +16,7 @@ from typing import Any
from .actions import SurfaceNotPermitted
from .drivers import Driver, realize_step
from .evidence import EvidencePack, Observation, Stratum
from .evidence import EvidencePack, Observation, Stratum, InvalidEvidence, json_value
from .observers import StateObserver
from .oracles import Judgment, Oracle, Verdict, overall
from .revisions import intent_revisions, scenario_revision
@ -134,6 +134,17 @@ class Runner:
def run(self, asset: VerificationAsset, *, evidence_store=None) -> RunResult:
scenario: Scenario = asset.scenario
# Validate the whole schedule before touching any actor or SUT.
step_ids = set()
for key, actor in self._world.cast.actors.items():
if not isinstance(key, str) or not key or actor.id != key:
raise ValueError("cast keys must match nonempty actor identities")
for step in scenario.steps:
if not isinstance(step.id, str) or not step.id or step.id in step_ids:
raise ValueError("scenario step ids must be nonempty and unique")
step_ids.add(step.id)
if not isinstance(step.actor_id, str) or step.actor_id not in self._world.cast.actors:
raise ValueError("scenario step names an unknown actor")
run_id = f"run-{uuid.uuid4().hex[:12]}"
pack = EvidencePack(
run_id=run_id,
@ -242,7 +253,19 @@ class Runner:
break
# --- S3: what is now true ------------------------------------
snapshot = self._observer.snapshot()
raw_snapshot = self._observer.snapshot()
try:
snapshot = json_value(raw_snapshot)
if type(snapshot) is not dict:
raise InvalidEvidence("observation snapshot must be a JSON object")
except InvalidEvidence as exc:
self._record(
pack, Stratum.JUDGMENT, self._observer.name, "evidence_failure",
{"reason": str(exc)}, step.id,
)
aborted = True
skip_remaining(step_index, "observation cannot be retained losslessly")
break
self._record(
pack, Stratum.JUDGMENT, self._observer.name,
"state_snapshot", dict(snapshot), step.id,

View file

@ -0,0 +1,88 @@
"""Reject unsafe run inputs before effects and lossy observation types before judgment."""
from dataclasses import replace
import json
import pytest
from scenarios.alice_bob_carol import build
from testdriver import Runner, Verdict
from testdriver.classification import classify
from testdriver.storage import EvidenceStore
from testdriver.variants import substitute
@pytest.mark.parametrize('bad', [('a', 'b'), {'nested': [('a', 'b')]}, {1: 'value'},
float('nan'), float('inf')])
def test_lossy_observations_never_produce_a_verdict_that_cannot_be_replayed(tmp_path, bad):
w, d, o, a, oracle = build()
calls = []
case = a.scenario.use_case
claim = replace(case.claims[0], predicate=lambda snapshot: snapshot['pair'] == ('a', 'b'))
a.scenario = replace(a.scenario, use_case=replace(case, claims=(claim, *case.claims[1:])))
class Observer:
name = o.name
def snapshot(self):
return {**o.snapshot(), 'pair': bad}
class Driver:
def realize(self, actor, action):
calls.append(action.name)
return d.realize(actor, action)
store = EvidenceStore(tmp_path)
result = Runner(w, Driver(), Observer(), oracle).run(a, evidence_store=store)
assert result.verdict is Verdict.INCONCLUSIVE
assert len(calls) == 1
assert all(j.verdict is Verdict.INCONCLUSIVE for j in result.judgments)
retained = store.load(result.run_id)
assert retained['run_verdict'] == 'INCONCLUSIVE'
assert any(o['kind'] == 'evidence_failure' for o in retained['observations'])
assert not classify(retained, retained).safe_to_accept
def test_json_native_observations_keep_live_and_replayed_verdicts(tmp_path):
w, d, o, a, oracle = build()
result = Runner(w, d, o, oracle).run(a, evidence_store=EvidenceStore(tmp_path))
retained = EvidenceStore(tmp_path).load(result.run_id)
snapshots = {o['step_id']: o['data'] for o in retained['observations']
if o['kind'] == 'state_snapshot'}
assertions = {c.id: c for c in (*a.scenario.use_case.claims, *a.scenario.use_case.invariants)}
for verdict in retained['verdicts']:
replayed = oracle.judge(assertions[verdict['assertion_id']], snapshots[verdict['step_id']], verdict['step_id'])
assert replayed.verdict.value == verdict['verdict']
assert result.verdict is Verdict.PASS
@pytest.mark.parametrize('bad', [('a', 'b'), {1: 'value'}])
def test_serializer_does_not_silently_convert_unsupported_types(bad):
w, d, o, a, oracle = build()
pack = Runner(w, d, o, oracle).run(a).evidence
pack.asset['invalid'] = bad
with pytest.raises(TypeError):
pack.to_json()
@pytest.mark.parametrize('damage', ['unknown-actor', 'duplicate-step', 'empty-step', 'cast-mismatch'])
def test_entire_schedule_is_validated_before_any_side_effect(tmp_path, damage):
w, d, o, a, oracle = build()
calls = []
if damage == 'unknown-actor':
a = substitute(a, variant_id='typo', step_id=a.scenario.steps[-1].id, actor_id='unknown')
elif damage == 'cast-mismatch':
w.cast['alice'].id = 'different'
else:
steps = list(a.scenario.steps)
steps[-1] = replace(steps[-1], id='' if damage == 'empty-step' else steps[0].id)
a.scenario = replace(a.scenario, steps=tuple(steps))
class Driver:
def realize(self, actor, action):
calls.append(action.name)
return d.realize(actor, action)
with pytest.raises(ValueError, match='scenario|actor|step|cast'):
Runner(w, Driver(), o, oracle).run(a, evidence_store=EvidenceStore(tmp_path))
assert calls == []
assert not w.sut.resources
assert not list(tmp_path.iterdir())

View file

@ -132,3 +132,18 @@ scope prose: the pilot remains live here, and model/browser/authoring experiment
remain TD-WP-0003-T01/T06/T07. Broader scheduling/lifecycle/adapter features remain
explicit scope limits to prioritize from actual pilot needs. No real-system,
paid-model or browser-engine execution was performed.
**2026-09-28 evidence/preflight follow-up — done (T02/T03).** Runner validates
JSON-native observations before evaluation/retention. Lossy types (including
tuples), invalid keys, non-finite numbers and cyclic data cannot silently change
meaning in JSON; a bad snapshot stops execution with an evidence-failure record
and pending assertions INCONCLUSIVE. Serialization uses the same strict value
contract without dataclass or custom-value coercion. Runner preflights every
scheduled actor, unique nonempty step id and cast key/id agreement before any
SUT action. Intentionally partial scenarios remain valid.
Validation: 12 new regressions, of which 10 failed before the fixes; **131 focused
tests passed**, and the full suite passed **440 tests in 169.07 seconds**.
`git diff --check` is clean. Decision: `aa6c3e31-614c-4030-990c-945acef08515`.
No new tasks/workplans; T05 remains waiting and this workplan remains blocked.