Reject lossy observations and validate schedules before execution
Assistant: codex Assistant-Model: gpt-6-astra Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
This commit is contained in:
parent
eaf5d348e4
commit
10077edb8c
8 changed files with 184 additions and 21 deletions
3
SCOPE.md
3
SCOPE.md
|
|
@ -42,7 +42,8 @@ time-window or independent cleanup adapters in this repository.
|
||||||
- Evidence storage is opt-in and local. It is not encryption, signing, automatic
|
- Evidence storage is opt-in and local. It is not encryption, signing, automatic
|
||||||
redaction, retention management or safe storage of arbitrary secrets. Unexpected
|
redaction, retention management or safe storage of arbitrary secrets. Unexpected
|
||||||
driver/observer exceptions still propagate; no finalized receipt is promised for
|
driver/observer exceptions still propagate; no finalized receipt is promised for
|
||||||
an interrupted/crashed process. Strict JSON rejects unsupported evidence values.
|
an interrupted/crashed process. JSON-native snapshots are validated before judgment; lossy values stop the run
|
||||||
|
with pending judgments INCONCLUSIVE. Full schedule/cast validation precedes actions.
|
||||||
- No automatic finding-to-regression pipeline, asset registry, maturity promotion,
|
- No automatic finding-to-regression pipeline, asset registry, maturity promotion,
|
||||||
thawing or retirement. Parent ids and mutation history provide limited lineage.
|
thawing or retirement. Parent ids and mutation history provide limited lineage.
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -31,8 +31,13 @@ publication of the same run yields one winner and FileExistsError for others.
|
||||||
Directory fsync uses the local POSIX filesystem API. Storage failures propagate;
|
Directory fsync uses the local POSIX filesystem API. Storage failures propagate;
|
||||||
a caller must not report retained evidence if a write failed.
|
a caller must not report retained evidence if a write failed.
|
||||||
|
|
||||||
Serialization rejects unsupported values, non-string mapping keys and non-finite
|
Observations must use JSON-native dictionaries with string keys, lists, strings,
|
||||||
numbers. Observations detach nested data when recorded, so later domain mutations
|
integers, finite floats, booleans and null. Tuples and custom value/container types
|
||||||
|
are rejected rather than coerced. Runner checks and detaches the snapshot before
|
||||||
|
postconditions or oracles evaluate it; invalid evidence stops further actions,
|
||||||
|
records an evidence failure and makes pending judgments INCONCLUSIVE. Already
|
||||||
|
observed failures remain FAIL. Serialization also rejects unsupported values,
|
||||||
|
non-string mapping keys, cycles and non-finite numbers. Observations detach nested data when recorded, so later domain mutations
|
||||||
do not rewrite earlier snapshots. Receipts include the final run verdict and
|
do not rewrite earlier snapshots. Receipts include the final run verdict and
|
||||||
asset id, parent, maturity, variant and adaptation history. Independent claim
|
asset id, parent, maturity, variant and adaptation history. Independent claim
|
||||||
revisions and scenario revisions bind evidence to the original run inputs.
|
revisions and scenario revisions bind evidence to the original run inputs.
|
||||||
|
|
@ -62,7 +67,9 @@ PY
|
||||||
```
|
```
|
||||||
|
|
||||||
`substitute(..., actor_id='other')` changes the scheduled actor; that actor must
|
`substitute(..., actor_id='other')` changes the scheduled actor; that actor must
|
||||||
exist in the execution world's cast. `arguments={...}` changes only existing
|
exist in the execution world's cast. Runner validates all scheduled actor
|
||||||
|
references, cast key/id agreement and nonempty unique step ids before any action;
|
||||||
|
invalid input raises ValueError with no execution or finalized receipt. `arguments={...}` changes only existing
|
||||||
argument keys and covers resource, tenant or privilege substitution without new
|
argument keys and covers resource, tenant or privilege substitution without new
|
||||||
framework concepts. Mutable argument data is copied. Unknown/duplicate step ids,
|
framework concepts. Mutable argument data is copied. Unknown/duplicate step ids,
|
||||||
unknown keys, empty actor ids and invalid variant ids are rejected. Callers choose
|
unknown keys, empty actor ids and invalid variant ids are rejected. Callers choose
|
||||||
|
|
@ -83,3 +90,5 @@ scenario intent and can still qualify as mechanical adaptation.
|
||||||
|
|
||||||
This API deliberately does not generate new assertions, skip/reorder steps,
|
This API deliberately does not generate new assertions, skip/reorder steps,
|
||||||
schedule races or infer required evidence from a successful implementation run.
|
schedule races or infer required evidence from a successful implementation run.
|
||||||
|
|
||||||
|
Evidence/preflight decision: `aa6c3e31-614c-4030-990c-945acef08515`.
|
||||||
|
|
|
||||||
|
|
@ -105,3 +105,16 @@ experiment records. The scope update and local implementation do not close those
|
||||||
validation gaps.
|
validation gaps.
|
||||||
|
|
||||||
Decision: `f3b35dce-6047-4a95-8083-4d7032d4e99f`.
|
Decision: `f3b35dce-6047-4a95-8083-4d7032d4e99f`.
|
||||||
|
|
||||||
|
|
||||||
|
## Evidence/preflight correction — 2026-09-28
|
||||||
|
|
||||||
|
A subsequent review reproduced tuple-to-list conversion changing an oracle result
|
||||||
|
on receipt replay, and a late unknown actor causing failure after earlier SUT
|
||||||
|
mutations. The follow-up under TD-WP-0004-T02/T03 rejects non-JSON-native snapshots
|
||||||
|
before judgment, retains a safe evidence-failure receipt with pending assertions
|
||||||
|
INCONCLUSIVE, and validates the whole schedule/cast before execution. This closes
|
||||||
|
the reproduced cases without changing claim definitions or expanding pilot scope.
|
||||||
|
|
||||||
|
Follow-up validation: **440 tests passed** in 169.07 seconds, including 12 new
|
||||||
|
regressions (10 reproduced failures before implementation).
|
||||||
|
|
|
||||||
|
|
@ -14,3 +14,4 @@ file is a pointer table, not a second source of truth.
|
||||||
| TD-WP-0003 T08 guard-integrity follow-up | Strict boolean judgments and intact isolation guards | `2ecac1c5-5254-4ce9-b092-b47c3e1283a9` | `docs/TestDriverGeneralisationReview.md` |
|
| TD-WP-0003 T08 guard-integrity follow-up | Strict boolean judgments and intact isolation guards | `2ecac1c5-5254-4ce9-b092-b47c3e1283a9` | `docs/TestDriverGeneralisationReview.md` |
|
||||||
| TD-WP-0003 T08 HTTP/surface follow-up | Origin-bound credentials and runner surface validation | `88490ee8-839d-40dc-affa-b00a1c68e3c5` | `docs/TestDriverGeneralisationReview.md` |
|
| TD-WP-0003 T08 HTTP/surface follow-up | Origin-bound credentials and runner surface validation | `88490ee8-839d-40dc-affa-b00a1c68e3c5` | `docs/TestDriverGeneralisationReview.md` |
|
||||||
| TD-WP-0004 | Evidence-backed scope, local receipts and intent-preserving variants | `f3b35dce-6047-4a95-8083-4d7032d4e99f` | `history/2026-09-28-121933-scope-intent-assessment.md` |
|
| TD-WP-0004 | Evidence-backed scope, local receipts and intent-preserving variants | `f3b35dce-6047-4a95-8083-4d7032d4e99f` | `history/2026-09-28-121933-scope-intent-assessment.md` |
|
||||||
|
| TD-WP-0004 evidence/preflight follow-up | Lossless observation types and schedule preflight | `aa6c3e31-614c-4030-990c-945acef08515` | `docs/TestDriverEvidenceAndVariants.md` |
|
||||||
|
|
|
||||||
|
|
@ -13,13 +13,35 @@ See docs/TestDriverClassificationDesign.md, Part A.
|
||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
|
|
||||||
import json
|
import json
|
||||||
|
import math
|
||||||
from copy import deepcopy
|
from copy import deepcopy
|
||||||
from dataclasses import dataclass, field, asdict, replace
|
from dataclasses import dataclass, field, fields, replace
|
||||||
from datetime import datetime, timezone
|
from datetime import datetime, timezone
|
||||||
from enum import Enum
|
from enum import Enum
|
||||||
from typing import Any
|
from typing import Any
|
||||||
|
|
||||||
|
|
||||||
|
class InvalidEvidence(TypeError):
|
||||||
|
"""Evidence cannot be represented losslessly as JSON-native values."""
|
||||||
|
|
||||||
|
|
||||||
|
def json_value(value, active=None):
|
||||||
|
"""Validate and detach without coercing tuples, keys or custom value types."""
|
||||||
|
if value is None or type(value) in (bool, int, str):
|
||||||
|
return value
|
||||||
|
if type(value) is float and math.isfinite(value):
|
||||||
|
return value
|
||||||
|
active = set() if active is None else active
|
||||||
|
if id(value) in active or len(active) >= 100:
|
||||||
|
raise InvalidEvidence("cyclic or excessively deep evidence")
|
||||||
|
nested = active | {id(value)}
|
||||||
|
if type(value) is list:
|
||||||
|
return [json_value(item, nested) for item in value]
|
||||||
|
if type(value) is dict and all(type(key) is str for key in value):
|
||||||
|
return {key: json_value(item, nested) for key, item in value.items()}
|
||||||
|
raise InvalidEvidence("evidence requires JSON-native values and string keys")
|
||||||
|
|
||||||
|
|
||||||
class Stratum(str, Enum):
|
class Stratum(str, Enum):
|
||||||
SURFACE = "S1"
|
SURFACE = "S1"
|
||||||
REALIZATION = "S2"
|
REALIZATION = "S2"
|
||||||
|
|
@ -81,19 +103,10 @@ class EvidencePack:
|
||||||
return [o for o in self.observations if o.stratum is stratum]
|
return [o for o in self.observations if o.stratum is stratum]
|
||||||
|
|
||||||
def to_json(self) -> str:
|
def to_json(self) -> str:
|
||||||
payload = asdict(self)
|
payload = {item.name: getattr(self, item.name) for item in fields(self)
|
||||||
|
if item.name != "observations"}
|
||||||
payload["observations"] = [
|
payload["observations"] = [
|
||||||
{**asdict(o), "stratum": o.stratum.value} for o in self.observations
|
{**{item.name: getattr(o, item.name) for item in fields(o)},
|
||||||
|
"stratum": o.stratum.value} for o in self.observations
|
||||||
]
|
]
|
||||||
def check_keys(value):
|
return json.dumps(json_value(payload), indent=2, sort_keys=True, allow_nan=False)
|
||||||
if isinstance(value, dict):
|
|
||||||
if any(not isinstance(key, str) for key in value):
|
|
||||||
raise TypeError("evidence objects require string keys")
|
|
||||||
for item in value.values():
|
|
||||||
check_keys(item)
|
|
||||||
elif isinstance(value, (list, tuple)):
|
|
||||||
for item in value:
|
|
||||||
check_keys(item)
|
|
||||||
|
|
||||||
check_keys(payload)
|
|
||||||
return json.dumps(payload, indent=2, sort_keys=True, allow_nan=False)
|
|
||||||
|
|
|
||||||
|
|
@ -16,7 +16,7 @@ from typing import Any
|
||||||
|
|
||||||
from .actions import SurfaceNotPermitted
|
from .actions import SurfaceNotPermitted
|
||||||
from .drivers import Driver, realize_step
|
from .drivers import Driver, realize_step
|
||||||
from .evidence import EvidencePack, Observation, Stratum
|
from .evidence import EvidencePack, Observation, Stratum, InvalidEvidence, json_value
|
||||||
from .observers import StateObserver
|
from .observers import StateObserver
|
||||||
from .oracles import Judgment, Oracle, Verdict, overall
|
from .oracles import Judgment, Oracle, Verdict, overall
|
||||||
from .revisions import intent_revisions, scenario_revision
|
from .revisions import intent_revisions, scenario_revision
|
||||||
|
|
@ -134,6 +134,17 @@ class Runner:
|
||||||
|
|
||||||
def run(self, asset: VerificationAsset, *, evidence_store=None) -> RunResult:
|
def run(self, asset: VerificationAsset, *, evidence_store=None) -> RunResult:
|
||||||
scenario: Scenario = asset.scenario
|
scenario: Scenario = asset.scenario
|
||||||
|
# Validate the whole schedule before touching any actor or SUT.
|
||||||
|
step_ids = set()
|
||||||
|
for key, actor in self._world.cast.actors.items():
|
||||||
|
if not isinstance(key, str) or not key or actor.id != key:
|
||||||
|
raise ValueError("cast keys must match nonempty actor identities")
|
||||||
|
for step in scenario.steps:
|
||||||
|
if not isinstance(step.id, str) or not step.id or step.id in step_ids:
|
||||||
|
raise ValueError("scenario step ids must be nonempty and unique")
|
||||||
|
step_ids.add(step.id)
|
||||||
|
if not isinstance(step.actor_id, str) or step.actor_id not in self._world.cast.actors:
|
||||||
|
raise ValueError("scenario step names an unknown actor")
|
||||||
run_id = f"run-{uuid.uuid4().hex[:12]}"
|
run_id = f"run-{uuid.uuid4().hex[:12]}"
|
||||||
pack = EvidencePack(
|
pack = EvidencePack(
|
||||||
run_id=run_id,
|
run_id=run_id,
|
||||||
|
|
@ -242,7 +253,19 @@ class Runner:
|
||||||
break
|
break
|
||||||
|
|
||||||
# --- S3: what is now true ------------------------------------
|
# --- S3: what is now true ------------------------------------
|
||||||
snapshot = self._observer.snapshot()
|
raw_snapshot = self._observer.snapshot()
|
||||||
|
try:
|
||||||
|
snapshot = json_value(raw_snapshot)
|
||||||
|
if type(snapshot) is not dict:
|
||||||
|
raise InvalidEvidence("observation snapshot must be a JSON object")
|
||||||
|
except InvalidEvidence as exc:
|
||||||
|
self._record(
|
||||||
|
pack, Stratum.JUDGMENT, self._observer.name, "evidence_failure",
|
||||||
|
{"reason": str(exc)}, step.id,
|
||||||
|
)
|
||||||
|
aborted = True
|
||||||
|
skip_remaining(step_index, "observation cannot be retained losslessly")
|
||||||
|
break
|
||||||
self._record(
|
self._record(
|
||||||
pack, Stratum.JUDGMENT, self._observer.name,
|
pack, Stratum.JUDGMENT, self._observer.name,
|
||||||
"state_snapshot", dict(snapshot), step.id,
|
"state_snapshot", dict(snapshot), step.id,
|
||||||
|
|
|
||||||
88
tests/test_run_preflight_and_json.py
Normal file
88
tests/test_run_preflight_and_json.py
Normal file
|
|
@ -0,0 +1,88 @@
|
||||||
|
"""Reject unsafe run inputs before effects and lossy observation types before judgment."""
|
||||||
|
from dataclasses import replace
|
||||||
|
import json
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
from scenarios.alice_bob_carol import build
|
||||||
|
from testdriver import Runner, Verdict
|
||||||
|
from testdriver.classification import classify
|
||||||
|
from testdriver.storage import EvidenceStore
|
||||||
|
from testdriver.variants import substitute
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize('bad', [('a', 'b'), {'nested': [('a', 'b')]}, {1: 'value'},
|
||||||
|
float('nan'), float('inf')])
|
||||||
|
def test_lossy_observations_never_produce_a_verdict_that_cannot_be_replayed(tmp_path, bad):
|
||||||
|
w, d, o, a, oracle = build()
|
||||||
|
calls = []
|
||||||
|
case = a.scenario.use_case
|
||||||
|
claim = replace(case.claims[0], predicate=lambda snapshot: snapshot['pair'] == ('a', 'b'))
|
||||||
|
a.scenario = replace(a.scenario, use_case=replace(case, claims=(claim, *case.claims[1:])))
|
||||||
|
|
||||||
|
class Observer:
|
||||||
|
name = o.name
|
||||||
|
def snapshot(self):
|
||||||
|
return {**o.snapshot(), 'pair': bad}
|
||||||
|
|
||||||
|
class Driver:
|
||||||
|
def realize(self, actor, action):
|
||||||
|
calls.append(action.name)
|
||||||
|
return d.realize(actor, action)
|
||||||
|
|
||||||
|
store = EvidenceStore(tmp_path)
|
||||||
|
result = Runner(w, Driver(), Observer(), oracle).run(a, evidence_store=store)
|
||||||
|
assert result.verdict is Verdict.INCONCLUSIVE
|
||||||
|
assert len(calls) == 1
|
||||||
|
assert all(j.verdict is Verdict.INCONCLUSIVE for j in result.judgments)
|
||||||
|
retained = store.load(result.run_id)
|
||||||
|
assert retained['run_verdict'] == 'INCONCLUSIVE'
|
||||||
|
assert any(o['kind'] == 'evidence_failure' for o in retained['observations'])
|
||||||
|
assert not classify(retained, retained).safe_to_accept
|
||||||
|
|
||||||
|
|
||||||
|
def test_json_native_observations_keep_live_and_replayed_verdicts(tmp_path):
|
||||||
|
w, d, o, a, oracle = build()
|
||||||
|
result = Runner(w, d, o, oracle).run(a, evidence_store=EvidenceStore(tmp_path))
|
||||||
|
retained = EvidenceStore(tmp_path).load(result.run_id)
|
||||||
|
snapshots = {o['step_id']: o['data'] for o in retained['observations']
|
||||||
|
if o['kind'] == 'state_snapshot'}
|
||||||
|
assertions = {c.id: c for c in (*a.scenario.use_case.claims, *a.scenario.use_case.invariants)}
|
||||||
|
for verdict in retained['verdicts']:
|
||||||
|
replayed = oracle.judge(assertions[verdict['assertion_id']], snapshots[verdict['step_id']], verdict['step_id'])
|
||||||
|
assert replayed.verdict.value == verdict['verdict']
|
||||||
|
assert result.verdict is Verdict.PASS
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize('bad', [('a', 'b'), {1: 'value'}])
|
||||||
|
def test_serializer_does_not_silently_convert_unsupported_types(bad):
|
||||||
|
w, d, o, a, oracle = build()
|
||||||
|
pack = Runner(w, d, o, oracle).run(a).evidence
|
||||||
|
pack.asset['invalid'] = bad
|
||||||
|
with pytest.raises(TypeError):
|
||||||
|
pack.to_json()
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize('damage', ['unknown-actor', 'duplicate-step', 'empty-step', 'cast-mismatch'])
|
||||||
|
def test_entire_schedule_is_validated_before_any_side_effect(tmp_path, damage):
|
||||||
|
w, d, o, a, oracle = build()
|
||||||
|
calls = []
|
||||||
|
if damage == 'unknown-actor':
|
||||||
|
a = substitute(a, variant_id='typo', step_id=a.scenario.steps[-1].id, actor_id='unknown')
|
||||||
|
elif damage == 'cast-mismatch':
|
||||||
|
w.cast['alice'].id = 'different'
|
||||||
|
else:
|
||||||
|
steps = list(a.scenario.steps)
|
||||||
|
steps[-1] = replace(steps[-1], id='' if damage == 'empty-step' else steps[0].id)
|
||||||
|
a.scenario = replace(a.scenario, steps=tuple(steps))
|
||||||
|
|
||||||
|
class Driver:
|
||||||
|
def realize(self, actor, action):
|
||||||
|
calls.append(action.name)
|
||||||
|
return d.realize(actor, action)
|
||||||
|
|
||||||
|
with pytest.raises(ValueError, match='scenario|actor|step|cast'):
|
||||||
|
Runner(w, Driver(), o, oracle).run(a, evidence_store=EvidenceStore(tmp_path))
|
||||||
|
assert calls == []
|
||||||
|
assert not w.sut.resources
|
||||||
|
assert not list(tmp_path.iterdir())
|
||||||
|
|
@ -132,3 +132,18 @@ scope prose: the pilot remains live here, and model/browser/authoring experiment
|
||||||
remain TD-WP-0003-T01/T06/T07. Broader scheduling/lifecycle/adapter features remain
|
remain TD-WP-0003-T01/T06/T07. Broader scheduling/lifecycle/adapter features remain
|
||||||
explicit scope limits to prioritize from actual pilot needs. No real-system,
|
explicit scope limits to prioritize from actual pilot needs. No real-system,
|
||||||
paid-model or browser-engine execution was performed.
|
paid-model or browser-engine execution was performed.
|
||||||
|
|
||||||
|
|
||||||
|
**2026-09-28 evidence/preflight follow-up — done (T02/T03).** Runner validates
|
||||||
|
JSON-native observations before evaluation/retention. Lossy types (including
|
||||||
|
tuples), invalid keys, non-finite numbers and cyclic data cannot silently change
|
||||||
|
meaning in JSON; a bad snapshot stops execution with an evidence-failure record
|
||||||
|
and pending assertions INCONCLUSIVE. Serialization uses the same strict value
|
||||||
|
contract without dataclass or custom-value coercion. Runner preflights every
|
||||||
|
scheduled actor, unique nonempty step id and cast key/id agreement before any
|
||||||
|
SUT action. Intentionally partial scenarios remain valid.
|
||||||
|
|
||||||
|
Validation: 12 new regressions, of which 10 failed before the fixes; **131 focused
|
||||||
|
tests passed**, and the full suite passed **440 tests in 169.07 seconds**.
|
||||||
|
`git diff --check` is clean. Decision: `aa6c3e31-614c-4030-990c-945acef08515`.
|
||||||
|
No new tasks/workplans; T05 remains waiting and this workplan remains blocked.
|
||||||
|
|
|
||||||
Loading…
Add table
Add a link
Reference in a new issue