Reject aborted runs and incomplete classification evidence

Assistant: codex
Assistant-Model: gpt-6-astra
Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
This commit is contained in:
tegwick 2026-09-28 12:17:13 +02:00
parent 081fb8c73c
commit 7779768058
7 changed files with 311 additions and 3 deletions

View file

@ -144,3 +144,45 @@ independent observation/cleanup receipts and appropriate approved access. A
successful historical receipt cannot substitute for those. The existing T01,
T06 and T07 tasks retain the experiment blockers; this review creates no
production implementation commitment or new work record.
## Safety follow-up: aborted runs and incomplete evidence
A post-closeout review reproduced two false-pass paths that the original suite
missed. A forbidden-surface exception after a passing prefix returned PASS even
though later claims had never been judged. Separately, deleting almost every
judgment and all state snapshots still yielded UNCHANGED/safe-to-accept.
The runner now records `scheduled_steps` and `expected_judgments` from the
scenario before executing any action. Only claims attached to scheduled steps
are included; the existing one-step browser scenario remains intentionally
narrow. On a surface violation, every remaining scheduled claim and invariant
gets an explicit INCONCLUSIVE judgment and no fabricated observation. A prior
FAIL still dominates. An abort after the last assertion is also INCONCLUSIVE,
because completing the assertions does not mean the scheduled run completed.
The classifier validates both packs against their input manifests. It requires
exactly one judgment for every expected assertion/step pair and exactly one S1
realization, S2 realization check and nonempty S3 state snapshot for every
scheduled step, with the correct strata. Missing, duplicate or invalid verdict
records, unfinished runs, and missing identity/version fields prevent automatic
acceptance. The candidate must retain the baseline's step and judgment coverage.
Evidence failure returns AMBIGUOUS before any adaptation rule is considered.
This adds two serialized EvidencePack fields. Historical packs without these
manifests remain historical artifacts but classify as AMBIGUOUS; rerun the
scenario to obtain evidence eligible for automatic acceptance. Do not infer a
manifest from surviving observations, since that would preserve the original
hole. These checks establish structural completeness, not authenticity of a
pack or correctness of arbitrary caller-supplied predicates. The existing
actor/observer separation and deterministic oracle remain required.
Regression coverage in `tests/test_evidence_completeness.py` includes aborts at
each step, abort after the final assertion, preservation of an earlier failure,
individual missing strata/records, duplicate records, and identical truncation
of both packs. Complete ordinary runs and intentionally narrow scenarios retain
their existing behavior. The two defects are handled under the existing T08
readiness review; the external T01/T06/T07 blockers remain unchanged.
Validation after the safety fixes: **299 tests passed** (`python3 -m pytest -q`).
The new regression module had 55 failures and five passing controls before the
fix; `git diff --check` passes after it.