Reject aborted runs and incomplete classification evidence
Assistant: codex Assistant-Model: gpt-6-astra Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
This commit is contained in:
parent
081fb8c73c
commit
7779768058
7 changed files with 311 additions and 3 deletions
|
|
@ -144,3 +144,45 @@ independent observation/cleanup receipts and appropriate approved access. A
|
|||
successful historical receipt cannot substitute for those. The existing T01,
|
||||
T06 and T07 tasks retain the experiment blockers; this review creates no
|
||||
production implementation commitment or new work record.
|
||||
|
||||
## Safety follow-up: aborted runs and incomplete evidence
|
||||
|
||||
A post-closeout review reproduced two false-pass paths that the original suite
|
||||
missed. A forbidden-surface exception after a passing prefix returned PASS even
|
||||
though later claims had never been judged. Separately, deleting almost every
|
||||
judgment and all state snapshots still yielded UNCHANGED/safe-to-accept.
|
||||
|
||||
The runner now records `scheduled_steps` and `expected_judgments` from the
|
||||
scenario before executing any action. Only claims attached to scheduled steps
|
||||
are included; the existing one-step browser scenario remains intentionally
|
||||
narrow. On a surface violation, every remaining scheduled claim and invariant
|
||||
gets an explicit INCONCLUSIVE judgment and no fabricated observation. A prior
|
||||
FAIL still dominates. An abort after the last assertion is also INCONCLUSIVE,
|
||||
because completing the assertions does not mean the scheduled run completed.
|
||||
|
||||
The classifier validates both packs against their input manifests. It requires
|
||||
exactly one judgment for every expected assertion/step pair and exactly one S1
|
||||
realization, S2 realization check and nonempty S3 state snapshot for every
|
||||
scheduled step, with the correct strata. Missing, duplicate or invalid verdict
|
||||
records, unfinished runs, and missing identity/version fields prevent automatic
|
||||
acceptance. The candidate must retain the baseline's step and judgment coverage.
|
||||
Evidence failure returns AMBIGUOUS before any adaptation rule is considered.
|
||||
|
||||
This adds two serialized EvidencePack fields. Historical packs without these
|
||||
manifests remain historical artifacts but classify as AMBIGUOUS; rerun the
|
||||
scenario to obtain evidence eligible for automatic acceptance. Do not infer a
|
||||
manifest from surviving observations, since that would preserve the original
|
||||
hole. These checks establish structural completeness, not authenticity of a
|
||||
pack or correctness of arbitrary caller-supplied predicates. The existing
|
||||
actor/observer separation and deterministic oracle remain required.
|
||||
|
||||
Regression coverage in `tests/test_evidence_completeness.py` includes aborts at
|
||||
each step, abort after the final assertion, preservation of an earlier failure,
|
||||
individual missing strata/records, duplicate records, and identical truncation
|
||||
of both packs. Complete ordinary runs and intentionally narrow scenarios retain
|
||||
their existing behavior. The two defects are handled under the existing T08
|
||||
readiness review; the external T01/T06/T07 blockers remain unchanged.
|
||||
|
||||
Validation after the safety fixes: **299 tests passed** (`python3 -m pytest -q`).
|
||||
The new regression module had 55 failures and five passing controls before the
|
||||
fix; `git diff --check` passes after it.
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue