Require qualified runs and intent revisions for automatic acceptance

Assistant: codex
Assistant-Model: gpt-6-astra
Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
This commit is contained in:
tegwick 2026-09-28 12:30:13 +02:00
parent 7779768058
commit a8bb787d12
10 changed files with 486 additions and 14 deletions

View file

@ -186,3 +186,61 @@ readiness review; the external T01/T06/T07 blockers remain unchanged.
Validation after the safety fixes: **299 tests passed** (`python3 -m pytest -q`).
The new regression module had 55 failures and five passing controls before the
fix; `git diff --check` passes after it.
## Admission follow-up: crystallization and intent revisions
The next boundary review reproduced three additional holes: repeated aborted
runs qualified as stable; a complete failing run compared with itself was
UNCHANGED/safe-to-accept; and changing a claim's text or predicate while retaining
its ID/provenance was invisible to the classifier. These fixes remain under T08.
Crystallization now uses the classifier's acceptance gate for every run against
the first run. All packs must have complete evidence, passing judgments, verified
postconditions, unchanged intent and successful realizations. Scenario/use-case
identity, scheduled steps and expected judgments must match across the window.
Repeated copies of one run receipt do not satisfy the independent-run count.
Only then does trajectory comparison decide stability. A refused window returns
no trajectories. An unchanged product failure is not a stable success.
The classifier requires a passing baseline. If the baseline contains FAIL or
INCONCLUSIVE, even a repaired candidate requires a new passing reference rather
than automatic acceptance against the failed baseline. Candidate failures against
a passing reference still report BEHAVIOUR_CHANGED. Baseline realization must
also be established before either automatic-acceptance outcome is returned;
missing or non-boolean postconditions cannot be treated as success.
Evidence packs now include `intent_revisions`: SHA-256 revisions recorded before
execution for the use case and each claim/invariant, including claims outside a
partial scenario's schedule. They include narrative or assertion text, source,
provenance, claim scheduling, and predicate implementation. The predicate hash
covers bytecode/constants, defaults, keyword defaults, closure values, referenced
Python helper functions and their referenced global values. It is bound to the
Python implementation/version; checkout filenames and line numbers are excluded.
Only hashes enter evidence, not captured values or code. Changed definitions
produce INTENT_CHANGED, which requests review and does not assert human approval.
An internally complete revised claim schedule is an intent change; missing
baseline coverage under unchanged intent remains incomplete evidence.
No Claim/Invariant constructor change or claim serialization language is added.
`revisions.py` is input-change detection for pure Python predicates. It does not
prove semantic equivalence, sandbox Python, attest a caller-supplied pack or make
mutable runtime behavior safe. Predicate dependencies must stay stable during a
run. Cyclic/opaque callables, module dependencies and dynamic builtin lookups
cannot be fingerprinted reliably and receive a null revision. They may still be
evaluated by the oracle, but their evidence cannot be automatically accepted or
crystallized. Refactor such predicates to consume plain independent snapshots.
Changing Python versions conservatively requires intent review and fresh evidence.
Older packs without usable intent revisions require a rerun, just as packs
without the earlier completeness manifest do. Existing historical experiment
receipts are not retroactively upgraded or rewritten. Tests in
`tests/test_acceptance_boundaries.py` cover the reproduced cases, helper/global
and captured-value changes, invariant/narrative/source/scheduling changes,
cross-process identity, source relocation, redaction of captured values, malformed
postconditions, and the complete-passing control. No live-model, browser-engine
or production readiness blocker is removed by this work.
Validation for this follow-up: initial reproduction had 20 failures and two
passing controls; full suite **327 passed**. The final dynamic-dependency and
strict-postcondition additions passed **32 boundary checks**, including four
new cases. These counts overlap; they are not separate independent samples.