Require qualified runs and intent revisions for automatic acceptance
Assistant: codex Assistant-Model: gpt-6-astra Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
This commit is contained in:
parent
7779768058
commit
a8bb787d12
10 changed files with 486 additions and 14 deletions
|
|
@ -186,3 +186,61 @@ readiness review; the external T01/T06/T07 blockers remain unchanged.
|
|||
Validation after the safety fixes: **299 tests passed** (`python3 -m pytest -q`).
|
||||
The new regression module had 55 failures and five passing controls before the
|
||||
fix; `git diff --check` passes after it.
|
||||
|
||||
## Admission follow-up: crystallization and intent revisions
|
||||
|
||||
The next boundary review reproduced three additional holes: repeated aborted
|
||||
runs qualified as stable; a complete failing run compared with itself was
|
||||
UNCHANGED/safe-to-accept; and changing a claim's text or predicate while retaining
|
||||
its ID/provenance was invisible to the classifier. These fixes remain under T08.
|
||||
|
||||
Crystallization now uses the classifier's acceptance gate for every run against
|
||||
the first run. All packs must have complete evidence, passing judgments, verified
|
||||
postconditions, unchanged intent and successful realizations. Scenario/use-case
|
||||
identity, scheduled steps and expected judgments must match across the window.
|
||||
Repeated copies of one run receipt do not satisfy the independent-run count.
|
||||
Only then does trajectory comparison decide stability. A refused window returns
|
||||
no trajectories. An unchanged product failure is not a stable success.
|
||||
|
||||
The classifier requires a passing baseline. If the baseline contains FAIL or
|
||||
INCONCLUSIVE, even a repaired candidate requires a new passing reference rather
|
||||
than automatic acceptance against the failed baseline. Candidate failures against
|
||||
a passing reference still report BEHAVIOUR_CHANGED. Baseline realization must
|
||||
also be established before either automatic-acceptance outcome is returned;
|
||||
missing or non-boolean postconditions cannot be treated as success.
|
||||
|
||||
Evidence packs now include `intent_revisions`: SHA-256 revisions recorded before
|
||||
execution for the use case and each claim/invariant, including claims outside a
|
||||
partial scenario's schedule. They include narrative or assertion text, source,
|
||||
provenance, claim scheduling, and predicate implementation. The predicate hash
|
||||
covers bytecode/constants, defaults, keyword defaults, closure values, referenced
|
||||
Python helper functions and their referenced global values. It is bound to the
|
||||
Python implementation/version; checkout filenames and line numbers are excluded.
|
||||
Only hashes enter evidence, not captured values or code. Changed definitions
|
||||
produce INTENT_CHANGED, which requests review and does not assert human approval.
|
||||
An internally complete revised claim schedule is an intent change; missing
|
||||
baseline coverage under unchanged intent remains incomplete evidence.
|
||||
|
||||
No Claim/Invariant constructor change or claim serialization language is added.
|
||||
`revisions.py` is input-change detection for pure Python predicates. It does not
|
||||
prove semantic equivalence, sandbox Python, attest a caller-supplied pack or make
|
||||
mutable runtime behavior safe. Predicate dependencies must stay stable during a
|
||||
run. Cyclic/opaque callables, module dependencies and dynamic builtin lookups
|
||||
cannot be fingerprinted reliably and receive a null revision. They may still be
|
||||
evaluated by the oracle, but their evidence cannot be automatically accepted or
|
||||
crystallized. Refactor such predicates to consume plain independent snapshots.
|
||||
Changing Python versions conservatively requires intent review and fresh evidence.
|
||||
|
||||
Older packs without usable intent revisions require a rerun, just as packs
|
||||
without the earlier completeness manifest do. Existing historical experiment
|
||||
receipts are not retroactively upgraded or rewritten. Tests in
|
||||
`tests/test_acceptance_boundaries.py` cover the reproduced cases, helper/global
|
||||
and captured-value changes, invariant/narrative/source/scheduling changes,
|
||||
cross-process identity, source relocation, redaction of captured values, malformed
|
||||
postconditions, and the complete-passing control. No live-model, browser-engine
|
||||
or production readiness blocker is removed by this work.
|
||||
|
||||
Validation for this follow-up: initial reproduction had 20 failures and two
|
||||
passing controls; full suite **327 passed**. The final dynamic-dependency and
|
||||
strict-postcondition additions passed **32 boundary checks**, including four
|
||||
new cases. These counts overlap; they are not separate independent samples.
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue