Require qualified runs and intent revisions for automatic acceptance

Assistant: codex
Assistant-Model: gpt-6-astra
Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
This commit is contained in:
tegwick 2026-09-28 12:30:13 +02:00
parent 7779768058
commit a8bb787d12
10 changed files with 486 additions and 14 deletions

View file

@ -186,3 +186,61 @@ readiness review; the external T01/T06/T07 blockers remain unchanged.
Validation after the safety fixes: **299 tests passed** (`python3 -m pytest -q`). Validation after the safety fixes: **299 tests passed** (`python3 -m pytest -q`).
The new regression module had 55 failures and five passing controls before the The new regression module had 55 failures and five passing controls before the
fix; `git diff --check` passes after it. fix; `git diff --check` passes after it.
## Admission follow-up: crystallization and intent revisions
The next boundary review reproduced three additional holes: repeated aborted
runs qualified as stable; a complete failing run compared with itself was
UNCHANGED/safe-to-accept; and changing a claim's text or predicate while retaining
its ID/provenance was invisible to the classifier. These fixes remain under T08.
Crystallization now uses the classifier's acceptance gate for every run against
the first run. All packs must have complete evidence, passing judgments, verified
postconditions, unchanged intent and successful realizations. Scenario/use-case
identity, scheduled steps and expected judgments must match across the window.
Repeated copies of one run receipt do not satisfy the independent-run count.
Only then does trajectory comparison decide stability. A refused window returns
no trajectories. An unchanged product failure is not a stable success.
The classifier requires a passing baseline. If the baseline contains FAIL or
INCONCLUSIVE, even a repaired candidate requires a new passing reference rather
than automatic acceptance against the failed baseline. Candidate failures against
a passing reference still report BEHAVIOUR_CHANGED. Baseline realization must
also be established before either automatic-acceptance outcome is returned;
missing or non-boolean postconditions cannot be treated as success.
Evidence packs now include `intent_revisions`: SHA-256 revisions recorded before
execution for the use case and each claim/invariant, including claims outside a
partial scenario's schedule. They include narrative or assertion text, source,
provenance, claim scheduling, and predicate implementation. The predicate hash
covers bytecode/constants, defaults, keyword defaults, closure values, referenced
Python helper functions and their referenced global values. It is bound to the
Python implementation/version; checkout filenames and line numbers are excluded.
Only hashes enter evidence, not captured values or code. Changed definitions
produce INTENT_CHANGED, which requests review and does not assert human approval.
An internally complete revised claim schedule is an intent change; missing
baseline coverage under unchanged intent remains incomplete evidence.
No Claim/Invariant constructor change or claim serialization language is added.
`revisions.py` is input-change detection for pure Python predicates. It does not
prove semantic equivalence, sandbox Python, attest a caller-supplied pack or make
mutable runtime behavior safe. Predicate dependencies must stay stable during a
run. Cyclic/opaque callables, module dependencies and dynamic builtin lookups
cannot be fingerprinted reliably and receive a null revision. They may still be
evaluated by the oracle, but their evidence cannot be automatically accepted or
crystallized. Refactor such predicates to consume plain independent snapshots.
Changing Python versions conservatively requires intent review and fresh evidence.
Older packs without usable intent revisions require a rerun, just as packs
without the earlier completeness manifest do. Existing historical experiment
receipts are not retroactively upgraded or rewritten. Tests in
`tests/test_acceptance_boundaries.py` cover the reproduced cases, helper/global
and captured-value changes, invariant/narrative/source/scheduling changes,
cross-process identity, source relocation, redaction of captured values, malformed
postconditions, and the complete-passing control. No live-model, browser-engine
or production readiness blocker is removed by this work.
Validation for this follow-up: initial reproduction had 20 failures and two
passing controls; full suite **327 passed**. The final dynamic-dependency and
strict-postcondition additions passed **32 boundary checks**, including four
new cases. These counts overlap; they are not separate independent samples.

View file

@ -64,3 +64,10 @@ judgments, realization metrics and lineage remain. See
Previous compression (T10): `Verdict.SUSPICIOUS`, `Step.expect_refusal`, Previous compression (T10): `Verdict.SUSPICIOUS`, `Step.expect_refusal`,
`ActorIsolationError`, `World.seed`, `EvidencePack.latest()`, `Trajectory.method`. `ActorIsolationError`, `World.seed`, `EvidencePack.latest()`, `Trajectory.method`.
T08 admission follow-up: `revisions.py` implements conservative intent-definition
identity alongside `intent.py`/`provenance.py`. `classification.py` requires a
passing reference and complete revisioned evidence; `crystallization.py` consumes
that same admission rule. These close local safety holes, not new concepts or an
increase in external-validation level. See the generalisation review.

View file

@ -9,3 +9,4 @@ file is a pointer table, not a second source of truth.
| D-01 … D-07 | Stratified evidence and claim provenance as the basis for adaptation safety | `fef5213f-ce9b-44c2-b327-a0b0ba4b6270` | `docs/TestDriverClassificationDesign.md` | | D-01 … D-07 | Stratified evidence and claim provenance as the basis for adaptation safety | `fef5213f-ce9b-44c2-b327-a0b0ba4b6270` | `docs/TestDriverClassificationDesign.md` |
| TD-WP-0003 T04/T05/T08 | Python claims, concept removal and not-ready assessment | `d80734d1-01b8-4a98-8e72-c86f1806d572` | `docs/TestDriverGeneralisationReview.md` | | TD-WP-0003 T04/T05/T08 | Python claims, concept removal and not-ready assessment | `d80734d1-01b8-4a98-8e72-c86f1806d572` | `docs/TestDriverGeneralisationReview.md` |
| TD-WP-0003 T08 safety follow-up | Run-input manifests and incomplete-run handling | `afb38ade-893b-48a2-b1b3-c55e49cc40ee` | `docs/TestDriverGeneralisationReview.md` | | TD-WP-0003 T08 safety follow-up | Run-input manifests and incomplete-run handling | `afb38ade-893b-48a2-b1b3-c55e49cc40ee` | `docs/TestDriverGeneralisationReview.md` |
| TD-WP-0003 T08 admission follow-up | Qualified crystallization, passing baselines and intent revisions | `88e51e10-cd32-44d2-a2d6-055675b5ce04` | `docs/TestDriverGeneralisationReview.md` |

View file

@ -17,9 +17,9 @@ M19 in the lab prove it), so no amount of evidence separates them. What the
classifier can honestly say is *"behaviour changed against intent that did not"*, classifier can honestly say is *"behaviour changed against intent that did not"*,
and hand that to a human. and hand that to a human.
Intent change **is** detectable, but only when a human has actually changed the Intent change **is** detectable when recorded definitions change: the intent
intent: the claim fingerprint moves. That is a fact about the recorded use case, revision moves. That calls for review; it does not establish who approved the
not an inference about behaviour. change and is not an inference about behaviour.
""" """
from __future__ import annotations from __future__ import annotations
@ -45,7 +45,7 @@ class Classification(str, Enum):
""" """
INTENT_CHANGED = "INTENT_CHANGED" INTENT_CHANGED = "INTENT_CHANGED"
"""The recorded claim set itself changed. A human has already acted.""" """Recorded intent changed. Review its definitions and approval provenance."""
REALIZATION_FAILED = "REALIZATION_FAILED" REALIZATION_FAILED = "REALIZATION_FAILED"
"""The action could not be performed at all. Not a verdict about the system.""" """The action could not be performed at all. Not a verdict about the system."""
@ -132,7 +132,8 @@ def _surface_fingerprint(pack: Mapping[str, Any]) -> str:
def _claim_fingerprint(pack: Mapping[str, Any]) -> str: def _claim_fingerprint(pack: Mapping[str, Any]) -> str:
return json.dumps(sorted(pack.get("provenance_index", {}).items())) return json.dumps({"provenance": pack["provenance_index"],
"revisions": pack["intent_revisions"]}, sort_keys=True)
def _has_complete_evidence(pack: Mapping[str, Any]) -> bool: def _has_complete_evidence(pack: Mapping[str, Any]) -> bool:
@ -180,7 +181,15 @@ def _has_complete_evidence(pack: Mapping[str, Any]) -> bool:
return False return False
if any(o["kind"] == "surface_violation" for o in observations): if any(o["kind"] == "surface_violation" for o in observations):
return False return False
return isinstance(pack["provenance_index"], Mapping) revisions = pack["intent_revisions"]
required_ids = {assertion for assertion, _ in expected_keys} | {pack["use_case_id"]}
return (
isinstance(pack["provenance_index"], Mapping)
and isinstance(revisions, Mapping) and required_ids <= revisions.keys()
and all(isinstance(revision, str) and len(revision) == 64
and all(c in "0123456789abcdef" for c in revision)
for revision in revisions.values())
)
except (KeyError, TypeError, ValueError): except (KeyError, TypeError, ValueError):
# Missing/malformed records are evidence failure, not a default pass. # Missing/malformed records are evidence failure, not a default pass.
return False return False
@ -194,6 +203,7 @@ def extract_signals(baseline: Mapping[str, Any], candidate: Mapping[str, Any]) -
evidence_incomplete=True, evidence_incomplete=True,
) )
before, after = _verdict_map(baseline), _verdict_map(candidate) before, after = _verdict_map(baseline), _verdict_map(candidate)
claims_differ = _claim_fingerprint(baseline) != _claim_fingerprint(candidate)
regressed = any( regressed = any(
after.get(key) == "FAIL" and verdict != "FAIL" for key, verdict in before.items() after.get(key) == "FAIL" and verdict != "FAIL" for key, verdict in before.items()
@ -212,10 +222,10 @@ def extract_signals(baseline: Mapping[str, Any], candidate: Mapping[str, Any]) -
met = None met = None
elif any(c.get("postcondition_met") is False for c in checks): elif any(c.get("postcondition_met") is False for c in checks):
met = False met = False
elif any(c.get("postcondition_met") is None for c in checks): elif all(c.get("postcondition_met") is True for c in checks):
met = None
else:
met = True met = True
else:
met = None
return Signals( return Signals(
surface_differs=_surface_fingerprint(baseline) != _surface_fingerprint(candidate), surface_differs=_surface_fingerprint(baseline) != _surface_fingerprint(candidate),
@ -223,10 +233,12 @@ def extract_signals(baseline: Mapping[str, Any], candidate: Mapping[str, Any]) -
postcondition_met=met, postcondition_met=met,
verdicts_regressed=regressed, verdicts_regressed=regressed,
verdicts_inconclusive=any(v == "INCONCLUSIVE" for v in after.values()), verdicts_inconclusive=any(v == "INCONCLUSIVE" for v in after.values()),
claims_differ=_claim_fingerprint(baseline) != _claim_fingerprint(candidate), claims_differ=claims_differ,
evidence_incomplete=( evidence_incomplete=(
not claims_differ and (
not before.keys() <= after.keys() not before.keys() <= after.keys()
or not set(baseline["scheduled_steps"]) <= set(candidate["scheduled_steps"]) or not set(baseline["scheduled_steps"]) <= set(candidate["scheduled_steps"])
)
), ),
) )
@ -258,6 +270,10 @@ def classify(baseline: Mapping[str, Any], candidate: Mapping[str, Any]) -> Outco
if signals.evidence_incomplete: if signals.evidence_incomplete:
return Outcome(Classification.AMBIGUOUS, return Outcome(Classification.AMBIGUOUS,
"the run did not retain enough evidence to classify", signals) "the run did not retain enough evidence to classify", signals)
if any(v["verdict"] != "PASS" for v in baseline["verdicts"]):
return Outcome(Classification.AMBIGUOUS,
"the baseline is not a passing reference; obtain an accepted baseline",
signals)
failures = regressions(baseline, candidate) failures = regressions(baseline, candidate)
if signals.verdicts_inconclusive: if signals.verdicts_inconclusive:
return Outcome(Classification.AMBIGUOUS, return Outcome(Classification.AMBIGUOUS,
@ -267,8 +283,8 @@ def classify(baseline: Mapping[str, Any], candidate: Mapping[str, Any]) -> Outco
# 2. Intent moving is a fact about the recorded use case, not an inference. # 2. Intent moving is a fact about the recorded use case, not an inference.
if signals.claims_differ: if signals.claims_differ:
return Outcome(Classification.INTENT_CHANGED, return Outcome(Classification.INTENT_CHANGED,
"the recorded claim set differs from the baseline; " "the recorded intent differs from the baseline; "
"a human changed what is being asserted", signals, failures) "review the changed definitions and provenance", signals, failures)
# 3. A regression is reported before anything is allowed to explain it away. # 3. A regression is reported before anything is allowed to explain it away.
# This is decision-table row 3: coincidence is not exoneration. # This is decision-table row 3: coincidence is not exoneration.
@ -295,6 +311,15 @@ def classify(baseline: Mapping[str, Any], candidate: Mapping[str, Any]) -> Outco
return Outcome(Classification.AMBIGUOUS, return Outcome(Classification.AMBIGUOUS,
"the action's postcondition could not be evaluated", signals) "the action's postcondition could not be evaluated", signals)
# A passing verdict alone does not prove the reference actions took effect.
baseline_checks = [o["data"] for o in baseline["observations"]
if o["kind"] == "realization_check"]
if (any(c.get("postcondition_met") is not True for c in baseline_checks)
or any(o["kind"] == "realization" and o["data"].get("raised")
for o in baseline["observations"])):
return Outcome(Classification.AMBIGUOUS,
"the baseline does not establish successful realization", signals)
# 6. Only now, with every claim intact and the action verified, may a # 6. Only now, with every claim intact and the action verified, may a
# surface difference be called an adaptation. # surface difference be called an adaptation.
if signals.surface_differs: if signals.surface_differs:

View file

@ -24,6 +24,7 @@ from datetime import datetime, timezone
from typing import Any, Mapping, Sequence from typing import Any, Mapping, Sequence
from .actions import SemanticAction from .actions import SemanticAction
from .classification import classify
from .drivers import Realization from .drivers import Realization
from .world import Actor from .world import Actor
@ -90,6 +91,20 @@ def assess_stability(packs: Sequence[Mapping[str, Any]], minimum: int = 3) -> St
False, len(packs), 0, False, len(packs), 0,
f"need at least {minimum} runs to judge stability, have {len(packs)}", f"need at least {minimum} runs to judge stability, have {len(packs)}",
) )
if not packs:
return StabilityReport(False, 0, 0, "no runs to judge stability")
for pack in packs:
outcome = classify(packs[0], pack)
if not outcome.safe_to_accept:
return StabilityReport(False, len(packs), 0,
f"run is not eligible for crystallization: {outcome.reason}")
if (pack["scenario_id"], pack["use_case_id"], pack["scheduled_steps"],
pack["expected_judgments"]) != (
packs[0]["scenario_id"], packs[0]["use_case_id"], packs[0]["scheduled_steps"],
packs[0]["expected_judgments"]):
return StabilityReport(False, len(packs), 0, "scenario scope differs across runs")
if len({pack["run_id"] for pack in packs}) != len(packs):
return StabilityReport(False, len(packs), 0, "need distinct runs, not repeated receipts")
captured = [capture(pack) for pack in packs] captured = [capture(pack) for pack in packs]
keys = {tuple(t.key() for t in trajectory) for trajectory in captured} keys = {tuple(t.key() for t in trajectory) for trajectory in captured}
if len(keys) != 1: if len(keys) != 1:

View file

@ -67,6 +67,7 @@ class EvidencePack:
# Run inputs, captured before execution. A surviving prefix is not a manifest. # Run inputs, captured before execution. A surviving prefix is not a manifest.
scheduled_steps: list[str] = field(default_factory=list) scheduled_steps: list[str] = field(default_factory=list)
expected_judgments: list[dict[str, str]] = field(default_factory=list) expected_judgments: list[dict[str, str]] = field(default_factory=list)
intent_revisions: dict[str, str | None] = field(default_factory=dict)
def record(self, observation: Observation) -> None: def record(self, observation: Observation) -> None:
self.observations.append(observation) self.observations.append(observation)

113
src/testdriver/revisions.py Normal file
View file

@ -0,0 +1,113 @@
"""Conservative, process-independent revisions of Python intent inputs.
Only hashes leave this module: captured values may be sensitive. Unsupported
runtime dependencies produce no revision, never a name-only fallback. This is
change detection for pure predicates, not a proof of semantic equivalence.
"""
from __future__ import annotations
import builtins
import dis
import hashlib
import json
import sys
from types import CodeType, FunctionType
from .intent import UseCase
# These can access dependencies that bytecode name inspection cannot enumerate.
_DYNAMIC_BUILTINS = frozenset({
"eval", "exec", "globals", "locals", "vars", "getattr", "setattr", "delattr",
"__import__", "open", "input",
})
def _dump(value) -> str:
return json.dumps(value, sort_keys=True, separators=(",", ":"))
def _global_names(code: CodeType) -> set[str]:
names = {i.argval for i in dis.get_instructions(code)
if i.opname in ("LOAD_GLOBAL", "LOAD_NAME")}
for constant in code.co_consts:
if isinstance(constant, CodeType):
names.update(_global_names(constant))
return names
def _describe(value, active: set[int]):
"""Do not execute predicates, getters, reprs or arbitrary serialization hooks."""
if value is None or type(value) in (bool, int, str):
return [type(value).__name__, value]
if type(value) is float:
return ["float", value.hex()]
if type(value) is bytes:
return ["bytes", value.hex()]
if value is Ellipsis:
return ["ellipsis"]
if len(active) >= 32 or id(value) in active:
raise ValueError("recursive or overly deep predicate dependency")
active = active | {id(value)}
if type(value) in (tuple, list, set, frozenset):
items = [_describe(item, active) for item in value]
if type(value) in (set, frozenset):
items.sort(key=_dump)
return [type(value).__name__, items]
if type(value) is dict:
items = [[_describe(k, active), _describe(v, active)] for k, v in value.items()]
return ["dict", sorted(items, key=_dump)]
if isinstance(value, CodeType):
return ["code", value.co_code.hex(), _describe(value.co_consts, active),
value.co_names, value.co_varnames, value.co_freevars, value.co_cellvars,
value.co_argcount, value.co_posonlyargcount, value.co_kwonlyargcount,
value.co_flags, value.co_exceptiontable.hex()]
if isinstance(value, FunctionType):
dependencies = {}
for name in sorted(_global_names(value.__code__)):
if name in value.__globals__:
dependency = value.__globals__[name]
else:
dependency = value.__builtins__[name]
# Builtins are interpreter-version-bound, not opaque user callables.
if name in vars(builtins) and dependency is vars(builtins)[name]:
if name in _DYNAMIC_BUILTINS:
raise ValueError("dynamic predicate dependency")
dependencies[name] = ["builtin", name]
else:
dependencies[name] = _describe(dependency, active)
return ["function", _describe(value.__code__, active),
_describe(value.__defaults__, active), _describe(value.__kwdefaults__, active),
[_describe(cell.cell_contents, active) for cell in value.__closure__ or ()],
dependencies]
raise ValueError("unsupported predicate dependency")
def _digest(value) -> str:
return hashlib.sha256(_dump(value).encode()).hexdigest()
def intent_revisions(case: UseCase) -> dict[str, str | None]:
"""Capture intent before execution, including unscheduled use-case claims.
IDs remain stable; revisions change with their protected meaning. Predicate
bytecode is bound to the Python implementation/version. Code location and
line numbers are deliberately excluded so checkout paths do not cause drift.
"""
revisions: dict[str, str | None] = {
case.id: _digest(["use-case", case.title, case.narrative,
case.provenance.value, case.source_ref]),
}
for assertion in (*case.claims, *case.invariants):
if assertion.id in revisions:
return {} # Ambiguous identities cannot produce an admissible revision.
try:
predicate = _describe(assertion.predicate, set())
except (KeyError, TypeError, ValueError, RecursionError):
revisions[assertion.id] = None
continue
revisions[assertion.id] = _digest([
"python-intent-v1", sys.implementation.name, list(sys.version_info[:3]),
type(assertion).__name__, assertion.text, assertion.provenance.value,
assertion.source_ref, getattr(assertion, "after_step", None), predicate,
])
return revisions

View file

@ -18,6 +18,7 @@ from .drivers import Driver
from .evidence import EvidencePack, Observation, Stratum from .evidence import EvidencePack, Observation, Stratum
from .observers import StateObserver from .observers import StateObserver
from .oracles import Judgment, Oracle, Verdict, overall from .oracles import Judgment, Oracle, Verdict, overall
from .revisions import intent_revisions
from .scenario import Scenario, VerificationAsset from .scenario import Scenario, VerificationAsset
from .world import World from .world import World
@ -118,6 +119,7 @@ class Runner:
scenario_id=scenario.id, scenario_id=scenario.id,
use_case_id=scenario.use_case.id, use_case_id=scenario.use_case.id,
sut_version=self._world.sut_version, sut_version=self._world.sut_version,
intent_revisions=intent_revisions(scenario.use_case),
scheduled_steps=[step.id for step in scenario.steps], scheduled_steps=[step.id for step in scenario.steps],
expected_judgments=[ expected_judgments=[
{"assertion_id": assertion.id, "step_id": step.id} {"assertion_id": assertion.id, "step_id": step.id}

View file

@ -0,0 +1,228 @@
"""Admission, crystallization and intent revisions must agree on safe evidence."""
from copy import deepcopy
from dataclasses import replace
import json
import pytest
from scenarios.alice_bob_carol import build
from testdriver import Runner
from testdriver.classification import Classification, classify
from testdriver.crystallization import assess_stability
def pack(*mutations, change=None, abort=False):
world, driver, observer, asset, oracle = build(*mutations)
if change:
asset.scenario = replace(asset.scenario, use_case=change(asset.scenario.use_case))
if abort:
steps = list(asset.scenario.steps)
steps[1] = replace(steps[1], action=replace(
steps[1].action, permitted_surfaces=frozenset({"browser"})))
asset.scenario = replace(asset.scenario, steps=tuple(steps))
return json.loads(Runner(world, driver, observer, oracle).run(asset).evidence.to_json())
def change_claim(**changes):
return lambda case: replace(case, claims=tuple(
replace(c, **changes) if c.id == "c-bob-revoked" else c for c in case.claims
))
@pytest.mark.parametrize("mutation", ["M17", "M15", "M16", "M18"])
def test_failing_baseline_cannot_authorize_same_failure_or_fixed_run(mutation):
failing = pack(mutation)
assert any(v["verdict"] == "FAIL" for v in failing["verdicts"])
for candidate in (deepcopy(failing), pack()):
outcome = classify(failing, candidate)
assert not outcome.safe_to_accept
assert outcome.classification is Classification.AMBIGUOUS
assert "baseline" in outcome.reason
@pytest.mark.parametrize("damage", ["abort", "snapshots", "verdict", "failure", "inconclusive"])
def test_repeated_invalid_runs_cannot_crystallize(damage):
packs = [pack("M17") if damage == "failure" else pack(abort=damage == "abort")
for _ in range(3)]
for candidate in packs:
if damage == "snapshots":
candidate["observations"] = [o for o in candidate["observations"]
if o["kind"] != "state_snapshot"]
elif damage == "verdict":
candidate["verdicts"].pop()
elif damage == "inconclusive":
candidate["verdicts"][0]["verdict"] = "INCONCLUSIVE"
report = assess_stability(packs)
assert not report.stable
assert not report.trajectories
assert report.reason
def test_one_bad_run_in_an_otherwise_stable_window_prevents_freezing():
packs = [pack(), pack(abort=True), pack()]
assert not assess_stability(packs).stable
def test_replaying_one_receipt_does_not_count_as_three_runs():
baseline = pack()
assert not assess_stability([deepcopy(baseline) for _ in range(3)]).stable
def test_complete_passing_runs_still_classify_and_crystallize():
packs = [pack() for _ in range(3)]
assert classify(packs[0], packs[1]).classification is Classification.UNCHANGED
assert assess_stability(packs).stable
@pytest.mark.parametrize("changes", [
{"text": "Revocation is optional"},
{"source_ref": "new-independent-spec/revocation-v2"},
{"predicate": lambda obs: True},
])
def test_same_id_does_not_hide_a_changed_claim(changes):
before, after = pack(), pack(change=change_claim(**changes))
assert classify(before, after).classification is Classification.INTENT_CHANGED
assert not assess_stability([before, after, pack()]).stable
def test_invariant_changes_are_intent_changes_too():
def change(case):
return replace(case, invariants=(replace(case.invariants[0], predicate=lambda obs: True),
*case.invariants[1:]))
assert classify(pack(), pack(change=change)).classification is Classification.INTENT_CHANGED
def test_narrative_changes_require_intent_review():
assert classify(pack(), pack(change=lambda c: replace(c, narrative="New policy"))) \
.classification is Classification.INTENT_CHANGED
def threshold(value):
return lambda obs: len(obs["audit:R"]) >= value
def default_threshold(value):
return lambda obs, floor=value: len(obs["audit:R"]) >= floor
@pytest.mark.parametrize("factory", [threshold, default_threshold])
def test_captured_predicate_values_are_part_of_the_revision(factory):
first = pack(change=change_claim(predicate=factory(0)))
second = pack(change=change_claim(predicate=factory(1)))
assert classify(first, second).classification is Classification.INTENT_CHANGED
assert classify(first, pack(change=change_claim(predicate=factory(0)))).safe_to_accept
POLICY_FLOOR = 0
def helper(obs):
return len(obs["audit:R"]) >= POLICY_FLOOR
def predicate_using_helper(obs):
return helper(obs)
def test_global_values_and_helpers_are_part_of_the_revision(monkeypatch):
first = pack(change=change_claim(predicate=predicate_using_helper))
monkeypatch.setitem(globals(), "POLICY_FLOOR", 1)
second = pack(change=change_claim(predicate=predicate_using_helper))
assert classify(first, second).classification is Classification.INTENT_CHANGED
monkeypatch.setitem(globals(), "helper", lambda obs: True)
third = pack(change=change_claim(predicate=predicate_using_helper))
assert classify(second, third).classification is Classification.INTENT_CHANGED
def test_unidentifiable_callable_cannot_be_automatically_accepted():
class Opaque:
def __call__(self, obs):
return True
packs = [pack(change=change_claim(predicate=Opaque())) for _ in range(3)]
assert not classify(packs[0], packs[1]).safe_to_accept
assert not assess_stability(packs).stable
def test_legacy_pack_without_intent_revisions_requires_rerun():
before, after = pack(), pack()
after.pop("intent_revisions", None)
assert not classify(before, after).safe_to_accept
assert not assess_stability([after, pack(), pack()]).stable
def test_predicate_location_does_not_change_revision():
from types import FunctionType
from scenarios.alice_bob_carol import USE_CASE
from testdriver.revisions import intent_revisions
original = USE_CASE.claims[0].predicate
moved = FunctionType(original.__code__.replace(
co_filename="/a/different/checkout/scenario.py", co_firstlineno=999,
), original.__globals__)
changed = replace(USE_CASE, claims=(replace(USE_CASE.claims[0], predicate=moved),
*USE_CASE.claims[1:]))
assert intent_revisions(USE_CASE) == intent_revisions(changed)
def test_revisions_are_stable_across_processes():
import os
import subprocess
import sys
from pathlib import Path
from scenarios.alice_bob_carol import USE_CASE
from testdriver.revisions import intent_revisions
root = Path(__file__).resolve().parents[1]
output = subprocess.check_output([
sys.executable, "-c",
"import json; from scenarios.alice_bob_carol import USE_CASE; "
"from testdriver.revisions import intent_revisions; "
"print(json.dumps(intent_revisions(USE_CASE)))",
], cwd=root, env={**os.environ, "PYTHONPATH": os.pathsep.join((str(root / "src"), str(root)))},
text=True)
assert json.loads(output) == intent_revisions(USE_CASE)
def test_captured_values_are_hashed_not_retained():
secret = "synthetic-private-dependency-do-not-retain"
predicate = lambda obs: bool(secret) and obs["probe_read:bob:R"] is False
result = pack(change=change_claim(predicate=predicate))
assert result["intent_revisions"]["c-bob-revoked"] is not None
assert secret not in json.dumps(result)
def test_schedule_change_is_an_intent_change():
before = pack()
after = pack(change=change_claim(after_step="s1-create", predicate=lambda obs: True))
assert classify(before, after).classification is Classification.INTENT_CHANGED
def test_failed_baseline_realization_does_not_authorize_automatic_acceptance():
before, after = pack(), pack()
next(o for o in before["observations"]
if o["kind"] == "realization")["data"]["raised"] = "Unrealized"
assert not classify(before, after).safe_to_accept
def test_unknown_baseline_postcondition_does_not_authorize_automatic_acceptance():
before, after = pack(), pack()
next(o for o in before["observations"]
if o["kind"] == "realization_check")["data"]["postcondition_met"] = None
assert not classify(before, after).safe_to_accept
def test_dynamic_global_lookup_is_not_given_a_misleading_revision():
def dynamic(obs):
return globals()["POLICY_FLOOR"] <= len(obs["audit:R"])
before, after = pack(change=change_claim(predicate=dynamic)), pack(change=change_claim(predicate=dynamic))
assert before["intent_revisions"]["c-bob-revoked"] is None
assert not classify(before, after).safe_to_accept
@pytest.mark.parametrize("invalid", [None, "true", 1])
def test_one_unverified_postcondition_prevents_acceptance_and_freezing(invalid):
runs = [pack() for _ in range(3)]
next(o for o in runs[1]["observations"]
if o["kind"] == "realization_check")["data"]["postcondition_met"] = invalid
assert not classify(runs[0], runs[1]).safe_to_accept
assert not assess_stability(runs).stable

View file

@ -315,3 +315,25 @@ Validation: the new regression module produced 55 failures before the fix; all
intentional partial scenarios. `git diff --check` is clean. Decision: intentional partial scenarios. `git diff --check` is clean. Decision:
`afb38ade-893b-48a2-b1b3-c55e49cc40ee`. No new task or workplan was created; `afb38ade-893b-48a2-b1b3-c55e49cc40ee`. No new task or workplan was created;
T01/T06/T07 and the blocked workplan state are unchanged. T01/T06/T07 and the blocked workplan state are unchanged.
**2026-09-28 admission follow-up — done.** Crystallization now consumes the
classifier's evidence/acceptance gate for every run and requires matching scope,
unchanged intent, successful realization and distinct receipts before assessing
trajectory stability. Failed or inconclusive baselines cannot authorize automatic
acceptance, including repeated identical failures or repaired candidates.
Evidence records conservative definition revisions for use cases, claims and
invariants, including text/source/provenance/scheduling and Python predicate
code/defaults/captures/helpers/globals. Changed definitions trigger intent review;
unsupported/dynamic predicate dependencies and legacy revision-free packs remain
ineligible for automatic acceptance. Revisions retain hashes, not captured values;
Python version changes require fresh evidence. Public Claim/Invariant signatures
remain unchanged. Decision: `88e51e10-cd32-44d2-a2d6-055675b5ce04`.
Validation: the initial regressions produced 20 failures and two passing controls.
The full suite then passed **327 tests**; final dynamic-dependency and strict
postcondition guards passed **32 boundary tests** (including four new cases).
The earlier classification/crystallization subset passed 69 tests. Process/path
stability and captured-value redaction are covered; `git diff --check` is clean.
No new task or workplan; T01/T06/T07 remain waiting and this workplan stays blocked.