diff --git a/docs/TestDriverGeneralisationReview.md b/docs/TestDriverGeneralisationReview.md index b24922d..7ae9bf5 100644 --- a/docs/TestDriverGeneralisationReview.md +++ b/docs/TestDriverGeneralisationReview.md @@ -186,3 +186,61 @@ readiness review; the external T01/T06/T07 blockers remain unchanged. Validation after the safety fixes: **299 tests passed** (`python3 -m pytest -q`). The new regression module had 55 failures and five passing controls before the fix; `git diff --check` passes after it. + +## Admission follow-up: crystallization and intent revisions + +The next boundary review reproduced three additional holes: repeated aborted +runs qualified as stable; a complete failing run compared with itself was +UNCHANGED/safe-to-accept; and changing a claim's text or predicate while retaining +its ID/provenance was invisible to the classifier. These fixes remain under T08. + +Crystallization now uses the classifier's acceptance gate for every run against +the first run. All packs must have complete evidence, passing judgments, verified +postconditions, unchanged intent and successful realizations. Scenario/use-case +identity, scheduled steps and expected judgments must match across the window. +Repeated copies of one run receipt do not satisfy the independent-run count. +Only then does trajectory comparison decide stability. A refused window returns +no trajectories. An unchanged product failure is not a stable success. + +The classifier requires a passing baseline. If the baseline contains FAIL or +INCONCLUSIVE, even a repaired candidate requires a new passing reference rather +than automatic acceptance against the failed baseline. Candidate failures against +a passing reference still report BEHAVIOUR_CHANGED. Baseline realization must +also be established before either automatic-acceptance outcome is returned; +missing or non-boolean postconditions cannot be treated as success. + +Evidence packs now include `intent_revisions`: SHA-256 revisions recorded before +execution for the use case and each claim/invariant, including claims outside a +partial scenario's schedule. They include narrative or assertion text, source, +provenance, claim scheduling, and predicate implementation. The predicate hash +covers bytecode/constants, defaults, keyword defaults, closure values, referenced +Python helper functions and their referenced global values. It is bound to the +Python implementation/version; checkout filenames and line numbers are excluded. +Only hashes enter evidence, not captured values or code. Changed definitions +produce INTENT_CHANGED, which requests review and does not assert human approval. +An internally complete revised claim schedule is an intent change; missing +baseline coverage under unchanged intent remains incomplete evidence. + +No Claim/Invariant constructor change or claim serialization language is added. +`revisions.py` is input-change detection for pure Python predicates. It does not +prove semantic equivalence, sandbox Python, attest a caller-supplied pack or make +mutable runtime behavior safe. Predicate dependencies must stay stable during a +run. Cyclic/opaque callables, module dependencies and dynamic builtin lookups +cannot be fingerprinted reliably and receive a null revision. They may still be +evaluated by the oracle, but their evidence cannot be automatically accepted or +crystallized. Refactor such predicates to consume plain independent snapshots. +Changing Python versions conservatively requires intent review and fresh evidence. + +Older packs without usable intent revisions require a rerun, just as packs +without the earlier completeness manifest do. Existing historical experiment +receipts are not retroactively upgraded or rewritten. Tests in +`tests/test_acceptance_boundaries.py` cover the reproduced cases, helper/global +and captured-value changes, invariant/narrative/source/scheduling changes, +cross-process identity, source relocation, redaction of captured values, malformed +postconditions, and the complete-passing control. No live-model, browser-engine +or production readiness blocker is removed by this work. + +Validation for this follow-up: initial reproduction had 20 failures and two +passing controls; full suite **327 passed**. The final dynamic-dependency and +strict-postcondition additions passed **32 boundary checks**, including four +new cases. These counts overlap; they are not separate independent samples. diff --git a/research/concepts/fitness-map.md b/research/concepts/fitness-map.md index 700d954..5a6a547 100644 --- a/research/concepts/fitness-map.md +++ b/research/concepts/fitness-map.md @@ -64,3 +64,10 @@ judgments, realization metrics and lineage remain. See Previous compression (T10): `Verdict.SUSPICIOUS`, `Step.expect_refusal`, `ActorIsolationError`, `World.seed`, `EvidencePack.latest()`, `Trajectory.method`. + + +T08 admission follow-up: `revisions.py` implements conservative intent-definition +identity alongside `intent.py`/`provenance.py`. `classification.py` requires a +passing reference and complete revisioned evidence; `crystallization.py` consumes +that same admission rule. These close local safety holes, not new concepts or an +increase in external-validation level. See the generalisation review. diff --git a/research/decisions/README.md b/research/decisions/README.md index d060457..68d7199 100644 --- a/research/decisions/README.md +++ b/research/decisions/README.md @@ -9,3 +9,4 @@ file is a pointer table, not a second source of truth. | D-01 … D-07 | Stratified evidence and claim provenance as the basis for adaptation safety | `fef5213f-ce9b-44c2-b327-a0b0ba4b6270` | `docs/TestDriverClassificationDesign.md` | | TD-WP-0003 T04/T05/T08 | Python claims, concept removal and not-ready assessment | `d80734d1-01b8-4a98-8e72-c86f1806d572` | `docs/TestDriverGeneralisationReview.md` | | TD-WP-0003 T08 safety follow-up | Run-input manifests and incomplete-run handling | `afb38ade-893b-48a2-b1b3-c55e49cc40ee` | `docs/TestDriverGeneralisationReview.md` | +| TD-WP-0003 T08 admission follow-up | Qualified crystallization, passing baselines and intent revisions | `88e51e10-cd32-44d2-a2d6-055675b5ce04` | `docs/TestDriverGeneralisationReview.md` | diff --git a/src/testdriver/classification.py b/src/testdriver/classification.py index 5f9c6af..ab395e3 100644 --- a/src/testdriver/classification.py +++ b/src/testdriver/classification.py @@ -17,9 +17,9 @@ M19 in the lab prove it), so no amount of evidence separates them. What the classifier can honestly say is *"behaviour changed against intent that did not"*, and hand that to a human. -Intent change **is** detectable, but only when a human has actually changed the -intent: the claim fingerprint moves. That is a fact about the recorded use case, -not an inference about behaviour. +Intent change **is** detectable when recorded definitions change: the intent +revision moves. That calls for review; it does not establish who approved the +change and is not an inference about behaviour. """ from __future__ import annotations @@ -45,7 +45,7 @@ class Classification(str, Enum): """ INTENT_CHANGED = "INTENT_CHANGED" - """The recorded claim set itself changed. A human has already acted.""" + """Recorded intent changed. Review its definitions and approval provenance.""" REALIZATION_FAILED = "REALIZATION_FAILED" """The action could not be performed at all. Not a verdict about the system.""" @@ -132,7 +132,8 @@ def _surface_fingerprint(pack: Mapping[str, Any]) -> str: def _claim_fingerprint(pack: Mapping[str, Any]) -> str: - return json.dumps(sorted(pack.get("provenance_index", {}).items())) + return json.dumps({"provenance": pack["provenance_index"], + "revisions": pack["intent_revisions"]}, sort_keys=True) def _has_complete_evidence(pack: Mapping[str, Any]) -> bool: @@ -180,7 +181,15 @@ def _has_complete_evidence(pack: Mapping[str, Any]) -> bool: return False if any(o["kind"] == "surface_violation" for o in observations): return False - return isinstance(pack["provenance_index"], Mapping) + revisions = pack["intent_revisions"] + required_ids = {assertion for assertion, _ in expected_keys} | {pack["use_case_id"]} + return ( + isinstance(pack["provenance_index"], Mapping) + and isinstance(revisions, Mapping) and required_ids <= revisions.keys() + and all(isinstance(revision, str) and len(revision) == 64 + and all(c in "0123456789abcdef" for c in revision) + for revision in revisions.values()) + ) except (KeyError, TypeError, ValueError): # Missing/malformed records are evidence failure, not a default pass. return False @@ -194,6 +203,7 @@ def extract_signals(baseline: Mapping[str, Any], candidate: Mapping[str, Any]) - evidence_incomplete=True, ) before, after = _verdict_map(baseline), _verdict_map(candidate) + claims_differ = _claim_fingerprint(baseline) != _claim_fingerprint(candidate) regressed = any( after.get(key) == "FAIL" and verdict != "FAIL" for key, verdict in before.items() @@ -212,10 +222,10 @@ def extract_signals(baseline: Mapping[str, Any], candidate: Mapping[str, Any]) - met = None elif any(c.get("postcondition_met") is False for c in checks): met = False - elif any(c.get("postcondition_met") is None for c in checks): - met = None - else: + elif all(c.get("postcondition_met") is True for c in checks): met = True + else: + met = None return Signals( surface_differs=_surface_fingerprint(baseline) != _surface_fingerprint(candidate), @@ -223,10 +233,12 @@ def extract_signals(baseline: Mapping[str, Any], candidate: Mapping[str, Any]) - postcondition_met=met, verdicts_regressed=regressed, verdicts_inconclusive=any(v == "INCONCLUSIVE" for v in after.values()), - claims_differ=_claim_fingerprint(baseline) != _claim_fingerprint(candidate), + claims_differ=claims_differ, evidence_incomplete=( - not before.keys() <= after.keys() - or not set(baseline["scheduled_steps"]) <= set(candidate["scheduled_steps"]) + not claims_differ and ( + not before.keys() <= after.keys() + or not set(baseline["scheduled_steps"]) <= set(candidate["scheduled_steps"]) + ) ), ) @@ -258,6 +270,10 @@ def classify(baseline: Mapping[str, Any], candidate: Mapping[str, Any]) -> Outco if signals.evidence_incomplete: return Outcome(Classification.AMBIGUOUS, "the run did not retain enough evidence to classify", signals) + if any(v["verdict"] != "PASS" for v in baseline["verdicts"]): + return Outcome(Classification.AMBIGUOUS, + "the baseline is not a passing reference; obtain an accepted baseline", + signals) failures = regressions(baseline, candidate) if signals.verdicts_inconclusive: return Outcome(Classification.AMBIGUOUS, @@ -267,8 +283,8 @@ def classify(baseline: Mapping[str, Any], candidate: Mapping[str, Any]) -> Outco # 2. Intent moving is a fact about the recorded use case, not an inference. if signals.claims_differ: return Outcome(Classification.INTENT_CHANGED, - "the recorded claim set differs from the baseline; " - "a human changed what is being asserted", signals, failures) + "the recorded intent differs from the baseline; " + "review the changed definitions and provenance", signals, failures) # 3. A regression is reported before anything is allowed to explain it away. # This is decision-table row 3: coincidence is not exoneration. @@ -295,6 +311,15 @@ def classify(baseline: Mapping[str, Any], candidate: Mapping[str, Any]) -> Outco return Outcome(Classification.AMBIGUOUS, "the action's postcondition could not be evaluated", signals) + # A passing verdict alone does not prove the reference actions took effect. + baseline_checks = [o["data"] for o in baseline["observations"] + if o["kind"] == "realization_check"] + if (any(c.get("postcondition_met") is not True for c in baseline_checks) + or any(o["kind"] == "realization" and o["data"].get("raised") + for o in baseline["observations"])): + return Outcome(Classification.AMBIGUOUS, + "the baseline does not establish successful realization", signals) + # 6. Only now, with every claim intact and the action verified, may a # surface difference be called an adaptation. if signals.surface_differs: diff --git a/src/testdriver/crystallization.py b/src/testdriver/crystallization.py index 6e90483..9e69689 100644 --- a/src/testdriver/crystallization.py +++ b/src/testdriver/crystallization.py @@ -24,6 +24,7 @@ from datetime import datetime, timezone from typing import Any, Mapping, Sequence from .actions import SemanticAction +from .classification import classify from .drivers import Realization from .world import Actor @@ -90,6 +91,20 @@ def assess_stability(packs: Sequence[Mapping[str, Any]], minimum: int = 3) -> St False, len(packs), 0, f"need at least {minimum} runs to judge stability, have {len(packs)}", ) + if not packs: + return StabilityReport(False, 0, 0, "no runs to judge stability") + for pack in packs: + outcome = classify(packs[0], pack) + if not outcome.safe_to_accept: + return StabilityReport(False, len(packs), 0, + f"run is not eligible for crystallization: {outcome.reason}") + if (pack["scenario_id"], pack["use_case_id"], pack["scheduled_steps"], + pack["expected_judgments"]) != ( + packs[0]["scenario_id"], packs[0]["use_case_id"], packs[0]["scheduled_steps"], + packs[0]["expected_judgments"]): + return StabilityReport(False, len(packs), 0, "scenario scope differs across runs") + if len({pack["run_id"] for pack in packs}) != len(packs): + return StabilityReport(False, len(packs), 0, "need distinct runs, not repeated receipts") captured = [capture(pack) for pack in packs] keys = {tuple(t.key() for t in trajectory) for trajectory in captured} if len(keys) != 1: diff --git a/src/testdriver/evidence.py b/src/testdriver/evidence.py index b7261f8..d8abbf1 100644 --- a/src/testdriver/evidence.py +++ b/src/testdriver/evidence.py @@ -67,6 +67,7 @@ class EvidencePack: # Run inputs, captured before execution. A surviving prefix is not a manifest. scheduled_steps: list[str] = field(default_factory=list) expected_judgments: list[dict[str, str]] = field(default_factory=list) + intent_revisions: dict[str, str | None] = field(default_factory=dict) def record(self, observation: Observation) -> None: self.observations.append(observation) diff --git a/src/testdriver/revisions.py b/src/testdriver/revisions.py new file mode 100644 index 0000000..1d81bcb --- /dev/null +++ b/src/testdriver/revisions.py @@ -0,0 +1,113 @@ +"""Conservative, process-independent revisions of Python intent inputs. + +Only hashes leave this module: captured values may be sensitive. Unsupported +runtime dependencies produce no revision, never a name-only fallback. This is +change detection for pure predicates, not a proof of semantic equivalence. +""" +from __future__ import annotations + +import builtins +import dis +import hashlib +import json +import sys +from types import CodeType, FunctionType + +from .intent import UseCase + +# These can access dependencies that bytecode name inspection cannot enumerate. +_DYNAMIC_BUILTINS = frozenset({ + "eval", "exec", "globals", "locals", "vars", "getattr", "setattr", "delattr", + "__import__", "open", "input", +}) + + +def _dump(value) -> str: + return json.dumps(value, sort_keys=True, separators=(",", ":")) + + +def _global_names(code: CodeType) -> set[str]: + names = {i.argval for i in dis.get_instructions(code) + if i.opname in ("LOAD_GLOBAL", "LOAD_NAME")} + for constant in code.co_consts: + if isinstance(constant, CodeType): + names.update(_global_names(constant)) + return names + + +def _describe(value, active: set[int]): + """Do not execute predicates, getters, reprs or arbitrary serialization hooks.""" + if value is None or type(value) in (bool, int, str): + return [type(value).__name__, value] + if type(value) is float: + return ["float", value.hex()] + if type(value) is bytes: + return ["bytes", value.hex()] + if value is Ellipsis: + return ["ellipsis"] + if len(active) >= 32 or id(value) in active: + raise ValueError("recursive or overly deep predicate dependency") + active = active | {id(value)} + if type(value) in (tuple, list, set, frozenset): + items = [_describe(item, active) for item in value] + if type(value) in (set, frozenset): + items.sort(key=_dump) + return [type(value).__name__, items] + if type(value) is dict: + items = [[_describe(k, active), _describe(v, active)] for k, v in value.items()] + return ["dict", sorted(items, key=_dump)] + if isinstance(value, CodeType): + return ["code", value.co_code.hex(), _describe(value.co_consts, active), + value.co_names, value.co_varnames, value.co_freevars, value.co_cellvars, + value.co_argcount, value.co_posonlyargcount, value.co_kwonlyargcount, + value.co_flags, value.co_exceptiontable.hex()] + if isinstance(value, FunctionType): + dependencies = {} + for name in sorted(_global_names(value.__code__)): + if name in value.__globals__: + dependency = value.__globals__[name] + else: + dependency = value.__builtins__[name] + # Builtins are interpreter-version-bound, not opaque user callables. + if name in vars(builtins) and dependency is vars(builtins)[name]: + if name in _DYNAMIC_BUILTINS: + raise ValueError("dynamic predicate dependency") + dependencies[name] = ["builtin", name] + else: + dependencies[name] = _describe(dependency, active) + return ["function", _describe(value.__code__, active), + _describe(value.__defaults__, active), _describe(value.__kwdefaults__, active), + [_describe(cell.cell_contents, active) for cell in value.__closure__ or ()], + dependencies] + raise ValueError("unsupported predicate dependency") + + +def _digest(value) -> str: + return hashlib.sha256(_dump(value).encode()).hexdigest() + + +def intent_revisions(case: UseCase) -> dict[str, str | None]: + """Capture intent before execution, including unscheduled use-case claims. + + IDs remain stable; revisions change with their protected meaning. Predicate + bytecode is bound to the Python implementation/version. Code location and + line numbers are deliberately excluded so checkout paths do not cause drift. + """ + revisions: dict[str, str | None] = { + case.id: _digest(["use-case", case.title, case.narrative, + case.provenance.value, case.source_ref]), + } + for assertion in (*case.claims, *case.invariants): + if assertion.id in revisions: + return {} # Ambiguous identities cannot produce an admissible revision. + try: + predicate = _describe(assertion.predicate, set()) + except (KeyError, TypeError, ValueError, RecursionError): + revisions[assertion.id] = None + continue + revisions[assertion.id] = _digest([ + "python-intent-v1", sys.implementation.name, list(sys.version_info[:3]), + type(assertion).__name__, assertion.text, assertion.provenance.value, + assertion.source_ref, getattr(assertion, "after_step", None), predicate, + ]) + return revisions diff --git a/src/testdriver/runner.py b/src/testdriver/runner.py index 614acd6..189ff8f 100644 --- a/src/testdriver/runner.py +++ b/src/testdriver/runner.py @@ -18,6 +18,7 @@ from .drivers import Driver from .evidence import EvidencePack, Observation, Stratum from .observers import StateObserver from .oracles import Judgment, Oracle, Verdict, overall +from .revisions import intent_revisions from .scenario import Scenario, VerificationAsset from .world import World @@ -118,6 +119,7 @@ class Runner: scenario_id=scenario.id, use_case_id=scenario.use_case.id, sut_version=self._world.sut_version, + intent_revisions=intent_revisions(scenario.use_case), scheduled_steps=[step.id for step in scenario.steps], expected_judgments=[ {"assertion_id": assertion.id, "step_id": step.id} diff --git a/tests/test_acceptance_boundaries.py b/tests/test_acceptance_boundaries.py new file mode 100644 index 0000000..258963f --- /dev/null +++ b/tests/test_acceptance_boundaries.py @@ -0,0 +1,228 @@ +"""Admission, crystallization and intent revisions must agree on safe evidence.""" +from copy import deepcopy +from dataclasses import replace +import json + +import pytest + +from scenarios.alice_bob_carol import build +from testdriver import Runner +from testdriver.classification import Classification, classify +from testdriver.crystallization import assess_stability + + +def pack(*mutations, change=None, abort=False): + world, driver, observer, asset, oracle = build(*mutations) + if change: + asset.scenario = replace(asset.scenario, use_case=change(asset.scenario.use_case)) + if abort: + steps = list(asset.scenario.steps) + steps[1] = replace(steps[1], action=replace( + steps[1].action, permitted_surfaces=frozenset({"browser"}))) + asset.scenario = replace(asset.scenario, steps=tuple(steps)) + return json.loads(Runner(world, driver, observer, oracle).run(asset).evidence.to_json()) + + +def change_claim(**changes): + return lambda case: replace(case, claims=tuple( + replace(c, **changes) if c.id == "c-bob-revoked" else c for c in case.claims + )) + + +@pytest.mark.parametrize("mutation", ["M17", "M15", "M16", "M18"]) +def test_failing_baseline_cannot_authorize_same_failure_or_fixed_run(mutation): + failing = pack(mutation) + assert any(v["verdict"] == "FAIL" for v in failing["verdicts"]) + for candidate in (deepcopy(failing), pack()): + outcome = classify(failing, candidate) + assert not outcome.safe_to_accept + assert outcome.classification is Classification.AMBIGUOUS + assert "baseline" in outcome.reason + + +@pytest.mark.parametrize("damage", ["abort", "snapshots", "verdict", "failure", "inconclusive"]) +def test_repeated_invalid_runs_cannot_crystallize(damage): + packs = [pack("M17") if damage == "failure" else pack(abort=damage == "abort") + for _ in range(3)] + for candidate in packs: + if damage == "snapshots": + candidate["observations"] = [o for o in candidate["observations"] + if o["kind"] != "state_snapshot"] + elif damage == "verdict": + candidate["verdicts"].pop() + elif damage == "inconclusive": + candidate["verdicts"][0]["verdict"] = "INCONCLUSIVE" + report = assess_stability(packs) + assert not report.stable + assert not report.trajectories + assert report.reason + + +def test_one_bad_run_in_an_otherwise_stable_window_prevents_freezing(): + packs = [pack(), pack(abort=True), pack()] + assert not assess_stability(packs).stable + + +def test_replaying_one_receipt_does_not_count_as_three_runs(): + baseline = pack() + assert not assess_stability([deepcopy(baseline) for _ in range(3)]).stable + + +def test_complete_passing_runs_still_classify_and_crystallize(): + packs = [pack() for _ in range(3)] + assert classify(packs[0], packs[1]).classification is Classification.UNCHANGED + assert assess_stability(packs).stable + + +@pytest.mark.parametrize("changes", [ + {"text": "Revocation is optional"}, + {"source_ref": "new-independent-spec/revocation-v2"}, + {"predicate": lambda obs: True}, +]) +def test_same_id_does_not_hide_a_changed_claim(changes): + before, after = pack(), pack(change=change_claim(**changes)) + assert classify(before, after).classification is Classification.INTENT_CHANGED + assert not assess_stability([before, after, pack()]).stable + + +def test_invariant_changes_are_intent_changes_too(): + def change(case): + return replace(case, invariants=(replace(case.invariants[0], predicate=lambda obs: True), + *case.invariants[1:])) + assert classify(pack(), pack(change=change)).classification is Classification.INTENT_CHANGED + + +def test_narrative_changes_require_intent_review(): + assert classify(pack(), pack(change=lambda c: replace(c, narrative="New policy"))) \ + .classification is Classification.INTENT_CHANGED + + +def threshold(value): + return lambda obs: len(obs["audit:R"]) >= value + + +def default_threshold(value): + return lambda obs, floor=value: len(obs["audit:R"]) >= floor + + +@pytest.mark.parametrize("factory", [threshold, default_threshold]) +def test_captured_predicate_values_are_part_of_the_revision(factory): + first = pack(change=change_claim(predicate=factory(0))) + second = pack(change=change_claim(predicate=factory(1))) + assert classify(first, second).classification is Classification.INTENT_CHANGED + assert classify(first, pack(change=change_claim(predicate=factory(0)))).safe_to_accept + + +POLICY_FLOOR = 0 + + +def helper(obs): + return len(obs["audit:R"]) >= POLICY_FLOOR + + +def predicate_using_helper(obs): + return helper(obs) + + +def test_global_values_and_helpers_are_part_of_the_revision(monkeypatch): + first = pack(change=change_claim(predicate=predicate_using_helper)) + monkeypatch.setitem(globals(), "POLICY_FLOOR", 1) + second = pack(change=change_claim(predicate=predicate_using_helper)) + assert classify(first, second).classification is Classification.INTENT_CHANGED + monkeypatch.setitem(globals(), "helper", lambda obs: True) + third = pack(change=change_claim(predicate=predicate_using_helper)) + assert classify(second, third).classification is Classification.INTENT_CHANGED + + +def test_unidentifiable_callable_cannot_be_automatically_accepted(): + class Opaque: + def __call__(self, obs): + return True + packs = [pack(change=change_claim(predicate=Opaque())) for _ in range(3)] + assert not classify(packs[0], packs[1]).safe_to_accept + assert not assess_stability(packs).stable + + +def test_legacy_pack_without_intent_revisions_requires_rerun(): + before, after = pack(), pack() + after.pop("intent_revisions", None) + assert not classify(before, after).safe_to_accept + assert not assess_stability([after, pack(), pack()]).stable + + +def test_predicate_location_does_not_change_revision(): + from types import FunctionType + from scenarios.alice_bob_carol import USE_CASE + from testdriver.revisions import intent_revisions + + original = USE_CASE.claims[0].predicate + moved = FunctionType(original.__code__.replace( + co_filename="/a/different/checkout/scenario.py", co_firstlineno=999, + ), original.__globals__) + changed = replace(USE_CASE, claims=(replace(USE_CASE.claims[0], predicate=moved), + *USE_CASE.claims[1:])) + assert intent_revisions(USE_CASE) == intent_revisions(changed) + + +def test_revisions_are_stable_across_processes(): + import os + import subprocess + import sys + from pathlib import Path + from scenarios.alice_bob_carol import USE_CASE + from testdriver.revisions import intent_revisions + + root = Path(__file__).resolve().parents[1] + output = subprocess.check_output([ + sys.executable, "-c", + "import json; from scenarios.alice_bob_carol import USE_CASE; " + "from testdriver.revisions import intent_revisions; " + "print(json.dumps(intent_revisions(USE_CASE)))", + ], cwd=root, env={**os.environ, "PYTHONPATH": os.pathsep.join((str(root / "src"), str(root)))}, + text=True) + assert json.loads(output) == intent_revisions(USE_CASE) + + +def test_captured_values_are_hashed_not_retained(): + secret = "synthetic-private-dependency-do-not-retain" + predicate = lambda obs: bool(secret) and obs["probe_read:bob:R"] is False + result = pack(change=change_claim(predicate=predicate)) + assert result["intent_revisions"]["c-bob-revoked"] is not None + assert secret not in json.dumps(result) + + +def test_schedule_change_is_an_intent_change(): + before = pack() + after = pack(change=change_claim(after_step="s1-create", predicate=lambda obs: True)) + assert classify(before, after).classification is Classification.INTENT_CHANGED + + +def test_failed_baseline_realization_does_not_authorize_automatic_acceptance(): + before, after = pack(), pack() + next(o for o in before["observations"] + if o["kind"] == "realization")["data"]["raised"] = "Unrealized" + assert not classify(before, after).safe_to_accept + + +def test_unknown_baseline_postcondition_does_not_authorize_automatic_acceptance(): + before, after = pack(), pack() + next(o for o in before["observations"] + if o["kind"] == "realization_check")["data"]["postcondition_met"] = None + assert not classify(before, after).safe_to_accept + + +def test_dynamic_global_lookup_is_not_given_a_misleading_revision(): + def dynamic(obs): + return globals()["POLICY_FLOOR"] <= len(obs["audit:R"]) + before, after = pack(change=change_claim(predicate=dynamic)), pack(change=change_claim(predicate=dynamic)) + assert before["intent_revisions"]["c-bob-revoked"] is None + assert not classify(before, after).safe_to_accept + + +@pytest.mark.parametrize("invalid", [None, "true", 1]) +def test_one_unverified_postcondition_prevents_acceptance_and_freezing(invalid): + runs = [pack() for _ in range(3)] + next(o for o in runs[1]["observations"] + if o["kind"] == "realization_check")["data"]["postcondition_met"] = invalid + assert not classify(runs[0], runs[1]).safe_to_accept + assert not assess_stability(runs).stable diff --git a/workplans/TD-WP-0003-generalise-and-settle.md b/workplans/TD-WP-0003-generalise-and-settle.md index e979fc4..436b3b3 100644 --- a/workplans/TD-WP-0003-generalise-and-settle.md +++ b/workplans/TD-WP-0003-generalise-and-settle.md @@ -315,3 +315,25 @@ Validation: the new regression module produced 55 failures before the fix; all intentional partial scenarios. `git diff --check` is clean. Decision: `afb38ade-893b-48a2-b1b3-c55e49cc40ee`. No new task or workplan was created; T01/T06/T07 and the blocked workplan state are unchanged. + + +**2026-09-28 admission follow-up — done.** Crystallization now consumes the +classifier's evidence/acceptance gate for every run and requires matching scope, +unchanged intent, successful realization and distinct receipts before assessing +trajectory stability. Failed or inconclusive baselines cannot authorize automatic +acceptance, including repeated identical failures or repaired candidates. + +Evidence records conservative definition revisions for use cases, claims and +invariants, including text/source/provenance/scheduling and Python predicate +code/defaults/captures/helpers/globals. Changed definitions trigger intent review; +unsupported/dynamic predicate dependencies and legacy revision-free packs remain +ineligible for automatic acceptance. Revisions retain hashes, not captured values; +Python version changes require fresh evidence. Public Claim/Invariant signatures +remain unchanged. Decision: `88e51e10-cd32-44d2-a2d6-055675b5ce04`. + +Validation: the initial regressions produced 20 failures and two passing controls. +The full suite then passed **327 tests**; final dynamic-dependency and strict +postcondition guards passed **32 boundary tests** (including four new cases). +The earlier classification/crystallization subset passed 69 tests. Process/path +stability and captured-value redaction are covered; `git diff --check` is clean. +No new task or workplan; T01/T06/T07 remain waiting and this workplan stays blocked.