Require qualified runs and intent revisions for automatic acceptance
Assistant: codex Assistant-Model: gpt-6-astra Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
This commit is contained in:
parent
7779768058
commit
a8bb787d12
10 changed files with 486 additions and 14 deletions
|
|
@ -186,3 +186,61 @@ readiness review; the external T01/T06/T07 blockers remain unchanged.
|
|||
Validation after the safety fixes: **299 tests passed** (`python3 -m pytest -q`).
|
||||
The new regression module had 55 failures and five passing controls before the
|
||||
fix; `git diff --check` passes after it.
|
||||
|
||||
## Admission follow-up: crystallization and intent revisions
|
||||
|
||||
The next boundary review reproduced three additional holes: repeated aborted
|
||||
runs qualified as stable; a complete failing run compared with itself was
|
||||
UNCHANGED/safe-to-accept; and changing a claim's text or predicate while retaining
|
||||
its ID/provenance was invisible to the classifier. These fixes remain under T08.
|
||||
|
||||
Crystallization now uses the classifier's acceptance gate for every run against
|
||||
the first run. All packs must have complete evidence, passing judgments, verified
|
||||
postconditions, unchanged intent and successful realizations. Scenario/use-case
|
||||
identity, scheduled steps and expected judgments must match across the window.
|
||||
Repeated copies of one run receipt do not satisfy the independent-run count.
|
||||
Only then does trajectory comparison decide stability. A refused window returns
|
||||
no trajectories. An unchanged product failure is not a stable success.
|
||||
|
||||
The classifier requires a passing baseline. If the baseline contains FAIL or
|
||||
INCONCLUSIVE, even a repaired candidate requires a new passing reference rather
|
||||
than automatic acceptance against the failed baseline. Candidate failures against
|
||||
a passing reference still report BEHAVIOUR_CHANGED. Baseline realization must
|
||||
also be established before either automatic-acceptance outcome is returned;
|
||||
missing or non-boolean postconditions cannot be treated as success.
|
||||
|
||||
Evidence packs now include `intent_revisions`: SHA-256 revisions recorded before
|
||||
execution for the use case and each claim/invariant, including claims outside a
|
||||
partial scenario's schedule. They include narrative or assertion text, source,
|
||||
provenance, claim scheduling, and predicate implementation. The predicate hash
|
||||
covers bytecode/constants, defaults, keyword defaults, closure values, referenced
|
||||
Python helper functions and their referenced global values. It is bound to the
|
||||
Python implementation/version; checkout filenames and line numbers are excluded.
|
||||
Only hashes enter evidence, not captured values or code. Changed definitions
|
||||
produce INTENT_CHANGED, which requests review and does not assert human approval.
|
||||
An internally complete revised claim schedule is an intent change; missing
|
||||
baseline coverage under unchanged intent remains incomplete evidence.
|
||||
|
||||
No Claim/Invariant constructor change or claim serialization language is added.
|
||||
`revisions.py` is input-change detection for pure Python predicates. It does not
|
||||
prove semantic equivalence, sandbox Python, attest a caller-supplied pack or make
|
||||
mutable runtime behavior safe. Predicate dependencies must stay stable during a
|
||||
run. Cyclic/opaque callables, module dependencies and dynamic builtin lookups
|
||||
cannot be fingerprinted reliably and receive a null revision. They may still be
|
||||
evaluated by the oracle, but their evidence cannot be automatically accepted or
|
||||
crystallized. Refactor such predicates to consume plain independent snapshots.
|
||||
Changing Python versions conservatively requires intent review and fresh evidence.
|
||||
|
||||
Older packs without usable intent revisions require a rerun, just as packs
|
||||
without the earlier completeness manifest do. Existing historical experiment
|
||||
receipts are not retroactively upgraded or rewritten. Tests in
|
||||
`tests/test_acceptance_boundaries.py` cover the reproduced cases, helper/global
|
||||
and captured-value changes, invariant/narrative/source/scheduling changes,
|
||||
cross-process identity, source relocation, redaction of captured values, malformed
|
||||
postconditions, and the complete-passing control. No live-model, browser-engine
|
||||
or production readiness blocker is removed by this work.
|
||||
|
||||
Validation for this follow-up: initial reproduction had 20 failures and two
|
||||
passing controls; full suite **327 passed**. The final dynamic-dependency and
|
||||
strict-postcondition additions passed **32 boundary checks**, including four
|
||||
new cases. These counts overlap; they are not separate independent samples.
|
||||
|
|
|
|||
|
|
@ -64,3 +64,10 @@ judgments, realization metrics and lineage remain. See
|
|||
|
||||
Previous compression (T10): `Verdict.SUSPICIOUS`, `Step.expect_refusal`,
|
||||
`ActorIsolationError`, `World.seed`, `EvidencePack.latest()`, `Trajectory.method`.
|
||||
|
||||
|
||||
T08 admission follow-up: `revisions.py` implements conservative intent-definition
|
||||
identity alongside `intent.py`/`provenance.py`. `classification.py` requires a
|
||||
passing reference and complete revisioned evidence; `crystallization.py` consumes
|
||||
that same admission rule. These close local safety holes, not new concepts or an
|
||||
increase in external-validation level. See the generalisation review.
|
||||
|
|
|
|||
|
|
@ -9,3 +9,4 @@ file is a pointer table, not a second source of truth.
|
|||
| D-01 … D-07 | Stratified evidence and claim provenance as the basis for adaptation safety | `fef5213f-ce9b-44c2-b327-a0b0ba4b6270` | `docs/TestDriverClassificationDesign.md` |
|
||||
| TD-WP-0003 T04/T05/T08 | Python claims, concept removal and not-ready assessment | `d80734d1-01b8-4a98-8e72-c86f1806d572` | `docs/TestDriverGeneralisationReview.md` |
|
||||
| TD-WP-0003 T08 safety follow-up | Run-input manifests and incomplete-run handling | `afb38ade-893b-48a2-b1b3-c55e49cc40ee` | `docs/TestDriverGeneralisationReview.md` |
|
||||
| TD-WP-0003 T08 admission follow-up | Qualified crystallization, passing baselines and intent revisions | `88e51e10-cd32-44d2-a2d6-055675b5ce04` | `docs/TestDriverGeneralisationReview.md` |
|
||||
|
|
|
|||
|
|
@ -17,9 +17,9 @@ M19 in the lab prove it), so no amount of evidence separates them. What the
|
|||
classifier can honestly say is *"behaviour changed against intent that did not"*,
|
||||
and hand that to a human.
|
||||
|
||||
Intent change **is** detectable, but only when a human has actually changed the
|
||||
intent: the claim fingerprint moves. That is a fact about the recorded use case,
|
||||
not an inference about behaviour.
|
||||
Intent change **is** detectable when recorded definitions change: the intent
|
||||
revision moves. That calls for review; it does not establish who approved the
|
||||
change and is not an inference about behaviour.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
|
@ -45,7 +45,7 @@ class Classification(str, Enum):
|
|||
"""
|
||||
|
||||
INTENT_CHANGED = "INTENT_CHANGED"
|
||||
"""The recorded claim set itself changed. A human has already acted."""
|
||||
"""Recorded intent changed. Review its definitions and approval provenance."""
|
||||
|
||||
REALIZATION_FAILED = "REALIZATION_FAILED"
|
||||
"""The action could not be performed at all. Not a verdict about the system."""
|
||||
|
|
@ -132,7 +132,8 @@ def _surface_fingerprint(pack: Mapping[str, Any]) -> str:
|
|||
|
||||
|
||||
def _claim_fingerprint(pack: Mapping[str, Any]) -> str:
|
||||
return json.dumps(sorted(pack.get("provenance_index", {}).items()))
|
||||
return json.dumps({"provenance": pack["provenance_index"],
|
||||
"revisions": pack["intent_revisions"]}, sort_keys=True)
|
||||
|
||||
|
||||
def _has_complete_evidence(pack: Mapping[str, Any]) -> bool:
|
||||
|
|
@ -180,7 +181,15 @@ def _has_complete_evidence(pack: Mapping[str, Any]) -> bool:
|
|||
return False
|
||||
if any(o["kind"] == "surface_violation" for o in observations):
|
||||
return False
|
||||
return isinstance(pack["provenance_index"], Mapping)
|
||||
revisions = pack["intent_revisions"]
|
||||
required_ids = {assertion for assertion, _ in expected_keys} | {pack["use_case_id"]}
|
||||
return (
|
||||
isinstance(pack["provenance_index"], Mapping)
|
||||
and isinstance(revisions, Mapping) and required_ids <= revisions.keys()
|
||||
and all(isinstance(revision, str) and len(revision) == 64
|
||||
and all(c in "0123456789abcdef" for c in revision)
|
||||
for revision in revisions.values())
|
||||
)
|
||||
except (KeyError, TypeError, ValueError):
|
||||
# Missing/malformed records are evidence failure, not a default pass.
|
||||
return False
|
||||
|
|
@ -194,6 +203,7 @@ def extract_signals(baseline: Mapping[str, Any], candidate: Mapping[str, Any]) -
|
|||
evidence_incomplete=True,
|
||||
)
|
||||
before, after = _verdict_map(baseline), _verdict_map(candidate)
|
||||
claims_differ = _claim_fingerprint(baseline) != _claim_fingerprint(candidate)
|
||||
|
||||
regressed = any(
|
||||
after.get(key) == "FAIL" and verdict != "FAIL" for key, verdict in before.items()
|
||||
|
|
@ -212,10 +222,10 @@ def extract_signals(baseline: Mapping[str, Any], candidate: Mapping[str, Any]) -
|
|||
met = None
|
||||
elif any(c.get("postcondition_met") is False for c in checks):
|
||||
met = False
|
||||
elif any(c.get("postcondition_met") is None for c in checks):
|
||||
met = None
|
||||
else:
|
||||
elif all(c.get("postcondition_met") is True for c in checks):
|
||||
met = True
|
||||
else:
|
||||
met = None
|
||||
|
||||
return Signals(
|
||||
surface_differs=_surface_fingerprint(baseline) != _surface_fingerprint(candidate),
|
||||
|
|
@ -223,10 +233,12 @@ def extract_signals(baseline: Mapping[str, Any], candidate: Mapping[str, Any]) -
|
|||
postcondition_met=met,
|
||||
verdicts_regressed=regressed,
|
||||
verdicts_inconclusive=any(v == "INCONCLUSIVE" for v in after.values()),
|
||||
claims_differ=_claim_fingerprint(baseline) != _claim_fingerprint(candidate),
|
||||
claims_differ=claims_differ,
|
||||
evidence_incomplete=(
|
||||
not claims_differ and (
|
||||
not before.keys() <= after.keys()
|
||||
or not set(baseline["scheduled_steps"]) <= set(candidate["scheduled_steps"])
|
||||
)
|
||||
),
|
||||
)
|
||||
|
||||
|
|
@ -258,6 +270,10 @@ def classify(baseline: Mapping[str, Any], candidate: Mapping[str, Any]) -> Outco
|
|||
if signals.evidence_incomplete:
|
||||
return Outcome(Classification.AMBIGUOUS,
|
||||
"the run did not retain enough evidence to classify", signals)
|
||||
if any(v["verdict"] != "PASS" for v in baseline["verdicts"]):
|
||||
return Outcome(Classification.AMBIGUOUS,
|
||||
"the baseline is not a passing reference; obtain an accepted baseline",
|
||||
signals)
|
||||
failures = regressions(baseline, candidate)
|
||||
if signals.verdicts_inconclusive:
|
||||
return Outcome(Classification.AMBIGUOUS,
|
||||
|
|
@ -267,8 +283,8 @@ def classify(baseline: Mapping[str, Any], candidate: Mapping[str, Any]) -> Outco
|
|||
# 2. Intent moving is a fact about the recorded use case, not an inference.
|
||||
if signals.claims_differ:
|
||||
return Outcome(Classification.INTENT_CHANGED,
|
||||
"the recorded claim set differs from the baseline; "
|
||||
"a human changed what is being asserted", signals, failures)
|
||||
"the recorded intent differs from the baseline; "
|
||||
"review the changed definitions and provenance", signals, failures)
|
||||
|
||||
# 3. A regression is reported before anything is allowed to explain it away.
|
||||
# This is decision-table row 3: coincidence is not exoneration.
|
||||
|
|
@ -295,6 +311,15 @@ def classify(baseline: Mapping[str, Any], candidate: Mapping[str, Any]) -> Outco
|
|||
return Outcome(Classification.AMBIGUOUS,
|
||||
"the action's postcondition could not be evaluated", signals)
|
||||
|
||||
# A passing verdict alone does not prove the reference actions took effect.
|
||||
baseline_checks = [o["data"] for o in baseline["observations"]
|
||||
if o["kind"] == "realization_check"]
|
||||
if (any(c.get("postcondition_met") is not True for c in baseline_checks)
|
||||
or any(o["kind"] == "realization" and o["data"].get("raised")
|
||||
for o in baseline["observations"])):
|
||||
return Outcome(Classification.AMBIGUOUS,
|
||||
"the baseline does not establish successful realization", signals)
|
||||
|
||||
# 6. Only now, with every claim intact and the action verified, may a
|
||||
# surface difference be called an adaptation.
|
||||
if signals.surface_differs:
|
||||
|
|
|
|||
|
|
@ -24,6 +24,7 @@ from datetime import datetime, timezone
|
|||
from typing import Any, Mapping, Sequence
|
||||
|
||||
from .actions import SemanticAction
|
||||
from .classification import classify
|
||||
from .drivers import Realization
|
||||
from .world import Actor
|
||||
|
||||
|
|
@ -90,6 +91,20 @@ def assess_stability(packs: Sequence[Mapping[str, Any]], minimum: int = 3) -> St
|
|||
False, len(packs), 0,
|
||||
f"need at least {minimum} runs to judge stability, have {len(packs)}",
|
||||
)
|
||||
if not packs:
|
||||
return StabilityReport(False, 0, 0, "no runs to judge stability")
|
||||
for pack in packs:
|
||||
outcome = classify(packs[0], pack)
|
||||
if not outcome.safe_to_accept:
|
||||
return StabilityReport(False, len(packs), 0,
|
||||
f"run is not eligible for crystallization: {outcome.reason}")
|
||||
if (pack["scenario_id"], pack["use_case_id"], pack["scheduled_steps"],
|
||||
pack["expected_judgments"]) != (
|
||||
packs[0]["scenario_id"], packs[0]["use_case_id"], packs[0]["scheduled_steps"],
|
||||
packs[0]["expected_judgments"]):
|
||||
return StabilityReport(False, len(packs), 0, "scenario scope differs across runs")
|
||||
if len({pack["run_id"] for pack in packs}) != len(packs):
|
||||
return StabilityReport(False, len(packs), 0, "need distinct runs, not repeated receipts")
|
||||
captured = [capture(pack) for pack in packs]
|
||||
keys = {tuple(t.key() for t in trajectory) for trajectory in captured}
|
||||
if len(keys) != 1:
|
||||
|
|
|
|||
|
|
@ -67,6 +67,7 @@ class EvidencePack:
|
|||
# Run inputs, captured before execution. A surviving prefix is not a manifest.
|
||||
scheduled_steps: list[str] = field(default_factory=list)
|
||||
expected_judgments: list[dict[str, str]] = field(default_factory=list)
|
||||
intent_revisions: dict[str, str | None] = field(default_factory=dict)
|
||||
|
||||
def record(self, observation: Observation) -> None:
|
||||
self.observations.append(observation)
|
||||
|
|
|
|||
113
src/testdriver/revisions.py
Normal file
113
src/testdriver/revisions.py
Normal file
|
|
@ -0,0 +1,113 @@
|
|||
"""Conservative, process-independent revisions of Python intent inputs.
|
||||
|
||||
Only hashes leave this module: captured values may be sensitive. Unsupported
|
||||
runtime dependencies produce no revision, never a name-only fallback. This is
|
||||
change detection for pure predicates, not a proof of semantic equivalence.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import builtins
|
||||
import dis
|
||||
import hashlib
|
||||
import json
|
||||
import sys
|
||||
from types import CodeType, FunctionType
|
||||
|
||||
from .intent import UseCase
|
||||
|
||||
# These can access dependencies that bytecode name inspection cannot enumerate.
|
||||
_DYNAMIC_BUILTINS = frozenset({
|
||||
"eval", "exec", "globals", "locals", "vars", "getattr", "setattr", "delattr",
|
||||
"__import__", "open", "input",
|
||||
})
|
||||
|
||||
|
||||
def _dump(value) -> str:
|
||||
return json.dumps(value, sort_keys=True, separators=(",", ":"))
|
||||
|
||||
|
||||
def _global_names(code: CodeType) -> set[str]:
|
||||
names = {i.argval for i in dis.get_instructions(code)
|
||||
if i.opname in ("LOAD_GLOBAL", "LOAD_NAME")}
|
||||
for constant in code.co_consts:
|
||||
if isinstance(constant, CodeType):
|
||||
names.update(_global_names(constant))
|
||||
return names
|
||||
|
||||
|
||||
def _describe(value, active: set[int]):
|
||||
"""Do not execute predicates, getters, reprs or arbitrary serialization hooks."""
|
||||
if value is None or type(value) in (bool, int, str):
|
||||
return [type(value).__name__, value]
|
||||
if type(value) is float:
|
||||
return ["float", value.hex()]
|
||||
if type(value) is bytes:
|
||||
return ["bytes", value.hex()]
|
||||
if value is Ellipsis:
|
||||
return ["ellipsis"]
|
||||
if len(active) >= 32 or id(value) in active:
|
||||
raise ValueError("recursive or overly deep predicate dependency")
|
||||
active = active | {id(value)}
|
||||
if type(value) in (tuple, list, set, frozenset):
|
||||
items = [_describe(item, active) for item in value]
|
||||
if type(value) in (set, frozenset):
|
||||
items.sort(key=_dump)
|
||||
return [type(value).__name__, items]
|
||||
if type(value) is dict:
|
||||
items = [[_describe(k, active), _describe(v, active)] for k, v in value.items()]
|
||||
return ["dict", sorted(items, key=_dump)]
|
||||
if isinstance(value, CodeType):
|
||||
return ["code", value.co_code.hex(), _describe(value.co_consts, active),
|
||||
value.co_names, value.co_varnames, value.co_freevars, value.co_cellvars,
|
||||
value.co_argcount, value.co_posonlyargcount, value.co_kwonlyargcount,
|
||||
value.co_flags, value.co_exceptiontable.hex()]
|
||||
if isinstance(value, FunctionType):
|
||||
dependencies = {}
|
||||
for name in sorted(_global_names(value.__code__)):
|
||||
if name in value.__globals__:
|
||||
dependency = value.__globals__[name]
|
||||
else:
|
||||
dependency = value.__builtins__[name]
|
||||
# Builtins are interpreter-version-bound, not opaque user callables.
|
||||
if name in vars(builtins) and dependency is vars(builtins)[name]:
|
||||
if name in _DYNAMIC_BUILTINS:
|
||||
raise ValueError("dynamic predicate dependency")
|
||||
dependencies[name] = ["builtin", name]
|
||||
else:
|
||||
dependencies[name] = _describe(dependency, active)
|
||||
return ["function", _describe(value.__code__, active),
|
||||
_describe(value.__defaults__, active), _describe(value.__kwdefaults__, active),
|
||||
[_describe(cell.cell_contents, active) for cell in value.__closure__ or ()],
|
||||
dependencies]
|
||||
raise ValueError("unsupported predicate dependency")
|
||||
|
||||
|
||||
def _digest(value) -> str:
|
||||
return hashlib.sha256(_dump(value).encode()).hexdigest()
|
||||
|
||||
|
||||
def intent_revisions(case: UseCase) -> dict[str, str | None]:
|
||||
"""Capture intent before execution, including unscheduled use-case claims.
|
||||
|
||||
IDs remain stable; revisions change with their protected meaning. Predicate
|
||||
bytecode is bound to the Python implementation/version. Code location and
|
||||
line numbers are deliberately excluded so checkout paths do not cause drift.
|
||||
"""
|
||||
revisions: dict[str, str | None] = {
|
||||
case.id: _digest(["use-case", case.title, case.narrative,
|
||||
case.provenance.value, case.source_ref]),
|
||||
}
|
||||
for assertion in (*case.claims, *case.invariants):
|
||||
if assertion.id in revisions:
|
||||
return {} # Ambiguous identities cannot produce an admissible revision.
|
||||
try:
|
||||
predicate = _describe(assertion.predicate, set())
|
||||
except (KeyError, TypeError, ValueError, RecursionError):
|
||||
revisions[assertion.id] = None
|
||||
continue
|
||||
revisions[assertion.id] = _digest([
|
||||
"python-intent-v1", sys.implementation.name, list(sys.version_info[:3]),
|
||||
type(assertion).__name__, assertion.text, assertion.provenance.value,
|
||||
assertion.source_ref, getattr(assertion, "after_step", None), predicate,
|
||||
])
|
||||
return revisions
|
||||
|
|
@ -18,6 +18,7 @@ from .drivers import Driver
|
|||
from .evidence import EvidencePack, Observation, Stratum
|
||||
from .observers import StateObserver
|
||||
from .oracles import Judgment, Oracle, Verdict, overall
|
||||
from .revisions import intent_revisions
|
||||
from .scenario import Scenario, VerificationAsset
|
||||
from .world import World
|
||||
|
||||
|
|
@ -118,6 +119,7 @@ class Runner:
|
|||
scenario_id=scenario.id,
|
||||
use_case_id=scenario.use_case.id,
|
||||
sut_version=self._world.sut_version,
|
||||
intent_revisions=intent_revisions(scenario.use_case),
|
||||
scheduled_steps=[step.id for step in scenario.steps],
|
||||
expected_judgments=[
|
||||
{"assertion_id": assertion.id, "step_id": step.id}
|
||||
|
|
|
|||
228
tests/test_acceptance_boundaries.py
Normal file
228
tests/test_acceptance_boundaries.py
Normal file
|
|
@ -0,0 +1,228 @@
|
|||
"""Admission, crystallization and intent revisions must agree on safe evidence."""
|
||||
from copy import deepcopy
|
||||
from dataclasses import replace
|
||||
import json
|
||||
|
||||
import pytest
|
||||
|
||||
from scenarios.alice_bob_carol import build
|
||||
from testdriver import Runner
|
||||
from testdriver.classification import Classification, classify
|
||||
from testdriver.crystallization import assess_stability
|
||||
|
||||
|
||||
def pack(*mutations, change=None, abort=False):
|
||||
world, driver, observer, asset, oracle = build(*mutations)
|
||||
if change:
|
||||
asset.scenario = replace(asset.scenario, use_case=change(asset.scenario.use_case))
|
||||
if abort:
|
||||
steps = list(asset.scenario.steps)
|
||||
steps[1] = replace(steps[1], action=replace(
|
||||
steps[1].action, permitted_surfaces=frozenset({"browser"})))
|
||||
asset.scenario = replace(asset.scenario, steps=tuple(steps))
|
||||
return json.loads(Runner(world, driver, observer, oracle).run(asset).evidence.to_json())
|
||||
|
||||
|
||||
def change_claim(**changes):
|
||||
return lambda case: replace(case, claims=tuple(
|
||||
replace(c, **changes) if c.id == "c-bob-revoked" else c for c in case.claims
|
||||
))
|
||||
|
||||
|
||||
@pytest.mark.parametrize("mutation", ["M17", "M15", "M16", "M18"])
|
||||
def test_failing_baseline_cannot_authorize_same_failure_or_fixed_run(mutation):
|
||||
failing = pack(mutation)
|
||||
assert any(v["verdict"] == "FAIL" for v in failing["verdicts"])
|
||||
for candidate in (deepcopy(failing), pack()):
|
||||
outcome = classify(failing, candidate)
|
||||
assert not outcome.safe_to_accept
|
||||
assert outcome.classification is Classification.AMBIGUOUS
|
||||
assert "baseline" in outcome.reason
|
||||
|
||||
|
||||
@pytest.mark.parametrize("damage", ["abort", "snapshots", "verdict", "failure", "inconclusive"])
|
||||
def test_repeated_invalid_runs_cannot_crystallize(damage):
|
||||
packs = [pack("M17") if damage == "failure" else pack(abort=damage == "abort")
|
||||
for _ in range(3)]
|
||||
for candidate in packs:
|
||||
if damage == "snapshots":
|
||||
candidate["observations"] = [o for o in candidate["observations"]
|
||||
if o["kind"] != "state_snapshot"]
|
||||
elif damage == "verdict":
|
||||
candidate["verdicts"].pop()
|
||||
elif damage == "inconclusive":
|
||||
candidate["verdicts"][0]["verdict"] = "INCONCLUSIVE"
|
||||
report = assess_stability(packs)
|
||||
assert not report.stable
|
||||
assert not report.trajectories
|
||||
assert report.reason
|
||||
|
||||
|
||||
def test_one_bad_run_in_an_otherwise_stable_window_prevents_freezing():
|
||||
packs = [pack(), pack(abort=True), pack()]
|
||||
assert not assess_stability(packs).stable
|
||||
|
||||
|
||||
def test_replaying_one_receipt_does_not_count_as_three_runs():
|
||||
baseline = pack()
|
||||
assert not assess_stability([deepcopy(baseline) for _ in range(3)]).stable
|
||||
|
||||
|
||||
def test_complete_passing_runs_still_classify_and_crystallize():
|
||||
packs = [pack() for _ in range(3)]
|
||||
assert classify(packs[0], packs[1]).classification is Classification.UNCHANGED
|
||||
assert assess_stability(packs).stable
|
||||
|
||||
|
||||
@pytest.mark.parametrize("changes", [
|
||||
{"text": "Revocation is optional"},
|
||||
{"source_ref": "new-independent-spec/revocation-v2"},
|
||||
{"predicate": lambda obs: True},
|
||||
])
|
||||
def test_same_id_does_not_hide_a_changed_claim(changes):
|
||||
before, after = pack(), pack(change=change_claim(**changes))
|
||||
assert classify(before, after).classification is Classification.INTENT_CHANGED
|
||||
assert not assess_stability([before, after, pack()]).stable
|
||||
|
||||
|
||||
def test_invariant_changes_are_intent_changes_too():
|
||||
def change(case):
|
||||
return replace(case, invariants=(replace(case.invariants[0], predicate=lambda obs: True),
|
||||
*case.invariants[1:]))
|
||||
assert classify(pack(), pack(change=change)).classification is Classification.INTENT_CHANGED
|
||||
|
||||
|
||||
def test_narrative_changes_require_intent_review():
|
||||
assert classify(pack(), pack(change=lambda c: replace(c, narrative="New policy"))) \
|
||||
.classification is Classification.INTENT_CHANGED
|
||||
|
||||
|
||||
def threshold(value):
|
||||
return lambda obs: len(obs["audit:R"]) >= value
|
||||
|
||||
|
||||
def default_threshold(value):
|
||||
return lambda obs, floor=value: len(obs["audit:R"]) >= floor
|
||||
|
||||
|
||||
@pytest.mark.parametrize("factory", [threshold, default_threshold])
|
||||
def test_captured_predicate_values_are_part_of_the_revision(factory):
|
||||
first = pack(change=change_claim(predicate=factory(0)))
|
||||
second = pack(change=change_claim(predicate=factory(1)))
|
||||
assert classify(first, second).classification is Classification.INTENT_CHANGED
|
||||
assert classify(first, pack(change=change_claim(predicate=factory(0)))).safe_to_accept
|
||||
|
||||
|
||||
POLICY_FLOOR = 0
|
||||
|
||||
|
||||
def helper(obs):
|
||||
return len(obs["audit:R"]) >= POLICY_FLOOR
|
||||
|
||||
|
||||
def predicate_using_helper(obs):
|
||||
return helper(obs)
|
||||
|
||||
|
||||
def test_global_values_and_helpers_are_part_of_the_revision(monkeypatch):
|
||||
first = pack(change=change_claim(predicate=predicate_using_helper))
|
||||
monkeypatch.setitem(globals(), "POLICY_FLOOR", 1)
|
||||
second = pack(change=change_claim(predicate=predicate_using_helper))
|
||||
assert classify(first, second).classification is Classification.INTENT_CHANGED
|
||||
monkeypatch.setitem(globals(), "helper", lambda obs: True)
|
||||
third = pack(change=change_claim(predicate=predicate_using_helper))
|
||||
assert classify(second, third).classification is Classification.INTENT_CHANGED
|
||||
|
||||
|
||||
def test_unidentifiable_callable_cannot_be_automatically_accepted():
|
||||
class Opaque:
|
||||
def __call__(self, obs):
|
||||
return True
|
||||
packs = [pack(change=change_claim(predicate=Opaque())) for _ in range(3)]
|
||||
assert not classify(packs[0], packs[1]).safe_to_accept
|
||||
assert not assess_stability(packs).stable
|
||||
|
||||
|
||||
def test_legacy_pack_without_intent_revisions_requires_rerun():
|
||||
before, after = pack(), pack()
|
||||
after.pop("intent_revisions", None)
|
||||
assert not classify(before, after).safe_to_accept
|
||||
assert not assess_stability([after, pack(), pack()]).stable
|
||||
|
||||
|
||||
def test_predicate_location_does_not_change_revision():
|
||||
from types import FunctionType
|
||||
from scenarios.alice_bob_carol import USE_CASE
|
||||
from testdriver.revisions import intent_revisions
|
||||
|
||||
original = USE_CASE.claims[0].predicate
|
||||
moved = FunctionType(original.__code__.replace(
|
||||
co_filename="/a/different/checkout/scenario.py", co_firstlineno=999,
|
||||
), original.__globals__)
|
||||
changed = replace(USE_CASE, claims=(replace(USE_CASE.claims[0], predicate=moved),
|
||||
*USE_CASE.claims[1:]))
|
||||
assert intent_revisions(USE_CASE) == intent_revisions(changed)
|
||||
|
||||
|
||||
def test_revisions_are_stable_across_processes():
|
||||
import os
|
||||
import subprocess
|
||||
import sys
|
||||
from pathlib import Path
|
||||
from scenarios.alice_bob_carol import USE_CASE
|
||||
from testdriver.revisions import intent_revisions
|
||||
|
||||
root = Path(__file__).resolve().parents[1]
|
||||
output = subprocess.check_output([
|
||||
sys.executable, "-c",
|
||||
"import json; from scenarios.alice_bob_carol import USE_CASE; "
|
||||
"from testdriver.revisions import intent_revisions; "
|
||||
"print(json.dumps(intent_revisions(USE_CASE)))",
|
||||
], cwd=root, env={**os.environ, "PYTHONPATH": os.pathsep.join((str(root / "src"), str(root)))},
|
||||
text=True)
|
||||
assert json.loads(output) == intent_revisions(USE_CASE)
|
||||
|
||||
|
||||
def test_captured_values_are_hashed_not_retained():
|
||||
secret = "synthetic-private-dependency-do-not-retain"
|
||||
predicate = lambda obs: bool(secret) and obs["probe_read:bob:R"] is False
|
||||
result = pack(change=change_claim(predicate=predicate))
|
||||
assert result["intent_revisions"]["c-bob-revoked"] is not None
|
||||
assert secret not in json.dumps(result)
|
||||
|
||||
|
||||
def test_schedule_change_is_an_intent_change():
|
||||
before = pack()
|
||||
after = pack(change=change_claim(after_step="s1-create", predicate=lambda obs: True))
|
||||
assert classify(before, after).classification is Classification.INTENT_CHANGED
|
||||
|
||||
|
||||
def test_failed_baseline_realization_does_not_authorize_automatic_acceptance():
|
||||
before, after = pack(), pack()
|
||||
next(o for o in before["observations"]
|
||||
if o["kind"] == "realization")["data"]["raised"] = "Unrealized"
|
||||
assert not classify(before, after).safe_to_accept
|
||||
|
||||
|
||||
def test_unknown_baseline_postcondition_does_not_authorize_automatic_acceptance():
|
||||
before, after = pack(), pack()
|
||||
next(o for o in before["observations"]
|
||||
if o["kind"] == "realization_check")["data"]["postcondition_met"] = None
|
||||
assert not classify(before, after).safe_to_accept
|
||||
|
||||
|
||||
def test_dynamic_global_lookup_is_not_given_a_misleading_revision():
|
||||
def dynamic(obs):
|
||||
return globals()["POLICY_FLOOR"] <= len(obs["audit:R"])
|
||||
before, after = pack(change=change_claim(predicate=dynamic)), pack(change=change_claim(predicate=dynamic))
|
||||
assert before["intent_revisions"]["c-bob-revoked"] is None
|
||||
assert not classify(before, after).safe_to_accept
|
||||
|
||||
|
||||
@pytest.mark.parametrize("invalid", [None, "true", 1])
|
||||
def test_one_unverified_postcondition_prevents_acceptance_and_freezing(invalid):
|
||||
runs = [pack() for _ in range(3)]
|
||||
next(o for o in runs[1]["observations"]
|
||||
if o["kind"] == "realization_check")["data"]["postcondition_met"] = invalid
|
||||
assert not classify(runs[0], runs[1]).safe_to_accept
|
||||
assert not assess_stability(runs).stable
|
||||
|
|
@ -315,3 +315,25 @@ Validation: the new regression module produced 55 failures before the fix; all
|
|||
intentional partial scenarios. `git diff --check` is clean. Decision:
|
||||
`afb38ade-893b-48a2-b1b3-c55e49cc40ee`. No new task or workplan was created;
|
||||
T01/T06/T07 and the blocked workplan state are unchanged.
|
||||
|
||||
|
||||
**2026-09-28 admission follow-up — done.** Crystallization now consumes the
|
||||
classifier's evidence/acceptance gate for every run and requires matching scope,
|
||||
unchanged intent, successful realization and distinct receipts before assessing
|
||||
trajectory stability. Failed or inconclusive baselines cannot authorize automatic
|
||||
acceptance, including repeated identical failures or repaired candidates.
|
||||
|
||||
Evidence records conservative definition revisions for use cases, claims and
|
||||
invariants, including text/source/provenance/scheduling and Python predicate
|
||||
code/defaults/captures/helpers/globals. Changed definitions trigger intent review;
|
||||
unsupported/dynamic predicate dependencies and legacy revision-free packs remain
|
||||
ineligible for automatic acceptance. Revisions retain hashes, not captured values;
|
||||
Python version changes require fresh evidence. Public Claim/Invariant signatures
|
||||
remain unchanged. Decision: `88e51e10-cd32-44d2-a2d6-055675b5ce04`.
|
||||
|
||||
Validation: the initial regressions produced 20 failures and two passing controls.
|
||||
The full suite then passed **327 tests**; final dynamic-dependency and strict
|
||||
postcondition guards passed **32 boundary tests** (including four new cases).
|
||||
The earlier classification/crystallization subset passed 69 tests. Process/path
|
||||
stability and captured-value redaction are covered; `git diff --check` is clean.
|
||||
No new task or workplan; T01/T06/T07 remain waiting and this workplan stays blocked.
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue