test-driver/tests/test_reference_scenario.py
tegwick 1b9860a8ee T10: gate review and first compression pass
All four gate criteria met. False Adaptation Rate 0/7 with 12 of 13 mechanical
mutations absorbed. 178 tests pass. TD-WP-0002 finished.

Fitness loop closed via F-0003: actor isolation was a property of scenarios
written to expose it, not of runs. Actors now carry an automatic private
marker and the runner examines all of them on every scenario, with two
permanent regressions behind it.

Compression - six abstractions removed, each declared and never used:
Verdict.SUSPICIOUS (a verdict no oracle could emit), Step.expect_refusal,
ActorIsolationError, World.seed, EvidencePack.latest, Trajectory.method.

F-0008: Temperature may be redundant. Crystallization was built without it
ever being consulted; measured stability of realization did the work, and is
observed rather than declared. Gated for removal alongside energy.py.

INTENT_CHANGED and REALIZATION_FAILED had never run. Both now have
purpose-built cases and a test that fails if a seventh outcome is added
without one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 1629012@bnt-lap001
Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39
2026-08-23 00:39:36 +02:00

84 lines
2.8 KiB
Python

"""The reference scenario must run deterministically and be replayable."""
from __future__ import annotations
import json
import pytest
from testdriver import Runner, Stratum, Verdict
from scenarios.alice_bob_carol import build
def run_once(*mutations: str):
world, driver, observer, asset, oracle = build(*mutations)
return Runner(world, driver, observer, oracle).run(asset), world
def test_reference_scenario_passes():
result, _ = run_once()
assert result.verdict is Verdict.PASS, [
j.as_dict() for j in result.judgments if j.verdict is not Verdict.PASS
]
def test_every_claim_is_judged():
result, _ = run_once()
judged = {j.assertion_id for j in result.judgments}
assert {"c-bob-reads", "c-carol-denied", "c-bob-revoked"} <= judged
def test_invariants_are_evaluated_after_every_step():
result, _ = run_once()
per_step = [j for j in result.judgments if j.assertion_id == "i-audit-append-only"]
assert len(per_step) == 3
def test_run_is_replayable_from_known_initial_state():
"""Two runs from the same seed produce identical judgments."""
first, _ = run_once()
second, _ = run_once()
assert [(j.assertion_id, j.verdict) for j in first.judgments] == [
(j.assertion_id, j.verdict) for j in second.judgments
]
assert first.run_id != second.run_id
def test_evidence_is_stratified_and_serializable():
result, _ = run_once()
pack = result.evidence
assert pack.of_stratum(Stratum.SURFACE)
assert pack.of_stratum(Stratum.REALIZATION)
assert pack.of_stratum(Stratum.JUDGMENT)
parsed = json.loads(pack.to_json())
assert parsed["run_id"] == result.run_id
assert parsed["sut_version"] == "lab-0.2.0-baseline"
def test_evidence_records_claim_provenance():
"""A verdict must be auditable for the independence of the claim behind it."""
result, _ = run_once()
assert result.evidence.provenance_index["c-bob-revoked"] == "human"
def test_actors_hold_isolated_credentials_and_memory():
_, world = run_once()
alice, bob = world.cast["alice"], world.cast["bob"]
assert alice.credentials["token"] != bob.credentials["token"]
alice.remember("secret", "only alice knows this")
assert bob.recall("secret") is None
# Every actor carries its own automatic private marker (F-0003) and nothing
# else it was not given.
assert bob.known_keys() == ("__canary__",)
assert alice.canary != bob.canary
def test_every_run_records_a_verdict_on_isolation():
"""F-0003: isolation is examined on every scenario, not only on ones
written to expose it."""
result, _ = run_once()
examined = [
obs for obs in result.evidence.observations if obs.kind == "actor_isolation"
]
assert len(examined) == 1
assert examined[0].data["violations"] == []