All four gate criteria met. False Adaptation Rate 0/7 with 12 of 13 mechanical mutations absorbed. 178 tests pass. TD-WP-0002 finished. Fitness loop closed via F-0003: actor isolation was a property of scenarios written to expose it, not of runs. Actors now carry an automatic private marker and the runner examines all of them on every scenario, with two permanent regressions behind it. Compression - six abstractions removed, each declared and never used: Verdict.SUSPICIOUS (a verdict no oracle could emit), Step.expect_refusal, ActorIsolationError, World.seed, EvidencePack.latest, Trajectory.method. F-0008: Temperature may be redundant. Crystallization was built without it ever being consulted; measured stability of realization did the work, and is observed rather than declared. Gated for removal alongside energy.py. INTENT_CHANGED and REALIZATION_FAILED had never run. Both now have purpose-built cases and a test that fails if a seventh outcome is added without one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Assistant: claude-code Assistant-Model: opus Assistant-Process: 1629012@bnt-lap001 Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39
1.5 KiB
1.5 KiB
test-driver
Agentic framework for integration, end-to-end, multi-user interaction and security testing, driven by use cases.
Tests mature alongside the software they protect: fluid and agentic while
behaviour is changing, deterministic once it settles. See INTENT.md for the
thesis and SCOPE.md for boundaries.
Status: research prototype. The deterministic kernel runs; agentic
realization, adaptation classification and crystallization are not built yet.
Current work: workplans/TD-WP-0002-vertical-spike-crystallization.md.
Run
python3 -m pytest -q # the whole suite
python3 -m pytest -q tests/test_reference_scenario.py
No third-party dependencies. Python ≥ 3.11, pytest for the suite.
Layout
src/testdriver/ the kernel — intent, world, actions, drivers,
observers, oracles, evidence, runner
usecases/ durable test intent, including not-yet-runnable use cases
lab/ the system under test
scenarios/ reference scenarios
research/ hypotheses, experiments, findings, fitness map
docs/ concept model, improvement loop, milestones, design notes
history/ assessments and completed-work write-ups
Reading order
INTENT.md— the thesisdocs/TestDriverConceptModel.md— canonical concept setdocs/TestDriverClassificationDesign.md— why adaptation cannot normalize a defect, and where model judgment is and is not permitteddocs/TestDriverInitialMilestones.md— canonical milestones M0–M10