test-driver/README.md
tegwick 04e9573b5a T04: deterministic semantic kernel
Alice/Bob/Carol runs end to end, deterministically, replayable from seed.
16 tests pass, no third-party dependencies.

- src/testdriver: intent, provenance, world, actions, drivers, observers,
  oracles, evidence, energy, scenario, runner
- lab/minimal.py: the SUT, exposing the independent observation channel
  required by D-07
- evidence is stratified S1/S2/S3; Runner refuses to attribute S2/S3 to an
  actor; claims are frozen and provenance-checked at construction
- missing evidence yields INCONCLUSIVE, which outranks PASS in the run verdict
- EnergyEvents captured, no scoring (H-005 dormant)

The observation channel records both stored state and an out-of-band
enforcement probe; their disagreement is an invariant and is what detects an
authorization defect that leaves the audit trail intact. A seeded
RevokeIsCosmetic lab fails the run via both the claim and that invariant.

Also closes TD-WP-0001-T02 (stack and commands now exist).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 1629012@bnt-lap001
Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39
2026-08-22 23:21:07 +02:00

1.4 KiB
Raw Blame History

test-driver

Agentic framework for integration, end-to-end, multi-user interaction and security testing, driven by use cases.

Tests mature alongside the software they protect: fluid and agentic while behaviour is changing, deterministic once it settles. See INTENT.md for the thesis and SCOPE.md for boundaries.

Status: research prototype. The deterministic kernel runs; agentic realization, adaptation classification and crystallization are not built yet. Current work: workplans/TD-WP-0002-vertical-spike-crystallization.md.

Run

python3 -m pytest -q          # the whole suite
python3 -m pytest -q tests/test_reference_scenario.py

No third-party dependencies. Python ≥ 3.11, pytest for the suite.

Layout

src/testdriver/   the kernel — intent, world, actions, drivers,
                  observers, oracles, evidence, runner
lab/              the system under test
scenarios/        reference scenarios
research/         hypotheses, experiments, findings, fitness map
docs/             concept model, improvement loop, milestones, design notes
history/          assessments and completed-work write-ups

Reading order

  1. INTENT.md — the thesis
  2. docs/TestDriverConceptModel.md — canonical concept set
  3. docs/TestDriverClassificationDesign.md — why adaptation cannot normalize a defect, and where model judgment is and is not permitted
  4. docs/TestDriverInitialMilestones.md — canonical milestones M0M10