Alice/Bob/Carol runs end to end, deterministically, replayable from seed. 16 tests pass, no third-party dependencies. - src/testdriver: intent, provenance, world, actions, drivers, observers, oracles, evidence, energy, scenario, runner - lab/minimal.py: the SUT, exposing the independent observation channel required by D-07 - evidence is stratified S1/S2/S3; Runner refuses to attribute S2/S3 to an actor; claims are frozen and provenance-checked at construction - missing evidence yields INCONCLUSIVE, which outranks PASS in the run verdict - EnergyEvents captured, no scoring (H-005 dormant) The observation channel records both stored state and an out-of-band enforcement probe; their disagreement is an invariant and is what detects an authorization defect that leaves the audit trail intact. A seeded RevokeIsCosmetic lab fails the run via both the claim and that invariant. Also closes TD-WP-0001-T02 (stack and commands now exist). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Assistant: claude-code Assistant-Model: opus Assistant-Process: 1629012@bnt-lap001 Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39
41 lines
1.4 KiB
Markdown
41 lines
1.4 KiB
Markdown
# test-driver
|
||
|
||
Agentic framework for integration, end-to-end, multi-user interaction and
|
||
security testing, driven by use cases.
|
||
|
||
Tests mature alongside the software they protect: fluid and agentic while
|
||
behaviour is changing, deterministic once it settles. See `INTENT.md` for the
|
||
thesis and `SCOPE.md` for boundaries.
|
||
|
||
**Status:** research prototype. The deterministic kernel runs; agentic
|
||
realization, adaptation classification and crystallization are not built yet.
|
||
Current work: `workplans/TD-WP-0002-vertical-spike-crystallization.md`.
|
||
|
||
## Run
|
||
|
||
```bash
|
||
python3 -m pytest -q # the whole suite
|
||
python3 -m pytest -q tests/test_reference_scenario.py
|
||
```
|
||
|
||
No third-party dependencies. Python ≥ 3.11, pytest for the suite.
|
||
|
||
## Layout
|
||
|
||
```
|
||
src/testdriver/ the kernel — intent, world, actions, drivers,
|
||
observers, oracles, evidence, runner
|
||
lab/ the system under test
|
||
scenarios/ reference scenarios
|
||
research/ hypotheses, experiments, findings, fitness map
|
||
docs/ concept model, improvement loop, milestones, design notes
|
||
history/ assessments and completed-work write-ups
|
||
```
|
||
|
||
## Reading order
|
||
|
||
1. `INTENT.md` — the thesis
|
||
2. `docs/TestDriverConceptModel.md` — canonical concept set
|
||
3. `docs/TestDriverClassificationDesign.md` — why adaptation cannot normalize a
|
||
defect, and where model judgment is and is not permitted
|
||
4. `docs/TestDriverInitialMilestones.md` — canonical milestones M0–M10
|