2026-08-22 18:48:27 +00:00
|
|
|
|
# test-driver
|
|
|
|
|
|
|
T04: deterministic semantic kernel
Alice/Bob/Carol runs end to end, deterministically, replayable from seed.
16 tests pass, no third-party dependencies.
- src/testdriver: intent, provenance, world, actions, drivers, observers,
oracles, evidence, energy, scenario, runner
- lab/minimal.py: the SUT, exposing the independent observation channel
required by D-07
- evidence is stratified S1/S2/S3; Runner refuses to attribute S2/S3 to an
actor; claims are frozen and provenance-checked at construction
- missing evidence yields INCONCLUSIVE, which outranks PASS in the run verdict
- EnergyEvents captured, no scoring (H-005 dormant)
The observation channel records both stored state and an out-of-band
enforcement probe; their disagreement is an invariant and is what detects an
authorization defect that leaves the audit trail intact. A seeded
RevokeIsCosmetic lab fails the run via both the claim and that invariant.
Also closes TD-WP-0001-T02 (stack and commands now exist).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 1629012@bnt-lap001
Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39
2026-08-22 23:21:07 +02:00
|
|
|
|
Agentic framework for integration, end-to-end, multi-user interaction and
|
|
|
|
|
|
security testing, driven by use cases.
|
|
|
|
|
|
|
|
|
|
|
|
Tests mature alongside the software they protect: fluid and agentic while
|
|
|
|
|
|
behaviour is changing, deterministic once it settles. See `INTENT.md` for the
|
|
|
|
|
|
thesis and `SCOPE.md` for boundaries.
|
|
|
|
|
|
|
|
|
|
|
|
**Status:** research prototype. The deterministic kernel runs; agentic
|
|
|
|
|
|
realization, adaptation classification and crystallization are not built yet.
|
|
|
|
|
|
Current work: `workplans/TD-WP-0002-vertical-spike-crystallization.md`.
|
|
|
|
|
|
|
|
|
|
|
|
## Run
|
|
|
|
|
|
|
|
|
|
|
|
```bash
|
|
|
|
|
|
python3 -m pytest -q # the whole suite
|
|
|
|
|
|
python3 -m pytest -q tests/test_reference_scenario.py
|
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
No third-party dependencies. Python ≥ 3.11, pytest for the suite.
|
|
|
|
|
|
|
|
|
|
|
|
## Layout
|
|
|
|
|
|
|
|
|
|
|
|
```
|
|
|
|
|
|
src/testdriver/ the kernel — intent, world, actions, drivers,
|
|
|
|
|
|
observers, oracles, evidence, runner
|
|
|
|
|
|
lab/ the system under test
|
|
|
|
|
|
scenarios/ reference scenarios
|
|
|
|
|
|
research/ hypotheses, experiments, findings, fitness map
|
|
|
|
|
|
docs/ concept model, improvement loop, milestones, design notes
|
|
|
|
|
|
history/ assessments and completed-work write-ups
|
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
## Reading order
|
|
|
|
|
|
|
|
|
|
|
|
1. `INTENT.md` — the thesis
|
|
|
|
|
|
2. `docs/TestDriverConceptModel.md` — canonical concept set
|
|
|
|
|
|
3. `docs/TestDriverClassificationDesign.md` — why adaptation cannot normalize a
|
|
|
|
|
|
defect, and where model judgment is and is not permitted
|
|
|
|
|
|
4. `docs/TestDriverInitialMilestones.md` — canonical milestones M0–M10
|