T04: deterministic semantic kernel
Alice/Bob/Carol runs end to end, deterministically, replayable from seed. 16 tests pass, no third-party dependencies. - src/testdriver: intent, provenance, world, actions, drivers, observers, oracles, evidence, energy, scenario, runner - lab/minimal.py: the SUT, exposing the independent observation channel required by D-07 - evidence is stratified S1/S2/S3; Runner refuses to attribute S2/S3 to an actor; claims are frozen and provenance-checked at construction - missing evidence yields INCONCLUSIVE, which outranks PASS in the run verdict - EnergyEvents captured, no scoring (H-005 dormant) The observation channel records both stored state and an out-of-band enforcement probe; their disagreement is an invariant and is what detects an authorization defect that leaves the audit trail intact. A seeded RevokeIsCosmetic lab fails the run via both the claim and that invariant. Also closes TD-WP-0001-T02 (stack and commands now exist). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Assistant: claude-code Assistant-Model: opus Assistant-Process: 1629012@bnt-lap001 Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39
This commit is contained in:
parent
8da4c5bf7a
commit
04e9573b5a
42 changed files with 1533 additions and 20 deletions
40
README.md
40
README.md
|
|
@ -1,3 +1,41 @@
|
|||
# test-driver
|
||||
|
||||
Agentic framework for integration, end2end, multiuserinteraction, security testing based on usecases.
|
||||
Agentic framework for integration, end-to-end, multi-user interaction and
|
||||
security testing, driven by use cases.
|
||||
|
||||
Tests mature alongside the software they protect: fluid and agentic while
|
||||
behaviour is changing, deterministic once it settles. See `INTENT.md` for the
|
||||
thesis and `SCOPE.md` for boundaries.
|
||||
|
||||
**Status:** research prototype. The deterministic kernel runs; agentic
|
||||
realization, adaptation classification and crystallization are not built yet.
|
||||
Current work: `workplans/TD-WP-0002-vertical-spike-crystallization.md`.
|
||||
|
||||
## Run
|
||||
|
||||
```bash
|
||||
python3 -m pytest -q # the whole suite
|
||||
python3 -m pytest -q tests/test_reference_scenario.py
|
||||
```
|
||||
|
||||
No third-party dependencies. Python ≥ 3.11, pytest for the suite.
|
||||
|
||||
## Layout
|
||||
|
||||
```
|
||||
src/testdriver/ the kernel — intent, world, actions, drivers,
|
||||
observers, oracles, evidence, runner
|
||||
lab/ the system under test
|
||||
scenarios/ reference scenarios
|
||||
research/ hypotheses, experiments, findings, fitness map
|
||||
docs/ concept model, improvement loop, milestones, design notes
|
||||
history/ assessments and completed-work write-ups
|
||||
```
|
||||
|
||||
## Reading order
|
||||
|
||||
1. `INTENT.md` — the thesis
|
||||
2. `docs/TestDriverConceptModel.md` — canonical concept set
|
||||
3. `docs/TestDriverClassificationDesign.md` — why adaptation cannot normalize a
|
||||
defect, and where model judgment is and is not permitted
|
||||
4. `docs/TestDriverInitialMilestones.md` — canonical milestones M0–M10
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue