T04: deterministic semantic kernel

Alice/Bob/Carol runs end to end, deterministically, replayable from seed.
16 tests pass, no third-party dependencies.

- src/testdriver: intent, provenance, world, actions, drivers, observers,
  oracles, evidence, energy, scenario, runner
- lab/minimal.py: the SUT, exposing the independent observation channel
  required by D-07
- evidence is stratified S1/S2/S3; Runner refuses to attribute S2/S3 to an
  actor; claims are frozen and provenance-checked at construction
- missing evidence yields INCONCLUSIVE, which outranks PASS in the run verdict
- EnergyEvents captured, no scoring (H-005 dormant)

The observation channel records both stored state and an out-of-band
enforcement probe; their disagreement is an invariant and is what detects an
authorization defect that leaves the audit trail intact. A seeded
RevokeIsCosmetic lab fails the run via both the claim and that invariant.

Also closes TD-WP-0001-T02 (stack and commands now exist).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 1629012@bnt-lap001
Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39
This commit is contained in:
tegwick 2026-08-22 23:21:07 +02:00
parent 8da4c5bf7a
commit 04e9573b5a
42 changed files with 1533 additions and 20 deletions

View file

@ -133,6 +133,37 @@ curl -s -X PATCH "http://127.0.0.1:8000/tasks/<task_id>" \
{CREDENTIAL_ROUTING}
<!-- REPO-AGENTS-EXTENSIONS -->
## Stack and commands
Python ≥ 3.11, stdlib only; pytest for the suite. No package manager step is
needed — `pyproject.toml` puts `src/` and the repo root on `pythonpath`.
```bash
python3 -m pytest -q # run everything
python3 -m pytest -q -k oracle # narrow
```
Deliberately boring by decision (`docs/TestDriverResearchPrototype.md`): one
process, one database, one browser engine, one application under test. Novelty
belongs in the verification model, never in the infrastructure. Do not add a
dependency without a stated reason in the workplan.
## Non-negotiables
These are architectural, not stylistic. Breaking one silently defeats the
framework's purpose — see `docs/TestDriverClassificationDesign.md`.
- **Claims and invariants are run inputs.** Never add a code path that lets
adaptation, retry, or a learned trajectory modify them (D-02).
- **S2/S3 evidence is never collected by an actor.** `Runner` enforces this;
do not weaken the check (D-01).
- **Model judgment is confined to S1** — locating controls, proposing paths.
Never verdicts, never claim evaluation (Concept Model § 2.3).
- **Missing evidence yields `INCONCLUSIVE`**, never a default pass or fail.
- **Claims require independent provenance.** `agent-from-implementation` output
is an exploratory hypothesis until a human promotes it (D-06).
<!-- Append repo-specific agent instructions below this marker.
The state-hub template sync preserves content after this line. -->