T04: deterministic semantic kernel

Alice/Bob/Carol runs end to end, deterministically, replayable from seed.
16 tests pass, no third-party dependencies.

- src/testdriver: intent, provenance, world, actions, drivers, observers,
  oracles, evidence, energy, scenario, runner
- lab/minimal.py: the SUT, exposing the independent observation channel
  required by D-07
- evidence is stratified S1/S2/S3; Runner refuses to attribute S2/S3 to an
  actor; claims are frozen and provenance-checked at construction
- missing evidence yields INCONCLUSIVE, which outranks PASS in the run verdict
- EnergyEvents captured, no scoring (H-005 dormant)

The observation channel records both stored state and an out-of-band
enforcement probe; their disagreement is an invariant and is what detects an
authorization defect that leaves the audit trail intact. A seeded
RevokeIsCosmetic lab fails the run via both the claim and that invariant.

Also closes TD-WP-0001-T02 (stack and commands now exist).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 1629012@bnt-lap001
Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39
This commit is contained in:
tegwick 2026-08-22 23:21:07 +02:00
parent 8da4c5bf7a
commit 04e9573b5a
42 changed files with 1533 additions and 20 deletions

View file

@ -32,7 +32,7 @@ Replace generated placeholders with repo-specific facts where needed.
```task
id: TD-WP-0001-T02
status: wait
status: done
priority: high
state_hub_task_id: "39da4237-aaa3-5b4f-93c4-9cbf794f0850"
```
@ -58,8 +58,9 @@ checkout:
statehub fix-consistency
```
Blocked until the stack exists: no code, no dependency manifest and no test
runner are present yet. Unblocks with TD-WP-0002-T04 (deterministic semantic
kernel), which introduces the first Python package and pytest configuration.
**Done 2026-08-22.** Unblocked by TD-WP-0002-T04. `pyproject.toml` added;
`python3 -m pytest -q` is the whole workflow — stdlib only, no install step.
Commands and the architectural non-negotiables are recorded in `AGENTS.md`
under the repo-extensions marker, and in `README.md`.
Seeded workplan: `workplans/TD-WP-0002-vertical-spike-crystallization.md`.

View file

@ -176,7 +176,7 @@ Three things worth carrying forward:
```task
id: TD-WP-0002-T04
status: todo
status: done
priority: high
state_hub_task_id: "ffbcc9e8-c1bd-5c50-b62d-495ad9e135a4"
```
@ -198,6 +198,26 @@ structured Evidence Pack; the scenario replays from known initial state.
Emit raw `EnergyEvent` records from this point onward. Implement no scoring.
**Done 2026-08-22.** `src/testdriver/` (11 modules), `lab/minimal.py`,
`scenarios/alice_bob_carol.py`, 16 passing tests. The reference scenario runs
end to end and replays identically from the same seed; evidence comes out
stratified 3/3/3 across S1/S2/S3.
Three things that came out of building it rather than designing it:
- **The observation channel needs two probes, not one.** Reading stored state
alone verifies test-driver's reimplementation of the rules rather than the
system's enforcement of them; probing enforcement alone cannot notice that
record and enforcement disagree. The lab exposes both, and their disagreement
is now an invariant (`i-enforcement-matches-record`). That invariant is what
catches an authorization defect which leaves the audit trail looking correct.
- **A seeded `RevokeIsCosmetic` lab already fails the run** — both the claim and
the independent invariant fire, and the claim set is provably untouched. Early
evidence for H-004, though not yet the experiment.
- **Scenarios are Python, not YAML.** Claims are predicates over observations; a
YAML dialect able to express them would be a programming language with worse
tooling. Revisit once we know which predicates actually recur.
## Test-driver lab with labelled ground truth
```task