Align scope with intent and add durable evidence and bounded variants

Assistant: codex
Assistant-Model: gpt-6-astra
Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
This commit is contained in:
tegwick 2026-09-28 14:31:22 +02:00
parent 8822480b6e
commit eaf5d348e4
14 changed files with 767 additions and 92 deletions

149
SCOPE.md
View file

@ -1,103 +1,80 @@
# SCOPE
> This file helps you quickly understand what this repository is about,
> when it is relevant, and when it is not.
Assessed 2026-09-28. `test-driver` is an executable **local research prototype**
for use-case-driven verification, adaptation classification and deterministic
replay. It is not yet a validated autonomous test platform for real systems.
---
## What it can do
## One-liner
| Capability | Implemented boundary and evidence |
|---|---|
| Deterministic verification | Python UseCase/Claim/Invariant inputs, ordered semantic steps, separate drivers and observers, explicit PASS/FAIL/INCONCLUSIVE. Non-boolean or missing predicate evidence is inconclusive. `src/testdriver/runner.py`, `oracles.py`; `tests/test_kernel_guarantees.py`. |
| Multi-user/domain examples | Sharing and revocation, delegated approval, and tenant deletion/recreation run against three synthetic lab domains. Actor stores/canaries and S2/S3 collector attribution are checked. These are in-process diagnostics, not process isolation. `scenarios/`, `tests/test_generalisation.py`. |
| Mechanical discovery | A deterministic heuristic finds forms in server-rendered HTML; a recorded-selector control provides comparison. This is not live-model inference or a JavaScript browser engine. `agentic.py`, `browser.py`. |
| Evidence-based classification | Compare complete passing reference runs with candidate evidence, explicit intent revisions and postconditions. Distinguish mechanical changes, observed behavior changes, changed intent, realization failures and ambiguity. Cannot infer whether a product behavior change was authorized. `classification.py`, `revisions.py`. |
| Crystallization | Check repeated eligible, distinct runs for stable paths; replay frozen paths by step; generate a single-action pytest module importing original predicates. One checked-in browser descendant runs without model involvement. No automatic full T0–T5 lifecycle or general multi-step code generator. `crystallization.py`, `crystallized/`. |
| Retained evidence and lineage | Optional local `EvidenceStore` writes versioned, checksummed JSON receipts atomically without replacing an existing run. Receipts contain asset/parent/maturity/variant, intent revisions, scheduled coverage, observations and final verdict. Reloaded packs can be classified. See [usage](docs/TestDriverEvidenceAndVariants.md). |
| Bounded adversarial variants | `variants.substitute` derives one actor or argument substitution, retaining the exact independent use case, assertions, postconditions and surface permissions. Resource, tenant and privilege values use the same operation. Scenario revisions prevent changed action intent from qualifying as mere mechanical adaptation. No automatic claim authoring. |
| Local research calibration | A 38-mutation catalogue, independent fixture judgments, mechanical/defect experiments and tests that deliberately break framework guarantees. Lab false-adaptation results are bounded experiment results, not a universal safety proof. `lab/GROUND-TRUTH.md`, `research/`, `tests/selfverification/`. |
| HTTP boundary checks | Browser and frozen/standalone HTTP bind bearer credentials to a configured origin, including redirects. Runner also validates reported surfaces. Drivers still need pre-action checks; a post-call check cannot undo effects or attest an untruthful report. |
`test-driver` is a use-case-driven verification framework whose tests mature
alongside the software they protect — fluid and agentic while behaviour is hot,
deterministic once it cools.
## Where it is useful
---
Use it to investigate whether independent Python claims survive changed
mechanics, to build small sequential multi-user domain fixtures, to calibrate
mutation detection, or to replay a demonstrated stable HTML action. A new domain
needs an explicitly authored scenario, driver and independent observation adapter.
## Core Idea
The audit-core E2 module is an approved-intent contract and fixture calibration,
**not** a runnable production engagement. It has no deployed custody, cluster,
time-window or independent cleanup adapters in this repository.
A test is treated not as code but as a **verification asset** with identity,
intent, evidence, lineage, maturity, temperature and energy. Use cases are the
primary behavioural source; integration, journey, multi-user, security and
resilience tests are projections of the same use case rather than separate
suites.
## Limits and readiness
Verification assets progress along a maturity continuum
(`T0 Exploratory → T5 Deterministic`) called **Crystallization**. Agents may
explore and realise semantic actions against unstable interfaces; oracles remain
deterministic and independent from actors, so the framework can distinguish a
legitimate mechanical change from a product defect rather than adapting to
whatever the implementation happens to do.
- No natural-language use-case parser or autonomous scenario planner. Intent is
explicitly authored Python; source/provenance labels are not external attestation.
- No live-model experiment or validated agentic cost benefit. The runtime protocol
exists; the demonstrated runtime is heuristic and token-free.
- No Playwright/browser engine, JavaScript execution, screenshots or DOM timing.
- Sequential execution only: no actual grant/revoke races, causal deadlines,
concurrency scheduler or general dependency-fault/resilience engine.
- No generic CLI, Kubernetes, message-bus or multi-service production adapters.
- Evidence storage is opt-in and local. It is not encryption, signing, automatic
redaction, retention management or safe storage of arbitrary secrets. Unexpected
driver/observer exceptions still propagate; no finalized receipt is promised for
an interrupted/crashed process. Strict JSON rejects unsupported evidence values.
- No automatic finding-to-regression pipeline, asset registry, maturity promotion,
thawing or retirement. Parent ids and mutation history provide limited lineage.
**Current state:** research prototype. Concept corpus is complete
(`INTENT.md`, `docs/`); implementation has not started. See
`history/2026-08-22-concept-assessment-swot.md` for the standing assessment and
the reasoning behind the current workplan sequence.
Temperature, Energy, Confidence, Campaign, Metabolism and automated Retirement
were **deliberately removed**, not postponed implementation requirements. The
superseding note in INTENT.md and [settlement](docs/TestDriverGeneralisationReview.md)
take precedence over its historical lifecycle sections.
---
Outstanding work is tracked in [TD-WP-0003](workplans/TD-WP-0003-generalise-and-settle.md)
and [TD-WP-0004](workplans/TD-WP-0004-scope-evidence-and-variants.md), with priorities
and evidence in the [scope/intent assessment](history/2026-09-28-121933-scope-intent-assessment.md).
## In Scope
## Non-goals
- The conceptual model: UseCase, Actor, Scenario, SemanticAction, Observation,
Oracle, Verdict, VerificationAsset, Finding, Adaptation, Crystallization.
- A deterministic semantic scenario kernel and its evidence format.
- The **test-driver lab** — a small mutable application under test carrying
labelled mechanical, semantic and defect mutations as ground truth.
- Adaptation detection and the defect-vs-adaptation classifier.
- Crystallization of agentic realisations into deterministic regression tests.
- Security testing expressed as mutation of ordinary use cases.
- Self-verification of the framework's own foundational guarantees.
- The research control plane: hypotheses, experiments, findings, fitness map.
Replacing pytest, browser engines or CI; a distributed test cloud; a vulnerability
scanner, load-testing or observability platform; automatic rewriting of semantic
requirements; model-based verdicts; exhaustive scenario permutations. These remain
outside the intended initial project.
---
## Actual stack and entry points
## Out of Scope
Python ≥3.11, dataclasses and the standard library; pytest for testing. Labs use
in-memory domain state and a local stdlib HTTP server. No runtime third-party
dependencies, database, SQLite, Pydantic, YAML parser or browser engine is used.
- Replacing unit-test frameworks, browser automation engines, or CI systems.
- Building a load-testing, fuzzing, vulnerability-scanning or observability
platform.
- Test-management SaaS, distributed test clouds, or multi-tenant hosting.
- Making all tests agentic, or using model judgment where a deterministic oracle
is available.
- Rewriting semantic requirements to match implementation behaviour.
- Scale and performance verification before the conceptual model is proven.
```bash
python3 -m pytest -q
python3 -m pytest -q tests/test_reference_scenario.py tests/test_scope_delivery.py
```
---
## Relevant When
- You need the test-driver conceptual vocabulary or its canonical concept set.
- You are working on the crystallization, adaptation-classification, or
semantic-action binding mechanisms.
- You are extending the lab or its mutation catalogue.
- You need the framework's hypotheses, fitness scorecard, or evidence format.
---
## Not Relevant When
- You need ordinary unit or component tests for another repo — use that repo's
own test stack.
- You are looking for fleet coordination or cross-repo memory — that is State
Hub.
---
## Getting Oriented
1. `INTENT.md` — purpose, thesis, design heuristics, non-goals.
2. `docs/TestDriverConceptModel.md` — the canonical concept set v0.1.
3. `docs/TestDriverImprovementLoop.md` — hypotheses, findings taxonomy,
self-improvement cycle.
4. `docs/TestDriverInitialMilestones.md` — M0–M10 and the prototype success gate.
5. `history/2026-08-22-concept-assessment-swot.md` — assessment and the reasons
the first workplan reorders those milestones into a vertical spike.
6. `workplans/` — current work. Agent instructions: `AGENTS.md`.
---
## Stack
Deliberately boring, per `docs/TestDriverResearchPrototype.md`: Python, pytest,
Playwright, Pydantic/dataclasses, YAML, SQLite. One process, one database, one
browser engine, one application under test. Novelty belongs in the verification
model, not the infrastructure.
Read `INTENT.md` for the thesis, this file for current capability,
`docs/TestDriverEvidenceAndVariants.md` for retained-run/variant examples,
`docs/TestDriverGeneralisationReview.md` for readiness constraints, and `workplans/`
for execution status. The concept corpus contains historical proposals and is not
an implementation inventory.