test-driver/SCOPE.md
tegwick 10077edb8c Reject lossy observations and validate schedules before execution
Assistant: codex
Assistant-Model: gpt-6-astra
Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
2026-09-28 14:50:27 +02:00

81 lines
6.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# SCOPE
Assessed 2026-09-28. `test-driver` is an executable **local research prototype**
for use-case-driven verification, adaptation classification and deterministic
replay. It is not yet a validated autonomous test platform for real systems.
## What it can do
| Capability | Implemented boundary and evidence |
|---|---|
| Deterministic verification | Python UseCase/Claim/Invariant inputs, ordered semantic steps, separate drivers and observers, explicit PASS/FAIL/INCONCLUSIVE. Non-boolean or missing predicate evidence is inconclusive. `src/testdriver/runner.py`, `oracles.py`; `tests/test_kernel_guarantees.py`. |
| Multi-user/domain examples | Sharing and revocation, delegated approval, and tenant deletion/recreation run against three synthetic lab domains. Actor stores/canaries and S2/S3 collector attribution are checked. These are in-process diagnostics, not process isolation. `scenarios/`, `tests/test_generalisation.py`. |
| Mechanical discovery | A deterministic heuristic finds forms in server-rendered HTML; a recorded-selector control provides comparison. This is not live-model inference or a JavaScript browser engine. `agentic.py`, `browser.py`. |
| Evidence-based classification | Compare complete passing reference runs with candidate evidence, explicit intent revisions and postconditions. Distinguish mechanical changes, observed behavior changes, changed intent, realization failures and ambiguity. Cannot infer whether a product behavior change was authorized. `classification.py`, `revisions.py`. |
| Crystallization | Check repeated eligible, distinct runs for stable paths; replay frozen paths by step; generate a single-action pytest module importing original predicates. One checked-in browser descendant runs without model involvement. No automatic full T0–T5 lifecycle or general multi-step code generator. `crystallization.py`, `crystallized/`. |
| Retained evidence and lineage | Optional local `EvidenceStore` writes versioned, checksummed JSON receipts atomically without replacing an existing run. Receipts contain asset/parent/maturity/variant, intent revisions, scheduled coverage, observations and final verdict. Reloaded packs can be classified. See [usage](docs/TestDriverEvidenceAndVariants.md). |
| Bounded adversarial variants | `variants.substitute` derives one actor or argument substitution, retaining the exact independent use case, assertions, postconditions and surface permissions. Resource, tenant and privilege values use the same operation. Scenario revisions prevent changed action intent from qualifying as mere mechanical adaptation. No automatic claim authoring. |
| Local research calibration | A 38-mutation catalogue, independent fixture judgments, mechanical/defect experiments and tests that deliberately break framework guarantees. Lab false-adaptation results are bounded experiment results, not a universal safety proof. `lab/GROUND-TRUTH.md`, `research/`, `tests/selfverification/`. |
| HTTP boundary checks | Browser and frozen/standalone HTTP bind bearer credentials to a configured origin, including redirects. Runner also validates reported surfaces. Drivers still need pre-action checks; a post-call check cannot undo effects or attest an untruthful report. |
## Where it is useful
Use it to investigate whether independent Python claims survive changed
mechanics, to build small sequential multi-user domain fixtures, to calibrate
mutation detection, or to replay a demonstrated stable HTML action. A new domain
needs an explicitly authored scenario, driver and independent observation adapter.
The audit-core E2 module is an approved-intent contract and fixture calibration,
**not** a runnable production engagement. It has no deployed custody, cluster,
time-window or independent cleanup adapters in this repository.
## Limits and readiness
- No natural-language use-case parser or autonomous scenario planner. Intent is
explicitly authored Python; source/provenance labels are not external attestation.
- No live-model experiment or validated agentic cost benefit. The runtime protocol
exists; the demonstrated runtime is heuristic and token-free.
- No Playwright/browser engine, JavaScript execution, screenshots or DOM timing.
- Sequential execution only: no actual grant/revoke races, causal deadlines,
concurrency scheduler or general dependency-fault/resilience engine.
- No generic CLI, Kubernetes, message-bus or multi-service production adapters.
- Evidence storage is opt-in and local. It is not encryption, signing, automatic
redaction, retention management or safe storage of arbitrary secrets. Unexpected
driver/observer exceptions still propagate; no finalized receipt is promised for
an interrupted/crashed process. JSON-native snapshots are validated before judgment; lossy values stop the run
with pending judgments INCONCLUSIVE. Full schedule/cast validation precedes actions.
- No automatic finding-to-regression pipeline, asset registry, maturity promotion,
thawing or retirement. Parent ids and mutation history provide limited lineage.
Temperature, Energy, Confidence, Campaign, Metabolism and automated Retirement
were **deliberately removed**, not postponed implementation requirements. The
superseding note in INTENT.md and [settlement](docs/TestDriverGeneralisationReview.md)
take precedence over its historical lifecycle sections.
Outstanding work is tracked in [TD-WP-0003](workplans/TD-WP-0003-generalise-and-settle.md)
and [TD-WP-0004](workplans/TD-WP-0004-scope-evidence-and-variants.md), with priorities
and evidence in the [scope/intent assessment](history/2026-09-28-121933-scope-intent-assessment.md).
## Non-goals
Replacing pytest, browser engines or CI; a distributed test cloud; a vulnerability
scanner, load-testing or observability platform; automatic rewriting of semantic
requirements; model-based verdicts; exhaustive scenario permutations. These remain
outside the intended initial project.
## Actual stack and entry points
Python ≥3.11, dataclasses and the standard library; pytest for testing. Labs use
in-memory domain state and a local stdlib HTTP server. No runtime third-party
dependencies, database, SQLite, Pydantic, YAML parser or browser engine is used.
```bash
python3 -m pytest -q
python3 -m pytest -q tests/test_reference_scenario.py tests/test_scope_delivery.py
```
Read `INTENT.md` for the thesis, this file for current capability,
`docs/TestDriverEvidenceAndVariants.md` for retained-run/variant examples,
`docs/TestDriverGeneralisationReview.md` for readiness constraints, and `workplans/`
for execution status. The concept corpus contains historical proposals and is not
an implementation inventory.