Align scope with intent and add durable evidence and bounded variants
Assistant: codex Assistant-Model: gpt-6-astra Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
This commit is contained in:
parent
8822480b6e
commit
eaf5d348e4
14 changed files with 767 additions and 92 deletions
149
SCOPE.md
149
SCOPE.md
|
|
@ -1,103 +1,80 @@
|
|||
# SCOPE
|
||||
|
||||
> This file helps you quickly understand what this repository is about,
|
||||
> when it is relevant, and when it is not.
|
||||
Assessed 2026-09-28. `test-driver` is an executable **local research prototype**
|
||||
for use-case-driven verification, adaptation classification and deterministic
|
||||
replay. It is not yet a validated autonomous test platform for real systems.
|
||||
|
||||
---
|
||||
## What it can do
|
||||
|
||||
## One-liner
|
||||
| Capability | Implemented boundary and evidence |
|
||||
|---|---|
|
||||
| Deterministic verification | Python UseCase/Claim/Invariant inputs, ordered semantic steps, separate drivers and observers, explicit PASS/FAIL/INCONCLUSIVE. Non-boolean or missing predicate evidence is inconclusive. `src/testdriver/runner.py`, `oracles.py`; `tests/test_kernel_guarantees.py`. |
|
||||
| Multi-user/domain examples | Sharing and revocation, delegated approval, and tenant deletion/recreation run against three synthetic lab domains. Actor stores/canaries and S2/S3 collector attribution are checked. These are in-process diagnostics, not process isolation. `scenarios/`, `tests/test_generalisation.py`. |
|
||||
| Mechanical discovery | A deterministic heuristic finds forms in server-rendered HTML; a recorded-selector control provides comparison. This is not live-model inference or a JavaScript browser engine. `agentic.py`, `browser.py`. |
|
||||
| Evidence-based classification | Compare complete passing reference runs with candidate evidence, explicit intent revisions and postconditions. Distinguish mechanical changes, observed behavior changes, changed intent, realization failures and ambiguity. Cannot infer whether a product behavior change was authorized. `classification.py`, `revisions.py`. |
|
||||
| Crystallization | Check repeated eligible, distinct runs for stable paths; replay frozen paths by step; generate a single-action pytest module importing original predicates. One checked-in browser descendant runs without model involvement. No automatic full T0–T5 lifecycle or general multi-step code generator. `crystallization.py`, `crystallized/`. |
|
||||
| Retained evidence and lineage | Optional local `EvidenceStore` writes versioned, checksummed JSON receipts atomically without replacing an existing run. Receipts contain asset/parent/maturity/variant, intent revisions, scheduled coverage, observations and final verdict. Reloaded packs can be classified. See [usage](docs/TestDriverEvidenceAndVariants.md). |
|
||||
| Bounded adversarial variants | `variants.substitute` derives one actor or argument substitution, retaining the exact independent use case, assertions, postconditions and surface permissions. Resource, tenant and privilege values use the same operation. Scenario revisions prevent changed action intent from qualifying as mere mechanical adaptation. No automatic claim authoring. |
|
||||
| Local research calibration | A 38-mutation catalogue, independent fixture judgments, mechanical/defect experiments and tests that deliberately break framework guarantees. Lab false-adaptation results are bounded experiment results, not a universal safety proof. `lab/GROUND-TRUTH.md`, `research/`, `tests/selfverification/`. |
|
||||
| HTTP boundary checks | Browser and frozen/standalone HTTP bind bearer credentials to a configured origin, including redirects. Runner also validates reported surfaces. Drivers still need pre-action checks; a post-call check cannot undo effects or attest an untruthful report. |
|
||||
|
||||
`test-driver` is a use-case-driven verification framework whose tests mature
|
||||
alongside the software they protect — fluid and agentic while behaviour is hot,
|
||||
deterministic once it cools.
|
||||
## Where it is useful
|
||||
|
||||
---
|
||||
Use it to investigate whether independent Python claims survive changed
|
||||
mechanics, to build small sequential multi-user domain fixtures, to calibrate
|
||||
mutation detection, or to replay a demonstrated stable HTML action. A new domain
|
||||
needs an explicitly authored scenario, driver and independent observation adapter.
|
||||
|
||||
## Core Idea
|
||||
The audit-core E2 module is an approved-intent contract and fixture calibration,
|
||||
**not** a runnable production engagement. It has no deployed custody, cluster,
|
||||
time-window or independent cleanup adapters in this repository.
|
||||
|
||||
A test is treated not as code but as a **verification asset** with identity,
|
||||
intent, evidence, lineage, maturity, temperature and energy. Use cases are the
|
||||
primary behavioural source; integration, journey, multi-user, security and
|
||||
resilience tests are projections of the same use case rather than separate
|
||||
suites.
|
||||
## Limits and readiness
|
||||
|
||||
Verification assets progress along a maturity continuum
|
||||
(`T0 Exploratory → T5 Deterministic`) called **Crystallization**. Agents may
|
||||
explore and realise semantic actions against unstable interfaces; oracles remain
|
||||
deterministic and independent from actors, so the framework can distinguish a
|
||||
legitimate mechanical change from a product defect rather than adapting to
|
||||
whatever the implementation happens to do.
|
||||
- No natural-language use-case parser or autonomous scenario planner. Intent is
|
||||
explicitly authored Python; source/provenance labels are not external attestation.
|
||||
- No live-model experiment or validated agentic cost benefit. The runtime protocol
|
||||
exists; the demonstrated runtime is heuristic and token-free.
|
||||
- No Playwright/browser engine, JavaScript execution, screenshots or DOM timing.
|
||||
- Sequential execution only: no actual grant/revoke races, causal deadlines,
|
||||
concurrency scheduler or general dependency-fault/resilience engine.
|
||||
- No generic CLI, Kubernetes, message-bus or multi-service production adapters.
|
||||
- Evidence storage is opt-in and local. It is not encryption, signing, automatic
|
||||
redaction, retention management or safe storage of arbitrary secrets. Unexpected
|
||||
driver/observer exceptions still propagate; no finalized receipt is promised for
|
||||
an interrupted/crashed process. Strict JSON rejects unsupported evidence values.
|
||||
- No automatic finding-to-regression pipeline, asset registry, maturity promotion,
|
||||
thawing or retirement. Parent ids and mutation history provide limited lineage.
|
||||
|
||||
**Current state:** research prototype. Concept corpus is complete
|
||||
(`INTENT.md`, `docs/`); implementation has not started. See
|
||||
`history/2026-08-22-concept-assessment-swot.md` for the standing assessment and
|
||||
the reasoning behind the current workplan sequence.
|
||||
Temperature, Energy, Confidence, Campaign, Metabolism and automated Retirement
|
||||
were **deliberately removed**, not postponed implementation requirements. The
|
||||
superseding note in INTENT.md and [settlement](docs/TestDriverGeneralisationReview.md)
|
||||
take precedence over its historical lifecycle sections.
|
||||
|
||||
---
|
||||
Outstanding work is tracked in [TD-WP-0003](workplans/TD-WP-0003-generalise-and-settle.md)
|
||||
and [TD-WP-0004](workplans/TD-WP-0004-scope-evidence-and-variants.md), with priorities
|
||||
and evidence in the [scope/intent assessment](history/2026-09-28-121933-scope-intent-assessment.md).
|
||||
|
||||
## In Scope
|
||||
## Non-goals
|
||||
|
||||
- The conceptual model: UseCase, Actor, Scenario, SemanticAction, Observation,
|
||||
Oracle, Verdict, VerificationAsset, Finding, Adaptation, Crystallization.
|
||||
- A deterministic semantic scenario kernel and its evidence format.
|
||||
- The **test-driver lab** — a small mutable application under test carrying
|
||||
labelled mechanical, semantic and defect mutations as ground truth.
|
||||
- Adaptation detection and the defect-vs-adaptation classifier.
|
||||
- Crystallization of agentic realisations into deterministic regression tests.
|
||||
- Security testing expressed as mutation of ordinary use cases.
|
||||
- Self-verification of the framework's own foundational guarantees.
|
||||
- The research control plane: hypotheses, experiments, findings, fitness map.
|
||||
Replacing pytest, browser engines or CI; a distributed test cloud; a vulnerability
|
||||
scanner, load-testing or observability platform; automatic rewriting of semantic
|
||||
requirements; model-based verdicts; exhaustive scenario permutations. These remain
|
||||
outside the intended initial project.
|
||||
|
||||
---
|
||||
## Actual stack and entry points
|
||||
|
||||
## Out of Scope
|
||||
Python ≥3.11, dataclasses and the standard library; pytest for testing. Labs use
|
||||
in-memory domain state and a local stdlib HTTP server. No runtime third-party
|
||||
dependencies, database, SQLite, Pydantic, YAML parser or browser engine is used.
|
||||
|
||||
- Replacing unit-test frameworks, browser automation engines, or CI systems.
|
||||
- Building a load-testing, fuzzing, vulnerability-scanning or observability
|
||||
platform.
|
||||
- Test-management SaaS, distributed test clouds, or multi-tenant hosting.
|
||||
- Making all tests agentic, or using model judgment where a deterministic oracle
|
||||
is available.
|
||||
- Rewriting semantic requirements to match implementation behaviour.
|
||||
- Scale and performance verification before the conceptual model is proven.
|
||||
```bash
|
||||
python3 -m pytest -q
|
||||
python3 -m pytest -q tests/test_reference_scenario.py tests/test_scope_delivery.py
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Relevant When
|
||||
|
||||
- You need the test-driver conceptual vocabulary or its canonical concept set.
|
||||
- You are working on the crystallization, adaptation-classification, or
|
||||
semantic-action binding mechanisms.
|
||||
- You are extending the lab or its mutation catalogue.
|
||||
- You need the framework's hypotheses, fitness scorecard, or evidence format.
|
||||
|
||||
---
|
||||
|
||||
## Not Relevant When
|
||||
|
||||
- You need ordinary unit or component tests for another repo — use that repo's
|
||||
own test stack.
|
||||
- You are looking for fleet coordination or cross-repo memory — that is State
|
||||
Hub.
|
||||
|
||||
---
|
||||
|
||||
## Getting Oriented
|
||||
|
||||
1. `INTENT.md` — purpose, thesis, design heuristics, non-goals.
|
||||
2. `docs/TestDriverConceptModel.md` — the canonical concept set v0.1.
|
||||
3. `docs/TestDriverImprovementLoop.md` — hypotheses, findings taxonomy,
|
||||
self-improvement cycle.
|
||||
4. `docs/TestDriverInitialMilestones.md` — M0–M10 and the prototype success gate.
|
||||
5. `history/2026-08-22-concept-assessment-swot.md` — assessment and the reasons
|
||||
the first workplan reorders those milestones into a vertical spike.
|
||||
6. `workplans/` — current work. Agent instructions: `AGENTS.md`.
|
||||
|
||||
---
|
||||
|
||||
## Stack
|
||||
|
||||
Deliberately boring, per `docs/TestDriverResearchPrototype.md`: Python, pytest,
|
||||
Playwright, Pydantic/dataclasses, YAML, SQLite. One process, one database, one
|
||||
browser engine, one application under test. Novelty belongs in the verification
|
||||
model, not the infrastructure.
|
||||
Read `INTENT.md` for the thesis, this file for current capability,
|
||||
`docs/TestDriverEvidenceAndVariants.md` for retained-run/variant examples,
|
||||
`docs/TestDriverGeneralisationReview.md` for readiness constraints, and `workplans/`
|
||||
for execution status. The concept corpus contains historical proposals and is not
|
||||
an implementation inventory.
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue