Align scope with intent and add durable evidence and bounded variants
Assistant: codex Assistant-Model: gpt-6-astra Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
This commit is contained in:
parent
8822480b6e
commit
eaf5d348e4
14 changed files with 767 additions and 92 deletions
|
|
@ -13,6 +13,10 @@ multi-user use cases and a 38-mutation catalogue exercise the model. Live-model
|
|||
economics, browser-engine coverage and independent authoring-cost measurement
|
||||
remain blocked in `workplans/TD-WP-0003-generalise-and-settle.md`.
|
||||
See [readiness and interface decisions](docs/TestDriverGeneralisationReview.md).
|
||||
Opt-in local evidence receipts and explicit actor/argument variants are available;
|
||||
see [usage](docs/TestDriverEvidenceAndVariants.md) and the
|
||||
[scope/intent assessment](history/2026-09-28-121933-scope-intent-assessment.md).
|
||||
An independently owned real-system pilot remains blocked in TD-WP-0004.
|
||||
|
||||
## Run
|
||||
|
||||
|
|
|
|||
149
SCOPE.md
149
SCOPE.md
|
|
@ -1,103 +1,80 @@
|
|||
# SCOPE
|
||||
|
||||
> This file helps you quickly understand what this repository is about,
|
||||
> when it is relevant, and when it is not.
|
||||
Assessed 2026-09-28. `test-driver` is an executable **local research prototype**
|
||||
for use-case-driven verification, adaptation classification and deterministic
|
||||
replay. It is not yet a validated autonomous test platform for real systems.
|
||||
|
||||
---
|
||||
## What it can do
|
||||
|
||||
## One-liner
|
||||
| Capability | Implemented boundary and evidence |
|
||||
|---|---|
|
||||
| Deterministic verification | Python UseCase/Claim/Invariant inputs, ordered semantic steps, separate drivers and observers, explicit PASS/FAIL/INCONCLUSIVE. Non-boolean or missing predicate evidence is inconclusive. `src/testdriver/runner.py`, `oracles.py`; `tests/test_kernel_guarantees.py`. |
|
||||
| Multi-user/domain examples | Sharing and revocation, delegated approval, and tenant deletion/recreation run against three synthetic lab domains. Actor stores/canaries and S2/S3 collector attribution are checked. These are in-process diagnostics, not process isolation. `scenarios/`, `tests/test_generalisation.py`. |
|
||||
| Mechanical discovery | A deterministic heuristic finds forms in server-rendered HTML; a recorded-selector control provides comparison. This is not live-model inference or a JavaScript browser engine. `agentic.py`, `browser.py`. |
|
||||
| Evidence-based classification | Compare complete passing reference runs with candidate evidence, explicit intent revisions and postconditions. Distinguish mechanical changes, observed behavior changes, changed intent, realization failures and ambiguity. Cannot infer whether a product behavior change was authorized. `classification.py`, `revisions.py`. |
|
||||
| Crystallization | Check repeated eligible, distinct runs for stable paths; replay frozen paths by step; generate a single-action pytest module importing original predicates. One checked-in browser descendant runs without model involvement. No automatic full T0–T5 lifecycle or general multi-step code generator. `crystallization.py`, `crystallized/`. |
|
||||
| Retained evidence and lineage | Optional local `EvidenceStore` writes versioned, checksummed JSON receipts atomically without replacing an existing run. Receipts contain asset/parent/maturity/variant, intent revisions, scheduled coverage, observations and final verdict. Reloaded packs can be classified. See [usage](docs/TestDriverEvidenceAndVariants.md). |
|
||||
| Bounded adversarial variants | `variants.substitute` derives one actor or argument substitution, retaining the exact independent use case, assertions, postconditions and surface permissions. Resource, tenant and privilege values use the same operation. Scenario revisions prevent changed action intent from qualifying as mere mechanical adaptation. No automatic claim authoring. |
|
||||
| Local research calibration | A 38-mutation catalogue, independent fixture judgments, mechanical/defect experiments and tests that deliberately break framework guarantees. Lab false-adaptation results are bounded experiment results, not a universal safety proof. `lab/GROUND-TRUTH.md`, `research/`, `tests/selfverification/`. |
|
||||
| HTTP boundary checks | Browser and frozen/standalone HTTP bind bearer credentials to a configured origin, including redirects. Runner also validates reported surfaces. Drivers still need pre-action checks; a post-call check cannot undo effects or attest an untruthful report. |
|
||||
|
||||
`test-driver` is a use-case-driven verification framework whose tests mature
|
||||
alongside the software they protect — fluid and agentic while behaviour is hot,
|
||||
deterministic once it cools.
|
||||
## Where it is useful
|
||||
|
||||
---
|
||||
Use it to investigate whether independent Python claims survive changed
|
||||
mechanics, to build small sequential multi-user domain fixtures, to calibrate
|
||||
mutation detection, or to replay a demonstrated stable HTML action. A new domain
|
||||
needs an explicitly authored scenario, driver and independent observation adapter.
|
||||
|
||||
## Core Idea
|
||||
The audit-core E2 module is an approved-intent contract and fixture calibration,
|
||||
**not** a runnable production engagement. It has no deployed custody, cluster,
|
||||
time-window or independent cleanup adapters in this repository.
|
||||
|
||||
A test is treated not as code but as a **verification asset** with identity,
|
||||
intent, evidence, lineage, maturity, temperature and energy. Use cases are the
|
||||
primary behavioural source; integration, journey, multi-user, security and
|
||||
resilience tests are projections of the same use case rather than separate
|
||||
suites.
|
||||
## Limits and readiness
|
||||
|
||||
Verification assets progress along a maturity continuum
|
||||
(`T0 Exploratory → T5 Deterministic`) called **Crystallization**. Agents may
|
||||
explore and realise semantic actions against unstable interfaces; oracles remain
|
||||
deterministic and independent from actors, so the framework can distinguish a
|
||||
legitimate mechanical change from a product defect rather than adapting to
|
||||
whatever the implementation happens to do.
|
||||
- No natural-language use-case parser or autonomous scenario planner. Intent is
|
||||
explicitly authored Python; source/provenance labels are not external attestation.
|
||||
- No live-model experiment or validated agentic cost benefit. The runtime protocol
|
||||
exists; the demonstrated runtime is heuristic and token-free.
|
||||
- No Playwright/browser engine, JavaScript execution, screenshots or DOM timing.
|
||||
- Sequential execution only: no actual grant/revoke races, causal deadlines,
|
||||
concurrency scheduler or general dependency-fault/resilience engine.
|
||||
- No generic CLI, Kubernetes, message-bus or multi-service production adapters.
|
||||
- Evidence storage is opt-in and local. It is not encryption, signing, automatic
|
||||
redaction, retention management or safe storage of arbitrary secrets. Unexpected
|
||||
driver/observer exceptions still propagate; no finalized receipt is promised for
|
||||
an interrupted/crashed process. Strict JSON rejects unsupported evidence values.
|
||||
- No automatic finding-to-regression pipeline, asset registry, maturity promotion,
|
||||
thawing or retirement. Parent ids and mutation history provide limited lineage.
|
||||
|
||||
**Current state:** research prototype. Concept corpus is complete
|
||||
(`INTENT.md`, `docs/`); implementation has not started. See
|
||||
`history/2026-08-22-concept-assessment-swot.md` for the standing assessment and
|
||||
the reasoning behind the current workplan sequence.
|
||||
Temperature, Energy, Confidence, Campaign, Metabolism and automated Retirement
|
||||
were **deliberately removed**, not postponed implementation requirements. The
|
||||
superseding note in INTENT.md and [settlement](docs/TestDriverGeneralisationReview.md)
|
||||
take precedence over its historical lifecycle sections.
|
||||
|
||||
---
|
||||
Outstanding work is tracked in [TD-WP-0003](workplans/TD-WP-0003-generalise-and-settle.md)
|
||||
and [TD-WP-0004](workplans/TD-WP-0004-scope-evidence-and-variants.md), with priorities
|
||||
and evidence in the [scope/intent assessment](history/2026-09-28-121933-scope-intent-assessment.md).
|
||||
|
||||
## In Scope
|
||||
## Non-goals
|
||||
|
||||
- The conceptual model: UseCase, Actor, Scenario, SemanticAction, Observation,
|
||||
Oracle, Verdict, VerificationAsset, Finding, Adaptation, Crystallization.
|
||||
- A deterministic semantic scenario kernel and its evidence format.
|
||||
- The **test-driver lab** — a small mutable application under test carrying
|
||||
labelled mechanical, semantic and defect mutations as ground truth.
|
||||
- Adaptation detection and the defect-vs-adaptation classifier.
|
||||
- Crystallization of agentic realisations into deterministic regression tests.
|
||||
- Security testing expressed as mutation of ordinary use cases.
|
||||
- Self-verification of the framework's own foundational guarantees.
|
||||
- The research control plane: hypotheses, experiments, findings, fitness map.
|
||||
Replacing pytest, browser engines or CI; a distributed test cloud; a vulnerability
|
||||
scanner, load-testing or observability platform; automatic rewriting of semantic
|
||||
requirements; model-based verdicts; exhaustive scenario permutations. These remain
|
||||
outside the intended initial project.
|
||||
|
||||
---
|
||||
## Actual stack and entry points
|
||||
|
||||
## Out of Scope
|
||||
Python ≥3.11, dataclasses and the standard library; pytest for testing. Labs use
|
||||
in-memory domain state and a local stdlib HTTP server. No runtime third-party
|
||||
dependencies, database, SQLite, Pydantic, YAML parser or browser engine is used.
|
||||
|
||||
- Replacing unit-test frameworks, browser automation engines, or CI systems.
|
||||
- Building a load-testing, fuzzing, vulnerability-scanning or observability
|
||||
platform.
|
||||
- Test-management SaaS, distributed test clouds, or multi-tenant hosting.
|
||||
- Making all tests agentic, or using model judgment where a deterministic oracle
|
||||
is available.
|
||||
- Rewriting semantic requirements to match implementation behaviour.
|
||||
- Scale and performance verification before the conceptual model is proven.
|
||||
```bash
|
||||
python3 -m pytest -q
|
||||
python3 -m pytest -q tests/test_reference_scenario.py tests/test_scope_delivery.py
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Relevant When
|
||||
|
||||
- You need the test-driver conceptual vocabulary or its canonical concept set.
|
||||
- You are working on the crystallization, adaptation-classification, or
|
||||
semantic-action binding mechanisms.
|
||||
- You are extending the lab or its mutation catalogue.
|
||||
- You need the framework's hypotheses, fitness scorecard, or evidence format.
|
||||
|
||||
---
|
||||
|
||||
## Not Relevant When
|
||||
|
||||
- You need ordinary unit or component tests for another repo — use that repo's
|
||||
own test stack.
|
||||
- You are looking for fleet coordination or cross-repo memory — that is State
|
||||
Hub.
|
||||
|
||||
---
|
||||
|
||||
## Getting Oriented
|
||||
|
||||
1. `INTENT.md` — purpose, thesis, design heuristics, non-goals.
|
||||
2. `docs/TestDriverConceptModel.md` — the canonical concept set v0.1.
|
||||
3. `docs/TestDriverImprovementLoop.md` — hypotheses, findings taxonomy,
|
||||
self-improvement cycle.
|
||||
4. `docs/TestDriverInitialMilestones.md` — M0–M10 and the prototype success gate.
|
||||
5. `history/2026-08-22-concept-assessment-swot.md` — assessment and the reasons
|
||||
the first workplan reorders those milestones into a vertical spike.
|
||||
6. `workplans/` — current work. Agent instructions: `AGENTS.md`.
|
||||
|
||||
---
|
||||
|
||||
## Stack
|
||||
|
||||
Deliberately boring, per `docs/TestDriverResearchPrototype.md`: Python, pytest,
|
||||
Playwright, Pydantic/dataclasses, YAML, SQLite. One process, one database, one
|
||||
browser engine, one application under test. Novelty belongs in the verification
|
||||
model, not the infrastructure.
|
||||
Read `INTENT.md` for the thesis, this file for current capability,
|
||||
`docs/TestDriverEvidenceAndVariants.md` for retained-run/variant examples,
|
||||
`docs/TestDriverGeneralisationReview.md` for readiness constraints, and `workplans/`
|
||||
for execution status. The concept corpus contains historical proposals and is not
|
||||
an implementation inventory.
|
||||
|
|
|
|||
|
|
@ -11,6 +11,7 @@
|
|||
| workplan | TD-WP-0001 | finished | — | workplans/TD-WP-0001-statehub-bootstrap.md |
|
||||
| workplan | TD-WP-0002 | finished | — | workplans/TD-WP-0002-vertical-spike-crystallization.md |
|
||||
| workplan | TD-WP-0003 | blocked | — | workplans/TD-WP-0003-generalise-and-settle.md |
|
||||
| workplan | TD-WP-0004 | blocked | — | workplans/TD-WP-0004-scope-evidence-and-variants.md |
|
||||
| task | TD-WP-0001-T01 | done | — | workplans/TD-WP-0001-statehub-bootstrap.md |
|
||||
| task | TD-WP-0001-T02 | done | — | workplans/TD-WP-0001-statehub-bootstrap.md |
|
||||
| task | TD-WP-0001-T03 | done | — | workplans/TD-WP-0001-statehub-bootstrap.md |
|
||||
|
|
@ -32,3 +33,8 @@
|
|||
| task | TD-WP-0003-T06 | wait | — | workplans/TD-WP-0003-generalise-and-settle.md |
|
||||
| task | TD-WP-0003-T07 | wait | — | workplans/TD-WP-0003-generalise-and-settle.md |
|
||||
| task | TD-WP-0003-T08 | done | — | workplans/TD-WP-0003-generalise-and-settle.md |
|
||||
| task | TD-WP-0004-T01 | done | — | workplans/TD-WP-0004-scope-evidence-and-variants.md |
|
||||
| task | TD-WP-0004-T02 | done | — | workplans/TD-WP-0004-scope-evidence-and-variants.md |
|
||||
| task | TD-WP-0004-T03 | done | — | workplans/TD-WP-0004-scope-evidence-and-variants.md |
|
||||
| task | TD-WP-0004-T04 | done | — | workplans/TD-WP-0004-scope-evidence-and-variants.md |
|
||||
| task | TD-WP-0004-T05 | wait | — | workplans/TD-WP-0004-scope-evidence-and-variants.md |
|
||||
|
|
|
|||
85
docs/TestDriverEvidenceAndVariants.md
Normal file
85
docs/TestDriverEvidenceAndVariants.md
Normal file
|
|
@ -0,0 +1,85 @@
|
|||
# Retained evidence and explicit adversarial variants
|
||||
|
||||
TD-WP-0004 adds two standard-library APIs. Use Python ≥3.11 from the repository;
|
||||
pytest remains the only dependency of the test suite.
|
||||
|
||||
## Retain, load and compare runs
|
||||
|
||||
```bash
|
||||
PYTHONPATH=src:. python3 - <<'PY'
|
||||
from scenarios.alice_bob_carol import build
|
||||
from testdriver import Runner
|
||||
from testdriver.storage import EvidenceStore
|
||||
from testdriver.classification import classify
|
||||
|
||||
store = EvidenceStore('/tmp/test-driver-evidence')
|
||||
receipts = []
|
||||
for _ in range(2):
|
||||
world, driver, observer, asset, oracle = build()
|
||||
result = Runner(world, driver, observer, oracle).run(asset, evidence_store=store)
|
||||
receipts.append(store.load(result.run_id))
|
||||
print(result.run_id, result.verdict.value)
|
||||
print(classify(*receipts).classification.value) # UNCHANGED
|
||||
PY
|
||||
```
|
||||
|
||||
`EvidenceStore.write(pack)` returns its path. `load(run_id)` returns the evidence
|
||||
mapping expected by classification/crystallization. The version-1 envelope wraps
|
||||
evidence with a SHA-256 checksum. Files publish atomically without overwrite,
|
||||
with mode 0600; newly created store directories request mode 0700. Concurrent
|
||||
publication of the same run yields one winner and FileExistsError for others.
|
||||
Directory fsync uses the local POSIX filesystem API. Storage failures propagate;
|
||||
a caller must not report retained evidence if a write failed.
|
||||
|
||||
Serialization rejects unsupported values, non-string mapping keys and non-finite
|
||||
numbers. Observations detach nested data when recorded, so later domain mutations
|
||||
do not rewrite earlier snapshots. Receipts include the final run verdict and
|
||||
asset id, parent, maturity, variant and adaptation history. Independent claim
|
||||
revisions and scenario revisions bind evidence to the original run inputs.
|
||||
|
||||
Retention is opt-in. It covers finalized returned runs, including FAIL and guard
|
||||
aborts; unexpected driver/observer exceptions or process termination do not promise
|
||||
a complete receipt. The checksum detects damage, not malicious replacement with
|
||||
a recomputed checksum. The store is not encryption, credential redaction, access
|
||||
custody or a retention service. Select observations safe to retain before enabling
|
||||
it; existing driver mechanics can contain action arguments.
|
||||
|
||||
## Derive a security question from the same intent
|
||||
|
||||
```bash
|
||||
PYTHONPATH=src:. python3 - <<'PY'
|
||||
from scenarios.alice_bob_carol import build
|
||||
from testdriver import Runner
|
||||
from testdriver.variants import substitute
|
||||
|
||||
world, driver, observer, parent, oracle = build()
|
||||
variant = substitute(parent, variant_id='grant-to-carol', step_id='s2-grant',
|
||||
arguments={'subject_id': 'carol'})
|
||||
assert variant.scenario.use_case is parent.scenario.use_case
|
||||
result = Runner(world, driver, observer, oracle).run(variant)
|
||||
print(variant.parent_id, result.verdict.value) # FAIL: original independent claims still apply
|
||||
PY
|
||||
```
|
||||
|
||||
`substitute(..., actor_id='other')` changes the scheduled actor; that actor must
|
||||
exist in the execution world's cast. `arguments={...}` changes only existing
|
||||
argument keys and covers resource, tenant or privilege substitution without new
|
||||
framework concepts. Mutable argument data is copied. Unknown/duplicate step ids,
|
||||
unknown keys, empty actor ids and invalid variant ids are rejected. Callers choose
|
||||
unique variant ids for distinct definitions; the API is not an asset registry.
|
||||
|
||||
The exact UseCase, claim/invariant predicates, postconditions, watches, step order
|
||||
and permitted surfaces remain unchanged. Parent identity and mutation metadata
|
||||
are retained without copying changed values into mutation history. A derived
|
||||
variant does not become an approved new requirement, and a denial is not a default
|
||||
PASS: it needs the independently authored assertions appropriate to the question.
|
||||
|
||||
Acceptance fingerprints include scheduled actor/action identities, argument values,
|
||||
order, surface permissions and postcondition definitions. Changed definitions
|
||||
require intent review, even when product verdicts happen to pass. Unsupported
|
||||
runtime dependencies or old packs without this revision cannot be accepted or
|
||||
crystallized automatically. Mechanical selector/endpoint choices remain outside
|
||||
scenario intent and can still qualify as mechanical adaptation.
|
||||
|
||||
This API deliberately does not generate new assertions, skip/reorder steps,
|
||||
schedule races or infer required evidence from a successful implementation run.
|
||||
107
history/2026-09-28-121933-scope-intent-assessment.md
Normal file
107
history/2026-09-28-121933-scope-intent-assessment.md
Normal file
|
|
@ -0,0 +1,107 @@
|
|||
# Scope against intent — 2026-09-28 12:19:33 UTC
|
||||
|
||||
## Assessment basis
|
||||
|
||||
Reviewed the repository at `e419bfe` and the delivery in TD-WP-0004. This report
|
||||
separates the baseline gaps from work implemented during this assessment. Sources:
|
||||
[INTENT](../INTENT.md), [scope](../SCOPE.md), [readiness settlement](../docs/TestDriverGeneralisationReview.md),
|
||||
[source](../src/testdriver/), [tests](../tests/),
|
||||
[38-mutation ground truth](../lab/GROUND-TRUTH.md),
|
||||
[E-001 results](../research/evidence/2026-09-28-e001.json),
|
||||
and the [audit-core contract](../usecases/audit_core_e2_tenant_boundary.py).
|
||||
No new production or paid-model experiment was performed.
|
||||
|
||||
The old SCOPE.md materially understated implementation ("implementation has not
|
||||
started") while overstating the stack and lifecycle model. The repo already had
|
||||
398 passing tests, three synthetic multi-user domains, qualified crystallization
|
||||
and extensive acceptance guards. It had no browser engine or database, and the
|
||||
Temperature/Energy/Campaign family had already been removed. SCOPE.md now describes
|
||||
implemented behavior and its limits rather than restating the thesis.
|
||||
|
||||
## Intent coverage
|
||||
|
||||
| INTENT area | Demonstrated before this work | Gap / disposition |
|
||||
|---|---|---|
|
||||
| Independent intent, claims and oracles | Frozen Python assertions; provenance admission; S1/S2/S3 separation; missing/invalid results inconclusive; definition revisions | Strong local coverage, but provenance labels and matching collector names are not external attestation. Domain adapters and independent authors remain necessary. |
|
||||
| Fluid-to-deterministic verification | Heuristic HTML discovery, stable-trajectory checks, frozen replay and one generated single-action test | Narrow demonstration, no live model, no general multi-step generator, no automatic full lifecycle. Keep claims bounded. |
|
||||
| Multi-user isolation | Separate actor objects/sessions; store/canary guards; tenant and delegation fixtures | No process security boundary, actual concurrent scheduler or causal/time-window enforcement. |
|
||||
| Security via use-case mutation | 38 seeded SUT mutations and hand-authored adversarial cases | No reusable scenario-side derivation at baseline. T03 delivers explicit actor/argument substitutions; skip/reorder/replay/concurrency and inferred claims are not added. |
|
||||
| Durable evidence and lineage | In-memory JSON-serializable evidence, experimental JSON files, asset and parent fields | Generic runs lacked a retained receipt API and asset lineage in the pack. T02 delivers opt-in local receipts and reload, preserving existing judgment behavior. |
|
||||
| Stable intent during adaptation | Claim/use-case revisions and recorded step coverage | Actor/action argument/permission/postcondition changes lacked a common intent revision. T03 adds scenario revisions and rejects changed intent for automatic acceptance/freezing. |
|
||||
| Real integration / end-to-end | Local stdlib HTTP journey, direct domain adapters; audit-core contract with fixture tests | No independently validated real-system adapter, custody/cleanup execution or operational observation cost evidence. T05 explicitly waits for an approved pilot. |
|
||||
| Resilience and security concurrency | Sequential mutations, expected refusals and missing-evidence checks | Not evidence of races, time-window handling or fault-tolerant external execution. Follow-on design depends on a real pilot's needs. |
|
||||
| Lineage and finding improvement loop | Research findings/decisions, parent ids, generated lineage comments | No automated finding-to-asset ingestion or asset registry. T02/T03 retain basic lineage, not an orchestration service. |
|
||||
| Energy/Temperature/Campaign/Retirement | Removed after generalisation found no decision-making use | Intent's superseding note governs. Deliberate de-scope, not a backlog to resurrect. |
|
||||
| Scale / distributed execution | None | Explicit initial non-goals, not missing milestone work. |
|
||||
|
||||
## Most relevant gaps, ranked
|
||||
|
||||
1. **Independent real-system validation.** This most limits confidence in the
|
||||
central thesis: synthetic labs do not measure the cost of independent
|
||||
observation, real requirements provenance, ambiguity or false-adaptation harm.
|
||||
Owner: **TD-WP-0004-T05**, waiting for an operator/target owner to select and
|
||||
approve a bounded target, observation path, fixtures and cleanup/custody plan.
|
||||
Implementing the existing audit-core contract without those inputs would not
|
||||
produce valid evidence. This task remains live and the workplan stays blocked.
|
||||
2. **Durable, reviewable run evidence and lineage.** High immediate value, small
|
||||
infrastructure cost, useful to every future adapter. **Delivered in T02**:
|
||||
versioned local receipts, checksum verification, no-overwrite atomic publish,
|
||||
strict JSON and recorded asset ancestry/variant/final verdict. Opt-in retention
|
||||
avoids silently persisting arbitrary adapter data. No signing/redaction claim.
|
||||
3. **Reusable adversarial variants with unchanged semantic expectations.** This
|
||||
is a direct part of INTENT's differentiation. **Delivered in T03**: explicit
|
||||
actor/argument substitution in two synthetic domains, preserved claim identity,
|
||||
descendant lineage and scenario-intent admission binding. It produces a test
|
||||
variant, never a new inferred oracle; an expected denial still needs independently
|
||||
authored expectations. A changed scenario requires review even if it passes.
|
||||
4. **Measured model/browser/authoring value.** Existing records remain canonical:
|
||||
**TD-WP-0003-T01** needs model, run count, spend ceiling and suite policy;
|
||||
**T06** needs browser setup and T01; **T07** needs a prospectively timed fresh
|
||||
independent author. This request does not supply those missing inputs. No
|
||||
duplicate records or fabricated economic/authoring measurements were created.
|
||||
5. **Broader scheduling and crystallization.** Actual concurrency/resilience and
|
||||
general multi-step code generation could extend the thesis, but should follow
|
||||
a pilot that demonstrates the requirement. They are explicitly outside this
|
||||
workplan's implementation commitment, not silently declared complete. T05's
|
||||
readiness review must decide which capability to fund before claiming the
|
||||
broader INTENT success criterion.
|
||||
|
||||
## Delivered behavior and assurance limits
|
||||
|
||||
[TD-WP-0004](../workplans/TD-WP-0004-scope-evidence-and-variants.md) was registered
|
||||
through `statehub fix-consistency`. [Evidence/variant usage](../docs/TestDriverEvidenceAndVariants.md)
|
||||
provides reproducible examples. Tests verify storage roundtrips, same-run collisions
|
||||
and concurrent publication, corruption/truncation, permissions, failing/aborted
|
||||
runs, unsupported JSON, stable retained nested observations, preserved assertion
|
||||
objects, two-domain adversarial outcomes and changed-scenario acceptance refusal.
|
||||
|
||||
Evidence is immutable by the store API, not tamper-proof: a party able to rewrite
|
||||
both receipt and checksum can forge it. No automatic credential redaction is
|
||||
promised; adapters remain responsible for retention-safe observations. Complete
|
||||
receipts cover normal and guard-aborted returns, not unexpected exceptions or
|
||||
process crashes. Historical packs lacking scenario revisions need rerunning for
|
||||
automatic acceptance. Existing partial scenarios remain supported; revisions do
|
||||
not manufacture missing assertions or evaluate unseen state.
|
||||
|
||||
## Readiness verdict
|
||||
|
||||
**Useful local research framework; not ready for autonomous real-system assurance.**
|
||||
The central safety model and local maturation example exist. Durable evidence
|
||||
and bounded variants improve reuse, but do not establish live-model economics,
|
||||
independent production provenance, browser coverage or concurrent correctness.
|
||||
The priority after local delivery is the approved independent pilot, together
|
||||
with the already recorded experiment choices—not rebuilding retired concepts.
|
||||
|
||||
## Validation and work status
|
||||
|
||||
The full suite passed **428 tests in 169.03 seconds**, including 30 new delivery
|
||||
regressions. An earlier existing-boundary subset passed 161 tests. Both documented
|
||||
examples executed successfully; local relative links and `git diff --check` pass.
|
||||
These counts describe regression coverage, not independent production trials.
|
||||
|
||||
TD-WP-0004-T01–T04 are complete. T05 is waiting and flagged for human input, so
|
||||
TD-WP-0004 remains **blocked**. TD-WP-0003-T01/T06/T07 remain the existing external
|
||||
experiment records. The scope update and local implementation do not close those
|
||||
validation gaps.
|
||||
|
||||
Decision: `f3b35dce-6047-4a95-8083-4d7032d4e99f`.
|
||||
|
|
@ -13,3 +13,4 @@ file is a pointer table, not a second source of truth.
|
|||
| TD-WP-0003 T08 isolation/replay follow-up | Isolation acceptance gate and step-bound frozen replay | `e8fabf6e-3e5e-424e-9a60-152bf3dec641` | `docs/TestDriverGeneralisationReview.md` |
|
||||
| TD-WP-0003 T08 guard-integrity follow-up | Strict boolean judgments and intact isolation guards | `2ecac1c5-5254-4ce9-b092-b47c3e1283a9` | `docs/TestDriverGeneralisationReview.md` |
|
||||
| TD-WP-0003 T08 HTTP/surface follow-up | Origin-bound credentials and runner surface validation | `88490ee8-839d-40dc-affa-b00a1c68e3c5` | `docs/TestDriverGeneralisationReview.md` |
|
||||
| TD-WP-0004 | Evidence-backed scope, local receipts and intent-preserving variants | `f3b35dce-6047-4a95-8083-4d7032d4e99f` | `history/2026-09-28-121933-scope-intent-assessment.md` |
|
||||
|
|
|
|||
|
|
@ -133,7 +133,8 @@ def _surface_fingerprint(pack: Mapping[str, Any]) -> str:
|
|||
|
||||
def _claim_fingerprint(pack: Mapping[str, Any]) -> str:
|
||||
return json.dumps({"provenance": pack["provenance_index"],
|
||||
"revisions": pack["intent_revisions"]}, sort_keys=True)
|
||||
"revisions": pack["intent_revisions"],
|
||||
"scenario_revision": pack["scenario_revision"]}, sort_keys=True)
|
||||
|
||||
|
||||
def _has_complete_evidence(pack: Mapping[str, Any]) -> bool:
|
||||
|
|
@ -189,6 +190,10 @@ def _has_complete_evidence(pack: Mapping[str, Any]) -> bool:
|
|||
return False
|
||||
if any(o["kind"] == "surface_violation" for o in observations):
|
||||
return False
|
||||
revision = pack.get("scenario_revision")
|
||||
if (not isinstance(revision, str) or len(revision) != 64
|
||||
or any(c not in "0123456789abcdef" for c in revision)):
|
||||
return False
|
||||
revisions = pack["intent_revisions"]
|
||||
required_ids = {assertion for assertion, _ in expected_keys} | {pack["use_case_id"]}
|
||||
return (
|
||||
|
|
|
|||
|
|
@ -13,7 +13,8 @@ See docs/TestDriverClassificationDesign.md, Part A.
|
|||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from dataclasses import dataclass, field, asdict
|
||||
from copy import deepcopy
|
||||
from dataclasses import dataclass, field, asdict, replace
|
||||
from datetime import datetime, timezone
|
||||
from enum import Enum
|
||||
from typing import Any
|
||||
|
|
@ -69,8 +70,12 @@ class EvidencePack:
|
|||
expected_judgments: list[dict[str, str]] = field(default_factory=list)
|
||||
intent_revisions: dict[str, str | None] = field(default_factory=dict)
|
||||
|
||||
scenario_revision: str | None = None
|
||||
asset: dict[str, Any] = field(default_factory=dict)
|
||||
run_verdict: str | None = None
|
||||
|
||||
def record(self, observation: Observation) -> None:
|
||||
self.observations.append(observation)
|
||||
self.observations.append(replace(observation, data=deepcopy(observation.data)))
|
||||
|
||||
def of_stratum(self, stratum: Stratum) -> list[Observation]:
|
||||
return [o for o in self.observations if o.stratum is stratum]
|
||||
|
|
@ -80,4 +85,15 @@ class EvidencePack:
|
|||
payload["observations"] = [
|
||||
{**asdict(o), "stratum": o.stratum.value} for o in self.observations
|
||||
]
|
||||
return json.dumps(payload, indent=2, sort_keys=True, default=str)
|
||||
def check_keys(value):
|
||||
if isinstance(value, dict):
|
||||
if any(not isinstance(key, str) for key in value):
|
||||
raise TypeError("evidence objects require string keys")
|
||||
for item in value.values():
|
||||
check_keys(item)
|
||||
elif isinstance(value, (list, tuple)):
|
||||
for item in value:
|
||||
check_keys(item)
|
||||
|
||||
check_keys(payload)
|
||||
return json.dumps(payload, indent=2, sort_keys=True, allow_nan=False)
|
||||
|
|
|
|||
|
|
@ -111,3 +111,17 @@ def intent_revisions(case: UseCase) -> dict[str, str | None]:
|
|||
assertion.source_ref, getattr(assertion, "after_step", None), predicate,
|
||||
])
|
||||
return revisions
|
||||
|
||||
|
||||
def scenario_revision(scenario) -> str | None:
|
||||
"""Bind action intent and execution order, excluding mechanical driver choices."""
|
||||
try:
|
||||
steps = [[step.id, step.actor_id, step.action.name,
|
||||
_describe(dict(step.action.args), set()),
|
||||
sorted(step.action.permitted_surfaces),
|
||||
_describe(step.action.postcondition, set())]
|
||||
for step in scenario.steps]
|
||||
return _digest(["python-scenario-v1", sys.implementation.name,
|
||||
list(sys.version_info[:3]), scenario.use_case.id, steps])
|
||||
except (KeyError, TypeError, ValueError, RecursionError):
|
||||
return None
|
||||
|
|
|
|||
|
|
@ -9,6 +9,7 @@ independence claim structural rather than procedural.
|
|||
from __future__ import annotations
|
||||
|
||||
import uuid
|
||||
from copy import deepcopy
|
||||
from dataclasses import dataclass
|
||||
from datetime import datetime, timezone
|
||||
from typing import Any
|
||||
|
|
@ -18,7 +19,7 @@ from .drivers import Driver, realize_step
|
|||
from .evidence import EvidencePack, Observation, Stratum
|
||||
from .observers import StateObserver
|
||||
from .oracles import Judgment, Oracle, Verdict, overall
|
||||
from .revisions import intent_revisions
|
||||
from .revisions import intent_revisions, scenario_revision
|
||||
from .scenario import Scenario, VerificationAsset
|
||||
from .world import World
|
||||
|
||||
|
|
@ -131,7 +132,7 @@ class Runner:
|
|||
|
||||
# -- execution --------------------------------------------------------
|
||||
|
||||
def run(self, asset: VerificationAsset) -> RunResult:
|
||||
def run(self, asset: VerificationAsset, *, evidence_store=None) -> RunResult:
|
||||
scenario: Scenario = asset.scenario
|
||||
run_id = f"run-{uuid.uuid4().hex[:12]}"
|
||||
pack = EvidencePack(
|
||||
|
|
@ -140,6 +141,10 @@ class Runner:
|
|||
use_case_id=scenario.use_case.id,
|
||||
sut_version=self._world.sut_version,
|
||||
intent_revisions=intent_revisions(scenario.use_case),
|
||||
scenario_revision=scenario_revision(scenario),
|
||||
asset={"id": asset.id, "parent_id": asset.parent_id,
|
||||
"maturity": asset.maturity, "variant": scenario.variant,
|
||||
"adaptation_history": deepcopy(asset.adaptation_history)},
|
||||
scheduled_steps=[step.id for step in scenario.steps],
|
||||
expected_judgments=[
|
||||
{"assertion_id": assertion.id, "step_id": step.id}
|
||||
|
|
@ -301,4 +306,7 @@ class Runner:
|
|||
# Preserve observed failures, but a passing prefix cannot certify an abort.
|
||||
if aborted and result_verdict is Verdict.PASS:
|
||||
result_verdict = Verdict.INCONCLUSIVE
|
||||
pack.run_verdict = result_verdict.value
|
||||
if evidence_store is not None:
|
||||
evidence_store.write(pack)
|
||||
return RunResult(run_id, result_verdict, judgments, pack)
|
||||
|
|
|
|||
72
src/testdriver/storage.py
Normal file
72
src/testdriver/storage.py
Normal file
|
|
@ -0,0 +1,72 @@
|
|||
"""Local, versioned evidence receipts. Checksums detect damage, not forgery."""
|
||||
import hashlib
|
||||
import json
|
||||
import os
|
||||
from pathlib import Path
|
||||
import re
|
||||
import tempfile
|
||||
|
||||
from .evidence import EvidencePack
|
||||
|
||||
|
||||
def _canonical(payload):
|
||||
return json.dumps(payload, sort_keys=True, separators=(',', ':'), allow_nan=False).encode()
|
||||
|
||||
|
||||
class EvidenceStore:
|
||||
"""Atomically publish one immutable-by-API receipt per run, without overwrites.
|
||||
|
||||
Evidence must already be suitable for retention. This is not a secret scrubber,
|
||||
authenticated signature, remote store or retention policy.
|
||||
"""
|
||||
|
||||
def __init__(self, directory):
|
||||
self.directory = Path(directory)
|
||||
|
||||
def _path(self, run_id):
|
||||
if not isinstance(run_id, str) or not re.fullmatch(r'[A-Za-z0-9][A-Za-z0-9_-]{0,127}', run_id):
|
||||
raise ValueError('invalid evidence run id')
|
||||
return self.directory / f'{run_id}.json'
|
||||
|
||||
def write(self, pack: EvidencePack) -> Path:
|
||||
path = self._path(pack.run_id)
|
||||
if not pack.finished_at or pack.run_verdict not in ('PASS', 'FAIL', 'INCONCLUSIVE'):
|
||||
raise ValueError('only finalized run evidence can be stored')
|
||||
payload = json.loads(pack.to_json()) # strict serialization before any write
|
||||
envelope = {'schema_version': 1, 'evidence': payload,
|
||||
'sha256': hashlib.sha256(_canonical(payload)).hexdigest()}
|
||||
encoded = _canonical(envelope) + b'\n'
|
||||
self.directory.mkdir(mode=0o700, parents=True, exist_ok=True)
|
||||
temporary = None
|
||||
try:
|
||||
with tempfile.NamedTemporaryFile(dir=self.directory, prefix='.evidence-', delete=False) as stream:
|
||||
temporary = Path(stream.name)
|
||||
stream.write(encoded)
|
||||
stream.flush()
|
||||
os.fsync(stream.fileno())
|
||||
# link publishes atomically and fails if the destination already exists.
|
||||
os.link(temporary, path)
|
||||
directory_fd = os.open(self.directory, os.O_RDONLY | os.O_DIRECTORY)
|
||||
try:
|
||||
os.fsync(directory_fd)
|
||||
finally:
|
||||
os.close(directory_fd)
|
||||
finally:
|
||||
if temporary is not None:
|
||||
temporary.unlink(missing_ok=True)
|
||||
return path
|
||||
|
||||
def load(self, run_id) -> dict:
|
||||
path = self._path(run_id)
|
||||
try:
|
||||
envelope = json.loads(path.read_text())
|
||||
if type(envelope['schema_version']) is not int or envelope['schema_version'] != 1:
|
||||
raise ValueError('unsupported evidence schema version')
|
||||
payload = envelope['evidence']
|
||||
if (hashlib.sha256(_canonical(payload)).hexdigest() != envelope['sha256']
|
||||
or payload['run_id'] != run_id or not payload['finished_at']
|
||||
or payload['run_verdict'] not in ('PASS', 'FAIL', 'INCONCLUSIVE')):
|
||||
raise ValueError('invalid evidence receipt or checksum')
|
||||
return payload
|
||||
except (KeyError, TypeError, json.JSONDecodeError) as exc:
|
||||
raise ValueError('malformed evidence receipt') from exc
|
||||
47
src/testdriver/variants.py
Normal file
47
src/testdriver/variants.py
Normal file
|
|
@ -0,0 +1,47 @@
|
|||
"""Explicit adversarial substitutions preserve the independent assertion set."""
|
||||
from copy import deepcopy
|
||||
from dataclasses import replace
|
||||
import re
|
||||
|
||||
from .scenario import VerificationAsset
|
||||
|
||||
|
||||
def substitute(asset: VerificationAsset, *, variant_id: str, step_id: str,
|
||||
actor_id: str | None = None, arguments: dict | None = None) -> VerificationAsset:
|
||||
"""Derive one actor/resource/tenant/privilege substitution, never its verdict.
|
||||
|
||||
New actor identities must exist in the execution world's cast. Argument keys
|
||||
must already exist on the action; domain values remain the author's choice.
|
||||
Expectations, surfaces and postconditions stay exactly as independently authored.
|
||||
"""
|
||||
if not isinstance(variant_id, str) or not re.fullmatch(r'[A-Za-z0-9][A-Za-z0-9_-]{0,63}', variant_id):
|
||||
raise ValueError('invalid variant id')
|
||||
if actor_id is not None and (not isinstance(actor_id, str) or not actor_id):
|
||||
raise ValueError('invalid actor id')
|
||||
if arguments is not None and not isinstance(arguments, dict):
|
||||
raise ValueError('argument substitutions must be a dictionary')
|
||||
matching = [step for step in asset.scenario.steps if step.id == step_id]
|
||||
if len(matching) != 1:
|
||||
raise ValueError('substitution requires exactly one matching step')
|
||||
target = matching[0]
|
||||
if set(arguments or {}) - target.action.args.keys():
|
||||
raise ValueError('substitution names an unknown argument')
|
||||
if actor_id is None and not arguments:
|
||||
raise ValueError('substitution requires an actor or argument change')
|
||||
steps = []
|
||||
for step in asset.scenario.steps:
|
||||
args = deepcopy(dict(step.action.args))
|
||||
if step.id == step_id:
|
||||
args.update(deepcopy(arguments or {}))
|
||||
steps.append(replace(step, actor_id=actor_id if step.id == step_id and actor_id is not None else step.actor_id,
|
||||
action=replace(step.action, args=args)))
|
||||
scenario = replace(asset.scenario, id=f'{asset.scenario.id}--{variant_id}',
|
||||
variant=variant_id, steps=tuple(steps))
|
||||
# Keep values out of lineage: mechanics/evidence may have their own retention policy.
|
||||
history = deepcopy(asset.adaptation_history) + [{
|
||||
'kind': 'adversarial-substitution', 'parent_id': asset.id,
|
||||
'variant': variant_id, 'step_id': step_id, 'actor_changed': actor_id is not None,
|
||||
'argument_keys': sorted(arguments or {}),
|
||||
}]
|
||||
return VerificationAsset(f'{asset.id}--{variant_id}', scenario, maturity=asset.maturity,
|
||||
parent_id=asset.id, adaptation_history=history)
|
||||
199
tests/test_scope_delivery.py
Normal file
199
tests/test_scope_delivery.py
Normal file
|
|
@ -0,0 +1,199 @@
|
|||
"""Intent preservation, durable receipts and reusable adversarial variants."""
|
||||
from concurrent.futures import ThreadPoolExecutor
|
||||
from dataclasses import replace
|
||||
import json
|
||||
import stat
|
||||
|
||||
import pytest
|
||||
|
||||
from scenarios.alice_bob_carol import build
|
||||
from scenarios.tenant_lifecycle import build as tenant
|
||||
from testdriver import Runner, Verdict
|
||||
from testdriver.classification import classify, Classification
|
||||
from testdriver.crystallization import assess_stability
|
||||
from testdriver.evidence import Observation, Stratum
|
||||
from testdriver.storage import EvidenceStore
|
||||
from testdriver.variants import substitute
|
||||
|
||||
|
||||
def run(builder=build, store=None):
|
||||
world, driver, observer, asset, oracle = builder()
|
||||
return Runner(world, driver, observer, oracle).run(asset, evidence_store=store)
|
||||
|
||||
|
||||
def test_store_roundtrip_retains_lineage_and_admission(tmp_path):
|
||||
store = EvidenceStore(tmp_path / 'evidence')
|
||||
results = [run(store=store) for _ in range(3)]
|
||||
packs = [store.load(result.run_id) for result in results]
|
||||
assert packs[0] == json.loads(results[0].evidence.to_json())
|
||||
assert packs[0]['asset']['id'] == build()[3].id
|
||||
assert packs[0]['run_verdict'] == 'PASS'
|
||||
assert classify(packs[0], packs[1]).safe_to_accept
|
||||
assert assess_stability(packs).stable
|
||||
assert stat.S_IMODE((store.directory / f'{results[0].run_id}.json').stat().st_mode) == 0o600
|
||||
with pytest.raises(FileExistsError):
|
||||
store.write(results[0].evidence)
|
||||
assert not list(store.directory.glob('.evidence-*'))
|
||||
|
||||
|
||||
def test_failed_and_guard_aborted_runs_are_retained(tmp_path):
|
||||
store = EvidenceStore(tmp_path)
|
||||
failing = run(lambda: build('M17'), store)
|
||||
assert store.load(failing.run_id)['run_verdict'] == 'FAIL'
|
||||
w, d, o, a, oracle = build()
|
||||
w.cast['bob']._memory = w.cast['alice']._memory
|
||||
aborted = Runner(w, d, o, oracle).run(a, evidence_store=store)
|
||||
assert store.load(aborted.run_id)['run_verdict'] == 'INCONCLUSIVE'
|
||||
|
||||
|
||||
@pytest.mark.parametrize('damage', ['checksum', 'schema', 'truncated', 'wrong-id'])
|
||||
def test_corrupt_receipts_are_rejected(tmp_path, damage):
|
||||
store = EvidenceStore(tmp_path)
|
||||
result = run(store=store)
|
||||
path = tmp_path / f'{result.run_id}.json'
|
||||
envelope = json.loads(path.read_text())
|
||||
if damage == 'checksum':
|
||||
envelope['evidence']['run_verdict'] = 'FAIL'
|
||||
elif damage == 'schema':
|
||||
envelope['schema_version'] = 999
|
||||
elif damage == 'wrong-id':
|
||||
other = tmp_path / 'other.json'
|
||||
other.write_text(path.read_text())
|
||||
with pytest.raises(ValueError):
|
||||
store.load('other')
|
||||
return
|
||||
path.write_text('{' if damage == 'truncated' else json.dumps(envelope))
|
||||
with pytest.raises(ValueError):
|
||||
store.load(result.run_id)
|
||||
|
||||
|
||||
@pytest.mark.parametrize('run_id', ['../escape', '/absolute', '.', '', 'a/b'])
|
||||
def test_store_rejects_unsafe_identifiers_without_writes(tmp_path, run_id):
|
||||
pack = run().evidence
|
||||
pack.run_id = run_id
|
||||
with pytest.raises(ValueError):
|
||||
EvidenceStore(tmp_path).write(pack)
|
||||
assert not list(tmp_path.iterdir())
|
||||
|
||||
|
||||
@pytest.mark.parametrize('bad', [object(), float('nan'), {1: 'coerced-key'}])
|
||||
def test_unserializable_evidence_is_not_silently_stringified(tmp_path, bad):
|
||||
pack = run().evidence
|
||||
pack.asset['unsupported'] = bad
|
||||
with pytest.raises((TypeError, ValueError)):
|
||||
EvidenceStore(tmp_path).write(pack)
|
||||
assert not list(tmp_path.iterdir())
|
||||
|
||||
|
||||
def test_concurrent_writers_publish_exactly_one_complete_receipt(tmp_path):
|
||||
store = EvidenceStore(tmp_path)
|
||||
pack = run().evidence
|
||||
def write():
|
||||
try:
|
||||
store.write(pack)
|
||||
return True
|
||||
except FileExistsError:
|
||||
return False
|
||||
with ThreadPoolExecutor(max_workers=4) as workers:
|
||||
assert sum(workers.map(lambda _: write(), range(4))) == 1
|
||||
assert store.load(pack.run_id)['run_id'] == pack.run_id
|
||||
assert not list(tmp_path.glob('.evidence-*'))
|
||||
|
||||
|
||||
def test_runner_propagates_storage_failure(tmp_path):
|
||||
path = tmp_path / 'file'
|
||||
path.write_text('not a directory')
|
||||
with pytest.raises(FileExistsError):
|
||||
run(store=EvidenceStore(path))
|
||||
|
||||
|
||||
def test_recorded_observations_do_not_change_with_live_nested_objects():
|
||||
pack = run().evidence
|
||||
snapshot = {'nested': {'values': [1]}}
|
||||
pack.record(Observation('custom', Stratum.JUDGMENT, 'observer', None, 'state', snapshot))
|
||||
snapshot['nested']['values'].append(2)
|
||||
assert pack.observations[-1].data == {'nested': {'values': [1]}}
|
||||
|
||||
|
||||
def test_variant_preserves_claims_surfaces_and_parent_and_copies_arguments(tmp_path):
|
||||
world, driver, observer, parent, oracle = build()
|
||||
step = parent.scenario.steps[1]
|
||||
variant = substitute(parent, variant_id='carol-read', step_id=step.id,
|
||||
arguments={'subject_id': 'carol'})
|
||||
assert variant.parent_id == parent.id
|
||||
assert variant.scenario.use_case is parent.scenario.use_case
|
||||
for original, changed in zip(parent.scenario.steps, variant.scenario.steps):
|
||||
assert original.action.postcondition is changed.action.postcondition
|
||||
assert original.action.permitted_surfaces == changed.action.permitted_surfaces
|
||||
assert original.action.args is not changed.action.args
|
||||
assert step.action.args['subject_id'] == 'bob'
|
||||
result = Runner(world, driver, observer, oracle).run(variant, evidence_store=EvidenceStore(tmp_path))
|
||||
assert result.verdict is Verdict.FAIL
|
||||
assert result.evidence.asset['parent_id'] == parent.id
|
||||
assert result.judgment('c-carol-denied').verdict is Verdict.FAIL
|
||||
|
||||
|
||||
def test_actor_substitution_applies_to_another_domain_without_new_claims():
|
||||
world, driver, observer, parent, oracle = tenant()
|
||||
variant = substitute(parent, variant_id='foreign-creator', step_id='create-a', actor_id='admin-b')
|
||||
result = Runner(world, driver, observer, oracle).run(variant)
|
||||
assert result.judgment('tenant-create-a').verdict is Verdict.FAIL
|
||||
assert variant.scenario.use_case is parent.scenario.use_case
|
||||
|
||||
|
||||
@pytest.mark.parametrize('changes', [
|
||||
{'step_id': 'absent', 'actor_id': 'bob'},
|
||||
{'step_id': 's2-grant', 'arguments': {'absent': 1}},
|
||||
{'step_id': 's2-grant', 'actor_id': ''},
|
||||
{'step_id': 's2-grant'},
|
||||
])
|
||||
def test_invalid_substitutions_do_not_modify_parent(changes):
|
||||
parent = build()[3]
|
||||
before = repr(parent)
|
||||
with pytest.raises(ValueError):
|
||||
substitute(parent, variant_id='invalid', **changes)
|
||||
assert repr(parent) == before
|
||||
|
||||
|
||||
@pytest.mark.parametrize('change', ['actor', 'argument', 'surface', 'postcondition', 'order'])
|
||||
def test_scenario_intent_changes_require_review_even_with_passing_verdicts(change):
|
||||
baseline = json.loads(run().evidence.to_json())
|
||||
world, driver, observer, asset, oracle = build()
|
||||
steps = list(asset.scenario.steps)
|
||||
step = steps[0]
|
||||
if change == 'actor':
|
||||
# Identical authorized operation via another identity, without changing claim definitions.
|
||||
driver._tokens['carol'] = driver._tokens['alice']
|
||||
steps[0] = replace(step, actor_id='carol')
|
||||
elif change == 'argument':
|
||||
steps[0] = replace(step, action=replace(step.action, args={**step.action.args, 'content': 'changed'}))
|
||||
elif change == 'surface':
|
||||
steps[0] = replace(step, action=replace(step.action, permitted_surfaces=frozenset({'api', 'browser'})))
|
||||
elif change == 'postcondition':
|
||||
steps[0] = replace(step, action=replace(step.action, postcondition=lambda obs: True))
|
||||
else:
|
||||
# A schedule-only assertion-free no-op pair can be reordered without changing outcomes.
|
||||
steps.extend([replace(step, id='extra-a'), replace(step, id='extra-b')])
|
||||
steps[-2:] = reversed(steps[-2:])
|
||||
asset.scenario = replace(asset.scenario, steps=tuple(steps))
|
||||
result = Runner(world, driver, observer, oracle).run(asset)
|
||||
pack = json.loads(result.evidence.to_json())
|
||||
outcome = classify(baseline, pack)
|
||||
if change != 'order':
|
||||
assert result.verdict is Verdict.PASS
|
||||
assert outcome.classification is Classification.INTENT_CHANGED
|
||||
assert not outcome.safe_to_accept
|
||||
assert outcome.classification in (Classification.INTENT_CHANGED, Classification.AMBIGUOUS)
|
||||
|
||||
|
||||
def test_legacy_packs_require_fresh_scenario_intent_evidence():
|
||||
pack = json.loads(run().evidence.to_json())
|
||||
del pack['scenario_revision']
|
||||
assert not classify(pack, pack).safe_to_accept
|
||||
|
||||
|
||||
def test_action_order_is_part_of_scenario_revision():
|
||||
from testdriver.revisions import scenario_revision
|
||||
scenario = build()[3].scenario
|
||||
assert scenario_revision(scenario) != scenario_revision(
|
||||
replace(scenario, steps=tuple(reversed(scenario.steps))))
|
||||
134
workplans/TD-WP-0004-scope-evidence-and-variants.md
Normal file
134
workplans/TD-WP-0004-scope-evidence-and-variants.md
Normal file
|
|
@ -0,0 +1,134 @@
|
|||
---
|
||||
id: TD-WP-0004
|
||||
type: workplan
|
||||
title: "Align scope with evidence and close durable-run and variant gaps"
|
||||
domain: infotech
|
||||
repo: test-driver
|
||||
status: blocked
|
||||
owner: codex
|
||||
topic_slug: custodian
|
||||
created: "2026-09-28"
|
||||
updated: "2026-09-28"
|
||||
state_hub_workstream_id: "efd2200b-4856-5655-a010-0c3b3a858391"
|
||||
---
|
||||
|
||||
# Scope, retained evidence and adversarial variants
|
||||
|
||||
User-authorized assessment and implementation following the scope review at
|
||||
`e419bfe`. No new dependencies. Keep the Python intent API, deterministic oracles,
|
||||
sequential execution and standard-library HTTP. Do not revive the lifecycle
|
||||
concepts retired by TD-WP-0003. Evidence is local and must not imply production
|
||||
readiness or proof of arbitrary actor isolation.
|
||||
|
||||
## Assess actual scope against intent
|
||||
|
||||
```task
|
||||
id: TD-WP-0004-T01
|
||||
status: done
|
||||
priority: high
|
||||
state_hub_task_id: "1c75b826-b6de-5aed-bd8b-40ea4217cdcd"
|
||||
```
|
||||
|
||||
Replace stale SCOPE.md claims with implementation-backed capabilities, limitations
|
||||
and actual stack. Write `history/2026-09-28-121933-scope-intent-assessment.md` with
|
||||
an intent/capability matrix, ranked gaps, evidence links and live work ownership.
|
||||
Distinguish missing capabilities from deliberately removed concepts.
|
||||
|
||||
## Retain complete run evidence and asset lineage
|
||||
|
||||
```task
|
||||
id: TD-WP-0004-T02
|
||||
status: done
|
||||
priority: high
|
||||
state_hub_task_id: "e6c957db-e53d-5bcf-957b-ff0fac4c2e77"
|
||||
```
|
||||
|
||||
Add an opt-in local evidence store with strict JSON, schema version, safe run-id
|
||||
filenames, atomic no-overwrite publication, restrictive file permissions, reload
|
||||
and tamper/corruption detection. Record asset id, parent, maturity and variant in
|
||||
run evidence. Runner can persist completed/guard-aborted runs on request; storage
|
||||
failure must be explicit. Do not serialize arbitrary Python objects as strings,
|
||||
or claim storage is encryption, redaction, signing or a replacement for custody.
|
||||
Validate roundtrip, collisions, corruption, unsupported data and failing runs.
|
||||
|
||||
## Derive bounded adversarial variants without rewriting claims
|
||||
|
||||
```task
|
||||
id: TD-WP-0004-T03
|
||||
status: done
|
||||
priority: high
|
||||
state_hub_task_id: "7ce7ca92-7fff-5d2b-8482-db64049c1429"
|
||||
```
|
||||
|
||||
Provide reusable actor and argument substitution (including resource, tenant and
|
||||
privilege values) for one scheduled step. Retain exact UseCase, predicates,
|
||||
postconditions, schedule and permitted surfaces; copy mutable argument data and
|
||||
record descendant identity/parent/mutation history. Reject unknown steps/keys and
|
||||
invalid identities. No automatic discovery, concurrent scheduler, skip/reorder
|
||||
or inferred security assertions.
|
||||
|
||||
Bind acceptance to a conservative revision of scheduled action intent (actor,
|
||||
order, args, permitted surfaces and postcondition). Changed/unsupported scenario
|
||||
intent must not be accepted/frozen as a mechanical change. Verify controls and
|
||||
adversarial behavior against at least two synthetic domains, without modifying
|
||||
lab implementations to manufacture the result.
|
||||
|
||||
## Publish reproducible usage and validate the implemented scope
|
||||
|
||||
```task
|
||||
id: TD-WP-0004-T04
|
||||
status: done
|
||||
priority: medium
|
||||
state_hub_task_id: "4d0be7bb-5c76-5d68-972c-97cda53e0621"
|
||||
```
|
||||
|
||||
Document executable evidence-store and variant examples, retained-data limits and
|
||||
new admission requirements. Run targeted regressions and the full suite. Update
|
||||
the assessment with delivered capability, remaining gaps and validation counts.
|
||||
Register/sync records and commit/push implementation and documents.
|
||||
|
||||
## Validate an independently owned real-system pilot
|
||||
|
||||
```task
|
||||
id: TD-WP-0004-T05
|
||||
status: wait
|
||||
priority: high
|
||||
state_hub_task_id: "aeacea6b-20db-5dd6-9585-4d7fefe412a2"
|
||||
```
|
||||
|
||||
Blocked pending an operator/target-owner-selected system and approved bounded
|
||||
engagement: independent requirements, observation adapter and cost, test fixture
|
||||
and cleanup authority, custody/expiry routing and a recorded tolerance for false
|
||||
adaptation. No production target or credential access is authorized by this
|
||||
repo-local request. Audit-core E2 is a contract/calibration candidate, not an
|
||||
executable pilot. When prerequisites exist, implement and execute its bounded
|
||||
adapter and retain independent results; until then this task/workplan stays
|
||||
blocked after local completion. Owner of selection/approval: Bernd Worsch and
|
||||
the target owner.
|
||||
|
||||
Existing external work remains in TD-WP-0003-T01 (model/run/budget choices), T06
|
||||
(browser setup and T01) and T07 (fresh independent timed authoring). Do not
|
||||
duplicate those tasks. Longer-term concurrency, resilience fault scheduling and
|
||||
multi-step artifact generation are follow-on decisions informed by the pilot,
|
||||
not implementation commitments in this bounded workplan.
|
||||
|
||||
|
||||
## Local implementation closeout — 2026-09-28
|
||||
|
||||
T01–T04 are done. SCOPE.md now reflects executable capability and actual stack;
|
||||
`history/2026-09-28-121933-scope-intent-assessment.md` ranks the intent gaps and
|
||||
records the delivered changes. EvidenceStore, detached/strict evidence with
|
||||
lineage, reusable substitutions and scenario-definition admission binding are
|
||||
implemented. Usage examples ran successfully; relative document links resolve.
|
||||
|
||||
Validation: 30 new delivery regressions passed; the existing acceptance,
|
||||
classification, generalisation and completeness subset passed 161 tests. The
|
||||
final full suite passed **428 tests in 169.03 seconds**. `git diff --check` is
|
||||
clean. Decision: `f3b35dce-6047-4a95-8083-4d7032d4e99f`.
|
||||
|
||||
T05 remains **wait**, flagged for human input with target/engagement prerequisites;
|
||||
therefore this workplan is **blocked**, not finished. No residual is hidden in
|
||||
scope prose: the pilot remains live here, and model/browser/authoring experiments
|
||||
remain TD-WP-0003-T01/T06/T07. Broader scheduling/lifecycle/adapter features remain
|
||||
explicit scope limits to prioritize from actual pilot needs. No real-system,
|
||||
paid-model or browser-engine execution was performed.
|
||||
Loading…
Add table
Add a link
Reference in a new issue