# Scope against intent — 2026-09-28 12:19:33 UTC ## Assessment basis Reviewed the repository at `e419bfe` and the delivery in TD-WP-0004. This report separates the baseline gaps from work implemented during this assessment. Sources: [INTENT](../INTENT.md), [scope](../SCOPE.md), [readiness settlement](../docs/TestDriverGeneralisationReview.md), [source](../src/testdriver/), [tests](../tests/), [38-mutation ground truth](../lab/GROUND-TRUTH.md), [E-001 results](../research/evidence/2026-09-28-e001.json), and the [audit-core contract](../usecases/audit_core_e2_tenant_boundary.py). No new production or paid-model experiment was performed. The old SCOPE.md materially understated implementation ("implementation has not started") while overstating the stack and lifecycle model. The repo already had 398 passing tests, three synthetic multi-user domains, qualified crystallization and extensive acceptance guards. It had no browser engine or database, and the Temperature/Energy/Campaign family had already been removed. SCOPE.md now describes implemented behavior and its limits rather than restating the thesis. ## Intent coverage | INTENT area | Demonstrated before this work | Gap / disposition | |---|---|---| | Independent intent, claims and oracles | Frozen Python assertions; provenance admission; S1/S2/S3 separation; missing/invalid results inconclusive; definition revisions | Strong local coverage, but provenance labels and matching collector names are not external attestation. Domain adapters and independent authors remain necessary. | | Fluid-to-deterministic verification | Heuristic HTML discovery, stable-trajectory checks, frozen replay and one generated single-action test | Narrow demonstration, no live model, no general multi-step generator, no automatic full lifecycle. Keep claims bounded. | | Multi-user isolation | Separate actor objects/sessions; store/canary guards; tenant and delegation fixtures | No process security boundary, actual concurrent scheduler or causal/time-window enforcement. | | Security via use-case mutation | 38 seeded SUT mutations and hand-authored adversarial cases | No reusable scenario-side derivation at baseline. T03 delivers explicit actor/argument substitutions; skip/reorder/replay/concurrency and inferred claims are not added. | | Durable evidence and lineage | In-memory JSON-serializable evidence, experimental JSON files, asset and parent fields | Generic runs lacked a retained receipt API and asset lineage in the pack. T02 delivers opt-in local receipts and reload, preserving existing judgment behavior. | | Stable intent during adaptation | Claim/use-case revisions and recorded step coverage | Actor/action argument/permission/postcondition changes lacked a common intent revision. T03 adds scenario revisions and rejects changed intent for automatic acceptance/freezing. | | Real integration / end-to-end | Local stdlib HTTP journey, direct domain adapters; audit-core contract with fixture tests | No independently validated real-system adapter, custody/cleanup execution or operational observation cost evidence. T05 explicitly waits for an approved pilot. | | Resilience and security concurrency | Sequential mutations, expected refusals and missing-evidence checks | Not evidence of races, time-window handling or fault-tolerant external execution. Follow-on design depends on a real pilot's needs. | | Lineage and finding improvement loop | Research findings/decisions, parent ids, generated lineage comments | No automated finding-to-asset ingestion or asset registry. T02/T03 retain basic lineage, not an orchestration service. | | Energy/Temperature/Campaign/Retirement | Removed after generalisation found no decision-making use | Intent's superseding note governs. Deliberate de-scope, not a backlog to resurrect. | | Scale / distributed execution | None | Explicit initial non-goals, not missing milestone work. | ## Most relevant gaps, ranked 1. **Independent real-system validation.** This most limits confidence in the central thesis: synthetic labs do not measure the cost of independent observation, real requirements provenance, ambiguity or false-adaptation harm. Owner: **TD-WP-0004-T05**, waiting for an operator/target owner to select and approve a bounded target, observation path, fixtures and cleanup/custody plan. Implementing the existing audit-core contract without those inputs would not produce valid evidence. This task remains live and the workplan stays blocked. 2. **Durable, reviewable run evidence and lineage.** High immediate value, small infrastructure cost, useful to every future adapter. **Delivered in T02**: versioned local receipts, checksum verification, no-overwrite atomic publish, strict JSON and recorded asset ancestry/variant/final verdict. Opt-in retention avoids silently persisting arbitrary adapter data. No signing/redaction claim. 3. **Reusable adversarial variants with unchanged semantic expectations.** This is a direct part of INTENT's differentiation. **Delivered in T03**: explicit actor/argument substitution in two synthetic domains, preserved claim identity, descendant lineage and scenario-intent admission binding. It produces a test variant, never a new inferred oracle; an expected denial still needs independently authored expectations. A changed scenario requires review even if it passes. 4. **Measured model/browser/authoring value.** Existing records remain canonical: **TD-WP-0003-T01** needs model, run count, spend ceiling and suite policy; **T06** needs browser setup and T01; **T07** needs a prospectively timed fresh independent author. This request does not supply those missing inputs. No duplicate records or fabricated economic/authoring measurements were created. 5. **Broader scheduling and crystallization.** Actual concurrency/resilience and general multi-step code generation could extend the thesis, but should follow a pilot that demonstrates the requirement. They are explicitly outside this workplan's implementation commitment, not silently declared complete. T05's readiness review must decide which capability to fund before claiming the broader INTENT success criterion. ## Delivered behavior and assurance limits [TD-WP-0004](../workplans/TD-WP-0004-scope-evidence-and-variants.md) was registered through `statehub fix-consistency`. [Evidence/variant usage](../docs/TestDriverEvidenceAndVariants.md) provides reproducible examples. Tests verify storage roundtrips, same-run collisions and concurrent publication, corruption/truncation, permissions, failing/aborted runs, unsupported JSON, stable retained nested observations, preserved assertion objects, two-domain adversarial outcomes and changed-scenario acceptance refusal. Evidence is immutable by the store API, not tamper-proof: a party able to rewrite both receipt and checksum can forge it. No automatic credential redaction is promised; adapters remain responsible for retention-safe observations. Complete receipts cover normal and guard-aborted returns, not unexpected exceptions or process crashes. Historical packs lacking scenario revisions need rerunning for automatic acceptance. Existing partial scenarios remain supported; revisions do not manufacture missing assertions or evaluate unseen state. ## Readiness verdict **Useful local research framework; not ready for autonomous real-system assurance.** The central safety model and local maturation example exist. Durable evidence and bounded variants improve reuse, but do not establish live-model economics, independent production provenance, browser coverage or concurrent correctness. The priority after local delivery is the approved independent pilot, together with the already recorded experiment choices—not rebuilding retired concepts. ## Validation and work status The full suite passed **428 tests in 169.03 seconds**, including 30 new delivery regressions. An earlier existing-boundary subset passed 161 tests. Both documented examples executed successfully; local relative links and `git diff --check` pass. These counts describe regression coverage, not independent production trials. TD-WP-0004-T01–T04 are complete. T05 is waiting and flagged for human input, so TD-WP-0004 remains **blocked**. TD-WP-0003-T01/T06/T07 remain the existing external experiment records. The scope update and local implementation do not close those validation gaps. Decision: `f3b35dce-6047-4a95-8083-4d7032d4e99f`. ## Evidence/preflight correction — 2026-09-28 A subsequent review reproduced tuple-to-list conversion changing an oracle result on receipt replay, and a late unknown actor causing failure after earlier SUT mutations. The follow-up under TD-WP-0004-T02/T03 rejects non-JSON-native snapshots before judgment, retains a safe evidence-failure receipt with pending assertions INCONCLUSIVE, and validates the whole schedule/cast before execution. This closes the reproduced cases without changing claim definitions or expanding pilot scope. Follow-up validation: **440 tests passed** in 169.07 seconds, including 12 new regressions (10 reproduced failures before implementation).