test-driver/history/2026-09-28-121933-scope-intent-assessment.md

108 lines
8.3 KiB
Markdown
Raw Normal View History

# Scope against intent — 2026-09-28 12:19:33 UTC
## Assessment basis
Reviewed the repository at `e419bfe` and the delivery in TD-WP-0004. This report
separates the baseline gaps from work implemented during this assessment. Sources:
[INTENT](../INTENT.md), [scope](../SCOPE.md), [readiness settlement](../docs/TestDriverGeneralisationReview.md),
[source](../src/testdriver/), [tests](../tests/),
[38-mutation ground truth](../lab/GROUND-TRUTH.md),
[E-001 results](../research/evidence/2026-09-28-e001.json),
and the [audit-core contract](../usecases/audit_core_e2_tenant_boundary.py).
No new production or paid-model experiment was performed.
The old SCOPE.md materially understated implementation ("implementation has not
started") while overstating the stack and lifecycle model. The repo already had
398 passing tests, three synthetic multi-user domains, qualified crystallization
and extensive acceptance guards. It had no browser engine or database, and the
Temperature/Energy/Campaign family had already been removed. SCOPE.md now describes
implemented behavior and its limits rather than restating the thesis.
## Intent coverage
| INTENT area | Demonstrated before this work | Gap / disposition |
|---|---|---|
| Independent intent, claims and oracles | Frozen Python assertions; provenance admission; S1/S2/S3 separation; missing/invalid results inconclusive; definition revisions | Strong local coverage, but provenance labels and matching collector names are not external attestation. Domain adapters and independent authors remain necessary. |
| Fluid-to-deterministic verification | Heuristic HTML discovery, stable-trajectory checks, frozen replay and one generated single-action test | Narrow demonstration, no live model, no general multi-step generator, no automatic full lifecycle. Keep claims bounded. |
| Multi-user isolation | Separate actor objects/sessions; store/canary guards; tenant and delegation fixtures | No process security boundary, actual concurrent scheduler or causal/time-window enforcement. |
| Security via use-case mutation | 38 seeded SUT mutations and hand-authored adversarial cases | No reusable scenario-side derivation at baseline. T03 delivers explicit actor/argument substitutions; skip/reorder/replay/concurrency and inferred claims are not added. |
| Durable evidence and lineage | In-memory JSON-serializable evidence, experimental JSON files, asset and parent fields | Generic runs lacked a retained receipt API and asset lineage in the pack. T02 delivers opt-in local receipts and reload, preserving existing judgment behavior. |
| Stable intent during adaptation | Claim/use-case revisions and recorded step coverage | Actor/action argument/permission/postcondition changes lacked a common intent revision. T03 adds scenario revisions and rejects changed intent for automatic acceptance/freezing. |
| Real integration / end-to-end | Local stdlib HTTP journey, direct domain adapters; audit-core contract with fixture tests | No independently validated real-system adapter, custody/cleanup execution or operational observation cost evidence. T05 explicitly waits for an approved pilot. |
| Resilience and security concurrency | Sequential mutations, expected refusals and missing-evidence checks | Not evidence of races, time-window handling or fault-tolerant external execution. Follow-on design depends on a real pilot's needs. |
| Lineage and finding improvement loop | Research findings/decisions, parent ids, generated lineage comments | No automated finding-to-asset ingestion or asset registry. T02/T03 retain basic lineage, not an orchestration service. |
| Energy/Temperature/Campaign/Retirement | Removed after generalisation found no decision-making use | Intent's superseding note governs. Deliberate de-scope, not a backlog to resurrect. |
| Scale / distributed execution | None | Explicit initial non-goals, not missing milestone work. |
## Most relevant gaps, ranked
1. **Independent real-system validation.** This most limits confidence in the
central thesis: synthetic labs do not measure the cost of independent
observation, real requirements provenance, ambiguity or false-adaptation harm.
Owner: **TD-WP-0004-T05**, waiting for an operator/target owner to select and
approve a bounded target, observation path, fixtures and cleanup/custody plan.
Implementing the existing audit-core contract without those inputs would not
produce valid evidence. This task remains live and the workplan stays blocked.
2. **Durable, reviewable run evidence and lineage.** High immediate value, small
infrastructure cost, useful to every future adapter. **Delivered in T02**:
versioned local receipts, checksum verification, no-overwrite atomic publish,
strict JSON and recorded asset ancestry/variant/final verdict. Opt-in retention
avoids silently persisting arbitrary adapter data. No signing/redaction claim.
3. **Reusable adversarial variants with unchanged semantic expectations.** This
is a direct part of INTENT's differentiation. **Delivered in T03**: explicit
actor/argument substitution in two synthetic domains, preserved claim identity,
descendant lineage and scenario-intent admission binding. It produces a test
variant, never a new inferred oracle; an expected denial still needs independently
authored expectations. A changed scenario requires review even if it passes.
4. **Measured model/browser/authoring value.** Existing records remain canonical:
**TD-WP-0003-T01** needs model, run count, spend ceiling and suite policy;
**T06** needs browser setup and T01; **T07** needs a prospectively timed fresh
independent author. This request does not supply those missing inputs. No
duplicate records or fabricated economic/authoring measurements were created.
5. **Broader scheduling and crystallization.** Actual concurrency/resilience and
general multi-step code generation could extend the thesis, but should follow
a pilot that demonstrates the requirement. They are explicitly outside this
workplan's implementation commitment, not silently declared complete. T05's
readiness review must decide which capability to fund before claiming the
broader INTENT success criterion.
## Delivered behavior and assurance limits
[TD-WP-0004](../workplans/TD-WP-0004-scope-evidence-and-variants.md) was registered
through `statehub fix-consistency`. [Evidence/variant usage](../docs/TestDriverEvidenceAndVariants.md)
provides reproducible examples. Tests verify storage roundtrips, same-run collisions
and concurrent publication, corruption/truncation, permissions, failing/aborted
runs, unsupported JSON, stable retained nested observations, preserved assertion
objects, two-domain adversarial outcomes and changed-scenario acceptance refusal.
Evidence is immutable by the store API, not tamper-proof: a party able to rewrite
both receipt and checksum can forge it. No automatic credential redaction is
promised; adapters remain responsible for retention-safe observations. Complete
receipts cover normal and guard-aborted returns, not unexpected exceptions or
process crashes. Historical packs lacking scenario revisions need rerunning for
automatic acceptance. Existing partial scenarios remain supported; revisions do
not manufacture missing assertions or evaluate unseen state.
## Readiness verdict
**Useful local research framework; not ready for autonomous real-system assurance.**
The central safety model and local maturation example exist. Durable evidence
and bounded variants improve reuse, but do not establish live-model economics,
independent production provenance, browser coverage or concurrent correctness.
The priority after local delivery is the approved independent pilot, together
with the already recorded experiment choices—not rebuilding retired concepts.
## Validation and work status
The full suite passed **428 tests in 169.03 seconds**, including 30 new delivery
regressions. An earlier existing-boundary subset passed 161 tests. Both documented
examples executed successfully; local relative links and `git diff --check` pass.
These counts describe regression coverage, not independent production trials.
TD-WP-0004-T01–T04 are complete. T05 is waiting and flagged for human input, so
TD-WP-0004 remains **blocked**. TD-WP-0003-T01/T06/T07 remain the existing external
experiment records. The scope update and local implementation do not close those
validation gaps.
Decision: `f3b35dce-6047-4a95-8083-4d7032d4e99f`.