108 lines
8.3 KiB
Markdown
108 lines
8.3 KiB
Markdown
|
|
# Scope against intent — 2026-09-28 12:19:33 UTC
|
|||
|
|
|
|||
|
|
## Assessment basis
|
|||
|
|
|
|||
|
|
Reviewed the repository at `e419bfe` and the delivery in TD-WP-0004. This report
|
|||
|
|
separates the baseline gaps from work implemented during this assessment. Sources:
|
|||
|
|
[INTENT](../INTENT.md), [scope](../SCOPE.md), [readiness settlement](../docs/TestDriverGeneralisationReview.md),
|
|||
|
|
[source](../src/testdriver/), [tests](../tests/),
|
|||
|
|
[38-mutation ground truth](../lab/GROUND-TRUTH.md),
|
|||
|
|
[E-001 results](../research/evidence/2026-09-28-e001.json),
|
|||
|
|
and the [audit-core contract](../usecases/audit_core_e2_tenant_boundary.py).
|
|||
|
|
No new production or paid-model experiment was performed.
|
|||
|
|
|
|||
|
|
The old SCOPE.md materially understated implementation ("implementation has not
|
|||
|
|
started") while overstating the stack and lifecycle model. The repo already had
|
|||
|
|
398 passing tests, three synthetic multi-user domains, qualified crystallization
|
|||
|
|
and extensive acceptance guards. It had no browser engine or database, and the
|
|||
|
|
Temperature/Energy/Campaign family had already been removed. SCOPE.md now describes
|
|||
|
|
implemented behavior and its limits rather than restating the thesis.
|
|||
|
|
|
|||
|
|
## Intent coverage
|
|||
|
|
|
|||
|
|
| INTENT area | Demonstrated before this work | Gap / disposition |
|
|||
|
|
|---|---|---|
|
|||
|
|
| Independent intent, claims and oracles | Frozen Python assertions; provenance admission; S1/S2/S3 separation; missing/invalid results inconclusive; definition revisions | Strong local coverage, but provenance labels and matching collector names are not external attestation. Domain adapters and independent authors remain necessary. |
|
|||
|
|
| Fluid-to-deterministic verification | Heuristic HTML discovery, stable-trajectory checks, frozen replay and one generated single-action test | Narrow demonstration, no live model, no general multi-step generator, no automatic full lifecycle. Keep claims bounded. |
|
|||
|
|
| Multi-user isolation | Separate actor objects/sessions; store/canary guards; tenant and delegation fixtures | No process security boundary, actual concurrent scheduler or causal/time-window enforcement. |
|
|||
|
|
| Security via use-case mutation | 38 seeded SUT mutations and hand-authored adversarial cases | No reusable scenario-side derivation at baseline. T03 delivers explicit actor/argument substitutions; skip/reorder/replay/concurrency and inferred claims are not added. |
|
|||
|
|
| Durable evidence and lineage | In-memory JSON-serializable evidence, experimental JSON files, asset and parent fields | Generic runs lacked a retained receipt API and asset lineage in the pack. T02 delivers opt-in local receipts and reload, preserving existing judgment behavior. |
|
|||
|
|
| Stable intent during adaptation | Claim/use-case revisions and recorded step coverage | Actor/action argument/permission/postcondition changes lacked a common intent revision. T03 adds scenario revisions and rejects changed intent for automatic acceptance/freezing. |
|
|||
|
|
| Real integration / end-to-end | Local stdlib HTTP journey, direct domain adapters; audit-core contract with fixture tests | No independently validated real-system adapter, custody/cleanup execution or operational observation cost evidence. T05 explicitly waits for an approved pilot. |
|
|||
|
|
| Resilience and security concurrency | Sequential mutations, expected refusals and missing-evidence checks | Not evidence of races, time-window handling or fault-tolerant external execution. Follow-on design depends on a real pilot's needs. |
|
|||
|
|
| Lineage and finding improvement loop | Research findings/decisions, parent ids, generated lineage comments | No automated finding-to-asset ingestion or asset registry. T02/T03 retain basic lineage, not an orchestration service. |
|
|||
|
|
| Energy/Temperature/Campaign/Retirement | Removed after generalisation found no decision-making use | Intent's superseding note governs. Deliberate de-scope, not a backlog to resurrect. |
|
|||
|
|
| Scale / distributed execution | None | Explicit initial non-goals, not missing milestone work. |
|
|||
|
|
|
|||
|
|
## Most relevant gaps, ranked
|
|||
|
|
|
|||
|
|
1. **Independent real-system validation.** This most limits confidence in the
|
|||
|
|
central thesis: synthetic labs do not measure the cost of independent
|
|||
|
|
observation, real requirements provenance, ambiguity or false-adaptation harm.
|
|||
|
|
Owner: **TD-WP-0004-T05**, waiting for an operator/target owner to select and
|
|||
|
|
approve a bounded target, observation path, fixtures and cleanup/custody plan.
|
|||
|
|
Implementing the existing audit-core contract without those inputs would not
|
|||
|
|
produce valid evidence. This task remains live and the workplan stays blocked.
|
|||
|
|
2. **Durable, reviewable run evidence and lineage.** High immediate value, small
|
|||
|
|
infrastructure cost, useful to every future adapter. **Delivered in T02**:
|
|||
|
|
versioned local receipts, checksum verification, no-overwrite atomic publish,
|
|||
|
|
strict JSON and recorded asset ancestry/variant/final verdict. Opt-in retention
|
|||
|
|
avoids silently persisting arbitrary adapter data. No signing/redaction claim.
|
|||
|
|
3. **Reusable adversarial variants with unchanged semantic expectations.** This
|
|||
|
|
is a direct part of INTENT's differentiation. **Delivered in T03**: explicit
|
|||
|
|
actor/argument substitution in two synthetic domains, preserved claim identity,
|
|||
|
|
descendant lineage and scenario-intent admission binding. It produces a test
|
|||
|
|
variant, never a new inferred oracle; an expected denial still needs independently
|
|||
|
|
authored expectations. A changed scenario requires review even if it passes.
|
|||
|
|
4. **Measured model/browser/authoring value.** Existing records remain canonical:
|
|||
|
|
**TD-WP-0003-T01** needs model, run count, spend ceiling and suite policy;
|
|||
|
|
**T06** needs browser setup and T01; **T07** needs a prospectively timed fresh
|
|||
|
|
independent author. This request does not supply those missing inputs. No
|
|||
|
|
duplicate records or fabricated economic/authoring measurements were created.
|
|||
|
|
5. **Broader scheduling and crystallization.** Actual concurrency/resilience and
|
|||
|
|
general multi-step code generation could extend the thesis, but should follow
|
|||
|
|
a pilot that demonstrates the requirement. They are explicitly outside this
|
|||
|
|
workplan's implementation commitment, not silently declared complete. T05's
|
|||
|
|
readiness review must decide which capability to fund before claiming the
|
|||
|
|
broader INTENT success criterion.
|
|||
|
|
|
|||
|
|
## Delivered behavior and assurance limits
|
|||
|
|
|
|||
|
|
[TD-WP-0004](../workplans/TD-WP-0004-scope-evidence-and-variants.md) was registered
|
|||
|
|
through `statehub fix-consistency`. [Evidence/variant usage](../docs/TestDriverEvidenceAndVariants.md)
|
|||
|
|
provides reproducible examples. Tests verify storage roundtrips, same-run collisions
|
|||
|
|
and concurrent publication, corruption/truncation, permissions, failing/aborted
|
|||
|
|
runs, unsupported JSON, stable retained nested observations, preserved assertion
|
|||
|
|
objects, two-domain adversarial outcomes and changed-scenario acceptance refusal.
|
|||
|
|
|
|||
|
|
Evidence is immutable by the store API, not tamper-proof: a party able to rewrite
|
|||
|
|
both receipt and checksum can forge it. No automatic credential redaction is
|
|||
|
|
promised; adapters remain responsible for retention-safe observations. Complete
|
|||
|
|
receipts cover normal and guard-aborted returns, not unexpected exceptions or
|
|||
|
|
process crashes. Historical packs lacking scenario revisions need rerunning for
|
|||
|
|
automatic acceptance. Existing partial scenarios remain supported; revisions do
|
|||
|
|
not manufacture missing assertions or evaluate unseen state.
|
|||
|
|
|
|||
|
|
## Readiness verdict
|
|||
|
|
|
|||
|
|
**Useful local research framework; not ready for autonomous real-system assurance.**
|
|||
|
|
The central safety model and local maturation example exist. Durable evidence
|
|||
|
|
and bounded variants improve reuse, but do not establish live-model economics,
|
|||
|
|
independent production provenance, browser coverage or concurrent correctness.
|
|||
|
|
The priority after local delivery is the approved independent pilot, together
|
|||
|
|
with the already recorded experiment choices—not rebuilding retired concepts.
|
|||
|
|
|
|||
|
|
## Validation and work status
|
|||
|
|
|
|||
|
|
The full suite passed **428 tests in 169.03 seconds**, including 30 new delivery
|
|||
|
|
regressions. An earlier existing-boundary subset passed 161 tests. Both documented
|
|||
|
|
examples executed successfully; local relative links and `git diff --check` pass.
|
|||
|
|
These counts describe regression coverage, not independent production trials.
|
|||
|
|
|
|||
|
|
TD-WP-0004-T01–T04 are complete. T05 is waiting and flagged for human input, so
|
|||
|
|
TD-WP-0004 remains **blocked**. TD-WP-0003-T01/T06/T07 remain the existing external
|
|||
|
|
experiment records. The scope update and local implementation do not close those
|
|||
|
|
validation gaps.
|
|||
|
|
|
|||
|
|
Decision: `f3b35dce-6047-4a95-8083-4d7032d4e99f`.
|