Assistant: codex Assistant-Model: gpt-6-astra Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
107 lines
8.3 KiB
Markdown
107 lines
8.3 KiB
Markdown
# Scope against intent — 2026-09-28 12:19:33 UTC
|
||
|
||
## Assessment basis
|
||
|
||
Reviewed the repository at `e419bfe` and the delivery in TD-WP-0004. This report
|
||
separates the baseline gaps from work implemented during this assessment. Sources:
|
||
[INTENT](../INTENT.md), [scope](../SCOPE.md), [readiness settlement](../docs/TestDriverGeneralisationReview.md),
|
||
[source](../src/testdriver/), [tests](../tests/),
|
||
[38-mutation ground truth](../lab/GROUND-TRUTH.md),
|
||
[E-001 results](../research/evidence/2026-09-28-e001.json),
|
||
and the [audit-core contract](../usecases/audit_core_e2_tenant_boundary.py).
|
||
No new production or paid-model experiment was performed.
|
||
|
||
The old SCOPE.md materially understated implementation ("implementation has not
|
||
started") while overstating the stack and lifecycle model. The repo already had
|
||
398 passing tests, three synthetic multi-user domains, qualified crystallization
|
||
and extensive acceptance guards. It had no browser engine or database, and the
|
||
Temperature/Energy/Campaign family had already been removed. SCOPE.md now describes
|
||
implemented behavior and its limits rather than restating the thesis.
|
||
|
||
## Intent coverage
|
||
|
||
| INTENT area | Demonstrated before this work | Gap / disposition |
|
||
|---|---|---|
|
||
| Independent intent, claims and oracles | Frozen Python assertions; provenance admission; S1/S2/S3 separation; missing/invalid results inconclusive; definition revisions | Strong local coverage, but provenance labels and matching collector names are not external attestation. Domain adapters and independent authors remain necessary. |
|
||
| Fluid-to-deterministic verification | Heuristic HTML discovery, stable-trajectory checks, frozen replay and one generated single-action test | Narrow demonstration, no live model, no general multi-step generator, no automatic full lifecycle. Keep claims bounded. |
|
||
| Multi-user isolation | Separate actor objects/sessions; store/canary guards; tenant and delegation fixtures | No process security boundary, actual concurrent scheduler or causal/time-window enforcement. |
|
||
| Security via use-case mutation | 38 seeded SUT mutations and hand-authored adversarial cases | No reusable scenario-side derivation at baseline. T03 delivers explicit actor/argument substitutions; skip/reorder/replay/concurrency and inferred claims are not added. |
|
||
| Durable evidence and lineage | In-memory JSON-serializable evidence, experimental JSON files, asset and parent fields | Generic runs lacked a retained receipt API and asset lineage in the pack. T02 delivers opt-in local receipts and reload, preserving existing judgment behavior. |
|
||
| Stable intent during adaptation | Claim/use-case revisions and recorded step coverage | Actor/action argument/permission/postcondition changes lacked a common intent revision. T03 adds scenario revisions and rejects changed intent for automatic acceptance/freezing. |
|
||
| Real integration / end-to-end | Local stdlib HTTP journey, direct domain adapters; audit-core contract with fixture tests | No independently validated real-system adapter, custody/cleanup execution or operational observation cost evidence. T05 explicitly waits for an approved pilot. |
|
||
| Resilience and security concurrency | Sequential mutations, expected refusals and missing-evidence checks | Not evidence of races, time-window handling or fault-tolerant external execution. Follow-on design depends on a real pilot's needs. |
|
||
| Lineage and finding improvement loop | Research findings/decisions, parent ids, generated lineage comments | No automated finding-to-asset ingestion or asset registry. T02/T03 retain basic lineage, not an orchestration service. |
|
||
| Energy/Temperature/Campaign/Retirement | Removed after generalisation found no decision-making use | Intent's superseding note governs. Deliberate de-scope, not a backlog to resurrect. |
|
||
| Scale / distributed execution | None | Explicit initial non-goals, not missing milestone work. |
|
||
|
||
## Most relevant gaps, ranked
|
||
|
||
1. **Independent real-system validation.** This most limits confidence in the
|
||
central thesis: synthetic labs do not measure the cost of independent
|
||
observation, real requirements provenance, ambiguity or false-adaptation harm.
|
||
Owner: **TD-WP-0004-T05**, waiting for an operator/target owner to select and
|
||
approve a bounded target, observation path, fixtures and cleanup/custody plan.
|
||
Implementing the existing audit-core contract without those inputs would not
|
||
produce valid evidence. This task remains live and the workplan stays blocked.
|
||
2. **Durable, reviewable run evidence and lineage.** High immediate value, small
|
||
infrastructure cost, useful to every future adapter. **Delivered in T02**:
|
||
versioned local receipts, checksum verification, no-overwrite atomic publish,
|
||
strict JSON and recorded asset ancestry/variant/final verdict. Opt-in retention
|
||
avoids silently persisting arbitrary adapter data. No signing/redaction claim.
|
||
3. **Reusable adversarial variants with unchanged semantic expectations.** This
|
||
is a direct part of INTENT's differentiation. **Delivered in T03**: explicit
|
||
actor/argument substitution in two synthetic domains, preserved claim identity,
|
||
descendant lineage and scenario-intent admission binding. It produces a test
|
||
variant, never a new inferred oracle; an expected denial still needs independently
|
||
authored expectations. A changed scenario requires review even if it passes.
|
||
4. **Measured model/browser/authoring value.** Existing records remain canonical:
|
||
**TD-WP-0003-T01** needs model, run count, spend ceiling and suite policy;
|
||
**T06** needs browser setup and T01; **T07** needs a prospectively timed fresh
|
||
independent author. This request does not supply those missing inputs. No
|
||
duplicate records or fabricated economic/authoring measurements were created.
|
||
5. **Broader scheduling and crystallization.** Actual concurrency/resilience and
|
||
general multi-step code generation could extend the thesis, but should follow
|
||
a pilot that demonstrates the requirement. They are explicitly outside this
|
||
workplan's implementation commitment, not silently declared complete. T05's
|
||
readiness review must decide which capability to fund before claiming the
|
||
broader INTENT success criterion.
|
||
|
||
## Delivered behavior and assurance limits
|
||
|
||
[TD-WP-0004](../workplans/TD-WP-0004-scope-evidence-and-variants.md) was registered
|
||
through `statehub fix-consistency`. [Evidence/variant usage](../docs/TestDriverEvidenceAndVariants.md)
|
||
provides reproducible examples. Tests verify storage roundtrips, same-run collisions
|
||
and concurrent publication, corruption/truncation, permissions, failing/aborted
|
||
runs, unsupported JSON, stable retained nested observations, preserved assertion
|
||
objects, two-domain adversarial outcomes and changed-scenario acceptance refusal.
|
||
|
||
Evidence is immutable by the store API, not tamper-proof: a party able to rewrite
|
||
both receipt and checksum can forge it. No automatic credential redaction is
|
||
promised; adapters remain responsible for retention-safe observations. Complete
|
||
receipts cover normal and guard-aborted returns, not unexpected exceptions or
|
||
process crashes. Historical packs lacking scenario revisions need rerunning for
|
||
automatic acceptance. Existing partial scenarios remain supported; revisions do
|
||
not manufacture missing assertions or evaluate unseen state.
|
||
|
||
## Readiness verdict
|
||
|
||
**Useful local research framework; not ready for autonomous real-system assurance.**
|
||
The central safety model and local maturation example exist. Durable evidence
|
||
and bounded variants improve reuse, but do not establish live-model economics,
|
||
independent production provenance, browser coverage or concurrent correctness.
|
||
The priority after local delivery is the approved independent pilot, together
|
||
with the already recorded experiment choices—not rebuilding retired concepts.
|
||
|
||
## Validation and work status
|
||
|
||
The full suite passed **428 tests in 169.03 seconds**, including 30 new delivery
|
||
regressions. An earlier existing-boundary subset passed 161 tests. Both documented
|
||
examples executed successfully; local relative links and `git diff --check` pass.
|
||
These counts describe regression coverage, not independent production trials.
|
||
|
||
TD-WP-0004-T01–T04 are complete. T05 is waiting and flagged for human input, so
|
||
TD-WP-0004 remains **blocked**. TD-WP-0003-T01/T06/T07 remain the existing external
|
||
experiment records. The scope update and local implementation do not close those
|
||
validation gaps.
|
||
|
||
Decision: `f3b35dce-6047-4a95-8083-4d7032d4e99f`.
|