Align scope with intent and add durable evidence and bounded variants
Assistant: codex Assistant-Model: gpt-6-astra Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
This commit is contained in:
parent
8822480b6e
commit
eaf5d348e4
14 changed files with 767 additions and 92 deletions
107
history/2026-09-28-121933-scope-intent-assessment.md
Normal file
107
history/2026-09-28-121933-scope-intent-assessment.md
Normal file
|
|
@ -0,0 +1,107 @@
|
|||
# Scope against intent — 2026-09-28 12:19:33 UTC
|
||||
|
||||
## Assessment basis
|
||||
|
||||
Reviewed the repository at `e419bfe` and the delivery in TD-WP-0004. This report
|
||||
separates the baseline gaps from work implemented during this assessment. Sources:
|
||||
[INTENT](../INTENT.md), [scope](../SCOPE.md), [readiness settlement](../docs/TestDriverGeneralisationReview.md),
|
||||
[source](../src/testdriver/), [tests](../tests/),
|
||||
[38-mutation ground truth](../lab/GROUND-TRUTH.md),
|
||||
[E-001 results](../research/evidence/2026-09-28-e001.json),
|
||||
and the [audit-core contract](../usecases/audit_core_e2_tenant_boundary.py).
|
||||
No new production or paid-model experiment was performed.
|
||||
|
||||
The old SCOPE.md materially understated implementation ("implementation has not
|
||||
started") while overstating the stack and lifecycle model. The repo already had
|
||||
398 passing tests, three synthetic multi-user domains, qualified crystallization
|
||||
and extensive acceptance guards. It had no browser engine or database, and the
|
||||
Temperature/Energy/Campaign family had already been removed. SCOPE.md now describes
|
||||
implemented behavior and its limits rather than restating the thesis.
|
||||
|
||||
## Intent coverage
|
||||
|
||||
| INTENT area | Demonstrated before this work | Gap / disposition |
|
||||
|---|---|---|
|
||||
| Independent intent, claims and oracles | Frozen Python assertions; provenance admission; S1/S2/S3 separation; missing/invalid results inconclusive; definition revisions | Strong local coverage, but provenance labels and matching collector names are not external attestation. Domain adapters and independent authors remain necessary. |
|
||||
| Fluid-to-deterministic verification | Heuristic HTML discovery, stable-trajectory checks, frozen replay and one generated single-action test | Narrow demonstration, no live model, no general multi-step generator, no automatic full lifecycle. Keep claims bounded. |
|
||||
| Multi-user isolation | Separate actor objects/sessions; store/canary guards; tenant and delegation fixtures | No process security boundary, actual concurrent scheduler or causal/time-window enforcement. |
|
||||
| Security via use-case mutation | 38 seeded SUT mutations and hand-authored adversarial cases | No reusable scenario-side derivation at baseline. T03 delivers explicit actor/argument substitutions; skip/reorder/replay/concurrency and inferred claims are not added. |
|
||||
| Durable evidence and lineage | In-memory JSON-serializable evidence, experimental JSON files, asset and parent fields | Generic runs lacked a retained receipt API and asset lineage in the pack. T02 delivers opt-in local receipts and reload, preserving existing judgment behavior. |
|
||||
| Stable intent during adaptation | Claim/use-case revisions and recorded step coverage | Actor/action argument/permission/postcondition changes lacked a common intent revision. T03 adds scenario revisions and rejects changed intent for automatic acceptance/freezing. |
|
||||
| Real integration / end-to-end | Local stdlib HTTP journey, direct domain adapters; audit-core contract with fixture tests | No independently validated real-system adapter, custody/cleanup execution or operational observation cost evidence. T05 explicitly waits for an approved pilot. |
|
||||
| Resilience and security concurrency | Sequential mutations, expected refusals and missing-evidence checks | Not evidence of races, time-window handling or fault-tolerant external execution. Follow-on design depends on a real pilot's needs. |
|
||||
| Lineage and finding improvement loop | Research findings/decisions, parent ids, generated lineage comments | No automated finding-to-asset ingestion or asset registry. T02/T03 retain basic lineage, not an orchestration service. |
|
||||
| Energy/Temperature/Campaign/Retirement | Removed after generalisation found no decision-making use | Intent's superseding note governs. Deliberate de-scope, not a backlog to resurrect. |
|
||||
| Scale / distributed execution | None | Explicit initial non-goals, not missing milestone work. |
|
||||
|
||||
## Most relevant gaps, ranked
|
||||
|
||||
1. **Independent real-system validation.** This most limits confidence in the
|
||||
central thesis: synthetic labs do not measure the cost of independent
|
||||
observation, real requirements provenance, ambiguity or false-adaptation harm.
|
||||
Owner: **TD-WP-0004-T05**, waiting for an operator/target owner to select and
|
||||
approve a bounded target, observation path, fixtures and cleanup/custody plan.
|
||||
Implementing the existing audit-core contract without those inputs would not
|
||||
produce valid evidence. This task remains live and the workplan stays blocked.
|
||||
2. **Durable, reviewable run evidence and lineage.** High immediate value, small
|
||||
infrastructure cost, useful to every future adapter. **Delivered in T02**:
|
||||
versioned local receipts, checksum verification, no-overwrite atomic publish,
|
||||
strict JSON and recorded asset ancestry/variant/final verdict. Opt-in retention
|
||||
avoids silently persisting arbitrary adapter data. No signing/redaction claim.
|
||||
3. **Reusable adversarial variants with unchanged semantic expectations.** This
|
||||
is a direct part of INTENT's differentiation. **Delivered in T03**: explicit
|
||||
actor/argument substitution in two synthetic domains, preserved claim identity,
|
||||
descendant lineage and scenario-intent admission binding. It produces a test
|
||||
variant, never a new inferred oracle; an expected denial still needs independently
|
||||
authored expectations. A changed scenario requires review even if it passes.
|
||||
4. **Measured model/browser/authoring value.** Existing records remain canonical:
|
||||
**TD-WP-0003-T01** needs model, run count, spend ceiling and suite policy;
|
||||
**T06** needs browser setup and T01; **T07** needs a prospectively timed fresh
|
||||
independent author. This request does not supply those missing inputs. No
|
||||
duplicate records or fabricated economic/authoring measurements were created.
|
||||
5. **Broader scheduling and crystallization.** Actual concurrency/resilience and
|
||||
general multi-step code generation could extend the thesis, but should follow
|
||||
a pilot that demonstrates the requirement. They are explicitly outside this
|
||||
workplan's implementation commitment, not silently declared complete. T05's
|
||||
readiness review must decide which capability to fund before claiming the
|
||||
broader INTENT success criterion.
|
||||
|
||||
## Delivered behavior and assurance limits
|
||||
|
||||
[TD-WP-0004](../workplans/TD-WP-0004-scope-evidence-and-variants.md) was registered
|
||||
through `statehub fix-consistency`. [Evidence/variant usage](../docs/TestDriverEvidenceAndVariants.md)
|
||||
provides reproducible examples. Tests verify storage roundtrips, same-run collisions
|
||||
and concurrent publication, corruption/truncation, permissions, failing/aborted
|
||||
runs, unsupported JSON, stable retained nested observations, preserved assertion
|
||||
objects, two-domain adversarial outcomes and changed-scenario acceptance refusal.
|
||||
|
||||
Evidence is immutable by the store API, not tamper-proof: a party able to rewrite
|
||||
both receipt and checksum can forge it. No automatic credential redaction is
|
||||
promised; adapters remain responsible for retention-safe observations. Complete
|
||||
receipts cover normal and guard-aborted returns, not unexpected exceptions or
|
||||
process crashes. Historical packs lacking scenario revisions need rerunning for
|
||||
automatic acceptance. Existing partial scenarios remain supported; revisions do
|
||||
not manufacture missing assertions or evaluate unseen state.
|
||||
|
||||
## Readiness verdict
|
||||
|
||||
**Useful local research framework; not ready for autonomous real-system assurance.**
|
||||
The central safety model and local maturation example exist. Durable evidence
|
||||
and bounded variants improve reuse, but do not establish live-model economics,
|
||||
independent production provenance, browser coverage or concurrent correctness.
|
||||
The priority after local delivery is the approved independent pilot, together
|
||||
with the already recorded experiment choices—not rebuilding retired concepts.
|
||||
|
||||
## Validation and work status
|
||||
|
||||
The full suite passed **428 tests in 169.03 seconds**, including 30 new delivery
|
||||
regressions. An earlier existing-boundary subset passed 161 tests. Both documented
|
||||
examples executed successfully; local relative links and `git diff --check` pass.
|
||||
These counts describe regression coverage, not independent production trials.
|
||||
|
||||
TD-WP-0004-T01–T04 are complete. T05 is waiting and flagged for human input, so
|
||||
TD-WP-0004 remains **blocked**. TD-WP-0003-T01/T06/T07 remain the existing external
|
||||
experiment records. The scope update and local implementation do not close those
|
||||
validation gaps.
|
||||
|
||||
Decision: `f3b35dce-6047-4a95-8083-4d7032d4e99f`.
|
||||
Loading…
Add table
Add a link
Reference in a new issue