# TD-WP-0003 generalisation and settlement — 2026-09-28 The kernel now expresses sharing/revocation, delegated approval and tenant lifecycle without additional framework concepts. This is evidence of local expressiveness, not readiness for autonomous verification of a real system. ## New use cases (T02) The agent wrote [synthetic contracts](../usecases/generalisation-spec.md) before the implementations. Claims are `agent-from-spec`, not human-authored. The same agent wrote the synthetic requirements and labs; this does not establish independence from a real product team or requirements quality. - `scenarios/delegated_approval.py`: premature execution, submission, delegation, rejected self-approval, rejected former-reviewer approval, delegate approval, requester execution. Three independently enabled defects must fail. - `scenarios/tenant_lifecycle.py`: two administrators use the same local id; foreign delete attempts fail; deletion and recreation preserve the other tenant. Four independently enabled defects must fail. `tests/test_generalisation.py` verifies baselines, each defect, missing-evidence INCONCLUSIVE, deterministic replay and verdict reproduction from serialized S3. Separate observers read state and probe enforcement; drivers only emit S1. No claims or invariants are learned from the lab's responses. Ordered `Step`s already express the required Schedule. `Scenario.variant` identifies each lab fault. A separate scheduler, Lens or Variant object was not needed. There is no concurrent scheduling or time-window evidence here. The resistance was in the adapters, not new concepts: `StateObserver`'s standard snapshot is specific to resource grants, so each domain needs its own collector implementing `snapshot()` and `name`. Expected refusal must be a completed protocol response with an independently observed receipt; setting `Realization.raised` would suppress dependent claims as INCONCLUSIVE. Unexpected exceptions still abort these fixture drivers; they are not silently swallowed. ## Structural experiment (T03) [Machine-readable results](../research/evidence/2026-09-28-e001.json) retain each arm's surface, judgments, metrics and the journey classification/signals. Reproduce with: ```bash PYTHONPATH=src:. python3 tools/measure_e001.py /tmp/e001.json ``` | Identifier axis | Discovery | Recorded selectors | |---|---:|---:| | Preserved | 17/17 | 17/17 | | Dropped | 9/10 | 0/10 | M25–M31 add reordered fields, dialog relocation, anchor affordances, fieldset grouping, textarea input, an unrelated preceding form, and sidebar relocation. M32–M38 pair those mechanisms with preserved identifiers. Existing M02/M21/M22 supply rewritten DOM, reworded controls and renamed fields. New labels were fixed from the synthetic contract before running and are agent-authored, not claimed as newly human-reviewed ground truth. The 38-entry catalogue contains 27 mechanical, four semantic and seven defect mutations. Mechanical acceptance is 26/27. M22 still defeats the heuristic. False adaptation is 0/7 on the existing defect set; adding mechanical variants does not enlarge that safety denominator. No claim/invariant index changed. M25 and M32 are UNCHANGED because the coarse surface fingerprint ignores field order and test ids; this is not a claim that their HTML is identical. These are selected mechanisms, not random independent samples. Paired variants are correlated; 9/10 is not a population reliability estimate. The control is a stable-test-id replay, not every possible conventional selector strategy. Both arms remain token-free and test structural HTML only. H-001 stays narrowed to identifier-loss cases, and F-0004/F-0005/F-0007's external validity limits remain. ## Concept settlement and compatibility (T04/T05) **Keep claims as Python predicates.** Across all three runnable use cases the repeated shapes are equality, membership and comparisons between enforcement and state. The audit-core consumer additionally compares sanitized absence surfaces, bounded execution, cleanup and receipt binding. A small vocabulary would either duplicate Python or need frequent escape hatches; serialization would create a second language and a second statement of intent. This publishes the interface decision requested by T05: the existing `Claim(id, text, provenance, predicate, after_step, source_ref=None)` and `Invariant(id, text, provenance, predicate, source_ref=None)` constructors stay compatible. Predicates receive independent observation mappings and return booleans. Required evidence must use explicit lookup, not a pass-producing fallback. The oracle converts missing keys or predicate exceptions to INCONCLUSIVE. Predicates must not reach into actors or the live SUT. Generated regression artifacts continue to import their original scenario claims; consumers must ship those modules with the artifact. Fully standalone claim serialization is deliberately rejected. The audit-core consumer needs no migration; its tests remain in the default suite. **Remove unused lifecycle concepts now.** Temperature, Energy, Confidence, Campaign, Metabolism and automated Retirement have informed no decision in these three use cases. They leave the current model, not another deferred gate. Trajectory stability continues to drive crystallization. H-005 becomes `DORMANT-INDEFINITE`; F-0008 is resolved. This is a research-prototype compatibility change: `testdriver.energy`, public `EnergyEvent`/`EnergyEventType` exports and the `EvidencePack.energy_events` field are removed. Consumers must stop importing or requiring them. Historical JSON may retain that field; the current classifier already ignores it. Execution identity, timestamps, judgments, S1 realization costs and lineage remain; none of these imply a computed lifecycle score. There is no dependency addition. Historical concept proposals carry a supersession notice. ## Authoring measurement (T07 remains waiting) [Timing receipt](../research/evidence/2026-09-28-authoring.json) preserves the actual measurements and line counts. The timer was placed inside the shell write operation, after code composition. The approval interval also included composition of the next case, and the tenant interval largely measured file writes and pytest. These intervals are **not valid authoring-cost measurements** and must not be compared as productivity numbers. Human maintenance effort and model latency/cost were not captured. The first attempt cannot be reconstructed. The observable concepts consulted and friction are recorded above. T07 retains the remaining work: an independent author must implement these contracts afresh with a timer started before reading/design, stopping after first passing defect checks, separately recording tooling wait and review. Do not count a rerun or copy of these implementations as a new authoring sample. No replacement task or workplan is created. ## Real-system readiness (T08) **Verdict: not ready for autonomous real-system application.** Local research and attended review of prepared claims can continue. Completing the review does not satisfy the workplan's live-model success gate or authorize a production run. | Required criterion | Current evidence | Assessment | |---|---|---| | Independent observation channel can read records and probe enforcement without trusting actors | Three labs; audit-core declares an independent observer but has no runnable driver/collector here | Not met outside synthetic labs; adapter construction, permissions and operating cost remain unmeasured | | Claims trace to independently approved intent predating the tested behavior | Local provenance guards; audit-core cites approved engagement and workplan, while its precedent receipt is calibration only | Structurally supported, operational chain must be checked for the exact next engagement; a ticket written after code is insufficient without human promotion | | Legitimate behavioral ambiguity is escalated without rewriting intent | M12/M19 are behaviorally indistinguishable; AMBIGUOUS and BEHAVIOUR_CHANGED require review | Demonstrated in lab, real backlog dispute resolution is still attended; missing evidence stays INCONCLUSIVE | | False-adaptation consequences are identified for the engagement | Audit-core tenant disclosure, forged foreign append or leftover identity could be concealed by a false pass | Zero tolerance for automatic acceptance of regression; 0/7 synthetic defects is no guarantee against unseen attacks | | Execution, custody, timing, cleanup and target revision are enforced | Audit-core's PhaseContract is durable intent, explicitly not a runnable Scenario; tests use supplied observations | Not met: no causal/time-window runtime, custody or cluster driver, independent cleanup collector or fresh admitted window here | | Capability, economics and maintenance costs justify adoption | Token-free structural discovery, deterministic crystallization and invalid first authoring timer | Not met; T01/T06/T07 retain the unanswered measurements | The audit-core material referenced here is the checked-in `usecases/audit_core_e2_tenant_boundary.py` contract and its fixture-based tests, not a fresh inspection of any deployment or a revalidation of the 2026-08-22 engagement. A real run needs the exact engagement's target-owner acceptance, independent observation/cleanup receipts and appropriate approved access. A successful historical receipt cannot substitute for those. The existing T01, T06 and T07 tasks retain the experiment blockers; this review creates no production implementation commitment or new work record. ## Safety follow-up: aborted runs and incomplete evidence A post-closeout review reproduced two false-pass paths that the original suite missed. A forbidden-surface exception after a passing prefix returned PASS even though later claims had never been judged. Separately, deleting almost every judgment and all state snapshots still yielded UNCHANGED/safe-to-accept. The runner now records `scheduled_steps` and `expected_judgments` from the scenario before executing any action. Only claims attached to scheduled steps are included; the existing one-step browser scenario remains intentionally narrow. On a surface violation, every remaining scheduled claim and invariant gets an explicit INCONCLUSIVE judgment and no fabricated observation. A prior FAIL still dominates. An abort after the last assertion is also INCONCLUSIVE, because completing the assertions does not mean the scheduled run completed. The classifier validates both packs against their input manifests. It requires exactly one judgment for every expected assertion/step pair and exactly one S1 realization, S2 realization check and nonempty S3 state snapshot for every scheduled step, with the correct strata. Missing, duplicate or invalid verdict records, unfinished runs, and missing identity/version fields prevent automatic acceptance. The candidate must retain the baseline's step and judgment coverage. Evidence failure returns AMBIGUOUS before any adaptation rule is considered. This adds two serialized EvidencePack fields. Historical packs without these manifests remain historical artifacts but classify as AMBIGUOUS; rerun the scenario to obtain evidence eligible for automatic acceptance. Do not infer a manifest from surviving observations, since that would preserve the original hole. These checks establish structural completeness, not authenticity of a pack or correctness of arbitrary caller-supplied predicates. The existing actor/observer separation and deterministic oracle remain required. Regression coverage in `tests/test_evidence_completeness.py` includes aborts at each step, abort after the final assertion, preservation of an earlier failure, individual missing strata/records, duplicate records, and identical truncation of both packs. Complete ordinary runs and intentionally narrow scenarios retain their existing behavior. The two defects are handled under the existing T08 readiness review; the external T01/T06/T07 blockers remain unchanged. Validation after the safety fixes: **299 tests passed** (`python3 -m pytest -q`). The new regression module had 55 failures and five passing controls before the fix; `git diff --check` passes after it. ## Admission follow-up: crystallization and intent revisions The next boundary review reproduced three additional holes: repeated aborted runs qualified as stable; a complete failing run compared with itself was UNCHANGED/safe-to-accept; and changing a claim's text or predicate while retaining its ID/provenance was invisible to the classifier. These fixes remain under T08. Crystallization now uses the classifier's acceptance gate for every run against the first run. All packs must have complete evidence, passing judgments, verified postconditions, unchanged intent and successful realizations. Scenario/use-case identity, scheduled steps and expected judgments must match across the window. Repeated copies of one run receipt do not satisfy the independent-run count. Only then does trajectory comparison decide stability. A refused window returns no trajectories. An unchanged product failure is not a stable success. The classifier requires a passing baseline. If the baseline contains FAIL or INCONCLUSIVE, even a repaired candidate requires a new passing reference rather than automatic acceptance against the failed baseline. Candidate failures against a passing reference still report BEHAVIOUR_CHANGED. Baseline realization must also be established before either automatic-acceptance outcome is returned; missing or non-boolean postconditions cannot be treated as success. Evidence packs now include `intent_revisions`: SHA-256 revisions recorded before execution for the use case and each claim/invariant, including claims outside a partial scenario's schedule. They include narrative or assertion text, source, provenance, claim scheduling, and predicate implementation. The predicate hash covers bytecode/constants, defaults, keyword defaults, closure values, referenced Python helper functions and their referenced global values. It is bound to the Python implementation/version; checkout filenames and line numbers are excluded. Only hashes enter evidence, not captured values or code. Changed definitions produce INTENT_CHANGED, which requests review and does not assert human approval. An internally complete revised claim schedule is an intent change; missing baseline coverage under unchanged intent remains incomplete evidence. No Claim/Invariant constructor change or claim serialization language is added. `revisions.py` is input-change detection for pure Python predicates. It does not prove semantic equivalence, sandbox Python, attest a caller-supplied pack or make mutable runtime behavior safe. Predicate dependencies must stay stable during a run. Cyclic/opaque callables, module dependencies and dynamic builtin lookups cannot be fingerprinted reliably and receive a null revision. They may still be evaluated by the oracle, but their evidence cannot be automatically accepted or crystallized. Refactor such predicates to consume plain independent snapshots. Changing Python versions conservatively requires intent review and fresh evidence. Older packs without usable intent revisions require a rerun, just as packs without the earlier completeness manifest do. Existing historical experiment receipts are not retroactively upgraded or rewritten. Tests in `tests/test_acceptance_boundaries.py` cover the reproduced cases, helper/global and captured-value changes, invariant/narrative/source/scheduling changes, cross-process identity, source relocation, redaction of captured values, malformed postconditions, and the complete-passing control. No live-model, browser-engine or production readiness blocker is removed by this work. Validation for this follow-up: initial reproduction had 20 failures and two passing controls; full suite **327 passed**. The final dynamic-dependency and strict-postcondition additions passed **32 boundary checks**, including four new cases. These counts overlap; they are not separate independent samples. ## Isolation and repeated-action replay follow-up (2026-09-28) Recorded actor isolation violations now stop the runner before further actions or judgments. Unreached scheduled assertions become INCONCLUSIVE; prior observed product failures remain FAIL. The framework finding remains separate S3 evidence, so it does not add or rewrite product claims. Isolation is examined at entry and after each action. Classification requires all these checks to be present, S3, and free of violations; crystallization inherits that gate. Older packs with only an entry check need rerunning before automatic acceptance. Canary checks remain a diagnostic at those boundaries, not proof against arbitrary memory access or transient leaks inside a driver. Runner dispatch now passes step identity to drivers that expose `realize_step`. Composite drivers forward it while ordinary action-only drivers keep their existing interface. Frozen replay selects by step and verifies the action name; repeated names no longer overwrite different targets. Unknown/mismatched steps, duplicate frozen step ids and ambiguous action-only calls cannot send a request. Unambiguous action-only calls remain supported. Replay has no advancing cursor, so the same driver can run the schedule repeatedly. Regression coverage includes entry/mid-run/final-action leaks, otherwise passing packs with invalid isolation evidence, distinct targets for repeated actions through direct and composite dispatch, driver reuse and invalid frozen lookup. Decision: `e8fabf6e-3e5e-424e-9a60-152bf3dec641`. Validation: the complete suite passed **343 tests** (153.67 seconds). ## Predicate and isolation guard integrity (2026-09-28) Oracle predicates must return actual booleans. None, numbers, strings, containers and objects now produce INCONCLUSIVE; their truth-value methods are never invoked. Existing True/False results retain PASS/FAIL semantics. Predicate results are not copied into the diagnostic, avoiding disclosure of their values. At each existing isolation boundary, the runner now checks for shared memory stores, missing or changed stored canaries, and invalid or duplicate actor canaries. Removing markers can no longer make shared actor memory appear clean. These findings use the existing abort/acceptance/crystallization gates and record actor identities without marker or private memory values. This remains a boundary diagnostic, not a sandbox against arbitrary or transient memory access. Twenty regressions cover claims and invariants, non-coercible objects, real boolean controls, invalid-result admission and crystallization, plus broken isolation guards before and during runs. Sixteen failed before implementation; the focused suite passed 69 tests afterward. Decision: `2ecac1c5-5254-4ce9-b092-b47c3e1283a9`. Full-suite validation: **363 tests passed** in 153.10 seconds. ## Authenticated HTTP and surface boundaries (2026-09-28) Authenticated HTTP now stays on the configured scheme, host and effective port. Absolute and protocol-relative targets are validated before adding credentials; redirects are validated before forwarding a request. Userinfo and unsupported schemes are rejected. Same-origin relative/absolute URLs and redirects continue to work. Browser sessions turn origin violations into existing surface findings. The generator embeds the same stdlib helper definitions in standalone descendants; the checked-in descendant is updated too. No runtime framework dependency is introduced into generated tests, and no package dependency is added. Runner also validates the realization's reported surface against the original scheduled action, so a driver that forgets its check cannot certify forbidden surface substitution. This is a post-call guard: drivers must still check before acting, and an actual side effect cannot be undone by the runner. It does not attest a dishonest driver's report. Origin restrictions likewise do not claim path-level authorization or protect against an already compromised allowed host. Tests use synthetic bearer credentials and local HTTP servers to verify that rejected targets/redirects receive no requests, across sessions, freshly generated modules and the checked-in descendant. Coverage includes five redirect codes, protocol-relative targets, scheme/userinfo rejection, normal origin equivalence, same-origin redirect chains and driver omission blocking acceptance/freezing. The focused subset passed **107 tests**. Decision: `88490ee8-839d-40dc-affa-b00a1c68e3c5`. Full-suite validation: **398 tests passed** in 165.81 seconds.