Assistant: codex Assistant-Model: gpt-6-astra Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
321 lines
20 KiB
Markdown
321 lines
20 KiB
Markdown
# TD-WP-0003 generalisation and settlement — 2026-09-28
|
||
|
||
The kernel now expresses sharing/revocation, delegated approval and tenant
|
||
lifecycle without additional framework concepts. This is evidence of local
|
||
expressiveness, not readiness for autonomous verification of a real system.
|
||
|
||
## New use cases (T02)
|
||
|
||
The agent wrote [synthetic contracts](../usecases/generalisation-spec.md) before
|
||
the implementations. Claims are `agent-from-spec`, not human-authored. The same
|
||
agent wrote the synthetic requirements and labs; this does not establish
|
||
independence from a real product team or requirements quality.
|
||
|
||
- `scenarios/delegated_approval.py`: premature execution, submission, delegation,
|
||
rejected self-approval, rejected former-reviewer approval, delegate approval,
|
||
requester execution. Three independently enabled defects must fail.
|
||
- `scenarios/tenant_lifecycle.py`: two administrators use the same local id;
|
||
foreign delete attempts fail; deletion and recreation preserve the other
|
||
tenant. Four independently enabled defects must fail.
|
||
|
||
`tests/test_generalisation.py` verifies baselines, each defect, missing-evidence
|
||
INCONCLUSIVE, deterministic replay and verdict reproduction from serialized S3.
|
||
Separate observers read state and probe enforcement; drivers only emit S1.
|
||
No claims or invariants are learned from the lab's responses.
|
||
|
||
Ordered `Step`s already express the required Schedule. `Scenario.variant`
|
||
identifies each lab fault. A separate scheduler, Lens or Variant object was not
|
||
needed. There is no concurrent scheduling or time-window evidence here.
|
||
|
||
The resistance was in the adapters, not new concepts: `StateObserver`'s standard
|
||
snapshot is specific to resource grants, so each domain needs its own collector
|
||
implementing `snapshot()` and `name`. Expected refusal must be a completed
|
||
protocol response with an independently observed receipt; setting
|
||
`Realization.raised` would suppress dependent claims as INCONCLUSIVE. Unexpected
|
||
exceptions still abort these fixture drivers; they are not silently swallowed.
|
||
|
||
## Structural experiment (T03)
|
||
|
||
[Machine-readable results](../research/evidence/2026-09-28-e001.json) retain each
|
||
arm's surface, judgments, metrics and the journey classification/signals.
|
||
Reproduce with:
|
||
|
||
```bash
|
||
PYTHONPATH=src:. python3 tools/measure_e001.py /tmp/e001.json
|
||
```
|
||
|
||
| Identifier axis | Discovery | Recorded selectors |
|
||
|---|---:|---:|
|
||
| Preserved | 17/17 | 17/17 |
|
||
| Dropped | 9/10 | 0/10 |
|
||
|
||
M25–M31 add reordered fields, dialog relocation, anchor affordances, fieldset
|
||
grouping, textarea input, an unrelated preceding form, and sidebar relocation.
|
||
M32–M38 pair those mechanisms with preserved identifiers. Existing M02/M21/M22
|
||
supply rewritten DOM, reworded controls and renamed fields. New labels were fixed
|
||
from the synthetic contract before running and are agent-authored, not claimed
|
||
as newly human-reviewed ground truth.
|
||
|
||
The 38-entry catalogue contains 27 mechanical, four semantic and seven defect
|
||
mutations. Mechanical acceptance is 26/27. M22 still defeats the heuristic.
|
||
False adaptation is 0/7 on the existing defect set; adding mechanical variants
|
||
does not enlarge that safety denominator. No claim/invariant index changed.
|
||
M25 and M32 are UNCHANGED because the coarse surface fingerprint ignores field
|
||
order and test ids; this is not a claim that their HTML is identical.
|
||
|
||
These are selected mechanisms, not random independent samples. Paired variants
|
||
are correlated; 9/10 is not a population reliability estimate. The control is a
|
||
stable-test-id replay, not every possible conventional selector strategy. Both
|
||
arms remain token-free and test structural HTML only. H-001 stays narrowed to
|
||
identifier-loss cases, and F-0004/F-0005/F-0007's external validity limits remain.
|
||
|
||
## Concept settlement and compatibility (T04/T05)
|
||
|
||
**Keep claims as Python predicates.** Across all three runnable use cases the
|
||
repeated shapes are equality, membership and comparisons between enforcement
|
||
and state. The audit-core consumer additionally compares sanitized absence
|
||
surfaces, bounded execution, cleanup and receipt binding. A small vocabulary
|
||
would either duplicate Python or need frequent escape hatches; serialization
|
||
would create a second language and a second statement of intent.
|
||
|
||
This publishes the interface decision requested by T05: the existing
|
||
`Claim(id, text, provenance, predicate, after_step, source_ref=None)` and
|
||
`Invariant(id, text, provenance, predicate, source_ref=None)` constructors stay
|
||
compatible. Predicates receive independent observation mappings and return
|
||
booleans. Required evidence must use explicit lookup, not a pass-producing
|
||
fallback. The oracle converts missing keys or predicate exceptions to
|
||
INCONCLUSIVE. Predicates must not reach into actors or the live SUT.
|
||
Generated regression artifacts continue to import their original scenario
|
||
claims; consumers must ship those modules with the artifact. Fully standalone
|
||
claim serialization is deliberately rejected. The audit-core consumer needs no
|
||
migration; its tests remain in the default suite.
|
||
|
||
**Remove unused lifecycle concepts now.** Temperature, Energy, Confidence,
|
||
Campaign, Metabolism and automated Retirement have informed no decision in
|
||
these three use cases. They leave the current model, not another deferred gate.
|
||
Trajectory stability continues to drive crystallization. H-005 becomes
|
||
`DORMANT-INDEFINITE`; F-0008 is resolved.
|
||
|
||
This is a research-prototype compatibility change: `testdriver.energy`, public
|
||
`EnergyEvent`/`EnergyEventType` exports and the `EvidencePack.energy_events`
|
||
field are removed. Consumers must stop importing or requiring them. Historical
|
||
JSON may retain that field; the current classifier already ignores it.
|
||
Execution identity, timestamps, judgments, S1 realization costs and lineage
|
||
remain; none of these imply a computed lifecycle score. There is no dependency
|
||
addition. Historical concept proposals carry a supersession notice.
|
||
|
||
## Authoring measurement (T07 remains waiting)
|
||
|
||
[Timing receipt](../research/evidence/2026-09-28-authoring.json) preserves the
|
||
actual measurements and line counts. The timer was placed inside the shell
|
||
write operation, after code composition. The approval interval also included
|
||
composition of the next case, and the tenant interval largely measured file
|
||
writes and pytest. These intervals are **not valid authoring-cost measurements**
|
||
and must not be compared as productivity numbers. Human maintenance effort and
|
||
model latency/cost were not captured. The first attempt cannot be reconstructed.
|
||
|
||
The observable concepts consulted and friction are recorded above. T07 retains
|
||
the remaining work: an independent author must implement these contracts afresh
|
||
with a timer started before reading/design, stopping after first passing defect
|
||
checks, separately recording tooling wait and review. Do not count a rerun or
|
||
copy of these implementations as a new authoring sample. No replacement task
|
||
or workplan is created.
|
||
|
||
## Real-system readiness (T08)
|
||
|
||
**Verdict: not ready for autonomous real-system application.** Local research and
|
||
attended review of prepared claims can continue. Completing the review does not
|
||
satisfy the workplan's live-model success gate or authorize a production run.
|
||
|
||
| Required criterion | Current evidence | Assessment |
|
||
|---|---|---|
|
||
| Independent observation channel can read records and probe enforcement without trusting actors | Three labs; audit-core declares an independent observer but has no runnable driver/collector here | Not met outside synthetic labs; adapter construction, permissions and operating cost remain unmeasured |
|
||
| Claims trace to independently approved intent predating the tested behavior | Local provenance guards; audit-core cites approved engagement and workplan, while its precedent receipt is calibration only | Structurally supported, operational chain must be checked for the exact next engagement; a ticket written after code is insufficient without human promotion |
|
||
| Legitimate behavioral ambiguity is escalated without rewriting intent | M12/M19 are behaviorally indistinguishable; AMBIGUOUS and BEHAVIOUR_CHANGED require review | Demonstrated in lab, real backlog dispute resolution is still attended; missing evidence stays INCONCLUSIVE |
|
||
| False-adaptation consequences are identified for the engagement | Audit-core tenant disclosure, forged foreign append or leftover identity could be concealed by a false pass | Zero tolerance for automatic acceptance of regression; 0/7 synthetic defects is no guarantee against unseen attacks |
|
||
| Execution, custody, timing, cleanup and target revision are enforced | Audit-core's PhaseContract is durable intent, explicitly not a runnable Scenario; tests use supplied observations | Not met: no causal/time-window runtime, custody or cluster driver, independent cleanup collector or fresh admitted window here |
|
||
| Capability, economics and maintenance costs justify adoption | Token-free structural discovery, deterministic crystallization and invalid first authoring timer | Not met; T01/T06/T07 retain the unanswered measurements |
|
||
|
||
The audit-core material referenced here is the checked-in
|
||
`usecases/audit_core_e2_tenant_boundary.py` contract and its fixture-based tests,
|
||
not a fresh inspection of any deployment or a revalidation of the 2026-08-22
|
||
engagement. A real run needs the exact engagement's target-owner acceptance,
|
||
independent observation/cleanup receipts and appropriate approved access. A
|
||
successful historical receipt cannot substitute for those. The existing T01,
|
||
T06 and T07 tasks retain the experiment blockers; this review creates no
|
||
production implementation commitment or new work record.
|
||
|
||
## Safety follow-up: aborted runs and incomplete evidence
|
||
|
||
A post-closeout review reproduced two false-pass paths that the original suite
|
||
missed. A forbidden-surface exception after a passing prefix returned PASS even
|
||
though later claims had never been judged. Separately, deleting almost every
|
||
judgment and all state snapshots still yielded UNCHANGED/safe-to-accept.
|
||
|
||
The runner now records `scheduled_steps` and `expected_judgments` from the
|
||
scenario before executing any action. Only claims attached to scheduled steps
|
||
are included; the existing one-step browser scenario remains intentionally
|
||
narrow. On a surface violation, every remaining scheduled claim and invariant
|
||
gets an explicit INCONCLUSIVE judgment and no fabricated observation. A prior
|
||
FAIL still dominates. An abort after the last assertion is also INCONCLUSIVE,
|
||
because completing the assertions does not mean the scheduled run completed.
|
||
|
||
The classifier validates both packs against their input manifests. It requires
|
||
exactly one judgment for every expected assertion/step pair and exactly one S1
|
||
realization, S2 realization check and nonempty S3 state snapshot for every
|
||
scheduled step, with the correct strata. Missing, duplicate or invalid verdict
|
||
records, unfinished runs, and missing identity/version fields prevent automatic
|
||
acceptance. The candidate must retain the baseline's step and judgment coverage.
|
||
Evidence failure returns AMBIGUOUS before any adaptation rule is considered.
|
||
|
||
This adds two serialized EvidencePack fields. Historical packs without these
|
||
manifests remain historical artifacts but classify as AMBIGUOUS; rerun the
|
||
scenario to obtain evidence eligible for automatic acceptance. Do not infer a
|
||
manifest from surviving observations, since that would preserve the original
|
||
hole. These checks establish structural completeness, not authenticity of a
|
||
pack or correctness of arbitrary caller-supplied predicates. The existing
|
||
actor/observer separation and deterministic oracle remain required.
|
||
|
||
Regression coverage in `tests/test_evidence_completeness.py` includes aborts at
|
||
each step, abort after the final assertion, preservation of an earlier failure,
|
||
individual missing strata/records, duplicate records, and identical truncation
|
||
of both packs. Complete ordinary runs and intentionally narrow scenarios retain
|
||
their existing behavior. The two defects are handled under the existing T08
|
||
readiness review; the external T01/T06/T07 blockers remain unchanged.
|
||
|
||
Validation after the safety fixes: **299 tests passed** (`python3 -m pytest -q`).
|
||
The new regression module had 55 failures and five passing controls before the
|
||
fix; `git diff --check` passes after it.
|
||
|
||
## Admission follow-up: crystallization and intent revisions
|
||
|
||
The next boundary review reproduced three additional holes: repeated aborted
|
||
runs qualified as stable; a complete failing run compared with itself was
|
||
UNCHANGED/safe-to-accept; and changing a claim's text or predicate while retaining
|
||
its ID/provenance was invisible to the classifier. These fixes remain under T08.
|
||
|
||
Crystallization now uses the classifier's acceptance gate for every run against
|
||
the first run. All packs must have complete evidence, passing judgments, verified
|
||
postconditions, unchanged intent and successful realizations. Scenario/use-case
|
||
identity, scheduled steps and expected judgments must match across the window.
|
||
Repeated copies of one run receipt do not satisfy the independent-run count.
|
||
Only then does trajectory comparison decide stability. A refused window returns
|
||
no trajectories. An unchanged product failure is not a stable success.
|
||
|
||
The classifier requires a passing baseline. If the baseline contains FAIL or
|
||
INCONCLUSIVE, even a repaired candidate requires a new passing reference rather
|
||
than automatic acceptance against the failed baseline. Candidate failures against
|
||
a passing reference still report BEHAVIOUR_CHANGED. Baseline realization must
|
||
also be established before either automatic-acceptance outcome is returned;
|
||
missing or non-boolean postconditions cannot be treated as success.
|
||
|
||
Evidence packs now include `intent_revisions`: SHA-256 revisions recorded before
|
||
execution for the use case and each claim/invariant, including claims outside a
|
||
partial scenario's schedule. They include narrative or assertion text, source,
|
||
provenance, claim scheduling, and predicate implementation. The predicate hash
|
||
covers bytecode/constants, defaults, keyword defaults, closure values, referenced
|
||
Python helper functions and their referenced global values. It is bound to the
|
||
Python implementation/version; checkout filenames and line numbers are excluded.
|
||
Only hashes enter evidence, not captured values or code. Changed definitions
|
||
produce INTENT_CHANGED, which requests review and does not assert human approval.
|
||
An internally complete revised claim schedule is an intent change; missing
|
||
baseline coverage under unchanged intent remains incomplete evidence.
|
||
|
||
No Claim/Invariant constructor change or claim serialization language is added.
|
||
`revisions.py` is input-change detection for pure Python predicates. It does not
|
||
prove semantic equivalence, sandbox Python, attest a caller-supplied pack or make
|
||
mutable runtime behavior safe. Predicate dependencies must stay stable during a
|
||
run. Cyclic/opaque callables, module dependencies and dynamic builtin lookups
|
||
cannot be fingerprinted reliably and receive a null revision. They may still be
|
||
evaluated by the oracle, but their evidence cannot be automatically accepted or
|
||
crystallized. Refactor such predicates to consume plain independent snapshots.
|
||
Changing Python versions conservatively requires intent review and fresh evidence.
|
||
|
||
Older packs without usable intent revisions require a rerun, just as packs
|
||
without the earlier completeness manifest do. Existing historical experiment
|
||
receipts are not retroactively upgraded or rewritten. Tests in
|
||
`tests/test_acceptance_boundaries.py` cover the reproduced cases, helper/global
|
||
and captured-value changes, invariant/narrative/source/scheduling changes,
|
||
cross-process identity, source relocation, redaction of captured values, malformed
|
||
postconditions, and the complete-passing control. No live-model, browser-engine
|
||
or production readiness blocker is removed by this work.
|
||
|
||
Validation for this follow-up: initial reproduction had 20 failures and two
|
||
passing controls; full suite **327 passed**. The final dynamic-dependency and
|
||
strict-postcondition additions passed **32 boundary checks**, including four
|
||
new cases. These counts overlap; they are not separate independent samples.
|
||
|
||
## Isolation and repeated-action replay follow-up (2026-09-28)
|
||
|
||
Recorded actor isolation violations now stop the runner before further actions
|
||
or judgments. Unreached scheduled assertions become INCONCLUSIVE; prior observed
|
||
product failures remain FAIL. The framework finding remains separate S3 evidence,
|
||
so it does not add or rewrite product claims. Isolation is examined at entry and
|
||
after each action. Classification requires all these checks to be present, S3,
|
||
and free of violations; crystallization inherits that gate. Older packs with
|
||
only an entry check need rerunning before automatic acceptance. Canary checks
|
||
remain a diagnostic at those boundaries, not proof against arbitrary memory
|
||
access or transient leaks inside a driver.
|
||
|
||
Runner dispatch now passes step identity to drivers that expose `realize_step`.
|
||
Composite drivers forward it while ordinary action-only drivers keep their
|
||
existing interface. Frozen replay selects by step and verifies the action name;
|
||
repeated names no longer overwrite different targets. Unknown/mismatched steps,
|
||
duplicate frozen step ids and ambiguous action-only calls cannot send a request.
|
||
Unambiguous action-only calls remain supported. Replay has no advancing cursor,
|
||
so the same driver can run the schedule repeatedly.
|
||
|
||
Regression coverage includes entry/mid-run/final-action leaks, otherwise passing
|
||
packs with invalid isolation evidence, distinct targets for repeated actions
|
||
through direct and composite dispatch, driver reuse and invalid frozen lookup.
|
||
Decision: `e8fabf6e-3e5e-424e-9a60-152bf3dec641`.
|
||
|
||
Validation: the complete suite passed **343 tests** (153.67 seconds).
|
||
|
||
## Predicate and isolation guard integrity (2026-09-28)
|
||
|
||
Oracle predicates must return actual booleans. None, numbers, strings,
|
||
containers and objects now produce INCONCLUSIVE; their truth-value methods are
|
||
never invoked. Existing True/False results retain PASS/FAIL semantics. Predicate
|
||
results are not copied into the diagnostic, avoiding disclosure of their values.
|
||
|
||
At each existing isolation boundary, the runner now checks for shared memory
|
||
stores, missing or changed stored canaries, and invalid or duplicate actor
|
||
canaries. Removing markers can no longer make shared actor memory appear clean.
|
||
These findings use the existing abort/acceptance/crystallization gates and record
|
||
actor identities without marker or private memory values. This remains a
|
||
boundary diagnostic, not a sandbox against arbitrary or transient memory access.
|
||
|
||
Twenty regressions cover claims and invariants, non-coercible objects, real
|
||
boolean controls, invalid-result admission and crystallization, plus broken
|
||
isolation guards before and during runs. Sixteen failed before implementation;
|
||
the focused suite passed 69 tests afterward. Decision: `2ecac1c5-5254-4ce9-b092-b47c3e1283a9`.
|
||
|
||
Full-suite validation: **363 tests passed** in 153.10 seconds.
|
||
|
||
## Authenticated HTTP and surface boundaries (2026-09-28)
|
||
|
||
Authenticated HTTP now stays on the configured scheme, host and effective port.
|
||
Absolute and protocol-relative targets are validated before adding credentials;
|
||
redirects are validated before forwarding a request. Userinfo and unsupported
|
||
schemes are rejected. Same-origin relative/absolute URLs and redirects continue
|
||
to work. Browser sessions turn origin violations into existing surface findings.
|
||
The generator embeds the same stdlib helper definitions in standalone descendants;
|
||
the checked-in descendant is updated too. No runtime framework dependency is
|
||
introduced into generated tests, and no package dependency is added.
|
||
|
||
Runner also validates the realization's reported surface against the original
|
||
scheduled action, so a driver that forgets its check cannot certify forbidden
|
||
surface substitution. This is a post-call guard: drivers must still check before
|
||
acting, and an actual side effect cannot be undone by the runner. It does not
|
||
attest a dishonest driver's report. Origin restrictions likewise do not claim
|
||
path-level authorization or protect against an already compromised allowed host.
|
||
|
||
Tests use synthetic bearer credentials and local HTTP servers to verify that
|
||
rejected targets/redirects receive no requests, across sessions, freshly generated
|
||
modules and the checked-in descendant. Coverage includes five redirect codes,
|
||
protocol-relative targets, scheme/userinfo rejection, normal origin equivalence,
|
||
same-origin redirect chains and driver omission blocking acceptance/freezing.
|
||
The focused subset passed **107 tests**. Decision: `88490ee8-839d-40dc-affa-b00a1c68e3c5`.
|
||
|
||
Full-suite validation: **398 tests passed** in 165.81 seconds.
|