Assistant: codex Assistant-Model: gpt-6-astra Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
146 lines
9.4 KiB
Markdown
146 lines
9.4 KiB
Markdown
# TD-WP-0003 generalisation and settlement — 2026-09-28
|
||
|
||
The kernel now expresses sharing/revocation, delegated approval and tenant
|
||
lifecycle without additional framework concepts. This is evidence of local
|
||
expressiveness, not readiness for autonomous verification of a real system.
|
||
|
||
## New use cases (T02)
|
||
|
||
The agent wrote [synthetic contracts](../usecases/generalisation-spec.md) before
|
||
the implementations. Claims are `agent-from-spec`, not human-authored. The same
|
||
agent wrote the synthetic requirements and labs; this does not establish
|
||
independence from a real product team or requirements quality.
|
||
|
||
- `scenarios/delegated_approval.py`: premature execution, submission, delegation,
|
||
rejected self-approval, rejected former-reviewer approval, delegate approval,
|
||
requester execution. Three independently enabled defects must fail.
|
||
- `scenarios/tenant_lifecycle.py`: two administrators use the same local id;
|
||
foreign delete attempts fail; deletion and recreation preserve the other
|
||
tenant. Four independently enabled defects must fail.
|
||
|
||
`tests/test_generalisation.py` verifies baselines, each defect, missing-evidence
|
||
INCONCLUSIVE, deterministic replay and verdict reproduction from serialized S3.
|
||
Separate observers read state and probe enforcement; drivers only emit S1.
|
||
No claims or invariants are learned from the lab's responses.
|
||
|
||
Ordered `Step`s already express the required Schedule. `Scenario.variant`
|
||
identifies each lab fault. A separate scheduler, Lens or Variant object was not
|
||
needed. There is no concurrent scheduling or time-window evidence here.
|
||
|
||
The resistance was in the adapters, not new concepts: `StateObserver`'s standard
|
||
snapshot is specific to resource grants, so each domain needs its own collector
|
||
implementing `snapshot()` and `name`. Expected refusal must be a completed
|
||
protocol response with an independently observed receipt; setting
|
||
`Realization.raised` would suppress dependent claims as INCONCLUSIVE. Unexpected
|
||
exceptions still abort these fixture drivers; they are not silently swallowed.
|
||
|
||
## Structural experiment (T03)
|
||
|
||
[Machine-readable results](../research/evidence/2026-09-28-e001.json) retain each
|
||
arm's surface, judgments, metrics and the journey classification/signals.
|
||
Reproduce with:
|
||
|
||
```bash
|
||
PYTHONPATH=src:. python3 tools/measure_e001.py /tmp/e001.json
|
||
```
|
||
|
||
| Identifier axis | Discovery | Recorded selectors |
|
||
|---|---:|---:|
|
||
| Preserved | 17/17 | 17/17 |
|
||
| Dropped | 9/10 | 0/10 |
|
||
|
||
M25–M31 add reordered fields, dialog relocation, anchor affordances, fieldset
|
||
grouping, textarea input, an unrelated preceding form, and sidebar relocation.
|
||
M32–M38 pair those mechanisms with preserved identifiers. Existing M02/M21/M22
|
||
supply rewritten DOM, reworded controls and renamed fields. New labels were fixed
|
||
from the synthetic contract before running and are agent-authored, not claimed
|
||
as newly human-reviewed ground truth.
|
||
|
||
The 38-entry catalogue contains 27 mechanical, four semantic and seven defect
|
||
mutations. Mechanical acceptance is 26/27. M22 still defeats the heuristic.
|
||
False adaptation is 0/7 on the existing defect set; adding mechanical variants
|
||
does not enlarge that safety denominator. No claim/invariant index changed.
|
||
M25 and M32 are UNCHANGED because the coarse surface fingerprint ignores field
|
||
order and test ids; this is not a claim that their HTML is identical.
|
||
|
||
These are selected mechanisms, not random independent samples. Paired variants
|
||
are correlated; 9/10 is not a population reliability estimate. The control is a
|
||
stable-test-id replay, not every possible conventional selector strategy. Both
|
||
arms remain token-free and test structural HTML only. H-001 stays narrowed to
|
||
identifier-loss cases, and F-0004/F-0005/F-0007's external validity limits remain.
|
||
|
||
## Concept settlement and compatibility (T04/T05)
|
||
|
||
**Keep claims as Python predicates.** Across all three runnable use cases the
|
||
repeated shapes are equality, membership and comparisons between enforcement
|
||
and state. The audit-core consumer additionally compares sanitized absence
|
||
surfaces, bounded execution, cleanup and receipt binding. A small vocabulary
|
||
would either duplicate Python or need frequent escape hatches; serialization
|
||
would create a second language and a second statement of intent.
|
||
|
||
This publishes the interface decision requested by T05: the existing
|
||
`Claim(id, text, provenance, predicate, after_step, source_ref=None)` and
|
||
`Invariant(id, text, provenance, predicate, source_ref=None)` constructors stay
|
||
compatible. Predicates receive independent observation mappings and return
|
||
booleans. Required evidence must use explicit lookup, not a pass-producing
|
||
fallback. The oracle converts missing keys or predicate exceptions to
|
||
INCONCLUSIVE. Predicates must not reach into actors or the live SUT.
|
||
Generated regression artifacts continue to import their original scenario
|
||
claims; consumers must ship those modules with the artifact. Fully standalone
|
||
claim serialization is deliberately rejected. The audit-core consumer needs no
|
||
migration; its tests remain in the default suite.
|
||
|
||
**Remove unused lifecycle concepts now.** Temperature, Energy, Confidence,
|
||
Campaign, Metabolism and automated Retirement have informed no decision in
|
||
these three use cases. They leave the current model, not another deferred gate.
|
||
Trajectory stability continues to drive crystallization. H-005 becomes
|
||
`DORMANT-INDEFINITE`; F-0008 is resolved.
|
||
|
||
This is a research-prototype compatibility change: `testdriver.energy`, public
|
||
`EnergyEvent`/`EnergyEventType` exports and the `EvidencePack.energy_events`
|
||
field are removed. Consumers must stop importing or requiring them. Historical
|
||
JSON may retain that field; the current classifier already ignores it.
|
||
Execution identity, timestamps, judgments, S1 realization costs and lineage
|
||
remain; none of these imply a computed lifecycle score. There is no dependency
|
||
addition. Historical concept proposals carry a supersession notice.
|
||
|
||
## Authoring measurement (T07 remains waiting)
|
||
|
||
[Timing receipt](../research/evidence/2026-09-28-authoring.json) preserves the
|
||
actual measurements and line counts. The timer was placed inside the shell
|
||
write operation, after code composition. The approval interval also included
|
||
composition of the next case, and the tenant interval largely measured file
|
||
writes and pytest. These intervals are **not valid authoring-cost measurements**
|
||
and must not be compared as productivity numbers. Human maintenance effort and
|
||
model latency/cost were not captured. The first attempt cannot be reconstructed.
|
||
|
||
The observable concepts consulted and friction are recorded above. T07 retains
|
||
the remaining work: an independent author must implement these contracts afresh
|
||
with a timer started before reading/design, stopping after first passing defect
|
||
checks, separately recording tooling wait and review. Do not count a rerun or
|
||
copy of these implementations as a new authoring sample. No replacement task
|
||
or workplan is created.
|
||
|
||
## Real-system readiness (T08)
|
||
|
||
**Verdict: not ready for autonomous real-system application.** Local research and
|
||
attended review of prepared claims can continue. Completing the review does not
|
||
satisfy the workplan's live-model success gate or authorize a production run.
|
||
|
||
| Required criterion | Current evidence | Assessment |
|
||
|---|---|---|
|
||
| Independent observation channel can read records and probe enforcement without trusting actors | Three labs; audit-core declares an independent observer but has no runnable driver/collector here | Not met outside synthetic labs; adapter construction, permissions and operating cost remain unmeasured |
|
||
| Claims trace to independently approved intent predating the tested behavior | Local provenance guards; audit-core cites approved engagement and workplan, while its precedent receipt is calibration only | Structurally supported, operational chain must be checked for the exact next engagement; a ticket written after code is insufficient without human promotion |
|
||
| Legitimate behavioral ambiguity is escalated without rewriting intent | M12/M19 are behaviorally indistinguishable; AMBIGUOUS and BEHAVIOUR_CHANGED require review | Demonstrated in lab, real backlog dispute resolution is still attended; missing evidence stays INCONCLUSIVE |
|
||
| False-adaptation consequences are identified for the engagement | Audit-core tenant disclosure, forged foreign append or leftover identity could be concealed by a false pass | Zero tolerance for automatic acceptance of regression; 0/7 synthetic defects is no guarantee against unseen attacks |
|
||
| Execution, custody, timing, cleanup and target revision are enforced | Audit-core's PhaseContract is durable intent, explicitly not a runnable Scenario; tests use supplied observations | Not met: no causal/time-window runtime, custody or cluster driver, independent cleanup collector or fresh admitted window here |
|
||
| Capability, economics and maintenance costs justify adoption | Token-free structural discovery, deterministic crystallization and invalid first authoring timer | Not met; T01/T06/T07 retain the unanswered measurements |
|
||
|
||
The audit-core material referenced here is the checked-in
|
||
`usecases/audit_core_e2_tenant_boundary.py` contract and its fixture-based tests,
|
||
not a fresh inspection of any deployment or a revalidation of the 2026-08-22
|
||
engagement. A real run needs the exact engagement's target-owner acceptance,
|
||
independent observation/cleanup receipts and appropriate approved access. A
|
||
successful historical receipt cannot substitute for those. The existing T01,
|
||
T06 and T07 tasks retain the experiment blockers; this review creates no
|
||
production implementation commitment or new work record.
|