147 lines
9.4 KiB
Markdown
147 lines
9.4 KiB
Markdown
|
|
# TD-WP-0003 generalisation and settlement — 2026-09-28
|
|||
|
|
|
|||
|
|
The kernel now expresses sharing/revocation, delegated approval and tenant
|
|||
|
|
lifecycle without additional framework concepts. This is evidence of local
|
|||
|
|
expressiveness, not readiness for autonomous verification of a real system.
|
|||
|
|
|
|||
|
|
## New use cases (T02)
|
|||
|
|
|
|||
|
|
The agent wrote [synthetic contracts](../usecases/generalisation-spec.md) before
|
|||
|
|
the implementations. Claims are `agent-from-spec`, not human-authored. The same
|
|||
|
|
agent wrote the synthetic requirements and labs; this does not establish
|
|||
|
|
independence from a real product team or requirements quality.
|
|||
|
|
|
|||
|
|
- `scenarios/delegated_approval.py`: premature execution, submission, delegation,
|
|||
|
|
rejected self-approval, rejected former-reviewer approval, delegate approval,
|
|||
|
|
requester execution. Three independently enabled defects must fail.
|
|||
|
|
- `scenarios/tenant_lifecycle.py`: two administrators use the same local id;
|
|||
|
|
foreign delete attempts fail; deletion and recreation preserve the other
|
|||
|
|
tenant. Four independently enabled defects must fail.
|
|||
|
|
|
|||
|
|
`tests/test_generalisation.py` verifies baselines, each defect, missing-evidence
|
|||
|
|
INCONCLUSIVE, deterministic replay and verdict reproduction from serialized S3.
|
|||
|
|
Separate observers read state and probe enforcement; drivers only emit S1.
|
|||
|
|
No claims or invariants are learned from the lab's responses.
|
|||
|
|
|
|||
|
|
Ordered `Step`s already express the required Schedule. `Scenario.variant`
|
|||
|
|
identifies each lab fault. A separate scheduler, Lens or Variant object was not
|
|||
|
|
needed. There is no concurrent scheduling or time-window evidence here.
|
|||
|
|
|
|||
|
|
The resistance was in the adapters, not new concepts: `StateObserver`'s standard
|
|||
|
|
snapshot is specific to resource grants, so each domain needs its own collector
|
|||
|
|
implementing `snapshot()` and `name`. Expected refusal must be a completed
|
|||
|
|
protocol response with an independently observed receipt; setting
|
|||
|
|
`Realization.raised` would suppress dependent claims as INCONCLUSIVE. Unexpected
|
|||
|
|
exceptions still abort these fixture drivers; they are not silently swallowed.
|
|||
|
|
|
|||
|
|
## Structural experiment (T03)
|
|||
|
|
|
|||
|
|
[Machine-readable results](../research/evidence/2026-09-28-e001.json) retain each
|
|||
|
|
arm's surface, judgments, metrics and the journey classification/signals.
|
|||
|
|
Reproduce with:
|
|||
|
|
|
|||
|
|
```bash
|
|||
|
|
PYTHONPATH=src:. python3 tools/measure_e001.py /tmp/e001.json
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
| Identifier axis | Discovery | Recorded selectors |
|
|||
|
|
|---|---:|---:|
|
|||
|
|
| Preserved | 17/17 | 17/17 |
|
|||
|
|
| Dropped | 9/10 | 0/10 |
|
|||
|
|
|
|||
|
|
M25–M31 add reordered fields, dialog relocation, anchor affordances, fieldset
|
|||
|
|
grouping, textarea input, an unrelated preceding form, and sidebar relocation.
|
|||
|
|
M32–M38 pair those mechanisms with preserved identifiers. Existing M02/M21/M22
|
|||
|
|
supply rewritten DOM, reworded controls and renamed fields. New labels were fixed
|
|||
|
|
from the synthetic contract before running and are agent-authored, not claimed
|
|||
|
|
as newly human-reviewed ground truth.
|
|||
|
|
|
|||
|
|
The 38-entry catalogue contains 27 mechanical, four semantic and seven defect
|
|||
|
|
mutations. Mechanical acceptance is 26/27. M22 still defeats the heuristic.
|
|||
|
|
False adaptation is 0/7 on the existing defect set; adding mechanical variants
|
|||
|
|
does not enlarge that safety denominator. No claim/invariant index changed.
|
|||
|
|
M25 and M32 are UNCHANGED because the coarse surface fingerprint ignores field
|
|||
|
|
order and test ids; this is not a claim that their HTML is identical.
|
|||
|
|
|
|||
|
|
These are selected mechanisms, not random independent samples. Paired variants
|
|||
|
|
are correlated; 9/10 is not a population reliability estimate. The control is a
|
|||
|
|
stable-test-id replay, not every possible conventional selector strategy. Both
|
|||
|
|
arms remain token-free and test structural HTML only. H-001 stays narrowed to
|
|||
|
|
identifier-loss cases, and F-0004/F-0005/F-0007's external validity limits remain.
|
|||
|
|
|
|||
|
|
## Concept settlement and compatibility (T04/T05)
|
|||
|
|
|
|||
|
|
**Keep claims as Python predicates.** Across all three runnable use cases the
|
|||
|
|
repeated shapes are equality, membership and comparisons between enforcement
|
|||
|
|
and state. The audit-core consumer additionally compares sanitized absence
|
|||
|
|
surfaces, bounded execution, cleanup and receipt binding. A small vocabulary
|
|||
|
|
would either duplicate Python or need frequent escape hatches; serialization
|
|||
|
|
would create a second language and a second statement of intent.
|
|||
|
|
|
|||
|
|
This publishes the interface decision requested by T05: the existing
|
|||
|
|
`Claim(id, text, provenance, predicate, after_step, source_ref=None)` and
|
|||
|
|
`Invariant(id, text, provenance, predicate, source_ref=None)` constructors stay
|
|||
|
|
compatible. Predicates receive independent observation mappings and return
|
|||
|
|
booleans. Required evidence must use explicit lookup, not a pass-producing
|
|||
|
|
fallback. The oracle converts missing keys or predicate exceptions to
|
|||
|
|
INCONCLUSIVE. Predicates must not reach into actors or the live SUT.
|
|||
|
|
Generated regression artifacts continue to import their original scenario
|
|||
|
|
claims; consumers must ship those modules with the artifact. Fully standalone
|
|||
|
|
claim serialization is deliberately rejected. The audit-core consumer needs no
|
|||
|
|
migration; its tests remain in the default suite.
|
|||
|
|
|
|||
|
|
**Remove unused lifecycle concepts now.** Temperature, Energy, Confidence,
|
|||
|
|
Campaign, Metabolism and automated Retirement have informed no decision in
|
|||
|
|
these three use cases. They leave the current model, not another deferred gate.
|
|||
|
|
Trajectory stability continues to drive crystallization. H-005 becomes
|
|||
|
|
`DORMANT-INDEFINITE`; F-0008 is resolved.
|
|||
|
|
|
|||
|
|
This is a research-prototype compatibility change: `testdriver.energy`, public
|
|||
|
|
`EnergyEvent`/`EnergyEventType` exports and the `EvidencePack.energy_events`
|
|||
|
|
field are removed. Consumers must stop importing or requiring them. Historical
|
|||
|
|
JSON may retain that field; the current classifier already ignores it.
|
|||
|
|
Execution identity, timestamps, judgments, S1 realization costs and lineage
|
|||
|
|
remain; none of these imply a computed lifecycle score. There is no dependency
|
|||
|
|
addition. Historical concept proposals carry a supersession notice.
|
|||
|
|
|
|||
|
|
## Authoring measurement (T07 remains waiting)
|
|||
|
|
|
|||
|
|
[Timing receipt](../research/evidence/2026-09-28-authoring.json) preserves the
|
|||
|
|
actual measurements and line counts. The timer was placed inside the shell
|
|||
|
|
write operation, after code composition. The approval interval also included
|
|||
|
|
composition of the next case, and the tenant interval largely measured file
|
|||
|
|
writes and pytest. These intervals are **not valid authoring-cost measurements**
|
|||
|
|
and must not be compared as productivity numbers. Human maintenance effort and
|
|||
|
|
model latency/cost were not captured. The first attempt cannot be reconstructed.
|
|||
|
|
|
|||
|
|
The observable concepts consulted and friction are recorded above. T07 retains
|
|||
|
|
the remaining work: an independent author must implement these contracts afresh
|
|||
|
|
with a timer started before reading/design, stopping after first passing defect
|
|||
|
|
checks, separately recording tooling wait and review. Do not count a rerun or
|
|||
|
|
copy of these implementations as a new authoring sample. No replacement task
|
|||
|
|
or workplan is created.
|
|||
|
|
|
|||
|
|
## Real-system readiness (T08)
|
|||
|
|
|
|||
|
|
**Verdict: not ready for autonomous real-system application.** Local research and
|
|||
|
|
attended review of prepared claims can continue. Completing the review does not
|
|||
|
|
satisfy the workplan's live-model success gate or authorize a production run.
|
|||
|
|
|
|||
|
|
| Required criterion | Current evidence | Assessment |
|
|||
|
|
|---|---|---|
|
|||
|
|
| Independent observation channel can read records and probe enforcement without trusting actors | Three labs; audit-core declares an independent observer but has no runnable driver/collector here | Not met outside synthetic labs; adapter construction, permissions and operating cost remain unmeasured |
|
|||
|
|
| Claims trace to independently approved intent predating the tested behavior | Local provenance guards; audit-core cites approved engagement and workplan, while its precedent receipt is calibration only | Structurally supported, operational chain must be checked for the exact next engagement; a ticket written after code is insufficient without human promotion |
|
|||
|
|
| Legitimate behavioral ambiguity is escalated without rewriting intent | M12/M19 are behaviorally indistinguishable; AMBIGUOUS and BEHAVIOUR_CHANGED require review | Demonstrated in lab, real backlog dispute resolution is still attended; missing evidence stays INCONCLUSIVE |
|
|||
|
|
| False-adaptation consequences are identified for the engagement | Audit-core tenant disclosure, forged foreign append or leftover identity could be concealed by a false pass | Zero tolerance for automatic acceptance of regression; 0/7 synthetic defects is no guarantee against unseen attacks |
|
|||
|
|
| Execution, custody, timing, cleanup and target revision are enforced | Audit-core's PhaseContract is durable intent, explicitly not a runnable Scenario; tests use supplied observations | Not met: no causal/time-window runtime, custody or cluster driver, independent cleanup collector or fresh admitted window here |
|
|||
|
|
| Capability, economics and maintenance costs justify adoption | Token-free structural discovery, deterministic crystallization and invalid first authoring timer | Not met; T01/T06/T07 retain the unanswered measurements |
|
|||
|
|
|
|||
|
|
The audit-core material referenced here is the checked-in
|
|||
|
|
`usecases/audit_core_e2_tenant_boundary.py` contract and its fixture-based tests,
|
|||
|
|
not a fresh inspection of any deployment or a revalidation of the 2026-08-22
|
|||
|
|
engagement. A real run needs the exact engagement's target-owner acceptance,
|
|||
|
|
independent observation/cleanup receipts and appropriate approved access. A
|
|||
|
|
successful historical receipt cannot substitute for those. The existing T01,
|
|||
|
|
T06 and T07 tasks retain the experiment blockers; this review creates no
|
|||
|
|
production implementation commitment or new work record.
|