Complete local generalisation tasks and document remaining experiment blockers
Assistant: codex Assistant-Model: gpt-6-astra Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
This commit is contained in:
parent
d54af628df
commit
cf682281d7
36 changed files with 5394 additions and 279 deletions
146
docs/TestDriverGeneralisationReview.md
Normal file
146
docs/TestDriverGeneralisationReview.md
Normal file
|
|
@ -0,0 +1,146 @@
|
|||
# TD-WP-0003 generalisation and settlement — 2026-09-28
|
||||
|
||||
The kernel now expresses sharing/revocation, delegated approval and tenant
|
||||
lifecycle without additional framework concepts. This is evidence of local
|
||||
expressiveness, not readiness for autonomous verification of a real system.
|
||||
|
||||
## New use cases (T02)
|
||||
|
||||
The agent wrote [synthetic contracts](../usecases/generalisation-spec.md) before
|
||||
the implementations. Claims are `agent-from-spec`, not human-authored. The same
|
||||
agent wrote the synthetic requirements and labs; this does not establish
|
||||
independence from a real product team or requirements quality.
|
||||
|
||||
- `scenarios/delegated_approval.py`: premature execution, submission, delegation,
|
||||
rejected self-approval, rejected former-reviewer approval, delegate approval,
|
||||
requester execution. Three independently enabled defects must fail.
|
||||
- `scenarios/tenant_lifecycle.py`: two administrators use the same local id;
|
||||
foreign delete attempts fail; deletion and recreation preserve the other
|
||||
tenant. Four independently enabled defects must fail.
|
||||
|
||||
`tests/test_generalisation.py` verifies baselines, each defect, missing-evidence
|
||||
INCONCLUSIVE, deterministic replay and verdict reproduction from serialized S3.
|
||||
Separate observers read state and probe enforcement; drivers only emit S1.
|
||||
No claims or invariants are learned from the lab's responses.
|
||||
|
||||
Ordered `Step`s already express the required Schedule. `Scenario.variant`
|
||||
identifies each lab fault. A separate scheduler, Lens or Variant object was not
|
||||
needed. There is no concurrent scheduling or time-window evidence here.
|
||||
|
||||
The resistance was in the adapters, not new concepts: `StateObserver`'s standard
|
||||
snapshot is specific to resource grants, so each domain needs its own collector
|
||||
implementing `snapshot()` and `name`. Expected refusal must be a completed
|
||||
protocol response with an independently observed receipt; setting
|
||||
`Realization.raised` would suppress dependent claims as INCONCLUSIVE. Unexpected
|
||||
exceptions still abort these fixture drivers; they are not silently swallowed.
|
||||
|
||||
## Structural experiment (T03)
|
||||
|
||||
[Machine-readable results](../research/evidence/2026-09-28-e001.json) retain each
|
||||
arm's surface, judgments, metrics and the journey classification/signals.
|
||||
Reproduce with:
|
||||
|
||||
```bash
|
||||
PYTHONPATH=src:. python3 tools/measure_e001.py /tmp/e001.json
|
||||
```
|
||||
|
||||
| Identifier axis | Discovery | Recorded selectors |
|
||||
|---|---:|---:|
|
||||
| Preserved | 17/17 | 17/17 |
|
||||
| Dropped | 9/10 | 0/10 |
|
||||
|
||||
M25–M31 add reordered fields, dialog relocation, anchor affordances, fieldset
|
||||
grouping, textarea input, an unrelated preceding form, and sidebar relocation.
|
||||
M32–M38 pair those mechanisms with preserved identifiers. Existing M02/M21/M22
|
||||
supply rewritten DOM, reworded controls and renamed fields. New labels were fixed
|
||||
from the synthetic contract before running and are agent-authored, not claimed
|
||||
as newly human-reviewed ground truth.
|
||||
|
||||
The 38-entry catalogue contains 27 mechanical, four semantic and seven defect
|
||||
mutations. Mechanical acceptance is 26/27. M22 still defeats the heuristic.
|
||||
False adaptation is 0/7 on the existing defect set; adding mechanical variants
|
||||
does not enlarge that safety denominator. No claim/invariant index changed.
|
||||
M25 and M32 are UNCHANGED because the coarse surface fingerprint ignores field
|
||||
order and test ids; this is not a claim that their HTML is identical.
|
||||
|
||||
These are selected mechanisms, not random independent samples. Paired variants
|
||||
are correlated; 9/10 is not a population reliability estimate. The control is a
|
||||
stable-test-id replay, not every possible conventional selector strategy. Both
|
||||
arms remain token-free and test structural HTML only. H-001 stays narrowed to
|
||||
identifier-loss cases, and F-0004/F-0005/F-0007's external validity limits remain.
|
||||
|
||||
## Concept settlement and compatibility (T04/T05)
|
||||
|
||||
**Keep claims as Python predicates.** Across all three runnable use cases the
|
||||
repeated shapes are equality, membership and comparisons between enforcement
|
||||
and state. The audit-core consumer additionally compares sanitized absence
|
||||
surfaces, bounded execution, cleanup and receipt binding. A small vocabulary
|
||||
would either duplicate Python or need frequent escape hatches; serialization
|
||||
would create a second language and a second statement of intent.
|
||||
|
||||
This publishes the interface decision requested by T05: the existing
|
||||
`Claim(id, text, provenance, predicate, after_step, source_ref=None)` and
|
||||
`Invariant(id, text, provenance, predicate, source_ref=None)` constructors stay
|
||||
compatible. Predicates receive independent observation mappings and return
|
||||
booleans. Required evidence must use explicit lookup, not a pass-producing
|
||||
fallback. The oracle converts missing keys or predicate exceptions to
|
||||
INCONCLUSIVE. Predicates must not reach into actors or the live SUT.
|
||||
Generated regression artifacts continue to import their original scenario
|
||||
claims; consumers must ship those modules with the artifact. Fully standalone
|
||||
claim serialization is deliberately rejected. The audit-core consumer needs no
|
||||
migration; its tests remain in the default suite.
|
||||
|
||||
**Remove unused lifecycle concepts now.** Temperature, Energy, Confidence,
|
||||
Campaign, Metabolism and automated Retirement have informed no decision in
|
||||
these three use cases. They leave the current model, not another deferred gate.
|
||||
Trajectory stability continues to drive crystallization. H-005 becomes
|
||||
`DORMANT-INDEFINITE`; F-0008 is resolved.
|
||||
|
||||
This is a research-prototype compatibility change: `testdriver.energy`, public
|
||||
`EnergyEvent`/`EnergyEventType` exports and the `EvidencePack.energy_events`
|
||||
field are removed. Consumers must stop importing or requiring them. Historical
|
||||
JSON may retain that field; the current classifier already ignores it.
|
||||
Execution identity, timestamps, judgments, S1 realization costs and lineage
|
||||
remain; none of these imply a computed lifecycle score. There is no dependency
|
||||
addition. Historical concept proposals carry a supersession notice.
|
||||
|
||||
## Authoring measurement (T07 remains waiting)
|
||||
|
||||
[Timing receipt](../research/evidence/2026-09-28-authoring.json) preserves the
|
||||
actual measurements and line counts. The timer was placed inside the shell
|
||||
write operation, after code composition. The approval interval also included
|
||||
composition of the next case, and the tenant interval largely measured file
|
||||
writes and pytest. These intervals are **not valid authoring-cost measurements**
|
||||
and must not be compared as productivity numbers. Human maintenance effort and
|
||||
model latency/cost were not captured. The first attempt cannot be reconstructed.
|
||||
|
||||
The observable concepts consulted and friction are recorded above. T07 retains
|
||||
the remaining work: an independent author must implement these contracts afresh
|
||||
with a timer started before reading/design, stopping after first passing defect
|
||||
checks, separately recording tooling wait and review. Do not count a rerun or
|
||||
copy of these implementations as a new authoring sample. No replacement task
|
||||
or workplan is created.
|
||||
|
||||
## Real-system readiness (T08)
|
||||
|
||||
**Verdict: not ready for autonomous real-system application.** Local research and
|
||||
attended review of prepared claims can continue. Completing the review does not
|
||||
satisfy the workplan's live-model success gate or authorize a production run.
|
||||
|
||||
| Required criterion | Current evidence | Assessment |
|
||||
|---|---|---|
|
||||
| Independent observation channel can read records and probe enforcement without trusting actors | Three labs; audit-core declares an independent observer but has no runnable driver/collector here | Not met outside synthetic labs; adapter construction, permissions and operating cost remain unmeasured |
|
||||
| Claims trace to independently approved intent predating the tested behavior | Local provenance guards; audit-core cites approved engagement and workplan, while its precedent receipt is calibration only | Structurally supported, operational chain must be checked for the exact next engagement; a ticket written after code is insufficient without human promotion |
|
||||
| Legitimate behavioral ambiguity is escalated without rewriting intent | M12/M19 are behaviorally indistinguishable; AMBIGUOUS and BEHAVIOUR_CHANGED require review | Demonstrated in lab, real backlog dispute resolution is still attended; missing evidence stays INCONCLUSIVE |
|
||||
| False-adaptation consequences are identified for the engagement | Audit-core tenant disclosure, forged foreign append or leftover identity could be concealed by a false pass | Zero tolerance for automatic acceptance of regression; 0/7 synthetic defects is no guarantee against unseen attacks |
|
||||
| Execution, custody, timing, cleanup and target revision are enforced | Audit-core's PhaseContract is durable intent, explicitly not a runnable Scenario; tests use supplied observations | Not met: no causal/time-window runtime, custody or cluster driver, independent cleanup collector or fresh admitted window here |
|
||||
| Capability, economics and maintenance costs justify adoption | Token-free structural discovery, deterministic crystallization and invalid first authoring timer | Not met; T01/T06/T07 retain the unanswered measurements |
|
||||
|
||||
The audit-core material referenced here is the checked-in
|
||||
`usecases/audit_core_e2_tenant_boundary.py` contract and its fixture-based tests,
|
||||
not a fresh inspection of any deployment or a revalidation of the 2026-08-22
|
||||
engagement. A real run needs the exact engagement's target-owner acceptance,
|
||||
independent observation/cleanup receipts and appropriate approved access. A
|
||||
successful historical receipt cannot substitute for those. The existing T01,
|
||||
T06 and T07 tasks retain the experiment blockers; this review creates no
|
||||
production implementation commitment or new work record.
|
||||
Loading…
Add table
Add a link
Reference in a new issue