test-driver/docs/TestDriverGeneralisationReview.md
tegwick cf682281d7 Complete local generalisation tasks and document remaining experiment blockers
Assistant: codex
Assistant-Model: gpt-6-astra
Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
2026-09-28 12:06:24 +02:00

9.4 KiB
Raw Blame History

TD-WP-0003 generalisation and settlement — 2026-09-28

The kernel now expresses sharing/revocation, delegated approval and tenant lifecycle without additional framework concepts. This is evidence of local expressiveness, not readiness for autonomous verification of a real system.

New use cases (T02)

The agent wrote synthetic contracts before the implementations. Claims are agent-from-spec, not human-authored. The same agent wrote the synthetic requirements and labs; this does not establish independence from a real product team or requirements quality.

  • scenarios/delegated_approval.py: premature execution, submission, delegation, rejected self-approval, rejected former-reviewer approval, delegate approval, requester execution. Three independently enabled defects must fail.
  • scenarios/tenant_lifecycle.py: two administrators use the same local id; foreign delete attempts fail; deletion and recreation preserve the other tenant. Four independently enabled defects must fail.

tests/test_generalisation.py verifies baselines, each defect, missing-evidence INCONCLUSIVE, deterministic replay and verdict reproduction from serialized S3. Separate observers read state and probe enforcement; drivers only emit S1. No claims or invariants are learned from the lab's responses.

Ordered Steps already express the required Schedule. Scenario.variant identifies each lab fault. A separate scheduler, Lens or Variant object was not needed. There is no concurrent scheduling or time-window evidence here.

The resistance was in the adapters, not new concepts: StateObserver's standard snapshot is specific to resource grants, so each domain needs its own collector implementing snapshot() and name. Expected refusal must be a completed protocol response with an independently observed receipt; setting Realization.raised would suppress dependent claims as INCONCLUSIVE. Unexpected exceptions still abort these fixture drivers; they are not silently swallowed.

Structural experiment (T03)

Machine-readable results retain each arm's surface, judgments, metrics and the journey classification/signals. Reproduce with:

PYTHONPATH=src:. python3 tools/measure_e001.py /tmp/e001.json
Identifier axis Discovery Recorded selectors
Preserved 17/17 17/17
Dropped 9/10 0/10

M25–M31 add reordered fields, dialog relocation, anchor affordances, fieldset grouping, textarea input, an unrelated preceding form, and sidebar relocation. M32–M38 pair those mechanisms with preserved identifiers. Existing M02/M21/M22 supply rewritten DOM, reworded controls and renamed fields. New labels were fixed from the synthetic contract before running and are agent-authored, not claimed as newly human-reviewed ground truth.

The 38-entry catalogue contains 27 mechanical, four semantic and seven defect mutations. Mechanical acceptance is 26/27. M22 still defeats the heuristic. False adaptation is 0/7 on the existing defect set; adding mechanical variants does not enlarge that safety denominator. No claim/invariant index changed. M25 and M32 are UNCHANGED because the coarse surface fingerprint ignores field order and test ids; this is not a claim that their HTML is identical.

These are selected mechanisms, not random independent samples. Paired variants are correlated; 9/10 is not a population reliability estimate. The control is a stable-test-id replay, not every possible conventional selector strategy. Both arms remain token-free and test structural HTML only. H-001 stays narrowed to identifier-loss cases, and F-0004/F-0005/F-0007's external validity limits remain.

Concept settlement and compatibility (T04/T05)

Keep claims as Python predicates. Across all three runnable use cases the repeated shapes are equality, membership and comparisons between enforcement and state. The audit-core consumer additionally compares sanitized absence surfaces, bounded execution, cleanup and receipt binding. A small vocabulary would either duplicate Python or need frequent escape hatches; serialization would create a second language and a second statement of intent.

This publishes the interface decision requested by T05: the existing Claim(id, text, provenance, predicate, after_step, source_ref=None) and Invariant(id, text, provenance, predicate, source_ref=None) constructors stay compatible. Predicates receive independent observation mappings and return booleans. Required evidence must use explicit lookup, not a pass-producing fallback. The oracle converts missing keys or predicate exceptions to INCONCLUSIVE. Predicates must not reach into actors or the live SUT. Generated regression artifacts continue to import their original scenario claims; consumers must ship those modules with the artifact. Fully standalone claim serialization is deliberately rejected. The audit-core consumer needs no migration; its tests remain in the default suite.

Remove unused lifecycle concepts now. Temperature, Energy, Confidence, Campaign, Metabolism and automated Retirement have informed no decision in these three use cases. They leave the current model, not another deferred gate. Trajectory stability continues to drive crystallization. H-005 becomes DORMANT-INDEFINITE; F-0008 is resolved.

This is a research-prototype compatibility change: testdriver.energy, public EnergyEvent/EnergyEventType exports and the EvidencePack.energy_events field are removed. Consumers must stop importing or requiring them. Historical JSON may retain that field; the current classifier already ignores it. Execution identity, timestamps, judgments, S1 realization costs and lineage remain; none of these imply a computed lifecycle score. There is no dependency addition. Historical concept proposals carry a supersession notice.

Authoring measurement (T07 remains waiting)

Timing receipt preserves the actual measurements and line counts. The timer was placed inside the shell write operation, after code composition. The approval interval also included composition of the next case, and the tenant interval largely measured file writes and pytest. These intervals are not valid authoring-cost measurements and must not be compared as productivity numbers. Human maintenance effort and model latency/cost were not captured. The first attempt cannot be reconstructed.

The observable concepts consulted and friction are recorded above. T07 retains the remaining work: an independent author must implement these contracts afresh with a timer started before reading/design, stopping after first passing defect checks, separately recording tooling wait and review. Do not count a rerun or copy of these implementations as a new authoring sample. No replacement task or workplan is created.

Real-system readiness (T08)

Verdict: not ready for autonomous real-system application. Local research and attended review of prepared claims can continue. Completing the review does not satisfy the workplan's live-model success gate or authorize a production run.

Required criterion Current evidence Assessment
Independent observation channel can read records and probe enforcement without trusting actors Three labs; audit-core declares an independent observer but has no runnable driver/collector here Not met outside synthetic labs; adapter construction, permissions and operating cost remain unmeasured
Claims trace to independently approved intent predating the tested behavior Local provenance guards; audit-core cites approved engagement and workplan, while its precedent receipt is calibration only Structurally supported, operational chain must be checked for the exact next engagement; a ticket written after code is insufficient without human promotion
Legitimate behavioral ambiguity is escalated without rewriting intent M12/M19 are behaviorally indistinguishable; AMBIGUOUS and BEHAVIOUR_CHANGED require review Demonstrated in lab, real backlog dispute resolution is still attended; missing evidence stays INCONCLUSIVE
False-adaptation consequences are identified for the engagement Audit-core tenant disclosure, forged foreign append or leftover identity could be concealed by a false pass Zero tolerance for automatic acceptance of regression; 0/7 synthetic defects is no guarantee against unseen attacks
Execution, custody, timing, cleanup and target revision are enforced Audit-core's PhaseContract is durable intent, explicitly not a runnable Scenario; tests use supplied observations Not met: no causal/time-window runtime, custody or cluster driver, independent cleanup collector or fresh admitted window here
Capability, economics and maintenance costs justify adoption Token-free structural discovery, deterministic crystallization and invalid first authoring timer Not met; T01/T06/T07 retain the unanswered measurements

The audit-core material referenced here is the checked-in usecases/audit_core_e2_tenant_boundary.py contract and its fixture-based tests, not a fresh inspection of any deployment or a revalidation of the 2026-08-22 engagement. A real run needs the exact engagement's target-owner acceptance, independent observation/cleanup receipts and appropriate approved access. A successful historical receipt cannot substitute for those. The existing T01, T06 and T07 tasks retain the experiment blockers; this review creates no production implementation commitment or new work record.