Complete local generalisation tasks and document remaining experiment blockers

Assistant: codex
Assistant-Model: gpt-6-astra
Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
This commit is contained in:
tegwick 2026-09-28 12:06:24 +02:00
parent d54af628df
commit cf682281d7
36 changed files with 5394 additions and 279 deletions

View file

@ -4,17 +4,25 @@ type: workplan
title: "Generalise the model and settle the open questions"
domain: infotech
repo: test-driver
status: proposed
flavor: planning
status: blocked
flavor: implementation
owner: codex
topic_slug: custodian
created: "2026-08-23"
updated: "2026-08-23"
updated: "2026-09-28"
state_hub_workstream_id: "36a082df-1288-567d-9a0b-f2e0f292786c"
---
# Generalise the model and settle the open questions
## Closeout status — 2026-09-28
T02/T03/T04/T05/T08 are done. T01/T06/T07 remain waiting for external experiment
choices or a valid independent authoring sample, so the workplan is **blocked**.
No new task or workplan was created. The live-model success gate is still unmet;
this is a partial closeout, not a declaration that the workplan succeeded.
Detailed evidence and interface changes: `docs/TestDriverGeneralisationReview.md`.
## Why this workplan exists
`TD-WP-0002` passed its gate: False Adaptation Rate 0/7, one asset crystallized
@ -76,7 +84,7 @@ consumes the kernel as a library. Two consequences for this workplan:
```task
id: TD-WP-0003-T01
status: todo
status: wait
priority: high
state_hub_task_id: "0090ce51-0086-5079-a094-fb63b3b53415"
```
@ -107,11 +115,13 @@ reliably, and the difference will not show in an average.
the live runtime ever runs in the default test suite (recommendation: no — keep
the suite deterministic and free, run the live arm on demand).
**2026-09-28 closeout — wait.** Await an explicit model, fixed run count, cost ceiling and suite policy from Bernd Worsch before the first live call. No live calls or costs were incurred; local heuristic evidence cannot close capability/economics.
## Two further use cases
```task
id: TD-WP-0003-T02
status: todo
status: done
priority: high
state_hub_task_id: "eedf5894-e540-5fa7-9dc3-7e2a5c892bd6"
```
@ -133,11 +143,13 @@ express them**, and every such growth is evidence:
Record the answer in the fitness map either way.
**2026-09-28 closeout — done.** Added delegated approval and tenant lifecycle from prewritten synthetic specifications, using existing kernel concepts and independent snapshot collectors. Baseline, seven seeded defects, missing evidence and S3 replay are covered in tests/test_generalisation.py. Same-agent synthetic authorship is explicitly limited, not human provenance.
## Strengthen the instrument on the deciding axis
```task
id: TD-WP-0003-T03
status: todo
status: done
priority: high
state_hub_task_id: "95406a5d-bcb6-5bcb-abed-834ad16aeb69"
```
@ -153,11 +165,13 @@ side proportionate so the comparison stays fair.
Then re-run E-001 and restate H-001 with a number that means something. Be
prepared for the honest outcome that the narrowed claim narrows further.
**2026-09-28 closeout — done.** Expanded the catalogue to 38 mutations, with ten dropped-identifier mechanisms and seven new paired preserved-id controls. Reproduced E-001: preserved 17/17 in both arms; dropped 9/10 discovery versus 0/10 recorded; full-journey mechanical acceptance 26/27, FAR 0/7. Raw receipt: research/evidence/2026-09-28-e001.json.
## Settle every gated concept
```task
id: TD-WP-0003-T04
status: todo
status: done
priority: medium
state_hub_task_id: "458dc9aa-3d0c-5e67-9f89-3ae8a5a5d297"
```
@ -176,11 +190,13 @@ evidence to say whether they are deferred or dead.
Removing a concept is a result. Extending a gate is not.
**2026-09-28 closeout — done.** Removed Temperature, Energy, Confidence, Campaign, Metabolism and automated Retirement from the current model; removed energy.py, exports and capture field. H-005 is dormant-indefinite; F-0008 resolved. Compatibility changes and superseded historical proposals are explicit.
## Decide how claims are expressed
```task
id: TD-WP-0003-T05
status: todo
status: done
priority: medium
state_hub_task_id: "34cc60cd-b966-54f8-a92d-cfe64587822f"
```
@ -202,6 +218,8 @@ Three candidate answers, and the middle one should not win by default:
Whatever is chosen, publish it as an interface change — the audit-core session
consumes `Claim` directly.
**2026-09-28 closeout — done.** Published the decision to retain Python predicates and imported original claims, after comparing all three runnable use cases and the audit-core consumer. Claim and Invariant signatures remain compatible; standalone serialization is deliberately rejected. See docs/TestDriverGeneralisationReview.md.
## A browser-engine driver, if it is warranted
```task
@ -226,11 +244,13 @@ as the default.
Do not start this before T01–T03. It widens the surface; those three settle what
the surface is for.
**2026-09-28 closeout — wait.** Still requires the existing Playwright setup/licence decision and T01 live-model prerequisite. T02/T03 are now complete. Structural HTML results do not satisfy the browser-engine/visual experiment.
## Measure the cost of expressing a use case
```task
id: TD-WP-0003-T07
status: todo
status: wait
priority: medium
state_hub_task_id: "62f08e27-f307-560f-9583-64c08835e29e"
```
@ -248,11 +268,13 @@ This feeds the metric the project ultimately cares about — *verified behaviour
unit of human maintenance effort* — and it is the number an adopter will ask for
first. It cannot be reconstructed later.
**2026-09-28 closeout — wait.** Partial evidence only: concepts, adapter friction, line counts and raw timer receipts recorded in docs/TestDriverGeneralisationReview.md and research/evidence/2026-09-28-authoring.json. Timer placement missed code composition, so first-authoring elapsed costs are invalid and cannot be reconstructed. Await an independently timed fresh implementation by an author who does not copy these fixtures; retain this task, no replacement record.
## Readiness review for real-system application
```task
id: TD-WP-0003-T08
status: todo
status: done
priority: high
state_hub_task_id: "9e040e94-7245-5b4b-b177-cc337f904ceb"
```
@ -275,3 +297,5 @@ At minimum the criteria must cover:
Then state a verdict, including the verdict "not yet, and here is what is missing".
A readiness review that cannot conclude *not ready* is not a review.
**2026-09-28 closeout — done.** Published explicit real-system gates and a NOT READY verdict for autonomous application, including D-07 cost, provenance from real backlogs, ambiguity, engagement-specific false-adaptation harm, custody/timing/cleanup and economics. The audit-core contract remains intent plus fixture calibration, not a runnable production engagement.