Complete local generalisation tasks and document remaining experiment blockers
Assistant: codex Assistant-Model: gpt-6-astra Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
This commit is contained in:
parent
d54af628df
commit
cf682281d7
36 changed files with 5394 additions and 279 deletions
|
|
@ -4,17 +4,25 @@ type: workplan
|
|||
title: "Generalise the model and settle the open questions"
|
||||
domain: infotech
|
||||
repo: test-driver
|
||||
status: proposed
|
||||
flavor: planning
|
||||
status: blocked
|
||||
flavor: implementation
|
||||
owner: codex
|
||||
topic_slug: custodian
|
||||
created: "2026-08-23"
|
||||
updated: "2026-08-23"
|
||||
updated: "2026-09-28"
|
||||
state_hub_workstream_id: "36a082df-1288-567d-9a0b-f2e0f292786c"
|
||||
---
|
||||
|
||||
# Generalise the model and settle the open questions
|
||||
|
||||
## Closeout status — 2026-09-28
|
||||
|
||||
T02/T03/T04/T05/T08 are done. T01/T06/T07 remain waiting for external experiment
|
||||
choices or a valid independent authoring sample, so the workplan is **blocked**.
|
||||
No new task or workplan was created. The live-model success gate is still unmet;
|
||||
this is a partial closeout, not a declaration that the workplan succeeded.
|
||||
Detailed evidence and interface changes: `docs/TestDriverGeneralisationReview.md`.
|
||||
|
||||
## Why this workplan exists
|
||||
|
||||
`TD-WP-0002` passed its gate: False Adaptation Rate 0/7, one asset crystallized
|
||||
|
|
@ -76,7 +84,7 @@ consumes the kernel as a library. Two consequences for this workplan:
|
|||
|
||||
```task
|
||||
id: TD-WP-0003-T01
|
||||
status: todo
|
||||
status: wait
|
||||
priority: high
|
||||
state_hub_task_id: "0090ce51-0086-5079-a094-fb63b3b53415"
|
||||
```
|
||||
|
|
@ -107,11 +115,13 @@ reliably, and the difference will not show in an average.
|
|||
the live runtime ever runs in the default test suite (recommendation: no — keep
|
||||
the suite deterministic and free, run the live arm on demand).
|
||||
|
||||
**2026-09-28 closeout — wait.** Await an explicit model, fixed run count, cost ceiling and suite policy from Bernd Worsch before the first live call. No live calls or costs were incurred; local heuristic evidence cannot close capability/economics.
|
||||
|
||||
## Two further use cases
|
||||
|
||||
```task
|
||||
id: TD-WP-0003-T02
|
||||
status: todo
|
||||
status: done
|
||||
priority: high
|
||||
state_hub_task_id: "eedf5894-e540-5fa7-9dc3-7e2a5c892bd6"
|
||||
```
|
||||
|
|
@ -133,11 +143,13 @@ express them**, and every such growth is evidence:
|
|||
|
||||
Record the answer in the fitness map either way.
|
||||
|
||||
**2026-09-28 closeout — done.** Added delegated approval and tenant lifecycle from prewritten synthetic specifications, using existing kernel concepts and independent snapshot collectors. Baseline, seven seeded defects, missing evidence and S3 replay are covered in tests/test_generalisation.py. Same-agent synthetic authorship is explicitly limited, not human provenance.
|
||||
|
||||
## Strengthen the instrument on the deciding axis
|
||||
|
||||
```task
|
||||
id: TD-WP-0003-T03
|
||||
status: todo
|
||||
status: done
|
||||
priority: high
|
||||
state_hub_task_id: "95406a5d-bcb6-5bcb-abed-834ad16aeb69"
|
||||
```
|
||||
|
|
@ -153,11 +165,13 @@ side proportionate so the comparison stays fair.
|
|||
Then re-run E-001 and restate H-001 with a number that means something. Be
|
||||
prepared for the honest outcome that the narrowed claim narrows further.
|
||||
|
||||
**2026-09-28 closeout — done.** Expanded the catalogue to 38 mutations, with ten dropped-identifier mechanisms and seven new paired preserved-id controls. Reproduced E-001: preserved 17/17 in both arms; dropped 9/10 discovery versus 0/10 recorded; full-journey mechanical acceptance 26/27, FAR 0/7. Raw receipt: research/evidence/2026-09-28-e001.json.
|
||||
|
||||
## Settle every gated concept
|
||||
|
||||
```task
|
||||
id: TD-WP-0003-T04
|
||||
status: todo
|
||||
status: done
|
||||
priority: medium
|
||||
state_hub_task_id: "458dc9aa-3d0c-5e67-9f89-3ae8a5a5d297"
|
||||
```
|
||||
|
|
@ -176,11 +190,13 @@ evidence to say whether they are deferred or dead.
|
|||
|
||||
Removing a concept is a result. Extending a gate is not.
|
||||
|
||||
**2026-09-28 closeout — done.** Removed Temperature, Energy, Confidence, Campaign, Metabolism and automated Retirement from the current model; removed energy.py, exports and capture field. H-005 is dormant-indefinite; F-0008 resolved. Compatibility changes and superseded historical proposals are explicit.
|
||||
|
||||
## Decide how claims are expressed
|
||||
|
||||
```task
|
||||
id: TD-WP-0003-T05
|
||||
status: todo
|
||||
status: done
|
||||
priority: medium
|
||||
state_hub_task_id: "34cc60cd-b966-54f8-a92d-cfe64587822f"
|
||||
```
|
||||
|
|
@ -202,6 +218,8 @@ Three candidate answers, and the middle one should not win by default:
|
|||
Whatever is chosen, publish it as an interface change — the audit-core session
|
||||
consumes `Claim` directly.
|
||||
|
||||
**2026-09-28 closeout — done.** Published the decision to retain Python predicates and imported original claims, after comparing all three runnable use cases and the audit-core consumer. Claim and Invariant signatures remain compatible; standalone serialization is deliberately rejected. See docs/TestDriverGeneralisationReview.md.
|
||||
|
||||
## A browser-engine driver, if it is warranted
|
||||
|
||||
```task
|
||||
|
|
@ -226,11 +244,13 @@ as the default.
|
|||
Do not start this before T01–T03. It widens the surface; those three settle what
|
||||
the surface is for.
|
||||
|
||||
**2026-09-28 closeout — wait.** Still requires the existing Playwright setup/licence decision and T01 live-model prerequisite. T02/T03 are now complete. Structural HTML results do not satisfy the browser-engine/visual experiment.
|
||||
|
||||
## Measure the cost of expressing a use case
|
||||
|
||||
```task
|
||||
id: TD-WP-0003-T07
|
||||
status: todo
|
||||
status: wait
|
||||
priority: medium
|
||||
state_hub_task_id: "62f08e27-f307-560f-9583-64c08835e29e"
|
||||
```
|
||||
|
|
@ -248,11 +268,13 @@ This feeds the metric the project ultimately cares about — *verified behaviour
|
|||
unit of human maintenance effort* — and it is the number an adopter will ask for
|
||||
first. It cannot be reconstructed later.
|
||||
|
||||
**2026-09-28 closeout — wait.** Partial evidence only: concepts, adapter friction, line counts and raw timer receipts recorded in docs/TestDriverGeneralisationReview.md and research/evidence/2026-09-28-authoring.json. Timer placement missed code composition, so first-authoring elapsed costs are invalid and cannot be reconstructed. Await an independently timed fresh implementation by an author who does not copy these fixtures; retain this task, no replacement record.
|
||||
|
||||
## Readiness review for real-system application
|
||||
|
||||
```task
|
||||
id: TD-WP-0003-T08
|
||||
status: todo
|
||||
status: done
|
||||
priority: high
|
||||
state_hub_task_id: "9e040e94-7245-5b4b-b177-cc337f904ceb"
|
||||
```
|
||||
|
|
@ -275,3 +297,5 @@ At minimum the criteria must cover:
|
|||
|
||||
Then state a verdict, including the verdict "not yet, and here is what is missing".
|
||||
A readiness review that cannot conclude *not ready* is not a review.
|
||||
|
||||
**2026-09-28 closeout — done.** Published explicit real-system gates and a NOT READY verdict for autonomous application, including D-07 cost, provenance from real backlogs, ambiguity, engagement-specific false-adaptation harm, custody/timing/cleanup and economics. The audit-core contract remains intent plus fixture calibration, not a runnable production engagement.
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue