Complete local generalisation tasks and document remaining experiment blockers

Assistant: codex
Assistant-Model: gpt-6-astra
Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
This commit is contained in:
tegwick 2026-09-28 12:06:24 +02:00
parent d54af628df
commit cf682281d7
36 changed files with 5394 additions and 279 deletions

View file

@ -1,6 +1,6 @@
# Concept ↔ Implementation Fitness Map
**Updated:** 2026-08-23 (TD-WP-0002-T10)
**Updated:** 2026-09-28 (TD-WP-0003-T02/T03/T04)
Traces each important concept to the implementation, experiment and evidence that
support it. **Unsupported entries are the point of this map** — a concept with no
@ -28,43 +28,39 @@ were aspirational, not evidenced.
| Concept | Level | Implementation | Experiment | Evidence | Open question |
|---|---|---|---|---|---|
| `C-use-case` | C1 | `intent.py` | — | — | Is a use case expressible without leaking mechanics? |
| `C-use-case` | C1 | `intent.py`, three reference scenarios | T02 | `tests/test_generalisation.py` | Three synthetic use cases fit; external authoring still unmeasured. |
| `C-actor-isolation` | **C2** | `world.py`, `runner.py` | E-001 | `td://self/actor-isolation`, examined every run | **F-0003 resolved** — automatic canaries; observable on every scenario. |
| `C-semantic-action` | **C2** | `actions.py`, `agentic.py` | E-001 (partial) | T07 arm comparison | **F-0005** — supported only where stable identifiers are absent. Narrower than the concept model claims. |
| `C-semantic-action` | **C2** | `actions.py`, `agentic.py` | E-001 | 2026-09-28 two-arm receipt | **F-0005** — supported only where stable identifiers are absent. Narrower than the concept model claims. |
| `C-oracle-independence` | **C2** | `runner.py`, `oracles.py` | E-001, E-003 | — | Independence of components ≠ independence of belief. (H-004) |
| `C-evidence-pack` | C1 | `evidence.py` | — | `td://self/evidence-reproducibility` | Verdicts are reproducible from S3 alone, on passing and failing runs. |
| `C-observation-channel` | C1 | `lab/app.py` | — | — | **D-07** — required of every system under test. Adoption cost unknown. |
| `C-adaptation` | **C2** | `classification.py`, `agentic.py` | E-001 | T08 matrix | 11/12 mechanical absorbed without a human. |
| `C-adaptation` | **C2** | `classification.py`, `agentic.py` | E-001 | T08 matrix | 26/27 mechanical absorbed; ten dropped-id cases now included. |
| `C-classification` | **C2** | `classification.py` | E-001, E-003 | T08 matrix, FAR 0/7 | **F-0006** — cannot infer SEMANTIC vs DEFECT; collapses to one escalating outcome. |
| `C-crystallization` | **C2** | `crystallization.py` | E-002 | T09 agreement matrix | Fidelity supported; **F-0007** — the economic case is unmeasurable with a token-free runtime. |
| `C-intent-provenance` | C1 | `provenance.py` | E-003 | `td://self/intent-independence` | Constrains provenance, not quality. Accepted residual. |
| `C-lineage` | C1 | `scenario.py`, generated headers | E-002 | generated module docstring | Parent pointer plus provenance in the artefact; no graph. |
| `C-energy` | C0 | `energy.py`, capture only | — | none | Dormant. **Gated**: if the next workplan ends with no decision having used the history, remove. |
| `C-temperature` | **C0 — under review** | — | — | none | **F-0008** — crystallization was built without it; measured stability did the work. Gated for removal. |
| `C-confidence` | C0 | — | — | none | Declared, unimplemented, never consulted. Same position as `C-temperature` without a built alternative. |
| `C-campaign` | C0 | — | — | — | Deferred. |
| `C-metabolism` | C0 | — | — | — | Deferred. Depends on C-energy. |
| `C-retirement` | C0 | — | — | — | Deferred. Depends on C-energy. |
| `C-energy` | Removed | — | T04 | Generalisation review | No demonstrated decision; not an implementation promise. |
| `C-temperature` | Removed | — | T04 | Generalisation review | No demonstrated decision; not an implementation promise. |
| `C-confidence` | Removed | — | T04 | Generalisation review | No demonstrated decision; not an implementation promise. |
| `C-campaign` | Removed | — | T04 | Generalisation review | No demonstrated decision; not an implementation promise. |
| `C-metabolism` | Removed | — | T04 | Generalisation review | No demonstrated decision; not an implementation promise. |
| `C-retirement` | Removed | — | T04 | Generalisation review | No demonstrated decision; not an implementation promise. |
| `C-security-mutation` | C1 | `lab/mutations.py` | E-003 | `lab/GROUND-TRUTH.md` | Catalogue is hand-written; no derivation mechanism from use cases yet. |
## Orphan check
## Generalisation and compression review — 2026-09-28
**Conceptual orphans** — concepts with no planned implementation in `TD-WP-0002`:
`C-temperature`, `C-confidence`, `C-campaign`, `C-metabolism`, `C-retirement`.
No new kernel concept was required for delegation/sequencing or tenant lifecycle.
Ordered `Step`s supply the Schedule; scenario `variant` identifies seeded defects.
No independent scheduler or Variant class was needed. Integration/security lenses
remain descriptive groupings; no Lens runtime object was consulted or validated.
The generic observer interface works, but each new domain needs a custom snapshot
collector: observation-channel adoption cost remains a real concern.
All five are deferred *by explicit decision*, not oversight. They are the group
most at risk of being built because they are easy and satisfying, and never
validated. They are revisited at T10, where the question is not "when do we build
these" but "does the evidence justify keeping them in the model at all".
Temperature, Energy, Confidence, Campaign, Metabolism and Retirement are removed,
not deferred to another gate. `energy.py`, its exports and evidence field are
removed; H-005 is dormant-indefinite and F-0008 is resolved. Ordinary evidence,
judgments, realization metrics and lineage remain. See
[review](../../docs/TestDriverGeneralisationReview.md) for compatibility and limits.
**Implementation orphans** — none. Every module in `src/testdriver/` traces to a
concept above.
**Removed at T10** (compression pass — see the gate review § 3):
`Verdict.SUSPICIOUS`, `Step.expect_refusal`, `ActorIsolationError`,
`World.seed`, `EvidencePack.latest()`, `Trajectory.method`. Each was declared and
never used; `SUSPICIOUS` was additionally a verdict no oracle could emit.
**Gated for removal**: `C-temperature` (F-0008) and `energy.py`. Both survive the
current review on the strength of being cheap, not of being used. If the next
workplan closes without a decision consulting either, they go.
Previous compression (T10): `Verdict.SUSPICIOUS`, `Step.expect_refusal`,
`ActorIsolationError`, `World.seed`, `EvidencePack.latest()`, `Trajectory.method`.

View file

@ -7,3 +7,4 @@ file is a pointer table, not a second source of truth.
| Ref | Title | Hub ID | Source |
|---|---|---|---|
| D-01 … D-07 | Stratified evidence and claim provenance as the basis for adaptation safety | `fef5213f-ce9b-44c2-b327-a0b0ba4b6270` | `docs/TestDriverClassificationDesign.md` |
| TD-WP-0003 T04/T05/T08 | Python claims, concept removal and not-ready assessment | `d80734d1-01b8-4a98-8e72-c86f1806d572` | `docs/TestDriverGeneralisationReview.md` |

View file

@ -0,0 +1,19 @@
{
"approval": {
"started": "2026-09-28T09:56:40.697848+00:00",
"first_passing_tests": "2026-09-28T09:57:56.524430+00:00",
"scenario_lines": 73,
"scenario": "scenarios/delegated_approval.py",
"limitation": "Timer started after code composition; interval is not authoring time."
},
"tenant": {
"started": "2026-09-28T09:57:56.524430+00:00",
"first_passing_tests": "2026-09-28T09:57:57.123098+00:00",
"scenario_lines": 70,
"scenario": "scenarios/tenant_lifecycle.py",
"limitation": "Timer started after code composition; interval is not authoring time."
},
"valid_authoring_measurement": false,
"shared_adapter_and_lab_lines": 134,
"shared_test_lines": 80
}

File diff suppressed because it is too large Load diff

View file

@ -1,7 +1,7 @@
---
id: E-001
title: Mechanical recovery and defect discrimination over the labelled mutation set
status: PLANNED
status: EXECUTED
hypotheses: [H-001, H-002, H-004]
task: TD-WP-0002-T08
created: "2026-08-22"
@ -37,3 +37,18 @@ tokens · wall time · retries.
`EXECUTED` 2026-08-22 (T08). Arm A run over 23 mutations; arm B over
the 12 mechanical ones. Results in H-001 and H-004. FAR 0/7.
## Expanded measurement — 2026-09-28 (TD-WP-0003-T03)
38 mutations, including ten that drop test ids and seven new paired structural
controls. Discovery and recorded selectors both pass 17/17 preserved-id cases;
on dropped ids discovery passes 9/10 and recorded selectors 0/10. M22 remains
unrealizable. The full journey accepts 26/27 mechanical cases; FAR remains 0/7
on the unchanged defect set. Claim/provenance indices remain unchanged.
Reproduce: `PYTHONPATH=src:. python3 tools/measure_e001.py /tmp/e001.json`.
[Raw result](../evidence/2026-09-28-e001.json).
These selected and correlated synthetic HTML cases do not establish population
reliability, visual/browser coverage, model capability or economics. See
[review](../../docs/TestDriverGeneralisationReview.md) for the mechanism list and
limits. H-001 remains supported only in the narrowed identifier-loss setting.

View file

@ -85,3 +85,10 @@ Fully standalone generation would need claims expressible in a serializable form
rather than as Python predicates. That is a real design question — it is the same
question as "should scenarios be YAML", deferred at T04 — and both should be
answered together at T10, with evidence about which predicates actually recur.
## Interface settlement — 2026-09-28
TD-WP-0003-T05 deliberately retains Python predicates and imported original
claims. Standalone serialization is not a promised deliverable. See
`docs/TestDriverGeneralisationReview.md` for the published contract. The economic
finding stays open under TD-WP-0003-T01; no live-model measurement was performed.

View file

@ -2,12 +2,12 @@
id: F-0008
type: framework-finding
class: UNNECESSARY_COMPLEXITY
status: open
status: resolved
discovered: "2026-08-23"
discovered_by: TD-WP-0002-T10
workplan: TD-WP-0002
task: TD-WP-0002-T10
carried_to: next workplan
carried_to: TD-WP-0003-T04
---
# F-0008 — Temperature may be redundant; measured stability did the work
@ -68,3 +68,9 @@ is currently the clearest instance in the corpus.
`Confidence`, `Campaign`, `Metabolism` and `Retirement` are in the same position —
declared, unimplemented, never consulted — but Temperature is the one with a
built alternative already doing its job, which makes it the decidable case.
## Resolution — 2026-09-28
T04 removes Temperature from the current concept set. Three synthetic use cases
and the expanded structural experiment require no temperature decision; observed
trajectory stability continues to govern crystallization. No replacement gate.

View file

@ -61,3 +61,18 @@ relayout is untested.
- 2026-08-22 `PROPOSED`. No evidence.
- 2026-08-22 `EXPERIMENTING`. Partial evidence from T07; awaiting E-001 in full.
## Expanded measurement — 2026-09-28 (TD-WP-0003-T03)
38 mutations, including ten that drop test ids and seven new paired structural
controls. Discovery and recorded selectors both pass 17/17 preserved-id cases;
on dropped ids discovery passes 9/10 and recorded selectors 0/10. M22 remains
unrealizable. The full journey accepts 26/27 mechanical cases; FAR remains 0/7
on the unchanged defect set. Claim/provenance indices remain unchanged.
Reproduce: `PYTHONPATH=src:. python3 tools/measure_e001.py /tmp/e001.json`.
[Raw result](../evidence/2026-09-28-e001.json).
These selected and correlated synthetic HTML cases do not establish population
reliability, visual/browser coverage, model capability or economics. See
[review](../../docs/TestDriverGeneralisationReview.md) for the mechanism list and
limits. H-001 remains supported only in the narrowed identifier-loss setting.

View file

@ -1,7 +1,7 @@
---
id: H-005
title: Verification Energy
status: PROPOSED
status: DORMANT-INDEFINITE
created: "2026-08-22"
experiments: []
concepts: [C-energy]
@ -28,9 +28,9 @@ dormant.** Validating it requires event history across many assets over months
history the spike will not accumulate. Implementing a scoring function now would
produce a number that cannot be checked, which is worse than no number.
`TD-WP-0002` therefore records **raw immutable `EnergyEvent`s from the first run**
and implements no scoring, decay, or selection logic. Events cannot be
reconstructed later; scores can always be computed later.
`TD-WP-0002` recorded raw events, but no decision used them. TD-WP-0003-T04
removed capture and scoring concepts on 2026-09-28. Ordinary run evidence and
verdicts remain; no energy history is promised for future reconstruction.
This is the hypothesis most likely to be **cheaply built and never validated**,
which is precisely why it is fenced off.
@ -38,3 +38,5 @@ which is precisely why it is fenced off.
## Status log
- 2026-08-22 `PROPOSED`, dormant by decision. Event capture only.
- 2026-09-28 `DORMANT-INDEFINITE`. Capture removed; no new implementation gate.