Complete local generalisation tasks and document remaining experiment blockers
Assistant: codex Assistant-Model: gpt-6-astra Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
This commit is contained in:
parent
d54af628df
commit
cf682281d7
36 changed files with 5394 additions and 279 deletions
|
|
@ -1,6 +1,6 @@
|
|||
# Concept ↔ Implementation Fitness Map
|
||||
|
||||
**Updated:** 2026-08-23 (TD-WP-0002-T10)
|
||||
**Updated:** 2026-09-28 (TD-WP-0003-T02/T03/T04)
|
||||
|
||||
Traces each important concept to the implementation, experiment and evidence that
|
||||
support it. **Unsupported entries are the point of this map** — a concept with no
|
||||
|
|
@ -28,43 +28,39 @@ were aspirational, not evidenced.
|
|||
|
||||
| Concept | Level | Implementation | Experiment | Evidence | Open question |
|
||||
|---|---|---|---|---|---|
|
||||
| `C-use-case` | C1 | `intent.py` | — | — | Is a use case expressible without leaking mechanics? |
|
||||
| `C-use-case` | C1 | `intent.py`, three reference scenarios | T02 | `tests/test_generalisation.py` | Three synthetic use cases fit; external authoring still unmeasured. |
|
||||
| `C-actor-isolation` | **C2** | `world.py`, `runner.py` | E-001 | `td://self/actor-isolation`, examined every run | **F-0003 resolved** — automatic canaries; observable on every scenario. |
|
||||
| `C-semantic-action` | **C2** | `actions.py`, `agentic.py` | E-001 (partial) | T07 arm comparison | **F-0005** — supported only where stable identifiers are absent. Narrower than the concept model claims. |
|
||||
| `C-semantic-action` | **C2** | `actions.py`, `agentic.py` | E-001 | 2026-09-28 two-arm receipt | **F-0005** — supported only where stable identifiers are absent. Narrower than the concept model claims. |
|
||||
| `C-oracle-independence` | **C2** | `runner.py`, `oracles.py` | E-001, E-003 | — | Independence of components ≠ independence of belief. (H-004) |
|
||||
| `C-evidence-pack` | C1 | `evidence.py` | — | `td://self/evidence-reproducibility` | Verdicts are reproducible from S3 alone, on passing and failing runs. |
|
||||
| `C-observation-channel` | C1 | `lab/app.py` | — | — | **D-07** — required of every system under test. Adoption cost unknown. |
|
||||
| `C-adaptation` | **C2** | `classification.py`, `agentic.py` | E-001 | T08 matrix | 11/12 mechanical absorbed without a human. |
|
||||
| `C-adaptation` | **C2** | `classification.py`, `agentic.py` | E-001 | T08 matrix | 26/27 mechanical absorbed; ten dropped-id cases now included. |
|
||||
| `C-classification` | **C2** | `classification.py` | E-001, E-003 | T08 matrix, FAR 0/7 | **F-0006** — cannot infer SEMANTIC vs DEFECT; collapses to one escalating outcome. |
|
||||
| `C-crystallization` | **C2** | `crystallization.py` | E-002 | T09 agreement matrix | Fidelity supported; **F-0007** — the economic case is unmeasurable with a token-free runtime. |
|
||||
| `C-intent-provenance` | C1 | `provenance.py` | E-003 | `td://self/intent-independence` | Constrains provenance, not quality. Accepted residual. |
|
||||
| `C-lineage` | C1 | `scenario.py`, generated headers | E-002 | generated module docstring | Parent pointer plus provenance in the artefact; no graph. |
|
||||
| `C-energy` | C0 | `energy.py`, capture only | — | none | Dormant. **Gated**: if the next workplan ends with no decision having used the history, remove. |
|
||||
| `C-temperature` | **C0 — under review** | — | — | none | **F-0008** — crystallization was built without it; measured stability did the work. Gated for removal. |
|
||||
| `C-confidence` | C0 | — | — | none | Declared, unimplemented, never consulted. Same position as `C-temperature` without a built alternative. |
|
||||
| `C-campaign` | C0 | — | — | — | Deferred. |
|
||||
| `C-metabolism` | C0 | — | — | — | Deferred. Depends on C-energy. |
|
||||
| `C-retirement` | C0 | — | — | — | Deferred. Depends on C-energy. |
|
||||
| `C-energy` | Removed | — | T04 | Generalisation review | No demonstrated decision; not an implementation promise. |
|
||||
| `C-temperature` | Removed | — | T04 | Generalisation review | No demonstrated decision; not an implementation promise. |
|
||||
| `C-confidence` | Removed | — | T04 | Generalisation review | No demonstrated decision; not an implementation promise. |
|
||||
| `C-campaign` | Removed | — | T04 | Generalisation review | No demonstrated decision; not an implementation promise. |
|
||||
| `C-metabolism` | Removed | — | T04 | Generalisation review | No demonstrated decision; not an implementation promise. |
|
||||
| `C-retirement` | Removed | — | T04 | Generalisation review | No demonstrated decision; not an implementation promise. |
|
||||
| `C-security-mutation` | C1 | `lab/mutations.py` | E-003 | `lab/GROUND-TRUTH.md` | Catalogue is hand-written; no derivation mechanism from use cases yet. |
|
||||
|
||||
## Orphan check
|
||||
## Generalisation and compression review — 2026-09-28
|
||||
|
||||
**Conceptual orphans** — concepts with no planned implementation in `TD-WP-0002`:
|
||||
`C-temperature`, `C-confidence`, `C-campaign`, `C-metabolism`, `C-retirement`.
|
||||
No new kernel concept was required for delegation/sequencing or tenant lifecycle.
|
||||
Ordered `Step`s supply the Schedule; scenario `variant` identifies seeded defects.
|
||||
No independent scheduler or Variant class was needed. Integration/security lenses
|
||||
remain descriptive groupings; no Lens runtime object was consulted or validated.
|
||||
The generic observer interface works, but each new domain needs a custom snapshot
|
||||
collector: observation-channel adoption cost remains a real concern.
|
||||
|
||||
All five are deferred *by explicit decision*, not oversight. They are the group
|
||||
most at risk of being built because they are easy and satisfying, and never
|
||||
validated. They are revisited at T10, where the question is not "when do we build
|
||||
these" but "does the evidence justify keeping them in the model at all".
|
||||
Temperature, Energy, Confidence, Campaign, Metabolism and Retirement are removed,
|
||||
not deferred to another gate. `energy.py`, its exports and evidence field are
|
||||
removed; H-005 is dormant-indefinite and F-0008 is resolved. Ordinary evidence,
|
||||
judgments, realization metrics and lineage remain. See
|
||||
[review](../../docs/TestDriverGeneralisationReview.md) for compatibility and limits.
|
||||
|
||||
**Implementation orphans** — none. Every module in `src/testdriver/` traces to a
|
||||
concept above.
|
||||
|
||||
**Removed at T10** (compression pass — see the gate review § 3):
|
||||
`Verdict.SUSPICIOUS`, `Step.expect_refusal`, `ActorIsolationError`,
|
||||
`World.seed`, `EvidencePack.latest()`, `Trajectory.method`. Each was declared and
|
||||
never used; `SUSPICIOUS` was additionally a verdict no oracle could emit.
|
||||
|
||||
**Gated for removal**: `C-temperature` (F-0008) and `energy.py`. Both survive the
|
||||
current review on the strength of being cheap, not of being used. If the next
|
||||
workplan closes without a decision consulting either, they go.
|
||||
Previous compression (T10): `Verdict.SUSPICIOUS`, `Step.expect_refusal`,
|
||||
`ActorIsolationError`, `World.seed`, `EvidencePack.latest()`, `Trajectory.method`.
|
||||
|
|
|
|||
|
|
@ -7,3 +7,4 @@ file is a pointer table, not a second source of truth.
|
|||
| Ref | Title | Hub ID | Source |
|
||||
|---|---|---|---|
|
||||
| D-01 … D-07 | Stratified evidence and claim provenance as the basis for adaptation safety | `fef5213f-ce9b-44c2-b327-a0b0ba4b6270` | `docs/TestDriverClassificationDesign.md` |
|
||||
| TD-WP-0003 T04/T05/T08 | Python claims, concept removal and not-ready assessment | `d80734d1-01b8-4a98-8e72-c86f1806d572` | `docs/TestDriverGeneralisationReview.md` |
|
||||
|
|
|
|||
19
research/evidence/2026-09-28-authoring.json
Normal file
19
research/evidence/2026-09-28-authoring.json
Normal file
|
|
@ -0,0 +1,19 @@
|
|||
{
|
||||
"approval": {
|
||||
"started": "2026-09-28T09:56:40.697848+00:00",
|
||||
"first_passing_tests": "2026-09-28T09:57:56.524430+00:00",
|
||||
"scenario_lines": 73,
|
||||
"scenario": "scenarios/delegated_approval.py",
|
||||
"limitation": "Timer started after code composition; interval is not authoring time."
|
||||
},
|
||||
"tenant": {
|
||||
"started": "2026-09-28T09:57:56.524430+00:00",
|
||||
"first_passing_tests": "2026-09-28T09:57:57.123098+00:00",
|
||||
"scenario_lines": 70,
|
||||
"scenario": "scenarios/tenant_lifecycle.py",
|
||||
"limitation": "Timer started after code composition; interval is not authoring time."
|
||||
},
|
||||
"valid_authoring_measurement": false,
|
||||
"shared_adapter_and_lab_lines": 134,
|
||||
"shared_test_lines": 80
|
||||
}
|
||||
4524
research/evidence/2026-09-28-e001.json
Normal file
4524
research/evidence/2026-09-28-e001.json
Normal file
File diff suppressed because it is too large
Load diff
|
|
@ -1,7 +1,7 @@
|
|||
---
|
||||
id: E-001
|
||||
title: Mechanical recovery and defect discrimination over the labelled mutation set
|
||||
status: PLANNED
|
||||
status: EXECUTED
|
||||
hypotheses: [H-001, H-002, H-004]
|
||||
task: TD-WP-0002-T08
|
||||
created: "2026-08-22"
|
||||
|
|
@ -37,3 +37,18 @@ tokens · wall time · retries.
|
|||
|
||||
`EXECUTED` 2026-08-22 (T08). Arm A run over 23 mutations; arm B over
|
||||
the 12 mechanical ones. Results in H-001 and H-004. FAR 0/7.
|
||||
|
||||
## Expanded measurement — 2026-09-28 (TD-WP-0003-T03)
|
||||
|
||||
38 mutations, including ten that drop test ids and seven new paired structural
|
||||
controls. Discovery and recorded selectors both pass 17/17 preserved-id cases;
|
||||
on dropped ids discovery passes 9/10 and recorded selectors 0/10. M22 remains
|
||||
unrealizable. The full journey accepts 26/27 mechanical cases; FAR remains 0/7
|
||||
on the unchanged defect set. Claim/provenance indices remain unchanged.
|
||||
|
||||
Reproduce: `PYTHONPATH=src:. python3 tools/measure_e001.py /tmp/e001.json`.
|
||||
[Raw result](../evidence/2026-09-28-e001.json).
|
||||
These selected and correlated synthetic HTML cases do not establish population
|
||||
reliability, visual/browser coverage, model capability or economics. See
|
||||
[review](../../docs/TestDriverGeneralisationReview.md) for the mechanism list and
|
||||
limits. H-001 remains supported only in the narrowed identifier-loss setting.
|
||||
|
|
|
|||
|
|
@ -85,3 +85,10 @@ Fully standalone generation would need claims expressible in a serializable form
|
|||
rather than as Python predicates. That is a real design question — it is the same
|
||||
question as "should scenarios be YAML", deferred at T04 — and both should be
|
||||
answered together at T10, with evidence about which predicates actually recur.
|
||||
|
||||
## Interface settlement — 2026-09-28
|
||||
|
||||
TD-WP-0003-T05 deliberately retains Python predicates and imported original
|
||||
claims. Standalone serialization is not a promised deliverable. See
|
||||
`docs/TestDriverGeneralisationReview.md` for the published contract. The economic
|
||||
finding stays open under TD-WP-0003-T01; no live-model measurement was performed.
|
||||
|
|
|
|||
|
|
@ -2,12 +2,12 @@
|
|||
id: F-0008
|
||||
type: framework-finding
|
||||
class: UNNECESSARY_COMPLEXITY
|
||||
status: open
|
||||
status: resolved
|
||||
discovered: "2026-08-23"
|
||||
discovered_by: TD-WP-0002-T10
|
||||
workplan: TD-WP-0002
|
||||
task: TD-WP-0002-T10
|
||||
carried_to: next workplan
|
||||
carried_to: TD-WP-0003-T04
|
||||
---
|
||||
|
||||
# F-0008 — Temperature may be redundant; measured stability did the work
|
||||
|
|
@ -68,3 +68,9 @@ is currently the clearest instance in the corpus.
|
|||
`Confidence`, `Campaign`, `Metabolism` and `Retirement` are in the same position —
|
||||
declared, unimplemented, never consulted — but Temperature is the one with a
|
||||
built alternative already doing its job, which makes it the decidable case.
|
||||
|
||||
## Resolution — 2026-09-28
|
||||
|
||||
T04 removes Temperature from the current concept set. Three synthetic use cases
|
||||
and the expanded structural experiment require no temperature decision; observed
|
||||
trajectory stability continues to govern crystallization. No replacement gate.
|
||||
|
|
|
|||
|
|
@ -61,3 +61,18 @@ relayout is untested.
|
|||
|
||||
- 2026-08-22 `PROPOSED`. No evidence.
|
||||
- 2026-08-22 `EXPERIMENTING`. Partial evidence from T07; awaiting E-001 in full.
|
||||
|
||||
## Expanded measurement — 2026-09-28 (TD-WP-0003-T03)
|
||||
|
||||
38 mutations, including ten that drop test ids and seven new paired structural
|
||||
controls. Discovery and recorded selectors both pass 17/17 preserved-id cases;
|
||||
on dropped ids discovery passes 9/10 and recorded selectors 0/10. M22 remains
|
||||
unrealizable. The full journey accepts 26/27 mechanical cases; FAR remains 0/7
|
||||
on the unchanged defect set. Claim/provenance indices remain unchanged.
|
||||
|
||||
Reproduce: `PYTHONPATH=src:. python3 tools/measure_e001.py /tmp/e001.json`.
|
||||
[Raw result](../evidence/2026-09-28-e001.json).
|
||||
These selected and correlated synthetic HTML cases do not establish population
|
||||
reliability, visual/browser coverage, model capability or economics. See
|
||||
[review](../../docs/TestDriverGeneralisationReview.md) for the mechanism list and
|
||||
limits. H-001 remains supported only in the narrowed identifier-loss setting.
|
||||
|
|
|
|||
|
|
@ -1,7 +1,7 @@
|
|||
---
|
||||
id: H-005
|
||||
title: Verification Energy
|
||||
status: PROPOSED
|
||||
status: DORMANT-INDEFINITE
|
||||
created: "2026-08-22"
|
||||
experiments: []
|
||||
concepts: [C-energy]
|
||||
|
|
@ -28,9 +28,9 @@ dormant.** Validating it requires event history across many assets over months
|
|||
history the spike will not accumulate. Implementing a scoring function now would
|
||||
produce a number that cannot be checked, which is worse than no number.
|
||||
|
||||
`TD-WP-0002` therefore records **raw immutable `EnergyEvent`s from the first run**
|
||||
and implements no scoring, decay, or selection logic. Events cannot be
|
||||
reconstructed later; scores can always be computed later.
|
||||
`TD-WP-0002` recorded raw events, but no decision used them. TD-WP-0003-T04
|
||||
removed capture and scoring concepts on 2026-09-28. Ordinary run evidence and
|
||||
verdicts remain; no energy history is promised for future reconstruction.
|
||||
|
||||
This is the hypothesis most likely to be **cheaply built and never validated**,
|
||||
which is precisely why it is fenced off.
|
||||
|
|
@ -38,3 +38,5 @@ which is precisely why it is fenced off.
|
|||
## Status log
|
||||
|
||||
- 2026-08-22 `PROPOSED`, dormant by decision. Event capture only.
|
||||
|
||||
- 2026-09-28 `DORMANT-INDEFINITE`. Capture removed; no new implementation gate.
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue