test-driver/research/concepts/fitness-map.md
tegwick eee7722714 T09: crystallization
A stable agentic realization becomes deterministic code. All four exit
criteria met; 163 tests pass.

- crystallization.py: trajectory capture, stability assessment requiring the
  same path across several runs, CrystallizedDriver, pytest codegen
- crystallized/test_grant_access.py: generated, runs with no model, carries
  its lineage in the docstring
- descendant preserves the ancestor's oracle set, agrees with it across five
  lab versions, and still catches a seeded defect
- reversibility shown both ways via new M24 (grant endpoint renamed): the
  frozen descendant fails loudly rather than searching, and the agentic
  ancestor recovers from the same mutation

F-0007 (open): the 54% cost reduction must not be quoted in support of the
thesis. The T07 runtime is token-free, so the measured saving is one page
fetch, one parse and a two-candidate scoring pass. The saving the concept
actually claims - tokens, latency, retry variance - is unmeasured. Together
with F-0005 this makes a bounded live-model experiment the highest-value next
investment.

Assertions in the generated test are imported rather than restated, so it is
not fully standalone. Deliberate: paraphrased claims would be a second
unverified statement of intent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 1629012@bnt-lap001
Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39
2026-08-23 00:21:48 +02:00

63 lines
4.5 KiB
Markdown

# Concept ↔ Implementation Fitness Map
**Updated:** 2026-08-23 (TD-WP-0002-T09)
Traces each important concept to the implementation, experiment and evidence that
support it. **Unsupported entries are the point of this map** — a concept with no
implementation and no evidence is not a gap to be embarrassed about, it is the
current honest state, and hiding it defeats the map's purpose.
Support levels follow `TestDriverImprovementLoop.md` §13:
`C0 Idea` · `C1 Hypothesis` · `C2 Experimentally Supported` ·
`C3 Practically Validated` · `C4 Architectural Invariant`
## Current state
`C-semantic-action` is the first concept to reach `C2`: it has an experiment
behind it (the T07 two-arm comparison), and that experiment narrowed the claim
rather than confirming it. Everything else still rests on unit tests.
The deterministic kernel exists (T04) and its guarantees are covered by unit
tests. **Levels do not move for those.** A passing unit test is not an experiment: it
shows the code does what its author intended, not that the concept holds under
the mutations it claims to survive. Levels rise when E-001/E-002/E-003 produce
evidence, not before. The implementation column below moves; the level column
does not. The initial classifications in §13 of the Improvement Loop
(`Actor Isolation C2`, `Independent Oracles C2`) are corrected downward here: they
were aspirational, not evidenced.
| Concept | Level | Implementation | Experiment | Evidence | Open question |
|---|---|---|---|---|---|
| `C-use-case` | C1 | `intent.py` | — | — | Is a use case expressible without leaking mechanics? |
| `C-actor-isolation` | C1 | `world.py` | E-001 | `td://self/actor-isolation` | **F-0003** — only observable when the scenario plants canaries. |
| `C-semantic-action` | **C2** | `actions.py`, `agentic.py` | E-001 (partial) | T07 arm comparison | **F-0005** — supported only where stable identifiers are absent. Narrower than the concept model claims. |
| `C-oracle-independence` | **C2** | `runner.py`, `oracles.py` | E-001, E-003 | — | Independence of components ≠ independence of belief. (H-004) |
| `C-evidence-pack` | C1 | `evidence.py` | — | `td://self/evidence-reproducibility` | Verdicts are reproducible from S3 alone, on passing and failing runs. |
| `C-observation-channel` | C1 | `lab/app.py` | — | — | **D-07** — required of every system under test. Adoption cost unknown. |
| `C-adaptation` | **C2** | `classification.py`, `agentic.py` | E-001 | T08 matrix | 11/12 mechanical absorbed without a human. |
| `C-classification` | **C2** | `classification.py` | E-001, E-003 | T08 matrix, FAR 0/7 | **F-0006** — cannot infer SEMANTIC vs DEFECT; collapses to one escalating outcome. |
| `C-crystallization` | **C2** | `crystallization.py` | E-002 | T09 agreement matrix | Fidelity supported; **F-0007** — the economic case is unmeasurable with a token-free runtime. |
| `C-intent-provenance` | C1 | `provenance.py` | E-003 | `td://self/intent-independence` | Constrains provenance, not quality. Accepted residual. |
| `C-lineage` | C1 | `scenario.py`, generated headers | E-002 | generated module docstring | Parent pointer plus provenance in the artefact; no graph. |
| `C-energy` | C0 | `energy.py`, capture only | — | — | Dormant by decision. (H-005) |
| `C-temperature` | C0 | — | — | — | Deferred. No implementation planned in TD-WP-0002. |
| `C-confidence` | C0 | — | — | — | Deferred. |
| `C-campaign` | C0 | — | — | — | Deferred. |
| `C-metabolism` | C0 | — | — | — | Deferred. Depends on C-energy. |
| `C-retirement` | C0 | — | — | — | Deferred. Depends on C-energy. |
| `C-security-mutation` | C1 | `lab/mutations.py` | E-003 | `lab/GROUND-TRUTH.md` | Catalogue is hand-written; no derivation mechanism from use cases yet. |
## Orphan check
**Conceptual orphans** — concepts with no planned implementation in `TD-WP-0002`:
`C-temperature`, `C-confidence`, `C-campaign`, `C-metabolism`, `C-retirement`.
All five are deferred *by explicit decision*, not oversight. They are the group
most at risk of being built because they are easy and satisfying, and never
validated. They are revisited at T10, where the question is not "when do we build
these" but "does the evidence justify keeping them in the model at all".
**Implementation orphans** — none. Every module in `src/testdriver/` traces to a
concept above. `energy.py` is the one to watch: it exists solely to capture
events for a dormant hypothesis, and if T10 finds no use for the history it
should be removed rather than kept out of sentiment.