test-driver/research/findings/F-0007-crystallization-cost-unmeasurable.md
tegwick eee7722714 T09: crystallization
A stable agentic realization becomes deterministic code. All four exit
criteria met; 163 tests pass.

- crystallization.py: trajectory capture, stability assessment requiring the
  same path across several runs, CrystallizedDriver, pytest codegen
- crystallized/test_grant_access.py: generated, runs with no model, carries
  its lineage in the docstring
- descendant preserves the ancestor's oracle set, agrees with it across five
  lab versions, and still catches a seeded defect
- reversibility shown both ways via new M24 (grant endpoint renamed): the
  frozen descendant fails loudly rather than searching, and the agentic
  ancestor recovers from the same mutation

F-0007 (open): the 54% cost reduction must not be quoted in support of the
thesis. The T07 runtime is token-free, so the measured saving is one page
fetch, one parse and a two-candidate scoring pass. The saving the concept
actually claims - tokens, latency, retry variance - is unmeasured. Together
with F-0005 this makes a bounded live-model experiment the highest-value next
investment.

Assertions in the generated test are imported rather than restated, so it is
not fully standalone. Deliberate: paraphrased claims would be a second
unverified statement of intent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 1629012@bnt-lap001
Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39
2026-08-23 00:21:48 +02:00

3.4 KiB

id type class status discovered discovered_by workplan task hypotheses carried_to
F-0007 framework-finding FRAMEWORK_LIMITATION open 2026-08-23 TD-WP-0002-T09 TD-WP-0002 TD-WP-0002-T09
H-003
TD-WP-0002-T10

F-0007 — Crystallization's economic case cannot be measured yet

The criterion

H-003 includes a cost clause, deliberately:

  • it costs no less to execute than the agentic ancestor.

with the note: if crystallization preserves semantics but saves nothing, the thesis is intact but the product rationale is not.

What was measured

median per run
agentic ancestor 6.87 ms
crystallized descendant 3.17 ms
reduction 53.9 %

The criterion is met — the descendant is measurably cheaper. But the number is close to meaningless as evidence for the thesis.

Why it is close to meaningless

The T07 runtime is a deterministic heuristic, chosen with the operator to avoid API cost and nondeterminism. It consumes zero tokens. So the entire measured saving is one page fetch, one HTML parse and a two-candidate scoring pass — a few milliseconds of local work.

The saving crystallization actually claims is of a different kind and two or three orders of magnitude larger: model tokens, model latency, and the variance that forces retries. None of those exist in this measurement, because none of those exist in this runtime.

So the honest statement is:

Crystallization is measurably cheaper than the ancestor it was frozen from. The measured 54 % is a floor produced by removing local discovery work, and it says nothing about the saving that motivates the concept.

Quoting "54 % cheaper" as support for the crystallization thesis would be misleading, and this finding exists so that nobody does.

What would make it measurable

A live-model runtime behind the same ActorRuntime interface. The RealizationMetrics fields (tokens_in, tokens_out, model, retries) were populated from the first run precisely so this comparison becomes a subtraction rather than a re-run of everything — see T07.

F-0005 already gives an independent reason to want one: M22 defeats the heuristic runtime while remaining solvable by reading a visible label. So a live model would settle two open questions at once — whether it adds capability (F-0005) and whether crystallization has an economic case (this finding).

That makes a bounded live-model experiment the highest-value next investment, above any further framework feature.

A second, smaller limitation

The generated test is not fully standalone. Its realization is plain urllib with no framework dependency, but its assertions are imported from the originating scenario module rather than restated.

That was the right call — a generated test that paraphrases its claims creates a second, unverified statement of intent, and drift between them would be silent. But it qualifies the adoption story in the workplan ("output that drops into a CI system which already exists"): what drops in is the realization, while the claims still require the use-case module on the path.

Fully standalone generation would need claims expressible in a serializable form rather than as Python predicates. That is a real design question — it is the same question as "should scenarios be YAML", deferred at T04 — and both should be answered together at T10, with evidence about which predicates actually recur.