test-driver/research
tegwick eee7722714 T09: crystallization
A stable agentic realization becomes deterministic code. All four exit
criteria met; 163 tests pass.

- crystallization.py: trajectory capture, stability assessment requiring the
  same path across several runs, CrystallizedDriver, pytest codegen
- crystallized/test_grant_access.py: generated, runs with no model, carries
  its lineage in the docstring
- descendant preserves the ancestor's oracle set, agrees with it across five
  lab versions, and still catches a seeded defect
- reversibility shown both ways via new M24 (grant endpoint renamed): the
  frozen descendant fails loudly rather than searching, and the agentic
  ancestor recovers from the same mutation

F-0007 (open): the 54% cost reduction must not be quoted in support of the
thesis. The T07 runtime is token-free, so the measured saving is one page
fetch, one parse and a two-candidate scoring pass. The saving the concept
actually claims - tokens, latency, retry variance - is unmeasured. Together
with F-0005 this makes a bounded live-model experiment the highest-value next
investment.

Assertions in the generated test are imported rather than restated, so it is
not fully standalone. Deliberate: paraphrased claims would be a second
unverified statement of intent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 1629012@bnt-lap001
Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39
2026-08-23 00:21:48 +02:00
..
concepts T09: crystallization 2026-08-23 00:21:48 +02:00
decisions T03: research control plane 2026-08-22 23:11:21 +02:00
experiments T09: crystallization 2026-08-23 00:21:48 +02:00
findings T09: crystallization 2026-08-23 00:21:48 +02:00
hypotheses T09: crystallization 2026-08-23 00:21:48 +02:00
README.md T03: research control plane 2026-08-22 23:11:21 +02:00

Research Control Plane

Deliberately small. This directory exists so that claims about test-driver can be falsified rather than accumulated. It is plain files — no CLI, no schema, no tooling — until there are enough readings to justify tooling.

research/
├── hypotheses/   H-NNN — a claim with a falsification condition
├── experiments/  E-NNN — a planned or executed test of a hypothesis
├── findings/     F-NNNN — findings about test-driver itself
├── concepts/     the Concept ↔ Implementation Fitness Map
└── decisions/    pointers to decisions recorded in State Hub

Identifier convention

Prefix Scope Example
H-NNN Hypothesis H-001
E-NNN Experiment E-001
F-NNNN Framework Finding F-0001
C-<slug> Concept in the fitness map C-actor-isolation
D-NN Decision, scoped to its design note D-07
TD-WP-NNNN-TNN Workplan task (State Hub) TD-WP-0002-T04

Identifiers are stable and never reused. A rejected hypothesis keeps its number.

Hypothesis lifecycle

PROPOSED → EXPERIMENTING → SUPPORTED → PRACTICALLY_VALIDATED → ARCHITECTURAL
                  └──────→ REJECTED

A hypothesis may be reopened if later evidence contradicts it. Reopening is recorded in the file, not by creating a new identifier.

Rules

  1. Every hypothesis states what would falsify it, in terms of an observable outcome, before any experiment runs. A hypothesis with no falsification condition is an opinion.
  2. Concepts with no supporting evidence are marked as such, not quietly retained. The fitness map is expected to contain unsupported entries; hiding them defeats its purpose.
  3. Subtraction counts as progress. A rejected hypothesis or a removed abstraction is a result, not a setback.