A stable agentic realization becomes deterministic code. All four exit criteria met; 163 tests pass. - crystallization.py: trajectory capture, stability assessment requiring the same path across several runs, CrystallizedDriver, pytest codegen - crystallized/test_grant_access.py: generated, runs with no model, carries its lineage in the docstring - descendant preserves the ancestor's oracle set, agrees with it across five lab versions, and still catches a seeded defect - reversibility shown both ways via new M24 (grant endpoint renamed): the frozen descendant fails loudly rather than searching, and the agentic ancestor recovers from the same mutation F-0007 (open): the 54% cost reduction must not be quoted in support of the thesis. The T07 runtime is token-free, so the measured saving is one page fetch, one parse and a two-candidate scoring pass. The saving the concept actually claims - tokens, latency, retry variance - is unmeasured. Together with F-0005 this makes a bounded live-model experiment the highest-value next investment. Assertions in the generated test are imported rather than restated, so it is not fully standalone. Deliberate: paraphrased claims would be a second unverified statement of intent. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Assistant: claude-code Assistant-Model: opus Assistant-Process: 1629012@bnt-lap001 Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39 |
||
|---|---|---|
| .. | ||
| concepts | ||
| decisions | ||
| experiments | ||
| findings | ||
| hypotheses | ||
| README.md | ||
Research Control Plane
Deliberately small. This directory exists so that claims about test-driver can be falsified rather than accumulated. It is plain files — no CLI, no schema, no tooling — until there are enough readings to justify tooling.
research/
├── hypotheses/ H-NNN — a claim with a falsification condition
├── experiments/ E-NNN — a planned or executed test of a hypothesis
├── findings/ F-NNNN — findings about test-driver itself
├── concepts/ the Concept ↔ Implementation Fitness Map
└── decisions/ pointers to decisions recorded in State Hub
Identifier convention
| Prefix | Scope | Example |
|---|---|---|
H-NNN |
Hypothesis | H-001 |
E-NNN |
Experiment | E-001 |
F-NNNN |
Framework Finding | F-0001 |
C-<slug> |
Concept in the fitness map | C-actor-isolation |
D-NN |
Decision, scoped to its design note | D-07 |
TD-WP-NNNN-TNN |
Workplan task (State Hub) | TD-WP-0002-T04 |
Identifiers are stable and never reused. A rejected hypothesis keeps its number.
Hypothesis lifecycle
PROPOSED → EXPERIMENTING → SUPPORTED → PRACTICALLY_VALIDATED → ARCHITECTURAL
└──────→ REJECTED
A hypothesis may be reopened if later evidence contradicts it. Reopening is recorded in the file, not by creating a new identifier.
Rules
- Every hypothesis states what would falsify it, in terms of an observable outcome, before any experiment runs. A hypothesis with no falsification condition is an opinion.
- Concepts with no supporting evidence are marked as such, not quietly retained. The fitness map is expected to contain unsupported entries; hiding them defeats its purpose.
- Subtraction counts as progress. A rejected hypothesis or a removed abstraction is a result, not a setback.