test-driver/research
tegwick 1b9860a8ee T10: gate review and first compression pass
All four gate criteria met. False Adaptation Rate 0/7 with 12 of 13 mechanical
mutations absorbed. 178 tests pass. TD-WP-0002 finished.

Fitness loop closed via F-0003: actor isolation was a property of scenarios
written to expose it, not of runs. Actors now carry an automatic private
marker and the runner examines all of them on every scenario, with two
permanent regressions behind it.

Compression - six abstractions removed, each declared and never used:
Verdict.SUSPICIOUS (a verdict no oracle could emit), Step.expect_refusal,
ActorIsolationError, World.seed, EvidencePack.latest, Trajectory.method.

F-0008: Temperature may be redundant. Crystallization was built without it
ever being consulted; measured stability of realization did the work, and is
observed rather than declared. Gated for removal alongside energy.py.

INTENT_CHANGED and REALIZATION_FAILED had never run. Both now have
purpose-built cases and a test that fails if a seventh outcome is added
without one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 1629012@bnt-lap001
Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39
2026-08-23 00:39:36 +02:00
..
concepts T10: gate review and first compression pass 2026-08-23 00:39:36 +02:00
decisions T03: research control plane 2026-08-22 23:11:21 +02:00
experiments T09: crystallization 2026-08-23 00:21:48 +02:00
findings T10: gate review and first compression pass 2026-08-23 00:39:36 +02:00
hypotheses T09: crystallization 2026-08-23 00:21:48 +02:00
README.md T03: research control plane 2026-08-22 23:11:21 +02:00

Research Control Plane

Deliberately small. This directory exists so that claims about test-driver can be falsified rather than accumulated. It is plain files — no CLI, no schema, no tooling — until there are enough readings to justify tooling.

research/
├── hypotheses/   H-NNN — a claim with a falsification condition
├── experiments/  E-NNN — a planned or executed test of a hypothesis
├── findings/     F-NNNN — findings about test-driver itself
├── concepts/     the Concept ↔ Implementation Fitness Map
└── decisions/    pointers to decisions recorded in State Hub

Identifier convention

Prefix Scope Example
H-NNN Hypothesis H-001
E-NNN Experiment E-001
F-NNNN Framework Finding F-0001
C-<slug> Concept in the fitness map C-actor-isolation
D-NN Decision, scoped to its design note D-07
TD-WP-NNNN-TNN Workplan task (State Hub) TD-WP-0002-T04

Identifiers are stable and never reused. A rejected hypothesis keeps its number.

Hypothesis lifecycle

PROPOSED → EXPERIMENTING → SUPPORTED → PRACTICALLY_VALIDATED → ARCHITECTURAL
                  └──────→ REJECTED

A hypothesis may be reopened if later evidence contradicts it. Reopening is recorded in the file, not by creating a new identifier.

Rules

  1. Every hypothesis states what would falsify it, in terms of an observable outcome, before any experiment runs. A hypothesis with no falsification condition is an opinion.
  2. Concepts with no supporting evidence are marked as such, not quietly retained. The fitness map is expected to contain unsupported entries; hiding them defeats its purpose.
  3. Subtraction counts as progress. A rejected hypothesis or a removed abstraction is a result, not a setback.