test-driver/research/concepts/fitness-map.md
tegwick 4ddb2f896c T05: the lab and its labelled mutation catalogue
lab/app.py (users, tenants, auth, resources, sharing, read/write, revoke,
audit), lab/http_api.py (JSON API + browser UI, stdlib only), 20 labelled
composable version-stamped mutations, ground-truth matrix. 48 tests pass.

Detection against the reference scenario: MECHANICAL 0/10 flagged (correct),
DEFECT 6/6, SEMANTIC 2/4 with both inert cases declared.

- F-0002: M16 and M18 initially escaped detection entirely. A use case
  protects exactly what it asserts. Resolved by adding two claims already
  stated as intent in INTENT.md; the six-mutation catalogue would never have
  surfaced this.
- test-id axis added: stable selectors survive most UI mutations, which would
  make H-001 trivially false. Mutations now vary on preserves_test_ids so the
  hypothesis is analysed split by that axis rather than rigged.
- M12 (semantic deferred revoke) and M19 (defect race) are behaviourally
  identical and asserted as such - the discrimination problem as a test.

lab/minimal.py removed; superseded by lab/app.py.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 1629012@bnt-lap001
Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39
2026-08-22 23:31:22 +02:00

3.7 KiB

Concept ↔ Implementation Fitness Map

Updated: 2026-08-22 (TD-WP-0002-T05)

Traces each important concept to the implementation, experiment and evidence that support it. Unsupported entries are the point of this map — a concept with no implementation and no evidence is not a gap to be embarrassed about, it is the current honest state, and hiding it defeats the map's purpose.

Support levels follow TestDriverImprovementLoop.md §13: C0 Idea · C1 Hypothesis · C2 Experimentally Supported · C3 Practically Validated · C4 Architectural Invariant

Current state

The deterministic kernel exists (T04) and its guarantees are covered by unit tests. Levels have not moved. A passing unit test is not an experiment: it shows the code does what its author intended, not that the concept holds under the mutations it claims to survive. Levels rise when E-001/E-002/E-003 produce evidence, not before. The implementation column below moves; the level column does not. The initial classifications in §13 of the Improvement Loop (Actor Isolation C2, Independent Oracles C2) are corrected downward here: they were aspirational, not evidenced.

Concept Level Implementation Experiment Evidence Open question
C-use-case C1 intent.py Is a use case expressible without leaking mechanics?
C-actor-isolation C1 world.py E-001 Isolation is asserted by construction; unverified.
C-semantic-action C1 actions.py E-001 Does identity survive restructuring better than a recorded sequence? (H-001)
C-oracle-independence C1 runner.py, oracles.py E-001, E-003 Independence of components ≠ independence of belief. (H-004)
C-evidence-pack C1 evidence.py What is the minimum sufficient for replay?
C-observation-channel C1 lab/app.py D-07 — required of every system under test. Adoption cost unknown.
C-adaptation C1 — (T08) E-001 (H-002)
C-classification C1 — (T08) E-001, E-003 Decision table is total on paper; unexercised.
C-crystallization C1 — (T09) E-002 (H-003)
C-intent-provenance C1 provenance.py E-003 Constrains provenance, not quality. Accepted residual.
C-lineage C0 Parent pointer only in the spike.
C-energy C0 energy.py, capture only Dormant by decision. (H-005)
C-temperature C0 Deferred. No implementation planned in TD-WP-0002.
C-confidence C0 Deferred.
C-campaign C0 Deferred.
C-metabolism C0 Deferred. Depends on C-energy.
C-retirement C0 Deferred. Depends on C-energy.
C-security-mutation C1 lab/mutations.py E-003 lab/GROUND-TRUTH.md Catalogue is hand-written; no derivation mechanism from use cases yet.

Orphan check

Conceptual orphans — concepts with no planned implementation in TD-WP-0002: C-temperature, C-confidence, C-campaign, C-metabolism, C-retirement.

All five are deferred by explicit decision, not oversight. They are the group most at risk of being built because they are easy and satisfying, and never validated. They are revisited at T10, where the question is not "when do we build these" but "does the evidence justify keeping them in the model at all".

Implementation orphans — none. Every module in src/testdriver/ traces to a concept above. energy.py is the one to watch: it exists solely to capture events for a dormant hypothesis, and if T10 finds no use for the history it should be removed rather than kept out of sentiment.