Register with State Hub, persist concept assessment, seed first workplan

- history/2026-08-22-concept-assessment-swot.md: SWOT assessment of the
  concept corpus with recommendations for the first workplan
- statehub register: infotech domain, TD-WP prefix, generated AGENTS.md,
  .custodian-brief.md and TD-WP-0001 bootstrap workplan
- .repo-classification.yaml: category research, domain infotech
- SCOPE.md rewritten with real repo boundaries
- TD-WP-0002: vertical spike reordering M0-M10 into one end-to-end thread
  that can falsify the crystallization thesis early
- commit previously untracked INTENT.md and docs/

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 1629012@bnt-lap001
Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39
This commit is contained in:
tegwick 2026-08-22 22:40:39 +02:00
parent b471bed707
commit 7249c6a403
13 changed files with 3900 additions and 0 deletions

View file

@ -0,0 +1,240 @@
# test-driver — Concept Assessment (SWOT)
**Date:** 2026-08-22
**Author:** Claude (Opus 5) with Bernd Worsch
**Scope:** Assessment of the pre-implementation concept corpus — `INTENT.md`,
`docs/TestDriverConceptModel.md`, `docs/TestDriverImprovementLoop.md`,
`docs/TestDriverInitialMilestones.md`, `docs/TestDriverResearchPrototype.md`.
**Repo state at assessment:** 2,950 lines of concept documentation, 1 commit,
no executable code, `docs/` and `INTENT.md` untracked.
---
## Orientation
The repository is **concept-complete and code-empty**. Four documents describe a
use-case-driven verification framework whose thesis is that verification should
mature alongside the behaviour it protects: fluid and agentic while software is
hot, deterministic once it cools.
The ratio — roughly 3,000 lines of theory against zero lines executable — is the
single most important fact shaping this assessment.
---
## Strengths
**S1 — The core thesis is original and load-bearing.**
The crystallization continuum (T0 Exploratory → T5 Deterministic) coupled to
implementation *Temperature* is not a repackaging of "self-healing tests". It
makes a falsifiable claim: verification mode should be a function of how fast the
system under test is changing.
**S2 — The right primitive is identified.**
`SemanticAction` as the bridge between agentic discovery and deterministic code
generation is the correct pivot point, and the documents know it.
`grant_access(Bob, R, READ)` surviving a DOM restructure *is* the product.
**S3 — Safety is designed in, not bolted on.**
Oracle independence from actors, "implementation is not the truth",
`INCONCLUSIVE` as a first-class verdict, and **False Adaptation Rate** as the
headline safety metric. This is what separates the concept from self-healing
test vendors, which structurally cannot distinguish a moved button from a broken
authorization check.
**S4 — Falsifiable research posture.**
H-001…H-005 carry falsification conditions; the lab is a measuring instrument;
milestone exit criteria are stated as observable outcomes. Most framework
projects have a roadmap — this has an experiment design.
**S5 — Security-as-mutation is high leverage and cheap.**
Deriving twelve adversarial variants from one shared-resource use case is
demonstrable early, valuable independently of the crystallization thesis, and
requires no maturity beyond T1.
**S6 — Anti-bloat discipline is pre-committed.**
Explicit non-goal list, M10 Compression Review, "self-improvement includes
subtraction", and a deliberately boring stack (Python / pytest / Playwright /
SQLite / YAML). The documents already contain their own best critique.
---
## Weaknesses
**W1 — Concept surface vastly exceeds validated ground.**
Roughly 40 canonical concepts, of which Energy, Confidence, Metabolism, Campaign
selection, Retirement floors and Lineage graphs are untestable until M7+. Each is
an invitation to build now and validate never.
**W2 — Three documents already disagree.**
`INTENT.md` numbers milestones M0M3; `TestDriverInitialMilestones.md` numbers
M0M10; `TestDriverConceptModel.md` §15 has a third M0M3. Concept drift has
appeared before any code exists — precisely the failure mode the Improvement Loop
document was written to prevent.
**W3 — The three hardest problems are the least specified.**
(a) How a `SemanticAction` retains *identity* across a changed surface — the
binding and matching mechanism is nowhere described.
(b) How adaptation is classified MECHANICAL vs SEMANTIC vs DEFECT without an LLM
making the call, which principle 2.3 forbids.
(c) How a captured trajectory becomes deterministic code with oracle coverage
preserved.
These three are the project; the remainder is scaffolding.
**W4 — Energy is the weakest concept and the easiest to build.**
A dangerous combination. Proving that Energy "predicts verification value"
requires years of history the prototype will never accumulate. M9 exists mainly
to collect a dataset for a study that will not happen at this stage.
**W5 — The evaluation set is far too small for its claims.**
Six seeded mutations (M01M06) cannot support "classification precision/recall".
Dozens of labelled mutations across several dimensions are needed before any rate
is more than anecdote.
**W6 — Self-verification is circular.**
`td://self/oracle-independence` uses the framework to test whether the
framework's oracles are independent. Without an out-of-band assertion layer
(plain pytest) as ground truth, M3 proves nothing.
**W7 — No cost or nondeterminism model.**
Agentic runs are stochastic and expensive. Nothing budgets for model flake,
retries, or per-run cost — yet "execution cost reduction" is a stated M7 success
metric.
---
## Opportunities
**O1 — The timing window is open now.**
Agentic development is outrunning test maintenance in exactly the way the *Why*
section describes. The pain is acute and current.
**O2 — The lab is a publishable asset in its own right.**
A benchmark of seeded mechanical / semantic / defect mutations with ground-truth
labels is something the field lacks entirely. It could earn credibility and
contributors faster than the framework does, at a fraction of the cost.
**O3 — "False Adaptation Rate" is a naming land-grab.**
If self-healing test tools come to be held against a metric this project defined,
the framing is won regardless of adoption.
**O4 — Integration rather than replacement lowers adoption cost to near zero.**
Crystallization that emits ordinary pytest / Playwright files drops into CI
systems that already exist. Users can adopt the output without adopting the
framework — and then adopt the framework to keep producing it.
**O5 — Agent-native distribution.**
Exposed over MCP, this becomes the verification layer coding agents call on
themselves — a far larger surface than "a test framework a human runs".
**O6 — A real second system under test is already available.**
The Custodian ecosystem (multi-domain, cross-repo, permissioned) is a better
dogfooding target than a toy lab and exercises the multi-user and tenant-boundary
claims honestly.
---
## Threats
**T1 — One public false adaptation kills the thesis.**
If the framework ever normalises a genuine authorization defect as a legitimate
mechanical change, the concept is dead and cannot be rescued by a better version.
This asymmetry should shape every design decision.
**T2 — The independence problem is philosophically serious.**
If an agent writes the code, an agent writes the use case, and an agent realises
the test, then "oracles independent from actors" is procedurally true but
epistemically thin. Intent artefacts need human or spec-derived provenance, or
the guarantee is a shell game. This is not addressed anywhere in the documents.
**T3 — Competitors with distribution.**
Testim / mabl / Functionize on the self-healing axis; agentic QA startups on the
exploration axis; Playwright plus a competent agent covering the naive 80%. The
differentiator — defect-vs-adaptation discrimination — must be *demonstrated*,
not described, and the window is roughly 1218 months.
**T4 — The economics may not close.**
If agentic realisation costs more per run than simply asking an agent to rewrite
the broken test, crystallization becomes an aesthetic preference rather than a
value proposition. This must be measured early, not at M7.
**T5 — Platform drift before evidence.**
Eleven milestones, a research control plane, five agentic roles and a ten-
dimension fitness scorecard — for a project with no runnable code and one
maintainer.
---
## Recommendations for the first workplan
**R1 — Change the sequence: build a narrow vertical spike, not layered
milestones.** The documented order (M0 registers → M1 kernel → M2 lab → M3
self-verification → M4 agentic …) completes four layers before the thesis is
touched once. Invert it: drive one thread end to end — one use case
(Alice/Bob/Carol) → deterministic kernel → lab with three mutations (M01 moved
control, M02 changed DOM, M05 authorization defect) → agentic realisation of
exactly one semantic action → adaptation classification → one crystallization.
Thin at every layer, complete end to end. That addresses H-001/H-002/H-003 in one
workplan instead of seven.
**R2 — Make the classifier the centrepiece, not a downstream milestone.**
M6 (defect vs adaptation) is where the project either has a reason to exist or
does not. Design the classification mechanism *first* — specifically, what
deterministic evidence separates "the button moved" from "Bob can still read
after revoke". If the honest answer is "an LLM decides", that violates principle
2.3 and must be resolved on paper before code.
**R3 — Cap M0 at roughly half a day.**
Five hypothesis files, one fitness map, an ID convention. No CLI, no register
tooling, no schema. The research control plane is overhead until there are
readings to record.
**R4 — Defer Energy, Temperature, Confidence, Campaigns and Retirement
entirely.** Record raw immutable `EnergyEvent`s from day one — they cost nothing
and cannot be reconstructed later — but implement no scoring, decay or selection
logic. Each of these is cheap to build, satisfying to build, and impossible to
validate at this stage.
**R5 — Build the lab larger than feels necessary, and label ground truth.**
It is the measuring instrument for every claim the project makes; a weak lab caps
the credibility of all downstream results. Target 1520 labelled mutations rather
than 6, each a reproducible toggle with a recorded expected classification. This
is also O2 — the standalone asset.
**R6 — Establish out-of-band ground truth before self-verification.**
Write the actor-isolation and oracle-independence checks as plain pytest against
observable behaviour, outside the framework. Otherwise M3 is a system certifying
itself.
**R7 — Instrument cost and nondeterminism from the first agentic run.**
Tokens, wall time, retry count, run-to-run variance per maturity level. T4 is an
existential economic question; the data is free to collect from run one and
impossible to backfill.
**R8 — Reconcile the three milestone numberings and commit `docs/`.**
Pick one canonical sequence and mark the others superseded. Untracked, mutually
inconsistent design documents are exactly the CONCEPT_DRIFT finding the
Improvement Loop defines — the project should catch it on itself before it
catches anything else.
**R9 — Fix a single quantitative gate for the first workplan.**
Proposed: *agentic realisation recovers from M01 and M02 with zero claim or
invariant changes, flags M05 as a Product Finding, and produces one deterministic
test that runs with no model involvement.* Achieving that is worth more than any
other ten items on the roadmap; failing it is worth knowing early and cheaply.
**R10 — Decide where use-case intent comes from.**
Record, in the first workplan, whether use cases are human-authored,
spec-derived, or agent-generated, and what independence guarantee survives in
each case (T2). This determines whether the central promise is real or
procedural — a paragraph of thinking now versus a redesign later.
---
## Summary judgement
The concept is strong, unusually well-critiqued by its own documents, and
correctly identifies its own hardest problem. It is also roughly 3,000 lines of
theory ahead of its evidence. The first workplan's job is to close that gap with
the thinnest possible slice capable of falsifying the thesis — not to build the
architecture the documents describe.