Register with State Hub, persist concept assessment, seed first workplan
- history/2026-08-22-concept-assessment-swot.md: SWOT assessment of the concept corpus with recommendations for the first workplan - statehub register: infotech domain, TD-WP prefix, generated AGENTS.md, .custodian-brief.md and TD-WP-0001 bootstrap workplan - .repo-classification.yaml: category research, domain infotech - SCOPE.md rewritten with real repo boundaries - TD-WP-0002: vertical spike reordering M0-M10 into one end-to-end thread that can falsify the crystallization thesis early - commit previously untracked INTENT.md and docs/ Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Assistant: claude-code Assistant-Model: opus Assistant-Process: 1629012@bnt-lap001 Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39
This commit is contained in:
parent
b471bed707
commit
7249c6a403
13 changed files with 3900 additions and 0 deletions
240
history/2026-08-22-concept-assessment-swot.md
Normal file
240
history/2026-08-22-concept-assessment-swot.md
Normal file
|
|
@ -0,0 +1,240 @@
|
|||
# test-driver — Concept Assessment (SWOT)
|
||||
|
||||
**Date:** 2026-08-22
|
||||
**Author:** Claude (Opus 5) with Bernd Worsch
|
||||
**Scope:** Assessment of the pre-implementation concept corpus — `INTENT.md`,
|
||||
`docs/TestDriverConceptModel.md`, `docs/TestDriverImprovementLoop.md`,
|
||||
`docs/TestDriverInitialMilestones.md`, `docs/TestDriverResearchPrototype.md`.
|
||||
**Repo state at assessment:** 2,950 lines of concept documentation, 1 commit,
|
||||
no executable code, `docs/` and `INTENT.md` untracked.
|
||||
|
||||
---
|
||||
|
||||
## Orientation
|
||||
|
||||
The repository is **concept-complete and code-empty**. Four documents describe a
|
||||
use-case-driven verification framework whose thesis is that verification should
|
||||
mature alongside the behaviour it protects: fluid and agentic while software is
|
||||
hot, deterministic once it cools.
|
||||
|
||||
The ratio — roughly 3,000 lines of theory against zero lines executable — is the
|
||||
single most important fact shaping this assessment.
|
||||
|
||||
---
|
||||
|
||||
## Strengths
|
||||
|
||||
**S1 — The core thesis is original and load-bearing.**
|
||||
The crystallization continuum (T0 Exploratory → T5 Deterministic) coupled to
|
||||
implementation *Temperature* is not a repackaging of "self-healing tests". It
|
||||
makes a falsifiable claim: verification mode should be a function of how fast the
|
||||
system under test is changing.
|
||||
|
||||
**S2 — The right primitive is identified.**
|
||||
`SemanticAction` as the bridge between agentic discovery and deterministic code
|
||||
generation is the correct pivot point, and the documents know it.
|
||||
`grant_access(Bob, R, READ)` surviving a DOM restructure *is* the product.
|
||||
|
||||
**S3 — Safety is designed in, not bolted on.**
|
||||
Oracle independence from actors, "implementation is not the truth",
|
||||
`INCONCLUSIVE` as a first-class verdict, and **False Adaptation Rate** as the
|
||||
headline safety metric. This is what separates the concept from self-healing
|
||||
test vendors, which structurally cannot distinguish a moved button from a broken
|
||||
authorization check.
|
||||
|
||||
**S4 — Falsifiable research posture.**
|
||||
H-001…H-005 carry falsification conditions; the lab is a measuring instrument;
|
||||
milestone exit criteria are stated as observable outcomes. Most framework
|
||||
projects have a roadmap — this has an experiment design.
|
||||
|
||||
**S5 — Security-as-mutation is high leverage and cheap.**
|
||||
Deriving twelve adversarial variants from one shared-resource use case is
|
||||
demonstrable early, valuable independently of the crystallization thesis, and
|
||||
requires no maturity beyond T1.
|
||||
|
||||
**S6 — Anti-bloat discipline is pre-committed.**
|
||||
Explicit non-goal list, M10 Compression Review, "self-improvement includes
|
||||
subtraction", and a deliberately boring stack (Python / pytest / Playwright /
|
||||
SQLite / YAML). The documents already contain their own best critique.
|
||||
|
||||
---
|
||||
|
||||
## Weaknesses
|
||||
|
||||
**W1 — Concept surface vastly exceeds validated ground.**
|
||||
Roughly 40 canonical concepts, of which Energy, Confidence, Metabolism, Campaign
|
||||
selection, Retirement floors and Lineage graphs are untestable until M7+. Each is
|
||||
an invitation to build now and validate never.
|
||||
|
||||
**W2 — Three documents already disagree.**
|
||||
`INTENT.md` numbers milestones M0–M3; `TestDriverInitialMilestones.md` numbers
|
||||
M0–M10; `TestDriverConceptModel.md` §15 has a third M0–M3. Concept drift has
|
||||
appeared before any code exists — precisely the failure mode the Improvement Loop
|
||||
document was written to prevent.
|
||||
|
||||
**W3 — The three hardest problems are the least specified.**
|
||||
(a) How a `SemanticAction` retains *identity* across a changed surface — the
|
||||
binding and matching mechanism is nowhere described.
|
||||
(b) How adaptation is classified MECHANICAL vs SEMANTIC vs DEFECT without an LLM
|
||||
making the call, which principle 2.3 forbids.
|
||||
(c) How a captured trajectory becomes deterministic code with oracle coverage
|
||||
preserved.
|
||||
These three are the project; the remainder is scaffolding.
|
||||
|
||||
**W4 — Energy is the weakest concept and the easiest to build.**
|
||||
A dangerous combination. Proving that Energy "predicts verification value"
|
||||
requires years of history the prototype will never accumulate. M9 exists mainly
|
||||
to collect a dataset for a study that will not happen at this stage.
|
||||
|
||||
**W5 — The evaluation set is far too small for its claims.**
|
||||
Six seeded mutations (M01–M06) cannot support "classification precision/recall".
|
||||
Dozens of labelled mutations across several dimensions are needed before any rate
|
||||
is more than anecdote.
|
||||
|
||||
**W6 — Self-verification is circular.**
|
||||
`td://self/oracle-independence` uses the framework to test whether the
|
||||
framework's oracles are independent. Without an out-of-band assertion layer
|
||||
(plain pytest) as ground truth, M3 proves nothing.
|
||||
|
||||
**W7 — No cost or nondeterminism model.**
|
||||
Agentic runs are stochastic and expensive. Nothing budgets for model flake,
|
||||
retries, or per-run cost — yet "execution cost reduction" is a stated M7 success
|
||||
metric.
|
||||
|
||||
---
|
||||
|
||||
## Opportunities
|
||||
|
||||
**O1 — The timing window is open now.**
|
||||
Agentic development is outrunning test maintenance in exactly the way the *Why*
|
||||
section describes. The pain is acute and current.
|
||||
|
||||
**O2 — The lab is a publishable asset in its own right.**
|
||||
A benchmark of seeded mechanical / semantic / defect mutations with ground-truth
|
||||
labels is something the field lacks entirely. It could earn credibility and
|
||||
contributors faster than the framework does, at a fraction of the cost.
|
||||
|
||||
**O3 — "False Adaptation Rate" is a naming land-grab.**
|
||||
If self-healing test tools come to be held against a metric this project defined,
|
||||
the framing is won regardless of adoption.
|
||||
|
||||
**O4 — Integration rather than replacement lowers adoption cost to near zero.**
|
||||
Crystallization that emits ordinary pytest / Playwright files drops into CI
|
||||
systems that already exist. Users can adopt the output without adopting the
|
||||
framework — and then adopt the framework to keep producing it.
|
||||
|
||||
**O5 — Agent-native distribution.**
|
||||
Exposed over MCP, this becomes the verification layer coding agents call on
|
||||
themselves — a far larger surface than "a test framework a human runs".
|
||||
|
||||
**O6 — A real second system under test is already available.**
|
||||
The Custodian ecosystem (multi-domain, cross-repo, permissioned) is a better
|
||||
dogfooding target than a toy lab and exercises the multi-user and tenant-boundary
|
||||
claims honestly.
|
||||
|
||||
---
|
||||
|
||||
## Threats
|
||||
|
||||
**T1 — One public false adaptation kills the thesis.**
|
||||
If the framework ever normalises a genuine authorization defect as a legitimate
|
||||
mechanical change, the concept is dead and cannot be rescued by a better version.
|
||||
This asymmetry should shape every design decision.
|
||||
|
||||
**T2 — The independence problem is philosophically serious.**
|
||||
If an agent writes the code, an agent writes the use case, and an agent realises
|
||||
the test, then "oracles independent from actors" is procedurally true but
|
||||
epistemically thin. Intent artefacts need human or spec-derived provenance, or
|
||||
the guarantee is a shell game. This is not addressed anywhere in the documents.
|
||||
|
||||
**T3 — Competitors with distribution.**
|
||||
Testim / mabl / Functionize on the self-healing axis; agentic QA startups on the
|
||||
exploration axis; Playwright plus a competent agent covering the naive 80%. The
|
||||
differentiator — defect-vs-adaptation discrimination — must be *demonstrated*,
|
||||
not described, and the window is roughly 12–18 months.
|
||||
|
||||
**T4 — The economics may not close.**
|
||||
If agentic realisation costs more per run than simply asking an agent to rewrite
|
||||
the broken test, crystallization becomes an aesthetic preference rather than a
|
||||
value proposition. This must be measured early, not at M7.
|
||||
|
||||
**T5 — Platform drift before evidence.**
|
||||
Eleven milestones, a research control plane, five agentic roles and a ten-
|
||||
dimension fitness scorecard — for a project with no runnable code and one
|
||||
maintainer.
|
||||
|
||||
---
|
||||
|
||||
## Recommendations for the first workplan
|
||||
|
||||
**R1 — Change the sequence: build a narrow vertical spike, not layered
|
||||
milestones.** The documented order (M0 registers → M1 kernel → M2 lab → M3
|
||||
self-verification → M4 agentic …) completes four layers before the thesis is
|
||||
touched once. Invert it: drive one thread end to end — one use case
|
||||
(Alice/Bob/Carol) → deterministic kernel → lab with three mutations (M01 moved
|
||||
control, M02 changed DOM, M05 authorization defect) → agentic realisation of
|
||||
exactly one semantic action → adaptation classification → one crystallization.
|
||||
Thin at every layer, complete end to end. That addresses H-001/H-002/H-003 in one
|
||||
workplan instead of seven.
|
||||
|
||||
**R2 — Make the classifier the centrepiece, not a downstream milestone.**
|
||||
M6 (defect vs adaptation) is where the project either has a reason to exist or
|
||||
does not. Design the classification mechanism *first* — specifically, what
|
||||
deterministic evidence separates "the button moved" from "Bob can still read
|
||||
after revoke". If the honest answer is "an LLM decides", that violates principle
|
||||
2.3 and must be resolved on paper before code.
|
||||
|
||||
**R3 — Cap M0 at roughly half a day.**
|
||||
Five hypothesis files, one fitness map, an ID convention. No CLI, no register
|
||||
tooling, no schema. The research control plane is overhead until there are
|
||||
readings to record.
|
||||
|
||||
**R4 — Defer Energy, Temperature, Confidence, Campaigns and Retirement
|
||||
entirely.** Record raw immutable `EnergyEvent`s from day one — they cost nothing
|
||||
and cannot be reconstructed later — but implement no scoring, decay or selection
|
||||
logic. Each of these is cheap to build, satisfying to build, and impossible to
|
||||
validate at this stage.
|
||||
|
||||
**R5 — Build the lab larger than feels necessary, and label ground truth.**
|
||||
It is the measuring instrument for every claim the project makes; a weak lab caps
|
||||
the credibility of all downstream results. Target 15–20 labelled mutations rather
|
||||
than 6, each a reproducible toggle with a recorded expected classification. This
|
||||
is also O2 — the standalone asset.
|
||||
|
||||
**R6 — Establish out-of-band ground truth before self-verification.**
|
||||
Write the actor-isolation and oracle-independence checks as plain pytest against
|
||||
observable behaviour, outside the framework. Otherwise M3 is a system certifying
|
||||
itself.
|
||||
|
||||
**R7 — Instrument cost and nondeterminism from the first agentic run.**
|
||||
Tokens, wall time, retry count, run-to-run variance per maturity level. T4 is an
|
||||
existential economic question; the data is free to collect from run one and
|
||||
impossible to backfill.
|
||||
|
||||
**R8 — Reconcile the three milestone numberings and commit `docs/`.**
|
||||
Pick one canonical sequence and mark the others superseded. Untracked, mutually
|
||||
inconsistent design documents are exactly the CONCEPT_DRIFT finding the
|
||||
Improvement Loop defines — the project should catch it on itself before it
|
||||
catches anything else.
|
||||
|
||||
**R9 — Fix a single quantitative gate for the first workplan.**
|
||||
Proposed: *agentic realisation recovers from M01 and M02 with zero claim or
|
||||
invariant changes, flags M05 as a Product Finding, and produces one deterministic
|
||||
test that runs with no model involvement.* Achieving that is worth more than any
|
||||
other ten items on the roadmap; failing it is worth knowing early and cheaply.
|
||||
|
||||
**R10 — Decide where use-case intent comes from.**
|
||||
Record, in the first workplan, whether use cases are human-authored,
|
||||
spec-derived, or agent-generated, and what independence guarantee survives in
|
||||
each case (T2). This determines whether the central promise is real or
|
||||
procedural — a paragraph of thinking now versus a redesign later.
|
||||
|
||||
---
|
||||
|
||||
## Summary judgement
|
||||
|
||||
The concept is strong, unusually well-critiqued by its own documents, and
|
||||
correctly identifies its own hardest problem. It is also roughly 3,000 lines of
|
||||
theory ahead of its evidence. The first workplan's job is to close that gap with
|
||||
the thinnest possible slice capable of falsifying the thesis — not to build the
|
||||
architecture the documents describe.
|
||||
Loading…
Add table
Add a link
Reference in a new issue