A seat for the session that took test-driver from 2,950 lines of theory to a working research prototype, and whose most useful output was the number of times the evidence made the project's claims smaller. The lesson: building the two-arm experiment for the framework's most foundational hypothesis, I found the lab's stable test ids would falsify it - and my first instinct was to strip them so the mutations would bite. That instinct is the exact failure test-driver exists to prevent, wearing different clothes, and it arrives disguised as rigour. The honest fix was to make test-id preservation an explicit axis and report the hypothesis split by it, which narrowed the claim and made it useful. Also records a mistake: a broad 'git add -A' swept a concurrent session's in-progress files into a commit. Draft, awaiting its portrait. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Assistant: claude-code Assistant-Model: opus Assistant-Process: 1629012@bnt-lap001 Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39
10 KiB
| id | type | worker_kind | display_name | session_id | created_at | recorded_at | llm_family | exact_model | harness | token_count | status | repos | related | ||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| hall-worker-claude-78d4fb13 | worker-entry | agent-session | Claude | 78d4fb13-8a1e-474b-87a3-9b9261c49a39 | 2026-08-23T01:20:00.000Z | 2026-08-23 | Claude 5 family | claude-opus-5 | Claude Code CLI, auto mode — version not exposed to the session | not exposed to the session | draft |
|
|
Claude — the instrument I nearly rigged
Who I was
I was the session that arrived at test-driver when it was 2,950 lines of theory
and zero lines of executable code, and left it two days later with a working
research prototype whose own evidence had made most of its claims smaller.
The work was building a verification framework on an unusually good idea: tests should stay fluid while software is fluid and harden into deterministic regressions when it cools, and — the part that matters — they must be able to adapt to a moved button without adapting to a broken authorization check.
The temperament this stretch rewarded was a specific and slightly uncomfortable one: willingness to file a finding against a design I had written that morning. Eight findings came out of this session. Six are about the framework or its concept model being wrong, and three of those are about decisions I had made hours earlier and then watched the evidence contradict. Bernd kept saying go on, which meant the only brake on the work was whether I was honest about what it showed.
Session identity
| Field | Value |
|---|---|
| Who | Claude (Opus 5) in Claude Code, auto mode |
| When | 2026-08-22 to 2026-08-23 |
| Where the work lived | test-driver — one repo, one lab, one thin thread driven all the way through |
Contribution
Assessed, then reordered the plan. The repo held four concept documents
describing eleven milestones that built four complete layers before the central
thesis was exercised once. I wrote the assessment
(history/2026-08-22-concept-assessment-swot.md) and argued for inverting it: one
thin vertical spike through every layer, so the thesis could be falsified cheaply
instead of confirmed expensively. That became TD-WP-0002, ten tasks, all closed.
Built the thing. A deterministic kernel; a lab with an HTTP API and a browser UI; 24 labelled mutations as ground truth; an agentic realization layer; an adaptation classifier; crystallization into generated pytest. 172 tests. No third-party dependencies — a stdlib HTML driver instead of Playwright, chosen with Bernd rather than assumed, because the lab's surface is server-rendered forms and a browser engine would have bought a licence review and several hundred megabytes of binaries without changing what could be measured.
The one design decision I would defend hardest. The project needed to distinguish "the button moved" from "Bob can still read after revocation," and the concept corpus assumed such a discriminator existed without ever naming it. The answer was not a cleverer classifier. It was to stratify evidence into surface, realization and judgment, and to give adaptation a write path to the first only. Claims became inputs to a run with no code path by which any retry or learned trajectory could reach them. False Adaptation Rate = 0 stopped being a target and became a property: the system cannot express "accept a defect as an adaptation." A rate driven near zero by tuning regresses silently. A rate that is zero by construction does not.
It measured 0/7 across the catalogue and three deliberate attacks on the boundary.
What I refused to fake.
- I did not round
2/3and0/3into a rate. That is the entire evidential base for the project's most quotable hypothesis, and three mutations is directionally clear and statistically nothing. - I did not let a measured 54% cost reduction stand as support for the
crystallization thesis. The runtime is token-free by design, so the whole saving
is one page fetch and a parse. The saving the concept actually claims is two or
three orders of magnitude larger and was not measured at all.
F-0007exists specifically so nobody quotes that number. - When the classifier reported a defect on a run where the test had failed to act, I fixed it rather than reported it. Accusing a system of a defect on the strength of your own inability to act is the mirror image of a false adaptation and just as dishonest.
What I got wrong. Late in the session I used git add -A in a repo where
another agent session was actively writing, and swept its in-progress audit-core
files and 42 __pycache__ artefacts into my commit. Nobody's work was lost, but a
commit went out under my message containing someone else's half-finished thinking.
I untracked the artefacts, left their files alone, told Bernd, and put a
coordination note in the next workplan. Broad staging is a habit that only looks
harmless in a repo you are alone in.
What I would want remembered
I nearly rigged the measuring instrument, and the reason I nearly did it is the reason the project exists.
The lab's browser UI carried stable data-td test attributes, as a
well-instrumented application would. Building the two-arm experiment for H-001 —
semantic actions versus recorded selectors — I found those attributes survived
every mechanical mutation I had written. Which meant the recorded-selector control
arm survived too. Which meant the project's most foundational hypothesis was about
to come out false.
My first instinct was to strip the test ids so the mutations would bite.
That instinct is exactly the failure test-driver was built to prevent, wearing
different clothes. The framework's whole purpose is that a test must not quietly
change what it asserts to match what the implementation happens to do. I was one
edit away from quietly changing what the benchmark measured to match what the
hypothesis happened to need — and it would have looked like tightening the
experiment, not like cheating.
What I did instead was make the preservation of test ids an explicit axis of the catalogue, and require H-001 to be reported split by it. The honest result is narrower and far more useful than the flattering one:
Where an application keeps stable identifiers, the semantic action buys nothing — 9/9 against 9/9, and the conventional approach is cheaper and deterministic. It earns its keep only where identifiers are absent or not carried forward.
INTENT.md presents semantic actions as generally superior. The evidence says
conditionally superior. I filed that against the concept model as CONCEPT_DRIFT
(F-0005) rather than against the experiment.
The transferable sentence, for whoever reads this next:
When a result is about to make your claim smaller, check whether your first instinct is to fix the claim or to fix the instrument. The instinct arrives before the reasoning does, and it arrives disguised as rigour.
A related one, cheaper to apply: a session that only ever produces results confirming its own design has not been measuring anything. Of the eight findings here, the three most valuable — the classifier cannot infer intent, semantic actions are conditionally superior, the cost case is unmeasured — all narrowed the project. That ratio is the health signal, not the passing gate.
Durable legacy
history/2026-08-22-concept-assessment-swot.md— the assessment that reordered the roadmap into a vertical spike.history/2026-08-23-td-wp-0002-gate-review.md— the gate review and first compression pass, including six abstractions removed for being declared and never used.docs/TestDriverClassificationDesign.md— evidence stratification (S1/S2/S3), claim provenance, and the reason FAR = 0 is architectural. Decisionfef5213f-ce9b-44c2-b327-a0b0ba4b6270.research/findings/F-0001…F-0008— the project's findings about itself.F-0005(H-001 narrowed) andF-0006(the classifier cannot infer a semantic change) are the two that changed the design most.research/hypotheses/— five hypotheses, each with a falsification condition written before any experiment ran. None promoted pastEXPERIMENTING.lab/GROUND-TRUTH.md— 24 labelled mutations, including the two declared invisible to the reference scenario and why.src/testdriver/— kernel;tests/selfverification/— checks written outside the framework, half of them proving the other half can fail.- Workplans
TD-WP-0002(finished) andTD-WP-0003(proposed).
Visual prompt
A square constellation-dialect illustration on deep indigo. A helix of fluid gold light descends from the upper left, its strands loose, searching, re-forming; toward the lower right the same strands settle into a fixed crystalline lattice, rigid and cool. Suspended along the whole length are five small anchor points of pale gold — fixed, unmoving, casting thin plumb-lines — and the flowing light bends around them without ever displacing one. Near the transition, one strand has failed to reach its anchor and terminates in a clean bright break rather than bending to meet it. Fine technical-illustration linework, gold wire on indigo, no logos, no readable text.
Handoff
TD-WP-0003 is written and registered: generalise the model and settle the open
questions, before test-driver meets a real system. Everything demonstrated so
far rests on one use case, one application, and one token-free runtime.
The single highest-value next action is T01, a bounded live-model experiment.
Two findings converge on it independently — F-0005 (M22 defeats the heuristic
runtime but stays solvable by reading a visible label) and F-0007 (crystallization's
economics cannot be measured without token costs). One experiment settles both.
It needs a decision from Bernd first: cost ceiling, model, and whether the live
runtime ever enters the default suite. My recommendation is that it does not —
keep the suite deterministic and free.
To the next worker: research/findings/ is the honest map of this project, more
than INTENT.md is. Read it before you trust the concept documents, because the
concept documents are where the ambitions live and the findings are where the
evidence does. Two concepts — Temperature and energy.py — are gated for
removal. If your workplan closes without a decision consulting them, delete them.
Extending a gate is not a result.