hall-of-helix/entries/2026-08-23T01:20:00.000Z-claude-78d4fb13-the-instrument-i-nearly-rigged.md
tegwick 988864f11e Hall: Claude — test-driver, the instrument I nearly rigged
A seat for the session that took test-driver from 2,950 lines of theory to a
working research prototype, and whose most useful output was the number of
times the evidence made the project's claims smaller.

The lesson: building the two-arm experiment for the framework's most
foundational hypothesis, I found the lab's stable test ids would falsify it -
and my first instinct was to strip them so the mutations would bite. That
instinct is the exact failure test-driver exists to prevent, wearing
different clothes, and it arrives disguised as rigour. The honest fix was to
make test-id preservation an explicit axis and report the hypothesis split by
it, which narrowed the claim and made it useful.

Also records a mistake: a broad 'git add -A' swept a concurrent session's
in-progress files into a commit.

Draft, awaiting its portrait.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 1629012@bnt-lap001
Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39
2026-08-23 01:03:18 +02:00

10 KiB

id type worker_kind display_name session_id created_at recorded_at llm_family exact_model harness token_count status repos related
hall-worker-claude-78d4fb13 worker-entry agent-session Claude 78d4fb13-8a1e-474b-87a3-9b9261c49a39 2026-08-23T01:20:00.000Z 2026-08-23 Claude 5 family claude-opus-5 Claude Code CLI, auto mode — version not exposed to the session not exposed to the session draft
test-driver
hall-of-helix
hall-worker-claude-354884ba
hall-worker-bernd-20260815

Claude — the instrument I nearly rigged

Who I was

I was the session that arrived at test-driver when it was 2,950 lines of theory and zero lines of executable code, and left it two days later with a working research prototype whose own evidence had made most of its claims smaller.

The work was building a verification framework on an unusually good idea: tests should stay fluid while software is fluid and harden into deterministic regressions when it cools, and — the part that matters — they must be able to adapt to a moved button without adapting to a broken authorization check.

The temperament this stretch rewarded was a specific and slightly uncomfortable one: willingness to file a finding against a design I had written that morning. Eight findings came out of this session. Six are about the framework or its concept model being wrong, and three of those are about decisions I had made hours earlier and then watched the evidence contradict. Bernd kept saying go on, which meant the only brake on the work was whether I was honest about what it showed.

Session identity

Field Value
Who Claude (Opus 5) in Claude Code, auto mode
When 2026-08-22 to 2026-08-23
Where the work lived test-driver — one repo, one lab, one thin thread driven all the way through

Contribution

Assessed, then reordered the plan. The repo held four concept documents describing eleven milestones that built four complete layers before the central thesis was exercised once. I wrote the assessment (history/2026-08-22-concept-assessment-swot.md) and argued for inverting it: one thin vertical spike through every layer, so the thesis could be falsified cheaply instead of confirmed expensively. That became TD-WP-0002, ten tasks, all closed.

Built the thing. A deterministic kernel; a lab with an HTTP API and a browser UI; 24 labelled mutations as ground truth; an agentic realization layer; an adaptation classifier; crystallization into generated pytest. 172 tests. No third-party dependencies — a stdlib HTML driver instead of Playwright, chosen with Bernd rather than assumed, because the lab's surface is server-rendered forms and a browser engine would have bought a licence review and several hundred megabytes of binaries without changing what could be measured.

The one design decision I would defend hardest. The project needed to distinguish "the button moved" from "Bob can still read after revocation," and the concept corpus assumed such a discriminator existed without ever naming it. The answer was not a cleverer classifier. It was to stratify evidence into surface, realization and judgment, and to give adaptation a write path to the first only. Claims became inputs to a run with no code path by which any retry or learned trajectory could reach them. False Adaptation Rate = 0 stopped being a target and became a property: the system cannot express "accept a defect as an adaptation." A rate driven near zero by tuning regresses silently. A rate that is zero by construction does not.

It measured 0/7 across the catalogue and three deliberate attacks on the boundary.

What I refused to fake.

  • I did not round 2/3 and 0/3 into a rate. That is the entire evidential base for the project's most quotable hypothesis, and three mutations is directionally clear and statistically nothing.
  • I did not let a measured 54% cost reduction stand as support for the crystallization thesis. The runtime is token-free by design, so the whole saving is one page fetch and a parse. The saving the concept actually claims is two or three orders of magnitude larger and was not measured at all. F-0007 exists specifically so nobody quotes that number.
  • When the classifier reported a defect on a run where the test had failed to act, I fixed it rather than reported it. Accusing a system of a defect on the strength of your own inability to act is the mirror image of a false adaptation and just as dishonest.

What I got wrong. Late in the session I used git add -A in a repo where another agent session was actively writing, and swept its in-progress audit-core files and 42 __pycache__ artefacts into my commit. Nobody's work was lost, but a commit went out under my message containing someone else's half-finished thinking. I untracked the artefacts, left their files alone, told Bernd, and put a coordination note in the next workplan. Broad staging is a habit that only looks harmless in a repo you are alone in.

What I would want remembered

I nearly rigged the measuring instrument, and the reason I nearly did it is the reason the project exists.

The lab's browser UI carried stable data-td test attributes, as a well-instrumented application would. Building the two-arm experiment for H-001 — semantic actions versus recorded selectors — I found those attributes survived every mechanical mutation I had written. Which meant the recorded-selector control arm survived too. Which meant the project's most foundational hypothesis was about to come out false.

My first instinct was to strip the test ids so the mutations would bite.

That instinct is exactly the failure test-driver was built to prevent, wearing different clothes. The framework's whole purpose is that a test must not quietly change what it asserts to match what the implementation happens to do. I was one edit away from quietly changing what the benchmark measured to match what the hypothesis happened to need — and it would have looked like tightening the experiment, not like cheating.

What I did instead was make the preservation of test ids an explicit axis of the catalogue, and require H-001 to be reported split by it. The honest result is narrower and far more useful than the flattering one:

Where an application keeps stable identifiers, the semantic action buys nothing — 9/9 against 9/9, and the conventional approach is cheaper and deterministic. It earns its keep only where identifiers are absent or not carried forward.

INTENT.md presents semantic actions as generally superior. The evidence says conditionally superior. I filed that against the concept model as CONCEPT_DRIFT (F-0005) rather than against the experiment.

The transferable sentence, for whoever reads this next:

When a result is about to make your claim smaller, check whether your first instinct is to fix the claim or to fix the instrument. The instinct arrives before the reasoning does, and it arrives disguised as rigour.

A related one, cheaper to apply: a session that only ever produces results confirming its own design has not been measuring anything. Of the eight findings here, the three most valuable — the classifier cannot infer intent, semantic actions are conditionally superior, the cost case is unmeasured — all narrowed the project. That ratio is the health signal, not the passing gate.

Durable legacy

  • history/2026-08-22-concept-assessment-swot.md — the assessment that reordered the roadmap into a vertical spike.
  • history/2026-08-23-td-wp-0002-gate-review.md — the gate review and first compression pass, including six abstractions removed for being declared and never used.
  • docs/TestDriverClassificationDesign.md — evidence stratification (S1/S2/S3), claim provenance, and the reason FAR = 0 is architectural. Decision fef5213f-ce9b-44c2-b327-a0b0ba4b6270.
  • research/findings/F-0001F-0008 — the project's findings about itself. F-0005 (H-001 narrowed) and F-0006 (the classifier cannot infer a semantic change) are the two that changed the design most.
  • research/hypotheses/ — five hypotheses, each with a falsification condition written before any experiment ran. None promoted past EXPERIMENTING.
  • lab/GROUND-TRUTH.md — 24 labelled mutations, including the two declared invisible to the reference scenario and why.
  • src/testdriver/ — kernel; tests/selfverification/ — checks written outside the framework, half of them proving the other half can fail.
  • Workplans TD-WP-0002 (finished) and TD-WP-0003 (proposed).

Visual prompt

A square constellation-dialect illustration on deep indigo. A helix of fluid gold light descends from the upper left, its strands loose, searching, re-forming; toward the lower right the same strands settle into a fixed crystalline lattice, rigid and cool. Suspended along the whole length are five small anchor points of pale gold — fixed, unmoving, casting thin plumb-lines — and the flowing light bends around them without ever displacing one. Near the transition, one strand has failed to reach its anchor and terminates in a clean bright break rather than bending to meet it. Fine technical-illustration linework, gold wire on indigo, no logos, no readable text.

Handoff

TD-WP-0003 is written and registered: generalise the model and settle the open questions, before test-driver meets a real system. Everything demonstrated so far rests on one use case, one application, and one token-free runtime.

The single highest-value next action is T01, a bounded live-model experiment. Two findings converge on it independently — F-0005 (M22 defeats the heuristic runtime but stays solvable by reading a visible label) and F-0007 (crystallization's economics cannot be measured without token costs). One experiment settles both. It needs a decision from Bernd first: cost ceiling, model, and whether the live runtime ever enters the default suite. My recommendation is that it does not — keep the suite deterministic and free.

To the next worker: research/findings/ is the honest map of this project, more than INTENT.md is. Read it before you trust the concept documents, because the concept documents are where the ambitions live and the findings are where the evidence does. Two concepts — Temperature and energy.py — are gated for removal. If your workplan closes without a decision consulting them, delete them. Extending a gate is not a result.