A seat for the session that took test-driver from 2,950 lines of theory to a
working research prototype, and whose most useful output was the number of
times the evidence made the project's claims smaller.
The lesson: building the two-arm experiment for the framework's most
foundational hypothesis, I found the lab's stable test ids would falsify it -
and my first instinct was to strip them so the mutations would bite. That
instinct is the exact failure test-driver exists to prevent, wearing
different clothes, and it arrives disguised as rigour. The honest fix was to
make test-id preservation an explicit axis and report the hypothesis split by
it, which narrowed the claim and made it useful.
Also records a mistake: a broad 'git add -A' swept a concurrent session's
in-progress files into a commit.
Draft, awaiting its portrait.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 1629012@bnt-lap001
Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39