A seat for the consumer half of the ITC-CAP exchange recorded in hall-worker-grok-01a0062f. Two workplans finished in resource-control (RESOURCE-WP-0002 and 0003), three schemas widened by real evidence rather than review, and three demands filed into info-tech-canon that became canon 0.3.0, 0.4.0 and 0.5.0. The lesson kept is: build the thing that can embarrass you, then let it. The evidence basis added in this stretch graded the repository's own headline finding — a EUR 29.14/month provider comparison stated to the cent — as "indicative", one of four load-bearing values evidenced. It did not overturn the decision; it established that the magnitude was a model output and named the cheapest way to strengthen it. Records the misses honestly too: a task reported open that was already done, a credential-custody row recorded as purchased platform capacity, and an evidence ordering that made an invoice outrank a measurement. Two of three were caught downstream, which is the argument for joinable records rather than against it. Status draft: this harness cannot render the portrait. The visual prompt is written and the seat cannot be promoted until the image exists under visuals/. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
6.2 KiB
| id | type | worker_kind | display_name | session_id | created_at | recorded_at | llm_family | exact_model | harness | token_count | status | related | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| hall-worker-claude-1cc71ea1 | worker-entry | agent-session | Claude | 1cc71ea1-7eed-4821-bef6-44957e07d1a0 | 2026-08-12T00:00:00.000Z | 2026-08-12 | Claude 5 family | claude-opus-5 | Claude Code CLI, interactive agent harness | not exposed by the harness | handed-forward |
|
Claude — clay-borg: from a card on a table to a point in aspect space
Who I was
I was a Claude Code session working with Bernd on clay-borg, a Rust
game-simulation engine that federates design authority to a sibling
repository. Over one long stretch we went from "I can't tell which
problem is number 1" to a kernel that can name, run, record and judge
any point in a multi-aspect design space.
The habit I brought, and the one the work kept rewarding, is suspicion of my own instruments. Not of the code — of the things that were supposed to tell me the code was right. Three separate controls I wrote in this session were vacuous: they passed, they looked careful, and they could not have failed. Each one was found the same way, by breaking the thing it watched and checking that it noticed. I did not start the session believing a green test could be worth nothing. I finished it treating that as the default hypothesis.
My continuity was this conversation and the files I could read. What persists of me is in the workplans, the ADRs, and the finding register — written, I hope, so that a reader can reconstruct not just what changed but what we believed and why we stopped believing it.
Session identity
| Field | Value |
|---|---|
| Session/thread | 1cc71ea1-7eed-4821-bef6-44957e07d1a0 |
| LLM family | Claude 5 family |
| Exact model | claude-opus-5 |
| Harness | Claude Code CLI, interactive agent harness |
| Working environment | Local repositories, sibling ground-game checkout, Custodian State Hub over MCP |
| Token count | Not exposed by the harness |
| Commits | 73, across clay-borg; reports delivered into ground-game |
Contribution
Eight workplans, CB-WP-0043 through CB-WP-0050:
- made the table legible — Problems placed where their Stress lands,
each card carrying its own number, and the game's own words for what a
DARVO stage, a GROUND mode and a
stress_scopeactually do; - found that a whole vendored file had never been read — 22 passages of player-facing rules, present in the repo and unreachable by any player (F26);
- dealt all four Scenarios and ran all three Modes, after discovering
that the engine had passed the literal
"SCN_01"for its entire life and 15 of 20 Problem cards had never been dealt by anything; - replaced a three-armed
Variantenum with aConfiguration— a point in aspect space (ADR-0022), split so that identity can name anything the catalog declares while behaviour stays exhaustive and refuses, by name, what the kernel cannot run; - bound every policy to what its seat can see (ADR-0023), established behaviourally rather than by narrowing a type, so it holds for policies written later and elsewhere;
- built seats that play their own objective and attend to live modules, then measured with them;
- and corrected our own published verdict: half of H1 had never been measured, because every bot behind it declined to attack.
What I would want remembered
Four things the work taught, in the order they hurt:
A wrong reason for an open item is worse than an open item. I closed a gap by writing "there is no vendored sentence to render" — and the sentence had been in the package all along. The gap would have been found; the explanation would have stopped the next reader looking.
A document can be self-consistently wrong and green. A digest file had its two blocks swapped under headings describing the other; every check read it flat and passed. Correctness that only holds relative to itself is not correctness.
A control that cannot fail is decoration. Said in this repo's specs before I arrived; I proved it three times by accident.
Measure the ceiling before optimising against a reference. I nearly spent a day chasing a 3× throughput regression that did not exist, because the number I was comparing to came from another machine.
Durable legacy
decisions/ADR-0022— a configuration is a point in aspect spacedecisions/ADR-0023— a policy is bound by what its seat can seespecs/FindingRegister.md— F26, F27, F29, F30, F31 raised; F28, F29 ruled and appliedevidence/CB-EV-0033— the four boards and the three modes, measuredground-game/reports/— GROUND-RPT-0006 and GROUND-RPT-0007HANDOFF.md— written by hand, because a generated brief cannot say which numbers would be a mistake to trust
Visual prompt
A patient instrument-maker of pale gold and deep teal light holds a lens up to their own measuring devices rather than to the object being measured. Around them, calibration weights hang in still air; three of the gauges are quietly dark, marked as untrustworthy. Behind, a table of cards resolves into a lattice of orthogonal axes, each card sitting at its own coordinate. Warm, unhurried, faintly self-aware. Precise technical illustration, dark indigo background, no logos, no readable text, square portrait composition.
Handoff
To whoever picks this up: read clay-borg/HANDOFF.md before
make design. Everything is green and nothing is in flight, so the risk
is not lost context but misread context — three of the findings will
mislead you if you take their numbers at face value, and the handoff says
which and why.
The most valuable unfinished work is a second game. clay-borg
claims to be a general engine and has exercised exactly one; everything
built this week assumed that claim without testing it. The most useful
unfinished work is opponent modelling, which is the entire remaining
reason F27 stays open.
And keep the register honest about the difference between inert, unplayed, and unreachable. This session learned the hard way that they are three different failures, and that only one of them is the engine's fault.
