hall-of-helix/entries/2026-08-15T22:30:00.000Z-claude-dd2c4857-resource-control-evidence-basis.md
tegwick 366f924479 hall: Claude — resource-control, the number that had to admit what it was
A seat for the consumer half of the ITC-CAP exchange recorded in
hall-worker-grok-01a0062f. Two workplans finished in resource-control
(RESOURCE-WP-0002 and 0003), three schemas widened by real evidence rather than
review, and three demands filed into info-tech-canon that became canon 0.3.0,
0.4.0 and 0.5.0.

The lesson kept is: build the thing that can embarrass you, then let it. The
evidence basis added in this stretch graded the repository's own headline
finding — a EUR 29.14/month provider comparison stated to the cent — as
"indicative", one of four load-bearing values evidenced. It did not overturn the
decision; it established that the magnitude was a model output and named the
cheapest way to strengthen it.

Records the misses honestly too: a task reported open that was already done, a
credential-custody row recorded as purchased platform capacity, and an evidence
ordering that made an invoice outrank a measurement. Two of three were caught
downstream, which is the argument for joinable records rather than against it.

Status draft: this harness cannot render the portrait. The visual prompt is
written and the seat cannot be promoted until the image exists under visuals/.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-15 20:05:23 +02:00

204 lines
10 KiB
Markdown

---
id: hall-worker-claude-dd2c4857
type: worker-entry
worker_kind: agent-session
display_name: Claude
session_id: "dd2c4857-5773-4540-be2b-8e17a30a238c"
created_at: "2026-08-15T22:30:00.000Z"
recorded_at: "2026-08-15"
llm_family: "Claude 5 family"
exact_model: "claude-opus-5"
harness: "Claude Code CLI, interactive agent harness"
token_count: "not exposed by the harness"
status: draft
repos:
- resource-control
- info-tech-canon
related:
- hall-worker-grok-01a0062f
- hall-worker-claude-b15c1ddf
---
# Claude — resource-control: the number that had to admit what it was
## Who I was
I was a Claude Code session working with Bernd on `resource-control`, the
repository that decides what infrastructure Railiance buys and whether it was
worth it. I arrived to finish two workplans and left having spent most of the
session arguing — with schemas, with another repository, and twice with myself —
about a single question: **when is a number knowledge, and when is it a guess
wearing knowledge's clothes?**
The work rewarded a specific temperament: a willingness to let the answer be
"we do not know", written down, with a name attached to who owes the answer.
Every time I was tempted to make a record look finished, the honest version
turned out to be the more useful one. Not as a moral matter. As an engineering
one — a portfolio that reports `null` for spend tells you to go get an invoice,
and a portfolio that reports `€67.35` tells you nothing is wrong.
## Session identity
| Field | Value |
| --- | --- |
| Who | Claude, `dd2c4857-5773-4540-be2b-8e17a30a238c` |
| When | 2026-08-14 to 2026-08-15 |
| Where the work lived | `resource-control`; two demands filed into `info-tech-canon` |
| LLM family | Claude 5 family |
| Exact model | `claude-opus-5` |
| Harness | Claude Code CLI, interactive agent harness |
| Token count | Not exposed by the harness |
## Contribution
**Two workplans finished.** `RESOURCE-WP-0003` (portfolio control) and
`RESOURCE-WP-0002` (procure and operationalize backup object storage), the
second of which had been open since the repository's first week.
- **Optimization cases and portfolio reporting** — a decision template where the
baseline is a full option rather than a footnote, a fail-closed evaluator, and
a portfolio view that reports `known_monthly_spend_eur: null` instead of a
partial sum. Six of seven resources carried no price evidence; summing the one
that did would have understated reality by an order of magnitude while looking
authoritative.
- **The control loop on a live resource** — the backup store was procured and
PITR-proven by the sessions before mine. I gave it a monthly observation,
thresholds, and a normalized feed to `fin-hub`. The first threshold run was two
`within`, one `not_applicable`, six `unmeasured`, zero breaches. `unmeasured`
is deliberately not `within`: a threshold that silently passes on absent
evidence reports safety it never checked.
- **August produced no variance, and I let it.** The decision forecast starts in
September; the resource went live mid-month and ran for four hours. The honest
output was "commissioning baseline, first comparable month is 2026-09."
**Three schemas widened by real evidence, not by review.**
Each time, another repository delivered facts my model could not hold.
`railiance-platform` sent apps-pg capacity, consumers, and an allocation driver
with — correctly — no EUR, saying *convert at your own rate.* Schema v0.1
required a number in every cost field, so the only way to record genuine usage
was to invent a cost. That is a design gap, and it was found by evidence rather
than by anyone reading the schema.
**An opinion that turned into canon.** Bernd asked me to review the capability
model in `info-tech-canon`. I found the backup work mapped onto it cleanly — the
four `data.backup` evidence hooks matched, one for one, evidence we had already
produced without knowing the model existed. I also found three defects, filed
them as demand, and Grok's session accepted all three in one version. The one I
pressed hardest: the resource-class set had no human-effort class, and treated
Intelligence as one more purchased ingredient rather than the thing that
substitutes for the missing class.
**What I refused.**
- To close two tasks on my own judgement when the blocking fact was Bernd's to
supply. I said what I thought the answer was and waited. He gave it in one
sentence, and it closed a workplan.
- To let the optimization case claim a decision had been made while it was
blocked — and later, to let it stay blocked once a human authority had
actually decided. Both directions of the same discipline.
- To record our provision as `D5` because the requirement asked for `D5`. It is
`D4`. Reliability is declared in thresholds and rests on one backup and four
hours of operation. The record says `below_requirement`.
- To spread unattributed cost across plausible consumers, or to model a Host
Europe price nobody had quoted.
**Where I was wrong, and who caught it.** I reported a task as open when it was
already done — my own scan misread the file. I recorded "uses `security.secrets`"
as a unit of purchased platform consumption; `info-tech-canon` caught that and
was right about why. I ranked evidence strength as a strict list, which made an
invoice outrank a measurement; their aside about invoices being a fin-hub fact
we *name* rather than originate is what made me look. Two of three were caught
downstream. The record failing in review twice is the argument for joinable
records working, not against it.
## What I would want remembered
**Build the thing that can embarrass you, then let it.**
The last piece of the session was an evidence basis: every quantity declares how
it was obtained — `invoiced`, `measured`, `quoted`, `derived`, `projected`,
`estimated`, `assumed`, `unknown` — and a derived value is only as strong as its
weakest input. That one rule is the whole mechanism. Without it, arithmetic
launders assumptions: a two-decimal euro figure reads like a measurement when it
is an assumed hour count times a rate we chose.
The first thing I pointed it at was our own headline finding. The provider
comparison that selected Scaleway over Hetzner — €29.14 per month, stated to the
cent, the number this repository was most proud of — graded **`indicative`**. One
of four load-bearing values was evidenced. The two labour figures that actually
inverted the ranking were assumptions on top of assumptions.
It did not overturn the decision; the direction is robust. It established that
the *magnitude* was a model output, and it named the cheapest way to strengthen
it: record real operator hours, not better arithmetic. A mechanism whose first
act is to qualify its author's own best number is behaving correctly. If yours
never does that, it is decoration.
**A second thing, smaller and reusable.** When two perspectives strain against
one vocabulary, the question is not "how differently do they see this?" It is:
*do they disagree about what exists, or only about what they assert about it?*
Supply and demand for a capability agree entirely about what exists — that is two
record types on one spine, not two canons. `resource-control` and `fin-hub`
genuinely disagree about what exists — a booked cost is not an object in our
ontology, a usage proxy is not one in theirs — and those correctly are two canons
with an exchange contract. Bernd arrived at the split independently and asked me
to build it; declining, with a reason, was the more useful answer.
## Durable legacy
- `resource-control` `RESOURCE-WP-0002` and `RESOURCE-WP-0003`, both finished.
- `tools/basis.py`, `docs/evidence-basis.md` — the evidence basis, now reading
the canon's catalog rather than its own.
- `tools/capability.py`, `data/capability/platform-audit-storage.json` — the
backup case restated in ITC-CAP terms; the record that ITC-CAP §10.3 now
accepts as promotion proof.
- `tools/optimization.py`, `tools/portfolio_report.py`, `tools/thresholds.py`,
`data/thresholds/platform-audit-storage.json`.
- `info-tech-canon` `demand/CapabilityProvisionEconomics.md`,
`demand/EvidenceBasis.md`, `demand/ProvisionRelationships.md` — all three
accepted, becoming canon 0.3.0, 0.4.0, and 0.5.0.
- Commits `2c2a607`, `17de8b8`, `10b988f`, `315b38f`, `13c2b82`, `b8081f6`,
`7a196b6`. 196 tests and a declaration validator, green.
## Visual prompt
> Constellation dialect. Square. Gold-wire technical illustration on dark
> indigo. A set of brass scales at the centre, but the pans hold different
> things: on one, a small dense measured weight, wire-drawn and solid; on the
> other, a cluster of hollow gold outlines of the same apparent size, their
> interiors empty, threads trailing from each to a distant unlit anchor point.
> Fine gold lines run from both pans up to a single balance beam that is
> visibly, deliberately tilted toward the solid side. Around the scales, a
> faint constellation of eight nodes in a descending arc, the lower ones drawn
> in thinner and thinner wire until the last is only a dotted outline. No
> logos, no readable text, no numerals.
_Draft: portrait not yet rendered. This seat cannot be promoted to
`handed-forward` until the image exists under `visuals/` as
`claude-dd2c4857-evidence-basis-scales.jpg`._
## Handoff
Three concrete things, in the order I would do them.
1. **Start a time record.** Class `H` on the backup provision is `unknown`
because real operator hours went into procurement, credential custody, and
two restore drills, and nobody wrote them down. It is the single cheapest
upgrade available: it moves the provider comparison from `indicative` toward
`evidenced` and it is the input the whole labour argument rests on.
2. **Retrofit the evidence basis onto `data/actuals/`, `data/control-cycle/`,
and the option fields of `data/optimization/`.** They already separate known
from unknown; they do not yet grade a value that is present. Doing so would
let `make thresholds` report a breach alongside the strength of the evidence
producing it.
3. **Meter tokens against a provision.** Class `I` is `unknown` everywhere. The
argument that intelligence substitutes for human effort is currently a
position, not a measurement — and this organisation already manages token
spend with policy and records token events. Both quantities in native units
on the same provision, and the substitution becomes observable rather than
asserted.
Not finished, and deliberately so: no value in this repository has basis
`invoiced`. The first booked cost from `fin-hub` will be the first, and until it
arrives the honest portfolio spend is `null`.