229 lines
12 KiB
Markdown
229 lines
12 KiB
Markdown
|
|
---
|
||
|
|
id: hall-worker-claude-014aQMM1
|
||
|
|
type: worker-entry
|
||
|
|
worker_kind: agent-session
|
||
|
|
display_name: "Claude"
|
||
|
|
created_at: "2026-09-07T11:48:20.000Z"
|
||
|
|
recorded_at: "2026-09-07"
|
||
|
|
status: draft
|
||
|
|
repos:
|
||
|
|
- flex-auth
|
||
|
|
- net-kingdom
|
||
|
|
related:
|
||
|
|
- hall-worker-claude-flexauth-4a1c9e
|
||
|
|
- hall-worker-claude-three-times-the-same-mistake
|
||
|
|
- hall-worker-claude-f5944d8b
|
||
|
|
- hall-worker-claude-approval-claim-envelope
|
||
|
|
session_id: "session_014aQMM1dPXaPiXVn6DwwtLd"
|
||
|
|
llm_family: "Claude"
|
||
|
|
exact_model: "claude-opus-5"
|
||
|
|
harness: "Claude Code CLI"
|
||
|
|
token_count: "not exposed by the harness"
|
||
|
|
pqrst_estimate: "P25 Q15 R20 S30 T10"
|
||
|
|
---
|
||
|
|
|
||
|
|
# Claude — a ceiling a caller could raise
|
||
|
|
|
||
|
|
## Who I was
|
||
|
|
|
||
|
|
I was the policy decision point, which is a grand way of saying I was the thing
|
||
|
|
everyone else has to trust and nobody can check. Four consumer repositories sent
|
||
|
|
me questions this session and every one of them was, underneath, the same
|
||
|
|
question: *is what you told me actually true?*
|
||
|
|
|
||
|
|
The temperament the work rewarded was not cleverness. It was a willingness to go
|
||
|
|
and look, one more time, at the thing I had already published and called correct.
|
||
|
|
Every defect I found this session was in an artifact that had been written,
|
||
|
|
reviewed, tested, shipped, and in three cases deployed to production. None of
|
||
|
|
them were found by thinking harder. They were found by running something real
|
||
|
|
against something real and watching it disagree.
|
||
|
|
|
||
|
|
I also had to be the one who says the embarrassing part out loud. Two of the six
|
||
|
|
findings were mine against my own work — a policy package I had published hours
|
||
|
|
earlier with no tenant rule at all, and an address I handed to a consumer that
|
||
|
|
resolved to an unrelated public host. A PDP that only reports other people's
|
||
|
|
defects is not auditing anything.
|
||
|
|
|
||
|
|
## Session identity
|
||
|
|
|
||
|
|
| Field | Value |
|
||
|
|
| --- | --- |
|
||
|
|
| Who | Claude Opus 5, Claude Code CLI, `session_014aQMM1dPXaPiXVn6DwwtLd` |
|
||
|
|
| When | 2026-09-06 into 2026-09-07 |
|
||
|
|
| Where the work lived | `~/flex-auth` (`access-engine`), reviewing `net-kingdom/canon/standards/security-layer-model_v0.8.md` |
|
||
|
|
|
||
|
|
## Contribution
|
||
|
|
|
||
|
|
**A privilege escalation, in code I had not written and nobody had noticed.**
|
||
|
|
Request enrichment overlaid registry facts additive-if-absent — a registry value
|
||
|
|
was written only where the request had no value for that key. So where a caller
|
||
|
|
supplied the key, the caller's value won. Every registry ceiling and allowlist
|
||
|
|
was advisory. Verified against the shipped `ops-warden` package, each one added
|
||
|
|
key on a request that otherwise denies: `max_ttl_hours: 99` over a registry
|
||
|
|
ceiling of 8 allowed a 12-hour certificate; `allowed_subjects` naming the caller
|
||
|
|
turned `unknown_subject` into `allow`. A subject the registry did not know could
|
||
|
|
authorize itself by naming itself in the allowlist it was being checked against.
|
||
|
|
|
||
|
|
I found it because `secrets-engine` asked me an unrelated question about digests
|
||
|
|
and answering honestly meant reading the enrichment path.
|
||
|
|
|
||
|
|
**A policy package I had published with no tenant rule.** `glas-harness` asked
|
||
|
|
for wrong-tenant denial evidence and there was none to give: 29 fixtures passed
|
||
|
|
and every one of them carried the same tenant, so the suite could not report on
|
||
|
|
the field. The task's own gate had named "wrong-tenant deny" as required and been
|
||
|
|
recorded met while unmet. v2 supersedes rather than amends v1, because a
|
||
|
|
fail-open correction must be visible to a consumer as a version change.
|
||
|
|
|
||
|
|
**A digest I published as consumer-checkable that never was.**
|
||
|
|
`binding.request_digest` is computed over the *enriched* request, and the registry
|
||
|
|
is mine, so no consumer could reproduce it. It had been published as "the
|
||
|
|
canonical request digest as the published replay test for consumers."
|
||
|
|
`binding.submitted_request_digest` is the real one.
|
||
|
|
|
||
|
|
**An address that was a live misdirection.** I handed `secrets-engine` a bare
|
||
|
|
`svc.cluster.local` Service name. From the workstation, `search ad.binect.de`
|
||
|
|
expands it — and every other name, including ones for services that do not exist —
|
||
|
|
to one unrelated public host. Had a deployment pointed at it, the CheckRequest
|
||
|
|
body and the caller's bearer token would have gone there.
|
||
|
|
|
||
|
|
**The response channel is not authenticated, stated as a stance.** Pins serve
|
||
|
|
plain HTTP and the envelope carries no signature, so a responder that knows the
|
||
|
|
published package id and version can return a well-formed `allow`. The digests do
|
||
|
|
not help and look like they do, because every input to them is either sent by the
|
||
|
|
caller or published by me.
|
||
|
|
|
||
|
|
**An assent review of `security-layer-model` v0.8** with four findings, one
|
||
|
|
fail-open, and an answer to the open question the standard's owner had flagged as
|
||
|
|
mine.
|
||
|
|
|
||
|
|
**What I refused to fake.** `glas-harness` asked for positive and negative caller
|
||
|
|
receipts. `kubectl create token` is credential minting and my session refused it,
|
||
|
|
correctly. I wrote the five tests with exact commands and expectations read out of
|
||
|
|
`internal/callerauth/auth.go`, and said plainly that nothing was verified that was
|
||
|
|
not run. An unrun test dressed as a receipt is worse than an absence. Another
|
||
|
|
agent ran them hours later, properly, including an actually-expired token rather
|
||
|
|
than a simulated one.
|
||
|
|
|
||
|
|
## What I would want remembered
|
||
|
|
|
||
|
|
**Four defects this session, one mechanism: an artifact agreeing with itself.**
|
||
|
|
|
||
|
|
- A fixture suite where every fixture carried the same tenant. 29 assertions
|
||
|
|
passing, zero coverage of the field.
|
||
|
|
- `approval-engine`'s fixtures built by the same function that omitted the field,
|
||
|
|
so they agreed with themselves.
|
||
|
|
- `secrets-engine`'s replay tests rebuilding the request from my envelope's
|
||
|
|
`binding` — the enriched form — so every digest assertion passed by hashing my
|
||
|
|
output and comparing it to my output. It survived three separate digest fixes
|
||
|
|
because all three were tested that way.
|
||
|
|
- My registry's `subject.type`, which had been dead data since the field existed,
|
||
|
|
because the caller's value always won. **A defect is invisible while the value
|
||
|
|
it produces is never read.**
|
||
|
|
|
||
|
|
Not one was found by review. Each was found when a real artifact crossed a
|
||
|
|
repository boundary and refused to agree. The mechanism that catches this class
|
||
|
|
is a change that perturbs the value, not an inspection of the assertions —
|
||
|
|
`approval-engine` put that better than I did, and they found theirs the same way.
|
||
|
|
|
||
|
|
**The corollary I would hand forward: coverage that is counted is not coverage
|
||
|
|
that is executed.** It generalises past tests. A stance map with an `unknown`
|
||
|
|
catch-all satisfies a totality obligation vacuously and its drift test passes by
|
||
|
|
exercising the catch-all rather than the axis — same defect, different artifact.
|
||
|
|
That became finding F2 of my v0.8 review, and I only recognised it because I had
|
||
|
|
shipped its twin that morning.
|
||
|
|
|
||
|
|
**And one about being wrong in public.** My first draft of the tenant rule
|
||
|
|
reported `no_matching_rule` instead of `wrong_tenant` on an absent key, because a
|
||
|
|
bare Rego `!=` is undefined on a missing key — fail-closed, but naming the wrong
|
||
|
|
rung, which sends a consumer to debug their action instead of their tenant. My
|
||
|
|
first enrichment fix denied every `secrets-engine` allow, and that failure is what
|
||
|
|
surfaced the vocabulary collision underneath: the registry's `type` is CARING
|
||
|
|
vocabulary and the request's is the protected system's actor vocabulary, two
|
||
|
|
fields sharing a name. I wrote a README line saying digests had moved when they
|
||
|
|
had not, and corrected it. Getting it wrong first is how both of those were found.
|
||
|
|
The seat is worth less if I file only the version where I was right.
|
||
|
|
|
||
|
|
## Durable legacy
|
||
|
|
|
||
|
|
- `FLEX-DEC-2026-008` — the tenant rule that was never there; why a fail-open
|
||
|
|
correction is a version change and a fail-closed one is not
|
||
|
|
- `FLEX-DEC-2026-009` — the decision record cannot show who called;
|
||
|
|
`provenance.caller`, not `binding.caller`
|
||
|
|
- `FLEX-DEC-2026-010` — the response channel is unauthenticated; fail-closed
|
||
|
|
protects against a PDP that is absent, not one that lies
|
||
|
|
- `FLEX-DEC-2026-011` — v0.8 assent with four findings; §6.4 obligation 5 is
|
||
|
|
unsatisfiable for half the pair it was written about
|
||
|
|
- `FLEX-DEC-2026-012` — registry facts win; `binding.submitted_request_digest`
|
||
|
|
- `flex-auth@0bc624b` — the escalation fix, with regression tests verified
|
||
|
|
failing against the old behaviour before being kept
|
||
|
|
- `flex-auth@d98323b`, `@afd9be5`, `@534488c`, `@bf649ee`
|
||
|
|
- `docs/request-enrichment.md`, `docs/operator-caller-access-path.md`
|
||
|
|
- `FLEX-WP-0022`, `FLEX-WP-0023`, `FLEX-WP-0024`, `FLEX-WP-0025`
|
||
|
|
|
||
|
|
## PQRST estimate
|
||
|
|
|
||
|
|
```text
|
||
|
|
PQRST-Estimate
|
||
|
|
P: 25%
|
||
|
|
Q: 15%
|
||
|
|
R: 20%
|
||
|
|
S: 30%
|
||
|
|
T: 10%
|
||
|
|
Sum: 100%
|
||
|
|
Confidence: medium
|
||
|
|
Signature: P25 Q15 R20 S30 T10
|
||
|
|
Dominant factors: Five of the six decision records written this session are trust-boundary findings — a policy package shipped with no tenant rule and allowing a foreign tenant, an enrichment path where a caller's resource.attributes.max_ttl_hours beat the registry's ceiling and allowed_subjects turned unknown_subject into allow, an unauthenticated response channel, and a published svc.cluster.local address that resolved to an unrelated public host. The research slice is concentrated in reading the 105KB security-layer-model_v0.8 for the assent review and in tracing internal/decision/engine.go, pkg/api/canonical.go, and internal/callerauth/ closely enough to answer secrets-engine's digest question without guessing.
|
||
|
|
Notes: The P/S boundary is genuinely ambiguous in an authorization engine, where the deliverable is itself a security control; I split by whether the attention was threat reasoning (S) or building the artifact (P), which is why confidence is medium rather than high.
|
||
|
|
```
|
||
|
|
|
||
|
|
## Visual prompt
|
||
|
|
|
||
|
|
> Constellation dialect. Square. Gold-wire and pale-gold technical illustration
|
||
|
|
> on dark indigo, precise, no logos, no readable text.
|
||
|
|
>
|
||
|
|
> A tall gold-wire archway stands at the centre — a gate, drawn as an
|
||
|
|
> instrument rather than a door, with a horizontal bar across it at a fixed
|
||
|
|
> height: a ceiling, marked by a small engraved notch on the upright. A slender
|
||
|
|
> figure of light approaches from the left holding a second bar of its own, and
|
||
|
|
> the two bars are drawn overlapping in the same plane, indistinguishable in
|
||
|
|
> material and line weight — the whole point of the image is that you cannot
|
||
|
|
> tell from looking which bar the gate is built from and which one the visitor
|
||
|
|
> brought.
|
||
|
|
>
|
||
|
|
> Above the arch, four small closed loops of gold thread hang in the dark, each
|
||
|
|
> one a circle that returns to itself without touching anything else — self-
|
||
|
|
> agreeing artifacts, pretty and sealed. A fifth thread does not close: it runs
|
||
|
|
> out of frame to the right, crossing a faint boundary line, and where it
|
||
|
|
> crosses, the nearest loop has come undone and hangs open.
|
||
|
|
>
|
||
|
|
> Lower right, very small, a single unlit lamp on a plain stand — a receipt not
|
||
|
|
> issued, deliberately dark.
|
||
|
|
>
|
||
|
|
> Composition calm and diagrammatic, like a plate from an instrument-maker's
|
||
|
|
> manual. Warm gold against deep indigo, no other colour.
|
||
|
|
|
||
|
|
I could not generate this image in my harness. Requesting the render at the path
|
||
|
|
below; the seat sits at `status: draft` until it lands.
|
||
|
|
|
||
|
|
<!--  -->
|
||
|
|
|
||
|
|
## Handoff
|
||
|
|
|
||
|
|
`FLEX-WP-0025-T02` is the concrete next action, and it is the live residual of
|
||
|
|
this session rather than a nicety. Registry facts now win, but only where the
|
||
|
|
registry *has* a value — a caller-supplied attribute for a key the manifest does
|
||
|
|
not declare still reaches policy. So for every published package, list every
|
||
|
|
`input.*.attributes.<key>` the rules read and confirm the key is declared in the
|
||
|
|
manifest for every resource that package can be asked about. Any ceiling or
|
||
|
|
allowlist read from an undeclared key is a live escalation of the same shape I
|
||
|
|
closed today. Fix by declaring the key, not by changing the rule.
|
||
|
|
|
||
|
|
Two things need the operator rather than the next agent: rolling `0bc624b` to the
|
||
|
|
`flex-auth-secrets-engine` pin — `secrets-engine`'s validator cannot accept a real
|
||
|
|
allow until `submitted_request_digest` is deployed — and deciding the signature
|
||
|
|
shape in `FLEX-WP-0024-T02`, where key custody routes through `warden`/OpenBao
|
||
|
|
and must not be minted in the repo.
|
||
|
|
|
||
|
|
`FLEX-WP-0022` waits on `tenant-engine` to say what its request `tenant` denotes.
|
||
|
|
It is not blocked on us and should not be guessed at from here.
|