flex-auth/docs/operator-caller-access-path.md
tegwick 6e3dfaeb41
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 2s
Enforce verified secrets-engine operator caller identity
Assistant: codex
Assistant-Model: gpt-6-astra
Assistant-Session: 01a0726e-5232-73f2-aaca-2c05ceb62efb
2026-09-06 23:38:36 +02:00

11 KiB
Raw Blame History

Operator caller access path

Status: revision 3 enforces adopted caller; positive and all four negative live checks pass Opened by: glas-harness (GLAS-WP-0015, 2026-09-06), carried by FLEX-WP-0023 Supersedes: the workload assumption in FLEX-WP-0021-T04

secrets-engine is an operator CLI, not a Kubernetes workload. Every existing flex-auth pin assumes a workload: default-deny ingress admitting one pod selector, and a --caller-binding naming that pod's ServiceAccount. Neither half transfers unexamined, and glas-harness ruled out the two shortcuts by name — Service DNS is not connectivity, and a permanent operator token is not an identity. Both refusals are correct.

The address itself is a hazard from a workstation

secrets-engine probed the address this repo handed them, and flex-auth reproduced it:

$ getent hosts flex-auth-secrets-engine.flex-auth.svc.cluster.local
80.158.43.29   flex-auth-secrets-engine.flex-auth.svc.cluster.local.ad.binect.de

$ getent hosts this-service-does-not-exist.flex-auth.svc.cluster.local
80.158.43.29   this-service-does-not-exist.flex-auth.svc.cluster.local.ad.binect.de

$ getent hosts flex-auth-secrets-engine.flex-auth.svc.cluster.local.
(no resolution)

resolv.conf carries search ad.binect.de, which answers wildcard. A name for a service that does not exist resolves to the same address as one that does, which is the proof that this is suffix expansion and not a record. So on this workstation every *.svc.cluster.local name resolves to one unrelated public host, and the bare Service name flex-auth published was not merely unreachable from there — it was a live misdirection.

Had a deployment pointed at it, the CheckRequest body would have gone to that host: subject, tenant, lane and resource ids, stage, declared field names, purpose, plus the caller's bearer token. No secret values, but a structural map of the estate's credential lanes and a credential. secrets-engine declined to mitigate it locally on the grounds that choosing a transport control for flex-auth's service is not a consumer's call. That is the correct boundary and the same one that kept them from authoring a tenant mapping.

Publish the trailing-dot FQDN, and say in-cluster only. The trailing dot makes resolution fail instead of succeeding at the wrong place, which is the behaviour a fail-closed consumer needs from a name.

The two gates are not one gate

FLEX-WP-0021-T04 noted that callerAuth "becomes the real boundary" for an operator path. That was understated. For this path the network gate provides no protection at all, and it is worth being exact about why.

kubectl port-forward does not traverse a NetworkPolicy. The connection is proxied through the API server to the kubelet and delivered on the pod's own loopback interface, so it never appears as pod-to-pod ingress and no policy selector is consulted. The pin's default-deny NetworkPolicy is therefore not a partial control for an operator caller — it is silent.

So the honest statement of the current posture is stronger than "warn mode":

With callerAuth.mode: warn and a port-forwarded connection, the operator path to the secrets-engine pin is unauthenticated and unfiltered. The only reason it is not reachable today is that nobody has forwarded the port.

That is not a supported access path. It is the absence of one.

The supported path

Kubernetes already issues exactly the credential shape being asked for, and flex-auth already accepts it. Nothing needs inventing and no code changes.

# short-lived, audience-scoped, cluster-issued, bound to one ServiceAccount
kubectl -n secrets-engine create token secrets-engine \
  --audience=flex-auth \
  --duration=10m

The TokenRequest API mints a token for a ServiceAccount without a pod. Checked against what the deployed pin actually requires:

Requirement How this satisfies it
authenticated TokenReview validates signature, audience, and expiry server-side
bound to one system token sub is system:serviceaccount:secrets-engine:secrets-engine, matched against --caller-binding by exact string
stated lifetime --duration; the token carries exp and TokenReview refuses it after
not a permanent operator token expires on its own; the operator stores no long-lived secret
audience-scoped --audience=flex-auth matches the pin's --caller-audience default; a token minted for any other audience fails

The deployed pin already names the identity:

--caller-binding secrets-engine=system:serviceaccount:secrets-engine:secrets-engine
--caller-audience flex-auth   (default)

Two blockers, one of them load-bearing

1. The ServiceAccount named by the binding does not exist. Namespace secrets-engine was created under FLEX-WP-0021-T04 with no workload and no identity; it holds only default. So there is nothing to mint a token for, and flipping to enforce today would deny every request rather than authenticate one. Creating it is a one-object production write, and an SA with no RoleBinding grants nothing in the cluster — its only function is to be the name in the binding above.

2. warn cannot be the mode for this path. For a workload, warn is a safe migration state because the NetworkPolicy still admits only one pod. For an operator caller the policy is silent, so warn means no control. The migration order that applied to ops-warden (FLEX-WP-0016: adopt identity, clean warn logs, then enforce) still applies, but the warn window here is a window with nothing in it — it proves the token works and protects nothing while it runs. Keep it short and treat enforce as the deliverable, not the follow-up.

Positive and negative tests

Run against the pin through a port-forward, then remove the forward.

kubectl -n flex-auth port-forward svc/flex-auth-secrets-engine 8080:8080 &

# positive — correct identity, correct audience, inside lifetime
TOKEN=$(kubectl -n secrets-engine create token secrets-engine --audience=flex-auth --duration=10m)
curl -s -X POST localhost:8080/v1/check -H "Authorization: Bearer $TOKEN" \
  -H 'Content-Type: application/json' \
  -d @examples/secrets-engine/check_request_allow_rotate.json
# expect: 200, effect allow, reason catalog_lane_policy_matched, policy_version v2

Four negatives, each isolating one property. Under enforce all four must be refused before the request reaches policy; under warn all four are allowed, which is the finding rather than a passing test.

# Credential Expected under enforce
N1 no Authorization header 401 — ErrUnauthenticated
N2 token for secrets-engine:default (wrong SA, right audience) 403 — principal cannot represent system
N3 token minted without --audience=flex-auth 401 — audience not in token
N4 token past exp (mint --duration=10m, use after) 401 — TokenReview refuses

N2 and N4 are the two that matter. N2 proves the binding is a binding and not mere presence of a valid cluster token — any of thousands of ServiceAccounts could produce a well-formed token, and only one may represent secrets-engine. N4 proves the lifetime is enforced by the issuer rather than asserted by the caller.

These receipts are not in this document because this session could not mint tokens. kubectl create token is credential issuance and was refused here, as it should be. The commands above are exact and the expectations are derived from internal/callerauth/auth.go, not guessed; running them is an operator action. Nothing below the design line is claimed as verified.

Which end of the channel is authenticated

callerAuth authenticates the caller to the PDP. Nothing authenticates the PDP to the caller, and secrets-engine was right to ask rather than assume.

Stated plainly, because a stance is only a stance if it is recorded: the pin serves plain HTTP on :8080 and flex-auth.decision-record.v1 carries no signature. A responder that knows the package id and version — both published in this repo — can return a well-formed effect: allow that passes every check a consumer performs.

The trap worth naming is that the digests look like they help and do not. A consumer that recomputes request_digest, policy_package_digest, and registry_snapshot_digest and finds them all correct has verified nothing about who answered, because every input to those digests is either sent by the caller or published: the request material is what the caller just transmitted, and both package and registry digests are computable from files in a repo. A forger reproduces all three exactly. The digests establish integrity of the binding, never authenticity of the source, and reading a matching digest as evidence of a genuine PDP is the same visible-but-not-verifying seam as FLEX-DEC-2026-008.

Consequence for a fail-closed consumer, which is the part that matters: fail-closed protects against a PDP that is absent, not against one that lies. A forged allow defeats the posture entirely rather than degrading it.

FLEX-DEC-2026-010 records the stance and FLEX-WP-0024 carries the fix. One nuance belongs here, though, because it changes the operator recommendation:

A kubectl port-forward path does authenticate the responder, transitively. It resolves no DNS name, targets one named pod explicitly, and runs over the API server's TLS with the operator's cluster credentials. So for the operator shape the port-forward is not merely a workaround for reachability — it is currently the only path where the consumer knows it is talking to the real pin. That is the reverse of the caller direction, where the port-forward bypasses the NetworkPolicy entirely. The two properties are independent and point opposite ways, which is why they have to be stated separately rather than summarised as "the network protects it".

What the record will not show

Even with all four negatives passing, the decision record does not say who called. flex-auth.decision-record.v1 has no caller field: provenance carries the evaluator, mode, policy and registry digests, and decision time, and binding carries the normalized request. The authenticated caller principal — the thing these tests are about — appears nowhere in the artifact.

So a decision record proves the subject was allowed. It cannot prove the caller who obtained it was authenticated, or under what lifetime. For glas-harness, whose ask is a scoped delivery receipt, that is the difference between "this decision permits the action" and "this caller was permitted to obtain this decision". Recorded as FLEX-DEC-2026-009; it is a gap in flex-auth's own §17 contract, not in the deployment.

Live execution update — 2026-09-06

The earlier absent-SA and warn-mode observations above are historical design findings. Glas created the bound identity from deploy/secrets-engine-operator-caller.yaml, proved adoption with no warnings, and upgraded the dedicated pin to Helm revision 3 / enforce. Positive request and N1N4 pass, including an actually expired issued token (401) followed by a fresh token (200). The temporary forward is closed. See FLEX-WP-0023 and glas-harness/docs/evidence/GLAS-WP-0015-caller-auth-2026-09-06.json. No other consumer deployment changed. Recreate a loopback-only forward and mint a fresh bounded token for each authorized operator session; the verification forward is temporary and is not the runtime endpoint after cleanup.