flex-auth/docs/operator-caller-access-path.md
tegwick 6e3dfaeb41
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 2s
Enforce verified secrets-engine operator caller identity
Assistant: codex
Assistant-Model: gpt-6-astra
Assistant-Session: 01a0726e-5232-73f2-aaca-2c05ceb62efb
2026-09-06 23:38:36 +02:00

221 lines
11 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Operator caller access path
**Status:** revision 3 enforces adopted caller; positive and all four negative live checks pass
**Opened by:** `glas-harness` (`GLAS-WP-0015`, 2026-09-06), carried by `FLEX-WP-0023`
**Supersedes:** the workload assumption in `FLEX-WP-0021-T04`
`secrets-engine` is an operator CLI, not a Kubernetes workload. Every existing
flex-auth pin assumes a workload: default-deny ingress admitting one pod
selector, and a `--caller-binding` naming that pod's ServiceAccount. Neither
half transfers unexamined, and `glas-harness` ruled out the two shortcuts by
name — **Service DNS is not connectivity, and a permanent operator token is not
an identity**. Both refusals are correct.
## The address itself is a hazard from a workstation
`secrets-engine` probed the address this repo handed them, and flex-auth
reproduced it:
```text
$ getent hosts flex-auth-secrets-engine.flex-auth.svc.cluster.local
80.158.43.29 flex-auth-secrets-engine.flex-auth.svc.cluster.local.ad.binect.de
$ getent hosts this-service-does-not-exist.flex-auth.svc.cluster.local
80.158.43.29 this-service-does-not-exist.flex-auth.svc.cluster.local.ad.binect.de
$ getent hosts flex-auth-secrets-engine.flex-auth.svc.cluster.local.
(no resolution)
```
`resolv.conf` carries `search ad.binect.de`, which answers wildcard. A name for a
service that does not exist resolves to the same address as one that does, which
is the proof that this is suffix expansion and not a record. So on this
workstation **every `*.svc.cluster.local` name resolves to one unrelated public
host**, and the bare Service name flex-auth published was not merely unreachable
from there — it was a live misdirection.
Had a deployment pointed at it, the CheckRequest body would have gone to that
host: subject, tenant, lane and resource ids, stage, declared field names,
purpose, plus the caller's bearer token. No secret values, but a structural map
of the estate's credential lanes and a credential. `secrets-engine` declined to
mitigate it locally on the grounds that choosing a transport control for
flex-auth's service is not a consumer's call. That is the correct boundary and
the same one that kept them from authoring a tenant mapping.
**Publish the trailing-dot FQDN, and say in-cluster only.** The trailing dot
makes resolution fail instead of succeeding at the wrong place, which is the
behaviour a fail-closed consumer needs from a name.
## The two gates are not one gate
`FLEX-WP-0021-T04` noted that `callerAuth` "becomes the real boundary" for an
operator path. That was understated. For this path the network gate provides
**no protection at all**, and it is worth being exact about why.
`kubectl port-forward` does not traverse a `NetworkPolicy`. The connection is
proxied through the API server to the kubelet and delivered on the pod's own
loopback interface, so it never appears as pod-to-pod ingress and no policy
selector is consulted. The pin's default-deny `NetworkPolicy` is therefore not
a partial control for an operator caller — it is silent.
So the honest statement of the current posture is stronger than "warn mode":
> With `callerAuth.mode: warn` and a port-forwarded connection, the operator
> path to the `secrets-engine` pin is **unauthenticated and unfiltered**. The
> only reason it is not reachable today is that nobody has forwarded the port.
That is not a supported access path. It is the absence of one.
## The supported path
Kubernetes already issues exactly the credential shape being asked for, and
flex-auth already accepts it. Nothing needs inventing and no code changes.
```bash
# short-lived, audience-scoped, cluster-issued, bound to one ServiceAccount
kubectl -n secrets-engine create token secrets-engine \
--audience=flex-auth \
--duration=10m
```
The `TokenRequest` API mints a token for a ServiceAccount **without a pod**.
Checked against what the deployed pin actually requires:
| Requirement | How this satisfies it |
| --- | --- |
| authenticated | `TokenReview` validates signature, audience, and expiry server-side |
| bound to one system | token `sub` is `system:serviceaccount:secrets-engine:secrets-engine`, matched against `--caller-binding` by exact string |
| stated lifetime | `--duration`; the token carries `exp` and `TokenReview` refuses it after |
| not a permanent operator token | expires on its own; the operator stores no long-lived secret |
| audience-scoped | `--audience=flex-auth` matches the pin's `--caller-audience` default; a token minted for any other audience fails |
The deployed pin already names the identity:
```text
--caller-binding secrets-engine=system:serviceaccount:secrets-engine:secrets-engine
--caller-audience flex-auth (default)
```
## Two blockers, one of them load-bearing
**1. The ServiceAccount named by the binding does not exist.** Namespace
`secrets-engine` was created under `FLEX-WP-0021-T04` with no workload and no
identity; it holds only `default`. So there is nothing to mint a token *for*,
and flipping to `enforce` today would deny every request rather than
authenticate one. Creating it is a one-object production write, and an SA with
no `RoleBinding` grants nothing in the cluster — its only function is to be the
name in the binding above.
**2. `warn` cannot be the mode for this path.** For a workload, `warn` is a safe
migration state because the `NetworkPolicy` still admits only one pod. For an
operator caller the policy is silent, so `warn` means *no* control. The
migration order that applied to `ops-warden` (`FLEX-WP-0016`: adopt identity,
clean warn logs, then enforce) still applies, but the warn window here is a
window with nothing in it — it proves the token works and protects nothing while
it runs. **Keep it short and treat `enforce` as the deliverable, not the
follow-up.**
## Positive and negative tests
Run against the pin through a port-forward, then remove the forward.
```bash
kubectl -n flex-auth port-forward svc/flex-auth-secrets-engine 8080:8080 &
# positive — correct identity, correct audience, inside lifetime
TOKEN=$(kubectl -n secrets-engine create token secrets-engine --audience=flex-auth --duration=10m)
curl -s -X POST localhost:8080/v1/check -H "Authorization: Bearer $TOKEN" \
-H 'Content-Type: application/json' \
-d @examples/secrets-engine/check_request_allow_rotate.json
# expect: 200, effect allow, reason catalog_lane_policy_matched, policy_version v2
```
Four negatives, each isolating one property. Under `enforce` all four must be
refused before the request reaches policy; under `warn` all four are *allowed*,
which is the finding rather than a passing test.
| # | Credential | Expected under `enforce` |
| --- | --- | --- |
| N1 | no `Authorization` header | 401 — `ErrUnauthenticated` |
| N2 | token for `secrets-engine:default` (wrong SA, right audience) | 403 — principal cannot represent system |
| N3 | token minted without `--audience=flex-auth` | 401 — audience not in token |
| N4 | token past `exp` (mint `--duration=10m`, use after) | 401 — `TokenReview` refuses |
N2 and N4 are the two that matter. N2 proves the binding is a binding and not
mere presence of a valid cluster token — any of thousands of ServiceAccounts
could produce a well-formed token, and only one may represent `secrets-engine`.
N4 proves the lifetime is enforced by the issuer rather than asserted by the
caller.
**These receipts are not in this document because this session could not mint
tokens.** `kubectl create token` is credential issuance and was refused here, as
it should be. The commands above are exact and the expectations are derived from
`internal/callerauth/auth.go`, not guessed; running them is an operator action.
Nothing below the design line is claimed as verified.
## Which end of the channel is authenticated
`callerAuth` authenticates the **caller to the PDP**. Nothing authenticates the
**PDP to the caller**, and `secrets-engine` was right to ask rather than assume.
Stated plainly, because a stance is only a stance if it is recorded: the pin
serves plain HTTP on `:8080` and `flex-auth.decision-record.v1` carries no
signature. **A responder that knows the package id and version — both published
in this repo — can return a well-formed `effect: allow` that passes every check
a consumer performs.**
The trap worth naming is that the digests look like they help and do not. A
consumer that recomputes `request_digest`, `policy_package_digest`, and
`registry_snapshot_digest` and finds them all correct has verified nothing about
who answered, because **every input to those digests is either sent by the
caller or published**: the request material is what the caller just transmitted,
and both package and registry digests are computable from files in a repo. A
forger reproduces all three exactly. The digests establish integrity of the
binding, never authenticity of the source, and reading a matching digest as
evidence of a genuine PDP is the same visible-but-not-verifying seam as
`FLEX-DEC-2026-008`.
Consequence for a fail-closed consumer, which is the part that matters:
**fail-closed protects against a PDP that is absent, not against one that
lies.** A forged allow defeats the posture entirely rather than degrading it.
`FLEX-DEC-2026-010` records the stance and `FLEX-WP-0024` carries the fix. One
nuance belongs here, though, because it changes the operator recommendation:
**A `kubectl port-forward` path does authenticate the responder, transitively.**
It resolves no DNS name, targets one named pod explicitly, and runs over the
API server's TLS with the operator's cluster credentials. So for the operator
shape the port-forward is not merely a workaround for reachability — it is
currently the *only* path where the consumer knows it is talking to the real
pin. That is the reverse of the caller direction, where the port-forward
bypasses the `NetworkPolicy` entirely. The two properties are independent and
point opposite ways, which is why they have to be stated separately rather than
summarised as "the network protects it".
## What the record will not show
Even with all four negatives passing, **the decision record does not say who
called.** `flex-auth.decision-record.v1` has no caller field: `provenance`
carries the evaluator, mode, policy and registry digests, and decision time, and
`binding` carries the normalized request. The authenticated caller principal —
the thing these tests are about — appears nowhere in the artifact.
So a decision record proves the *subject* was allowed. It cannot prove the
*caller* who obtained it was authenticated, or under what lifetime. For
`glas-harness`, whose ask is a scoped delivery receipt, that is the difference
between "this decision permits the action" and "this caller was permitted to
obtain this decision". Recorded as `FLEX-DEC-2026-009`; it is a gap in
flex-auth's own §17 contract, not in the deployment.
## Live execution update — 2026-09-06
The earlier absent-SA and warn-mode observations above are historical design
findings. Glas created the bound identity from
`deploy/secrets-engine-operator-caller.yaml`, proved adoption with no warnings,
and upgraded the dedicated pin to Helm revision 3 / enforce. Positive request and N1N4 pass, including an actually expired issued token
(401) followed by a fresh token (200). The temporary forward is closed.
See FLEX-WP-0023 and glas-harness/docs/evidence/GLAS-WP-0015-caller-auth-2026-09-06.json.
No other consumer deployment changed. Recreate a loopback-only forward and mint
a fresh bounded token for each authorized operator session; the verification
forward is temporary and is not the runtime endpoint after cleanup.