flex-auth had no image workflow, so its production images were hand-built on a workstation and pushed with workstation credentials -- an artifact whose provenance was someone's working tree rather than a forge revision. Ports activity-core's canonical image.yaml: the container-build runner fetches a tarball of the pushed commit, builds, and pushes :latest and :main-<short-sha>, then reports the immutable digest for the rollout step. examples/** is a build-trigger path on purpose -- policy packages are COPYed into the image and there is no hot reload, so a policy change is an image change. Runbook updated to say plainly that images are not built on workstations. FLEX-WP-0011 filed for the rest of the gap: flex-auth runs two production Deployments on railiance01 but has never been brought under the staged-promotion contract (RAIL-BS-WP-0006). No railiance/app.toml, no stage commands, no canary or approval evidence; the deploy/ directory is a rescue of specs that existed nowhere, not the sanctioned overlay shape. Pre-existing debt found during the FLEX-WP-0010 rollout, not a regression from it, and not a blocker for routine policy rollouts -- flex-auth fails closed. T03 also flags that the coulombcore drain plan still lists flex-auth on a host it no longer runs on. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| flex-auth-tenant-engine.yaml | ||
| flex-auth-user-engine.yaml | ||
| README.md | ||
flex-auth production deployment
Manifests for the two cluster-local flex-auth policy-decision services.
Both were previously applied with kubectl apply from a file that lived
outside this repo; these were recovered from the live objects'
kubectl.kubernetes.io/last-applied-configuration on 2026-08-11 and
committed so that a rollback does not depend on a cluster annotation.
| File | Deployment | Consumer | Service DNS |
|---|---|---|---|
flex-auth-tenant-engine.yaml |
flex-auth-tenant-engine |
tenant-engine write API | flex-auth-tenant-engine.flex-auth.svc.cluster.local:8080 |
flex-auth-user-engine.yaml |
flex-auth-user-engine |
user-engine portal | flex-auth-user-engine.flex-auth.svc.cluster.local:8080 |
Each file is a three-document manifest: Deployment, Service, and a
default-deny NetworkPolicy whose ingress is restricted to the one approved
consumer workload and which permits no egress.
A harmless diff on apply
kubectl apply reports the two NetworkPolicies as configured rather than
unchanged, every time. That is not drift: the manifests carry an explicit
egress: [], which the API server normalises away on read. With
policyTypes: [Ingress, Egress] and no egress rules, deny-all egress holds
either way. The empty list is kept because it states the intent to a reader
instead of leaving it implicit. Deployments and Services do round-trip as
unchanged.
One image, two deployments
Both Deployments run the same image repository and differ only in their
--registry / --policy arguments. The Containerfile does
COPY examples /opt/flex-auth/examples, so every image contains every
consumer's policy package; the arguments select which one that instance
serves.
Consequence worth remembering: rebuilding to pick up one consumer's policy change also re-bakes every other consumer's policy into the new image. The two Deployments are pinned to different digests precisely so that one can be rolled without moving the other. Roll only the Deployment whose policy actually changed.
Rolling out a policy change
The policy packages are baked into the image, not mounted from a ConfigMap, so a policy change requires a rebuild — there is no hot reload.
Do not build images on a workstation. The fleet builds in CI so that an
artifact's provenance is a forge revision rather than someone's working tree.
.forgejo/workflows/image.yaml handles it, and examples/** is one of its
trigger paths precisely because policy changes are image changes.
# 1. Push the commit you intend to ship; CI builds it on the container-build
# runner and pushes :latest and :main-<short-sha>
git push origin main
# 2. Take the immutable digest from the workflow's "Report immutable digest"
# step -- deploy by digest, never by tag
# 3. Edit the image digest in the relevant manifest, then apply
kubectl apply -f deploy/flex-auth-<consumer>.yaml
kubectl -n flex-auth rollout status deploy/flex-auth-<consumer> --timeout=120s
# 4. Verify the new policy actually took effect, from outside the cluster
kubectl -n flex-auth port-forward svc/flex-auth-<consumer> 19099:8080 &
curl -s -X POST http://127.0.0.1:19099/v1/check \
-H 'Content-Type: application/json' \
-d @examples/<consumer>/<a request that exercises the change>.json
Step 4 is not optional. Because the policy ships inside the image, a
successful rollout status only proves the container started — it says
nothing about which policy revision is being served.
Rollback
kubectl -n flex-auth rollout undo deploy/flex-auth-<consumer>
If the ReplicaSet history has been pruned, re-apply the manifest with the last-known-good digest below.
| Deployment | Last-known-good digest | Policy state |
|---|---|---|
flex-auth-tenant-engine |
sha256:c25fc34a6cd7e64d955f8723ec70e176a583d5ae71d76280c4e2d89fba0fe0aa |
four-action policy, pre-FLEX-WP-0010 (lifecycle actions deny unknown_action) |
flex-auth-user-engine |
sha256:a31961c45215aa6baf3bc748c6741ab703c2c8325e61aa7983a355026195e51b |
FLEX-WP-0009-T03, six fixtures verified live 2026-08-10 |
Rolling back the tenant-engine service to c25fc34a… restores fail-closed
behaviour for the lifecycle actions — tenant-engine's lifecycle endpoints
return 403 write_denied rather than writing. That is a safe failure mode,
not an outage of the older four actions, which keep working.