--- id: FLEX-WP-0011 type: workplan title: "Bring flex-auth under the railiance staged-promotion contract" domain: infotech repo: flex-auth status: finished owner: codex topic_slug: netkingdom planning_priority: P2 planning_order: 110 depends_on_workplans: - FLEX-WP-0009 related_workplans: - RAIL-BS-WP-0006 created: "2026-08-11" updated: "2026-08-16" state_hub_workstream_id: "deda35b4-f41d-559e-954e-a75f29237f68" --- # FLEX-WP-0011 - Bring flex-auth under the railiance staged-promotion contract flex-auth runs two production Deployments on **railiance01** (`flex-auth-tenant-engine`, `flex-auth-user-engine`) but has never been brought under the staged-promotion contract (`RAIL-BS-WP-0006`) that defines what counts as production on that host. This is pre-existing debt discovered during the `FLEX-WP-0010` rollout, not a regression introduced by it. ## What is missing The contract's gate table requires: | Gate | Requirement | flex-auth today | |---|---|---| | Overlay repo | `railiance//` with `app.toml` and stage manifests | **missing** — manifests live in `deploy/*.yaml`, a parallel convention | | Stage commands | `stage deploy`, `stage observe`, `stage promote`, `stage rollback` proven | **missing** | | Evidence | Backup/restore drill, canary observation, operator approval recorded | partial — probe evidence exists in workplans, no canary or approval record | | Registry | Image in forge OCI registry with immutable tag | **met** — Deployments pin by digest | Two related facts found at the same time: - The Deployments were originally created by `kubectl-client-side-apply` from a file that existed in no repository. `FLEX-WP-0010` recovered them into `deploy/` from the live objects' `last-applied-configuration` so that rollback no longer depends on a cluster annotation. That directory is a rescue, not the sanctioned shape, and this workplan should absorb it. - CI image publishing was added in `FLEX-WP-0010` (`.forgejo/workflows/image.yaml`, modelled on activity-core). Before that, images were hand-built on a workstation with no forge-verifiable provenance. The registry gate is now met by construction. ## Not a blocker for existing traffic flex-auth is already serving production decisions and its failure mode is fail-closed: consumers receive `deny`/`403` rather than an unauthorized allow. This workplan closes a governance and recoverability gap, not an active safety incident. It should not be used as a reason to hold routine policy rollouts. ## T01 - Author the railiance overlay ```task id: FLEX-WP-0011-T01 status: done priority: medium state_hub_task_id: "75748946-68c0-5f33-abe1-810fcd55e93f" ``` Write `railiance/app.toml` against schema `railiance.app.v1`, using `qonto-assistant/railiance/app.toml` as the reference. Declare `[app]` with an honest `criticality` (flex-auth is an authorization PDP whose outage blocks consumer writes fail-closed), `[source]` with `digest_policy = "required"`, `[rollback]` naming the concrete rollback command, and `observability.health_endpoints` for both Services' `/healthz`. Model stage1/stage2/stage3 on the two-Deployment reality: the two consumers are independently pinned and must be independently rollable. Do not collapse them into one release. Reshape `deploy/*.yaml` into the overlay layout (compare `activity-core/k8s/railiance/`) and keep `deploy/README.md`'s rollout and rollback runbook content rather than discarding it. Done when the overlay renders and a server-side dry run of the full manifest set is accepted. Done 2026-08-16: `railiance/app.toml` validates against `railiance.app.v1` (`criticality = high`, `digest_policy = required`, both production `/healthz` endpoints plus isolated canary). One chart, three independently rollable values files (`values/tenant-engine.yaml`, `values/user-engine.yaml`, `values/stage2-canary.yaml`). `tests/stage1.sh` renders all five values files. `bin/railiance run .` passed (stage1 result `/tmp/flex-auth-stage1-result.json`). Server-side dry run of canary + both production pins accepted (`kubectl apply --dry-run=server --server-side` and `helm upgrade --dry-run=server` for the canary). `deploy/README.md` kept as the emergency kubectl path and points at the overlay. Live discovery while pinning values: `flex-auth-tenant-engine` is back on `sha256:c25fc34a…` (four-action policy). `tenant.retire` / `tenant.update` / `tenant.reactivate` deny `unknown_action`. The FLEX-WP-0010 image `sha256:9320df39…` is a ReplicaSet at desired=0. Overlay pins match live; tenant-engine was notified (`50a48026-a75b-4b1f-97f6-7c60063a30b8`). This workplan does not re-promote that image. ## T02 - Prove the stage commands ```task id: FLEX-WP-0011-T02 status: done priority: medium state_hub_task_id: "4becc09f-2331-5487-a794-643c748dd1bd" ``` Prove `stage deploy`, `stage observe`, `stage promote`, and `stage rollback` against a real candidate image, with a canary observation window and a recorded operator approval. Record the previous stable digest as the rollback target before promoting. Done when a full promote-then-rollback cycle has been demonstrated and the evidence ids are recorded here. Done 2026-08-16 on railiance01, isolated release `flex-auth-canary` only. Production pins were not moved. | Step | Result | Evidence | | --- | --- | --- | | Approval | isolated-canary-only | progress `6b07bee3-0fd6-4b08-bb06-5dddac4130d9` | | `stage deploy` apply | Helm rev 1, deploy 1/1 | stage2-deploy-result status `applied` | | `stage observe` live | rollout/pods passed; metrics optional unavailable | stage2-observe-result status `passed` | | Canary `/healthz` | 200 `{"status":"ok"}` | port-forward 19101 | | Canary `me.read` | allow / `self_service` | `decision:27cbab2e5533508a` | | `stage promote` apply | Helm rev 2 | stage3-promote-result status `applied` | | `stage rollback` apply `--revision 1` | Helm rev 3 "Rollback to 1" | stage3-rollback-result status `applied` | | Cleanup | `helm uninstall flex-auth-canary` | production deploys unchanged | Previous stable recorded before promote: `flex-auth-user-engine@sha256:1f5290376dc5fcf456dc7a785e394d8b90949dabecd1d3e856f38557149bb5f4` `flex-auth-tenant-engine@sha256:c25fc34a6cd7e64d955f8723ec70e176a583d5ae71d76280c4e2d89fba0fe0aa` Candidate image: `sha256:1f5290376dc5fcf456dc7a785e394d8b90949dabecd1d3e856f38557149bb5f4`. ## T03 - Correct the stale drain-plan row ```task id: FLEX-WP-0011-T03 status: done priority: low state_hub_task_id: "abb788fe-22a5-5226-9f08-a43e20a47ce9" ``` `the-custodian/docs/coulombcore-drain-placement-plan.md` row 23 lists flex-auth as living on coulombcore, drain wave 7, status `grandfathered`, target railiance01. flex-auth in fact already runs on railiance01 (`92.205.62.239`); the row is stale and understates progress. This is custodian canon, not a flex-auth file — raise it with the custodian owner rather than editing it from this repo. Confirm at the same time whether a coulombcore flex-auth deployment still exists and needs retiring, or whether wave 7 can be closed. Done when the row reflects reality or a custodian decision id explains why it stands. Raised 2026-08-16 with `the-custodian` (message `7219e675-fd39-42ca-a302-8722a895531d`). Confirmed the same day: railiance01 namespace `flex-auth` is Active (7d+); coulombcore has k3s but no `flex-auth` namespace and no flex-auth Deployments. Wave 7.3 has nothing left to drain. T03 waits on the canon edit or a decision id. Done 2026-08-16: operator authorized the canon edit. Row 23 in `the-custodian/docs/coulombcore-drain-placement-plan.md` now lists current = railiance01 `flex-auth` ns, status **on target**, owner `FLEX-WP-0011`. Wave step 7.3 is restated as identity-endpoint follow-up only; host drain is complete. Custodian commit `c377f0d`.