Publish images via CI; file staged-promotion overlay debt
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 1s
Build and Publish Container Image / build-and-push (push) Successful in 52s

flex-auth had no image workflow, so its production images were hand-built
on a workstation and pushed with workstation credentials -- an artifact
whose provenance was someone's working tree rather than a forge revision.
Ports activity-core's canonical image.yaml: the container-build runner
fetches a tarball of the pushed commit, builds, and pushes :latest and
:main-<short-sha>, then reports the immutable digest for the rollout step.

examples/** is a build-trigger path on purpose -- policy packages are
COPYed into the image and there is no hot reload, so a policy change is
an image change.

Runbook updated to say plainly that images are not built on workstations.

FLEX-WP-0011 filed for the rest of the gap: flex-auth runs two production
Deployments on railiance01 but has never been brought under the
staged-promotion contract (RAIL-BS-WP-0006). No railiance/app.toml, no
stage commands, no canary or approval evidence; the deploy/ directory is
a rescue of specs that existed nowhere, not the sanctioned overlay shape.
Pre-existing debt found during the FLEX-WP-0010 rollout, not a regression
from it, and not a blocker for routine policy rollouts -- flex-auth fails
closed. T03 also flags that the coulombcore drain plan still lists
flex-auth on a host it no longer runs on.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
tegwick 2026-08-11 10:35:34 +02:00
parent 2456287e8c
commit e9911eb77f
3 changed files with 196 additions and 7 deletions

View file

@ -0,0 +1,65 @@
name: Build and Publish Container Image
# Modelled on activity-core/.forgejo/workflows/image.yaml -- the fleet's
# canonical image-publish pattern. Images are built by CI from a tarball of
# the pushed commit, never from a workstation working tree, so the artifact's
# provenance is a forge revision.
#
# `examples/**` is a build-trigger path on purpose: the policy packages are
# COPYed into the image (see Containerfile) and there is no hot reload, so a
# policy change is a new image.
on:
push:
branches:
- main
paths:
- ".forgejo/workflows/image.yaml"
- "Containerfile"
- "cmd/**"
- "internal/**"
- "pkg/**"
- "examples/**"
- "go.mod"
- "go.sum"
workflow_dispatch:
env:
REGISTRY: forgejo.coulomb.social
IMAGE_NAME: coulomb/flex-auth
DOCKER_HOST: tcp://127.0.0.1:2375
jobs:
build-and-push:
runs-on: container-build
steps:
- name: Build and push image
env:
REGISTRY_USER: ${{ secrets.REGISTRY_USER }}
REGISTRY_TOKEN: ${{ secrets.REGISTRY_TOKEN }}
run: |
set -eu
REF="${GITHUB_SHA:-main}"
SHORT="${REF:0:7}"
mkdir -p buildctx "${HOME}/bin"
wget -qO /tmp/repo.tar.gz \
"https://forgejo.coulomb.social/${GITHUB_REPOSITORY}/archive/${SHORT}.tar.gz"
tar xzf /tmp/repo.tar.gz -C buildctx --strip-components=1
wget -qO- https://download.docker.com/linux/static/stable/x86_64/docker-27.3.1.tgz \
| tar xz --strip-components=1 -C "${HOME}/bin" docker/docker
export PATH="${HOME}/bin:${PATH}"
echo "${REGISTRY_TOKEN}" | docker login "${REGISTRY}" -u "${REGISTRY_USER}" --password-stdin
IMAGE="${REGISTRY}/${IMAGE_NAME}"
docker build -f buildctx/Containerfile -t "${IMAGE}:latest" -t "${IMAGE}:main-${SHORT}" buildctx
docker push "${IMAGE}:latest"
docker push "${IMAGE}:main-${SHORT}"
echo "pushed ${IMAGE}:latest and ${IMAGE}:main-${SHORT}"
- name: Report immutable digest
run: |
set -eu
export PATH="${HOME}/bin:${PATH}"
IMAGE="${REGISTRY}/${IMAGE_NAME}"
SHORT="${GITHUB_SHA:0:7}"
# Deployments pin by digest, never by tag -- print it for the rollout step.
docker inspect --format='{{index .RepoDigests 0}}' "${IMAGE}:main-${SHORT}"

View file

@ -44,14 +44,18 @@ actually changed.
The policy packages are baked into the image, not mounted from a ConfigMap, The policy packages are baked into the image, not mounted from a ConfigMap,
so a policy change requires a rebuild — there is no hot reload. so a policy change requires a rebuild — there is no hot reload.
```bash **Do not build images on a workstation.** The fleet builds in CI so that an
# 1. Build and push from a clean checkout of the commit you intend to ship artifact's provenance is a forge revision rather than someone's working tree.
docker build -t forgejo.coulomb.social/coulomb/flex-auth:<tag> -f Containerfile . `.forgejo/workflows/image.yaml` handles it, and `examples/**` is one of its
docker push forgejo.coulomb.social/coulomb/flex-auth:<tag> trigger paths precisely because policy changes are image changes.
# 2. Resolve the digest and pin it — deploy by digest, never by tag ```bash
docker inspect --format='{{index .RepoDigests 0}}' \ # 1. Push the commit you intend to ship; CI builds it on the container-build
forgejo.coulomb.social/coulomb/flex-auth:<tag> # runner and pushes :latest and :main-<short-sha>
git push origin main
# 2. Take the immutable digest from the workflow's "Report immutable digest"
# step -- deploy by digest, never by tag
# 3. Edit the image digest in the relevant manifest, then apply # 3. Edit the image digest in the relevant manifest, then apply
kubectl apply -f deploy/flex-auth-<consumer>.yaml kubectl apply -f deploy/flex-auth-<consumer>.yaml

View file

@ -0,0 +1,120 @@
---
id: FLEX-WP-0011
type: workplan
title: "Bring flex-auth under the railiance staged-promotion contract"
domain: infotech
repo: flex-auth
status: proposed
owner: codex
topic_slug: netkingdom
planning_priority: P2
planning_order: 110
depends_on_workplans:
- FLEX-WP-0009
related_workplans:
- RAIL-BS-WP-0006
created: "2026-08-11"
updated: "2026-08-11"
---
# FLEX-WP-0011 - Bring flex-auth under the railiance staged-promotion contract
flex-auth runs two production Deployments on **railiance01**
(`flex-auth-tenant-engine`, `flex-auth-user-engine`) but has never been
brought under the staged-promotion contract (`RAIL-BS-WP-0006`) that defines
what counts as production on that host. This is pre-existing debt discovered
during the `FLEX-WP-0010` rollout, not a regression introduced by it.
## What is missing
The contract's gate table requires:
| Gate | Requirement | flex-auth today |
|---|---|---|
| Overlay repo | `railiance/<app>/` with `app.toml` and stage manifests | **missing** — manifests live in `deploy/*.yaml`, a parallel convention |
| Stage commands | `stage deploy`, `stage observe`, `stage promote`, `stage rollback` proven | **missing** |
| Evidence | Backup/restore drill, canary observation, operator approval recorded | partial — probe evidence exists in workplans, no canary or approval record |
| Registry | Image in forge OCI registry with immutable tag | **met** — Deployments pin by digest |
Two related facts found at the same time:
- The Deployments were originally created by `kubectl-client-side-apply` from
a file that existed in no repository. `FLEX-WP-0010` recovered them into
`deploy/` from the live objects' `last-applied-configuration` so that
rollback no longer depends on a cluster annotation. That directory is a
rescue, not the sanctioned shape, and this workplan should absorb it.
- CI image publishing was added in `FLEX-WP-0010`
(`.forgejo/workflows/image.yaml`, modelled on activity-core). Before that,
images were hand-built on a workstation with no forge-verifiable
provenance. The registry gate is now met by construction.
## Not a blocker for existing traffic
flex-auth is already serving production decisions and its failure mode is
fail-closed: consumers receive `deny`/`403` rather than an unauthorized
allow. This workplan closes a governance and recoverability gap, not an
active safety incident. It should not be used as a reason to hold routine
policy rollouts.
## T01 - Author the railiance overlay
```task
id: FLEX-WP-0011-T01
status: todo
priority: medium
```
Write `railiance/app.toml` against schema `railiance.app.v1`, using
`qonto-assistant/railiance/app.toml` as the reference. Declare `[app]` with
an honest `criticality` (flex-auth is an authorization PDP whose outage
blocks consumer writes fail-closed), `[source]` with `digest_policy =
"required"`, `[rollback]` naming the concrete rollback command, and
`observability.health_endpoints` for both Services' `/healthz`.
Model stage1/stage2/stage3 on the two-Deployment reality: the two consumers
are independently pinned and must be independently rollable. Do not collapse
them into one release.
Reshape `deploy/*.yaml` into the overlay layout (compare
`activity-core/k8s/railiance/`) and keep `deploy/README.md`'s rollout and
rollback runbook content rather than discarding it.
Done when the overlay renders and a server-side dry run of the full manifest
set is accepted.
## T02 - Prove the stage commands
```task
id: FLEX-WP-0011-T02
status: todo
priority: medium
```
Prove `stage deploy`, `stage observe`, `stage promote`, and `stage rollback`
against a real candidate image, with a canary observation window and a
recorded operator approval. Record the previous stable digest as the rollback
target before promoting.
Done when a full promote-then-rollback cycle has been demonstrated and the
evidence ids are recorded here.
## T03 - Correct the stale drain-plan row
```task
id: FLEX-WP-0011-T03
status: todo
priority: low
```
`the-custodian/docs/coulombcore-drain-placement-plan.md` row 23 lists
flex-auth as living on coulombcore, drain wave 7, status `grandfathered`,
target railiance01. flex-auth in fact already runs on railiance01
(`92.205.62.239`); the row is stale and understates progress.
This is custodian canon, not a flex-auth file — raise it with the custodian
owner rather than editing it from this repo. Confirm at the same time whether
a coulombcore flex-auth deployment still exists and needs retiring, or
whether wave 7 can be closed.
Done when the row reflects reality or a custodian decision id explains why it
stands.