rapp-canned-prompts/manifests/migration.yaml
tegwick 1b86e872a4 Correct a premature verified, and survive credential rotation
I recorded readiness_state: verified after a passing smoke run and found the
deployment 0/1 eight hours later.

Platform credentials are 30-minute leases, not passwords. The service read the
mounted URL once at start-up and never again, so External Secrets kept the file
current while the engine held the URL it booted with, and every reconnection
after the first lease expiry used a credential the database had already
revoked.

/readyz reported it accurately — "database unreachable: OperationalError" — and
the pod sat unready for eight hours without being restarted, because liveness
is deliberately independent of the database. That separation behaved exactly as
designed: the process was alive and could not serve, and it said so.

Fixed in 0.1.5: the engine re-reads the credential for every new connection and
pool_recycle is 900s, well inside the lease. Only username and password come
from the refreshed URL; host, port and database come from the engine, so a
malformed refresh cannot silently redirect the service.

readiness_state back to deployed. The bar for verified is now observation
across a full lease rotation, because a smoke run inside the first window
cannot distinguish a service that works from one that works once — an
availability property is not provable by a single sample.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-08 09:01:10 +02:00

73 lines
2.7 KiB
YAML

# Schema migration as a Job, not an init container and not a start-up hook.
# Running migrations at start-up races between replicas and couples a rollback
# of the code to a rollback of the schema. The Job name carries the target
# revision so a re-apply at the same revision is a no-op rather than a rerun.
apiVersion: batch/v1
kind: Job
metadata:
name: canned-prompts-schema-migration-0002
namespace: canned-prompts
labels:
app.kubernetes.io/name: canned-prompts-migration
app.kubernetes.io/component: migration
spec:
backoffLimit: 2
ttlSecondsAfterFinished: 86400
template:
metadata:
labels:
app.kubernetes.io/name: canned-prompts-migration
# Carries part-of so the egress NetworkPolicy selects this Job too.
# Without it the Job matches only the default-deny and cannot reach
# PostgreSQL or DNS.
app.kubernetes.io/part-of: canned-prompts
spec:
automountServiceAccountToken: false
restartPolicy: Never
securityContext:
runAsNonRoot: true
runAsUser: 10001
runAsGroup: 10001
fsGroup: 10001
seccompProfile:
type: RuntimeDefault
containers:
- name: migrate
image: forgejo.coulomb.social/coulomb/canned-prompts@sha256:14c7b92f20d63f2e70ea17b0b45d3c483521fbbaf14ee5bc277fa541e10452e2
command: ["alembic"]
args: ["upgrade", "head"]
workingDir: /app
env:
# The migration role owns the schema; the runtime role does not.
- name: CANNED_PROMPTS_DATABASE_URL_FILE
value: /var/run/secrets/postgres-migration/url
# Authenticate as the leased migration login, then SET ROLE to the
# durable owner before creating anything. Leases are revoked; an
# object owned by a dead login has to be normalized afterwards.
# Required by the rapp-postgres database-owner boundary.
- name: CANNED_PROMPTS_MIGRATION_ROLE
value: canned_prompts_owner
securityContext:
allowPrivilegeEscalation: false
capabilities:
drop: ["ALL"]
readOnlyRootFilesystem: true
resources:
requests:
cpu: 25m
memory: 64Mi
limits:
cpu: 500m
memory: 256Mi
volumeMounts:
- name: postgres-migration
mountPath: /var/run/secrets/postgres-migration
readOnly: true
volumes:
- name: postgres-migration
secret:
defaultMode: 0440
secretName: canned-prompts-postgres-migration
items:
- key: url
path: url