rapp-canned-prompts/manifests/migration.yaml
tegwick bdbef58259 Lease-watch verdict must check what it claims
Version 3 reported 'survived past the lease TTL' after a clean 12-minute window
on a 13-minute-old pod. Every measurement in it was accurate; the conclusion
was not, because nothing checked that the observation window had actually
exceeded the 30-minute lease it claimed to have outlasted.

Fifth defect in this instrument, and the first to err toward reassurance.
Versions 1 and 2 cried wolf, which provokes investigation. This one would have
been believed, and readiness_state: verified recorded on it — the same way
live-image-digest-match would have been believed. A check reporting success it
has not established is indistinguishable from one that works, until it matters.

The verdict now requires uptime > lease TTL and reports INCONCLUSIVE when a
window is clean but too short. 'Clean' and 'proven' are different claims and
only one of them was being measured.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-08 10:48:35 +02:00

73 lines
2.7 KiB
YAML

# Schema migration as a Job, not an init container and not a start-up hook.
# Running migrations at start-up races between replicas and couples a rollback
# of the code to a rollback of the schema. The Job name carries the target
# revision so a re-apply at the same revision is a no-op rather than a rerun.
apiVersion: batch/v1
kind: Job
metadata:
name: canned-prompts-schema-migration-0003
namespace: canned-prompts
labels:
app.kubernetes.io/name: canned-prompts-migration
app.kubernetes.io/component: migration
spec:
backoffLimit: 2
ttlSecondsAfterFinished: 86400
template:
metadata:
labels:
app.kubernetes.io/name: canned-prompts-migration
# Carries part-of so the egress NetworkPolicy selects this Job too.
# Without it the Job matches only the default-deny and cannot reach
# PostgreSQL or DNS.
app.kubernetes.io/part-of: canned-prompts
spec:
automountServiceAccountToken: false
restartPolicy: Never
securityContext:
runAsNonRoot: true
runAsUser: 10001
runAsGroup: 10001
fsGroup: 10001
seccompProfile:
type: RuntimeDefault
containers:
- name: migrate
image: forgejo.coulomb.social/coulomb/canned-prompts@sha256:d5de508ea66d3ad9e354f05ccb6f764bd3b01d290631558f0e866059e8a705a6
command: ["alembic"]
args: ["upgrade", "head"]
workingDir: /app
env:
# The migration role owns the schema; the runtime role does not.
- name: CANNED_PROMPTS_DATABASE_URL_FILE
value: /var/run/secrets/postgres-migration/url
# Authenticate as the leased migration login, then SET ROLE to the
# durable owner before creating anything. Leases are revoked; an
# object owned by a dead login has to be normalized afterwards.
# Required by the rapp-postgres database-owner boundary.
- name: CANNED_PROMPTS_MIGRATION_ROLE
value: canned_prompts_owner
securityContext:
allowPrivilegeEscalation: false
capabilities:
drop: ["ALL"]
readOnlyRootFilesystem: true
resources:
requests:
cpu: 25m
memory: 64Mi
limits:
cpu: 500m
memory: 256Mi
volumeMounts:
- name: postgres-migration
mountPath: /var/run/secrets/postgres-migration
readOnly: true
volumes:
- name: postgres-migration
secret:
defaultMode: 0440
secretName: canned-prompts-postgres-migration
items:
- key: url
path: url