rapp-canned-prompts/manifests/migration.yaml

74 lines
2.7 KiB
YAML
Raw Normal View History

Package canned-prompts for Railiance Registers the repo with State Hub (agents / practice, prefix RCP-WP) and fills in the rapp shape. declarations/rapp.yaml declares a manifest-managed platform service owned by canned-prompts, bound to rail-kubernetes and reef-railiance, with rollout, smoke and rollback contracts. The image pin says `pending-publication` rather than carrying a placeholder digest. The image builds and was verified locally (canned-prompts CANP-WP-0006-T06) but has never been pushed, so no registry digest exists. A placeholder shaped like a real digest would be worse than a sentinel: it could be mistaken for something deployable. manifests/ follows the rapp-sbom-nexus shape: namespace labelled for the postgres client, external secrets from OpenBao, a migration Job, and the runtime Deployment with a ClusterIP-only Service, dedicated ServiceAccount and default-deny plus runtime NetworkPolicies. Three choices worth stating. Credentials arrive as mounted files, never env vars — an env var holding a password is visible in kubectl describe, in crash dumps, and to anything that can read /proc. Migrations run as a Job rather than at start-up, so a schema rollback stays separate from a code rollback and replicas do not race. Liveness points at /healthz, which checks only that the process is up: pointing it at a database-dependent path would restart every replica during a database blip. Egress is PostgreSQL and DNS only. A package arrives by publish; the registry never reaches out, so it is given no path to. tools/smoke.sh checks what only the cluster can answer and calls canned-prompts' service/tools/smoke.py for health and migration head, rather than holding a second opinion about whether the service is healthy. RCP-WP-0002 carries the two operator actions that block a first rollout — publishing the image and provisioning database roles — and records per-publisher identity as a decision belonging upstream, which this repo must not paper over with cluster configuration implying finer control than exists. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 21:44:17 +02:00
# Schema migration as a Job, not an init container and not a start-up hook.
# Running migrations at start-up races between replicas and couples a rollback
# of the code to a rollback of the schema. The Job name carries the target
# revision so a re-apply at the same revision is a no-op rather than a rerun.
apiVersion: batch/v1
kind: Job
metadata:
name: canned-prompts-schema-migration-0002
namespace: canned-prompts
labels:
app.kubernetes.io/name: canned-prompts-migration
app.kubernetes.io/component: migration
spec:
backoffLimit: 2
ttlSecondsAfterFinished: 86400
template:
metadata:
labels:
app.kubernetes.io/name: canned-prompts-migration
First deployment: verified on railiance01 RCP-WP-0002 T02-T04 done, readiness_state verified with evidence attached rather than ahead of it. rapp-postgres provisioned canned_prompts on platform-pg-2 and sent a database-owner receipt with 12 checks proven live. creds/canned-prompts-publish was deliberately not issued, so the service runs read-only and POST /packages returns 503 explaining why — the intended posture, not a gap. Four defects surfaced that only a real rollout could expose, two of them silent: - SET ROLE opened an implicit transaction that Alembic nested inside rather than owning, so every revision logged as applied and was rolled back. Alembic reported success against an empty database. - The egress NetworkPolicy selected app.kubernetes.io/name, which the migration Job does not carry. The Job matched only the default-deny and succeeded exactly once, because it ran before the policies existed; the next migration would have failed on DNS. Now selects part-of, with ingress split into its own policy so the Job is never reachable. - env.py read database_url rather than resolved_database_url, so the migration could never run where the credential is a mounted file. - live-image-digest-match extracted the pin with a line-offset grep, which returned empty once comments were added above `version:`. The check degraded to reporting "not pinned yet" while a digest was pinned — it could not have passed for any pin. Now parsed as YAML. A verification step that cannot fail is worth less than none, because it is trusted. Evidence: docs/evidence/RCP-WP-0002-T04-first-deployment-2026-09-08.md Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-08 08:56:02 +02:00
# Carries part-of so the egress NetworkPolicy selects this Job too.
# Without it the Job matches only the default-deny and cannot reach
# PostgreSQL or DNS.
app.kubernetes.io/part-of: canned-prompts
Package canned-prompts for Railiance Registers the repo with State Hub (agents / practice, prefix RCP-WP) and fills in the rapp shape. declarations/rapp.yaml declares a manifest-managed platform service owned by canned-prompts, bound to rail-kubernetes and reef-railiance, with rollout, smoke and rollback contracts. The image pin says `pending-publication` rather than carrying a placeholder digest. The image builds and was verified locally (canned-prompts CANP-WP-0006-T06) but has never been pushed, so no registry digest exists. A placeholder shaped like a real digest would be worse than a sentinel: it could be mistaken for something deployable. manifests/ follows the rapp-sbom-nexus shape: namespace labelled for the postgres client, external secrets from OpenBao, a migration Job, and the runtime Deployment with a ClusterIP-only Service, dedicated ServiceAccount and default-deny plus runtime NetworkPolicies. Three choices worth stating. Credentials arrive as mounted files, never env vars — an env var holding a password is visible in kubectl describe, in crash dumps, and to anything that can read /proc. Migrations run as a Job rather than at start-up, so a schema rollback stays separate from a code rollback and replicas do not race. Liveness points at /healthz, which checks only that the process is up: pointing it at a database-dependent path would restart every replica during a database blip. Egress is PostgreSQL and DNS only. A package arrives by publish; the registry never reaches out, so it is given no path to. tools/smoke.sh checks what only the cluster can answer and calls canned-prompts' service/tools/smoke.py for health and migration head, rather than holding a second opinion about whether the service is healthy. RCP-WP-0002 carries the two operator actions that block a first rollout — publishing the image and provisioning database roles — and records per-publisher identity as a decision belonging upstream, which this repo must not paper over with cluster configuration implying finer control than exists. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 21:44:17 +02:00
spec:
automountServiceAccountToken: false
restartPolicy: Never
securityContext:
runAsNonRoot: true
runAsUser: 10001
runAsGroup: 10001
fsGroup: 10001
seccompProfile:
type: RuntimeDefault
containers:
- name: migrate
Correct a premature `verified`, and survive credential rotation I recorded readiness_state: verified after a passing smoke run and found the deployment 0/1 eight hours later. Platform credentials are 30-minute leases, not passwords. The service read the mounted URL once at start-up and never again, so External Secrets kept the file current while the engine held the URL it booted with, and every reconnection after the first lease expiry used a credential the database had already revoked. /readyz reported it accurately — "database unreachable: OperationalError" — and the pod sat unready for eight hours without being restarted, because liveness is deliberately independent of the database. That separation behaved exactly as designed: the process was alive and could not serve, and it said so. Fixed in 0.1.5: the engine re-reads the credential for every new connection and pool_recycle is 900s, well inside the lease. Only username and password come from the refreshed URL; host, port and database come from the engine, so a malformed refresh cannot silently redirect the service. readiness_state back to deployed. The bar for verified is now observation across a full lease rotation, because a smoke run inside the first window cannot distinguish a service that works from one that works once — an availability property is not provable by a single sample. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-08 09:01:10 +02:00
image: forgejo.coulomb.social/coulomb/canned-prompts@sha256:14c7b92f20d63f2e70ea17b0b45d3c483521fbbaf14ee5bc277fa541e10452e2
Package canned-prompts for Railiance Registers the repo with State Hub (agents / practice, prefix RCP-WP) and fills in the rapp shape. declarations/rapp.yaml declares a manifest-managed platform service owned by canned-prompts, bound to rail-kubernetes and reef-railiance, with rollout, smoke and rollback contracts. The image pin says `pending-publication` rather than carrying a placeholder digest. The image builds and was verified locally (canned-prompts CANP-WP-0006-T06) but has never been pushed, so no registry digest exists. A placeholder shaped like a real digest would be worse than a sentinel: it could be mistaken for something deployable. manifests/ follows the rapp-sbom-nexus shape: namespace labelled for the postgres client, external secrets from OpenBao, a migration Job, and the runtime Deployment with a ClusterIP-only Service, dedicated ServiceAccount and default-deny plus runtime NetworkPolicies. Three choices worth stating. Credentials arrive as mounted files, never env vars — an env var holding a password is visible in kubectl describe, in crash dumps, and to anything that can read /proc. Migrations run as a Job rather than at start-up, so a schema rollback stays separate from a code rollback and replicas do not race. Liveness points at /healthz, which checks only that the process is up: pointing it at a database-dependent path would restart every replica during a database blip. Egress is PostgreSQL and DNS only. A package arrives by publish; the registry never reaches out, so it is given no path to. tools/smoke.sh checks what only the cluster can answer and calls canned-prompts' service/tools/smoke.py for health and migration head, rather than holding a second opinion about whether the service is healthy. RCP-WP-0002 carries the two operator actions that block a first rollout — publishing the image and provisioning database roles — and records per-publisher identity as a decision belonging upstream, which this repo must not paper over with cluster configuration implying finer control than exists. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 21:44:17 +02:00
command: ["alembic"]
args: ["upgrade", "head"]
workingDir: /app
env:
# The migration role owns the schema; the runtime role does not.
- name: CANNED_PROMPTS_DATABASE_URL_FILE
value: /var/run/secrets/postgres-migration/url
First deployment: verified on railiance01 RCP-WP-0002 T02-T04 done, readiness_state verified with evidence attached rather than ahead of it. rapp-postgres provisioned canned_prompts on platform-pg-2 and sent a database-owner receipt with 12 checks proven live. creds/canned-prompts-publish was deliberately not issued, so the service runs read-only and POST /packages returns 503 explaining why — the intended posture, not a gap. Four defects surfaced that only a real rollout could expose, two of them silent: - SET ROLE opened an implicit transaction that Alembic nested inside rather than owning, so every revision logged as applied and was rolled back. Alembic reported success against an empty database. - The egress NetworkPolicy selected app.kubernetes.io/name, which the migration Job does not carry. The Job matched only the default-deny and succeeded exactly once, because it ran before the policies existed; the next migration would have failed on DNS. Now selects part-of, with ingress split into its own policy so the Job is never reachable. - env.py read database_url rather than resolved_database_url, so the migration could never run where the credential is a mounted file. - live-image-digest-match extracted the pin with a line-offset grep, which returned empty once comments were added above `version:`. The check degraded to reporting "not pinned yet" while a digest was pinned — it could not have passed for any pin. Now parsed as YAML. A verification step that cannot fail is worth less than none, because it is trusted. Evidence: docs/evidence/RCP-WP-0002-T04-first-deployment-2026-09-08.md Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-08 08:56:02 +02:00
# Authenticate as the leased migration login, then SET ROLE to the
# durable owner before creating anything. Leases are revoked; an
# object owned by a dead login has to be normalized afterwards.
# Required by the rapp-postgres database-owner boundary.
- name: CANNED_PROMPTS_MIGRATION_ROLE
value: canned_prompts_owner
Package canned-prompts for Railiance Registers the repo with State Hub (agents / practice, prefix RCP-WP) and fills in the rapp shape. declarations/rapp.yaml declares a manifest-managed platform service owned by canned-prompts, bound to rail-kubernetes and reef-railiance, with rollout, smoke and rollback contracts. The image pin says `pending-publication` rather than carrying a placeholder digest. The image builds and was verified locally (canned-prompts CANP-WP-0006-T06) but has never been pushed, so no registry digest exists. A placeholder shaped like a real digest would be worse than a sentinel: it could be mistaken for something deployable. manifests/ follows the rapp-sbom-nexus shape: namespace labelled for the postgres client, external secrets from OpenBao, a migration Job, and the runtime Deployment with a ClusterIP-only Service, dedicated ServiceAccount and default-deny plus runtime NetworkPolicies. Three choices worth stating. Credentials arrive as mounted files, never env vars — an env var holding a password is visible in kubectl describe, in crash dumps, and to anything that can read /proc. Migrations run as a Job rather than at start-up, so a schema rollback stays separate from a code rollback and replicas do not race. Liveness points at /healthz, which checks only that the process is up: pointing it at a database-dependent path would restart every replica during a database blip. Egress is PostgreSQL and DNS only. A package arrives by publish; the registry never reaches out, so it is given no path to. tools/smoke.sh checks what only the cluster can answer and calls canned-prompts' service/tools/smoke.py for health and migration head, rather than holding a second opinion about whether the service is healthy. RCP-WP-0002 carries the two operator actions that block a first rollout — publishing the image and provisioning database roles — and records per-publisher identity as a decision belonging upstream, which this repo must not paper over with cluster configuration implying finer control than exists. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 21:44:17 +02:00
securityContext:
allowPrivilegeEscalation: false
capabilities:
drop: ["ALL"]
readOnlyRootFilesystem: true
resources:
requests:
cpu: 25m
memory: 64Mi
limits:
cpu: 500m
memory: 256Mi
volumeMounts:
- name: postgres-migration
mountPath: /var/run/secrets/postgres-migration
readOnly: true
volumes:
- name: postgres-migration
secret:
defaultMode: 0440
secretName: canned-prompts-postgres-migration
items:
- key: url
path: url