rapp-canned-prompts/manifests/runtime.yaml
tegwick 1b86e872a4 Correct a premature verified, and survive credential rotation
I recorded readiness_state: verified after a passing smoke run and found the
deployment 0/1 eight hours later.

Platform credentials are 30-minute leases, not passwords. The service read the
mounted URL once at start-up and never again, so External Secrets kept the file
current while the engine held the URL it booted with, and every reconnection
after the first lease expiry used a credential the database had already
revoked.

/readyz reported it accurately — "database unreachable: OperationalError" — and
the pod sat unready for eight hours without being restarted, because liveness
is deliberately independent of the database. That separation behaved exactly as
designed: the process was alive and could not serve, and it said so.

Fixed in 0.1.5: the engine re-reads the credential for every new connection and
pool_recycle is 900s, well inside the lease. Only username and password come
from the refreshed URL; host, port and database come from the engine, so a
malformed refresh cannot silently redirect the service.

readiness_state back to deployed. The bar for verified is now observation
across a full lease rotation, because a smoke run inside the first window
cannot distinguish a service that works from one that works once — an
availability property is not provable by a single sample.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-08 09:01:10 +02:00

186 lines
5.4 KiB
YAML

apiVersion: apps/v1
kind: Deployment
metadata:
name: canned-prompts
namespace: canned-prompts
labels:
app.kubernetes.io/name: canned-prompts
spec:
replicas: 1
selector:
matchLabels:
app.kubernetes.io/name: canned-prompts
template:
metadata:
labels:
app.kubernetes.io/name: canned-prompts
app.kubernetes.io/part-of: canned-prompts
spec:
automountServiceAccountToken: false
serviceAccountName: canned-prompts
securityContext:
runAsNonRoot: true
runAsUser: 10001
runAsGroup: 10001
fsGroup: 10001
seccompProfile:
type: RuntimeDefault
containers:
- name: canned-prompts
image: forgejo.coulomb.social/coulomb/canned-prompts@sha256:14c7b92f20d63f2e70ea17b0b45d3c483521fbbaf14ee5bc277fa541e10452e2
imagePullPolicy: IfNotPresent
ports:
- name: http
containerPort: 8000
env:
# The URL arrives as a mounted file, never as an env var: an env var
# holding a password shows up in `kubectl describe`, in crash dumps,
# and to anything that can read /proc.
- name: CANNED_PROMPTS_DATABASE_URL_FILE
value: /var/run/secrets/postgres-runtime/url
- name: CANNED_PROMPTS_PUBLISH_TOKEN_FILE
value: /var/run/secrets/publish/token
- name: CANNED_PROMPTS_TENANT
value: railiance
securityContext:
allowPrivilegeEscalation: false
capabilities:
drop: ["ALL"]
readOnlyRootFilesystem: true
readinessProbe:
httpGet:
path: /readyz
port: http
periodSeconds: 5
livenessProbe:
# /healthz deliberately checks only that the process is up. Pointing
# liveness at a database-dependent path would restart every replica
# during a database blip, turning a brief outage into an outage plus
# a thundering herd.
httpGet:
path: /healthz
port: http
periodSeconds: 20
startupProbe:
httpGet:
path: /readyz
port: http
failureThreshold: 30
periodSeconds: 2
resources:
requests:
cpu: 25m
memory: 96Mi
limits:
cpu: 500m
memory: 384Mi
volumeMounts:
- name: postgres-runtime
mountPath: /var/run/secrets/postgres-runtime
readOnly: true
- name: publish
mountPath: /var/run/secrets/publish
readOnly: true
volumes:
- name: postgres-runtime
secret:
defaultMode: 0440
secretName: canned-prompts-postgres-runtime
items:
- key: url
path: url
- name: publish
secret:
defaultMode: 0440
secretName: canned-prompts-publish-token
optional: true
items:
- key: token
path: token
---
apiVersion: v1
kind: Service
metadata:
name: canned-prompts
namespace: canned-prompts
spec:
# ClusterIP only. No Ingress, no LoadBalancer: this registry is reachable
# from inside the cluster and nowhere else, which is what
# `private-service-only` in the smoke contract asserts.
type: ClusterIP
selector:
app.kubernetes.io/name: canned-prompts
ports:
- name: http
port: 8000
targetPort: http
---
apiVersion: v1
kind: ServiceAccount
metadata:
name: canned-prompts
namespace: canned-prompts
automountServiceAccountToken: false
---
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: canned-prompts-default-deny
namespace: canned-prompts
spec:
podSelector: {}
policyTypes: [Ingress, Egress]
---
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: canned-prompts-ingress
namespace: canned-prompts
spec:
# Ingress is for the serving pods only. The migration Job serves nothing and
# must not be reachable.
podSelector:
matchLabels:
app.kubernetes.io/name: canned-prompts
policyTypes: [Ingress]
ingress:
- from:
- namespaceSelector: {}
ports:
- protocol: TCP
port: 8000
---
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: canned-prompts-egress
namespace: canned-prompts
spec:
# Selects on part-of, not name, so it covers the migration Job as well as the
# serving pods. Selecting only `name: canned-prompts` left the Job matched by
# nothing but the default-deny — it succeeded once purely because it ran
# before these policies existed, and the next migration would have failed.
podSelector:
matchLabels:
app.kubernetes.io/part-of: canned-prompts
policyTypes: [Egress]
egress:
# PostgreSQL and DNS only. The service fetches nothing: a package arrives
# by publish, never by the registry reaching out, so it needs no egress to
# the internet and is not given any.
- to:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: databases
ports:
- protocol: TCP
port: 5432
- to:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: kube-system
ports:
- protocol: UDP
port: 53
- protocol: TCP
port: 53