rapp-canned-prompts/docs/evidence/RCP-WP-0002-T04-first-deployment-2026-09-08.md
tegwick 1b86e872a4 Correct a premature verified, and survive credential rotation
I recorded readiness_state: verified after a passing smoke run and found the
deployment 0/1 eight hours later.

Platform credentials are 30-minute leases, not passwords. The service read the
mounted URL once at start-up and never again, so External Secrets kept the file
current while the engine held the URL it booted with, and every reconnection
after the first lease expiry used a credential the database had already
revoked.

/readyz reported it accurately — "database unreachable: OperationalError" — and
the pod sat unready for eight hours without being restarted, because liveness
is deliberately independent of the database. That separation behaved exactly as
designed: the process was alive and could not serve, and it said so.

Fixed in 0.1.5: the engine re-reads the credential for every new connection and
pool_recycle is 900s, well inside the lease. Only username and password come
from the refreshed URL; host, port and database come from the engine, so a
malformed refresh cannot silently redirect the service.

readiness_state back to deployed. The bar for verified is now observation
across a full lease rotation, because a smoke run inside the first window
cannot distinguish a service that works from one that works once — an
availability property is not provable by a single sample.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-08 09:01:10 +02:00

3.8 KiB

RCP-WP-0002-T04 — first deployment evidence

Date: 2026-09-08 Image: sha256:14c7b92f20d63f2e70ea17b0b45d3c483521fbbaf14ee5bc277fa541e10452e2 (tag 0.1.5) Cluster: railiance01 / k3s, namespace canned-prompts Schema: alembic revision 0002, applied by the migration Job as canned_prompts_migrate with SET ROLE canned_prompts_owner

tools/smoke.sh

PASS  external-secrets-ready:canned-prompts-postgres-runtime
PASS  external-secrets-ready:canned-prompts-postgres-migration
PASS  private-service-only:type
PASS  private-service-only:no-ingress
PASS  networkpolicies-present
PASS  live-image-digest-match
--- service-level (/home/worsch/canned-prompts/service/tools/smoke.py) ---
PASS  liveness-ok         200 {'status': 'ok'}
PASS  readiness-ok        200 ok
PASS  state-health-ok     200 connected
PASS  migration-at-head   running 0002, expected 0002
PASS  index-queryable     200
PASS  registry-queryable  200

all 6 checks passed

all deployment checks passed

Read-only posture verified

POST /packages returns 503, not a 500 and not an acceptance:

{"detail":"publishing is not configured: this service has no publisher identity, so it cannot tell who is calling and refuses writes rather than accepting anonymous publishes (§ 20.1)"}

This is the intended state. creds/canned-prompts-publish was deliberately not issued (rapp-postgres receipt, 2026-09-08). The read surface answers normally: /packages and /index both return empty result sets rather than errors.

Defects found and fixed during this rollout

Defect Consequence had it shipped
env.py read database_url rather than resolved_database_url Migration could never run in the cluster, where the credential is a mounted file
SET ROLE opened an implicit transaction Alembic then nested inside Every migration logged as applied and was silently rolled back — an empty database reported as success
Egress NetworkPolicy selected name, not part-of Migration Job matched only the default-deny; it succeeded once purely because it ran before the policies existed
Missing optional publish-token file treated as a hard failure The documented read-only posture returned 500 instead of an explanatory 503
smoke.sh extracted the digest with a line-offset grep live-image-digest-match silently degraded to "not pinned" and could never have passed

Correction: verified was recorded too early

I set readiness_state: verified after a passing smoke run, then found the deployment 0/1 eight hours later. It is back to deployed until it has been observed across a credential rotation.

Cause. Platform credentials are 30-minute leases, not passwords. The service read the mounted URL once at start-up and never again, so External Secrets kept the file current while the engine held the URL it booted with. Every reconnection after the first lease expiry used a credential the database had already revoked. /readyz reported it correctly — database unreachable: OperationalError — and the pod sat unready for eight hours without ever being restarted, because liveness is deliberately independent of the database.

Fix (0.1.5). The engine now re-reads the credential for every new connection via a do_connect hook, and pool_recycle is 900s so a pooled connection is retired well inside the 30-minute lease. Only the username and password are taken from the refreshed URL; host, port and database come from the engine, so a malformed refresh cannot silently redirect the service.

What this says about the verification itself. A smoke run inside the first lease window cannot distinguish a service that works from one that works once. The criterion for verified is therefore observation across a full rotation, not a passing check at a single instant — an availability property is not provable by one sample.