# RCP-WP-0002-T04 — first deployment evidence **Date:** 2026-09-08 **Image:** `sha256:14c7b92f20d63f2e70ea17b0b45d3c483521fbbaf14ee5bc277fa541e10452e2` (tag 0.1.5) **Cluster:** railiance01 / k3s, namespace `canned-prompts` **Schema:** alembic revision 0002, applied by the migration Job as `canned_prompts_migrate` with `SET ROLE canned_prompts_owner` ## `tools/smoke.sh` ```text PASS external-secrets-ready:canned-prompts-postgres-runtime PASS external-secrets-ready:canned-prompts-postgres-migration PASS private-service-only:type PASS private-service-only:no-ingress PASS networkpolicies-present PASS live-image-digest-match --- service-level (/home/worsch/canned-prompts/service/tools/smoke.py) --- PASS liveness-ok 200 {'status': 'ok'} PASS readiness-ok 200 ok PASS state-health-ok 200 connected PASS migration-at-head running 0002, expected 0002 PASS index-queryable 200 PASS registry-queryable 200 all 6 checks passed all deployment checks passed ``` ## Read-only posture verified `POST /packages` returns **503**, not a 500 and not an acceptance: ```json {"detail":"publishing is not configured: this service has no publisher identity, so it cannot tell who is calling and refuses writes rather than accepting anonymous publishes (§ 20.1)"} ``` This is the intended state. `creds/canned-prompts-publish` was deliberately not issued (rapp-postgres receipt, 2026-09-08). The read surface answers normally: `/packages` and `/index` both return empty result sets rather than errors. ## Defects found and fixed during this rollout | Defect | Consequence had it shipped | |---|---| | `env.py` read `database_url` rather than `resolved_database_url` | Migration could never run in the cluster, where the credential is a mounted file | | `SET ROLE` opened an implicit transaction Alembic then nested inside | Every migration logged as applied and was silently rolled back — an empty database reported as success | | Egress NetworkPolicy selected `name`, not `part-of` | Migration Job matched only the default-deny; it succeeded once purely because it ran before the policies existed | | Missing optional publish-token file treated as a hard failure | The documented read-only posture returned 500 instead of an explanatory 503 | | `smoke.sh` extracted the digest with a line-offset `grep` | `live-image-digest-match` silently degraded to "not pinned" and could never have passed | ## Correction: `verified` was recorded too early I set `readiness_state: verified` after a passing smoke run, then found the deployment `0/1` eight hours later. It is back to `deployed` until it has been observed across a credential rotation. **Cause.** Platform credentials are 30-minute leases, not passwords. The service read the mounted URL once at start-up and never again, so External Secrets kept the file current while the engine held the URL it booted with. Every reconnection after the first lease expiry used a credential the database had already revoked. `/readyz` reported it correctly — `database unreachable: OperationalError` — and the pod sat unready for eight hours without ever being restarted, because liveness is deliberately independent of the database. **Fix (0.1.5).** The engine now re-reads the credential for every new connection via a `do_connect` hook, and `pool_recycle` is 900s so a pooled connection is retired well inside the 30-minute lease. Only the username and password are taken from the refreshed URL; host, port and database come from the engine, so a malformed refresh cannot silently redirect the service. **What this says about the verification itself.** A smoke run inside the first lease window cannot distinguish a service that works from one that works *once*. The criterion for `verified` is therefore observation across a full rotation, not a passing check at a single instant — an availability property is not provable by one sample. ## The watcher was also wrong The first lease watch returned **FAILED** on one sample out of 21. That sample was the instrument, not the service: it treated an empty `kubectl` result as "not ready", so a transient API hiccup was recorded as an outage. The service had returned **zero** 503s across the whole window. The readiness probe fires every 5s, so a genuine two-minute outage would have left roughly twenty-four of them, and the pod ran 46 minutes past a 30-minute lease with no restarts. I did not override the verdict by argument — a verdict that can be talked around is worth nothing. The watcher is fixed (`tools/lease-watch.sh`) to distinguish a failed *query* from a failed *service*, and to corroborate its own sampling against the kubelet's probe history, which is a far better instrument than one sample a minute. Then re-measured. **Three checks in this rollout were themselves defective**, all the same shape — unable to distinguish their own failure from the failure they were watching for: | Check | How it lied | |---|---| | `live-image-digest-match` | line-offset `grep` returned empty; reported "not pinned yet" while a digest was pinned | | `check_readiness` (earlier) | one `except` reported "database unreachable" for a reachable but unmigrated database | | lease watcher v1 | empty `kubectl` output counted as an outage | A verification step that cannot fail correctly is worse than none, because it is trusted. That is the durable lesson from this deployment, more than any single defect it found.