rapp-canned-prompts/docs/evidence/RCP-WP-0002-T04-first-deployment-2026-09-08.md
tegwick caf3912d1b Fix the lease watcher, and record that three checks were themselves wrong
The first lease watch returned FAILED on one sample of 21. That sample was the
instrument: it treated an empty kubectl result as not-ready, so a transient API
hiccup was recorded as an outage. The service returned zero 503s across the
window, with a readiness probe every 5s and 46 minutes of uptime past a
30-minute lease.

Not overriding the verdict by argument — a verdict that can be talked around is
worth nothing. The watcher now distinguishes a failed query from a failed
service and corroborates against the kubelet's probe history, which samples far
more often than once a minute. Promoted from scratch into tools/ so it is
reviewable and re-runnable.

Three checks in this rollout were defective in the same way: live-image-digest-match
degrading to 'not pinned' while a digest was pinned, check_readiness reporting
'database unreachable' for a reachable but unmigrated database, and this
watcher. A verification step that cannot fail correctly is worse than none,
because it is trusted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-08 09:48:06 +02:00

5.3 KiB

RCP-WP-0002-T04 — first deployment evidence

Date: 2026-09-08 Image: sha256:14c7b92f20d63f2e70ea17b0b45d3c483521fbbaf14ee5bc277fa541e10452e2 (tag 0.1.5) Cluster: railiance01 / k3s, namespace canned-prompts Schema: alembic revision 0002, applied by the migration Job as canned_prompts_migrate with SET ROLE canned_prompts_owner

tools/smoke.sh

PASS  external-secrets-ready:canned-prompts-postgres-runtime
PASS  external-secrets-ready:canned-prompts-postgres-migration
PASS  private-service-only:type
PASS  private-service-only:no-ingress
PASS  networkpolicies-present
PASS  live-image-digest-match
--- service-level (/home/worsch/canned-prompts/service/tools/smoke.py) ---
PASS  liveness-ok         200 {'status': 'ok'}
PASS  readiness-ok        200 ok
PASS  state-health-ok     200 connected
PASS  migration-at-head   running 0002, expected 0002
PASS  index-queryable     200
PASS  registry-queryable  200

all 6 checks passed

all deployment checks passed

Read-only posture verified

POST /packages returns 503, not a 500 and not an acceptance:

{"detail":"publishing is not configured: this service has no publisher identity, so it cannot tell who is calling and refuses writes rather than accepting anonymous publishes (§ 20.1)"}

This is the intended state. creds/canned-prompts-publish was deliberately not issued (rapp-postgres receipt, 2026-09-08). The read surface answers normally: /packages and /index both return empty result sets rather than errors.

Defects found and fixed during this rollout

Defect Consequence had it shipped
env.py read database_url rather than resolved_database_url Migration could never run in the cluster, where the credential is a mounted file
SET ROLE opened an implicit transaction Alembic then nested inside Every migration logged as applied and was silently rolled back — an empty database reported as success
Egress NetworkPolicy selected name, not part-of Migration Job matched only the default-deny; it succeeded once purely because it ran before the policies existed
Missing optional publish-token file treated as a hard failure The documented read-only posture returned 500 instead of an explanatory 503
smoke.sh extracted the digest with a line-offset grep live-image-digest-match silently degraded to "not pinned" and could never have passed

Correction: verified was recorded too early

I set readiness_state: verified after a passing smoke run, then found the deployment 0/1 eight hours later. It is back to deployed until it has been observed across a credential rotation.

Cause. Platform credentials are 30-minute leases, not passwords. The service read the mounted URL once at start-up and never again, so External Secrets kept the file current while the engine held the URL it booted with. Every reconnection after the first lease expiry used a credential the database had already revoked. /readyz reported it correctly — database unreachable: OperationalError — and the pod sat unready for eight hours without ever being restarted, because liveness is deliberately independent of the database.

Fix (0.1.5). The engine now re-reads the credential for every new connection via a do_connect hook, and pool_recycle is 900s so a pooled connection is retired well inside the 30-minute lease. Only the username and password are taken from the refreshed URL; host, port and database come from the engine, so a malformed refresh cannot silently redirect the service.

What this says about the verification itself. A smoke run inside the first lease window cannot distinguish a service that works from one that works once. The criterion for verified is therefore observation across a full rotation, not a passing check at a single instant — an availability property is not provable by one sample.

The watcher was also wrong

The first lease watch returned FAILED on one sample out of 21. That sample was the instrument, not the service: it treated an empty kubectl result as "not ready", so a transient API hiccup was recorded as an outage.

The service had returned zero 503s across the whole window. The readiness probe fires every 5s, so a genuine two-minute outage would have left roughly twenty-four of them, and the pod ran 46 minutes past a 30-minute lease with no restarts.

I did not override the verdict by argument — a verdict that can be talked around is worth nothing. The watcher is fixed (tools/lease-watch.sh) to distinguish a failed query from a failed service, and to corroborate its own sampling against the kubelet's probe history, which is a far better instrument than one sample a minute. Then re-measured.

Three checks in this rollout were themselves defective, all the same shape — unable to distinguish their own failure from the failure they were watching for:

Check How it lied
live-image-digest-match line-offset grep returned empty; reported "not pinned yet" while a digest was pinned
check_readiness (earlier) one except reported "database unreachable" for a reachable but unmigrated database
lease watcher v1 empty kubectl output counted as an outage

A verification step that cannot fail correctly is worse than none, because it is trusted. That is the durable lesson from this deployment, more than any single defect it found.