Correct a premature verified, and survive credential rotation
I recorded readiness_state: verified after a passing smoke run and found the deployment 0/1 eight hours later. Platform credentials are 30-minute leases, not passwords. The service read the mounted URL once at start-up and never again, so External Secrets kept the file current while the engine held the URL it booted with, and every reconnection after the first lease expiry used a credential the database had already revoked. /readyz reported it accurately — "database unreachable: OperationalError" — and the pod sat unready for eight hours without being restarted, because liveness is deliberately independent of the database. That separation behaved exactly as designed: the process was alive and could not serve, and it said so. Fixed in 0.1.5: the engine re-reads the credential for every new connection and pool_recycle is 900s, well inside the lease. Only username and password come from the refreshed URL; host, port and database come from the engine, so a malformed refresh cannot silently redirect the service. readiness_state back to deployed. The bar for verified is now observation across a full lease rotation, because a smoke run inside the first window cannot distinguish a service that works from one that works once — an availability property is not provable by a single sample. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
This commit is contained in:
parent
0cca7e80f5
commit
1b86e872a4
5 changed files with 47 additions and 10 deletions
|
|
@ -1,7 +1,7 @@
|
|||
# RCP-WP-0002-T04 — first deployment evidence
|
||||
|
||||
**Date:** 2026-09-08
|
||||
**Image:** `sha256:2fbac3c0d1609a76d9478e540ae5450c5425e8c04e0565cccc6a4eb0180ebcaf` (tag 0.1.4)
|
||||
**Image:** `sha256:14c7b92f20d63f2e70ea17b0b45d3c483521fbbaf14ee5bc277fa541e10452e2` (tag 0.1.5)
|
||||
**Cluster:** railiance01 / k3s, namespace `canned-prompts`
|
||||
**Schema:** alembic revision 0002, applied by the migration Job as `canned_prompts_migrate` with `SET ROLE canned_prompts_owner`
|
||||
|
||||
|
|
@ -46,3 +46,31 @@ This is the intended state. `creds/canned-prompts-publish` was deliberately not
|
|||
| Egress NetworkPolicy selected `name`, not `part-of` | Migration Job matched only the default-deny; it succeeded once purely because it ran before the policies existed |
|
||||
| Missing optional publish-token file treated as a hard failure | The documented read-only posture returned 500 instead of an explanatory 503 |
|
||||
| `smoke.sh` extracted the digest with a line-offset `grep` | `live-image-digest-match` silently degraded to "not pinned" and could never have passed |
|
||||
|
||||
|
||||
## Correction: `verified` was recorded too early
|
||||
|
||||
I set `readiness_state: verified` after a passing smoke run, then found the
|
||||
deployment `0/1` eight hours later. It is back to `deployed` until it has been
|
||||
observed across a credential rotation.
|
||||
|
||||
**Cause.** Platform credentials are 30-minute leases, not passwords. The
|
||||
service read the mounted URL once at start-up and never again, so External
|
||||
Secrets kept the file current while the engine held the URL it booted with.
|
||||
Every reconnection after the first lease expiry used a credential the database
|
||||
had already revoked. `/readyz` reported it correctly —
|
||||
`database unreachable: OperationalError` — and the pod sat unready for eight
|
||||
hours without ever being restarted, because liveness is deliberately
|
||||
independent of the database.
|
||||
|
||||
**Fix (0.1.5).** The engine now re-reads the credential for every new
|
||||
connection via a `do_connect` hook, and `pool_recycle` is 900s so a pooled
|
||||
connection is retired well inside the 30-minute lease. Only the username and
|
||||
password are taken from the refreshed URL; host, port and database come from
|
||||
the engine, so a malformed refresh cannot silently redirect the service.
|
||||
|
||||
**What this says about the verification itself.** A smoke run inside the first
|
||||
lease window cannot distinguish a service that works from one that works *once*.
|
||||
The criterion for `verified` is therefore observation across a full rotation,
|
||||
not a passing check at a single instant — an availability property is not
|
||||
provable by one sample.
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue