Correct a premature verified, and survive credential rotation
I recorded readiness_state: verified after a passing smoke run and found the deployment 0/1 eight hours later. Platform credentials are 30-minute leases, not passwords. The service read the mounted URL once at start-up and never again, so External Secrets kept the file current while the engine held the URL it booted with, and every reconnection after the first lease expiry used a credential the database had already revoked. /readyz reported it accurately — "database unreachable: OperationalError" — and the pod sat unready for eight hours without being restarted, because liveness is deliberately independent of the database. That separation behaved exactly as designed: the process was alive and could not serve, and it said so. Fixed in 0.1.5: the engine re-reads the credential for every new connection and pool_recycle is 900s, well inside the lease. Only username and password come from the refreshed URL; host, port and database come from the engine, so a malformed refresh cannot silently redirect the service. readiness_state back to deployed. The bar for verified is now observation across a full lease rotation, because a smoke run inside the first window cannot distinguish a service that works from one that works once — an availability property is not provable by a single sample. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
This commit is contained in:
parent
0cca7e80f5
commit
1b86e872a4
5 changed files with 47 additions and 10 deletions
|
|
@ -4,7 +4,7 @@ type: workplan
|
|||
title: "First deployment of canned-prompts on Railiance"
|
||||
domain: agents
|
||||
repo: rapp-canned-prompts
|
||||
status: finished
|
||||
status: active
|
||||
owner: codex
|
||||
topic_slug: practice
|
||||
created: "2026-09-06"
|
||||
|
|
@ -147,7 +147,7 @@ re-run ownership reconciliation. Notified.
|
|||
|
||||
```task
|
||||
id: RCP-WP-0002-T04
|
||||
status: done
|
||||
status: progress
|
||||
priority: high
|
||||
state_hub_task_id: "6d9eb97c-57e2-5b74-b6eb-455076713417"
|
||||
```
|
||||
|
|
@ -164,8 +164,17 @@ Record the output as evidence, then move `readiness_state` to `deployed`, and to
|
|||
**Done 2026-09-08.** All six deployment checks and all six service-level checks
|
||||
pass; evidence at
|
||||
`docs/evidence/RCP-WP-0002-T04-first-deployment-2026-09-08.md`.
|
||||
`readiness_state` is `verified`, with the evidence attached rather than ahead of
|
||||
it.
|
||||
**Then corrected.** I recorded `verified` after a passing smoke run and found
|
||||
the deployment `0/1` eight hours later: platform credentials are 30-minute
|
||||
leases, and the service read the mounted URL once at start-up. External Secrets
|
||||
kept the file current; the engine held the URL it booted with. Fixed in 0.1.5 by
|
||||
re-reading the credential for every new connection, with `pool_recycle` inside
|
||||
the lease TTL.
|
||||
|
||||
`readiness_state` is back to `deployed`. The bar for `verified` is now
|
||||
observation across a full lease rotation, because a smoke run inside the first
|
||||
window cannot tell a service that works from one that works *once* — an
|
||||
availability property is not provable by a single sample.
|
||||
|
||||
**The check that could never have passed.** `live-image-digest-match` read the
|
||||
pin with a line-offset `grep`, which returned empty once comments were added
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue