Correct a premature verified, and survive credential rotation

I recorded readiness_state: verified after a passing smoke run and found the
deployment 0/1 eight hours later.

Platform credentials are 30-minute leases, not passwords. The service read the
mounted URL once at start-up and never again, so External Secrets kept the file
current while the engine held the URL it booted with, and every reconnection
after the first lease expiry used a credential the database had already
revoked.

/readyz reported it accurately — "database unreachable: OperationalError" — and
the pod sat unready for eight hours without being restarted, because liveness
is deliberately independent of the database. That separation behaved exactly as
designed: the process was alive and could not serve, and it said so.

Fixed in 0.1.5: the engine re-reads the credential for every new connection and
pool_recycle is 900s, well inside the lease. Only username and password come
from the refreshed URL; host, port and database come from the engine, so a
malformed refresh cannot silently redirect the service.

readiness_state back to deployed. The bar for verified is now observation
across a full lease rotation, because a smoke run inside the first window
cannot distinguish a service that works from one that works once — an
availability property is not provable by a single sample.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
This commit is contained in:
tegwick 2026-09-08 09:01:10 +02:00
parent 0cca7e80f5
commit 1b86e872a4
5 changed files with 47 additions and 10 deletions

View file

@ -4,7 +4,7 @@ rapp_id: rapp-canned-prompts
repo: rapp-canned-prompts
ownership_repo: canned-prompts
contract_version: 1.0.0
readiness_state: verified
readiness_state: deployed
workload_identity:
name: canned-prompts
package_type: manifest-managed-platform-service
@ -34,11 +34,11 @@ composition:
upstream_components:
- name: canned-prompts
source: forgejo.coulomb.social/coulomb/canned-prompts
# Published 2026-09-07 from canned-prompts service/Dockerfile, tag 0.1.4.
# Published 2026-09-07 from canned-prompts service/Dockerfile, tag 0.1.5.
# Pinned by digest rather than tag: a tag can be moved, and
# live-image-digest-match would then pass against something that is no
# longer what this repo reviewed.
version: sha256:2fbac3c0d1609a76d9478e540ae5450c5425e8c04e0565cccc6a4eb0180ebcaf
version: sha256:14c7b92f20d63f2e70ea17b0b45d3c483521fbbaf14ee5bc277fa541e10452e2
rollout_contract:
default_mode: kubectl-server-side-apply
smoke_contract:

View file

@ -1,7 +1,7 @@
# RCP-WP-0002-T04 — first deployment evidence
**Date:** 2026-09-08
**Image:** `sha256:2fbac3c0d1609a76d9478e540ae5450c5425e8c04e0565cccc6a4eb0180ebcaf` (tag 0.1.4)
**Image:** `sha256:14c7b92f20d63f2e70ea17b0b45d3c483521fbbaf14ee5bc277fa541e10452e2` (tag 0.1.5)
**Cluster:** railiance01 / k3s, namespace `canned-prompts`
**Schema:** alembic revision 0002, applied by the migration Job as `canned_prompts_migrate` with `SET ROLE canned_prompts_owner`
@ -46,3 +46,31 @@ This is the intended state. `creds/canned-prompts-publish` was deliberately not
| Egress NetworkPolicy selected `name`, not `part-of` | Migration Job matched only the default-deny; it succeeded once purely because it ran before the policies existed |
| Missing optional publish-token file treated as a hard failure | The documented read-only posture returned 500 instead of an explanatory 503 |
| `smoke.sh` extracted the digest with a line-offset `grep` | `live-image-digest-match` silently degraded to "not pinned" and could never have passed |
## Correction: `verified` was recorded too early
I set `readiness_state: verified` after a passing smoke run, then found the
deployment `0/1` eight hours later. It is back to `deployed` until it has been
observed across a credential rotation.
**Cause.** Platform credentials are 30-minute leases, not passwords. The
service read the mounted URL once at start-up and never again, so External
Secrets kept the file current while the engine held the URL it booted with.
Every reconnection after the first lease expiry used a credential the database
had already revoked. `/readyz` reported it correctly —
`database unreachable: OperationalError` — and the pod sat unready for eight
hours without ever being restarted, because liveness is deliberately
independent of the database.
**Fix (0.1.5).** The engine now re-reads the credential for every new
connection via a `do_connect` hook, and `pool_recycle` is 900s so a pooled
connection is retired well inside the 30-minute lease. Only the username and
password are taken from the refreshed URL; host, port and database come from
the engine, so a malformed refresh cannot silently redirect the service.
**What this says about the verification itself.** A smoke run inside the first
lease window cannot distinguish a service that works from one that works *once*.
The criterion for `verified` is therefore observation across a full rotation,
not a passing check at a single instant — an availability property is not
provable by one sample.

View file

@ -33,7 +33,7 @@ spec:
type: RuntimeDefault
containers:
- name: migrate
image: forgejo.coulomb.social/coulomb/canned-prompts@sha256:2fbac3c0d1609a76d9478e540ae5450c5425e8c04e0565cccc6a4eb0180ebcaf
image: forgejo.coulomb.social/coulomb/canned-prompts@sha256:14c7b92f20d63f2e70ea17b0b45d3c483521fbbaf14ee5bc277fa541e10452e2
command: ["alembic"]
args: ["upgrade", "head"]
workingDir: /app

View file

@ -27,7 +27,7 @@ spec:
type: RuntimeDefault
containers:
- name: canned-prompts
image: forgejo.coulomb.social/coulomb/canned-prompts@sha256:2fbac3c0d1609a76d9478e540ae5450c5425e8c04e0565cccc6a4eb0180ebcaf
image: forgejo.coulomb.social/coulomb/canned-prompts@sha256:14c7b92f20d63f2e70ea17b0b45d3c483521fbbaf14ee5bc277fa541e10452e2
imagePullPolicy: IfNotPresent
ports:
- name: http

View file

@ -4,7 +4,7 @@ type: workplan
title: "First deployment of canned-prompts on Railiance"
domain: agents
repo: rapp-canned-prompts
status: finished
status: active
owner: codex
topic_slug: practice
created: "2026-09-06"
@ -147,7 +147,7 @@ re-run ownership reconciliation. Notified.
```task
id: RCP-WP-0002-T04
status: done
status: progress
priority: high
state_hub_task_id: "6d9eb97c-57e2-5b74-b6eb-455076713417"
```
@ -164,8 +164,17 @@ Record the output as evidence, then move `readiness_state` to `deployed`, and to
**Done 2026-09-08.** All six deployment checks and all six service-level checks
pass; evidence at
`docs/evidence/RCP-WP-0002-T04-first-deployment-2026-09-08.md`.
`readiness_state` is `verified`, with the evidence attached rather than ahead of
it.
**Then corrected.** I recorded `verified` after a passing smoke run and found
the deployment `0/1` eight hours later: platform credentials are 30-minute
leases, and the service read the mounted URL once at start-up. External Secrets
kept the file current; the engine held the URL it booted with. Fixed in 0.1.5 by
re-reading the credential for every new connection, with `pool_recycle` inside
the lease TTL.
`readiness_state` is back to `deployed`. The bar for `verified` is now
observation across a full lease rotation, because a smoke run inside the first
window cannot tell a service that works from one that works *once* — an
availability property is not provable by a single sample.
**The check that could never have passed.** `live-image-digest-match` read the
pin with a line-offset `grep`, which returned empty once comments were added