First deployment: verified on railiance01
RCP-WP-0002 T02-T04 done, readiness_state verified with evidence attached
rather than ahead of it.
rapp-postgres provisioned canned_prompts on platform-pg-2 and sent a
database-owner receipt with 12 checks proven live. creds/canned-prompts-publish
was deliberately not issued, so the service runs read-only and POST /packages
returns 503 explaining why — the intended posture, not a gap.
Four defects surfaced that only a real rollout could expose, two of them
silent:
- SET ROLE opened an implicit transaction that Alembic nested inside rather
than owning, so every revision logged as applied and was rolled back.
Alembic reported success against an empty database.
- The egress NetworkPolicy selected app.kubernetes.io/name, which the
migration Job does not carry. The Job matched only the default-deny and
succeeded exactly once, because it ran before the policies existed; the next
migration would have failed on DNS. Now selects part-of, with ingress split
into its own policy so the Job is never reachable.
- env.py read database_url rather than resolved_database_url, so the migration
could never run where the credential is a mounted file.
- live-image-digest-match extracted the pin with a line-offset grep, which
returned empty once comments were added above `version:`. The check degraded
to reporting "not pinned yet" while a digest was pinned — it could not have
passed for any pin. Now parsed as YAML. A verification step that cannot fail
is worth less than none, because it is trusted.
Evidence: docs/evidence/RCP-WP-0002-T04-first-deployment-2026-09-08.md
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM
Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-08 08:56:02 +02:00
# RCP-WP-0002-T04 — first deployment evidence
**Date:** 2026-09-08
Correct a premature `verified`, and survive credential rotation
I recorded readiness_state: verified after a passing smoke run and found the
deployment 0/1 eight hours later.
Platform credentials are 30-minute leases, not passwords. The service read the
mounted URL once at start-up and never again, so External Secrets kept the file
current while the engine held the URL it booted with, and every reconnection
after the first lease expiry used a credential the database had already
revoked.
/readyz reported it accurately — "database unreachable: OperationalError" — and
the pod sat unready for eight hours without being restarted, because liveness
is deliberately independent of the database. That separation behaved exactly as
designed: the process was alive and could not serve, and it said so.
Fixed in 0.1.5: the engine re-reads the credential for every new connection and
pool_recycle is 900s, well inside the lease. Only username and password come
from the refreshed URL; host, port and database come from the engine, so a
malformed refresh cannot silently redirect the service.
readiness_state back to deployed. The bar for verified is now observation
across a full lease rotation, because a smoke run inside the first window
cannot distinguish a service that works from one that works once — an
availability property is not provable by a single sample.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM
Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-08 09:01:10 +02:00
**Image:** `sha256:14c7b92f20d63f2e70ea17b0b45d3c483521fbbaf14ee5bc277fa541e10452e2` (tag 0.1.5)
First deployment: verified on railiance01
RCP-WP-0002 T02-T04 done, readiness_state verified with evidence attached
rather than ahead of it.
rapp-postgres provisioned canned_prompts on platform-pg-2 and sent a
database-owner receipt with 12 checks proven live. creds/canned-prompts-publish
was deliberately not issued, so the service runs read-only and POST /packages
returns 503 explaining why — the intended posture, not a gap.
Four defects surfaced that only a real rollout could expose, two of them
silent:
- SET ROLE opened an implicit transaction that Alembic nested inside rather
than owning, so every revision logged as applied and was rolled back.
Alembic reported success against an empty database.
- The egress NetworkPolicy selected app.kubernetes.io/name, which the
migration Job does not carry. The Job matched only the default-deny and
succeeded exactly once, because it ran before the policies existed; the next
migration would have failed on DNS. Now selects part-of, with ingress split
into its own policy so the Job is never reachable.
- env.py read database_url rather than resolved_database_url, so the migration
could never run where the credential is a mounted file.
- live-image-digest-match extracted the pin with a line-offset grep, which
returned empty once comments were added above `version:`. The check degraded
to reporting "not pinned yet" while a digest was pinned — it could not have
passed for any pin. Now parsed as YAML. A verification step that cannot fail
is worth less than none, because it is trusted.
Evidence: docs/evidence/RCP-WP-0002-T04-first-deployment-2026-09-08.md
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM
Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-08 08:56:02 +02:00
**Cluster:** railiance01 / k3s, namespace `canned-prompts`
**Schema:** alembic revision 0002, applied by the migration Job as `canned_prompts_migrate` with `SET ROLE canned_prompts_owner`
## `tools/smoke.sh`
```text
PASS external-secrets-ready:canned-prompts-postgres-runtime
PASS external-secrets-ready:canned-prompts-postgres-migration
PASS private-service-only:type
PASS private-service-only:no-ingress
PASS networkpolicies-present
PASS live-image-digest-match
--- service-level (/home/worsch/canned-prompts/service/tools/smoke.py) ---
PASS liveness-ok 200 {'status': 'ok'}
PASS readiness-ok 200 ok
PASS state-health-ok 200 connected
PASS migration-at-head running 0002, expected 0002
PASS index-queryable 200
PASS registry-queryable 200
all 6 checks passed
all deployment checks passed
```
## Read-only posture verified
`POST /packages` returns **503** , not a 500 and not an acceptance:
```json
{"detail":"publishing is not configured: this service has no publisher identity, so it cannot tell who is calling and refuses writes rather than accepting anonymous publishes (§ 20.1)"}
```
This is the intended state. `creds/canned-prompts-publish` was deliberately not issued (rapp-postgres receipt, 2026-09-08). The read surface answers normally: `/packages` and `/index` both return empty result sets rather than errors.
## Defects found and fixed during this rollout
| Defect | Consequence had it shipped |
|---|---|
| `env.py` read `database_url` rather than `resolved_database_url` | Migration could never run in the cluster, where the credential is a mounted file |
| `SET ROLE` opened an implicit transaction Alembic then nested inside | Every migration logged as applied and was silently rolled back — an empty database reported as success |
| Egress NetworkPolicy selected `name` , not `part-of` | Migration Job matched only the default-deny; it succeeded once purely because it ran before the policies existed |
| Missing optional publish-token file treated as a hard failure | The documented read-only posture returned 500 instead of an explanatory 503 |
| `smoke.sh` extracted the digest with a line-offset `grep` | `live-image-digest-match` silently degraded to "not pinned" and could never have passed |
Correct a premature `verified`, and survive credential rotation
I recorded readiness_state: verified after a passing smoke run and found the
deployment 0/1 eight hours later.
Platform credentials are 30-minute leases, not passwords. The service read the
mounted URL once at start-up and never again, so External Secrets kept the file
current while the engine held the URL it booted with, and every reconnection
after the first lease expiry used a credential the database had already
revoked.
/readyz reported it accurately — "database unreachable: OperationalError" — and
the pod sat unready for eight hours without being restarted, because liveness
is deliberately independent of the database. That separation behaved exactly as
designed: the process was alive and could not serve, and it said so.
Fixed in 0.1.5: the engine re-reads the credential for every new connection and
pool_recycle is 900s, well inside the lease. Only username and password come
from the refreshed URL; host, port and database come from the engine, so a
malformed refresh cannot silently redirect the service.
readiness_state back to deployed. The bar for verified is now observation
across a full lease rotation, because a smoke run inside the first window
cannot distinguish a service that works from one that works once — an
availability property is not provable by a single sample.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM
Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-08 09:01:10 +02:00
## Correction: `verified` was recorded too early
I set `readiness_state: verified` after a passing smoke run, then found the
deployment `0/1` eight hours later. It is back to `deployed` until it has been
observed across a credential rotation.
**Cause.** Platform credentials are 30-minute leases, not passwords. The
service read the mounted URL once at start-up and never again, so External
Secrets kept the file current while the engine held the URL it booted with.
Every reconnection after the first lease expiry used a credential the database
had already revoked. `/readyz` reported it correctly —
`database unreachable: OperationalError` — and the pod sat unready for eight
hours without ever being restarted, because liveness is deliberately
independent of the database.
**Fix (0.1.5).** The engine now re-reads the credential for every new
connection via a `do_connect` hook, and `pool_recycle` is 900s so a pooled
connection is retired well inside the 30-minute lease. Only the username and
password are taken from the refreshed URL; host, port and database come from
the engine, so a malformed refresh cannot silently redirect the service.
**What this says about the verification itself.** A smoke run inside the first
lease window cannot distinguish a service that works from one that works *once* .
The criterion for `verified` is therefore observation across a full rotation,
not a passing check at a single instant — an availability property is not
provable by one sample.
2026-09-08 09:48:06 +02:00
## The watcher was also wrong
The first lease watch returned **FAILED** on one sample out of 21. That sample
was the instrument, not the service: it treated an empty `kubectl` result as
"not ready", so a transient API hiccup was recorded as an outage.
The service had returned **zero** 503s across the whole window. The readiness
probe fires every 5s, so a genuine two-minute outage would have left roughly
twenty-four of them, and the pod ran 46 minutes past a 30-minute lease with no
restarts.
I did not override the verdict by argument — a verdict that can be talked around
is worth nothing. The watcher is fixed (`tools/lease-watch.sh` ) to distinguish
a failed *query* from a failed *service* , and to corroborate its own sampling
against the kubelet's probe history, which is a far better instrument than one
sample a minute. Then re-measured.
**Three checks in this rollout were themselves defective**, all the same shape —
unable to distinguish their own failure from the failure they were watching for:
| Check | How it lied |
|---|---|
| `live-image-digest-match` | line-offset `grep` returned empty; reported "not pinned yet" while a digest was pinned |
| `check_readiness` (earlier) | one `except` reported "database unreachable" for a reachable but unmigrated database |
| lease watcher v1 | empty `kubectl` output counted as an outage |
2026-09-08 10:09:35 +02:00
| lease watcher v2 | `grep -c` exits 1 on zero matches, so `\|\| echo "?"` appended a marker to a legitimate `0` and the verdict could never read clean |
2026-09-08 09:48:06 +02:00
A verification step that cannot fail correctly is worse than none, because it is
trusted. That is the durable lesson from this deployment, more than any single
defect it found.
2026-09-08 10:09:35 +02:00
The watcher needed three attempts, and the third failure is the most instructive
of the set. Version 2 reported `not_ready=0 query_failed=0` and `503s: 0` — every
measurement clean — and still printed `FAILED` , because `grep -c` exits non-zero
when it counts zero and the guard appended an error marker on top of the correct
answer. The check was wrong *in the direction of alarm* , which is the survivable
direction; the same class of bug pointing the other way is what let
`live-image-digest-match` report success it had not established.
So the verdict logic is now itself tested — the fix was accompanied by feeding
the counter a matching line and a non-matching one and confirming it returns 1
and 0 — rather than assumed correct because it looked right.