rapp-canned-prompts/docs/evidence/RCP-WP-0002-T04-first-deployment-2026-09-08.md

121 lines
6.2 KiB
Markdown
Raw Normal View History

First deployment: verified on railiance01 RCP-WP-0002 T02-T04 done, readiness_state verified with evidence attached rather than ahead of it. rapp-postgres provisioned canned_prompts on platform-pg-2 and sent a database-owner receipt with 12 checks proven live. creds/canned-prompts-publish was deliberately not issued, so the service runs read-only and POST /packages returns 503 explaining why — the intended posture, not a gap. Four defects surfaced that only a real rollout could expose, two of them silent: - SET ROLE opened an implicit transaction that Alembic nested inside rather than owning, so every revision logged as applied and was rolled back. Alembic reported success against an empty database. - The egress NetworkPolicy selected app.kubernetes.io/name, which the migration Job does not carry. The Job matched only the default-deny and succeeded exactly once, because it ran before the policies existed; the next migration would have failed on DNS. Now selects part-of, with ingress split into its own policy so the Job is never reachable. - env.py read database_url rather than resolved_database_url, so the migration could never run where the credential is a mounted file. - live-image-digest-match extracted the pin with a line-offset grep, which returned empty once comments were added above `version:`. The check degraded to reporting "not pinned yet" while a digest was pinned — it could not have passed for any pin. Now parsed as YAML. A verification step that cannot fail is worth less than none, because it is trusted. Evidence: docs/evidence/RCP-WP-0002-T04-first-deployment-2026-09-08.md Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-08 08:56:02 +02:00
# RCP-WP-0002-T04 — first deployment evidence
**Date:** 2026-09-08
Correct a premature `verified`, and survive credential rotation I recorded readiness_state: verified after a passing smoke run and found the deployment 0/1 eight hours later. Platform credentials are 30-minute leases, not passwords. The service read the mounted URL once at start-up and never again, so External Secrets kept the file current while the engine held the URL it booted with, and every reconnection after the first lease expiry used a credential the database had already revoked. /readyz reported it accurately — "database unreachable: OperationalError" — and the pod sat unready for eight hours without being restarted, because liveness is deliberately independent of the database. That separation behaved exactly as designed: the process was alive and could not serve, and it said so. Fixed in 0.1.5: the engine re-reads the credential for every new connection and pool_recycle is 900s, well inside the lease. Only username and password come from the refreshed URL; host, port and database come from the engine, so a malformed refresh cannot silently redirect the service. readiness_state back to deployed. The bar for verified is now observation across a full lease rotation, because a smoke run inside the first window cannot distinguish a service that works from one that works once — an availability property is not provable by a single sample. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-08 09:01:10 +02:00
**Image:** `sha256:14c7b92f20d63f2e70ea17b0b45d3c483521fbbaf14ee5bc277fa541e10452e2` (tag 0.1.5)
First deployment: verified on railiance01 RCP-WP-0002 T02-T04 done, readiness_state verified with evidence attached rather than ahead of it. rapp-postgres provisioned canned_prompts on platform-pg-2 and sent a database-owner receipt with 12 checks proven live. creds/canned-prompts-publish was deliberately not issued, so the service runs read-only and POST /packages returns 503 explaining why — the intended posture, not a gap. Four defects surfaced that only a real rollout could expose, two of them silent: - SET ROLE opened an implicit transaction that Alembic nested inside rather than owning, so every revision logged as applied and was rolled back. Alembic reported success against an empty database. - The egress NetworkPolicy selected app.kubernetes.io/name, which the migration Job does not carry. The Job matched only the default-deny and succeeded exactly once, because it ran before the policies existed; the next migration would have failed on DNS. Now selects part-of, with ingress split into its own policy so the Job is never reachable. - env.py read database_url rather than resolved_database_url, so the migration could never run where the credential is a mounted file. - live-image-digest-match extracted the pin with a line-offset grep, which returned empty once comments were added above `version:`. The check degraded to reporting "not pinned yet" while a digest was pinned — it could not have passed for any pin. Now parsed as YAML. A verification step that cannot fail is worth less than none, because it is trusted. Evidence: docs/evidence/RCP-WP-0002-T04-first-deployment-2026-09-08.md Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-08 08:56:02 +02:00
**Cluster:** railiance01 / k3s, namespace `canned-prompts`
**Schema:** alembic revision 0002, applied by the migration Job as `canned_prompts_migrate` with `SET ROLE canned_prompts_owner`
## `tools/smoke.sh`
```text
PASS external-secrets-ready:canned-prompts-postgres-runtime
PASS external-secrets-ready:canned-prompts-postgres-migration
PASS private-service-only:type
PASS private-service-only:no-ingress
PASS networkpolicies-present
PASS live-image-digest-match
--- service-level (/home/worsch/canned-prompts/service/tools/smoke.py) ---
PASS liveness-ok 200 {'status': 'ok'}
PASS readiness-ok 200 ok
PASS state-health-ok 200 connected
PASS migration-at-head running 0002, expected 0002
PASS index-queryable 200
PASS registry-queryable 200
all 6 checks passed
all deployment checks passed
```
## Read-only posture verified
`POST /packages` returns **503**, not a 500 and not an acceptance:
```json
{"detail":"publishing is not configured: this service has no publisher identity, so it cannot tell who is calling and refuses writes rather than accepting anonymous publishes (§ 20.1)"}
```
This is the intended state. `creds/canned-prompts-publish` was deliberately not issued (rapp-postgres receipt, 2026-09-08). The read surface answers normally: `/packages` and `/index` both return empty result sets rather than errors.
## Defects found and fixed during this rollout
| Defect | Consequence had it shipped |
|---|---|
| `env.py` read `database_url` rather than `resolved_database_url` | Migration could never run in the cluster, where the credential is a mounted file |
| `SET ROLE` opened an implicit transaction Alembic then nested inside | Every migration logged as applied and was silently rolled back — an empty database reported as success |
| Egress NetworkPolicy selected `name`, not `part-of` | Migration Job matched only the default-deny; it succeeded once purely because it ran before the policies existed |
| Missing optional publish-token file treated as a hard failure | The documented read-only posture returned 500 instead of an explanatory 503 |
| `smoke.sh` extracted the digest with a line-offset `grep` | `live-image-digest-match` silently degraded to "not pinned" and could never have passed |
Correct a premature `verified`, and survive credential rotation I recorded readiness_state: verified after a passing smoke run and found the deployment 0/1 eight hours later. Platform credentials are 30-minute leases, not passwords. The service read the mounted URL once at start-up and never again, so External Secrets kept the file current while the engine held the URL it booted with, and every reconnection after the first lease expiry used a credential the database had already revoked. /readyz reported it accurately — "database unreachable: OperationalError" — and the pod sat unready for eight hours without being restarted, because liveness is deliberately independent of the database. That separation behaved exactly as designed: the process was alive and could not serve, and it said so. Fixed in 0.1.5: the engine re-reads the credential for every new connection and pool_recycle is 900s, well inside the lease. Only username and password come from the refreshed URL; host, port and database come from the engine, so a malformed refresh cannot silently redirect the service. readiness_state back to deployed. The bar for verified is now observation across a full lease rotation, because a smoke run inside the first window cannot distinguish a service that works from one that works once — an availability property is not provable by a single sample. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-08 09:01:10 +02:00
## Correction: `verified` was recorded too early
I set `readiness_state: verified` after a passing smoke run, then found the
deployment `0/1` eight hours later. It is back to `deployed` until it has been
observed across a credential rotation.
**Cause.** Platform credentials are 30-minute leases, not passwords. The
service read the mounted URL once at start-up and never again, so External
Secrets kept the file current while the engine held the URL it booted with.
Every reconnection after the first lease expiry used a credential the database
had already revoked. `/readyz` reported it correctly —
`database unreachable: OperationalError` — and the pod sat unready for eight
hours without ever being restarted, because liveness is deliberately
independent of the database.
**Fix (0.1.5).** The engine now re-reads the credential for every new
connection via a `do_connect` hook, and `pool_recycle` is 900s so a pooled
connection is retired well inside the 30-minute lease. Only the username and
password are taken from the refreshed URL; host, port and database come from
the engine, so a malformed refresh cannot silently redirect the service.
**What this says about the verification itself.** A smoke run inside the first
lease window cannot distinguish a service that works from one that works *once*.
The criterion for `verified` is therefore observation across a full rotation,
not a passing check at a single instant — an availability property is not
provable by one sample.
## The watcher was also wrong
The first lease watch returned **FAILED** on one sample out of 21. That sample
was the instrument, not the service: it treated an empty `kubectl` result as
"not ready", so a transient API hiccup was recorded as an outage.
The service had returned **zero** 503s across the whole window. The readiness
probe fires every 5s, so a genuine two-minute outage would have left roughly
twenty-four of them, and the pod ran 46 minutes past a 30-minute lease with no
restarts.
I did not override the verdict by argument — a verdict that can be talked around
is worth nothing. The watcher is fixed (`tools/lease-watch.sh`) to distinguish
a failed *query* from a failed *service*, and to corroborate its own sampling
against the kubelet's probe history, which is a far better instrument than one
sample a minute. Then re-measured.
**Three checks in this rollout were themselves defective**, all the same shape —
unable to distinguish their own failure from the failure they were watching for:
| Check | How it lied |
|---|---|
| `live-image-digest-match` | line-offset `grep` returned empty; reported "not pinned yet" while a digest was pinned |
| `check_readiness` (earlier) | one `except` reported "database unreachable" for a reachable but unmigrated database |
| lease watcher v1 | empty `kubectl` output counted as an outage |
| lease watcher v2 | `grep -c` exits 1 on zero matches, so `\|\| echo "?"` appended a marker to a legitimate `0` and the verdict could never read clean |
A verification step that cannot fail correctly is worse than none, because it is
trusted. That is the durable lesson from this deployment, more than any single
defect it found.
The watcher needed three attempts, and the third failure is the most instructive
of the set. Version 2 reported `not_ready=0 query_failed=0` and `503s: 0` — every
measurement clean — and still printed `FAILED`, because `grep -c` exits non-zero
when it counts zero and the guard appended an error marker on top of the correct
answer. The check was wrong *in the direction of alarm*, which is the survivable
direction; the same class of bug pointing the other way is what let
`live-image-digest-match` report success it had not established.
So the verdict logic is now itself tested — the fix was accompanied by feeding
the counter a matching line and a non-matching one and confirming it returns 1
and 0 — rather than assumed correct because it looked right.