From 1b86e872a4c0f334cdf2fed9649a5c5ad9db2282 Mon Sep 17 00:00:00 2001 From: tegwick Date: Tue, 8 Sep 2026 09:01:10 +0200 Subject: [PATCH] Correct a premature `verified`, and survive credential rotation MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit I recorded readiness_state: verified after a passing smoke run and found the deployment 0/1 eight hours later. Platform credentials are 30-minute leases, not passwords. The service read the mounted URL once at start-up and never again, so External Secrets kept the file current while the engine held the URL it booted with, and every reconnection after the first lease expiry used a credential the database had already revoked. /readyz reported it accurately — "database unreachable: OperationalError" — and the pod sat unready for eight hours without being restarted, because liveness is deliberately independent of the database. That separation behaved exactly as designed: the process was alive and could not serve, and it said so. Fixed in 0.1.5: the engine re-reads the credential for every new connection and pool_recycle is 900s, well inside the lease. Only username and password come from the refreshed URL; host, port and database come from the engine, so a malformed refresh cannot silently redirect the service. readiness_state back to deployed. The bar for verified is now observation across a full lease rotation, because a smoke run inside the first window cannot distinguish a service that works from one that works once — an availability property is not provable by a single sample. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502 --- declarations/rapp.yaml | 6 ++-- ...WP-0002-T04-first-deployment-2026-09-08.md | 30 ++++++++++++++++++- manifests/migration.yaml | 2 +- manifests/runtime.yaml | 2 +- workplans/RCP-WP-0002-first-deployment.md | 17 ++++++++--- 5 files changed, 47 insertions(+), 10 deletions(-) diff --git a/declarations/rapp.yaml b/declarations/rapp.yaml index 5d55cf6..d6890bd 100644 --- a/declarations/rapp.yaml +++ b/declarations/rapp.yaml @@ -4,7 +4,7 @@ rapp_id: rapp-canned-prompts repo: rapp-canned-prompts ownership_repo: canned-prompts contract_version: 1.0.0 -readiness_state: verified +readiness_state: deployed workload_identity: name: canned-prompts package_type: manifest-managed-platform-service @@ -34,11 +34,11 @@ composition: upstream_components: - name: canned-prompts source: forgejo.coulomb.social/coulomb/canned-prompts - # Published 2026-09-07 from canned-prompts service/Dockerfile, tag 0.1.4. + # Published 2026-09-07 from canned-prompts service/Dockerfile, tag 0.1.5. # Pinned by digest rather than tag: a tag can be moved, and # live-image-digest-match would then pass against something that is no # longer what this repo reviewed. - version: sha256:2fbac3c0d1609a76d9478e540ae5450c5425e8c04e0565cccc6a4eb0180ebcaf + version: sha256:14c7b92f20d63f2e70ea17b0b45d3c483521fbbaf14ee5bc277fa541e10452e2 rollout_contract: default_mode: kubectl-server-side-apply smoke_contract: diff --git a/docs/evidence/RCP-WP-0002-T04-first-deployment-2026-09-08.md b/docs/evidence/RCP-WP-0002-T04-first-deployment-2026-09-08.md index ff60f7a..fa0d0df 100644 --- a/docs/evidence/RCP-WP-0002-T04-first-deployment-2026-09-08.md +++ b/docs/evidence/RCP-WP-0002-T04-first-deployment-2026-09-08.md @@ -1,7 +1,7 @@ # RCP-WP-0002-T04 — first deployment evidence **Date:** 2026-09-08 -**Image:** `sha256:2fbac3c0d1609a76d9478e540ae5450c5425e8c04e0565cccc6a4eb0180ebcaf` (tag 0.1.4) +**Image:** `sha256:14c7b92f20d63f2e70ea17b0b45d3c483521fbbaf14ee5bc277fa541e10452e2` (tag 0.1.5) **Cluster:** railiance01 / k3s, namespace `canned-prompts` **Schema:** alembic revision 0002, applied by the migration Job as `canned_prompts_migrate` with `SET ROLE canned_prompts_owner` @@ -46,3 +46,31 @@ This is the intended state. `creds/canned-prompts-publish` was deliberately not | Egress NetworkPolicy selected `name`, not `part-of` | Migration Job matched only the default-deny; it succeeded once purely because it ran before the policies existed | | Missing optional publish-token file treated as a hard failure | The documented read-only posture returned 500 instead of an explanatory 503 | | `smoke.sh` extracted the digest with a line-offset `grep` | `live-image-digest-match` silently degraded to "not pinned" and could never have passed | + + +## Correction: `verified` was recorded too early + +I set `readiness_state: verified` after a passing smoke run, then found the +deployment `0/1` eight hours later. It is back to `deployed` until it has been +observed across a credential rotation. + +**Cause.** Platform credentials are 30-minute leases, not passwords. The +service read the mounted URL once at start-up and never again, so External +Secrets kept the file current while the engine held the URL it booted with. +Every reconnection after the first lease expiry used a credential the database +had already revoked. `/readyz` reported it correctly — +`database unreachable: OperationalError` — and the pod sat unready for eight +hours without ever being restarted, because liveness is deliberately +independent of the database. + +**Fix (0.1.5).** The engine now re-reads the credential for every new +connection via a `do_connect` hook, and `pool_recycle` is 900s so a pooled +connection is retired well inside the 30-minute lease. Only the username and +password are taken from the refreshed URL; host, port and database come from +the engine, so a malformed refresh cannot silently redirect the service. + +**What this says about the verification itself.** A smoke run inside the first +lease window cannot distinguish a service that works from one that works *once*. +The criterion for `verified` is therefore observation across a full rotation, +not a passing check at a single instant — an availability property is not +provable by one sample. diff --git a/manifests/migration.yaml b/manifests/migration.yaml index 6c3bbf7..5c78fcf 100644 --- a/manifests/migration.yaml +++ b/manifests/migration.yaml @@ -33,7 +33,7 @@ spec: type: RuntimeDefault containers: - name: migrate - image: forgejo.coulomb.social/coulomb/canned-prompts@sha256:2fbac3c0d1609a76d9478e540ae5450c5425e8c04e0565cccc6a4eb0180ebcaf + image: forgejo.coulomb.social/coulomb/canned-prompts@sha256:14c7b92f20d63f2e70ea17b0b45d3c483521fbbaf14ee5bc277fa541e10452e2 command: ["alembic"] args: ["upgrade", "head"] workingDir: /app diff --git a/manifests/runtime.yaml b/manifests/runtime.yaml index de102f1..2bababb 100644 --- a/manifests/runtime.yaml +++ b/manifests/runtime.yaml @@ -27,7 +27,7 @@ spec: type: RuntimeDefault containers: - name: canned-prompts - image: forgejo.coulomb.social/coulomb/canned-prompts@sha256:2fbac3c0d1609a76d9478e540ae5450c5425e8c04e0565cccc6a4eb0180ebcaf + image: forgejo.coulomb.social/coulomb/canned-prompts@sha256:14c7b92f20d63f2e70ea17b0b45d3c483521fbbaf14ee5bc277fa541e10452e2 imagePullPolicy: IfNotPresent ports: - name: http diff --git a/workplans/RCP-WP-0002-first-deployment.md b/workplans/RCP-WP-0002-first-deployment.md index 1978050..e7ef435 100644 --- a/workplans/RCP-WP-0002-first-deployment.md +++ b/workplans/RCP-WP-0002-first-deployment.md @@ -4,7 +4,7 @@ type: workplan title: "First deployment of canned-prompts on Railiance" domain: agents repo: rapp-canned-prompts -status: finished +status: active owner: codex topic_slug: practice created: "2026-09-06" @@ -147,7 +147,7 @@ re-run ownership reconciliation. Notified. ```task id: RCP-WP-0002-T04 -status: done +status: progress priority: high state_hub_task_id: "6d9eb97c-57e2-5b74-b6eb-455076713417" ``` @@ -164,8 +164,17 @@ Record the output as evidence, then move `readiness_state` to `deployed`, and to **Done 2026-09-08.** All six deployment checks and all six service-level checks pass; evidence at `docs/evidence/RCP-WP-0002-T04-first-deployment-2026-09-08.md`. -`readiness_state` is `verified`, with the evidence attached rather than ahead of -it. +**Then corrected.** I recorded `verified` after a passing smoke run and found +the deployment `0/1` eight hours later: platform credentials are 30-minute +leases, and the service read the mounted URL once at start-up. External Secrets +kept the file current; the engine held the URL it booted with. Fixed in 0.1.5 by +re-reading the credential for every new connection, with `pool_recycle` inside +the lease TTL. + +`readiness_state` is back to `deployed`. The bar for `verified` is now +observation across a full lease rotation, because a smoke run inside the first +window cannot tell a service that works from one that works *once* — an +availability property is not provable by a single sample. **The check that could never have passed.** `live-image-digest-match` read the pin with a line-offset `grep`, which returned empty once comments were added