Version 2 measured everything correctly — not_ready=0, query_failed=0, zero
503s — and still printed FAILED. grep -c exits 1 when it counts zero matches,
so the '|| echo "?"' guard appended a marker on top of the legitimate 0 and
the verdict test could never match.
Third defect in the same instrument. The verdict logic is now tested rather
than assumed: the counter was fed a matching and a non-matching line and
confirmed to return 1 and 0.
Worth noting which direction each failure pointed. This one erred toward alarm,
which is survivable. live-image-digest-match erred the other way and reported
success it had not established — that is the one that would have shipped.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM
Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
The first lease watch returned FAILED on one sample of 21. That sample was the
instrument: it treated an empty kubectl result as not-ready, so a transient API
hiccup was recorded as an outage. The service returned zero 503s across the
window, with a readiness probe every 5s and 46 minutes of uptime past a
30-minute lease.
Not overriding the verdict by argument — a verdict that can be talked around is
worth nothing. The watcher now distinguishes a failed query from a failed
service and corroborates against the kubelet's probe history, which samples far
more often than once a minute. Promoted from scratch into tools/ so it is
reviewable and re-runnable.
Three checks in this rollout were defective in the same way: live-image-digest-match
degrading to 'not pinned' while a digest was pinned, check_readiness reporting
'database unreachable' for a reachable but unmigrated database, and this
watcher. A verification step that cannot fail correctly is worse than none,
because it is trusted.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM
Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
RCP-WP-0002 T02-T04 done, readiness_state verified with evidence attached
rather than ahead of it.
rapp-postgres provisioned canned_prompts on platform-pg-2 and sent a
database-owner receipt with 12 checks proven live. creds/canned-prompts-publish
was deliberately not issued, so the service runs read-only and POST /packages
returns 503 explaining why — the intended posture, not a gap.
Four defects surfaced that only a real rollout could expose, two of them
silent:
- SET ROLE opened an implicit transaction that Alembic nested inside rather
than owning, so every revision logged as applied and was rolled back.
Alembic reported success against an empty database.
- The egress NetworkPolicy selected app.kubernetes.io/name, which the
migration Job does not carry. The Job matched only the default-deny and
succeeded exactly once, because it ran before the policies existed; the next
migration would have failed on DNS. Now selects part-of, with ingress split
into its own policy so the Job is never reachable.
- env.py read database_url rather than resolved_database_url, so the migration
could never run where the credential is a mounted file.
- live-image-digest-match extracted the pin with a line-offset grep, which
returned empty once comments were added above `version:`. The check degraded
to reporting "not pinned yet" while a digest was pinned — it could not have
passed for any pin. Now parsed as YAML. A verification step that cannot fail
is worth less than none, because it is trusted.
Evidence: docs/evidence/RCP-WP-0002-T04-first-deployment-2026-09-08.md
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM
Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
Registers the repo with State Hub (agents / practice, prefix RCP-WP) and fills
in the rapp shape.
declarations/rapp.yaml declares a manifest-managed platform service owned by
canned-prompts, bound to rail-kubernetes and reef-railiance, with rollout,
smoke and rollback contracts.
The image pin says `pending-publication` rather than carrying a placeholder
digest. The image builds and was verified locally (canned-prompts
CANP-WP-0006-T06) but has never been pushed, so no registry digest exists. A
placeholder shaped like a real digest would be worse than a sentinel: it could
be mistaken for something deployable.
manifests/ follows the rapp-sbom-nexus shape: namespace labelled for the
postgres client, external secrets from OpenBao, a migration Job, and the
runtime Deployment with a ClusterIP-only Service, dedicated ServiceAccount and
default-deny plus runtime NetworkPolicies.
Three choices worth stating. Credentials arrive as mounted files, never env
vars — an env var holding a password is visible in kubectl describe, in crash
dumps, and to anything that can read /proc. Migrations run as a Job rather than
at start-up, so a schema rollback stays separate from a code rollback and
replicas do not race. Liveness points at /healthz, which checks only that the
process is up: pointing it at a database-dependent path would restart every
replica during a database blip.
Egress is PostgreSQL and DNS only. A package arrives by publish; the registry
never reaches out, so it is given no path to.
tools/smoke.sh checks what only the cluster can answer and calls
canned-prompts' service/tools/smoke.py for health and migration head, rather
than holding a second opinion about whether the service is healthy.
RCP-WP-0002 carries the two operator actions that block a first rollout —
publishing the image and provisioning database roles — and records
per-publisher identity as a decision belonging upstream, which this repo must
not paper over with cluster configuration implying finer control than exists.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM
Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502