Commit graph

3 commits

Author SHA1 Message Date
caf3912d1b Fix the lease watcher, and record that three checks were themselves wrong
The first lease watch returned FAILED on one sample of 21. That sample was the
instrument: it treated an empty kubectl result as not-ready, so a transient API
hiccup was recorded as an outage. The service returned zero 503s across the
window, with a readiness probe every 5s and 46 minutes of uptime past a
30-minute lease.

Not overriding the verdict by argument — a verdict that can be talked around is
worth nothing. The watcher now distinguishes a failed query from a failed
service and corroborates against the kubelet's probe history, which samples far
more often than once a minute. Promoted from scratch into tools/ so it is
reviewable and re-runnable.

Three checks in this rollout were defective in the same way: live-image-digest-match
degrading to 'not pinned' while a digest was pinned, check_readiness reporting
'database unreachable' for a reachable but unmigrated database, and this
watcher. A verification step that cannot fail correctly is worse than none,
because it is trusted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-08 09:48:06 +02:00
4d5af7698d First deployment: verified on railiance01
RCP-WP-0002 T02-T04 done, readiness_state verified with evidence attached
rather than ahead of it.

rapp-postgres provisioned canned_prompts on platform-pg-2 and sent a
database-owner receipt with 12 checks proven live. creds/canned-prompts-publish
was deliberately not issued, so the service runs read-only and POST /packages
returns 503 explaining why — the intended posture, not a gap.

Four defects surfaced that only a real rollout could expose, two of them
silent:

- SET ROLE opened an implicit transaction that Alembic nested inside rather
  than owning, so every revision logged as applied and was rolled back.
  Alembic reported success against an empty database.
- The egress NetworkPolicy selected app.kubernetes.io/name, which the
  migration Job does not carry. The Job matched only the default-deny and
  succeeded exactly once, because it ran before the policies existed; the next
  migration would have failed on DNS. Now selects part-of, with ingress split
  into its own policy so the Job is never reachable.
- env.py read database_url rather than resolved_database_url, so the migration
  could never run where the credential is a mounted file.
- live-image-digest-match extracted the pin with a line-offset grep, which
  returned empty once comments were added above `version:`. The check degraded
  to reporting "not pinned yet" while a digest was pinned — it could not have
  passed for any pin. Now parsed as YAML. A verification step that cannot fail
  is worth less than none, because it is trusted.

Evidence: docs/evidence/RCP-WP-0002-T04-first-deployment-2026-09-08.md

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-08 08:56:02 +02:00
5a0f4cb5c8 Package canned-prompts for Railiance
Registers the repo with State Hub (agents / practice, prefix RCP-WP) and fills
in the rapp shape.

declarations/rapp.yaml declares a manifest-managed platform service owned by
canned-prompts, bound to rail-kubernetes and reef-railiance, with rollout,
smoke and rollback contracts.

The image pin says `pending-publication` rather than carrying a placeholder
digest. The image builds and was verified locally (canned-prompts
CANP-WP-0006-T06) but has never been pushed, so no registry digest exists. A
placeholder shaped like a real digest would be worse than a sentinel: it could
be mistaken for something deployable.

manifests/ follows the rapp-sbom-nexus shape: namespace labelled for the
postgres client, external secrets from OpenBao, a migration Job, and the
runtime Deployment with a ClusterIP-only Service, dedicated ServiceAccount and
default-deny plus runtime NetworkPolicies.

Three choices worth stating. Credentials arrive as mounted files, never env
vars — an env var holding a password is visible in kubectl describe, in crash
dumps, and to anything that can read /proc. Migrations run as a Job rather than
at start-up, so a schema rollback stays separate from a code rollback and
replicas do not race. Liveness points at /healthz, which checks only that the
process is up: pointing it at a database-dependent path would restart every
replica during a database blip.

Egress is PostgreSQL and DNS only. A package arrives by publish; the registry
never reaches out, so it is given no path to.

tools/smoke.sh checks what only the cluster can answer and calls
canned-prompts' service/tools/smoke.py for health and migration head, rather
than holding a second opinion about whether the service is healthy.

RCP-WP-0002 carries the two operator actions that block a first rollout —
publishing the image and provisioning database roles — and records
per-publisher identity as a decision belonging upstream, which this repo must
not paper over with cluster configuration implying finer control than exists.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 21:44:17 +02:00