Commit graph

4 commits

Author SHA1 Message Date
90ebdb7e7d Fix the watcher verdict: grep -c exit status poisoned a clean result
Version 2 measured everything correctly — not_ready=0, query_failed=0, zero
503s — and still printed FAILED. grep -c exits 1 when it counts zero matches,
so the '|| echo "?"' guard appended a marker on top of the legitimate 0 and
the verdict test could never match.

Third defect in the same instrument. The verdict logic is now tested rather
than assumed: the counter was fed a matching and a non-matching line and
confirmed to return 1 and 0.

Worth noting which direction each failure pointed. This one erred toward alarm,
which is survivable. live-image-digest-match erred the other way and reported
success it had not established — that is the one that would have shipped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-08 10:09:35 +02:00
caf3912d1b Fix the lease watcher, and record that three checks were themselves wrong
The first lease watch returned FAILED on one sample of 21. That sample was the
instrument: it treated an empty kubectl result as not-ready, so a transient API
hiccup was recorded as an outage. The service returned zero 503s across the
window, with a readiness probe every 5s and 46 minutes of uptime past a
30-minute lease.

Not overriding the verdict by argument — a verdict that can be talked around is
worth nothing. The watcher now distinguishes a failed query from a failed
service and corroborates against the kubelet's probe history, which samples far
more often than once a minute. Promoted from scratch into tools/ so it is
reviewable and re-runnable.

Three checks in this rollout were defective in the same way: live-image-digest-match
degrading to 'not pinned' while a digest was pinned, check_readiness reporting
'database unreachable' for a reachable but unmigrated database, and this
watcher. A verification step that cannot fail correctly is worse than none,
because it is trusted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-08 09:48:06 +02:00
4d5af7698d First deployment: verified on railiance01
RCP-WP-0002 T02-T04 done, readiness_state verified with evidence attached
rather than ahead of it.

rapp-postgres provisioned canned_prompts on platform-pg-2 and sent a
database-owner receipt with 12 checks proven live. creds/canned-prompts-publish
was deliberately not issued, so the service runs read-only and POST /packages
returns 503 explaining why — the intended posture, not a gap.

Four defects surfaced that only a real rollout could expose, two of them
silent:

- SET ROLE opened an implicit transaction that Alembic nested inside rather
  than owning, so every revision logged as applied and was rolled back.
  Alembic reported success against an empty database.
- The egress NetworkPolicy selected app.kubernetes.io/name, which the
  migration Job does not carry. The Job matched only the default-deny and
  succeeded exactly once, because it ran before the policies existed; the next
  migration would have failed on DNS. Now selects part-of, with ingress split
  into its own policy so the Job is never reachable.
- env.py read database_url rather than resolved_database_url, so the migration
  could never run where the credential is a mounted file.
- live-image-digest-match extracted the pin with a line-offset grep, which
  returned empty once comments were added above `version:`. The check degraded
  to reporting "not pinned yet" while a digest was pinned — it could not have
  passed for any pin. Now parsed as YAML. A verification step that cannot fail
  is worth less than none, because it is trusted.

Evidence: docs/evidence/RCP-WP-0002-T04-first-deployment-2026-09-08.md

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-08 08:56:02 +02:00
5a0f4cb5c8 Package canned-prompts for Railiance
Registers the repo with State Hub (agents / practice, prefix RCP-WP) and fills
in the rapp shape.

declarations/rapp.yaml declares a manifest-managed platform service owned by
canned-prompts, bound to rail-kubernetes and reef-railiance, with rollout,
smoke and rollback contracts.

The image pin says `pending-publication` rather than carrying a placeholder
digest. The image builds and was verified locally (canned-prompts
CANP-WP-0006-T06) but has never been pushed, so no registry digest exists. A
placeholder shaped like a real digest would be worse than a sentinel: it could
be mistaken for something deployable.

manifests/ follows the rapp-sbom-nexus shape: namespace labelled for the
postgres client, external secrets from OpenBao, a migration Job, and the
runtime Deployment with a ClusterIP-only Service, dedicated ServiceAccount and
default-deny plus runtime NetworkPolicies.

Three choices worth stating. Credentials arrive as mounted files, never env
vars — an env var holding a password is visible in kubectl describe, in crash
dumps, and to anything that can read /proc. Migrations run as a Job rather than
at start-up, so a schema rollback stays separate from a code rollback and
replicas do not race. Liveness points at /healthz, which checks only that the
process is up: pointing it at a database-dependent path would restart every
replica during a database blip.

Egress is PostgreSQL and DNS only. A package arrives by publish; the registry
never reaches out, so it is given no path to.

tools/smoke.sh checks what only the cluster can answer and calls
canned-prompts' service/tools/smoke.py for health and migration head, rather
than holding a second opinion about whether the service is healthy.

RCP-WP-0002 carries the two operator actions that block a first rollout —
publishing the image and provisioning database roles — and records
per-publisher identity as a decision belonging upstream, which this repo must
not paper over with cluster configuration implying finer control than exists.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 21:44:17 +02:00