uptime 36m > lease TTL 30m, 21 samples, zero not-ready, zero query failures,
zero kubelet readiness 503s, no restarts. The verdict now names the uptime it
checked, so the claim can be audited rather than taken.
T05 done upstream in canned-prompts 0.2.0 / migration 0003. DR-3 had already
resolved the identity question as app-local accounts with OIDC demand-gated;
research found that rather than my judgement supplying it.
Left open deliberately: the NetworkPolicy ingress rule still admits any
namespace. Publisher identity now gates writes so it is no longer the only
control, and it should narrow once the legitimate callers are known — recorded
rather than tightened on a guess.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM
Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
Version 3 reported 'survived past the lease TTL' after a clean 12-minute window
on a 13-minute-old pod. Every measurement in it was accurate; the conclusion
was not, because nothing checked that the observation window had actually
exceeded the 30-minute lease it claimed to have outlasted.
Fifth defect in this instrument, and the first to err toward reassurance.
Versions 1 and 2 cried wolf, which provokes investigation. This one would have
been believed, and readiness_state: verified recorded on it — the same way
live-image-digest-match would have been believed. A check reporting success it
has not established is indistinguishable from one that works, until it matters.
The verdict now requires uptime > lease TTL and reports INCONCLUSIVE when a
window is clean but too short. 'Clean' and 'proven' are different claims and
only one of them was being measured.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM
Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
Continuously ready across a full credential lease rotation: 11 samples, zero
not-ready, zero query failures, zero kubelet readiness 503s with a probe every
5 seconds, no restarts, pod uptime well past the 30-minute lease TTL.
Recorded from a working instrument on the third attempt. The first two verdicts
were FAILED and both were the watcher's own defects — an empty kubectl result
counted as an outage, then grep -c's exit status poisoning a clean count. The
verdict was not talked around; the instrument was fixed and the measurement
repeated.
Evidence: docs/evidence/RCP-WP-0002-T04-lease-rotation-2026-09-08.log
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM
Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
Version 2 measured everything correctly — not_ready=0, query_failed=0, zero
503s — and still printed FAILED. grep -c exits 1 when it counts zero matches,
so the '|| echo "?"' guard appended a marker on top of the legitimate 0 and
the verdict test could never match.
Third defect in the same instrument. The verdict logic is now tested rather
than assumed: the counter was fed a matching and a non-matching line and
confirmed to return 1 and 0.
Worth noting which direction each failure pointed. This one erred toward alarm,
which is survivable. live-image-digest-match erred the other way and reported
success it had not established — that is the one that would have shipped.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM
Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
The first lease watch returned FAILED on one sample of 21. That sample was the
instrument: it treated an empty kubectl result as not-ready, so a transient API
hiccup was recorded as an outage. The service returned zero 503s across the
window, with a readiness probe every 5s and 46 minutes of uptime past a
30-minute lease.
Not overriding the verdict by argument — a verdict that can be talked around is
worth nothing. The watcher now distinguishes a failed query from a failed
service and corroborates against the kubelet's probe history, which samples far
more often than once a minute. Promoted from scratch into tools/ so it is
reviewable and re-runnable.
Three checks in this rollout were defective in the same way: live-image-digest-match
degrading to 'not pinned' while a digest was pinned, check_readiness reporting
'database unreachable' for a reachable but unmigrated database, and this
watcher. A verification step that cannot fail correctly is worse than none,
because it is trusted.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM
Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
I recorded readiness_state: verified after a passing smoke run and found the
deployment 0/1 eight hours later.
Platform credentials are 30-minute leases, not passwords. The service read the
mounted URL once at start-up and never again, so External Secrets kept the file
current while the engine held the URL it booted with, and every reconnection
after the first lease expiry used a credential the database had already
revoked.
/readyz reported it accurately — "database unreachable: OperationalError" — and
the pod sat unready for eight hours without being restarted, because liveness
is deliberately independent of the database. That separation behaved exactly as
designed: the process was alive and could not serve, and it said so.
Fixed in 0.1.5: the engine re-reads the credential for every new connection and
pool_recycle is 900s, well inside the lease. Only username and password come
from the refreshed URL; host, port and database come from the engine, so a
malformed refresh cannot silently redirect the service.
readiness_state back to deployed. The bar for verified is now observation
across a full lease rotation, because a smoke run inside the first window
cannot distinguish a service that works from one that works once — an
availability property is not provable by a single sample.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM
Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
RCP-WP-0002 T02-T04 done, readiness_state verified with evidence attached
rather than ahead of it.
rapp-postgres provisioned canned_prompts on platform-pg-2 and sent a
database-owner receipt with 12 checks proven live. creds/canned-prompts-publish
was deliberately not issued, so the service runs read-only and POST /packages
returns 503 explaining why — the intended posture, not a gap.
Four defects surfaced that only a real rollout could expose, two of them
silent:
- SET ROLE opened an implicit transaction that Alembic nested inside rather
than owning, so every revision logged as applied and was rolled back.
Alembic reported success against an empty database.
- The egress NetworkPolicy selected app.kubernetes.io/name, which the
migration Job does not carry. The Job matched only the default-deny and
succeeded exactly once, because it ran before the policies existed; the next
migration would have failed on DNS. Now selects part-of, with ingress split
into its own policy so the Job is never reachable.
- env.py read database_url rather than resolved_database_url, so the migration
could never run where the credential is a mounted file.
- live-image-digest-match extracted the pin with a line-offset grep, which
returned empty once comments were added above `version:`. The check degraded
to reporting "not pinned yet" while a digest was pinned — it could not have
passed for any pin. Now parsed as YAML. A verification step that cannot fail
is worth less than none, because it is trusted.
Evidence: docs/evidence/RCP-WP-0002-T04-first-deployment-2026-09-08.md
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM
Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502