Version 3 reported 'survived past the lease TTL' after a clean 12-minute window
on a 13-minute-old pod. Every measurement in it was accurate; the conclusion
was not, because nothing checked that the observation window had actually
exceeded the 30-minute lease it claimed to have outlasted.
Fifth defect in this instrument, and the first to err toward reassurance.
Versions 1 and 2 cried wolf, which provokes investigation. This one would have
been believed, and readiness_state: verified recorded on it — the same way
live-image-digest-match would have been believed. A check reporting success it
has not established is indistinguishable from one that works, until it matters.
The verdict now requires uptime > lease TTL and reports INCONCLUSIVE when a
window is clean but too short. 'Clean' and 'proven' are different claims and
only one of them was being measured.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM
Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
Version 2 measured everything correctly — not_ready=0, query_failed=0, zero
503s — and still printed FAILED. grep -c exits 1 when it counts zero matches,
so the '|| echo "?"' guard appended a marker on top of the legitimate 0 and
the verdict test could never match.
Third defect in the same instrument. The verdict logic is now tested rather
than assumed: the counter was fed a matching and a non-matching line and
confirmed to return 1 and 0.
Worth noting which direction each failure pointed. This one erred toward alarm,
which is survivable. live-image-digest-match erred the other way and reported
success it had not established — that is the one that would have shipped.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM
Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
The first lease watch returned FAILED on one sample of 21. That sample was the
instrument: it treated an empty kubectl result as not-ready, so a transient API
hiccup was recorded as an outage. The service returned zero 503s across the
window, with a readiness probe every 5s and 46 minutes of uptime past a
30-minute lease.
Not overriding the verdict by argument — a verdict that can be talked around is
worth nothing. The watcher now distinguishes a failed query from a failed
service and corroborates against the kubelet's probe history, which samples far
more often than once a minute. Promoted from scratch into tools/ so it is
reviewable and re-runnable.
Three checks in this rollout were defective in the same way: live-image-digest-match
degrading to 'not pinned' while a digest was pinned, check_readiness reporting
'database unreachable' for a reachable but unmigrated database, and this
watcher. A verification step that cannot fail correctly is worse than none,
because it is trusted.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM
Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502