The first lease watch returned FAILED on one sample of 21. That sample was the instrument: it treated an empty kubectl result as not-ready, so a transient API hiccup was recorded as an outage. The service returned zero 503s across the window, with a readiness probe every 5s and 46 minutes of uptime past a 30-minute lease. Not overriding the verdict by argument — a verdict that can be talked around is worth nothing. The watcher now distinguishes a failed query from a failed service and corroborates against the kubelet's probe history, which samples far more often than once a minute. Promoted from scratch into tools/ so it is reviewable and re-runnable. Three checks in this rollout were defective in the same way: live-image-digest-match degrading to 'not pinned' while a digest was pinned, check_readiness reporting 'database unreachable' for a reachable but unmigrated database, and this watcher. A verification step that cannot fail correctly is worse than none, because it is trusted. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502 |
||
|---|---|---|
| .. | ||
| lease-watch.sh | ||
| smoke.sh | ||