Fix the lease watcher, and record that three checks were themselves wrong

The first lease watch returned FAILED on one sample of 21. That sample was the
instrument: it treated an empty kubectl result as not-ready, so a transient API
hiccup was recorded as an outage. The service returned zero 503s across the
window, with a readiness probe every 5s and 46 minutes of uptime past a
30-minute lease.

Not overriding the verdict by argument — a verdict that can be talked around is
worth nothing. The watcher now distinguishes a failed query from a failed
service and corroborates against the kubelet's probe history, which samples far
more often than once a minute. Promoted from scratch into tools/ so it is
reviewable and re-runnable.

Three checks in this rollout were defective in the same way: live-image-digest-match
degrading to 'not pinned' while a digest was pinned, check_readiness reporting
'database unreachable' for a reachable but unmigrated database, and this
watcher. A verification step that cannot fail correctly is worse than none,
because it is trusted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
This commit is contained in:
tegwick 2026-09-08 09:48:06 +02:00
parent 371662aa8a
commit caf3912d1b
3 changed files with 79 additions and 2 deletions

View file

@ -74,3 +74,34 @@ lease window cannot distinguish a service that works from one that works *once*.
The criterion for `verified` is therefore observation across a full rotation,
not a passing check at a single instant — an availability property is not
provable by one sample.
## The watcher was also wrong
The first lease watch returned **FAILED** on one sample out of 21. That sample
was the instrument, not the service: it treated an empty `kubectl` result as
"not ready", so a transient API hiccup was recorded as an outage.
The service had returned **zero** 503s across the whole window. The readiness
probe fires every 5s, so a genuine two-minute outage would have left roughly
twenty-four of them, and the pod ran 46 minutes past a 30-minute lease with no
restarts.
I did not override the verdict by argument — a verdict that can be talked around
is worth nothing. The watcher is fixed (`tools/lease-watch.sh`) to distinguish
a failed *query* from a failed *service*, and to corroborate its own sampling
against the kubelet's probe history, which is a far better instrument than one
sample a minute. Then re-measured.
**Three checks in this rollout were themselves defective**, all the same shape —
unable to distinguish their own failure from the failure they were watching for:
| Check | How it lied |
|---|---|
| `live-image-digest-match` | line-offset `grep` returned empty; reported "not pinned yet" while a digest was pinned |
| `check_readiness` (earlier) | one `except` reported "database unreachable" for a reachable but unmigrated database |
| lease watcher v1 | empty `kubectl` output counted as an outage |
A verification step that cannot fail correctly is worse than none, because it is
trusted. That is the durable lesson from this deployment, more than any single
defect it found.