Fix the watcher verdict: grep -c exit status poisoned a clean result

Version 2 measured everything correctly — not_ready=0, query_failed=0, zero
503s — and still printed FAILED. grep -c exits 1 when it counts zero matches,
so the '|| echo "?"' guard appended a marker on top of the legitimate 0 and
the verdict test could never match.

Third defect in the same instrument. The verdict logic is now tested rather
than assumed: the counter was fed a matching and a non-matching line and
confirmed to return 1 and 0.

Worth noting which direction each failure pointed. This one erred toward alarm,
which is survivable. live-image-digest-match erred the other way and reported
success it had not established — that is the one that would have shipped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
This commit is contained in:
tegwick 2026-09-08 10:09:35 +02:00
parent caf3912d1b
commit 90ebdb7e7d
2 changed files with 23 additions and 1 deletions

View file

@ -101,7 +101,20 @@ unable to distinguish their own failure from the failure they were watching for:
| `live-image-digest-match` | line-offset `grep` returned empty; reported "not pinned yet" while a digest was pinned |
| `check_readiness` (earlier) | one `except` reported "database unreachable" for a reachable but unmigrated database |
| lease watcher v1 | empty `kubectl` output counted as an outage |
| lease watcher v2 | `grep -c` exits 1 on zero matches, so `\|\| echo "?"` appended a marker to a legitimate `0` and the verdict could never read clean |
A verification step that cannot fail correctly is worse than none, because it is
trusted. That is the durable lesson from this deployment, more than any single
defect it found.
The watcher needed three attempts, and the third failure is the most instructive
of the set. Version 2 reported `not_ready=0 query_failed=0` and `503s: 0` — every
measurement clean — and still printed `FAILED`, because `grep -c` exits non-zero
when it counts zero and the guard appended an error marker on top of the correct
answer. The check was wrong *in the direction of alarm*, which is the survivable
direction; the same class of bug pointing the other way is what let
`live-image-digest-match` report success it had not established.
So the verdict logic is now itself tested — the fix was accompanied by feeding
the counter a matching line and a non-matching one and confirming it returns 1
and 0 — rather than assumed correct because it looked right.

View file

@ -30,7 +30,16 @@ while [ "$(date +%s)" -lt "$DEADLINE" ]; do
done
# Corroborate with the kubelet, which probes every 5s and is a far better
# instrument than this loop's 1/min sampling.
P503=$(kubectl -n canned-prompts logs deploy/canned-prompts --since="${MINUTES}m" 2>/dev/null | grep -c '" 503' || echo "?")
# grep -c exits 1 when the count is zero, so `|| echo "?"` appended a marker on
# top of a legitimate "0" and the verdict test never matched. Capture the logs
# first: a kubectl failure is then distinguishable from an absence of 503s,
# which is the whole point.
if LOGS=$(kubectl -n canned-prompts logs deploy/canned-prompts --since="${MINUTES}m" 2>/dev/null); then
P503=$(printf '%s' "$LOGS" | grep -c '" 503' || true)
P503=${P503:-0}
else
P503="unavailable"
fi
UP=$(kubectl -n canned-prompts get pod -l app.kubernetes.io/name=canned-prompts -o jsonpath='{.items[0].status.startTime}' 2>/dev/null)
R=$(kubectl -n canned-prompts get pod -l app.kubernetes.io/name=canned-prompts -o jsonpath='{.items[0].status.containerStatuses[0].restartCount}' 2>/dev/null)
{