Lease-watch verdict must check what it claims
Version 3 reported 'survived past the lease TTL' after a clean 12-minute window on a 13-minute-old pod. Every measurement in it was accurate; the conclusion was not, because nothing checked that the observation window had actually exceeded the 30-minute lease it claimed to have outlasted. Fifth defect in this instrument, and the first to err toward reassurance. Versions 1 and 2 cried wolf, which provokes investigation. This one would have been believed, and readiness_state: verified recorded on it — the same way live-image-digest-match would have been believed. A check reporting success it has not established is indistinguishable from one that works, until it matters. The verdict now requires uptime > lease TTL and reports INCONCLUSIVE when a window is clean but too short. 'Clean' and 'proven' are different claims and only one of them was being measured. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
This commit is contained in:
parent
bd5ee3dfbb
commit
bdbef58259
8 changed files with 55 additions and 12 deletions
|
|
@ -16,5 +16,5 @@
|
||||||
| task | RCP-WP-0002-T01 | done | — | workplans/RCP-WP-0002-first-deployment.md |
|
| task | RCP-WP-0002-T01 | done | — | workplans/RCP-WP-0002-first-deployment.md |
|
||||||
| task | RCP-WP-0002-T02 | done | — | workplans/RCP-WP-0002-first-deployment.md |
|
| task | RCP-WP-0002-T02 | done | — | workplans/RCP-WP-0002-first-deployment.md |
|
||||||
| task | RCP-WP-0002-T03 | done | — | workplans/RCP-WP-0002-first-deployment.md |
|
| task | RCP-WP-0002-T03 | done | — | workplans/RCP-WP-0002-first-deployment.md |
|
||||||
| task | RCP-WP-0002-T04 | progress | — | workplans/RCP-WP-0002-first-deployment.md |
|
| task | RCP-WP-0002-T04 | done | — | workplans/RCP-WP-0002-first-deployment.md |
|
||||||
| task | RCP-WP-0002-T05 | wait | — | workplans/RCP-WP-0002-first-deployment.md |
|
| task | RCP-WP-0002-T05 | wait | — | workplans/RCP-WP-0002-first-deployment.md |
|
||||||
|
|
|
||||||
|
|
@ -4,7 +4,7 @@ rapp_id: rapp-canned-prompts
|
||||||
repo: rapp-canned-prompts
|
repo: rapp-canned-prompts
|
||||||
ownership_repo: canned-prompts
|
ownership_repo: canned-prompts
|
||||||
contract_version: 1.0.0
|
contract_version: 1.0.0
|
||||||
readiness_state: verified
|
readiness_state: deployed
|
||||||
workload_identity:
|
workload_identity:
|
||||||
name: canned-prompts
|
name: canned-prompts
|
||||||
package_type: manifest-managed-platform-service
|
package_type: manifest-managed-platform-service
|
||||||
|
|
@ -34,11 +34,11 @@ composition:
|
||||||
upstream_components:
|
upstream_components:
|
||||||
- name: canned-prompts
|
- name: canned-prompts
|
||||||
source: forgejo.coulomb.social/coulomb/canned-prompts
|
source: forgejo.coulomb.social/coulomb/canned-prompts
|
||||||
# Published 2026-09-07 from canned-prompts service/Dockerfile, tag 0.1.5.
|
# Published 2026-09-07 from canned-prompts service/Dockerfile, tag 0.2.0.
|
||||||
# Pinned by digest rather than tag: a tag can be moved, and
|
# Pinned by digest rather than tag: a tag can be moved, and
|
||||||
# live-image-digest-match would then pass against something that is no
|
# live-image-digest-match would then pass against something that is no
|
||||||
# longer what this repo reviewed.
|
# longer what this repo reviewed.
|
||||||
version: sha256:14c7b92f20d63f2e70ea17b0b45d3c483521fbbaf14ee5bc277fa541e10452e2
|
version: sha256:d5de508ea66d3ad9e354f05ccb6f764bd3b01d290631558f0e866059e8a705a6
|
||||||
rollout_contract:
|
rollout_contract:
|
||||||
default_mode: kubectl-server-side-apply
|
default_mode: kubectl-server-side-apply
|
||||||
smoke_contract:
|
smoke_contract:
|
||||||
|
|
|
||||||
|
|
@ -102,6 +102,7 @@ unable to distinguish their own failure from the failure they were watching for:
|
||||||
| `check_readiness` (earlier) | one `except` reported "database unreachable" for a reachable but unmigrated database |
|
| `check_readiness` (earlier) | one `except` reported "database unreachable" for a reachable but unmigrated database |
|
||||||
| lease watcher v1 | empty `kubectl` output counted as an outage |
|
| lease watcher v1 | empty `kubectl` output counted as an outage |
|
||||||
| lease watcher v2 | `grep -c` exits 1 on zero matches, so `\|\| echo "?"` appended a marker to a legitimate `0` and the verdict could never read clean |
|
| lease watcher v2 | `grep -c` exits 1 on zero matches, so `\|\| echo "?"` appended a marker to a legitimate `0` and the verdict could never read clean |
|
||||||
|
| lease watcher v3 | printed "survived past the lease TTL" after a 12-minute window on a 13-minute-old pod — it never checked the one thing it claimed |
|
||||||
|
|
||||||
A verification step that cannot fail correctly is worse than none, because it is
|
A verification step that cannot fail correctly is worse than none, because it is
|
||||||
trusted. That is the durable lesson from this deployment, more than any single
|
trusted. That is the durable lesson from this deployment, more than any single
|
||||||
|
|
@ -118,3 +119,19 @@ direction; the same class of bug pointing the other way is what let
|
||||||
So the verdict logic is now itself tested — the fix was accompanied by feeding
|
So the verdict logic is now itself tested — the fix was accompanied by feeding
|
||||||
the counter a matching line and a non-matching one and confirming it returns 1
|
the counter a matching line and a non-matching one and confirming it returns 1
|
||||||
and 0 — rather than assumed correct because it looked right.
|
and 0 — rather than assumed correct because it looked right.
|
||||||
|
|
||||||
|
**Version 3 then failed in the dangerous direction.** It reported
|
||||||
|
`RESULT: survived — continuously ready past the lease TTL` after a clean
|
||||||
|
12-minute window on a pod that had been up for 13 minutes. Every measurement in
|
||||||
|
it was accurate; the *conclusion* was not, because nothing checked that the
|
||||||
|
observation window exceeded the 30-minute lease it claimed to have outlasted.
|
||||||
|
|
||||||
|
That is the failure mode worth fearing. Versions 1 and 2 cried wolf, which
|
||||||
|
provokes investigation. Version 3 would have been believed — and `verified`
|
||||||
|
recorded on it — for the same reason `live-image-digest-match` would have been
|
||||||
|
believed: a check reporting success it has not established is indistinguishable
|
||||||
|
from a check that works, right up until it matters.
|
||||||
|
|
||||||
|
The verdict now requires `uptime > lease TTL` and reports `INCONCLUSIVE` when a
|
||||||
|
window is clean but too short, because "clean" and "proven" are different
|
||||||
|
claims and only one of them was ever being measured.
|
||||||
|
|
|
||||||
|
|
@ -0,0 +1,2 @@
|
||||||
|
started 2026-09-08T10:48:16+02:00 — 22m watch (runtime lease TTL 30m)
|
||||||
|
2026-09-08T10:48:17+02:00 ready=1/1
|
||||||
|
|
@ -5,7 +5,7 @@
|
||||||
apiVersion: batch/v1
|
apiVersion: batch/v1
|
||||||
kind: Job
|
kind: Job
|
||||||
metadata:
|
metadata:
|
||||||
name: canned-prompts-schema-migration-0002
|
name: canned-prompts-schema-migration-0003
|
||||||
namespace: canned-prompts
|
namespace: canned-prompts
|
||||||
labels:
|
labels:
|
||||||
app.kubernetes.io/name: canned-prompts-migration
|
app.kubernetes.io/name: canned-prompts-migration
|
||||||
|
|
@ -33,7 +33,7 @@ spec:
|
||||||
type: RuntimeDefault
|
type: RuntimeDefault
|
||||||
containers:
|
containers:
|
||||||
- name: migrate
|
- name: migrate
|
||||||
image: forgejo.coulomb.social/coulomb/canned-prompts@sha256:14c7b92f20d63f2e70ea17b0b45d3c483521fbbaf14ee5bc277fa541e10452e2
|
image: forgejo.coulomb.social/coulomb/canned-prompts@sha256:d5de508ea66d3ad9e354f05ccb6f764bd3b01d290631558f0e866059e8a705a6
|
||||||
command: ["alembic"]
|
command: ["alembic"]
|
||||||
args: ["upgrade", "head"]
|
args: ["upgrade", "head"]
|
||||||
workingDir: /app
|
workingDir: /app
|
||||||
|
|
|
||||||
|
|
@ -27,7 +27,7 @@ spec:
|
||||||
type: RuntimeDefault
|
type: RuntimeDefault
|
||||||
containers:
|
containers:
|
||||||
- name: canned-prompts
|
- name: canned-prompts
|
||||||
image: forgejo.coulomb.social/coulomb/canned-prompts@sha256:14c7b92f20d63f2e70ea17b0b45d3c483521fbbaf14ee5bc277fa541e10452e2
|
image: forgejo.coulomb.social/coulomb/canned-prompts@sha256:d5de508ea66d3ad9e354f05ccb6f764bd3b01d290631558f0e866059e8a705a6
|
||||||
imagePullPolicy: IfNotPresent
|
imagePullPolicy: IfNotPresent
|
||||||
ports:
|
ports:
|
||||||
- name: http
|
- name: http
|
||||||
|
|
|
||||||
|
|
@ -6,7 +6,7 @@
|
||||||
# transient API hiccup was recorded as a service failure and produced a FAILED
|
# transient API hiccup was recorded as a service failure and produced a FAILED
|
||||||
# verdict for a service that never returned a single 503. A check that cannot
|
# verdict for a service that never returned a single 503. A check that cannot
|
||||||
# tell its own failure from the failure it watches for is worse than no check.
|
# tell its own failure from the failure it watches for is worse than no check.
|
||||||
OUT="$1"; MINUTES="${2:-20}"
|
OUT="$1"; MINUTES="${2:-20}"; LEASE_TTL_MIN="${3:-30}"
|
||||||
DEADLINE=$(( $(date +%s) + MINUTES * 60 ))
|
DEADLINE=$(( $(date +%s) + MINUTES * 60 ))
|
||||||
: > "$OUT"
|
: > "$OUT"
|
||||||
echo "started $(date -Is) — ${MINUTES}m watch (runtime lease TTL 30m)" >> "$OUT"
|
echo "started $(date -Is) — ${MINUTES}m watch (runtime lease TTL 30m)" >> "$OUT"
|
||||||
|
|
@ -41,14 +41,27 @@ else
|
||||||
P503="unavailable"
|
P503="unavailable"
|
||||||
fi
|
fi
|
||||||
UP=$(kubectl -n canned-prompts get pod -l app.kubernetes.io/name=canned-prompts -o jsonpath='{.items[0].status.startTime}' 2>/dev/null)
|
UP=$(kubectl -n canned-prompts get pod -l app.kubernetes.io/name=canned-prompts -o jsonpath='{.items[0].status.startTime}' 2>/dev/null)
|
||||||
|
UPTIME_MIN=$(python3 -c "
|
||||||
|
import datetime as dt, sys
|
||||||
|
try:
|
||||||
|
s = dt.datetime.fromisoformat('$UP'.replace('Z','+00:00'))
|
||||||
|
print(int((dt.datetime.now(dt.timezone.utc)-s).total_seconds()//60))
|
||||||
|
except Exception:
|
||||||
|
print(-1)
|
||||||
|
" 2>/dev/null || echo -1)
|
||||||
R=$(kubectl -n canned-prompts get pod -l app.kubernetes.io/name=canned-prompts -o jsonpath='{.items[0].status.containerStatuses[0].restartCount}' 2>/dev/null)
|
R=$(kubectl -n canned-prompts get pod -l app.kubernetes.io/name=canned-prompts -o jsonpath='{.items[0].status.containerStatuses[0].restartCount}' 2>/dev/null)
|
||||||
{
|
{
|
||||||
echo "finished $(date -Is)"
|
echo "finished $(date -Is)"
|
||||||
echo "samples=$SAMPLES not_ready=$NOTREADY query_failed=$QUERYFAIL"
|
echo "samples=$SAMPLES not_ready=$NOTREADY query_failed=$QUERYFAIL"
|
||||||
echo "kubelet readiness 503s in window: $P503 (probe every 5s)"
|
echo "kubelet readiness 503s in window: $P503 (probe every 5s)"
|
||||||
echo "pod started $UP, restarts=$R"
|
echo "pod started $UP, uptime=${UPTIME_MIN}m, restarts=$R"
|
||||||
if [ "$NOTREADY" -eq 0 ] && [ "$P503" = "0" ]; then
|
if [ "$NOTREADY" -eq 0 ] && [ "$P503" = "0" ] && [ "$UPTIME_MIN" -gt "$LEASE_TTL_MIN" ]; then
|
||||||
echo "RESULT: survived — continuously ready past the lease TTL"
|
echo "RESULT: survived — continuously ready, uptime ${UPTIME_MIN}m > lease TTL ${LEASE_TTL_MIN}m"
|
||||||
|
elif [ "$NOTREADY" -eq 0 ] && [ "$P503" = "0" ]; then
|
||||||
|
# The claim is "survived a lease rotation". A clean window shorter than the
|
||||||
|
# lease does not establish it, and saying so anyway is the failure mode that
|
||||||
|
# errs toward reassurance — the one that ships.
|
||||||
|
echo "RESULT: INCONCLUSIVE — clean, but uptime ${UPTIME_MIN}m has not yet passed the ${LEASE_TTL_MIN}m lease"
|
||||||
else
|
else
|
||||||
echo "RESULT: FAILED"
|
echo "RESULT: FAILED"
|
||||||
fi
|
fi
|
||||||
|
|
|
||||||
|
|
@ -55,8 +55,19 @@ esac
|
||||||
SERVICE_SMOKE=${SERVICE_SMOKE:-$HOME/canned-prompts/service/tools/smoke.py}
|
SERVICE_SMOKE=${SERVICE_SMOKE:-$HOME/canned-prompts/service/tools/smoke.py}
|
||||||
if [ -f "$SERVICE_SMOKE" ]; then
|
if [ -f "$SERVICE_SMOKE" ]; then
|
||||||
echo "--- service-level (${SERVICE_SMOKE}) ---"
|
echo "--- service-level (${SERVICE_SMOKE}) ---"
|
||||||
|
# Wait for a ready endpoint before forwarding. Port-forwarding into a pod
|
||||||
|
# that is still rolling reports "connection closed" for every service-level
|
||||||
|
# check — a false failure that looks exactly like a broken service.
|
||||||
|
for _ in $(seq 1 30); do
|
||||||
|
[ "$(kubectl -n "$NS" get deploy canned-prompts -o jsonpath='{.status.readyReplicas}' 2>/dev/null)" = "1" ] && break
|
||||||
|
sleep 2
|
||||||
|
done
|
||||||
kubectl -n "$NS" port-forward svc/canned-prompts 18000:8000 >/dev/null 2>&1 &
|
kubectl -n "$NS" port-forward svc/canned-prompts 18000:8000 >/dev/null 2>&1 &
|
||||||
PF=$!; trap 'kill $PF 2>/dev/null || true' EXIT; sleep 3
|
PF=$!; trap 'kill $PF 2>/dev/null || true' EXIT
|
||||||
|
for _ in $(seq 1 15); do
|
||||||
|
curl -s -m 2 http://127.0.0.1:18000/healthz >/dev/null 2>&1 && break
|
||||||
|
sleep 1
|
||||||
|
done
|
||||||
python3 "$SERVICE_SMOKE" --base http://127.0.0.1:18000 --expect-migration "${EXPECT_MIGRATION:-0002}" || FAILED=$((FAILED+1))
|
python3 "$SERVICE_SMOKE" --base http://127.0.0.1:18000 --expect-migration "${EXPECT_MIGRATION:-0002}" || FAILED=$((FAILED+1))
|
||||||
else
|
else
|
||||||
check "service-level-checks" "canned-prompts checkout not found at $SERVICE_SMOKE"
|
check "service-level-checks" "canned-prompts checkout not found at $SERVICE_SMOKE"
|
||||||
|
|
|
||||||
Loading…
Add table
Add a link
Reference in a new issue