# Availability and recovery **Status:** V1 not yet evidenced. Dependency enumeration complete; exercise pending a live window on railiance01. **Framework:** NetKingdom Tenancy Posture v0.1 (draft-8) §4.6, §13 V-row. **Declared:** `tenancy.yaml` — `current.V: 0`, `target.V: 1`. ## Why V0 today `V0` is "no availability or recovery position; recovery is untested or depends on improvisation". That is accurate. Nothing in this repo exercises recovery of the complete audit path and records a measured recovery time. What exists is an *observation*, not an exercise: after the railiance01 node reboot on 2026-08-16, `/readyz` failed for roughly 40 seconds before the pod went Ready, recorded in `docs/operator-runbook.md`. That is useful and it is not V1 evidence — it enumerated nothing, measured nothing deliberately, and happened to us rather than being performed. Decision 4.6.1 is explicit that a replica count or a status page is not evidence of a level. audit-core runs one replica; that fact argues for neither V0 nor V1. Only the exercise settles it. ## Critical dependency enumeration (Decision 4.6.1) V is **end-to-end**: the minimum across the components and synchronous providers required to serve the operation. An application that restarts in five seconds over a database that takes ninety is not V1 at five seconds. The operation being declared is **accept an audit event** (`POST /v1/events`). Read surfaces are operator-facing and are not the availability claim. | # | Dependency | Role in the accept path | Failure shows as | |---|---|---|---| | 1 | `audit-core` Deployment (1 replica, namespace `audit-core`) | Serves the receiver | Connection refused; Service has no endpoint | | 2 | `platform-pg` (CNPG, `instances: 1`) | Custody. Synchronous — an accept is not accepted until committed | `/readyz` 503, request 503 `unavailable` | | 3 | Secret `audit-core-database` at `/etc/audit-core/db` | Runtime lease, re-read per connection | 503 `unavailable` on lease expiry or revocation | | 4 | ESO + OpenBao (`database/creds/audit-core-runtime`) | Refreshes 3 every 15 min | Delayed: works until the current lease expires | | 5 | Secret `audit-core-senders` + scope ConfigMap | Sender authentication; read at process start | 401 for every sender after a restart with a bad Secret | | 6 | CoreDNS | Resolves `platform-pg` | 503 `unavailable`, indistinguishable from 2 without logs | | 7 | NetworkPolicy (default-deny + sender ingress) | Admits `user-engine` | Sender timeouts; receiver sees nothing at all | **The binding constraint is 2.** `platform-pg` runs `instances: 1` — no HA, single node. §17 of the framework already records that no P1 tenant has HA. audit-core cannot exceed `platform-pg`'s availability level regardless of what it does to its own Deployment, so **V2 is not reachable from P1 as built** and is not a target. V1 is. Dependency 4 is worth separating: OpenBao being down does not stop accepts. It stops *renewal*, so the failure is deferred to lease expiry. An exercise that takes OpenBao down and declares success after two minutes has proved nothing — this is the invisible-failure shape §12 warns about. ## The V1 exercise §13's V1 row asks for three things, all mechanical: dependencies enumerated (above), restart/recreate recovery exercised, interruption and measured recovery time recorded. **Rules for the run.** Measure from the last successful accept to the next successful accept, not from pod restart to Ready — Ready is our own opinion of ourselves. Drive a real `POST /v1/events` on a fixed interval throughout and count what the sender sees, because the sender's experience is the availability being claimed. Record wall-clock times, not durations someone computed later. | # | Scenario | Action | Record | |---|---|---|---| | E1 | Receiver restart | `kubectl -n audit-core rollout restart deploy/audit-core` | Accepts lost; time to first successful accept | | E2 | Receiver recreate | Delete the pod | As E1, plus whether any accepted event is missing after recovery | | E3 | Custody restart | Restart the `platform-pg` primary | Time to first successful accept; 503 count; chain intact after | | E4 | Lease expiry under load | Revoke the runtime lease, wait for ESO refresh | Time to recovery **without** a pod restart — the property `credentials.py` exists to provide | | E5 | Node reboot | Reboot railiance01 | End-to-end recovery; compare against the 2026-08-16 observation of ~40s | **Integrity is part of the pass condition, not a separate check.** After every scenario, `python -m audit_core verify-chain` must report intact, and every event the driver believes was accepted must be present. A recovery that loses an accepted event, or forks the chain, is a custody defect and fails the exercise regardless of how fast it was. **RPO.** An accept is committed before the sender is told, so the expected RPO for accepted events is zero. The exercise tests that claim rather than assuming it; E2 and E3 are where it would break. ## Results *Not yet run.* Needs a live window on railiance01 and coordination with `user-engine`, since E4 and E5 are visible to the sender. On completion, record the measured recovery time per scenario here with dates, raise `tenancy.yaml` `current.V` to 1, note the exercise as the evidence, and set a review date — a recovery exercise from a year ago describes a system that no longer exists.