93 lines
5.3 KiB
Markdown
93 lines
5.3 KiB
Markdown
|
|
# Availability and recovery
|
||
|
|
|
||
|
|
**Status:** V1 not yet evidenced. Dependency enumeration complete; exercise
|
||
|
|
pending a live window on railiance01.
|
||
|
|
**Framework:** NetKingdom Tenancy Posture v0.1 (draft-8) §4.6, §13 V-row.
|
||
|
|
**Declared:** `tenancy.yaml` — `current.V: 0`, `target.V: 1`.
|
||
|
|
|
||
|
|
## Why V0 today
|
||
|
|
|
||
|
|
`V0` is "no availability or recovery position; recovery is untested or depends
|
||
|
|
on improvisation". That is accurate. Nothing in this repo exercises recovery of
|
||
|
|
the complete audit path and records a measured recovery time.
|
||
|
|
|
||
|
|
What exists is an *observation*, not an exercise: after the railiance01 node
|
||
|
|
reboot on 2026-08-16, `/readyz` failed for roughly 40 seconds before the pod
|
||
|
|
went Ready, recorded in `docs/operator-runbook.md`. That is useful and it is not
|
||
|
|
V1 evidence — it enumerated nothing, measured nothing deliberately, and happened
|
||
|
|
to us rather than being performed.
|
||
|
|
|
||
|
|
Decision 4.6.1 is explicit that a replica count or a status page is not
|
||
|
|
evidence of a level. audit-core runs one replica; that fact argues for neither
|
||
|
|
V0 nor V1. Only the exercise settles it.
|
||
|
|
|
||
|
|
## Critical dependency enumeration (Decision 4.6.1)
|
||
|
|
|
||
|
|
V is **end-to-end**: the minimum across the components and synchronous providers
|
||
|
|
required to serve the operation. An application that restarts in five seconds
|
||
|
|
over a database that takes ninety is not V1 at five seconds.
|
||
|
|
|
||
|
|
The operation being declared is **accept an audit event** (`POST /v1/events`).
|
||
|
|
Read surfaces are operator-facing and are not the availability claim.
|
||
|
|
|
||
|
|
| # | Dependency | Role in the accept path | Failure shows as |
|
||
|
|
|---|---|---|---|
|
||
|
|
| 1 | `audit-core` Deployment (1 replica, namespace `audit-core`) | Serves the receiver | Connection refused; Service has no endpoint |
|
||
|
|
| 2 | `platform-pg` (CNPG, `instances: 1`) | Custody. Synchronous — an accept is not accepted until committed | `/readyz` 503, request 503 `unavailable` |
|
||
|
|
| 3 | Secret `audit-core-database` at `/etc/audit-core/db` | Runtime lease, re-read per connection | 503 `unavailable` on lease expiry or revocation |
|
||
|
|
| 4 | ESO + OpenBao (`database/creds/audit-core-runtime`) | Refreshes 3 every 15 min | Delayed: works until the current lease expires |
|
||
|
|
| 5 | Secret `audit-core-senders` + scope ConfigMap | Sender authentication; read at process start | 401 for every sender after a restart with a bad Secret |
|
||
|
|
| 6 | CoreDNS | Resolves `platform-pg` | 503 `unavailable`, indistinguishable from 2 without logs |
|
||
|
|
| 7 | NetworkPolicy (default-deny + sender ingress) | Admits `user-engine` | Sender timeouts; receiver sees nothing at all |
|
||
|
|
|
||
|
|
**The binding constraint is 2.** `platform-pg` runs `instances: 1` — no HA, single
|
||
|
|
node. §17 of the framework already records that no P1 tenant has HA. audit-core
|
||
|
|
cannot exceed `platform-pg`'s availability level regardless of what it does to
|
||
|
|
its own Deployment, so **V2 is not reachable from P1 as built** and is not a
|
||
|
|
target. V1 is.
|
||
|
|
|
||
|
|
Dependency 4 is worth separating: OpenBao being down does not stop accepts. It
|
||
|
|
stops *renewal*, so the failure is deferred to lease expiry. An exercise that
|
||
|
|
takes OpenBao down and declares success after two minutes has proved nothing —
|
||
|
|
this is the invisible-failure shape §12 warns about.
|
||
|
|
|
||
|
|
## The V1 exercise
|
||
|
|
|
||
|
|
§13's V1 row asks for three things, all mechanical: dependencies enumerated
|
||
|
|
(above), restart/recreate recovery exercised, interruption and measured recovery
|
||
|
|
time recorded.
|
||
|
|
|
||
|
|
**Rules for the run.** Measure from the last successful accept to the next
|
||
|
|
successful accept, not from pod restart to Ready — Ready is our own opinion of
|
||
|
|
ourselves. Drive a real `POST /v1/events` on a fixed interval throughout and
|
||
|
|
count what the sender sees, because the sender's experience is the availability
|
||
|
|
being claimed. Record wall-clock times, not durations someone computed later.
|
||
|
|
|
||
|
|
| # | Scenario | Action | Record |
|
||
|
|
|---|---|---|---|
|
||
|
|
| E1 | Receiver restart | `kubectl -n audit-core rollout restart deploy/audit-core` | Accepts lost; time to first successful accept |
|
||
|
|
| E2 | Receiver recreate | Delete the pod | As E1, plus whether any accepted event is missing after recovery |
|
||
|
|
| E3 | Custody restart | Restart the `platform-pg` primary | Time to first successful accept; 503 count; chain intact after |
|
||
|
|
| E4 | Lease expiry under load | Revoke the runtime lease, wait for ESO refresh | Time to recovery **without** a pod restart — the property `credentials.py` exists to provide |
|
||
|
|
| E5 | Node reboot | Reboot railiance01 | End-to-end recovery; compare against the 2026-08-16 observation of ~40s |
|
||
|
|
|
||
|
|
**Integrity is part of the pass condition, not a separate check.** After every
|
||
|
|
scenario, `python -m audit_core verify-chain` must report intact, and every
|
||
|
|
event the driver believes was accepted must be present. A recovery that loses an
|
||
|
|
accepted event, or forks the chain, is a custody defect and fails the exercise
|
||
|
|
regardless of how fast it was.
|
||
|
|
|
||
|
|
**RPO.** An accept is committed before the sender is told, so the expected RPO
|
||
|
|
for accepted events is zero. The exercise tests that claim rather than assuming
|
||
|
|
it; E2 and E3 are where it would break.
|
||
|
|
|
||
|
|
## Results
|
||
|
|
|
||
|
|
*Not yet run.* Needs a live window on railiance01 and coordination with
|
||
|
|
`user-engine`, since E4 and E5 are visible to the sender.
|
||
|
|
|
||
|
|
On completion, record the measured recovery time per scenario here with dates,
|
||
|
|
raise `tenancy.yaml` `current.V` to 1, note the exercise as the evidence, and
|
||
|
|
set a review date — a recovery exercise from a year ago describes a system that
|
||
|
|
no longer exists.
|