ops-warden/docs/evidence/WARDEN-WP-0027-T02-drill-scenario-2026-08-22.md
tegwick 5ac47520b1
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 1s
docs: establish attended recovery drill scenario
Assistant: codex
Assistant-Model: gpt-5.6-sol
Assistant-Session: 01a0290b-3241-74c3-b868-6049545af836
2026-08-22 22:00:35 +02:00

3.2 KiB

WARDEN-WP-0027-T02 attended drill scenario

Status: preparing — live execution is prohibited.

Window

  • Scenario/window id: WARDEN-WP-0027-T02-DRILL-20260822-01
  • Preparation approval: Warden Desk approve recorded at 2026-08-22T19:22:28Z (metadata only)
  • Approval expires: 2026-08-23T20:00:00Z
  • Live window opens only when the operator gives the exact final go/no-go for this scenario after the owner preflight reports ready_for_live_execution: true
  • Maximum live-window duration after GO: 45 minutes
  • A NO-GO, missing gate, changed cluster identity, or expired approval closes this scenario without mutation

Assigned roles

Responsibility Assigned owner Acceptance
Preparation coordinator and hold-point enforcement ops-warden accepted
OpenBao snapshot, seal/unseal driver, post-unseal verification railiance-platform pending owner receipt
Independent provider console and distinct abort authority railiance-infra pending owner receipt
Two distinct 2-of-3 share custodians available out of band railiance-master custody authority pending quorum receipt
Final live GO or NO-GO human operator deliberately not requested yet

Owner procedure approval is already complete for audit-core, rapp-postgres, railiance-platform, railiance-cluster, and railiance-infra under the RAILIANCE-WP-0024 contract. Those receipts approve the procedure, not this live window.

Current value-safe baseline

The railiance-platform node-reboot preflight at 2026-08-22T19:23:55Z reported:

  • automated_checks_passed: true
  • one Ready node with active k3s
  • platform-pg healthy 1/1 with continuous archiving and a successful backup 17.15 hours old
  • OpenBao initialized and unsealed, Shamir shares=3, threshold=2
  • required ExternalSecret stores Valid and projections SecretSynced
  • audit-core at the reviewed digest, 1/1 Ready, zero restarts
  • secret_values_observed: false

This baseline is not reusable as the final hold-point result. The platform owner must rerun it against the current state and fresh snapshot receipt.

Pending receipts

  • railiance-platform: fresh encrypted, verified, off-host OpenBao Raft snapshot receipt matching the live cluster id and possible applied index; acceptance of the driver role
  • railiance-infra: independent provider-console access verified; acceptance of the distinct abort role
  • railiance-master: two distinct share custodians explicitly available through the approved out-of-band custody paths; no identities or share values in the receipt
  • Fully parameterized read-only owner preflight returns ready_for_live_execution: true and secret_values_observed: false

Final hold point

Once all pending receipts validate, ops-warden presents only the scenario id, bounded duration, owner roles, current preflight result, and stop conditions to the human operator. The live step requires an explicit GO for this exact scenario. Any other response is NO-GO.

No unseal share, token, provider credential, Secret data, decrypted snapshot, or value-derived fingerprint belongs in Git, State Hub, shell history, logs, or chat.