disaster-control/workplans/DISCTL-WP-0002-resilience-foundation.md
tegwick 2f6f489289 Bootstrap resilience control plane: intent, risk analysis, first workplan
- INTENT.md: purpose and boundaries of the resilience control plane
- specs/ResilienceControlSetup.md: design exploration (tiers, RTO/RPO, 3-2-1)
- history/2026-07-04-initial-risk-analysis.md: 12 failure modes ranked, mapped to artifacts
- workplans/DISCTL-WP-0002-resilience-foundation.md: registers, policy, first tested restore
- DISCTL-WP-0001 T03 (seed first real workplan) done

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-04 12:45:16 +02:00

120 lines
4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
id: DISCTL-WP-0002
type: workplan
title: "Resilience foundation: registers, policy, first tested restore"
domain: infotech
repo: disaster-control
status: ready
owner: claude-code
topic_slug: custodian
created: "2026-07-04"
updated: "2026-07-04"
---
# Resilience foundation: registers, policy, first tested restore
Stand up the minimum viable Resilience Control Plane for Coulomb Social /
Helix Forge: know what exists (registers), what must survive (policy, RTO/RPO),
how it comes back (runbook, panic sheet), and prove it once (drill + evidence).
Driven by `history/2026-07-04-initial-risk-analysis.md` — priority order R2
(untested restores), R9 (silent backup failure), R3 (backup-deleting breach),
R4/R5 (identity & secret recovery), R10 (control plane single-homed), R1
(full server loss). Design source: `specs/ResilienceControlSetup.md`.
## Mirror disaster-control off the primary Gitea (R10)
```task
id: DISCTL-WP-0002-T01
status: todo
priority: high
```
This repo must survive the disasters it describes. Add at least one mirror on
infrastructure independent of CoulombCore/Gitea (e.g. GitHub private mirror or
secondary machine), document the mirror location in README, and note how the
mirror is refreshed.
## Asset and data registers with Tier-0 recovery paths (R4, R7, R8, R12)
```task
id: DISCTL-WP-0002-T02
status: todo
priority: high
```
Write `AssetRegister.yaml` and `DataRegister.yaml` seeded with the current
stack: domains/DNS/registrar, email, Stripe, Bubble.io, OpenRouter, Gitea,
CoulombCore content, k3s, databases, backup storage, secrets/OpenBao. Every T0
entry must record its recovery path (backup admin, MFA recovery, break-glass).
Every T1 entry must name its source of truth and backup mechanism (or mark it
`backup: none` honestly). Use the tier model from the spec (§2).
## ResiliencePolicy and RTO/RPO matrix (R1, R6)
```task
id: DISCTL-WP-0002-T03
status: todo
priority: high
```
Write `ResiliencePolicy.md` (operating principle, tier definitions, drill
cadence, evidence requirement) and `RTO-RPO-Matrix.md` with explicit per-system
targets, starting from the spec's suggested values (§3): RPO ≤ 24h everywhere,
RTO ≤ 1 day full recovery, RTO ≤ 48h for code/content/control plane.
## BackupPolicy with credential separation and monitoring (R3, R9)
```task
id: DISCTL-WP-0002-T04
status: todo
priority: high
```
Write `BackupPolicy.md`: 3-2-1 + immutable/offline model, per-stream mechanisms
(gitea dump + mirrors, restic, k3s etcd snapshots, Bubble/Stripe exports,
SOPS/age secret vault), and two hard rules with verification steps —
(a) backup-storage credentials never live on the production host with delete
permission; (b) backup job success and repository integrity are checked on a
defined schedule with a named alerting path.
## Full-server-loss restore runbook and panic sheet (R1, R11)
```task
id: DISCTL-WP-0002-T05
status: todo
priority: medium
```
Write `RestoreRunbooks/01-full-server-loss.md` (new machine → Ansible → secret
recovery → data restore → k3s redeploy → validate → DNS cutover → evidence
report) and the one-page panic sheet: where backups are, who has access, how to
rebuild, how to rotate secrets, how to switch DNS — written so a competent
stranger could execute it.
## First restore drill with evidence file (R2)
```task
id: DISCTL-WP-0002-T06
status: todo
priority: high
```
Perform one real restore to a clean target — smallest meaningful scope (e.g.
one Gitea repo from a `gitea dump`, or the secret vault decrypted from cold) —
and write `TestEvidence/restore-drill-2026-MM-DD.md` with measured RTO/RPO,
result, issues found, and follow-up actions. This task is the exit criterion
for the workplan: no drill evidence, not finished.
## Incident playbooks (R3, R4, R5, R6, R7)
```task
id: DISCTL-WP-0002-T07
status: todo
priority: medium
```
Write the five playbooks under `IncidentPlaybooks/`: ransomware/hostile access,
credential leak, accidental delete / data corruption, provider outage
(Bubble/Stripe/OpenRouter/DNS/email), loss of admin identity. Keep each to one
page of numbered steps per the spec (§6).