- INTENT.md: purpose and boundaries of the resilience control plane - specs/ResilienceControlSetup.md: design exploration (tiers, RTO/RPO, 3-2-1) - history/2026-07-04-initial-risk-analysis.md: 12 failure modes ranked, mapped to artifacts - workplans/DISCTL-WP-0002-resilience-foundation.md: registers, policy, first tested restore - DISCTL-WP-0001 T03 (seed first real workplan) done Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
4 KiB
| id | type | title | domain | repo | status | owner | topic_slug | created | updated |
|---|---|---|---|---|---|---|---|---|---|
| DISCTL-WP-0002 | workplan | Resilience foundation: registers, policy, first tested restore | infotech | disaster-control | ready | claude-code | custodian | 2026-07-04 | 2026-07-04 |
Resilience foundation: registers, policy, first tested restore
Stand up the minimum viable Resilience Control Plane for Coulomb Social / Helix Forge: know what exists (registers), what must survive (policy, RTO/RPO), how it comes back (runbook, panic sheet), and prove it once (drill + evidence).
Driven by history/2026-07-04-initial-risk-analysis.md — priority order R2
(untested restores), R9 (silent backup failure), R3 (backup-deleting breach),
R4/R5 (identity & secret recovery), R10 (control plane single-homed), R1
(full server loss). Design source: specs/ResilienceControlSetup.md.
Mirror disaster-control off the primary Gitea (R10)
id: DISCTL-WP-0002-T01
status: todo
priority: high
This repo must survive the disasters it describes. Add at least one mirror on infrastructure independent of CoulombCore/Gitea (e.g. GitHub private mirror or secondary machine), document the mirror location in README, and note how the mirror is refreshed.
Asset and data registers with Tier-0 recovery paths (R4, R7, R8, R12)
id: DISCTL-WP-0002-T02
status: todo
priority: high
Write AssetRegister.yaml and DataRegister.yaml seeded with the current
stack: domains/DNS/registrar, email, Stripe, Bubble.io, OpenRouter, Gitea,
CoulombCore content, k3s, databases, backup storage, secrets/OpenBao. Every T0
entry must record its recovery path (backup admin, MFA recovery, break-glass).
Every T1 entry must name its source of truth and backup mechanism (or mark it
backup: none honestly). Use the tier model from the spec (§2).
ResiliencePolicy and RTO/RPO matrix (R1, R6)
id: DISCTL-WP-0002-T03
status: todo
priority: high
Write ResiliencePolicy.md (operating principle, tier definitions, drill
cadence, evidence requirement) and RTO-RPO-Matrix.md with explicit per-system
targets, starting from the spec's suggested values (§3): RPO ≤ 24h everywhere,
RTO ≤ 1 day full recovery, RTO ≤ 4–8h for code/content/control plane.
BackupPolicy with credential separation and monitoring (R3, R9)
id: DISCTL-WP-0002-T04
status: todo
priority: high
Write BackupPolicy.md: 3-2-1 + immutable/offline model, per-stream mechanisms
(gitea dump + mirrors, restic, k3s etcd snapshots, Bubble/Stripe exports,
SOPS/age secret vault), and two hard rules with verification steps —
(a) backup-storage credentials never live on the production host with delete
permission; (b) backup job success and repository integrity are checked on a
defined schedule with a named alerting path.
Full-server-loss restore runbook and panic sheet (R1, R11)
id: DISCTL-WP-0002-T05
status: todo
priority: medium
Write RestoreRunbooks/01-full-server-loss.md (new machine → Ansible → secret
recovery → data restore → k3s redeploy → validate → DNS cutover → evidence
report) and the one-page panic sheet: where backups are, who has access, how to
rebuild, how to rotate secrets, how to switch DNS — written so a competent
stranger could execute it.
First restore drill with evidence file (R2)
id: DISCTL-WP-0002-T06
status: todo
priority: high
Perform one real restore to a clean target — smallest meaningful scope (e.g.
one Gitea repo from a gitea dump, or the secret vault decrypted from cold) —
and write TestEvidence/restore-drill-2026-MM-DD.md with measured RTO/RPO,
result, issues found, and follow-up actions. This task is the exit criterion
for the workplan: no drill evidence, not finished.
Incident playbooks (R3, R4, R5, R6, R7)
id: DISCTL-WP-0002-T07
status: todo
priority: medium
Write the five playbooks under IncidentPlaybooks/: ransomware/hostile access,
credential leak, accidental delete / data corruption, provider outage
(Bubble/Stripe/OpenRouter/DNS/email), loss of admin identity. Keep each to one
page of numbered steps per the spec (§6).