Bootstrap resilience control plane: intent, risk analysis, first workplan

- INTENT.md: purpose and boundaries of the resilience control plane
- specs/ResilienceControlSetup.md: design exploration (tiers, RTO/RPO, 3-2-1)
- history/2026-07-04-initial-risk-analysis.md: 12 failure modes ranked, mapped to artifacts
- workplans/DISCTL-WP-0002-resilience-foundation.md: registers, policy, first tested restore
- DISCTL-WP-0001 T03 (seed first real workplan) done

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
tegwick 2026-07-04 12:45:16 +02:00
parent 96afc72b8a
commit 2f6f489289
5 changed files with 650 additions and 0 deletions

View file

@ -0,0 +1,294 @@
Resilience Control Setup Exploration
## Introduction
This is a rough outline for how we will secure Coulomb Social / Helix Forge against disasters.
We will **not** start with “buy a backup tool.” instead we will set up a small **Resilience Control Plane**:
a documented, automated, tested system that knows what exists, what must survive, how it is rebuilt, and whether restore actually works.
The useful framing is: **business continuity + disaster recovery + cyber recovery + evidence**. ISO 22301 is the standard family for business-continuity management, while NIST CSF 2.0 frames cyber resilience around Govern, Identify, Protect, Detect, Respond, and Recover. NIST SP 800-34 remains a practical reference for IT contingency planning, and CISA explicitly recommends offline, encrypted backups that are regularly tested. ([ISO][1])
Remark: The Resilience Control Plane might evolve into a product of it's own later. The repository disaster-control for now is about resilience for the coulomb project and adjacent endevors. Extracting a product is for a later day.
## 1. Treat Coulomb / Helix as a recoverable organization
Given your current shape — Bubble.io legacy app, Stripe memberships, OpenRouter usage, CoulombCore on Ubuntu, Gitea, Decap CMS, k3s, Ansible, AI/webhook services, many repos — create one repo, for example:
`disaster-control`
Inside it:
```text
disaster-control/
INTENT.md
ResiliencePolicy.md
AssetRegister.yaml
DataRegister.yaml
ServiceDependencyMap.md
RTO-RPO-Matrix.md
BackupPolicy.md
RestoreRunbooks/
01-full-server-loss.md
02-gitea-restore.md
03-k3s-restore.md
04-bubble-export-restore.md
05-security-breach-recovery.md
06-secret-rotation.md
IncidentPlaybooks/
ransomware.md
credential-leak.md
accidental-delete.md
provider-outage.md
data-corruption.md
TestEvidence/
restore-drill-YYYY-MM-DD.md
```
This becomes the “single source of truth” for recovery, independent of any one server.
## 2. Classify assets by recovery importance
Use tiers. This prevents over-engineering everything.
| Tier | Asset type | Examples | Target |
| ------------------------------------ | ----------------------------------------------------------------------------------------- | ------------------------------------------ | ------ |
| **T0 Identity / trust / money** | Domains, DNS, registrar, email, Stripe, OAuth, admin accounts, secrets, SSH keys | Must be recoverable even after breach | |
| **T1 Irreplaceable source of truth** | Gitea repos, CoulombCore content, user/member data, product registry, paid customer state | Strong backups, tested restore | |
| **T2 Runtime platforms** | Ubuntu server, k3s, Gitea instance, Decap CMS, webhooks, databases | Rebuild from IaC + restore data | |
| **T3 Derived / rebuildable** | caches, generated artifacts, embeddings, logs beyond retention, local model outputs | Recompute where possible | |
| **T4 External dependencies** | Bubble.io, Stripe, OpenRouter, domain/DNS provider, email provider | Export, monitor, fallback, document limits | |
The key idea: **servers are disposable; data, identity, secrets, evidence, and payment/account records are not.**
## 3. Define RTO and RPO per system
Use explicit recovery objectives:
| System | Suggested RPO | Suggested RTO | Notes |
| ------------------------------ | -----------------------------------------: | ------------: | ------------------------------------------------------------------------------------------------------ |
| Gitea repos | 1h or better | 48h | Mirror repos plus Gitea dump |
| CoulombCore content | 1h6h | 48h | Back up content store + DB |
| k3s cluster state | 1h24h | 8h | Prefer rebuildable cluster + snapshots |
| Bubble.io database | 24h initially | 24h | Export regularly; Bubble has database restore/export capabilities ([manual.bubble.io][2]) |
| Stripe financial/customer data | 24h7d | 24h | Stripe can export/sync data to storage/warehouse destinations ([Stripe Dokumentation][3]) |
| Secrets | immediate rotation plan | 24h | Back up encrypted recovery material, not plaintext secrets |
| DNS/domains | manual export after changes | 2h | Registrar lock + MFA + backup admin |
| OpenRouter dependency | no data RPO unless storing prompts/results | 14h fallback | Track API keys, quotas, usage; OpenRouter exposes key/limit/credit info via API/docs ([OpenRouter][4]) |
For an early-stage environment, I would optimize for **RPO ≤ 24h everywhere, RTO ≤ 1 day for full recovery, and RTO ≤ 48h for code/content/control-plane recovery**.
## 4. Implement the backup architecture
Use a practical **3-2-1 + immutable/offline + tested restore** model: at least three copies, on two storage types, one offsite; plus one copy that an attacker cannot easily delete or encrypt. The UK NCSC describes 3-2-1 as a common resilient-backup pattern, and CISA emphasizes offline, encrypted, regularly tested backups for ransomware resilience. ([National Cyber Security Centre][5])
For Coulomb / Helix, I would set it up like this:
```text
Production server
├─ local snapshots for fast rollback
├─ encrypted daily backup to offsite object storage
├─ immutable/offline weekly backup
└─ monthly restore drill to clean machine
```
Concrete backup streams:
| Stream | Mechanism |
| ----------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Gitea** | Scheduled `gitea dump` plus repo mirrors. Giteas official docs describe `dump` as creating a ZIP backup usable for restore. ([Gitea Dokumentation][6]) |
| **Git repos** | Mirror clone to separate location/provider. GitHubs docs recommend `git clone --mirror` for repository backup including history. ([GitHub Docs][7]) |
| **k3s** | k3s datastore snapshots. k3s documents `etcd-snapshot` for managing and restoring embedded etcd snapshots. ([K3s][8]) |
| **Kubernetes resources/PVs** | Velero for cluster resources and persistent volumes. Velero is designed for Kubernetes backup, restore, migration, and cluster replication. ([Velero][9]) |
| **Files / content / configs** | Restic or Borg-style encrypted backups. Restic supports encrypted backups to local, self-hosted, and online storage. ([Restic][10]) |
| **Bubble.io** | Scheduled app/data exports and documented restore workflow. Bubble documents database restore and export/import tooling. ([manual.bubble.io][2]) |
| **Stripe** | Periodic CSV exports or Stripe Data Pipeline/cloud storage export for accounting/customer continuity. ([Stripe Dokumentation][3]) |
| **Secrets** | SOPS/age-encrypted sealed backup of recovery secrets, plus printed/offline recovery material where appropriate. After breach: rotate, dont blindly restore. |
| **DNS / domains** | Export zone records, registrar settings, recovery codes, admin contacts, and MFA recovery process. |
## 5. Make rebuild the default, not server resurrection
For the Ubuntu/k3s/Gitea environment, the best recovery pattern is:
```text
New clean machine
→ install base OS
→ apply Ansible
→ restore secrets from encrypted recovery vault
→ restore Gitea / DB / content
→ restore k3s state or redeploy manifests
→ validate services
→ switch DNS
```
This means your real disaster-recovery asset is not the server image. It is:
1. **Ansible/IaC**
2. **data backups**
3. **secret recovery**
4. **DNS cutover**
5. **tested runbooks**
For a security breach, never restore into the compromised machine. Build a clean environment, restore only known-good data, rotate credentials, and preserve evidence.
## 6. Create incident playbooks
You need at least five.
### A. Full server loss
Trigger: disk failure, VPS deletion, hardware loss.
Recovery:
1. Provision new server.
2. Apply Ansible baseline.
3. Restore Gitea dump and repo mirrors.
4. Restore databases/content.
5. Restore k3s from snapshot or redeploy workloads.
6. Validate health checks.
7. Move DNS.
8. Write evidence report.
### B. Accidental deletion or bad migration
Trigger: deleted repo, corrupted database, bad script.
Recovery:
1. Stop writes.
2. Clone latest backup into staging.
3. Identify last good point.
4. Restore selectively.
5. Diff recovered data against production.
6. Resume writes.
### C. Ransomware / hostile access
Trigger: suspicious encryption, unknown admin access, altered repos, leaked token.
Recovery:
1. Isolate systems.
2. Preserve logs/snapshots for evidence.
3. Revoke exposed credentials.
4. Build clean replacement environment.
5. Restore from backup before compromise time.
6. Rotate all secrets.
7. Force password/session reset where needed.
8. Publish internal incident note.
### D. Provider outage
Trigger: Bubble, Stripe, OpenRouter, DNS provider, email provider unavailable.
Recovery:
1. Switch status page / static fallback.
2. Disable dependent features gracefully.
3. Queue work where possible.
4. Use alternate LLM provider path where feasible.
5. Reconcile once provider returns.
### E. Loss of admin identity
Trigger: locked out of registrar, Gitea, Stripe, Bubble, email, cloud.
Recovery:
1. Use documented break-glass account.
2. Use offline recovery codes.
3. Verify ownership records.
4. Rotate compromised accounts.
5. Update evidence log.
## 7. Separate backup credentials from production credentials
A common failure mode: the attacker gets production admin access and deletes the backups too. So:
* Backup storage credentials must **not** live on the main server with delete permissions.
* Use append-only or immutable storage where possible.
* Use separate admin accounts for production, backup, registrar, and payment systems.
* Enable MFA everywhere.
* Keep offline recovery codes.
* Keep a minimal “break-glass” process documented.
## 8. Add monitoring and evidence
You want the system to prove itself.
Minimum checks:
| Check | Frequency |
| --------------------------------- | --------------------------: |
| Backup job succeeded | daily |
| Backup repository integrity check | weekly |
| Restore one repo to staging | weekly |
| Restore Gitea dump to staging | monthly |
| Rebuild full server from Ansible | quarterly |
| Rotate critical secrets | quarterly or after incident |
| Review asset register | monthly |
| Review external dependencies | monthly |
Every drill should produce a short file:
```text
TestEvidence/restore-drill-2026-07-04.md
```
With:
```markdown
# Restore Drill: Gitea
Date: 2026-07-04
Backup used: gitea-dump-...
Target: clean staging server
Result: success / partial / failed
RTO measured: ...
RPO measured: ...
Issues found:
- ...
Actions:
- ...
```
This turns resilience into a learning loop.
## 9. Minimal viable setup
For the first version, I would implement this:
1. **Create `coulomb-resilience-control` repo.**
2. **Write AssetRegister.yaml** for domains, Gitea, Bubble, Stripe, OpenRouter, k3s, databases, storage, secrets.
3. **Automate Gitea backup** using `gitea dump` plus repo mirrors.
4. **Automate server/content backups** with encrypted restic/Borg-style offsite backup.
5. **Add k3s etcd snapshots** if k3s state matters; otherwise prefer full redeploy from manifests/Ansible.
6. **Export Bubble data** on a fixed schedule while Bubble remains in the architecture.
7. **Export Stripe/accounting data** to storage/warehouse or at least scheduled CSV snapshots.
8. **Create one clean restore machine** and perform a monthly restore drill.
9. **Document secret rotation and breach recovery.**
10. **Create a one-page “panic sheet”**: where backups are, who has access, how to rebuild, how to rotate secrets, how to switch DNS.
## 10. Suggested operating principle
For Coulomb / Helix Forge, I would use this rule:
> **Every capability that becomes valuable must declare its source of truth, recovery method, backup schedule, RTO, RPO, owner, and restore test evidence.**
That fits naturally with your capability-registry thinking. A repo or service is not “production-worthy” until its recovery story is known.
A good next artifact would be a `ResiliencePolicy.md` plus `AssetRegister.yaml` seed for your current Coulomb/Helix stack.
[1]: https://www.iso.org/standard/75106.html?utm_source=chatgpt.com "ISO 22301:2019 - Business continuity management systems"
[2]: https://manual.bubble.io/help-guides/maintaining-an-application/database-maintenance/database-copy-and-backups?utm_source=chatgpt.com "Restoring database backups"
[3]: https://docs.stripe.com/data/data-pipeline?utm_source=chatgpt.com "Data Pipeline | Stripe Documentation"
[4]: https://openrouter.ai/docs/api/reference/limits?utm_source=chatgpt.com "API Rate Limits - Manage Model Usage and Quotas"
[5]: https://www.ncsc.gov.uk/blog-post/offline-backups-in-an-online-world?utm_source=chatgpt.com "Offline backups in an online world"
[6]: https://docs.gitea.com/administration/backup-and-restore?utm_source=chatgpt.com "Backup and Restore"
[7]: https://docs.github.com/en/repositories/archiving-a-github-repository/backing-up-a-repository?utm_source=chatgpt.com "Backing up a repository"
[8]: https://docs.k3s.io/cli/etcd-snapshot?utm_source=chatgpt.com "k3s etcd-snapshot"
[9]: https://velero.io/docs/main/?utm_source=chatgpt.com "Velero Docs - Overview"
[10]: https://restic.net/?utm_source=chatgpt.com "restic · Backups done right!"
xxx