# Restore Runbook: Full Server Loss **Trigger:** disk failure, VPS deletion, hardware loss, unrecoverable OS corruption. **RTO target:** ≤ 1 day full recovery; ≤ 4–8h for code/content/control plane. **Owner:** Bernd Do **not** restore into a compromised machine. This runbook assumes hardware loss, not active breach (see `IncidentPlaybooks/ransomware.md` for breach). --- ## Prerequisites - [ ] Password-manager copy of age backup keys - [ ] OpenBao unseal material (if restoring railiance01 secrets) - [ ] Nextcloud admin or local backup cache access - [ ] IONOS DNS access - [ ] `warden sign` SSH access to new host - [ ] This repo cloned (from mirror once T01 complete; until then: any copy) --- ## Phase 1 — Provision clean host 1. Provision new Ubuntu server (or recover WSL2 workstation for dev lane only). 2. Apply SSH hardening via Ansible (`railiance-infra`). 3. Install tooling: `age`, `docker`, `k3s`, `helm`, `kubectl`, `ansible`. 4. Record new host identity in `AssetRegister.yaml` after recovery. --- ## Phase 2 — Secret recovery 1. Retrieve OpenBao unseal keys from sealed recovery vault (password manager). 2. If railiance01: follow `railiance-platform` OpenBao deploy/unseal runbooks. 3. Retrieve backup upload credentials from OpenBao if re-wiring backup jobs. 4. **Do not** copy plaintext secrets into Git, chat, or State Hub. --- ## Phase 3 — Data restore (by lane) ### Dev workstation (bnt-lap001) 1. Download latest `db-*.sql.age` and `config-*.tar.gz.age` from Nextcloud. 2. `age -d -i ~/.config/age/railiance-backup.key -o dump.sql db-*.sql.age` 3. Restore PostgreSQL `custodian` database into `infra-postgres-1`. 4. Decrypt and unpack config tar to `$HOME`. 5. Verify: `bin/railiance preflight` exits 0. Reference: `railiance-cluster/docs/backup-restore.md` ### railiance01 — S1 OS config (optional) 1. Decrypt `os-config-*.tar.gz.age` from `/opt/backup/railiance/infra/` (if disk survived) or offsite copy when wired. 2. Restore files or re-apply Ansible playbooks (preferred). ### railiance01 — S2 k8s runtime 1. If rebuilding k3s: install k3s on clean host. 2. Decrypt `helm-values-*.tar.gz.age` — redeploy each release with saved values. 3. Decrypt `kubeconfig-*.yaml.age` if needed for admin access. 4. k3s `state.db` hot backup: prefer **redeploy from Git/Helm** over raw DB restore. Reference: `railiance-cluster/tools/cmd/railiance-backup-s2` ### Gitea (coulombcore) 1. If coulombcore lost: provision host, install Gitea via Ansible. 2. Restore from latest `gitea dump` or git mirrors of critical repos. 3. Verify: clone `coulomb/disaster-control`, API health. ### Forgejo (railiance01) 1. If automated backup wired: download latest dump from Nextcloud. 2. Run `railiance-infra/tools/forgejo-restore-drill.sh` (adapt for production namespace). 3. Verify: `GET /api/v1/version`, org repos visible. Evidence: `TestEvidence/restore-drill-2026-07-04-forgejo.md` ### Platform databases Blocked until `BackupPickupQueue.md` items 3–4 close (platform tool missing). --- ## Phase 4 — Validate services | Check | Command / endpoint | | --- | --- | | k3s nodes | `kubectl get nodes` | | Forgejo health | `curl -fsS https://forgejo.coulomb.social/api/v1/version` | | Gitea health | `curl -fsS https://gitea.coulomb.social/api/v1/version` | | State Hub | `curl -fsS http://127.0.0.1:8000/state/summary` | | Representative app | App-specific smoke (e.g. vergabe health endpoint) | --- ## Phase 5 — DNS cutover 1. Confirm all services healthy on new infrastructure. 2. Update IONOS A records per `PanicSheet.md` §5. 3. Wait for TTL propagation; verify externally (`dig @8.8.8.8`). 4. Monitor for 24h. --- ## Phase 6 — Evidence report Write `TestEvidence/restore-drill-YYYY-MM-DD-full-server-loss.md`: - Start/end times (RTO measured) - Backup artifacts used (RPO achieved) - Systems restored vs redeployed - Issues and follow-up actions - Update `AssetRegister.yaml` if host identities changed --- ## References - `PanicSheet.md` - `BackupPolicy.md` - `AssetRegister.yaml` - `RTO-RPO-Matrix.md`