disaster-control/history/2026-07-04-forgejo-backup-strategy-assessment.md
tegwick 96afc72b8a Record Forgejo backup strategy assessment in history/
Persist 2026-07-04 post-drill findings: restore proven, scheduled backups and
off-cluster durability open. Establishes history/ for dated resilience snapshots.
2026-07-04 12:36:50 +02:00

151 lines
No EOL
7.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Forgejo / Railiance Backup Strategy Assessment
Date: 2026-07-04
Author: grok (session on `the-custodian`)
`no_secret_material_recorded: true`
Status: **assessment recorded — pickup from disaster-control**
## Context
Assessment performed after **RAIL-HO-WP-0005-T09** isolated restore drill (2026-07-04).
Restore path is proven; scheduled backups and off-cluster durability are not.
Related workplans and evidence:
| Artifact | Location |
| --- | --- |
| Forgejo migration umbrella | `railiance-infra/workplans/RAIL-HO-WP-0005-forgejo-production-migration.md` |
| Restore drill evidence | `railiance-infra/docs/forgejo-restore-drill-evidence.md` |
| Restore orchestration | `railiance-infra/tools/forgejo-restore-drill.sh` |
| Open production decisions | `the-custodian/docs/forgejo-production-decisions.md` (decision #7 TBD) |
| App/platform handoff | `railiance-apps/docs/app-data-backup-restore-handoff.md` |
## Executive summary
Backup posture is **layered on paper, thin in practice** for Forgejo and most
platform data. We can recover Forgejo if an operator manually runs `forgejo dump`
and knows the restore script. We do **not** yet have a production-grade backup
strategy: scheduled jobs, encrypted off-cluster storage, retention policy, and
Gitea parity are all open before tier-3 (`state-hub`) or production cutover.
**Honest grade:** restore procedure **B+**; backup automation **D**;
off-cluster durability **F**; Gitea (canonical source forge) **at risk**.
## Layer inventory
| Layer | What it should cover | Current state |
| --- | --- | --- |
| **Restore drill** | Prove Forgejo recovery | **Done** (2026-07-04) — manual `forgejo dump` restored in `forgejo-restore-drill`; health 200; `glas-harness`, `key-cape`, 3 org repos verified |
| **Forgejo production** | DB + git repos + packages + attachments | **Manual only** — no schedule, no off-cluster target |
| **CNPG (`forgejo-db`, `apps-pg`)** | DB base backups + WAL | **Not configured** — no `Backup` / `ScheduledBackup` CRs on railiance01 |
| **Cluster runtime (S2)** | k3s state, Helm values, kubeconfig | **Works locally**`railiance-cluster/tools/cmd/railiance-backup-s2` → age-encrypted files in `/opt/backup/railiance/cluster/`, keep 7; no network upload |
| **Platform (S3)** | Legacy PostgreSQL HA, Valkey → Nextcloud | **Documented, not wired**`railiance-platform/Makefile` `backup` target references `tools/cmd/railiance-backup`, but that tool is absent from the repo (relocation planned/cancelled) |
| **Gitea (coulombcore)** | Canonical source forge today | **Weak** — bundled PG + `gitea-shared-storage` PVC; prior audit found no Gitea backup CronJobs; package blob backup not evidenced |
| **Adjacent SSO/MFA** | privacyidea, authelia, lldap | **Scheduled** — daily backup CronJobs exist on railiance01 |
Live cluster check (2026-07-04): `kubectl get backup,scheduledbackup -A` → no
resources. Backup-related CronJobs: `privacyidea-backup`, `authelia-backup`,
`lldap-backup` only.
## What T09 proved (restore, not backup)
The drill validated **recovery mechanics**, not operational backup coverage:
1. `forgejo dump` from production pod (~11.7 MiB: `repos/`, `data/`,
`forgejo-db.sql`, `app.ini`)
2. Manual unzip + file copy + `psql` import (Forgejo 11.0.3 has no
`forgejo restore` CLI)
3. Isolated Helm deploy in `forgejo-restore-drill`
4. API checks: health 200; `coulomb/glas-harness` and `coulomb/key-cape` on
`main`; 3 org repos visible
| Metric | Observed |
| --- | --- |
| RPO (operational today) | Point-in-time of last manual `forgejo dump` |
| RTO (isolated restore) | ~35 minutes (CNPG ready + import + Helm on railiance01) |
| Production impact during drill | None — read-only dump; separate namespace |
Re-run hygiene: `DRILL_CLEAN=1 ./tools/forgejo-restore-drill.sh` (SQL import is
not idempotent).
## Gaps blocking production cutover
1. **No scheduled Forgejo backups** — no CronJob/operator for `forgejo dump`
2. **No CNPG-native backups** for `forgejo-db` — cluster manifest has storage
only (`railiance-platform/helm/forgejo-db-cluster.yaml`)
3. **No approved off-cluster target** — drill archive on workstation `/tmp`
only; encryption/retention undecided (T02 decision #7)
4. **Platform backup tool missing** — S3 `make backup` target without
implementation
5. **Gitea canonical data** — safety contract requires fresh Gitea backup before
each migration wave; no automated Gitea backup on coulombcore
6. **Split-brain restore risk** — restoring only CNPG DB would miss git blobs on
the Forgejo app PVC; full `forgejo dump` is the coherent unit today
## Target architecture (intended, not implemented)
Per `RAIL-HO-WP-0005` repo boundaries:
```text
S2 railiance-cluster
etcd/state, helm values, kubeconfig
local age-encrypted, 7-day retention
S3 railiance-platform
CNPG ScheduledBackup → off-cluster target (MinIO/S3/Nextcloud — TBD)
implement or restore railiance-backup tool
S5 railiance-apps
scheduled forgejo dump (repos + data + SQL + app.ini)
include OCI package blobs; age-encrypt; upload to approved target
Evidence loop (this repo)
quarterly restore drill → isolated namespace
dated evidence under history/ or TestEvidence/
```
Ownership: `railiance-platform` owns backup mechanism; `railiance-apps` owns
readiness gates and restore evidence per
`docs/app-data-backup-restore-handoff.md`.
## Pickup queue (for disaster-control)
Track and drive closure of these items; implementation spans multiple repos.
| # | Item | Owner repo | Blocker |
| --- | --- | --- | --- |
| 1 | Operator decision: backup destination + retention | `railiance-infra` / T02 | Human — decision #7 in `forgejo-production-decisions.md` |
| 2 | CNPG `ScheduledBackup` on `forgejo-db` | `railiance-platform` | #1 + backup target wiring |
| 3 | Scheduled `forgejo dump` → age-encrypt → upload | `railiance-apps` | #1 |
| 4 | Implement or restore `tools/cmd/railiance-backup` | `railiance-platform` | Relocation never completed |
| 5 | Gitea pre-cutover snapshot runbook + evidence | `railiance-infra` / `railiance-apps` | No automated Gitea backup today |
| 6 | `apps-pg` backup coverage | `railiance-platform` | Documented follow-up in `docs/apps-pg.md` |
| 7 | Quarterly restore drill cadence + evidence | `disaster-control` + `railiance-infra` | Calendar + RTO/RPO matrix not written |
| 8 | Populate `RTO-RPO-Matrix.md`, `BackupPolicy.md`, registers | `disaster-control` | Per `INTENT.md` / `specs/ResilienceControlSetup.md` |
**Gate:** tier-3 production repos (`state-hub`, etc.) remain blocked until T09
scheduled automation and T02 backup target are closed (`RAIL-HO-WP-0005`,
`CUST-WP-0054`).
## Suggested RTO/RPO targets (draft — not adopted)
Aligns with `specs/ResilienceControlSetup.md` tier **T1** (irreplaceable source
of truth) until operator confirms:
| System | Draft RPO | Draft RTO | Notes |
| --- | --- | --- | --- |
| Forgejo git + packages | 24h | 48h | Daily `forgejo dump` + CNPG backup |
| Gitea (canonical until cutover) | 24h | 48h | Mirror + dump before each migration wave |
| `forgejo-db` CNPG | 1h24h | 24h | WAL + base backup to off-cluster target |
| k3s cluster state (S2) | 24h | 48h | Existing `railiance-backup-s2` local retention |
## References
- `railiance-infra/docs/forgejo-restore-drill-evidence.md`
- `railiance-infra/tools/forgejo-restore-drill.sh`
- `railiance-platform/helm/forgejo-db-cluster.yaml`
- `railiance-platform/docs/apps-pg.md` (backup follow-up note)
- `railiance-platform/docs/postgresql-ha.md` (intended platform backup)
- `the-custodian/docs/forgejo-production-decisions.md`
- `disaster-control/specs/ResilienceControlSetup.md`
- `disaster-control/INTENT.md`