Persist 2026-07-04 post-drill findings: restore proven, scheduled backups and off-cluster durability open. Establishes history/ for dated resilience snapshots.
7.6 KiB
Forgejo / Railiance Backup Strategy Assessment
Date: 2026-07-04
Author: grok (session on the-custodian)
no_secret_material_recorded: true
Status: assessment recorded — pickup from disaster-control
Context
Assessment performed after RAIL-HO-WP-0005-T09 isolated restore drill (2026-07-04). Restore path is proven; scheduled backups and off-cluster durability are not.
Related workplans and evidence:
| Artifact | Location |
|---|---|
| Forgejo migration umbrella | railiance-infra/workplans/RAIL-HO-WP-0005-forgejo-production-migration.md |
| Restore drill evidence | railiance-infra/docs/forgejo-restore-drill-evidence.md |
| Restore orchestration | railiance-infra/tools/forgejo-restore-drill.sh |
| Open production decisions | the-custodian/docs/forgejo-production-decisions.md (decision #7 TBD) |
| App/platform handoff | railiance-apps/docs/app-data-backup-restore-handoff.md |
Executive summary
Backup posture is layered on paper, thin in practice for Forgejo and most
platform data. We can recover Forgejo if an operator manually runs forgejo dump
and knows the restore script. We do not yet have a production-grade backup
strategy: scheduled jobs, encrypted off-cluster storage, retention policy, and
Gitea parity are all open before tier-3 (state-hub) or production cutover.
Honest grade: restore procedure B+; backup automation D; off-cluster durability F; Gitea (canonical source forge) at risk.
Layer inventory
| Layer | What it should cover | Current state |
|---|---|---|
| Restore drill | Prove Forgejo recovery | Done (2026-07-04) — manual forgejo dump restored in forgejo-restore-drill; health 200; glas-harness, key-cape, 3 org repos verified |
| Forgejo production | DB + git repos + packages + attachments | Manual only — no schedule, no off-cluster target |
CNPG (forgejo-db, apps-pg) |
DB base backups + WAL | Not configured — no Backup / ScheduledBackup CRs on railiance01 |
| Cluster runtime (S2) | k3s state, Helm values, kubeconfig | Works locally — railiance-cluster/tools/cmd/railiance-backup-s2 → age-encrypted files in /opt/backup/railiance/cluster/, keep 7; no network upload |
| Platform (S3) | Legacy PostgreSQL HA, Valkey → Nextcloud | Documented, not wired — railiance-platform/Makefile backup target references tools/cmd/railiance-backup, but that tool is absent from the repo (relocation planned/cancelled) |
| Gitea (coulombcore) | Canonical source forge today | Weak — bundled PG + gitea-shared-storage PVC; prior audit found no Gitea backup CronJobs; package blob backup not evidenced |
| Adjacent SSO/MFA | privacyidea, authelia, lldap | Scheduled — daily backup CronJobs exist on railiance01 |
Live cluster check (2026-07-04): kubectl get backup,scheduledbackup -A → no
resources. Backup-related CronJobs: privacyidea-backup, authelia-backup,
lldap-backup only.
What T09 proved (restore, not backup)
The drill validated recovery mechanics, not operational backup coverage:
forgejo dumpfrom production pod (~11.7 MiB:repos/,data/,forgejo-db.sql,app.ini)- Manual unzip + file copy +
psqlimport (Forgejo 11.0.3 has noforgejo restoreCLI) - Isolated Helm deploy in
forgejo-restore-drill - API checks: health 200;
coulomb/glas-harnessandcoulomb/key-capeonmain; 3 org repos visible
| Metric | Observed |
|---|---|
| RPO (operational today) | Point-in-time of last manual forgejo dump |
| RTO (isolated restore) | ~3–5 minutes (CNPG ready + import + Helm on railiance01) |
| Production impact during drill | None — read-only dump; separate namespace |
Re-run hygiene: DRILL_CLEAN=1 ./tools/forgejo-restore-drill.sh (SQL import is
not idempotent).
Gaps blocking production cutover
- No scheduled Forgejo backups — no CronJob/operator for
forgejo dump - No CNPG-native backups for
forgejo-db— cluster manifest has storage only (railiance-platform/helm/forgejo-db-cluster.yaml) - No approved off-cluster target — drill archive on workstation
/tmponly; encryption/retention undecided (T02 decision #7) - Platform backup tool missing — S3
make backuptarget without implementation - Gitea canonical data — safety contract requires fresh Gitea backup before each migration wave; no automated Gitea backup on coulombcore
- Split-brain restore risk — restoring only CNPG DB would miss git blobs on
the Forgejo app PVC; full
forgejo dumpis the coherent unit today
Target architecture (intended, not implemented)
Per RAIL-HO-WP-0005 repo boundaries:
S2 railiance-cluster
etcd/state, helm values, kubeconfig
local age-encrypted, 7-day retention
S3 railiance-platform
CNPG ScheduledBackup → off-cluster target (MinIO/S3/Nextcloud — TBD)
implement or restore railiance-backup tool
S5 railiance-apps
scheduled forgejo dump (repos + data + SQL + app.ini)
include OCI package blobs; age-encrypt; upload to approved target
Evidence loop (this repo)
quarterly restore drill → isolated namespace
dated evidence under history/ or TestEvidence/
Ownership: railiance-platform owns backup mechanism; railiance-apps owns
readiness gates and restore evidence per
docs/app-data-backup-restore-handoff.md.
Pickup queue (for disaster-control)
Track and drive closure of these items; implementation spans multiple repos.
| # | Item | Owner repo | Blocker |
|---|---|---|---|
| 1 | Operator decision: backup destination + retention | railiance-infra / T02 |
Human — decision #7 in forgejo-production-decisions.md |
| 2 | CNPG ScheduledBackup on forgejo-db |
railiance-platform |
#1 + backup target wiring |
| 3 | Scheduled forgejo dump → age-encrypt → upload |
railiance-apps |
#1 |
| 4 | Implement or restore tools/cmd/railiance-backup |
railiance-platform |
Relocation never completed |
| 5 | Gitea pre-cutover snapshot runbook + evidence | railiance-infra / railiance-apps |
No automated Gitea backup today |
| 6 | apps-pg backup coverage |
railiance-platform |
Documented follow-up in docs/apps-pg.md |
| 7 | Quarterly restore drill cadence + evidence | disaster-control + railiance-infra |
Calendar + RTO/RPO matrix not written |
| 8 | Populate RTO-RPO-Matrix.md, BackupPolicy.md, registers |
disaster-control |
Per INTENT.md / specs/ResilienceControlSetup.md |
Gate: tier-3 production repos (state-hub, etc.) remain blocked until T09
scheduled automation and T02 backup target are closed (RAIL-HO-WP-0005,
CUST-WP-0054).
Suggested RTO/RPO targets (draft — not adopted)
Aligns with specs/ResilienceControlSetup.md tier T1 (irreplaceable source
of truth) until operator confirms:
| System | Draft RPO | Draft RTO | Notes |
|---|---|---|---|
| Forgejo git + packages | 24h | 4–8h | Daily forgejo dump + CNPG backup |
| Gitea (canonical until cutover) | 24h | 4–8h | Mirror + dump before each migration wave |
forgejo-db CNPG |
1h–24h | 2–4h | WAL + base backup to off-cluster target |
| k3s cluster state (S2) | 24h | 4–8h | Existing railiance-backup-s2 local retention |
References
railiance-infra/docs/forgejo-restore-drill-evidence.mdrailiance-infra/tools/forgejo-restore-drill.shrailiance-platform/helm/forgejo-db-cluster.yamlrailiance-platform/docs/apps-pg.md(backup follow-up note)railiance-platform/docs/postgresql-ha.md(intended platform backup)the-custodian/docs/forgejo-production-decisions.mddisaster-control/specs/ResilienceControlSetup.mddisaster-control/INTENT.md