disaster-control/history/2026-07-04-forgejo-backup-strategy-assessment.md
tegwick 96afc72b8a Record Forgejo backup strategy assessment in history/
Persist 2026-07-04 post-drill findings: restore proven, scheduled backups and
off-cluster durability open. Establishes history/ for dated resilience snapshots.
2026-07-04 12:36:50 +02:00

7.6 KiB
Raw Permalink Blame History

Forgejo / Railiance Backup Strategy Assessment

Date: 2026-07-04
Author: grok (session on the-custodian)
no_secret_material_recorded: true
Status: assessment recorded — pickup from disaster-control

Context

Assessment performed after RAIL-HO-WP-0005-T09 isolated restore drill (2026-07-04). Restore path is proven; scheduled backups and off-cluster durability are not.

Related workplans and evidence:

Artifact Location
Forgejo migration umbrella railiance-infra/workplans/RAIL-HO-WP-0005-forgejo-production-migration.md
Restore drill evidence railiance-infra/docs/forgejo-restore-drill-evidence.md
Restore orchestration railiance-infra/tools/forgejo-restore-drill.sh
Open production decisions the-custodian/docs/forgejo-production-decisions.md (decision #7 TBD)
App/platform handoff railiance-apps/docs/app-data-backup-restore-handoff.md

Executive summary

Backup posture is layered on paper, thin in practice for Forgejo and most platform data. We can recover Forgejo if an operator manually runs forgejo dump and knows the restore script. We do not yet have a production-grade backup strategy: scheduled jobs, encrypted off-cluster storage, retention policy, and Gitea parity are all open before tier-3 (state-hub) or production cutover.

Honest grade: restore procedure B+; backup automation D; off-cluster durability F; Gitea (canonical source forge) at risk.

Layer inventory

Layer What it should cover Current state
Restore drill Prove Forgejo recovery Done (2026-07-04) — manual forgejo dump restored in forgejo-restore-drill; health 200; glas-harness, key-cape, 3 org repos verified
Forgejo production DB + git repos + packages + attachments Manual only — no schedule, no off-cluster target
CNPG (forgejo-db, apps-pg) DB base backups + WAL Not configured — no Backup / ScheduledBackup CRs on railiance01
Cluster runtime (S2) k3s state, Helm values, kubeconfig Works locallyrailiance-cluster/tools/cmd/railiance-backup-s2 → age-encrypted files in /opt/backup/railiance/cluster/, keep 7; no network upload
Platform (S3) Legacy PostgreSQL HA, Valkey → Nextcloud Documented, not wiredrailiance-platform/Makefile backup target references tools/cmd/railiance-backup, but that tool is absent from the repo (relocation planned/cancelled)
Gitea (coulombcore) Canonical source forge today Weak — bundled PG + gitea-shared-storage PVC; prior audit found no Gitea backup CronJobs; package blob backup not evidenced
Adjacent SSO/MFA privacyidea, authelia, lldap Scheduled — daily backup CronJobs exist on railiance01

Live cluster check (2026-07-04): kubectl get backup,scheduledbackup -A → no resources. Backup-related CronJobs: privacyidea-backup, authelia-backup, lldap-backup only.

What T09 proved (restore, not backup)

The drill validated recovery mechanics, not operational backup coverage:

  1. forgejo dump from production pod (~11.7 MiB: repos/, data/, forgejo-db.sql, app.ini)
  2. Manual unzip + file copy + psql import (Forgejo 11.0.3 has no forgejo restore CLI)
  3. Isolated Helm deploy in forgejo-restore-drill
  4. API checks: health 200; coulomb/glas-harness and coulomb/key-cape on main; 3 org repos visible
Metric Observed
RPO (operational today) Point-in-time of last manual forgejo dump
RTO (isolated restore) ~35 minutes (CNPG ready + import + Helm on railiance01)
Production impact during drill None — read-only dump; separate namespace

Re-run hygiene: DRILL_CLEAN=1 ./tools/forgejo-restore-drill.sh (SQL import is not idempotent).

Gaps blocking production cutover

  1. No scheduled Forgejo backups — no CronJob/operator for forgejo dump
  2. No CNPG-native backups for forgejo-db — cluster manifest has storage only (railiance-platform/helm/forgejo-db-cluster.yaml)
  3. No approved off-cluster target — drill archive on workstation /tmp only; encryption/retention undecided (T02 decision #7)
  4. Platform backup tool missing — S3 make backup target without implementation
  5. Gitea canonical data — safety contract requires fresh Gitea backup before each migration wave; no automated Gitea backup on coulombcore
  6. Split-brain restore risk — restoring only CNPG DB would miss git blobs on the Forgejo app PVC; full forgejo dump is the coherent unit today

Target architecture (intended, not implemented)

Per RAIL-HO-WP-0005 repo boundaries:

S2 railiance-cluster
  etcd/state, helm values, kubeconfig
  local age-encrypted, 7-day retention

S3 railiance-platform
  CNPG ScheduledBackup → off-cluster target (MinIO/S3/Nextcloud — TBD)
  implement or restore railiance-backup tool

S5 railiance-apps
  scheduled forgejo dump (repos + data + SQL + app.ini)
  include OCI package blobs; age-encrypt; upload to approved target

Evidence loop (this repo)
  quarterly restore drill → isolated namespace
  dated evidence under history/ or TestEvidence/

Ownership: railiance-platform owns backup mechanism; railiance-apps owns readiness gates and restore evidence per docs/app-data-backup-restore-handoff.md.

Pickup queue (for disaster-control)

Track and drive closure of these items; implementation spans multiple repos.

# Item Owner repo Blocker
1 Operator decision: backup destination + retention railiance-infra / T02 Human — decision #7 in forgejo-production-decisions.md
2 CNPG ScheduledBackup on forgejo-db railiance-platform #1 + backup target wiring
3 Scheduled forgejo dump → age-encrypt → upload railiance-apps #1
4 Implement or restore tools/cmd/railiance-backup railiance-platform Relocation never completed
5 Gitea pre-cutover snapshot runbook + evidence railiance-infra / railiance-apps No automated Gitea backup today
6 apps-pg backup coverage railiance-platform Documented follow-up in docs/apps-pg.md
7 Quarterly restore drill cadence + evidence disaster-control + railiance-infra Calendar + RTO/RPO matrix not written
8 Populate RTO-RPO-Matrix.md, BackupPolicy.md, registers disaster-control Per INTENT.md / specs/ResilienceControlSetup.md

Gate: tier-3 production repos (state-hub, etc.) remain blocked until T09 scheduled automation and T02 backup target are closed (RAIL-HO-WP-0005, CUST-WP-0054).

Suggested RTO/RPO targets (draft — not adopted)

Aligns with specs/ResilienceControlSetup.md tier T1 (irreplaceable source of truth) until operator confirms:

System Draft RPO Draft RTO Notes
Forgejo git + packages 24h 48h Daily forgejo dump + CNPG backup
Gitea (canonical until cutover) 24h 48h Mirror + dump before each migration wave
forgejo-db CNPG 1h24h 24h WAL + base backup to off-cluster target
k3s cluster state (S2) 24h 48h Existing railiance-backup-s2 local retention

References

  • railiance-infra/docs/forgejo-restore-drill-evidence.md
  • railiance-infra/tools/forgejo-restore-drill.sh
  • railiance-platform/helm/forgejo-db-cluster.yaml
  • railiance-platform/docs/apps-pg.md (backup follow-up note)
  • railiance-platform/docs/postgresql-ha.md (intended platform backup)
  • the-custodian/docs/forgejo-production-decisions.md
  • disaster-control/specs/ResilienceControlSetup.md
  • disaster-control/INTENT.md