diff --git a/BackupPickupQueue.md b/BackupPickupQueue.md new file mode 100644 index 0000000..5d67215 --- /dev/null +++ b/BackupPickupQueue.md @@ -0,0 +1,82 @@ +# Backup Pickup Queue + +**Updated:** 2026-07-07 +**Owner:** Bernd (`disaster-control` tracks; implementation in owner repos) +**Source:** `history/2026-07-04-forgejo-backup-strategy-assessment.md` pickup queue +**Review cadence:** Monthly, or when a cross-repo workplan closes + +When an item reaches `done`, update `AssetRegister.yaml`, `DataRegister.yaml`, +`BackupPolicy.md`, and `RTO-RPO-Matrix.md` to reflect the new posture. + +--- + +## Queue + +| # | Item | Owner repo | Blocker | Status | Last checked | Evidence | +| --- | --- | --- | --- | --- | --- | --- | +| 1 | Backup destination + retention decision | `railiance-infra` / T02 | — | **done** | 2026-07-09 | `the-custodian/docs/forgejo-production-decisions.md` — Option A: Nextcloud WebDAV, 14 daily + 4 weekly, age encrypt | +| 2 | CNPG `ScheduledBackup` on `forgejo-db` | `railiance-platform` | #1 done; Phase 1 uses logical pg_dump per decision (WAL deferred) | **open** | 2026-07-07 | No `ScheduledBackup` CR found; `helm/forgejo-db-cluster.yaml` has storage only | +| 3 | Scheduled `forgejo dump` → age → upload | `railiance-platform` | #1 done; needs cron + Nextcloud wiring | **open** | 2026-07-07 | Restore proven (T09); no CronJob. Owner shifted to platform per 2026-07-09 decision | +| 4 | Restore `tools/cmd/railiance-backup` to platform | `railiance-platform` | Relocation cancelled (RAIL-PL-WP-0001-T05); needs re-implementation | **open** | 2026-07-07 | `make backup` references missing `tools/cmd/railiance-backup` | +| 5 | Gitea pre-cutover snapshot runbook + evidence | `railiance-infra` / `railiance-apps` | No automated Gitea backup today | **open** | 2026-07-07 | Safety contract requires fresh dump before each migration wave; no runbook evidenced | +| 6 | `apps-pg` backup coverage | `railiance-platform` | CNPG ScheduledBackup not configured | **open** | 2026-07-07 | `railiance-platform/docs/apps-pg.md` — follow-up note; blocks app production-trust gates | + +--- + +## Status legend + +| Status | Meaning | +| --- | --- | +| `open` | Not started or no implementation evidenced | +| `in_progress` | Active workplan or PR in flight | +| `done` | Implemented, verified, registers updated | +| `deferred` | Explicitly postponed with decision record | + +--- + +## Promotion gate (from BackupPolicy.md §7) + +Tier-3 repo cutovers remain blocked until: + +1. Items **3** and **4** (or equivalent platform lane) deliver automated daily backups. +2. **7 consecutive days** of successful daily runs. +3. One restore drill uses a **Nextcloud artifact** (not workstation `/tmp`). + +Current gate status: **not met** — item 3 open, item 4 open. + +--- + +## Notes + +### Item 1 — closed 2026-07-09 + +Operator chose Option A: extend the existing Railiance platform backup lane +(age + Nextcloud WebDAV). Recorded in `BackupPolicy.md` §7 and +`RTO-RPO-Matrix.md`. + +### Item 2 — scope clarification + +2026-07-09 decision explicitly defers CNPG WAL/PITR to Phase 2. Phase 1 uses +daily logical `pg_dump` from `forgejo-db` (overlaps item 3). A +`ScheduledBackup` CR may still be valuable later but is not the Phase 1 gate. + +### Item 3 — owner correction + +Assessment originally assigned scheduled `forgejo dump` to `railiance-apps`. +2026-07-09 decision assigns cron, encrypt, upload, and retention to +`railiance-platform`. `railiance-apps` constraint: no backup secrets in Git. + +### Item 4 — re-implementation path + +The old `railiance-backup` script still exists in `railiance-cluster/tools/cmd/` +(dev-workstation scope). Platform needs a new or relocated tool covering +PostgreSQL HA, Valkey, and Forgejo lane uploads. See +`railiance-platform/workplans/archived/RAIL-PL-WP-0001-platform-baseline.md` T05. + +--- + +## Review log + +| Date | Reviewer | Notes | +| --- | --- | --- | +| 2026-07-07 | grok | Queue created (DISCTL-WP-0003-T04). Item 1 done. Items 2–6 open. | \ No newline at end of file diff --git a/RTO-RPO-Matrix.md b/RTO-RPO-Matrix.md new file mode 100644 index 0000000..c1e6ecc --- /dev/null +++ b/RTO-RPO-Matrix.md @@ -0,0 +1,95 @@ +# RTO / RPO Matrix + +**Updated:** 2026-07-07 +**Owner:** Bernd +**Scope:** Backup lanes — initial seed (DISCTL-WP-0003-T03) +**Related:** `AssetRegister.yaml` · `BackupPolicy.md` · `DataRegister.yaml` + +Recovery Time Objective (RTO): maximum acceptable downtime. +Recovery Point Objective (RPO): maximum acceptable data loss (time since last good backup). + +**Status legend:** + +| Status | Meaning | +| --- | --- | +| `adopted` | Operator-confirmed target; policy applies | +| `draft` | Proposed from assessment; pending operator confirmation | +| `degraded` | Target adopted but current posture cannot meet it | + +Global baseline (from `specs/ResilienceControlSetup.md` §3): **RPO ≤ 24h everywhere**; +**RTO ≤ 1 day** for full recovery; **RTO ≤ 4–8h** for code/content/control plane. + +--- + +## Backup lane targets + +| System | Register ID | RPO target | RTO target | Status | Evidence basis | Current posture | +| --- | --- | --- | --- | --- | --- | --- | +| Forgejo git + packages | `backup-forgejo-production` | 24h | 4h | **adopted** | Decision #7 (2026-07-09); T09 drill | **degraded** — manual dump only; automation not wired | +| `forgejo-db` CNPG | `backup-forgejo-production` | 24h (Phase 1 logical dump) | 4h | **adopted** | Decision #7 (2026-07-09) | **degraded** — no scheduled pg_dump | +| Gitea (canonical until cutover) | `backup-gitea-canonical` | 24h | 4–8h | draft | Assessment §suggested targets | **degraded** — no automated backup | +| k3s cluster state (S2) | `backup-s2-k8s-runtime` | 24h | 4–8h | draft | `railiance-backup-s2` local retention | Manual backup; on-node only | +| S1 OS config | `backup-s1-os-config` | 24h | 4–8h | draft | Manual `make backup`; no schedule | Manual backup; on-node only | +| Dev workstation custodian DB | `backup-dev-workstation` | 24h | 4h | draft | Daily cron 02:00; preflight gate | Operational — meets target if cron healthy | +| Dev workstation operator config | `backup-dev-workstation` | 24h | 4h | draft | Daily cron 02:00; preflight gate | Operational — meets target if cron healthy | +| S3 platform PostgreSQL | `backup-s3-platform` | 24h | 4h | **adopted** | Decision #7 (extends platform lane) | **degraded** — tool missing | +| `apps-pg` CNPG | `backup-apps-pg` | 24h | 4–8h | draft | `railiance-platform/docs/apps-pg.md` | **degraded** — no ScheduledBackup CR | +| Adjacent SSO/MFA stack | `backup-adjacent-sso` | 24h | 4–8h | draft | Daily CronJobs on railiance01 | Operational — scheduled daily | + +--- + +## Measured results (restore drills) + +| Date | System | RPO achieved | RTO achieved | Result | Evidence | +| --- | --- | --- | --- | --- | --- | +| 2026-07-04 | Forgejo (isolated restore) | Point-in-time of manual `forgejo dump` | ~3–5 min (isolated namespace) | success | `TestEvidence/restore-drill-2026-07-04-forgejo.md` | + +No other systems have measured drill results yet. Next drills should produce +dated files under `TestEvidence/` (DISCTL-WP-0003-T05). + +--- + +## Adopted vs operational gap + +Systems marked **adopted** with **degraded** posture are policy violations until +the promotion gate in `BackupPolicy.md` §7 is met: + +1. Automated daily backups succeed for 7 consecutive days. +2. One restore drill uses a Nextcloud artifact (not workstation `/tmp`). + +| System | Adopted RPO/RTO | Blocker to operational | +| --- | --- | --- | +| Forgejo git + packages | 24h / 4h | No scheduled `forgejo dump` CronJob | +| `forgejo-db` CNPG | 24h / 4h | No daily pg_dump automation | +| S3 platform PostgreSQL | 24h / 4h | `railiance-backup` tool absent from platform | + +--- + +## Draft targets pending operator confirmation + +The following rows use assessment-suggested values. Confirm or adjust: + +| System | Proposed RPO | Proposed RTO | Question for operator | +| --- | --- | --- | --- | +| Gitea canonical | 24h | 4–8h | Accept 24h RPO until cutover, or require mirror + hourly? | +| k3s cluster state (S2) | 24h | 4–8h | Accept manual-only S2 backup, or require daily cron + offsite? | +| S1 OS config | 24h | 4–8h | Accept manual-only, or require scheduled + offsite? | +| Dev workstation | 24h | 4h | Confirm daily cron meets RPO; any need for sub-daily? | +| `apps-pg` | 24h | 4–8h | Confirm before app production-trust gates | +| Adjacent SSO | 24h | 4–8h | Confirm daily CronJob retention and restore path | + +--- + +## Cross-references + +| Artifact | Link | +| --- | --- | +| Stream inventory | `AssetRegister.yaml` → `backup_streams` | +| Data classes | `DataRegister.yaml` → `data_classes` | +| Backup mechanisms | `BackupPolicy.md` §6–§7 | +| Forgejo adoption | `the-custodian/docs/forgejo-production-decisions.md` | +| Gap assessment | `history/2026-07-04-forgejo-backup-strategy-assessment.md` | + +**Follow-up:** `ResiliencePolicy.md` (operating principle, drill cadence) remains +in `DISCTL-WP-0002-T03`. Broader non-backup systems (Bubble, Stripe, DNS) will +be added when `DISCTL-WP-0002-T02` expands the registers. \ No newline at end of file diff --git a/TestEvidence/README.md b/TestEvidence/README.md new file mode 100644 index 0000000..6501286 --- /dev/null +++ b/TestEvidence/README.md @@ -0,0 +1,54 @@ +# Test Evidence + +Dated restore-drill reports proving RTO/RPO targets in `RTO-RPO-Matrix.md` are +actually met — not just documented. + +## Naming convention + +```text +TestEvidence/restore-drill-YYYY-MM-DD-.md +``` + +Examples: `restore-drill-2026-07-04-forgejo.md` + +## Required fields + +Every evidence file must include: + +| Field | Description | +| --- | --- | +| Date | Drill execution date (UTC) | +| System | What was restored (register ID or system name) | +| Backup artifact | Filename, size, timestamp, source (manual/automated) | +| Target environment | Isolated namespace, staging host, clean VM, etc. | +| Operator | Who ran the drill | +| Result | `success` / `partial` / `failed` | +| RPO achieved | Actual recovery point (how stale was the backup?) | +| RTO achieved | Wall-clock time from start to validated service | +| Checks performed | API calls, smoke tests, representative workflow | +| Issues found | Anything that blocked or degraded restore | +| Follow-up actions | Workplan items, pickup queue refs, register updates | +| Canonical evidence | Link to owner-repo evidence if drill ran elsewhere | + +## Drill cadence + +| Tier | Minimum cadence | Owner | +| --- | --- | --- | +| T1 irreplaceable source of truth | Quarterly | Bernd | +| T0 identity / trust | Quarterly (cold decrypt included) | Bernd | +| T2 runtime platforms | Semi-annual | Bernd | + +Quarterly Forgejo drill gate: re-run `forgejo-restore-drill.sh` from an +**automated** Nextcloud artifact once daily backups are wired (promotion gate +in `BackupPolicy.md` §7). + +## Next untested targets + +Priority order after Forgejo manual drill (2026-07-04): + +1. **S2 k8s runtime** — decrypt `railiance-backup-s2` artifact on clean target; + validate Helm values redeploy +2. **Dev workstation cold restore** — decrypt Nextcloud artifact without using + production key on same machine (password-manager key only) +3. **Gitea dump** — manual `gitea dump` restore to staging before next migration wave +4. **Secret vault cold decrypt** — SOPS/age recovery material (R5 mitigation) \ No newline at end of file diff --git a/TestEvidence/restore-drill-2026-07-04-forgejo.md b/TestEvidence/restore-drill-2026-07-04-forgejo.md new file mode 100644 index 0000000..cbc9970 --- /dev/null +++ b/TestEvidence/restore-drill-2026-07-04-forgejo.md @@ -0,0 +1,96 @@ +# Restore Drill: Forgejo production + +**Date:** 2026-07-04 +**System:** `backup-forgejo-production` (Forgejo git + packages + `forgejo-db`) +**Operator:** Bernd (via grok session, `railiance-infra`) +**Result:** success +**Workplan:** RAIL-HO-WP-0005-T09 + +`no_secret_material_recorded: true` + +## Summary + +Proved that a production `forgejo dump` can be restored into an isolated +namespace and serve repository metadata without touching production Forgejo or +Gitea. This validates **recovery mechanics** — not operational backup coverage. + +## Backup artifact + +| Field | Value | +| --- | --- | +| Method | `forgejo dump` from production pod (manual) | +| Production pod | `forgejo-gitea-64c5b57684-ph9vt` (namespace `forgejo`) | +| Archive path | `/tmp/forgejo-drill/forgejo-drill-backup.zip` (workstation) | +| Archive size | 12,284,847 bytes (~11.7 MiB) | +| Archive timestamp | 2026-07-04 11:20 +0200 | +| Contents | `repos/`, `data/`, `forgejo-db.sql`, `app.ini` | +| Repos in dump | `forgejo-actions-probe`, `glas-harness`, `key-cape` (`coulomb/`) | +| Automated / offsite | **No** — workstation `/tmp` only; not from Nextcloud | + +## Restore target + +| Field | Value | +| --- | --- | +| Namespace | `forgejo-restore-drill` | +| Database | CNPG `forgejo-db-restore` (isolated, 1 instance) | +| App data PVC | `forgejo-restore-data` (`local-path`, 10Gi) | +| Helm release | `forgejo-restore` (`gitea-charts/gitea` 12.5.0) | +| Orchestration | `railiance-infra/tools/forgejo-restore-drill.sh` | + +Restore steps (Forgejo 11.0.3 — no `forgejo restore` CLI): + +1. Unzip dump into import pod staging area. +2. Copy `repos/` → `/data/git/gitea-repositories/`. +3. Copy `data/` → `/data/` (packages, attachments, avatars). +4. Import `forgejo-db.sql` via `psql` into `forgejo-db-restore`. +5. Deploy isolated Helm release bound to restored PVC + restore DB host. + +## Validation checks + +Port-forward: `svc/forgejo-restore-gitea-http` → `127.0.0.1:13000` + +| Check | Result | +| --- | --- | +| `GET /` health | HTTP 200 | +| `GET /api/v1/repos/coulomb/glas-harness` | `full_name=coulomb/glas-harness`, `default_branch=main` | +| `GET /api/v1/repos/coulomb/key-cape` | `full_name=coulomb/key-cape`, `default_branch=main` | +| `GET /api/v1/orgs/coulomb/repos` | 3 repos visible | + +Script exit marker: `restore-drill-complete` + +## RPO / RTO measured + +| Metric | Target (adopted) | Achieved | +| --- | --- | --- | +| RPO | 24h | Point-in-time of manual dump — **no scheduled backup** | +| RTO | 4h | ~3–5 min (isolated namespace on railiance01) | +| Production impact | None | Read-only dump; separate namespace | + +Posture: RTO target met. RPO target **not operational** until daily automated +backups run (see `BackupPickupQueue.md` items 3–4). + +## Issues found + +- No scheduled backups; CNPG `Backup` CRs not configured. +- Dump stored on workstation only — no encryption-at-rest on approved offsite target. +- SQL import not idempotent — re-runs need `DRILL_CLEAN=1`. + +## Follow-up actions + +| Action | Owner | Queue / workplan | +| --- | --- | --- | +| Wire daily `forgejo dump` + Nextcloud upload | `railiance-platform` | Pickup #3 | +| Restore platform `railiance-backup` tool | `railiance-platform` | Pickup #4 | +| Re-drill from Nextcloud artifact | `railiance-infra` | Promotion gate (`BackupPolicy.md` §7) | +| 7 consecutive daily backup successes | `railiance-platform` | Promotion gate | + +## Canonical evidence + +Full drill record (owner repo): + +`railiance-infra/docs/forgejo-restore-drill-evidence.md` + +Supporting artifacts: + +- `railiance-infra/tools/forgejo-restore-drill.sh` +- `railiance-infra/infra/forgejo-restore-drill/` \ No newline at end of file diff --git a/workplans/DISCTL-WP-0003-railiance-backup-governance.md b/workplans/DISCTL-WP-0003-railiance-backup-governance.md index 17aa7c1..2deb242 100644 --- a/workplans/DISCTL-WP-0003-railiance-backup-governance.md +++ b/workplans/DISCTL-WP-0003-railiance-backup-governance.md @@ -100,7 +100,7 @@ Feeds `DISCTL-WP-0002-T04`. ```task id: DISCTL-WP-0003-T03 -status: progress +status: done priority: high state_hub_task_id: "cec8f15d-471d-4bbb-ad01-e25bc73c66e3" ``` @@ -126,7 +126,7 @@ Feeds `DISCTL-WP-0002-T03` (partial — ResiliencePolicy.md remains in WP-0002). ```task id: DISCTL-WP-0003-T04 -status: todo +status: done priority: medium state_hub_task_id: "484e433c-4840-41f7-8eda-fd6c0520dd1d" ``` @@ -153,7 +153,7 @@ Feeds `DISCTL-WP-0002-T08`. ```task id: DISCTL-WP-0003-T05 -status: todo +status: done priority: high state_hub_task_id: "e28afd71-7716-4623-b2ae-0e8675d6a60f" ```