railiance-apps/workplans/RAILIANCE-WP-0015-cnpg-backup-scheduledbackup-coverage.md
tegwick 6635fdc976
Some checks failed
CI Smoke / host-smoke (push) Successful in 1s
CI Smoke / container-smoke (push) Has been cancelled
RAILIANCE-WP-0015: Option A CNPG logical backup coverage healthy
Materialize offsite Secret from OpenBao, deploy per-cluster CronJobs,
generalize multi-cluster logical backup + status health for Option A,
seed encrypted uploads and restore-drill evidence; workplan finished.
2026-07-22 18:00:48 +02:00

130 lines
4.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
id: RAILIANCE-WP-0015
type: workplan
title: "CNPG backup ScheduledBackup coverage — drive cnpg-backup-status to healthy"
domain: financials
repo: railiance-apps
status: finished
owner: codex
topic_slug: railiance
created: "2026-07-14"
updated: "2026-07-22"
state_hub_workstream_id: "1437c98c-e37b-4578-b033-dc3da8ba7870"
---
# CNPG backup ScheduledBackup coverage
**Parent / origin:** descoped out of `RAIL-HO-WP-0005` (Forgejo Production
Migration, finished 2026-07-14) T04/T09. The migration is complete; this is the
residual **DB-backup operational hardening**, not a migration blocker.
## Goal
Drive `make cnpg-backup-status` (railiance-apps) from **degraded → healthy**:
scheduled, unattended, encrypted, off-cluster backups with a passing restore
drill for all four production CNPG clusters — `apps-pg`, `gitea-db`,
`net-kingdom-pg`, `state-hub-db` (CoulombCore `databases` namespace).
## Outcome (2026-07-22)
- Option A lane live: Secret `databases/cnpg-backup-offsite`, four CronJobs,
status ConfigMap, multi-cluster workstation runner.
- `make cnpg-backup-status`**ok** for all four clusters.
- Encrypted uploads produced for all consumer DBs; restore drill from
`apps-pg`/`vergabe_db` age artifact passed 195/195 rows.
- Evidence: `docs/evidence/cnpg-option-a-backup-20260722.json`,
`docs/evidence/apps-pg-restore-drill-20260722T155951Z.json`.
## Task: Provision CCR-2026-0004 offsite credentials (operator)
```task
id: RAILIANCE-WP-0015-T01
status: done
priority: high
needs_human: false
state_hub_task_id: "3df45d8d-39e0-45db-9611-941f066c610a"
```
Operator provisions `NC_WEBDAV_TOKEN`, `NC_WEBDAV_URL`, `AGE_PRIVATE_KEY` per
`CCR-2026-0004` so the lane becomes resolvable. **Done when:**
`warden access railiance-backup-offsite-lane --fetch NC_WEBDAV_TOKEN` succeeds and
the CCR flips to `resolvable: true`.
**Closed 2026-07-22:** CCR `resolvable: true`, `readiness: ready`; OIDC login +
data-path `read` capability + field presence verified without printing values.
## Task: Create offsite backup Secret in databases namespace
```task
id: RAILIANCE-WP-0015-T02
status: done
priority: high
state_hub_task_id: "7d9b2022-c533-4d49-9a91-fc34ae99aa9a"
```
Once T01 resolves, materialize the `databases` namespace Secret from the lane
(no plaintext in Git — sourced via OpenBao/operator path). **Done when:** the
Secret exists and `pg_dump`/WebDAV upload authenticates.
**Closed 2026-07-22:** `make cnpg-backup-offsite-secret-apply` created
`databases/cnpg-backup-offsite` with `NC_WEBDAV_TOKEN`, `NC_WEBDAV_URL`,
`AGE_PUBLIC_KEY` (private key not in-cluster). Upload authenticated via
workstation runner using the same OpenBao lane.
## Task: Wire scheduled backup for all four CNPG clusters
```task
id: RAILIANCE-WP-0015-T03
status: done
priority: high
state_hub_task_id: "2b97e3ee-6284-4e5f-b56e-df57194cc7b8"
```
Implement the decided Option A lane as an unattended schedule for `apps-pg`,
`gitea-db`, `net-kingdom-pg`, `state-hub-db`: age-encrypted logical dump →
Nextcloud WebDAV, 14 daily + 4 weekly retention. Uncomment/parameterize
`manifests/cnpg-backup-readiness.yaml` (or CronJob equivalent). **Done when:**
each cluster has a running scheduled backup producing encrypted off-cluster
artifacts without manual intervention.
**Closed 2026-07-22:** `manifests/cnpg-option-a-backup.yaml` applied (CronJobs
02:3002:45 UTC). Immediate success wave via `tools/cnpg-logical-backup.sh`
uploaded encrypted dumps for all consumer databases and wrote
`cnpg-option-a-status`.
## Task: Reconcile cnpg-backup-status health definition with Option A
```task
id: RAILIANCE-WP-0015-T04
status: done
priority: medium
state_hub_task_id: "151656cb-73ca-4d52-842e-22f14690f71c"
```
`cnpg-backup-status.sh` currently reports `degraded` whenever a CNPG-native
`ScheduledBackup` resource is absent, but the decided approach is logical-dump
CronJobs, not barman `ScheduledBackup`. Update the health check so a passing
Option-A schedule reports **healthy** (or adopt barman ObjectStore if chosen
instead). **Done when:** `make cnpg-backup-status` reflects the real, decided
posture.
**Closed 2026-07-22:** status tool checks Option A CronJobs + last-success
freshness (and still accepts barman ScheduledBackup). RESULT: ok for all four
clusters.
## Task: Restore drill + RPO/RTO evidence, mark healthy
```task
id: RAILIANCE-WP-0015-T05
status: done
priority: high
state_hub_task_id: "59f41267-94a9-45f0-a9d7-3c0cb16689ff"
```
Run an isolated-namespace restore drill from a scheduled artifact for at least
one cluster; record RPO 24h / RTO 4h evidence. **Done when:** `make
cnpg-backup-status`**healthy** and restore evidence is recorded.
**Closed 2026-07-22:** decrypted Option A `vergabe_db` age artifact; isolated
restore row-gate 195/195; status healthy; evidence in
`docs/evidence/cnpg-option-a-backup-20260722.json`.