railiance-apps/workplans/RAPPS-WP-0001-cnpg-backup-scheduledbackup-coverage.md
codex dada84cf51
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 1s
fix(workplans): migrate active workplans off the retired RAILIANCE-WP prefix
RAILIANCE-WP is a family name, not a repository (ADR-007, and the prefix
registry already lists it retired). Three repositories independently used one
number space for unrelated work — RAILIANCE-WP-0012 was openbao extraction here,
a cnpg backup in railiance-apps and a deploy-verify in railiance-cluster. This
repository also carried two files both numbered 0016.

Active workplans move to the successor prefix and are renumbered from 0001 in
historical order. Archived workplans keep their historical identifiers.

Projection UUIDs are re-derived from the new canonical ids. Records already
registered under the old identifiers leave orphaned hub rows behind; that debt
is recorded in CUST-WP-0068 and clears when ADR-012's reset-from-forge lands.

Refs CUST-WP-0068-T03

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 2583210@bnt-lap001
Assistant-Session: f2bff2d5-e9b2-4338-92ca-10282a927006
2026-08-25 22:58:36 +02:00

4.7 KiB
Raw Permalink Blame History

id type title domain repo status owner topic_slug created updated state_hub_workstream_id
RAPPS-WP-0001 workplan CNPG backup ScheduledBackup coverage — drive cnpg-backup-status to healthy financials railiance-apps finished codex railiance 2026-07-14 2026-07-22 7d4701c2-4c1c-517e-bf74-7dd34a6ab420

CNPG backup ScheduledBackup coverage

Parent / origin: descoped out of RAIL-HO-WP-0005 (Forgejo Production Migration, finished 2026-07-14) T04/T09. The migration is complete; this is the residual DB-backup operational hardening, not a migration blocker.

Goal

Drive make cnpg-backup-status (railiance-apps) from degraded → healthy: scheduled, unattended, encrypted, off-cluster backups with a passing restore drill for all four production CNPG clusters — apps-pg, gitea-db, net-kingdom-pg, state-hub-db (CoulombCore databases namespace).

Outcome (2026-07-22)

  • Option A lane live: Secret databases/cnpg-backup-offsite, four CronJobs, status ConfigMap, multi-cluster workstation runner.
  • make cnpg-backup-statusok for all four clusters.
  • Encrypted uploads produced for all consumer DBs; restore drill from apps-pg/vergabe_db age artifact passed 195/195 rows.
  • Evidence: docs/evidence/cnpg-option-a-backup-20260722.json, docs/evidence/apps-pg-restore-drill-20260722T155951Z.json.

Task: Provision CCR-2026-0004 offsite credentials (operator)

id: RAPPS-WP-0001-T01
status: done
priority: high
needs_human: false
state_hub_task_id: "a21d47fe-e2dc-5037-a698-82d59fe2b18a"

Operator provisions NC_WEBDAV_TOKEN, NC_WEBDAV_URL, AGE_PRIVATE_KEY per CCR-2026-0004 so the lane becomes resolvable. Done when: warden access railiance-backup-offsite-lane --fetch NC_WEBDAV_TOKEN succeeds and the CCR flips to resolvable: true.

Closed 2026-07-22: CCR resolvable: true, readiness: ready; OIDC login + data-path read capability + field presence verified without printing values.

Task: Create offsite backup Secret in databases namespace

id: RAPPS-WP-0001-T02
status: done
priority: high
state_hub_task_id: "498309e0-c29d-56e3-a4af-ab9723e057c6"

Once T01 resolves, materialize the databases namespace Secret from the lane (no plaintext in Git — sourced via OpenBao/operator path). Done when: the Secret exists and pg_dump/WebDAV upload authenticates.

Closed 2026-07-22: make cnpg-backup-offsite-secret-apply created databases/cnpg-backup-offsite with NC_WEBDAV_TOKEN, NC_WEBDAV_URL, AGE_PUBLIC_KEY (private key not in-cluster). Upload authenticated via workstation runner using the same OpenBao lane.

Task: Wire scheduled backup for all four CNPG clusters

id: RAPPS-WP-0001-T03
status: done
priority: high
state_hub_task_id: "dfccabf9-93ab-5893-b647-68dacc50c129"

Implement the decided Option A lane as an unattended schedule for apps-pg, gitea-db, net-kingdom-pg, state-hub-db: age-encrypted logical dump → Nextcloud WebDAV, 14 daily + 4 weekly retention. Uncomment/parameterize manifests/cnpg-backup-readiness.yaml (or CronJob equivalent). Done when: each cluster has a running scheduled backup producing encrypted off-cluster artifacts without manual intervention.

Closed 2026-07-22: manifests/cnpg-option-a-backup.yaml applied (CronJobs 02:3002:45 UTC). Immediate success wave via tools/cnpg-logical-backup.sh uploaded encrypted dumps for all consumer databases and wrote cnpg-option-a-status.

Task: Reconcile cnpg-backup-status health definition with Option A

id: RAPPS-WP-0001-T04
status: done
priority: medium
state_hub_task_id: "28496364-2784-5d94-bb3b-bb5f9cf6d747"

cnpg-backup-status.sh currently reports degraded whenever a CNPG-native ScheduledBackup resource is absent, but the decided approach is logical-dump CronJobs, not barman ScheduledBackup. Update the health check so a passing Option-A schedule reports healthy (or adopt barman ObjectStore if chosen instead). Done when: make cnpg-backup-status reflects the real, decided posture.

Closed 2026-07-22: status tool checks Option A CronJobs + last-success freshness (and still accepts barman ScheduledBackup). RESULT: ok for all four clusters.

Task: Restore drill + RPO/RTO evidence, mark healthy

id: RAPPS-WP-0001-T05
status: done
priority: high
state_hub_task_id: "11a05a83-e7ba-5d9e-a463-0f34a8f48017"

Run an isolated-namespace restore drill from a scheduled artifact for at least one cluster; record RPO 24h / RTO 4h evidence. Done when: make cnpg-backup-statushealthy and restore evidence is recorded.

Closed 2026-07-22: decrypted Option A vergabe_db age artifact; isolated restore row-gate 195/195; status healthy; evidence in docs/evidence/cnpg-option-a-backup-20260722.json.