railiance-apps/workplans/RAILIANCE-WP-0015-cnpg-backup-scheduledbackup-coverage.md
codex b5151225ba
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 1s
fix(workplans): adopt ADR-007 derived identifiers for unregistered records
These workplans exist only in the retired local hub. Their random pre-ADR-007
identifiers are refused by C-06 as stale references, so they cannot be
registered. Deriving from the canonical record id takes no identity from
anything: central does not hold them and the old ids die with the cache.

Records central already holds were deliberately left untouched.

Refs CUST-WP-0068-T06

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 2583210@bnt-lap001
Assistant-Session: f2bff2d5-e9b2-4338-92ca-10282a927006
2026-08-25 20:18:05 +02:00

4.8 KiB
Raw Blame History

id type title domain repo status owner topic_slug created updated state_hub_workstream_id
RAILIANCE-WP-0015 workplan CNPG backup ScheduledBackup coverage — drive cnpg-backup-status to healthy financials railiance-apps finished codex railiance 2026-07-14 2026-07-22 1c73af72-fd4b-5ab6-8ff7-a3c302bf55a5

CNPG backup ScheduledBackup coverage

Parent / origin: descoped out of RAIL-HO-WP-0005 (Forgejo Production Migration, finished 2026-07-14) T04/T09. The migration is complete; this is the residual DB-backup operational hardening, not a migration blocker.

Goal

Drive make cnpg-backup-status (railiance-apps) from degraded → healthy: scheduled, unattended, encrypted, off-cluster backups with a passing restore drill for all four production CNPG clusters — apps-pg, gitea-db, net-kingdom-pg, state-hub-db (CoulombCore databases namespace).

Outcome (2026-07-22)

  • Option A lane live: Secret databases/cnpg-backup-offsite, four CronJobs, status ConfigMap, multi-cluster workstation runner.
  • make cnpg-backup-statusok for all four clusters.
  • Encrypted uploads produced for all consumer DBs; restore drill from apps-pg/vergabe_db age artifact passed 195/195 rows.
  • Evidence: docs/evidence/cnpg-option-a-backup-20260722.json, docs/evidence/apps-pg-restore-drill-20260722T155951Z.json.

Task: Provision CCR-2026-0004 offsite credentials (operator)

id: RAILIANCE-WP-0015-T01
status: done
priority: high
needs_human: false
state_hub_task_id: "37117b07-196d-5095-ba6b-f2123db4bf3b"

Operator provisions NC_WEBDAV_TOKEN, NC_WEBDAV_URL, AGE_PRIVATE_KEY per CCR-2026-0004 so the lane becomes resolvable. Done when: warden access railiance-backup-offsite-lane --fetch NC_WEBDAV_TOKEN succeeds and the CCR flips to resolvable: true.

Closed 2026-07-22: CCR resolvable: true, readiness: ready; OIDC login + data-path read capability + field presence verified without printing values.

Task: Create offsite backup Secret in databases namespace

id: RAILIANCE-WP-0015-T02
status: done
priority: high
state_hub_task_id: "a71c1bdd-4bc7-5547-bc00-0cdc7d459a4b"

Once T01 resolves, materialize the databases namespace Secret from the lane (no plaintext in Git — sourced via OpenBao/operator path). Done when: the Secret exists and pg_dump/WebDAV upload authenticates.

Closed 2026-07-22: make cnpg-backup-offsite-secret-apply created databases/cnpg-backup-offsite with NC_WEBDAV_TOKEN, NC_WEBDAV_URL, AGE_PUBLIC_KEY (private key not in-cluster). Upload authenticated via workstation runner using the same OpenBao lane.

Task: Wire scheduled backup for all four CNPG clusters

id: RAILIANCE-WP-0015-T03
status: done
priority: high
state_hub_task_id: "b5cc6fe2-a9c2-5f27-bcb7-a2e22e96b04b"

Implement the decided Option A lane as an unattended schedule for apps-pg, gitea-db, net-kingdom-pg, state-hub-db: age-encrypted logical dump → Nextcloud WebDAV, 14 daily + 4 weekly retention. Uncomment/parameterize manifests/cnpg-backup-readiness.yaml (or CronJob equivalent). Done when: each cluster has a running scheduled backup producing encrypted off-cluster artifacts without manual intervention.

Closed 2026-07-22: manifests/cnpg-option-a-backup.yaml applied (CronJobs 02:3002:45 UTC). Immediate success wave via tools/cnpg-logical-backup.sh uploaded encrypted dumps for all consumer databases and wrote cnpg-option-a-status.

Task: Reconcile cnpg-backup-status health definition with Option A

id: RAILIANCE-WP-0015-T04
status: done
priority: medium
state_hub_task_id: "3bfe6fba-6e7a-52f1-8b7f-906ba1e065cd"

cnpg-backup-status.sh currently reports degraded whenever a CNPG-native ScheduledBackup resource is absent, but the decided approach is logical-dump CronJobs, not barman ScheduledBackup. Update the health check so a passing Option-A schedule reports healthy (or adopt barman ObjectStore if chosen instead). Done when: make cnpg-backup-status reflects the real, decided posture.

Closed 2026-07-22: status tool checks Option A CronJobs + last-success freshness (and still accepts barman ScheduledBackup). RESULT: ok for all four clusters.

Task: Restore drill + RPO/RTO evidence, mark healthy

id: RAILIANCE-WP-0015-T05
status: done
priority: high
state_hub_task_id: "d4223b6e-4e46-5c4c-9d1d-94812b2e8bf6"

Run an isolated-namespace restore drill from a scheduled artifact for at least one cluster; record RPO 24h / RTO 4h evidence. Done when: make cnpg-backup-statushealthy and restore evidence is recorded.

Closed 2026-07-22: decrypted Option A vergabe_db age artifact; isolated restore row-gate 195/195; status healthy; evidence in docs/evidence/cnpg-option-a-backup-20260722.json.