RAILIANCE-WP-0016: promote active; inventory and activity-core cutover prep
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 2s

Mark workplan active with T01/T02 done, document topology, extend status
for activity-core mode, and wire Make targets to the platform multi-host
backup CLI. T03 remains operator-blocked on ESO token.
This commit is contained in:
tegwick 2026-07-22 19:50:59 +02:00
parent 04626b6335
commit c202fbf7be
6 changed files with 180 additions and 309 deletions

View file

@ -4,7 +4,7 @@ type: workplan
title: "railiance01 + activity-core unattended CNPG/Forgejo backup automation"
domain: financials
repo: railiance-apps
status: proposed
status: active
owner: codex
topic_slug: railiance
created: "2026-07-22"
@ -19,248 +19,138 @@ state_hub_workstream_id: "0c8ddbc9-3740-4e0a-acb2-901e050ab242"
depends on a **workstation** (crontab + interactive OpenBao OIDC + laptop
uptime). That is not an acceptable durability control plane.
**Goal:** Daily (and weekly retention) encrypted offsite backups for production
CNPG/Forgejo data run **without any workstation** — scheduled and observed by
**activity-core on railiance01**, following the proven
`weekly-forgejo-package-prune` pattern.
**Goal:** Daily encrypted offsite backups for production CNPG data run **without
any workstation** — scheduled and observed by **activity-core on railiance01**.
## Context and constraints
## Implementation status (2026-07-22)
### Topology (2026-07-22)
| Host | Clusters / data |
| Area | State |
| --- | --- |
| CoulombCore | `apps-pg` (vergage, core_hub, …), `gitea-db`, `net-kingdom-pg`, `state-hub-db` |
| railiance01 | `forgejo-db`, `net-kingdom-pg`, `state-hub-db` (+ activity-core, Forgejo) |
| Topology inventory | **done**`docs/cnpg-backup-topology-inventory.md` |
| Platform CLI + vendor age | **done**`railiance-platform/tools/cmd/cnpg-option-a-backup` |
| Dry-run forgejo-db | **pass** (JSON overall=ok) |
| activity-core resolver + definition | **landed** (`enabled: false` until creds+kube) |
| ESO ExternalSecret | **applied**; **SecretSyncedError** until ESO token re-mint |
| Worker RBAC databases exec | **applied** on railiance01 |
| Worker kubeconfig hostPath | **in git** (`20-runtime.yaml`); apply/restart pending |
| Schedule enable + soak | **blocked** on operator T03/T04 |
NK and state-hub exist on **both** hosts — WP must inventory which instance is
production-of-record per consumer before scheduling dumps.
### Operator unblock (next human steps)
### Why workstation fails as control plane
```bash
# 1) Re-mint ESO token with backup lane policy (operator OpenBao token)
cd ~/activity-core
OPENBAO_ACTIVITY_CORE_POLICIES="workload-kv-read-issue-core-runtime workload-kv-read-forgejo-admin workload-kv-read-railiance-backup-offsite-lane" \
scripts/openbao-eso-token-apply.sh
kubectl -n activity-core get externalsecret actcore-backup-offsite # expect SecretSynced
- Laptop sleep / travel / network breaks RPO 24h.
- OIDC `bao login` is interactive; agents and workers need non-interactive
OpenBao or ESO.
- WP-0015 in-cluster CronJobs are **suspended**: cluster pods cannot reach
GitHub or Alpine CDN to obtain `age`.
# 2) On railiance01 host as tegwick: vendor kubectl + ensure kubeconfigs
cd ~/railiance-platform && tools/cmd/install-cnpg-backup-vendor-tools
# Place CoulombCore kubeconfig at ~/.kube/config (R01 already config-hosteurope)
### Pattern to copy (ACTIVITY-WP-0020)
# 3) Roll worker with updated mounts + new image/code
kubectl apply -f k8s/railiance/20-runtime.yaml # or patch
# rebuild activity-core image with resolver if not live-mounted
```text
ActivityDefinition (cron)
→ shell context_source on activity-core worker
→ hostPath /opt/railiance-platform (or apps) CLI
→ credentials via actcore-runtime-secret / ESO
→ State Hub progress evidence sink
→ make automation-status / inventory can see the schedule
# 4) Sync definition, enable, force run
# set enabled: true on daily-cnpg-option-a-backup; make sync-activity-definitions sync-schedules
```
Reference: `activity-core/activity-definitions/weekly-forgejo-package-prune.md`.
### Target architecture
```text
activity-core Temporal schedule (railiance01)
shell resolver: cnpg_option_a_backup / forgejo_backup
hostPath tools (static age + kubectl + curl OR prebuilt image side-car script)
+ kubeconfig: railiance01 (local) + CoulombCore (remote cert/tunnel)
+ OpenBao/ESO: NC_WEBDAV_TOKEN, NC_WEBDAV_URL, AGE_PUBLIC_KEY
├─► pg_dump / forgejo dump → age → Nextcloud WebDAV
└─► patch ConfigMap cnpg-option-a-status (+ optional State Hub event)
```
**Acceptance for the workplan as a whole:** RPO 24h holds when the operator
workstation is powered off; `make cnpg-backup-status` remains healthy; schedule
appears in activity-core automation inventory; restore drill from an
activity-core-produced artifact passes.
### Repo ownership (do not invent ownership)
| Concern | Owner repo |
| --- | --- |
| S5 handoff, status Make targets, app evidence, this workplan | **railiance-apps** |
| CNPG/Forgejo dump CLI, hostPath packaging, OpenBao lane policy for workers | **railiance-platform** |
| ActivityDefinition, shell resolver, Temporal schedule, worker mounts/secrets | **activity-core** |
| SSH certs for remote kube if needed | **ops-warden** (sign only) |
| Tunnels | **ops-bridge** |
This workplan **coordinates** cross-repo work; implement in the owning repo and
cite commits/evidence here. Do **not** hand-register hub tasks — write files and
`statehub fix-consistency`.
---
## Task: Inventory production-of-record DB topology
```task
id: RAILIANCE-WP-0016-T01
status: todo
status: done
priority: high
state_hub_task_id: "acf23292-438b-44c0-a6cd-87327b54d743"
```
Document which CNPG instances are production-of-record for each consumer DB
across CoulombCore vs railiance01 (especially `net-kingdom-pg` and
`state-hub-db` dual hosting). Produce a table in `docs/` or STATE.md: cluster,
host, databases, consumer, backup required Y/N.
**Done when:** signed-off inventory exists; no ambiguous double-backup or missed
instance remains.
**Done 2026-07-22:** `docs/cnpg-backup-topology-inventory.md` — POR table;
dual-host NK/state-hub both backed up until DSN audit.
## Task: Package offline-capable backup runner for railiance01
```task
id: RAILIANCE-WP-0016-T02
status: todo
status: done
priority: high
state_hub_task_id: "35e81dc9-4f86-4a8f-8857-faef4117eeb6"
```
Deliver a runner that does **not** require workstation or internet package
install at job time:
- Static `age` (and friends) vendored under hostPath tools tree and/or a
private OCI image on Forgejo registry with `kubectl`+`age`+`curl`+`bash`.
- CLI entrypoints covering: multi-cluster logical dump (Option A), forgejo dump
path (platform), status CM update, non-secret JSON evidence.
- Prefer extending `railiance-platform` lib/tools so activity-core only shells out
(mirror `forgejo-package-prune`).
**Done when:** on railiance01 (or a dry-run host with same mounts), one command
produces age-encrypted artifacts **without** `apk`/`github.com` at runtime.
**Done 2026-07-22:** platform CLI + static `age` under `tools/vendor/bin`;
`install-cnpg-backup-vendor-tools` for kubectl; dry-run JSON ok for
`r01-forgejo-db`. Docs: `railiance-platform/docs/cnpg-option-a-backup.md`.
## Task: Non-interactive offsite credentials for activity-core worker
```task
id: RAILIANCE-WP-0016-T03
status: todo
status: wait
priority: high
needs_human: true
state_hub_task_id: "3bcb6a70-7b74-47a3-8afe-3921fdccd893"
```
Wire `NC_WEBDAV_TOKEN` / `NC_WEBDAV_URL` / `AGE_PUBLIC_KEY` into the activity-core
worker on railiance01 without interactive OIDC:
- Prefer ExternalSecret → `actcore-runtime-secret` (pattern: forgejo-admin ESO).
- OpenBao policy for the worker role; **do not** put `AGE_PRIVATE_KEY` on the
worker (recovery escrow stays operator/OpenBao only).
- Document rotation via CCR / warden route; no secrets in Git.
**Done when:** worker pod can upload a tiny probe object to Nextcloud and env
lengths are non-zero without operator login.
ESO manifest `actcore-backup-offsite` applied; status **SecretSyncedError**
(provider deny) until ESO token includes
`workload-kv-read-railiance-backup-offsite-lane`. Script defaults + combined
policy updated. **Blocked on operator OpenBao token mint.**
## Task: CoulombCore API reachability from railiance01 worker
```task
id: RAILIANCE-WP-0016-T04
status: todo
status: progress
priority: high
state_hub_task_id: "b8d2dd05-31a0-4fd7-bdef-27eb686e8f8a"
```
Backup of CoulombCore clusters (`apps-pg`, …) from railiance01 requires a
stable kubeconfig + auth (ops-warden cert + ops-bridge tunnel, or in-cluster
equivalent). Options to evaluate (pick one, record decision):
1. Worker host kubeconfig for CoulombCore with long-lived cert refresh hook.
2. activity-core (or a thin agent) also on CoulombCore for local dumps only.
3. CNPG-native path later (out of scope unless re-decided).
**Done when:** non-interactive `kubectl get cluster -n databases` against
CoulombCore succeeds from the same environment that will run the schedule.
Decision: hostPath `~/.kube``/kube` on worker; `KUBECONFIG_CORE=/kube/config`,
`KUBECONFIG_R01=/kube/config-hosteurope`. Git runtime yaml updated. Remaining:
ensure CoulombCore kubeconfig + certs on railiance01 host; rollout worker;
verify `kubectl get cluster` from inside worker for both hosts.
## Task: activity-core ActivityDefinition + shell resolver
```task
id: RAILIANCE-WP-0016-T05
status: todo
status: progress
priority: high
state_hub_task_id: "547b3acb-94c2-46e0-a7da-ef51ee0c3c8d"
```
In **activity-core**:
1. Shell resolver e.g. `cnpg_option_a_backup` (and optionally `forgejo_backup`)
analogous to `forgejo_package_prune`.
2. ActivityDefinition markdown with `trigger.type: cron` (target ~02:30 UTC
daily; weekly retention on Sundays).
3. `make sync` / schedule reconcile on railiance01; enable only after T02T04.
4. Evidence sink → State Hub progress (non-secret summary).
5. Visible in `make automation-status` / inventory.
**Done when:** Temporal schedule exists, one forced run succeeds, inventory lists
the automation as enabled.
Landed: `cnpg_option_a_backup` resolver, evidence sink, ops console allowlist,
`activity-definitions/daily-cnpg-option-a-backup.md` (**enabled: false**), unit
tests (3 passed). Remaining: image/deploy + enable after T03/T04.
## Task: Cut over health definition and retire workstation dependency
```task
id: RAILIANCE-WP-0016-T06
status: todo
status: progress
priority: medium
state_hub_task_id: "d6cab5a4-3cdd-4033-b21e-88db02197472"
```
In **railiance-apps**:
- Set `cnpg-option-a-schedule` `mode` to `activity-core` (or dual-run during
soak).
- `cnpg-backup-status` accepts activity-core last-success (status CM and/or
hub evidence age).
- Document that workstation cron is **stopgap only**; remove from operator
“required” path after soak.
- Optionally un-suspend in-cluster CronJobs **only if** T02 image path lands
in-cluster; otherwise leave suspended.
**Done when:** docs + status tool describe activity-core as primary; workstation
optional.
`cnpg-backup-status` accepts `mode=activity-core`. Full cutover after first
successful scheduled runs (T07). Workstation remains documented stopgap.
## Task: Soak, restore drill from automated artifact, finish
```task
id: RAILIANCE-WP-0016-T07
status: todo
status: wait
priority: high
state_hub_task_id: "f331bdcd-0e76-4bc0-b85a-1bb1227a876b"
```
- ≥3 consecutive successful scheduled runs (or 7 if aligning forgejo promotion
gate language).
- Restore drill from an artifact produced by activity-core (not workstation
one-off), row-count or equivalent gate, evidence JSON under `docs/evidence/`.
- Update `STATE.md`; archive/finish this workplan; optional message to
activity-core / platform owners if residual tasks remain in their repos.
**Done when:** RPO/RTO evidence recorded; workstation powered-off scenario is
credible; workplan `finished`.
Blocked on T03T05 enable path.
---
## Out of scope
- Switching Option A → barman ObjectStore / CNPG `ScheduledBackup` (Phase 2;
only if explicitly re-decided).
- Media/PVC app data backups (vergage media still disabled).
- Rotating EXPOSED taint history on the offsite lane (platform CCR ops), except
as a dependency note if worker fetch is blocked.
## Suggested implementation order
```text
T01 inventory → T02 runner package → T03 creds + T04 kube reachability (parallel)
→ T05 activity-core definition → T06 cutover docs/status → T07 soak + drill
```
## References
- `STATE.md` (this repo) — current posture
- `workplans/RAILIANCE-WP-0015-…` — interim Option A
- `docs/evidence/cnpg-option-a-backup-20260722.json`
- `activity-core/activity-definitions/weekly-forgejo-package-prune.md`
- `railiance-platform/docs/forgejo-backup.md`
- `activity-core/docs/runbook.md` (hostPath + ESO patterns)
- `STATE.md`
- `docs/cnpg-backup-topology-inventory.md`
- `railiance-platform/tools/cmd/cnpg-option-a-backup`
- `activity-core/activity-definitions/daily-cnpg-option-a-backup.md`