docs: STATE.md assessment; propose RAILIANCE-WP-0016 activity-core backups
All checks were successful
CI Smoke / host-smoke (push) Successful in 1s
CI Smoke / container-smoke (push) Successful in 3s

Capture post-WP-0015 posture (workstation SPOF) and register a workplan to
move unattended Option A + Forgejo backup automation to railiance01 via
activity-core without a laptop control plane.
This commit is contained in:
tegwick 2026-07-22 19:35:25 +02:00
parent 1376b4f34c
commit 7bc76dc0c6
2 changed files with 425 additions and 0 deletions

167
STATE.md Normal file
View file

@ -0,0 +1,167 @@
# STATE — railiance-apps
**Updated:** 2026-07-22
**Domain:** financials · **Repo:** railiance-apps
**Topic ID:** `ca369340-a64e-442e-98f1-a4fa7dc74a38`
This file is the human/agent snapshot of **where the repo stands**. Prefer it over
chat history for orientation. Generated custodian brief is `.custodian-brief.md`
(do not edit). Formal work structure lives in `workplans/`.
---
## One-line posture
S5 app releases (vergabe, hubs, Forgejo consumer surface) are operational; **CNPG
Option A logical backups exist and currently report healthy**, but unattended
execution still depends on a **workstation path** and must move to **railiance01 +
activity-core** so the laptop is not a durability SPOF.
---
## Clusters (as of 2026-07-22)
| Host | Kubeconfig | CNPG clusters in `databases` | Role |
| --- | --- | --- | --- |
| **CoulombCore** | `~/.kube/config` | `apps-pg`, `gitea-db`, `net-kingdom-pg`, `state-hub-db` | S5 apps DBs, legacy gitea-db, NK + state-hub instances |
| **railiance01** | `~/.kube/config-hosteurope` | `forgejo-db`, `net-kingdom-pg`, `state-hub-db` | Forgejo DB; NK + state-hub instances (may differ from Core) |
`PRODUCTION_KUBECONFIG` defaults to CoulombCore. Forgejo tooling uses railiance01.
---
## Backup durability (RAILIANCE-WP-0015 — **finished**)
### Goal met
Drive `make cnpg-backup-status` **degraded → healthy** with Option A (age-encrypted
logical dump → Nextcloud WebDAV), not barman `ScheduledBackup`.
### What is live today
| Item | State |
| --- | --- |
| OpenBao lane `CCR-2026-0004` / `railiance-backup-offsite-lane` | **resolvable**, fields present |
| Secret `databases/cnpg-backup-offsite` (CoulombCore) | **present** (`NC_WEBDAV_*`, `AGE_PUBLIC_KEY` only) |
| ConfigMap `cnpg-option-a-schedule` | **mode=workstation-cron** @ `30 2 * * *` |
| ConfigMap `cnpg-option-a-status` | last-success timestamps for all four Core clusters (2026-07-22) |
| CronJobs `cnpg-logical-backup-*` | **suspended** (cluster egress cannot install `age`) |
| Workstation runner `tools/cnpg-logical-backup.sh` | **works** (upload + status CM) |
| Restore drill | apps-pg/`vergage_db` age artifact **195/195** |
| Evidence | `docs/evidence/cnpg-option-a-backup-20260722.json` |
### What is *not* good enough
1. **Workstation SPOF** — daily unattended path is “someones machine has crontab +
bao OIDC + kubeconfig”. Laptop sleep/off breaks RPO.
2. **In-cluster CronJobs blocked** — pods cannot reach GitHub or Alpine CDN to get
`age`; no prebuilt image with `kubectl`+`age`+`curl` yet.
3. **Split brain topology** — Option A runner today targets **CoulombCore** four
clusters only; **forgejo-db** (railiance01) still uses platform `make forgejo-backup`
(also historically workstation-cron).
4. **No activity-core schedule** — unlike weekly Forgejo package prune, backups are
not Temporal/`ActivityDefinition` managed, so automation inventory and
`automation-status` cannot own them.
5. **Credential delivery** — offsite lane is human OIDC on workstation; worker needs
non-interactive OpenBao/ESO (or sealed secret) path.
### Operator residual (manual)
```cron
# Not yet proven installed on a always-on host — install on railiance01 or leave as stopgap
30 2 * * * cd $HOME/railiance-apps && make cnpg-logical-backup >>$HOME/.cache/railiance/backups/cnpg/cron.log 2>&1
```
---
## Adjacent finished work
| Workplan | Outcome |
| --- | --- |
| RAILIANCE-WP-0013 | Phase 1 apps-pg/`vergage_db` backup + restore drill (interim S5 gate) |
| RAILIANCE-WP-0015 | Fleet Option A tooling + healthy status on CoulombCore (workstation-backed) |
| RAIL-HO-WP-0005 | Forgejo prod migration; descoped residual backup → WP-0015 |
| ACTIVITY-WP-0020 | Weekly forgejo package prune via **activity-core** on railiance01 (pattern to copy) |
---
## Next major strand
**Move unattended Option A (+ forgejo-backup) to railiance01, scheduled by
activity-core — no workstation required.**
Tracking workplan: **`RAILIANCE-WP-0016`**
`workplans/RAILIANCE-WP-0016-railiance01-activity-core-backup-automation.md`
Design sketch (see workplan for tasks):
```text
┌─────────────────────────────┐
│ activity-core (railiance01) │
│ Temporal cron ActivityDef │
└─────────────┬───────────────┘
│ shell context
┌─────────────────────────────┐
│ hostPath /opt/railiance-* │
│ backup CLI (age+kubectl) │
│ OpenBao/ESO offsite secrets │
└─────────────┬───────────────┘
┌──────────────────┼──────────────────┐
▼ ▼ ▼
CoulombCore API railiance01 API Nextcloud WebDAV
(apps-pg, …) (forgejo-db, …) (age artifacts)
```
**Proven pattern:** `weekly-forgejo-package-prune` ActivityDefinition +
`/opt/railiance-platform` hostPath + ESO credentials + State Hub evidence sink.
---
## Active / proposed workplans
| ID | Status | Title |
| --- | --- | --- |
| RAILIANCE-WP-0015 | **finished** | CNPG Option A coverage → healthy (workstation interim) |
| RAILIANCE-WP-0016 | **proposed** | railiance01 + activity-core unattended backup automation |
(No other active workplans in hub as of last brief.)
---
## Quick operator commands
```bash
# Health (CoulombCore)
make cnpg-backup-status
# Manual Option A run (still workstation / OpenBao OIDC today)
bao login -method=oidc -path=netkingdom role=railiance-backup-workload-kv-read
make cnpg-logical-backup
# Secret re-materialize
make cnpg-backup-offsite-secret-apply
```
---
## Risks / open questions
| Risk | Mitigation direction (WP-0016) |
| --- | --- |
| Dual NK/state-hub DB instances on two hosts | Inventory which is production-of-record per consumer |
| CoulombCore access from railiance01 | Dedicated kubeconfig + ops-warden cert / tunnel for worker |
| Age binary in restricted network | Vendor static binary in image or hostPath tools tree |
| High-risk WebDAV + age keys in worker | ESO from OpenBao; never private key in Git; agent boundary |
| RPO if activity-core worker down | Temporal schedule health + status CM freshness alerts |
---
## Related docs
- `docs/app-data-backup-restore-handoff.md`
- `docs/credential-routing-railiance-apps.md`
- `manifests/cnpg-option-a-backup.yaml`
- `docs/evidence/cnpg-option-a-backup-20260722.json`
- activity-core: `activity-definitions/weekly-forgejo-package-prune.md`, `docs/runbook.md`
- platform: `docs/forgejo-backup.md`

View file

@ -0,0 +1,258 @@
---
id: RAILIANCE-WP-0016
type: workplan
title: "railiance01 + activity-core unattended CNPG/Forgejo backup automation"
domain: financials
repo: railiance-apps
status: proposed
owner: codex
topic_slug: railiance
created: "2026-07-22"
updated: "2026-07-22"
---
# railiance01 + activity-core unattended backup automation
**Predecessor:** `RAILIANCE-WP-0015` (finished) delivered Option A tooling and a
**healthy** `cnpg-backup-status` on CoulombCore, but unattended execution still
depends on a **workstation** (crontab + interactive OpenBao OIDC + laptop
uptime). That is not an acceptable durability control plane.
**Goal:** Daily (and weekly retention) encrypted offsite backups for production
CNPG/Forgejo data run **without any workstation** — scheduled and observed by
**activity-core on railiance01**, following the proven
`weekly-forgejo-package-prune` pattern.
## Context and constraints
### Topology (2026-07-22)
| Host | Clusters / data |
| --- | --- |
| CoulombCore | `apps-pg` (vergage, core_hub, …), `gitea-db`, `net-kingdom-pg`, `state-hub-db` |
| railiance01 | `forgejo-db`, `net-kingdom-pg`, `state-hub-db` (+ activity-core, Forgejo) |
NK and state-hub exist on **both** hosts — WP must inventory which instance is
production-of-record per consumer before scheduling dumps.
### Why workstation fails as control plane
- Laptop sleep / travel / network breaks RPO 24h.
- OIDC `bao login` is interactive; agents and workers need non-interactive
OpenBao or ESO.
- WP-0015 in-cluster CronJobs are **suspended**: cluster pods cannot reach
GitHub or Alpine CDN to obtain `age`.
### Pattern to copy (ACTIVITY-WP-0020)
```text
ActivityDefinition (cron)
→ shell context_source on activity-core worker
→ hostPath /opt/railiance-platform (or apps) CLI
→ credentials via actcore-runtime-secret / ESO
→ State Hub progress evidence sink
→ make automation-status / inventory can see the schedule
```
Reference: `activity-core/activity-definitions/weekly-forgejo-package-prune.md`.
### Target architecture
```text
activity-core Temporal schedule (railiance01)
shell resolver: cnpg_option_a_backup / forgejo_backup
hostPath tools (static age + kubectl + curl OR prebuilt image side-car script)
+ kubeconfig: railiance01 (local) + CoulombCore (remote cert/tunnel)
+ OpenBao/ESO: NC_WEBDAV_TOKEN, NC_WEBDAV_URL, AGE_PUBLIC_KEY
├─► pg_dump / forgejo dump → age → Nextcloud WebDAV
└─► patch ConfigMap cnpg-option-a-status (+ optional State Hub event)
```
**Acceptance for the workplan as a whole:** RPO 24h holds when the operator
workstation is powered off; `make cnpg-backup-status` remains healthy; schedule
appears in activity-core automation inventory; restore drill from an
activity-core-produced artifact passes.
### Repo ownership (do not invent ownership)
| Concern | Owner repo |
| --- | --- |
| S5 handoff, status Make targets, app evidence, this workplan | **railiance-apps** |
| CNPG/Forgejo dump CLI, hostPath packaging, OpenBao lane policy for workers | **railiance-platform** |
| ActivityDefinition, shell resolver, Temporal schedule, worker mounts/secrets | **activity-core** |
| SSH certs for remote kube if needed | **ops-warden** (sign only) |
| Tunnels | **ops-bridge** |
This workplan **coordinates** cross-repo work; implement in the owning repo and
cite commits/evidence here. Do **not** hand-register hub tasks — write files and
`statehub fix-consistency`.
---
## Task: Inventory production-of-record DB topology
```task
id: RAILIANCE-WP-0016-T01
status: todo
priority: high
```
Document which CNPG instances are production-of-record for each consumer DB
across CoulombCore vs railiance01 (especially `net-kingdom-pg` and
`state-hub-db` dual hosting). Produce a table in `docs/` or STATE.md: cluster,
host, databases, consumer, backup required Y/N.
**Done when:** signed-off inventory exists; no ambiguous double-backup or missed
instance remains.
## Task: Package offline-capable backup runner for railiance01
```task
id: RAILIANCE-WP-0016-T02
status: todo
priority: high
```
Deliver a runner that does **not** require workstation or internet package
install at job time:
- Static `age` (and friends) vendored under hostPath tools tree and/or a
private OCI image on Forgejo registry with `kubectl`+`age`+`curl`+`bash`.
- CLI entrypoints covering: multi-cluster logical dump (Option A), forgejo dump
path (platform), status CM update, non-secret JSON evidence.
- Prefer extending `railiance-platform` lib/tools so activity-core only shells out
(mirror `forgejo-package-prune`).
**Done when:** on railiance01 (or a dry-run host with same mounts), one command
produces age-encrypted artifacts **without** `apk`/`github.com` at runtime.
## Task: Non-interactive offsite credentials for activity-core worker
```task
id: RAILIANCE-WP-0016-T03
status: todo
priority: high
needs_human: true
```
Wire `NC_WEBDAV_TOKEN` / `NC_WEBDAV_URL` / `AGE_PUBLIC_KEY` into the activity-core
worker on railiance01 without interactive OIDC:
- Prefer ExternalSecret → `actcore-runtime-secret` (pattern: forgejo-admin ESO).
- OpenBao policy for the worker role; **do not** put `AGE_PRIVATE_KEY` on the
worker (recovery escrow stays operator/OpenBao only).
- Document rotation via CCR / warden route; no secrets in Git.
**Done when:** worker pod can upload a tiny probe object to Nextcloud and env
lengths are non-zero without operator login.
## Task: CoulombCore API reachability from railiance01 worker
```task
id: RAILIANCE-WP-0016-T04
status: todo
priority: high
```
Backup of CoulombCore clusters (`apps-pg`, …) from railiance01 requires a
stable kubeconfig + auth (ops-warden cert + ops-bridge tunnel, or in-cluster
equivalent). Options to evaluate (pick one, record decision):
1. Worker host kubeconfig for CoulombCore with long-lived cert refresh hook.
2. activity-core (or a thin agent) also on CoulombCore for local dumps only.
3. CNPG-native path later (out of scope unless re-decided).
**Done when:** non-interactive `kubectl get cluster -n databases` against
CoulombCore succeeds from the same environment that will run the schedule.
## Task: activity-core ActivityDefinition + shell resolver
```task
id: RAILIANCE-WP-0016-T05
status: todo
priority: high
```
In **activity-core**:
1. Shell resolver e.g. `cnpg_option_a_backup` (and optionally `forgejo_backup`)
analogous to `forgejo_package_prune`.
2. ActivityDefinition markdown with `trigger.type: cron` (target ~02:30 UTC
daily; weekly retention on Sundays).
3. `make sync` / schedule reconcile on railiance01; enable only after T02T04.
4. Evidence sink → State Hub progress (non-secret summary).
5. Visible in `make automation-status` / inventory.
**Done when:** Temporal schedule exists, one forced run succeeds, inventory lists
the automation as enabled.
## Task: Cut over health definition and retire workstation dependency
```task
id: RAILIANCE-WP-0016-T06
status: todo
priority: medium
```
In **railiance-apps**:
- Set `cnpg-option-a-schedule` `mode` to `activity-core` (or dual-run during
soak).
- `cnpg-backup-status` accepts activity-core last-success (status CM and/or
hub evidence age).
- Document that workstation cron is **stopgap only**; remove from operator
“required” path after soak.
- Optionally un-suspend in-cluster CronJobs **only if** T02 image path lands
in-cluster; otherwise leave suspended.
**Done when:** docs + status tool describe activity-core as primary; workstation
optional.
## Task: Soak, restore drill from automated artifact, finish
```task
id: RAILIANCE-WP-0016-T07
status: todo
priority: high
```
- ≥3 consecutive successful scheduled runs (or 7 if aligning forgejo promotion
gate language).
- Restore drill from an artifact produced by activity-core (not workstation
one-off), row-count or equivalent gate, evidence JSON under `docs/evidence/`.
- Update `STATE.md`; archive/finish this workplan; optional message to
activity-core / platform owners if residual tasks remain in their repos.
**Done when:** RPO/RTO evidence recorded; workstation powered-off scenario is
credible; workplan `finished`.
---
## Out of scope
- Switching Option A → barman ObjectStore / CNPG `ScheduledBackup` (Phase 2;
only if explicitly re-decided).
- Media/PVC app data backups (vergage media still disabled).
- Rotating EXPOSED taint history on the offsite lane (platform CCR ops), except
as a dependency note if worker fetch is blocked.
## Suggested implementation order
```text
T01 inventory → T02 runner package → T03 creds + T04 kube reachability (parallel)
→ T05 activity-core definition → T06 cutover docs/status → T07 soak + drill
```
## References
- `STATE.md` (this repo) — current posture
- `workplans/RAILIANCE-WP-0015-…` — interim Option A
- `docs/evidence/cnpg-option-a-backup-20260722.json`
- `activity-core/activity-definitions/weekly-forgejo-package-prune.md`
- `railiance-platform/docs/forgejo-backup.md`
- `activity-core/docs/runbook.md` (hostPath + ESO patterns)