railiance-apps/STATE.md
tegwick 7bc76dc0c6
All checks were successful
CI Smoke / host-smoke (push) Successful in 1s
CI Smoke / container-smoke (push) Successful in 3s
docs: STATE.md assessment; propose RAILIANCE-WP-0016 activity-core backups
Capture post-WP-0015 posture (workstation SPOF) and register a workplan to
move unattended Option A + Forgejo backup automation to railiance01 via
activity-core without a laptop control plane.
2026-07-22 19:35:25 +02:00

167 lines
6.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# STATE — railiance-apps
**Updated:** 2026-07-22
**Domain:** financials · **Repo:** railiance-apps
**Topic ID:** `ca369340-a64e-442e-98f1-a4fa7dc74a38`
This file is the human/agent snapshot of **where the repo stands**. Prefer it over
chat history for orientation. Generated custodian brief is `.custodian-brief.md`
(do not edit). Formal work structure lives in `workplans/`.
---
## One-line posture
S5 app releases (vergabe, hubs, Forgejo consumer surface) are operational; **CNPG
Option A logical backups exist and currently report healthy**, but unattended
execution still depends on a **workstation path** and must move to **railiance01 +
activity-core** so the laptop is not a durability SPOF.
---
## Clusters (as of 2026-07-22)
| Host | Kubeconfig | CNPG clusters in `databases` | Role |
| --- | --- | --- | --- |
| **CoulombCore** | `~/.kube/config` | `apps-pg`, `gitea-db`, `net-kingdom-pg`, `state-hub-db` | S5 apps DBs, legacy gitea-db, NK + state-hub instances |
| **railiance01** | `~/.kube/config-hosteurope` | `forgejo-db`, `net-kingdom-pg`, `state-hub-db` | Forgejo DB; NK + state-hub instances (may differ from Core) |
`PRODUCTION_KUBECONFIG` defaults to CoulombCore. Forgejo tooling uses railiance01.
---
## Backup durability (RAILIANCE-WP-0015 — **finished**)
### Goal met
Drive `make cnpg-backup-status` **degraded → healthy** with Option A (age-encrypted
logical dump → Nextcloud WebDAV), not barman `ScheduledBackup`.
### What is live today
| Item | State |
| --- | --- |
| OpenBao lane `CCR-2026-0004` / `railiance-backup-offsite-lane` | **resolvable**, fields present |
| Secret `databases/cnpg-backup-offsite` (CoulombCore) | **present** (`NC_WEBDAV_*`, `AGE_PUBLIC_KEY` only) |
| ConfigMap `cnpg-option-a-schedule` | **mode=workstation-cron** @ `30 2 * * *` |
| ConfigMap `cnpg-option-a-status` | last-success timestamps for all four Core clusters (2026-07-22) |
| CronJobs `cnpg-logical-backup-*` | **suspended** (cluster egress cannot install `age`) |
| Workstation runner `tools/cnpg-logical-backup.sh` | **works** (upload + status CM) |
| Restore drill | apps-pg/`vergage_db` age artifact **195/195** |
| Evidence | `docs/evidence/cnpg-option-a-backup-20260722.json` |
### What is *not* good enough
1. **Workstation SPOF** — daily unattended path is “someones machine has crontab +
bao OIDC + kubeconfig”. Laptop sleep/off breaks RPO.
2. **In-cluster CronJobs blocked** — pods cannot reach GitHub or Alpine CDN to get
`age`; no prebuilt image with `kubectl`+`age`+`curl` yet.
3. **Split brain topology** — Option A runner today targets **CoulombCore** four
clusters only; **forgejo-db** (railiance01) still uses platform `make forgejo-backup`
(also historically workstation-cron).
4. **No activity-core schedule** — unlike weekly Forgejo package prune, backups are
not Temporal/`ActivityDefinition` managed, so automation inventory and
`automation-status` cannot own them.
5. **Credential delivery** — offsite lane is human OIDC on workstation; worker needs
non-interactive OpenBao/ESO (or sealed secret) path.
### Operator residual (manual)
```cron
# Not yet proven installed on a always-on host — install on railiance01 or leave as stopgap
30 2 * * * cd $HOME/railiance-apps && make cnpg-logical-backup >>$HOME/.cache/railiance/backups/cnpg/cron.log 2>&1
```
---
## Adjacent finished work
| Workplan | Outcome |
| --- | --- |
| RAILIANCE-WP-0013 | Phase 1 apps-pg/`vergage_db` backup + restore drill (interim S5 gate) |
| RAILIANCE-WP-0015 | Fleet Option A tooling + healthy status on CoulombCore (workstation-backed) |
| RAIL-HO-WP-0005 | Forgejo prod migration; descoped residual backup → WP-0015 |
| ACTIVITY-WP-0020 | Weekly forgejo package prune via **activity-core** on railiance01 (pattern to copy) |
---
## Next major strand
**Move unattended Option A (+ forgejo-backup) to railiance01, scheduled by
activity-core — no workstation required.**
Tracking workplan: **`RAILIANCE-WP-0016`**
`workplans/RAILIANCE-WP-0016-railiance01-activity-core-backup-automation.md`
Design sketch (see workplan for tasks):
```text
┌─────────────────────────────┐
│ activity-core (railiance01) │
│ Temporal cron ActivityDef │
└─────────────┬───────────────┘
│ shell context
┌─────────────────────────────┐
│ hostPath /opt/railiance-* │
│ backup CLI (age+kubectl) │
│ OpenBao/ESO offsite secrets │
└─────────────┬───────────────┘
┌──────────────────┼──────────────────┐
▼ ▼ ▼
CoulombCore API railiance01 API Nextcloud WebDAV
(apps-pg, …) (forgejo-db, …) (age artifacts)
```
**Proven pattern:** `weekly-forgejo-package-prune` ActivityDefinition +
`/opt/railiance-platform` hostPath + ESO credentials + State Hub evidence sink.
---
## Active / proposed workplans
| ID | Status | Title |
| --- | --- | --- |
| RAILIANCE-WP-0015 | **finished** | CNPG Option A coverage → healthy (workstation interim) |
| RAILIANCE-WP-0016 | **proposed** | railiance01 + activity-core unattended backup automation |
(No other active workplans in hub as of last brief.)
---
## Quick operator commands
```bash
# Health (CoulombCore)
make cnpg-backup-status
# Manual Option A run (still workstation / OpenBao OIDC today)
bao login -method=oidc -path=netkingdom role=railiance-backup-workload-kv-read
make cnpg-logical-backup
# Secret re-materialize
make cnpg-backup-offsite-secret-apply
```
---
## Risks / open questions
| Risk | Mitigation direction (WP-0016) |
| --- | --- |
| Dual NK/state-hub DB instances on two hosts | Inventory which is production-of-record per consumer |
| CoulombCore access from railiance01 | Dedicated kubeconfig + ops-warden cert / tunnel for worker |
| Age binary in restricted network | Vendor static binary in image or hostPath tools tree |
| High-risk WebDAV + age keys in worker | ESO from OpenBao; never private key in Git; agent boundary |
| RPO if activity-core worker down | Temporal schedule health + status CM freshness alerts |
---
## Related docs
- `docs/app-data-backup-restore-handoff.md`
- `docs/credential-routing-railiance-apps.md`
- `manifests/cnpg-option-a-backup.yaml`
- `docs/evidence/cnpg-option-a-backup-20260722.json`
- activity-core: `activity-definitions/weekly-forgejo-package-prune.md`, `docs/runbook.md`
- platform: `docs/forgejo-backup.md`