12 KiB
| id | type | title | domain | repo | status | owner | topic_slug | created | updated | related | origin | origin_ref | state_hub_workstream_id | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| RPF-WP-0019 | workplan | apps-pg: backup, per-consumer controls, and the isolation probes they make possible | financials | railiance-platform | finished | codex | railiance | 2026-08-17 | 2026-08-20 |
|
residual | RPF-WP-0018 | 160e226d-27b0-4c94-9e99-331f4354dd12 |
RPF-WP-0019 — apps-pg recoverability and per-consumer controls
Goal
Close the three defects RPF-WP-0018 surfaced in apps-pg by writing its
quota disclosure. Documentation found them; only this workplan fixes them.
Why this is separate from RPF-WP-0018
That workplan declared posture and policy. This one changes a live cluster. Keeping them apart matters: a declaration workplan that quietly starts editing production is how "we wrote it down" becomes indistinguishable from "we fixed it". The declaration is published as-is, with the defects visible, and this is the record of closing them.
The three defects
D1 — apps-pg has no backup. No barmanObjectStore, no retentionPolicy,
nothing. This is not a short retention window; it is no recovery path at all,
on a cluster holding two S5 application databases. Its R level is R0 and
R0 here means unrecoverable, not merely un-erasable.
D2 — no per-consumer controls. The connection pool is unpartitioned, so
one consumer can exhaust the cluster while staying politely inside its own
expectations. No statement_timeout, no idle_in_transaction_session_timeout,
no CPU or memory limits — the pod is BestEffort QoS and is the first thing
evicted under node pressure.
D3 — no isolation probes, so the P1 levels recorded for vergabe and
coulomb_social in docs/placement-policy.md §3.1 are provisioning
declarations without the §13 artifact.
Sequencing, and why it is not the obvious one
D1 first. It is the only one whose failure is unrecoverable. A cluster with no backup is one bad afternoon from data loss that no amount of isolation evidence compensates for.
D2 before D3, necessarily. Tenancy Posture §13.4: an artifact must assert something achievable. The noisy-neighbour artifact requires showing the governance controls bind. With no controls there is nothing to bind, so a probe written now could only demonstrate degradation — an artifact that "can only fail, or that passes by being run gently enough". Writing the probe first would produce an overclaim wearing the costume of evidence.
D1 also depends on a backup target, which is resource-control's bucket and
the platform-pg-backup-s3 credential lane — the same handoff
make postgres-backup-deploy waits on. Check whether that is now live before
assuming this is blocked.
Status 2026-08-18 — repository-complete, live-blocked
Everything this repo can do without touching the cluster is done and committed. What remains on T01, T02 and T04 is a single operator window against a live shared rail, in this order:
make apps-pg-deploy— Cluster reconcile: role connection limits, Burstable requests/limits, explicit aggregate parameters.- Apply
helm/apps-pg-consumer-controls.sql— the two 15s role timeouts. Idempotent; CNPG 1.28 has no managed-role settings field, so this is operator SQL by necessity, not by preference. make apps-pg-backup-deploy— the ScheduledBackup, once the governed Secret is confirmed live.- Capture
LastBackupSucceeded=Trueand a scratch restore. Until both exist,apps-pgR stays 0 — declared configuration is not a §13 artifact, and this workplan exists because that distinction was missed once already. - T04's probes, in an announced window, after 1–3 have settled.
Blocker resolved 2026-08-18, and it was ours. The earlier note recorded
this as kubectl returning Unauthorized — an access problem outside the
repo. That was wrong. railiance01 is reachable and healthy
(~/.kube/config-railiance01, k3s v1.35.1, apps-pg 9d, both consumers
present). What failed was our own wiring: KUBECONFIG is an environment
variable, so the Makefile's ?= default never applied, and
RAILIANCE01_KUBECONFIG pointed at config-hosteurope — a different cluster
that happens to be unauthorized from here.
The near-miss is the finding. Two reachable clusters each carry a CNPG
Cluster named apps-pg in a namespace named databases. The other one
(k3s v1.30.3) holds gitea-db and only one apps-pg consumer. Had KUBECONFIG
pointed there instead of at an unauthorized file, make apps-pg-deploy would
have applied this workplan's connection limits, role timeouts and backup
configuration to the wrong cluster, and reported success. The
Unauthorized error was the only thing that prevented it.
Fixed by pinning cluster identity rather than kubeconfig filename:
railiance01-guard compares the live kube-system namespace UID against
RAILIANCE01_CLUSTER_UID and fails closed on mismatch or unreachability. It
gates apps-pg-deploy, apps-pg-backup-deploy, apps-pg-overflow-dry-run,
apps-pg-status and apps-pg-shell. Filename selection could not have
protected against this: both kubeconfigs resolve to a 127.0.0.1 tunnel port,
and the environment overrides the default either way. make cluster-id prints
what is currently selected.
The guard is deliberately not global. db-status legitimately targets the
other cluster for gitea-db, so a blanket guard would break a working target
and teach people to bypass it.
Do not treat the rollout as evidence. T04's P1 claim and the R-axis both
need artifacts produced after application, and docs/placement-policy.md
§3.1 and tenancy.yaml should be updated only then.
Status 2026-08-20 — recoverability and controls live
The guarded live rollout completed against railiance01. The pod reconciled from BestEffort to Burstable QoS, both consumer roles now have a 20-connection limit, the explicit aggregate/logging parameters bind, and controlled SQL applied and verified both 15-second role timeouts. T02 is complete.
The first WAL archive attempt exposed a repository defect before any recovery
claim was made: the runtime identity is intentionally restricted by bucket
policy to platform-pg/*, while the reviewed manifest used sibling prefix
apps-pg/. Scaleway correctly denied PutObject. Desired state now keeps the
cell distinct beneath the governed prefix at platform-pg/apps-pg/ (and the
unprovisioned overflow at platform-pg/apps-pg-2/). Continuous archiving then
became healthy, the immediate base backup completed in eight seconds, and a
separate scratch Cluster restored both consumer databases in 56 seconds.
Evidence is docs/evidence/RPF-WP-0019-backup-restore-2026-08-20.md. T01 and
T02 are complete. T04 subsequently passed 14/14 and the workplan is finished.
Tasks
id: RPF-WP-0019-T01
status: done
priority: high
state_hub_task_id: "0f5175c7-bde3-46a0-87a8-b50b5eeac32c"
Establish a backup target for apps-pg. Confirm the state of the
resource-control bucket and the platform-pg-backup-s3 OpenBao Secret; if
live, configure barmanObjectStore and a retentionPolicy on the cluster. If
not live, record the dependency and say so — do not leave the absence
undocumented a second time.
2026-08-18 repository readiness: the governed Secret existed live and source
carried 30-day retention, continuous WAL and a daily 02:15 backup. At that
point it incorrectly used sibling prefix apps-pg/; the 2026-08-20 live run
proved the bucket policy denied it and corrected the path beneath
platform-pg/. The ScheduledBackup had not yet been applied and no successful
backup/restore evidence existed, so T01 correctly remained progress then.
Completed live 2026-08-20. The governed path is
platform-pg/apps-pg/, continuous archiving is healthy, Backup
apps-pg-daily-20260820204148 completed in eight seconds, and a separately
named scratch Cluster restored and matched both consumer databases in 56
seconds. See docs/evidence/RPF-WP-0019-backup-restore-2026-08-20.md.
id: RPF-WP-0019-T02
status: done
priority: high
state_hub_task_id: "c5b0b2ac-b9f5-42fb-8e75-2fac2ab1e852"
Declare and enforce per-consumer controls. Per-consumer connection
allowance, statement_timeout, idle_in_transaction_session_timeout, and
pod resource requests/limits to lift apps-pg off BestEffort QoS. Publish
every value in docs/s3-consumer-interfaces.md before it takes effect —
§10.2 is a disclosure rule, and applying a timeout consumers learn about by
hitting it would breach the rule while implementing it.
2026-08-18 repository readiness: both roles declare a 20-connection limit,
the pod has Burstable requests/limits, aggregate/logging parameters are
explicit, and controlled operator SQL sets both 15s role timeouts. Every value
was published in docs/s3-consumer-interfaces.md before application. At that
point, live SQL and Cluster reconciliation still required an operator window.
Completed live 2026-08-20. The railiance01 identity guard passed, the single
instance reconciled healthy with Burstable QoS, and PostgreSQL reported both
roles at connection limit 20 with statement_timeout=15s and
idle_in_transaction_session_timeout=15s. The declared aggregate and logging
parameters were also verified live.
id: RPF-WP-0019-T03
status: done
priority: medium
state_hub_task_id: "736cbc11-1992-4f1b-9ff0-cb4622ff39a5"
Declare the ceiling and overflow target. Owed under this repo's own
Rule P-4.1 before apps-pg's third consumer; it is at two. Name the binding
resource per Rule P-4.2 — memory or connections — and a named overflow
substrate per P-4.3.
Completed 2026-08-18. The declared ceiling is three, memory is the binding
constraint, and apps-pg-2 is a named, source-provisionable overflow cell with
a distinct credential and backup prefix. make apps-pg-verify-capacity
rejects a fourth consumer per cell and unbounded/duplicate roles. The cell
intentionally remains absent until a fourth consumer is approved.
id: RPF-WP-0019-T04
status: done
priority: medium
state_hub_task_id: "2fccdd26-c9c6-43cb-b069-27ed5668afac"
Isolation probes, after T02. Consumer-boundary probes on the
rapp-postgres model, then the §13 noisy-neighbour artifact: per-consumer
baseline, saturation run, evidence the controls bind, degradation measured and
judged against each consumer's declared service class. Update
docs/placement-policy.md §3.1 and docs/tenancy-posture.md when the P1
claims become evidenced.
Waiting on T02 live application and an announced probe window. No saturation or destructive recovery experiment is run against the shared production rail as part of repository preparation.
Started 2026-08-20 after T02 completed. The first privilege preflight found
that all three databases still inherited PostgreSQL's default PUBLIC
CONNECT and TEMPORARY grants, so either consumer could connect to the
other's database even though relation grants remained separate. The probe did
not launder that into a P1 pass. The controlled SQL and published interface now
revoke the defaults and grant each role access only to its own database; live
boundary and noisy-neighbour evidence follows that enforcement.
Completed 2026-08-20. make apps-pg-isolation-probe passed 14/14 using the
real consumer credentials. The greedy role bound at 20 connections and the
next connection was denied; five peer queries stayed available, with measured
five-query wall time increasing from 5,285ms to 8,146ms. Both consumers are
interactive but declare no numeric database latency objective, so this proves
P1 boundary and continued service, not a latency SLO or resource fairness.
Evidence: docs/evidence/RPF-WP-0019-isolation-2026-08-20.md.
Boundaries
apps-pgonly.platform-pg's equivalents arerapp-postgres's.- No consumer is migrated. This changes the cluster, not who is on it.
- Values are published before they are enforced, never after.
Risks
Applying limits to a live cluster breaks a consumer that was relying on
their absence. Most likely with statement_timeout. Mitigation is the
disclosure-first ordering in T02, which gives consumers a window to object.
T01 stays blocked on a handoff outside this repo and the cluster keeps no backup meanwhile. Mitigation is that T01 requires the dependency be recorded explicitly rather than left as a silent absence — which is exactly how D1 survived this long.