2026-08-17 23:01:57 +02:00
|
|
|
|
---
|
|
|
|
|
|
id: RPF-WP-0019
|
|
|
|
|
|
type: workplan
|
|
|
|
|
|
title: "apps-pg: backup, per-consumer controls, and the isolation probes they make possible"
|
|
|
|
|
|
domain: financials
|
|
|
|
|
|
repo: railiance-platform
|
2026-08-20 22:58:45 +02:00
|
|
|
|
status: finished
|
2026-08-17 23:01:57 +02:00
|
|
|
|
owner: codex
|
|
|
|
|
|
topic_slug: railiance
|
|
|
|
|
|
created: "2026-08-17"
|
2026-08-20 22:58:45 +02:00
|
|
|
|
updated: "2026-08-20"
|
2026-08-17 23:01:57 +02:00
|
|
|
|
related:
|
|
|
|
|
|
- RPF-WP-0018
|
|
|
|
|
|
origin: residual
|
|
|
|
|
|
origin_ref: RPF-WP-0018
|
2026-08-25 20:20:45 +02:00
|
|
|
|
state_hub_workstream_id: "6ded9d76-e3a8-5d52-9221-2c7935f3b364"
|
2026-08-17 23:01:57 +02:00
|
|
|
|
---
|
|
|
|
|
|
|
|
|
|
|
|
# RPF-WP-0019 — apps-pg recoverability and per-consumer controls
|
|
|
|
|
|
|
|
|
|
|
|
## Goal
|
|
|
|
|
|
|
|
|
|
|
|
Close the three defects `RPF-WP-0018` surfaced in `apps-pg` by writing its
|
|
|
|
|
|
quota disclosure. Documentation found them; only this workplan fixes them.
|
|
|
|
|
|
|
|
|
|
|
|
## Why this is separate from RPF-WP-0018
|
|
|
|
|
|
|
|
|
|
|
|
That workplan declared posture and policy. This one changes a live cluster.
|
|
|
|
|
|
Keeping them apart matters: a declaration workplan that quietly starts editing
|
|
|
|
|
|
production is how "we wrote it down" becomes indistinguishable from "we fixed
|
|
|
|
|
|
it". The declaration is published as-is, with the defects visible, and this is
|
|
|
|
|
|
the record of closing them.
|
|
|
|
|
|
|
|
|
|
|
|
## The three defects
|
|
|
|
|
|
|
|
|
|
|
|
**D1 — `apps-pg` has no backup.** No `barmanObjectStore`, no `retentionPolicy`,
|
|
|
|
|
|
nothing. This is not a short retention window; it is no recovery path at all,
|
|
|
|
|
|
on a cluster holding two S5 application databases. Its R level is `R0` and
|
|
|
|
|
|
`R0` here means unrecoverable, not merely un-erasable.
|
|
|
|
|
|
|
|
|
|
|
|
**D2 — no per-consumer controls.** The connection pool is unpartitioned, so
|
|
|
|
|
|
one consumer can exhaust the cluster while staying politely inside its own
|
|
|
|
|
|
expectations. No `statement_timeout`, no `idle_in_transaction_session_timeout`,
|
|
|
|
|
|
no CPU or memory limits — the pod is BestEffort QoS and is the first thing
|
|
|
|
|
|
evicted under node pressure.
|
|
|
|
|
|
|
|
|
|
|
|
**D3 — no isolation probes**, so the `P1` levels recorded for `vergabe` and
|
|
|
|
|
|
`coulomb_social` in `docs/placement-policy.md` §3.1 are provisioning
|
|
|
|
|
|
declarations without the §13 artifact.
|
|
|
|
|
|
|
|
|
|
|
|
## Sequencing, and why it is not the obvious one
|
|
|
|
|
|
|
|
|
|
|
|
**D1 first.** It is the only one whose failure is unrecoverable. A cluster
|
|
|
|
|
|
with no backup is one bad afternoon from data loss that no amount of isolation
|
|
|
|
|
|
evidence compensates for.
|
|
|
|
|
|
|
|
|
|
|
|
**D2 before D3, necessarily.** Tenancy Posture §13.4: an artifact must assert
|
|
|
|
|
|
something achievable. The noisy-neighbour artifact requires showing the
|
|
|
|
|
|
governance controls *bind*. With no controls there is nothing to bind, so a
|
|
|
|
|
|
probe written now could only demonstrate degradation — an artifact that "can
|
|
|
|
|
|
only fail, or that passes by being run gently enough". Writing the probe first
|
|
|
|
|
|
would produce an overclaim wearing the costume of evidence.
|
|
|
|
|
|
|
|
|
|
|
|
D1 also depends on a backup target, which is `resource-control`'s bucket and
|
|
|
|
|
|
the `platform-pg-backup-s3` credential lane — the same handoff
|
|
|
|
|
|
`make postgres-backup-deploy` waits on. Check whether that is now live before
|
|
|
|
|
|
assuming this is blocked.
|
|
|
|
|
|
|
2026-08-18 13:35:04 +02:00
|
|
|
|
## Status 2026-08-18 — repository-complete, live-blocked
|
|
|
|
|
|
|
|
|
|
|
|
Everything this repo can do without touching the cluster is done and
|
|
|
|
|
|
committed. What remains on T01, T02 and T04 is a single operator window
|
|
|
|
|
|
against a live shared rail, in this order:
|
|
|
|
|
|
|
|
|
|
|
|
1. `make apps-pg-deploy` — Cluster reconcile: role connection limits,
|
|
|
|
|
|
Burstable requests/limits, explicit aggregate parameters.
|
|
|
|
|
|
2. Apply `helm/apps-pg-consumer-controls.sql` — the two 15s role timeouts.
|
|
|
|
|
|
Idempotent; CNPG 1.28 has no managed-role settings field, so this is
|
|
|
|
|
|
operator SQL by necessity, not by preference.
|
|
|
|
|
|
3. `make apps-pg-backup-deploy` — the ScheduledBackup, once the governed
|
|
|
|
|
|
Secret is confirmed live.
|
|
|
|
|
|
4. Capture `LastBackupSucceeded=True` and a scratch restore. **Until both
|
|
|
|
|
|
exist, `apps-pg` R stays 0** — declared configuration is not a §13
|
|
|
|
|
|
artifact, and this workplan exists because that distinction was missed once
|
|
|
|
|
|
already.
|
|
|
|
|
|
5. T04's probes, in an announced window, after 1–3 have settled.
|
|
|
|
|
|
|
Pin apps-pg targets to railiance01 by cluster identity; seed RPF-WP-0020
Two reachable clusters each carry a CNPG Cluster named apps-pg in a
namespace named databases. KUBECONFIG is an environment variable, so the
Makefile ?= default never applied, and RAILIANCE01_KUBECONFIG pointed at
config-hosteurope - a different cluster. Had the environment pointed at the
other reachable cluster instead of an unauthorized one, make apps-pg-deploy
would have applied RPF-WP-0019 connection limits, role timeouts and backup
config to the wrong cluster and reported success. The Unauthorized error was
the only thing that prevented it.
Filename selection cannot protect against this: both kubeconfigs resolve to
a 127.0.0.1 tunnel port and the environment wins either way. railiance01-guard
pins identity instead, comparing the live kube-system namespace UID against
RAILIANCE01_CLUSTER_UID, and fails closed on mismatch or unreachability. It
gates apps-pg deploy, backup-deploy, overflow-dry-run, status and shell.
Verified refusing on the wrong cluster, refusing when unreachable, and
passing on railiance01. Not global: db-status legitimately targets the other
cluster for gitea-db.
RPF-WP-0019 blocker note corrected - the cluster was never unreachable, our
wiring was wrong.
RPF-WP-0020 seeded for the pre-existing CCR test failure, which is two
unrelated problems: CCR-2026-0010 is an active lane missing its whole
openbao.auth block, and CCR-2026-0011 is an honest in-flight draft the suite
has no way to express.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 15:18:37 +02:00
|
|
|
|
**Blocker resolved 2026-08-18, and it was ours.** The earlier note recorded
|
|
|
|
|
|
this as `kubectl` returning `Unauthorized` — an access problem outside the
|
|
|
|
|
|
repo. That was wrong. railiance01 is reachable and healthy
|
|
|
|
|
|
(`~/.kube/config-railiance01`, k3s v1.35.1, `apps-pg` 9d, both consumers
|
|
|
|
|
|
present). What failed was our own wiring: `KUBECONFIG` is an environment
|
|
|
|
|
|
variable, so the Makefile's `?=` default never applied, and
|
|
|
|
|
|
`RAILIANCE01_KUBECONFIG` pointed at `config-hosteurope` — a different cluster
|
|
|
|
|
|
that happens to be unauthorized from here.
|
|
|
|
|
|
|
|
|
|
|
|
**The near-miss is the finding.** Two reachable clusters each carry a CNPG
|
|
|
|
|
|
`Cluster` named `apps-pg` in a namespace named `databases`. The other one
|
|
|
|
|
|
(k3s v1.30.3) holds `gitea-db` and only one apps-pg consumer. Had `KUBECONFIG`
|
|
|
|
|
|
pointed there instead of at an unauthorized file, `make apps-pg-deploy` would
|
|
|
|
|
|
have applied this workplan's connection limits, role timeouts and backup
|
|
|
|
|
|
configuration **to the wrong cluster, and reported success.** The
|
|
|
|
|
|
`Unauthorized` error was the only thing that prevented it.
|
|
|
|
|
|
|
|
|
|
|
|
Fixed by pinning cluster *identity* rather than kubeconfig *filename*:
|
|
|
|
|
|
`railiance01-guard` compares the live `kube-system` namespace UID against
|
|
|
|
|
|
`RAILIANCE01_CLUSTER_UID` and fails closed on mismatch or unreachability. It
|
|
|
|
|
|
gates `apps-pg-deploy`, `apps-pg-backup-deploy`, `apps-pg-overflow-dry-run`,
|
|
|
|
|
|
`apps-pg-status` and `apps-pg-shell`. Filename selection could not have
|
|
|
|
|
|
protected against this: both kubeconfigs resolve to a `127.0.0.1` tunnel port,
|
|
|
|
|
|
and the environment overrides the default either way. `make cluster-id` prints
|
|
|
|
|
|
what is currently selected.
|
|
|
|
|
|
|
|
|
|
|
|
The guard is deliberately **not** global. `db-status` legitimately targets the
|
|
|
|
|
|
other cluster for `gitea-db`, so a blanket guard would break a working target
|
|
|
|
|
|
and teach people to bypass it.
|
|
|
|
|
|
|
2026-08-18 13:35:04 +02:00
|
|
|
|
|
|
|
|
|
|
**Do not treat the rollout as evidence.** T04's P1 claim and the R-axis both
|
|
|
|
|
|
need artifacts produced *after* application, and `docs/placement-policy.md`
|
|
|
|
|
|
§3.1 and `tenancy.yaml` should be updated only then.
|
|
|
|
|
|
|
2026-08-20 22:58:45 +02:00
|
|
|
|
## Status 2026-08-20 — recoverability and controls live
|
|
|
|
|
|
|
|
|
|
|
|
The guarded live rollout completed against railiance01. The pod reconciled
|
|
|
|
|
|
from BestEffort to Burstable QoS, both consumer roles now have a 20-connection
|
|
|
|
|
|
limit, the explicit aggregate/logging parameters bind, and controlled SQL
|
|
|
|
|
|
applied and verified both 15-second role timeouts. T02 is complete.
|
|
|
|
|
|
|
|
|
|
|
|
The first WAL archive attempt exposed a repository defect before any recovery
|
|
|
|
|
|
claim was made: the runtime identity is intentionally restricted by bucket
|
|
|
|
|
|
policy to `platform-pg/*`, while the reviewed manifest used sibling prefix
|
|
|
|
|
|
`apps-pg/`. Scaleway correctly denied `PutObject`. Desired state now keeps the
|
|
|
|
|
|
cell distinct beneath the governed prefix at `platform-pg/apps-pg/` (and the
|
|
|
|
|
|
unprovisioned overflow at `platform-pg/apps-pg-2/`). Continuous archiving then
|
|
|
|
|
|
became healthy, the immediate base backup completed in eight seconds, and a
|
|
|
|
|
|
separate scratch Cluster restored both consumer databases in 56 seconds.
|
|
|
|
|
|
Evidence is `docs/evidence/RPF-WP-0019-backup-restore-2026-08-20.md`. T01 and
|
|
|
|
|
|
T02 are complete. T04 subsequently passed 14/14 and the workplan is finished.
|
|
|
|
|
|
|
2026-08-17 23:01:57 +02:00
|
|
|
|
## Tasks
|
|
|
|
|
|
|
|
|
|
|
|
```task
|
|
|
|
|
|
id: RPF-WP-0019-T01
|
2026-08-20 22:58:45 +02:00
|
|
|
|
status: done
|
2026-08-17 23:01:57 +02:00
|
|
|
|
priority: high
|
2026-08-25 20:20:45 +02:00
|
|
|
|
state_hub_task_id: "9f0c7e4f-6351-51c1-8c5e-39f770668605"
|
2026-08-17 23:01:57 +02:00
|
|
|
|
```
|
|
|
|
|
|
**Establish a backup target for `apps-pg`.** Confirm the state of the
|
|
|
|
|
|
`resource-control` bucket and the `platform-pg-backup-s3` OpenBao Secret; if
|
|
|
|
|
|
live, configure `barmanObjectStore` and a `retentionPolicy` on the cluster. If
|
|
|
|
|
|
not live, record the dependency and say so — do not leave the absence
|
|
|
|
|
|
undocumented a second time.
|
|
|
|
|
|
|
2026-08-20 22:58:45 +02:00
|
|
|
|
2026-08-18 repository readiness: the governed Secret existed live and source
|
|
|
|
|
|
carried 30-day retention, continuous WAL and a daily 02:15 backup. At that
|
|
|
|
|
|
point it incorrectly used sibling prefix `apps-pg/`; the 2026-08-20 live run
|
|
|
|
|
|
proved the bucket policy denied it and corrected the path beneath
|
|
|
|
|
|
`platform-pg/`. The ScheduledBackup had not yet been applied and no successful
|
|
|
|
|
|
backup/restore evidence existed, so T01 correctly remained progress then.
|
|
|
|
|
|
|
|
|
|
|
|
Completed live 2026-08-20. The governed path is
|
|
|
|
|
|
`platform-pg/apps-pg/`, continuous archiving is healthy, Backup
|
|
|
|
|
|
`apps-pg-daily-20260820204148` completed in eight seconds, and a separately
|
|
|
|
|
|
named scratch Cluster restored and matched both consumer databases in 56
|
|
|
|
|
|
seconds. See `docs/evidence/RPF-WP-0019-backup-restore-2026-08-20.md`.
|
2026-08-18 13:35:04 +02:00
|
|
|
|
|
2026-08-17 23:01:57 +02:00
|
|
|
|
```task
|
|
|
|
|
|
id: RPF-WP-0019-T02
|
2026-08-20 22:58:45 +02:00
|
|
|
|
status: done
|
2026-08-17 23:01:57 +02:00
|
|
|
|
priority: high
|
2026-08-25 20:20:45 +02:00
|
|
|
|
state_hub_task_id: "8c95e2b5-884c-5504-9998-5bdd8ae64b5d"
|
2026-08-17 23:01:57 +02:00
|
|
|
|
```
|
|
|
|
|
|
**Declare and enforce per-consumer controls.** Per-consumer connection
|
|
|
|
|
|
allowance, `statement_timeout`, `idle_in_transaction_session_timeout`, and
|
|
|
|
|
|
pod resource requests/limits to lift `apps-pg` off BestEffort QoS. Publish
|
|
|
|
|
|
every value in `docs/s3-consumer-interfaces.md` before it takes effect —
|
|
|
|
|
|
§10.2 is a disclosure rule, and applying a timeout consumers learn about by
|
|
|
|
|
|
hitting it would breach the rule while implementing it.
|
|
|
|
|
|
|
2026-08-18 13:35:04 +02:00
|
|
|
|
2026-08-18 repository readiness: both roles declare a 20-connection limit,
|
|
|
|
|
|
the pod has Burstable requests/limits, aggregate/logging parameters are
|
|
|
|
|
|
explicit, and controlled operator SQL sets both 15s role timeouts. Every value
|
2026-08-20 22:58:45 +02:00
|
|
|
|
was published in `docs/s3-consumer-interfaces.md` before application. At that
|
|
|
|
|
|
point, live SQL and Cluster reconciliation still required an operator window.
|
|
|
|
|
|
|
|
|
|
|
|
Completed live 2026-08-20. The railiance01 identity guard passed, the single
|
|
|
|
|
|
instance reconciled healthy with Burstable QoS, and PostgreSQL reported both
|
|
|
|
|
|
roles at connection limit 20 with `statement_timeout=15s` and
|
|
|
|
|
|
`idle_in_transaction_session_timeout=15s`. The declared aggregate and logging
|
|
|
|
|
|
parameters were also verified live.
|
2026-08-18 13:35:04 +02:00
|
|
|
|
|
2026-08-17 23:01:57 +02:00
|
|
|
|
```task
|
|
|
|
|
|
id: RPF-WP-0019-T03
|
2026-08-18 13:35:04 +02:00
|
|
|
|
status: done
|
2026-08-17 23:01:57 +02:00
|
|
|
|
priority: medium
|
2026-08-25 20:20:45 +02:00
|
|
|
|
state_hub_task_id: "9d8c534d-72b7-5cac-ada1-72273fb3ab01"
|
2026-08-17 23:01:57 +02:00
|
|
|
|
```
|
|
|
|
|
|
**Declare the ceiling and overflow target.** Owed under this repo's own
|
|
|
|
|
|
Rule P-4.1 before `apps-pg`'s third consumer; it is at two. Name the binding
|
|
|
|
|
|
resource per Rule P-4.2 — memory or connections — and a named overflow
|
|
|
|
|
|
substrate per P-4.3.
|
|
|
|
|
|
|
2026-08-18 13:35:04 +02:00
|
|
|
|
Completed 2026-08-18. The declared ceiling is three, memory is the binding
|
|
|
|
|
|
constraint, and `apps-pg-2` is a named, source-provisionable overflow cell with
|
|
|
|
|
|
a distinct credential and backup prefix. `make apps-pg-verify-capacity`
|
|
|
|
|
|
rejects a fourth consumer per cell and unbounded/duplicate roles. The cell
|
|
|
|
|
|
intentionally remains absent until a fourth consumer is approved.
|
|
|
|
|
|
|
2026-08-17 23:01:57 +02:00
|
|
|
|
```task
|
|
|
|
|
|
id: RPF-WP-0019-T04
|
2026-08-20 22:58:45 +02:00
|
|
|
|
status: done
|
2026-08-17 23:01:57 +02:00
|
|
|
|
priority: medium
|
2026-08-25 20:20:45 +02:00
|
|
|
|
state_hub_task_id: "88ef2973-98cd-508a-9b10-73efcf6dce14"
|
2026-08-17 23:01:57 +02:00
|
|
|
|
```
|
|
|
|
|
|
**Isolation probes, after T02.** Consumer-boundary probes on the
|
|
|
|
|
|
`rapp-postgres` model, then the §13 noisy-neighbour artifact: per-consumer
|
|
|
|
|
|
baseline, saturation run, evidence the controls bind, degradation measured and
|
|
|
|
|
|
judged against each consumer's declared service class. Update
|
|
|
|
|
|
`docs/placement-policy.md` §3.1 and `docs/tenancy-posture.md` when the P1
|
|
|
|
|
|
claims become evidenced.
|
|
|
|
|
|
|
2026-08-18 13:35:04 +02:00
|
|
|
|
Waiting on T02 live application and an announced probe window. No saturation
|
|
|
|
|
|
or destructive recovery experiment is run against the shared production rail
|
|
|
|
|
|
as part of repository preparation.
|
|
|
|
|
|
|
2026-08-20 22:58:45 +02:00
|
|
|
|
Started 2026-08-20 after T02 completed. The first privilege preflight found
|
|
|
|
|
|
that all three databases still inherited PostgreSQL's default `PUBLIC`
|
|
|
|
|
|
`CONNECT` and `TEMPORARY` grants, so either consumer could connect to the
|
|
|
|
|
|
other's database even though relation grants remained separate. The probe did
|
|
|
|
|
|
not launder that into a P1 pass. The controlled SQL and published interface now
|
|
|
|
|
|
revoke the defaults and grant each role access only to its own database; live
|
|
|
|
|
|
boundary and noisy-neighbour evidence follows that enforcement.
|
|
|
|
|
|
|
|
|
|
|
|
Completed 2026-08-20. `make apps-pg-isolation-probe` passed 14/14 using the
|
|
|
|
|
|
real consumer credentials. The greedy role bound at 20 connections and the
|
|
|
|
|
|
next connection was denied; five peer queries stayed available, with measured
|
|
|
|
|
|
five-query wall time increasing from 5,285ms to 8,146ms. Both consumers are
|
|
|
|
|
|
interactive but declare no numeric database latency objective, so this proves
|
|
|
|
|
|
P1 boundary and continued service, not a latency SLO or resource fairness.
|
|
|
|
|
|
Evidence: `docs/evidence/RPF-WP-0019-isolation-2026-08-20.md`.
|
|
|
|
|
|
|
2026-08-17 23:01:57 +02:00
|
|
|
|
## Boundaries
|
|
|
|
|
|
|
|
|
|
|
|
- `apps-pg` only. `platform-pg`'s equivalents are `rapp-postgres`'s.
|
|
|
|
|
|
- No consumer is migrated. This changes the cluster, not who is on it.
|
|
|
|
|
|
- Values are published before they are enforced, never after.
|
|
|
|
|
|
|
|
|
|
|
|
## Risks
|
|
|
|
|
|
|
|
|
|
|
|
**Applying limits to a live cluster breaks a consumer that was relying on
|
|
|
|
|
|
their absence.** Most likely with `statement_timeout`. Mitigation is the
|
|
|
|
|
|
disclosure-first ordering in T02, which gives consumers a window to object.
|
|
|
|
|
|
|
|
|
|
|
|
**T01 stays blocked on a handoff outside this repo and the cluster keeps no
|
|
|
|
|
|
backup meanwhile.** Mitigation is that T01 requires the dependency be recorded
|
|
|
|
|
|
explicitly rather than left as a silent absence — which is exactly how D1
|
|
|
|
|
|
survived this long.
|