2026-08-17 23:01:57 +02:00
|
|
|
|
---
|
|
|
|
|
|
id: RPF-WP-0019
|
|
|
|
|
|
type: workplan
|
|
|
|
|
|
title: "apps-pg: backup, per-consumer controls, and the isolation probes they make possible"
|
|
|
|
|
|
domain: financials
|
|
|
|
|
|
repo: railiance-platform
|
2026-08-18 13:35:04 +02:00
|
|
|
|
status: active
|
2026-08-17 23:01:57 +02:00
|
|
|
|
owner: codex
|
|
|
|
|
|
topic_slug: railiance
|
|
|
|
|
|
created: "2026-08-17"
|
2026-08-18 13:35:04 +02:00
|
|
|
|
updated: "2026-08-18"
|
2026-08-17 23:01:57 +02:00
|
|
|
|
related:
|
|
|
|
|
|
- RPF-WP-0018
|
|
|
|
|
|
origin: residual
|
|
|
|
|
|
origin_ref: RPF-WP-0018
|
2026-08-18 13:35:04 +02:00
|
|
|
|
state_hub_workstream_id: "160e226d-27b0-4c94-9e99-331f4354dd12"
|
2026-08-17 23:01:57 +02:00
|
|
|
|
---
|
|
|
|
|
|
|
|
|
|
|
|
# RPF-WP-0019 — apps-pg recoverability and per-consumer controls
|
|
|
|
|
|
|
|
|
|
|
|
## Goal
|
|
|
|
|
|
|
|
|
|
|
|
Close the three defects `RPF-WP-0018` surfaced in `apps-pg` by writing its
|
|
|
|
|
|
quota disclosure. Documentation found them; only this workplan fixes them.
|
|
|
|
|
|
|
|
|
|
|
|
## Why this is separate from RPF-WP-0018
|
|
|
|
|
|
|
|
|
|
|
|
That workplan declared posture and policy. This one changes a live cluster.
|
|
|
|
|
|
Keeping them apart matters: a declaration workplan that quietly starts editing
|
|
|
|
|
|
production is how "we wrote it down" becomes indistinguishable from "we fixed
|
|
|
|
|
|
it". The declaration is published as-is, with the defects visible, and this is
|
|
|
|
|
|
the record of closing them.
|
|
|
|
|
|
|
|
|
|
|
|
## The three defects
|
|
|
|
|
|
|
|
|
|
|
|
**D1 — `apps-pg` has no backup.** No `barmanObjectStore`, no `retentionPolicy`,
|
|
|
|
|
|
nothing. This is not a short retention window; it is no recovery path at all,
|
|
|
|
|
|
on a cluster holding two S5 application databases. Its R level is `R0` and
|
|
|
|
|
|
`R0` here means unrecoverable, not merely un-erasable.
|
|
|
|
|
|
|
|
|
|
|
|
**D2 — no per-consumer controls.** The connection pool is unpartitioned, so
|
|
|
|
|
|
one consumer can exhaust the cluster while staying politely inside its own
|
|
|
|
|
|
expectations. No `statement_timeout`, no `idle_in_transaction_session_timeout`,
|
|
|
|
|
|
no CPU or memory limits — the pod is BestEffort QoS and is the first thing
|
|
|
|
|
|
evicted under node pressure.
|
|
|
|
|
|
|
|
|
|
|
|
**D3 — no isolation probes**, so the `P1` levels recorded for `vergabe` and
|
|
|
|
|
|
`coulomb_social` in `docs/placement-policy.md` §3.1 are provisioning
|
|
|
|
|
|
declarations without the §13 artifact.
|
|
|
|
|
|
|
|
|
|
|
|
## Sequencing, and why it is not the obvious one
|
|
|
|
|
|
|
|
|
|
|
|
**D1 first.** It is the only one whose failure is unrecoverable. A cluster
|
|
|
|
|
|
with no backup is one bad afternoon from data loss that no amount of isolation
|
|
|
|
|
|
evidence compensates for.
|
|
|
|
|
|
|
|
|
|
|
|
**D2 before D3, necessarily.** Tenancy Posture §13.4: an artifact must assert
|
|
|
|
|
|
something achievable. The noisy-neighbour artifact requires showing the
|
|
|
|
|
|
governance controls *bind*. With no controls there is nothing to bind, so a
|
|
|
|
|
|
probe written now could only demonstrate degradation — an artifact that "can
|
|
|
|
|
|
only fail, or that passes by being run gently enough". Writing the probe first
|
|
|
|
|
|
would produce an overclaim wearing the costume of evidence.
|
|
|
|
|
|
|
|
|
|
|
|
D1 also depends on a backup target, which is `resource-control`'s bucket and
|
|
|
|
|
|
the `platform-pg-backup-s3` credential lane — the same handoff
|
|
|
|
|
|
`make postgres-backup-deploy` waits on. Check whether that is now live before
|
|
|
|
|
|
assuming this is blocked.
|
|
|
|
|
|
|
2026-08-18 13:35:04 +02:00
|
|
|
|
## Status 2026-08-18 — repository-complete, live-blocked
|
|
|
|
|
|
|
|
|
|
|
|
Everything this repo can do without touching the cluster is done and
|
|
|
|
|
|
committed. What remains on T01, T02 and T04 is a single operator window
|
|
|
|
|
|
against a live shared rail, in this order:
|
|
|
|
|
|
|
|
|
|
|
|
1. `make apps-pg-deploy` — Cluster reconcile: role connection limits,
|
|
|
|
|
|
Burstable requests/limits, explicit aggregate parameters.
|
|
|
|
|
|
2. Apply `helm/apps-pg-consumer-controls.sql` — the two 15s role timeouts.
|
|
|
|
|
|
Idempotent; CNPG 1.28 has no managed-role settings field, so this is
|
|
|
|
|
|
operator SQL by necessity, not by preference.
|
|
|
|
|
|
3. `make apps-pg-backup-deploy` — the ScheduledBackup, once the governed
|
|
|
|
|
|
Secret is confirmed live.
|
|
|
|
|
|
4. Capture `LastBackupSucceeded=True` and a scratch restore. **Until both
|
|
|
|
|
|
exist, `apps-pg` R stays 0** — declared configuration is not a §13
|
|
|
|
|
|
artifact, and this workplan exists because that distinction was missed once
|
|
|
|
|
|
already.
|
|
|
|
|
|
5. T04's probes, in an announced window, after 1–3 have settled.
|
|
|
|
|
|
|
|
|
|
|
|
**Blocker as of this session:** the cluster is unreachable from the
|
|
|
|
|
|
workstation — `kubectl` returns `Unauthorized` against
|
|
|
|
|
|
`config-hosteurope`. Credentials or the tunnel need attention before any of
|
|
|
|
|
|
the above runs. This is a session-local access problem, not a defect in the
|
|
|
|
|
|
manifests: `make apps-pg-verify-capacity` passes and the capacity tests are
|
|
|
|
|
|
green.
|
|
|
|
|
|
|
|
|
|
|
|
**Do not treat the rollout as evidence.** T04's P1 claim and the R-axis both
|
|
|
|
|
|
need artifacts produced *after* application, and `docs/placement-policy.md`
|
|
|
|
|
|
§3.1 and `tenancy.yaml` should be updated only then.
|
|
|
|
|
|
|
2026-08-17 23:01:57 +02:00
|
|
|
|
## Tasks
|
|
|
|
|
|
|
|
|
|
|
|
```task
|
|
|
|
|
|
id: RPF-WP-0019-T01
|
2026-08-18 13:35:04 +02:00
|
|
|
|
status: progress
|
2026-08-17 23:01:57 +02:00
|
|
|
|
priority: high
|
2026-08-18 13:35:04 +02:00
|
|
|
|
state_hub_task_id: "0f5175c7-bde3-46a0-87a8-b50b5eeac32c"
|
2026-08-17 23:01:57 +02:00
|
|
|
|
```
|
|
|
|
|
|
**Establish a backup target for `apps-pg`.** Confirm the state of the
|
|
|
|
|
|
`resource-control` bucket and the `platform-pg-backup-s3` OpenBao Secret; if
|
|
|
|
|
|
live, configure `barmanObjectStore` and a `retentionPolicy` on the cluster. If
|
|
|
|
|
|
not live, record the dependency and say so — do not leave the absence
|
|
|
|
|
|
undocumented a second time.
|
|
|
|
|
|
|
2026-08-18 13:35:04 +02:00
|
|
|
|
2026-08-18 repository readiness: the governed Secret exists live and reviewed
|
|
|
|
|
|
desired state now carries a distinct `apps-pg/` prefix, 30-day retention,
|
|
|
|
|
|
continuous WAL and a daily 02:15 backup. NetworkPolicy admits the DNS/HTTPS
|
|
|
|
|
|
path archival needs. The ScheduledBackup is not applied and no successful
|
|
|
|
|
|
backup/restore evidence exists, so T01 remains progress.
|
|
|
|
|
|
|
2026-08-17 23:01:57 +02:00
|
|
|
|
```task
|
|
|
|
|
|
id: RPF-WP-0019-T02
|
2026-08-18 13:35:04 +02:00
|
|
|
|
status: progress
|
2026-08-17 23:01:57 +02:00
|
|
|
|
priority: high
|
2026-08-18 13:35:04 +02:00
|
|
|
|
state_hub_task_id: "c5b0b2ac-b9f5-42fb-8e75-2fac2ab1e852"
|
2026-08-17 23:01:57 +02:00
|
|
|
|
```
|
|
|
|
|
|
**Declare and enforce per-consumer controls.** Per-consumer connection
|
|
|
|
|
|
allowance, `statement_timeout`, `idle_in_transaction_session_timeout`, and
|
|
|
|
|
|
pod resource requests/limits to lift `apps-pg` off BestEffort QoS. Publish
|
|
|
|
|
|
every value in `docs/s3-consumer-interfaces.md` before it takes effect —
|
|
|
|
|
|
§10.2 is a disclosure rule, and applying a timeout consumers learn about by
|
|
|
|
|
|
hitting it would breach the rule while implementing it.
|
|
|
|
|
|
|
2026-08-18 13:35:04 +02:00
|
|
|
|
2026-08-18 repository readiness: both roles declare a 20-connection limit,
|
|
|
|
|
|
the pod has Burstable requests/limits, aggregate/logging parameters are
|
|
|
|
|
|
explicit, and controlled operator SQL sets both 15s role timeouts. Every value
|
|
|
|
|
|
was published in `docs/s3-consumer-interfaces.md` before application. Live SQL
|
|
|
|
|
|
and Cluster reconciliation remain an operator window, so T02 remains progress.
|
|
|
|
|
|
|
2026-08-17 23:01:57 +02:00
|
|
|
|
```task
|
|
|
|
|
|
id: RPF-WP-0019-T03
|
2026-08-18 13:35:04 +02:00
|
|
|
|
status: done
|
2026-08-17 23:01:57 +02:00
|
|
|
|
priority: medium
|
2026-08-18 13:35:04 +02:00
|
|
|
|
state_hub_task_id: "736cbc11-1992-4f1b-9ff0-cb4622ff39a5"
|
2026-08-17 23:01:57 +02:00
|
|
|
|
```
|
|
|
|
|
|
**Declare the ceiling and overflow target.** Owed under this repo's own
|
|
|
|
|
|
Rule P-4.1 before `apps-pg`'s third consumer; it is at two. Name the binding
|
|
|
|
|
|
resource per Rule P-4.2 — memory or connections — and a named overflow
|
|
|
|
|
|
substrate per P-4.3.
|
|
|
|
|
|
|
2026-08-18 13:35:04 +02:00
|
|
|
|
Completed 2026-08-18. The declared ceiling is three, memory is the binding
|
|
|
|
|
|
constraint, and `apps-pg-2` is a named, source-provisionable overflow cell with
|
|
|
|
|
|
a distinct credential and backup prefix. `make apps-pg-verify-capacity`
|
|
|
|
|
|
rejects a fourth consumer per cell and unbounded/duplicate roles. The cell
|
|
|
|
|
|
intentionally remains absent until a fourth consumer is approved.
|
|
|
|
|
|
|
2026-08-17 23:01:57 +02:00
|
|
|
|
```task
|
|
|
|
|
|
id: RPF-WP-0019-T04
|
2026-08-18 13:35:04 +02:00
|
|
|
|
status: wait
|
2026-08-17 23:01:57 +02:00
|
|
|
|
priority: medium
|
2026-08-18 13:35:04 +02:00
|
|
|
|
state_hub_task_id: "2fccdd26-c9c6-43cb-b069-27ed5668afac"
|
2026-08-17 23:01:57 +02:00
|
|
|
|
```
|
|
|
|
|
|
**Isolation probes, after T02.** Consumer-boundary probes on the
|
|
|
|
|
|
`rapp-postgres` model, then the §13 noisy-neighbour artifact: per-consumer
|
|
|
|
|
|
baseline, saturation run, evidence the controls bind, degradation measured and
|
|
|
|
|
|
judged against each consumer's declared service class. Update
|
|
|
|
|
|
`docs/placement-policy.md` §3.1 and `docs/tenancy-posture.md` when the P1
|
|
|
|
|
|
claims become evidenced.
|
|
|
|
|
|
|
2026-08-18 13:35:04 +02:00
|
|
|
|
Waiting on T02 live application and an announced probe window. No saturation
|
|
|
|
|
|
or destructive recovery experiment is run against the shared production rail
|
|
|
|
|
|
as part of repository preparation.
|
|
|
|
|
|
|
2026-08-17 23:01:57 +02:00
|
|
|
|
## Boundaries
|
|
|
|
|
|
|
|
|
|
|
|
- `apps-pg` only. `platform-pg`'s equivalents are `rapp-postgres`'s.
|
|
|
|
|
|
- No consumer is migrated. This changes the cluster, not who is on it.
|
|
|
|
|
|
- Values are published before they are enforced, never after.
|
|
|
|
|
|
|
|
|
|
|
|
## Risks
|
|
|
|
|
|
|
|
|
|
|
|
**Applying limits to a live cluster breaks a consumer that was relying on
|
|
|
|
|
|
their absence.** Most likely with `statement_timeout`. Mitigation is the
|
|
|
|
|
|
disclosure-first ordering in T02, which gives consumers a window to object.
|
|
|
|
|
|
|
|
|
|
|
|
**T01 stays blocked on a handoff outside this repo and the cluster keeps no
|
|
|
|
|
|
backup meanwhile.** Mitigation is that T01 requires the dependency be recorded
|
|
|
|
|
|
explicitly rather than left as a silent absence — which is exactly how D1
|
|
|
|
|
|
survived this long.
|