Placement policy 3.1 records that the P1 levels for vergabe and coulomb_social are provisioning declarations without the SS13 isolation probes - apps-pg has none. 3.2 records that apps-pg cannot produce the noisy-neighbour artifact at all until per-consumer controls exist, per SS13.4: with nothing to bind, a probe could only demonstrate degradation. RPF-WP-0019 seeded for the three live defects the disclosure work found: no backup, no per-consumer controls, no isolation probes. Kept separate from RPF-WP-0018 so declaring is not confused with fixing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
4.9 KiB
| id | type | title | domain | repo | status | owner | topic_slug | created | updated | related | origin | origin_ref | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| RPF-WP-0019 | workplan | apps-pg: backup, per-consumer controls, and the isolation probes they make possible | financials | railiance-platform | proposed | codex | railiance | 2026-08-17 | 2026-08-17 |
|
residual | RPF-WP-0018 |
RPF-WP-0019 — apps-pg recoverability and per-consumer controls
Goal
Close the three defects RPF-WP-0018 surfaced in apps-pg by writing its
quota disclosure. Documentation found them; only this workplan fixes them.
Why this is separate from RPF-WP-0018
That workplan declared posture and policy. This one changes a live cluster. Keeping them apart matters: a declaration workplan that quietly starts editing production is how "we wrote it down" becomes indistinguishable from "we fixed it". The declaration is published as-is, with the defects visible, and this is the record of closing them.
The three defects
D1 — apps-pg has no backup. No barmanObjectStore, no retentionPolicy,
nothing. This is not a short retention window; it is no recovery path at all,
on a cluster holding two S5 application databases. Its R level is R0 and
R0 here means unrecoverable, not merely un-erasable.
D2 — no per-consumer controls. The connection pool is unpartitioned, so
one consumer can exhaust the cluster while staying politely inside its own
expectations. No statement_timeout, no idle_in_transaction_session_timeout,
no CPU or memory limits — the pod is BestEffort QoS and is the first thing
evicted under node pressure.
D3 — no isolation probes, so the P1 levels recorded for vergabe and
coulomb_social in docs/placement-policy.md §3.1 are provisioning
declarations without the §13 artifact.
Sequencing, and why it is not the obvious one
D1 first. It is the only one whose failure is unrecoverable. A cluster with no backup is one bad afternoon from data loss that no amount of isolation evidence compensates for.
D2 before D3, necessarily. Tenancy Posture §13.4: an artifact must assert something achievable. The noisy-neighbour artifact requires showing the governance controls bind. With no controls there is nothing to bind, so a probe written now could only demonstrate degradation — an artifact that "can only fail, or that passes by being run gently enough". Writing the probe first would produce an overclaim wearing the costume of evidence.
D1 also depends on a backup target, which is resource-control's bucket and
the platform-pg-backup-s3 credential lane — the same handoff
make postgres-backup-deploy waits on. Check whether that is now live before
assuming this is blocked.
Tasks
id: RPF-WP-0019-T01
status: todo
priority: high
Establish a backup target for apps-pg. Confirm the state of the
resource-control bucket and the platform-pg-backup-s3 OpenBao Secret; if
live, configure barmanObjectStore and a retentionPolicy on the cluster. If
not live, record the dependency and say so — do not leave the absence
undocumented a second time.
id: RPF-WP-0019-T02
status: todo
priority: high
Declare and enforce per-consumer controls. Per-consumer connection
allowance, statement_timeout, idle_in_transaction_session_timeout, and
pod resource requests/limits to lift apps-pg off BestEffort QoS. Publish
every value in docs/s3-consumer-interfaces.md before it takes effect —
§10.2 is a disclosure rule, and applying a timeout consumers learn about by
hitting it would breach the rule while implementing it.
id: RPF-WP-0019-T03
status: todo
priority: medium
Declare the ceiling and overflow target. Owed under this repo's own
Rule P-4.1 before apps-pg's third consumer; it is at two. Name the binding
resource per Rule P-4.2 — memory or connections — and a named overflow
substrate per P-4.3.
id: RPF-WP-0019-T04
status: todo
priority: medium
Isolation probes, after T02. Consumer-boundary probes on the
rapp-postgres model, then the §13 noisy-neighbour artifact: per-consumer
baseline, saturation run, evidence the controls bind, degradation measured and
judged against each consumer's declared service class. Update
docs/placement-policy.md §3.1 and docs/tenancy-posture.md when the P1
claims become evidenced.
Boundaries
apps-pgonly.platform-pg's equivalents arerapp-postgres's.- No consumer is migrated. This changes the cluster, not who is on it.
- Values are published before they are enforced, never after.
Risks
Applying limits to a live cluster breaks a consumer that was relying on
their absence. Most likely with statement_timeout. Mitigation is the
disclosure-first ordering in T02, which gives consumers a window to object.
T01 stays blocked on a handoff outside this repo and the cluster keeps no backup meanwhile. Mitigation is that T01 requires the dependency be recorded explicitly rather than left as a silent absence — which is exactly how D1 survived this long.