RPF-WP-0018: mark unevidenced P1 claims; seed RPF-WP-0019 for apps-pg
Placement policy 3.1 records that the P1 levels for vergabe and coulomb_social are provisioning declarations without the SS13 isolation probes - apps-pg has none. 3.2 records that apps-pg cannot produce the noisy-neighbour artifact at all until per-consumer controls exist, per SS13.4: with nothing to bind, a probe could only demonstrate degradation. RPF-WP-0019 seeded for the three live defects the disclosure work found: no backup, no per-consumer controls, no isolation probes. Kept separate from RPF-WP-0018 so declaring is not confused with fixing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
cc4e659a9e
commit
b15215d191
2 changed files with 174 additions and 9 deletions
|
|
@ -70,15 +70,50 @@ has an answer that is not archaeology. The **placement owner** decides
|
|||
placement changes for that workload and is accountable for watching its
|
||||
triggers. It is not the same role as the operator of the cluster.
|
||||
|
||||
| Workload | Substrate | Provided level | Placement owner | Co-signer |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| `audit-core` | `platform-pg` | P1 | `audit-core` | this repo |
|
||||
| `tenant-engine` | `platform-pg` | P1 | `tenant-engine` | this repo |
|
||||
| `user-engine` | `user-engine-pg` | P2 | `net-kingdom` | this repo |
|
||||
| `target-revenue` | `target-revenue-pg` | P2 | `target-revenue` | this repo |
|
||||
| `forgejo` | `forgejo-db` | P2 | this repo | — |
|
||||
| `vergabe` | `apps-pg` | P1 | `vergabe-teilnahme` | this repo |
|
||||
| `coulomb_social` | `apps-pg` | P1 | `coulomb-social` | this repo |
|
||||
| Workload | Substrate | Provided level | Placement owner | Co-signer | §13 evidence |
|
||||
| --- | --- | --- | --- | --- | --- |
|
||||
| `audit-core` | `platform-pg` | P1 | `audit-core` | this repo | probes exist |
|
||||
| `tenant-engine` | `platform-pg` | P1 | `tenant-engine` | this repo | probes exist |
|
||||
| `user-engine` | `user-engine-pg` | P2 | `net-kingdom` | this repo | not assessed here |
|
||||
| `target-revenue` | `target-revenue-pg` | P2 | `target-revenue` | this repo | not assessed here |
|
||||
| `forgejo` | `forgejo-db` | P2 | this repo | — | single consumer |
|
||||
| `vergabe` | `apps-pg` | P1 | `vergabe-teilnahme` | this repo | **absent** |
|
||||
| `coulomb_social` | `apps-pg` | P1 | `coulomb-social` | this repo | **absent** |
|
||||
|
||||
### 3.1 Two of these P1 levels are asserted without evidence
|
||||
|
||||
§13.1 claims a level only with its artifact present. The P1–P4 artifact is
|
||||
*"provisioning declaration plus the platform's isolation probes"*.
|
||||
|
||||
`platform-pg` has them — `rapp-postgres` runs 15 adversarial probes against the
|
||||
consumer boundary. **`apps-pg` has none.** `scripts/capture-apps-pg-evidence.py`
|
||||
captures *capacity* evidence for `resource-control`; it makes no isolation
|
||||
assertion. So the P1 levels recorded above for `vergabe` and `coulomb_social`
|
||||
are, as of 2026-08-17, **provisioning declarations without probes** — stated
|
||||
here rather than quietly carried, because a table of levels that mixes
|
||||
evidenced and unevidenced entries without saying which is which is the failure
|
||||
§13.1 exists to prevent.
|
||||
|
||||
Recorded as owed by this repo. Until the probes exist, those two rows should be
|
||||
read as *intended* placement, not *verified* placement.
|
||||
|
||||
### 3.2 `apps-pg` cannot produce the noisy-neighbour artifact at all
|
||||
|
||||
Worth separating from the gap above, because it is not a matter of writing a
|
||||
missing test. §13's P1–P2 noisy-neighbour artifact requires a per-consumer
|
||||
usage baseline, a saturation run, and **evidence that the governance controls
|
||||
bind — that the greedy consumer is held at its limits.**
|
||||
|
||||
`apps-pg` has no per-consumer governance controls. Its connection pool is not
|
||||
partitioned, there is no per-consumer connection allowance, no statement
|
||||
timeout and no resource limit on the pod. There is nothing that could be shown
|
||||
to bind, so a run against it could only demonstrate that a saturating consumer
|
||||
degrades its co-residents — which §13.4 explicitly rejects as an artifact that
|
||||
"can only fail, or that passes by being run gently enough".
|
||||
|
||||
The honest sequence is therefore: **controls first, artifact second.** Writing
|
||||
the probe before the limits exist would produce an overclaim wearing the
|
||||
costume of evidence. Tracked as `RPF-WP-0019`.
|
||||
|
||||
**The consuming repo is the placement owner; this repo co-signs.** The
|
||||
consumer knows its latency requirement, its erasure horizon and its compliance
|
||||
|
|
|
|||
130
workplans/RPF-WP-0019-apps-pg-recoverability-and-controls.md
Normal file
130
workplans/RPF-WP-0019-apps-pg-recoverability-and-controls.md
Normal file
|
|
@ -0,0 +1,130 @@
|
|||
---
|
||||
id: RPF-WP-0019
|
||||
type: workplan
|
||||
title: "apps-pg: backup, per-consumer controls, and the isolation probes they make possible"
|
||||
domain: financials
|
||||
repo: railiance-platform
|
||||
status: proposed
|
||||
owner: codex
|
||||
topic_slug: railiance
|
||||
created: "2026-08-17"
|
||||
updated: "2026-08-17"
|
||||
related:
|
||||
- RPF-WP-0018
|
||||
origin: residual
|
||||
origin_ref: RPF-WP-0018
|
||||
---
|
||||
|
||||
# RPF-WP-0019 — apps-pg recoverability and per-consumer controls
|
||||
|
||||
## Goal
|
||||
|
||||
Close the three defects `RPF-WP-0018` surfaced in `apps-pg` by writing its
|
||||
quota disclosure. Documentation found them; only this workplan fixes them.
|
||||
|
||||
## Why this is separate from RPF-WP-0018
|
||||
|
||||
That workplan declared posture and policy. This one changes a live cluster.
|
||||
Keeping them apart matters: a declaration workplan that quietly starts editing
|
||||
production is how "we wrote it down" becomes indistinguishable from "we fixed
|
||||
it". The declaration is published as-is, with the defects visible, and this is
|
||||
the record of closing them.
|
||||
|
||||
## The three defects
|
||||
|
||||
**D1 — `apps-pg` has no backup.** No `barmanObjectStore`, no `retentionPolicy`,
|
||||
nothing. This is not a short retention window; it is no recovery path at all,
|
||||
on a cluster holding two S5 application databases. Its R level is `R0` and
|
||||
`R0` here means unrecoverable, not merely un-erasable.
|
||||
|
||||
**D2 — no per-consumer controls.** The connection pool is unpartitioned, so
|
||||
one consumer can exhaust the cluster while staying politely inside its own
|
||||
expectations. No `statement_timeout`, no `idle_in_transaction_session_timeout`,
|
||||
no CPU or memory limits — the pod is BestEffort QoS and is the first thing
|
||||
evicted under node pressure.
|
||||
|
||||
**D3 — no isolation probes**, so the `P1` levels recorded for `vergabe` and
|
||||
`coulomb_social` in `docs/placement-policy.md` §3.1 are provisioning
|
||||
declarations without the §13 artifact.
|
||||
|
||||
## Sequencing, and why it is not the obvious one
|
||||
|
||||
**D1 first.** It is the only one whose failure is unrecoverable. A cluster
|
||||
with no backup is one bad afternoon from data loss that no amount of isolation
|
||||
evidence compensates for.
|
||||
|
||||
**D2 before D3, necessarily.** Tenancy Posture §13.4: an artifact must assert
|
||||
something achievable. The noisy-neighbour artifact requires showing the
|
||||
governance controls *bind*. With no controls there is nothing to bind, so a
|
||||
probe written now could only demonstrate degradation — an artifact that "can
|
||||
only fail, or that passes by being run gently enough". Writing the probe first
|
||||
would produce an overclaim wearing the costume of evidence.
|
||||
|
||||
D1 also depends on a backup target, which is `resource-control`'s bucket and
|
||||
the `platform-pg-backup-s3` credential lane — the same handoff
|
||||
`make postgres-backup-deploy` waits on. Check whether that is now live before
|
||||
assuming this is blocked.
|
||||
|
||||
## Tasks
|
||||
|
||||
```task
|
||||
id: RPF-WP-0019-T01
|
||||
status: todo
|
||||
priority: high
|
||||
```
|
||||
**Establish a backup target for `apps-pg`.** Confirm the state of the
|
||||
`resource-control` bucket and the `platform-pg-backup-s3` OpenBao Secret; if
|
||||
live, configure `barmanObjectStore` and a `retentionPolicy` on the cluster. If
|
||||
not live, record the dependency and say so — do not leave the absence
|
||||
undocumented a second time.
|
||||
|
||||
```task
|
||||
id: RPF-WP-0019-T02
|
||||
status: todo
|
||||
priority: high
|
||||
```
|
||||
**Declare and enforce per-consumer controls.** Per-consumer connection
|
||||
allowance, `statement_timeout`, `idle_in_transaction_session_timeout`, and
|
||||
pod resource requests/limits to lift `apps-pg` off BestEffort QoS. Publish
|
||||
every value in `docs/s3-consumer-interfaces.md` before it takes effect —
|
||||
§10.2 is a disclosure rule, and applying a timeout consumers learn about by
|
||||
hitting it would breach the rule while implementing it.
|
||||
|
||||
```task
|
||||
id: RPF-WP-0019-T03
|
||||
status: todo
|
||||
priority: medium
|
||||
```
|
||||
**Declare the ceiling and overflow target.** Owed under this repo's own
|
||||
Rule P-4.1 before `apps-pg`'s third consumer; it is at two. Name the binding
|
||||
resource per Rule P-4.2 — memory or connections — and a named overflow
|
||||
substrate per P-4.3.
|
||||
|
||||
```task
|
||||
id: RPF-WP-0019-T04
|
||||
status: todo
|
||||
priority: medium
|
||||
```
|
||||
**Isolation probes, after T02.** Consumer-boundary probes on the
|
||||
`rapp-postgres` model, then the §13 noisy-neighbour artifact: per-consumer
|
||||
baseline, saturation run, evidence the controls bind, degradation measured and
|
||||
judged against each consumer's declared service class. Update
|
||||
`docs/placement-policy.md` §3.1 and `docs/tenancy-posture.md` when the P1
|
||||
claims become evidenced.
|
||||
|
||||
## Boundaries
|
||||
|
||||
- `apps-pg` only. `platform-pg`'s equivalents are `rapp-postgres`'s.
|
||||
- No consumer is migrated. This changes the cluster, not who is on it.
|
||||
- Values are published before they are enforced, never after.
|
||||
|
||||
## Risks
|
||||
|
||||
**Applying limits to a live cluster breaks a consumer that was relying on
|
||||
their absence.** Most likely with `statement_timeout`. Mitigation is the
|
||||
disclosure-first ordering in T02, which gives consumers a window to object.
|
||||
|
||||
**T01 stays blocked on a handoff outside this repo and the cluster keeps no
|
||||
backup meanwhile.** Mitigation is that T01 requires the dependency be recorded
|
||||
explicitly rather than left as a silent absence — which is exactly how D1
|
||||
survived this long.
|
||||
Loading…
Add table
Add a link
Reference in a new issue