--- id: ADR-008 type: architecture-decision-record title: "Tenancy Posture: Five Planes, Graduated Levels, Declared Conformance" status: proposed decided_by: Bernd Worsch date: "2026-08-17" revision: "draft-3" tags: ["architecture", "multi-tenancy", "isolation", "placement", "retention", "maturity", "tenant-engine", "flex-auth", "rapp-postgres", "scaling"] --- # ADR-008: Tenancy Posture — Five Planes, Graduated Levels, Declared Conformance ## Status **Proposed, draft-3.** - **draft-1** proposed a single model with fixed characteristics. Rejected: it could not describe a repo that is not there yet. - **draft-2** reframed to graduated levels per plane. Externally corroborated (§16), but four of its statements were wrong and one thing it needed was missing. - **draft-3** applies those corrections, adds the retention plane, and records an adoption stance. It is informed by four external research digests, one per original plane, in `research/2026-08-17-adr008-*`. Reviewed by nobody yet. §19 lists what each owner is being asked to accept. ## 1. Context The estate has been building multi-tenancy for months and has never written down what it is building. Five documents each cover a slice: | Document | Covers | Status | |---|---|---| | `iam-profile_v0.3` (NetKingdom) | Tenant identifier shape, `tenant_roles` claim, staleness rules | Ratified | | `tenant-engine-boundary-contract_v0.1` (NetKingdom) | Who owns tenant records, roles, plan assignment | Ratified | | `business-app-service-contract_v0.1` §1 (Custodian) | Business apps: instance-per-client, tenant-keyed data | Ratified | | `rapp-postgres` ADR-0001 | Consumer + tenant isolation in PostgreSQL | Proposed, governs one repo | | `rapp-postgres` ADR-0002 | Per-consumer retention and the erasure horizon | Proposed, governs one repo | | `shared-platform-relational-storage_v0.1` | The stacked-boundary gap | Routed 2026-08-10, **still unratified** | Four failures follow. **The gap was diagnosed once and the fix stalled.** The v0.1 draft was written to fill this hole and has sat unratified in neither canon directory. §20 attaches a ratification path so this one does not join it. **Placement is owned by nobody.** `user-engine-pg` and `target-revenue-pg` are dedicated; `apps-pg`, `net-kingdom-pg`, `platform-pg`, `state-hub-db` and `forgejo-db` are shared. Both live, neither written down. `tenant-engine` raised this with `railiance-platform` on 2026-08-16; unanswered. **Two contradictory defaults are already ratified.** Business apps get instance-per-client; platform services pool. Nothing says which shape a new service takes, and no definition separates the categories. **There is no honest way to describe a repo that is not there yet.** The estate absorbs repos with weak or absent tenant separation. Today such a repo is simply non-conformant, leaving it two bad options: misrepresent its posture, or stay outside the framework. ## 2. What this document is **A framework, not a model.** It specifies no single correct implementation. It supplies terminology (§3, §4), a declaration (§5), a conformance rule (§6), methodology (§12), and evidence definitions (§13). A service is conformant when its declared posture is accurate and its trajectory recorded. A service is non-conformant when it claims a level it cannot evidence — regardless of how high or low that level is. ## 3. Five orthogonal planes "Is this multi-tenant?" is treated as one question. It is five, and they are independent: | Plane | Question | Vocabulary owner | |---|---|---| | **Identity (I)** | How is a tenant named and validated? | `tenant-engine` / IAM Profile | | **Authorization (A)** | How is a request bound to the tenants it may act for? | `flex-auth` | | **Enforcement (E)** | Where, mechanically, is the tenant boundary enforced? | This framework | | **Placement (P)** | Which substrate holds a tenant's data? | `railiance-platform` | | **Retention (R)** | How long does data persist, and how is it erased? | The storage platform; policy by the consumer | Conflation produces errors today. `rapp-postgres`'s `PostgresConsumer` carries `tenantIsolation: consumer-service-boundary` — an **E**-plane fact in a **P**-plane artifact, reading as though storage enforces something it does not. The "dedicated versus shared" argument mixes P (capacity, blast radius) with E (correctness). The planes are separated *precisely so each may sit at a different level*. **Decision 3.1:** every document, declaration and plan tier that says "isolation" MUST name which plane it means. **Decision 3.2:** the planes couple at their tops and the couplings MUST be stated where they apply, not used to argue the planes are one: - `E4` is reachable only at `P3` or above. - `R`'s erasure horizon is bounded below by `P` — on shared substrate, a consumer's horizon is the instance maximum (§4.5). **Decision 3.3 — scope.** The P and R ladders describe a service's **primary datastore**. Caches, search indices, message queues and background jobs are named leak surfaces in the external baselines and are assessed separately, not covered by a posture vector. Saying so is honest; implying the vector covers them would not be. ## 4. Graduated levels Each plane carries an ordered ladder. Higher is stronger, not better: the right level is the one a service can evidence and its risk warrants. ### 4.1 Identity (I) | Level | State | |---|---| | **I0** | No tenant concept. Data not attributable to a tenant. | | **I1** | A local tenant notion exists but is not canonical, **or** the tenant is taken from the request rather than from a verified token. | | **I2** | Canonical identifiers, bound at the identity provider and carried as a verified claim; `tenant-engine` is the source of existence. | | **I3** | I2 plus capability roles honoured, with live `tenant-engine` re-query for privileged, destructive, credential-vending or `aal2`-class decisions. | I1 now explicitly absorbs request-supplied tenant identifiers. "Never trust client-supplied tenant IDs without validation" is a named anti-pattern; a service reading the tenant from a header is at I1 however canonical the string. `business-app-service-contract` §2.1 sets app-local accounts as the v1 baseline for business apps — a sanctioned low level with recorded triggers for moving up. That is the pattern this framework generalises. ### 4.2 Authorization (A) | Level | State | |---|---| | **A0** | No authorization, or tenant context not carried. | | **A1** | Ad-hoc checks scattered through handlers. | | **A2** | A single local authorization boundary; tenant context bound once, centrally. | | **A3** | Decisions delegated to `flex-auth` as PDP, with live re-query where the IAM Profile requires it. | | **A4** | A3 over a **standard** PDP interface (OpenID AuthZEN Authorization API 1.0), so the decision point is swappable and the enforcement point is not coupled to one engine's request shape. | A4 is new. `flex-auth` uses a bespoke `CheckRequest` and a bespoke action vocabulary, with action strings copied verbatim between repos to avoid re-derivation — exactly the coupling AuthZEN removes. The specification reached Final in January 2026 and Keycloak shipped experimental support in May. We are not wrong, we are pre-standard, and the ladder should have somewhere to go. **Internal service-to-service calls are in scope for this plane.** "Skipping tenant validation for internal services" is a named anti-pattern, and our estate is mostly internal calls — `flex-auth` calls `tenant-engine` synchronously on the authorization path. A service identity acting on behalf of a tenant must carry and revalidate tenant context to claim A2 or above. ### 4.3 Enforcement (E) | Level | Mechanism | |---|---| | **E0** | None. Data not tenant-keyed; separation incidental or absent. | | **E1** | Data tenant-keyed, filtering applied per query at call sites. | | **E2** | Filtering centralised at a single service-side choke point binding authenticated identity to permitted tenants. | | **E3** | E2 **plus** platform-assisted filtering: row-level security keyed on a tenant GUC set transaction-locally, or an equivalent enforced data-access layer. | | **E4** | Structural: the credential a workload holds cannot address another tenant's data at all. Requires per-tenant credentials and per-tenant substrate. | **Correction from draft-2.** Draft-2 described E3 as something "the application cannot trivially route around". That is false and it was this document overclaiming in exactly the way §6 prohibits. Any session can re-issue `SET` on a custom GUC, so an attacker with SQL execution can reset the tenant and read across the boundary. What E3 buys is precise, and the ladder must say so: | Threat | E1 | E2 | E3 | E4 | |---|:--:|:--:|:--:|:--:| | A developer forgets a tenant predicate | ✗ | ✓ | ✓ | ✓ | | A new code path bypasses the choke point | ✗ | ✗ | ✓ | ✓ | | SQL injection reaching the connection | ✗ | ✗ | ✗ | ✓ | | The application process is compromised | ✗ | ✗ | ✗ | ✓ | E3 is a strong control against **accident** — the common case, and the one that causes real breaches — and no control at all against **compromise**. Only E4 holds against both, because the credential itself cannot address another tenant's data. **Correction: E3 layers on E2, it does not replace it.** External practice treats application-layer and database-layer filtering as complementary. A service that dropped its choke point on reaching E3 would be *worse* off, since E3 fails open under injection. Claiming E3 therefore requires the E2 evidence artifact as well. **Correction: the GUC is set transaction-locally.** Draft-2 said "at pool checkout", which is session scope and the wrong instrument. Under a pooler in statement mode, `SET` leaks between clients and returns other tenants' rows — a failure that appears only under production concurrency and produces no error. Use `SET LOCAL` inside an explicit transaction. **Platform enforcement is a platform obligation.** Reaching E3 requires the storage platform to *offer* the mechanism: provisioned policies, a documented GUC contract, and a probe. Where a consumer wants E3 and the platform has not supplied it, the gap is the platform's. §19.6 asks `rapp-postgres` to define that contract, which must carry `FORCE ROW LEVEL SECURITY` on every tenant table (without it the table owner bypasses policies silently, and ADR-0001 already established that our migration role owns the tables it creates), no `BYPASSRLS` on leased roles, `SECURITY INVOKER` for ordinary logic, and an `EXPLAIN` comparison because RLS disables functional indexes built on non-leakproof functions. **Default expectation** for a new platform service: E2 at first serve, E3 recorded as target. Services whose cross-tenant exposure would be a reportable breach SHOULD target E3 or above. ### 4.4 Placement (P) | Level | Shape | Live occupants | |---|---|---| | **P0** | Shares a database with another consumer. | None sanctioned; the state absorbed repos arrive in. | | **P1** | Database per consumer, shared cluster. | `audit-core`, `tenant-engine` on `platform-pg` | | **P2** | Dedicated cluster per consumer. | `user-engine-pg`, `target-revenue-pg` | | **P3** | Dedicated cluster per tenant. | Business apps per `business-app-service-contract` §1.2 | | **P4** | P3 plus separate region or jurisdiction. | None | **P0 → P1 → P2 does not raise the E level.** Those steps buy consumer isolation, capacity predictability, independent retention and a smaller operational blast radius. Only P3 makes E4 reachable. This is the most misusable fact in the framework and §11 governs how it may be described. **Decision 4.4.1:** P1 is the default for platform services; P3 for client-facing business apps, as already ratified. A service unsure which it is must resolve that first (§19.4). **Decision 4.4.2 — placement scopes to data substrate.** Identity-provider placement (realm-per-tenant versus Organizations) is the same silo/pool decision on a different substrate, is live in our estate, and is undecided. Realm-per-tenant carries a stated ceiling around 5–20 tenants, far below our target. Recorded here as a parallel question (§19.7), not folded into P. ### 4.5 Retention and erasure (R) New in draft-3. Implemented abstractly by the storage platform for any dataset; policy is built on top of that interface by the consumer or its governance layer. Reference implementation: `rapp-postgres` ADR-0002. | Level | State | |---|---| | **R0** | No retention or deletion position. Data kept indefinitely by default; no deletion path exists. | | **R1** | Platform default retention applies (N=30 days). The consumer has declared no requirement. | | **R2** | Retention declared as N days per dataset; the **erasure horizon** is published, and the consumer makes no promise shorter than it. | | **R3** | Policy-driven deletion: the consumer or its governance layer declares what is due, the platform sweeps whole datasets on that instruction and evidences each run. | | **R4** | Verified erasure: deletion proven complete across live data, backups and derived copies within the published horizon. | Three properties. **The erasure horizon is the interval between deleting data and it ceasing to be recoverable from anything the platform holds.** Deleting a row does not remove it from yesterday's backup. With an N-day window, deleted data remains recoverable for N days. That is the difference between "deleted" and "erased" and the estate had never written it down. **On shared substrate, retention is not per-consumer.** Physical backup is instance-wide — one WAL stream, one window — so the instance retention is *derived* as the maximum across co-resident consumers, and every consumer's horizon is that maximum. A consumer declaring 7 days beside one declaring 90 gets 90. This is the retention analogue of ADR-0001's blast-radius disclosure: state the coupling rather than imply an isolation that is not there. **Retention is therefore a placement trigger.** A consumer needing a horizon shorter than the instance floor cannot have one at P1. It moves to P2 for a reason with nothing to do with performance — which is exactly why it needs recording, since nobody looks for a retention argument when reviewing placement. Deletion splits mechanism from policy. The platform deletes whole **datasets** on instruction and records an opaque policy reference it never interprets, so every deletion traces to what authorised it. Rows are not a dataset: row expiry is the consumer's own DML under its migration lease. Dropping a consumer's whole database is an operator-gated offboarding step, never a scheduled one. ## 5. The posture vector A service states one level per plane, plus a target, a date, and any placement exceptions: ```yaml tenancy: current: { I: 2, A: 3, E: 2, P: 1, R: 1 } target: { I: 2, A: 3, E: 3, P: 1, R: 2 } reviewed: "2026-08-17" gap: E: "Choke point exists and is tested; RLS not provisioned. Blocked on rapp-postgres publishing the GUC contract. Target Q4." R: "Retention declared; erasure horizon not yet published to consumers." ``` **Placement exceptions.** Draft-2 assigned one P level per service, which cannot express the vertically partitioned model — most tenants pooled, some dedicated — that §11's isolation tiers require. A tier requiring `P2` bought by three tenants would put the service at two levels at once, forcing an over- or under-claim. Placement is therefore declared as a default plus exceptions: ```yaml placement_exceptions: - tenants: ["tenant:enterprise:*"] P: 3 reason: "isolation tier; see adaptive-pricing tier definition" ``` A service with exceptions must be able to say which tenants are on which substrate. That mapping is a first-class artifact, not archaeology. Worked examples, best-effort and subject to owner correction: | Service | Current | Notes | |---|---|---| | `tenant-engine` | `I2 A3 E2 P1 R1` | Moving to P1 under TEN-WP-0009; retention declared, horizon not yet published. | | `audit-core` | `I2 A3 E2 P1 R1` | Holds audit evidence, so both E3 and R2 are urgent targets. | | A newly absorbed repo | `I1 A1 E1 P0 R0` | Conformant **if declared**, with a recorded path. | **Decision 5.1:** the posture vector is declared in the repo, not in the hub, consistent with local-files-are-source-of-truth. ## 6. Conformance is accuracy, not altitude > **A service is conformant when its declared posture is accurate, its target > is recorded, and it does not claim a level it cannot evidence. It is > non-conformant when it overclaims — at any altitude.** - Declaring `E0` is conformant. Concealing `E0` is not. - A repo may be absorbed at any posture. It may not be absorbed silently. - No service is blocked from the estate for being low on a ladder. Services MAY be blocked from *specific work* — serving a tenant grouping, holding a data class, carrying a plan tier — by requirements expressed as minimum levels. - Downgrading is permitted and must be declared. A regression found by guarding is a defect; a regression declared in advance is a decision. Without the plane separation, "not rigorous about tenant separation" is one verdict a repo passes or fails. With it, the same repo is `I1 A1 E1 P0 R0` with a path — a plan, not an indictment. ## 7. Portability across placement levels Movement between P levels must be operational, not a rebuild: - Connect by injected credential only — no cluster, host, namespace or database name in source. - Own a whole database, never tables inside someone else's. - Idempotent schema creation. - No cross-database joins or co-location assumptions. **Decision 7.1:** mandatory at P1 and above. At P3, SHOULD rather than MUST — a per-client instance that never moves is not misconformant for naming its own database. ## 8. Placement triggers Recorded at provisioning time: noisy neighbour on a latency-critical path; a compliance or residency requirement; a plan tier requiring a higher minimum; an erasure horizon that no longer fits (§4.5); connection or memory ceiling reached. **Decision 8.1:** triggers MUST be *monitored*, not merely recorded. A trigger in a YAML comment nobody re-reads is documentation, not control. **Decision 8.2:** placement policy ownership is proposed to `railiance-platform`, **co-signed by `adaptive-pricing`**. Tenancy model selection is a commercial decision as much as a technical one; an operations-shaped repo should not hold it alone. ## 9. Credentials as a tenancy control Short-lived leased credentials re-read at connection checkout, with overlap-first rotation, bound the residual risk at every E level below E4: a leaked credential expires rather than persisting. Stronger than the industry norm of a long-lived per-service secret. **Decision 9.1:** static long-lived database credentials are not a sanctioned path for any service above E0. ## 10. Blast radius must be published **Decision 10.1:** every platform holding consumer data MUST publish, in concrete terms, what a leaked runtime credential can and cannot reach at the levels it operates. `rapp-postgres` ADR-0001 §5 is the reference. Where the model cannot provide a guarantee, the platform says so and names the escalation. ## 11. Commercial expression - **11.1** Plan tiers are expressed *internally* as minimum levels. A tier may require `E3 P2 R2`; it need not print that anywhere customer-facing. - **11.2** Marketing and product language is free. No requirement to expose level labels or this document. "Dedicated infrastructure", "isolated tenancy", "private instance" all remain available. - **11.3** The constraint is on **evidence, not vocabulary**. A customer-facing isolation, availability or retention claim must map to a minimum level the delivering service actually holds, recorded once when the tier is defined. The review is internal and happens at tier definition — not per campaign. - **11.4** Two hard lines, because these reach contracts and compliance questionnaires: - A claim that another tenant **cannot** reach the customer's data requires **E4**. - A claim that deleted data **is gone** requires **R4**, or an erasure horizon disclosed alongside it. ## 12. Methodology — analyze, establish, improve, guard **Analyze.** Assess a repo against the ladders; produce `tenancy.current` with reasoning recorded. Applies to new and absorbed services alike. **Establish.** Declare the target and gap. The target is set by data class, tenant groupings served and plan tiers carried — not by ambition. **Improve.** Move one plane at a time. Raising P while leaving E untouched is the characteristic misstep. **Guard.** Verify continuously that the declared posture holds — **against the service's own declaration**, not a universal maximum. Nobody must prove every service is at E4; the check is that none is below what it declared. Regression found by guarding is a defect; regression declared in advance is a decision. The estate has been bitten twice by silent pin rollbacks producing ordinary-looking 403s and 404s rather than errors. Posture regression looks the same — an RLS context leak returns correct-looking rows for the wrong tenant. Guarding must be designed for invisible failure, not for crashes. ## 13. Evidence per level **Decision 13.1:** a level is claimed only with its evidence artifact present. This turns §6's accuracy rule from an honour system into a check. **Decision 13.2 — evidence is of two kinds, and conflating them is an overclaim.** *Mechanical* evidence is a structural assertion a machine can make and belongs in CI. *Adversarial* evidence is semantic, requires setting up separate tenant contexts and comparing responses, and carries a review date rather than a green build. Cross-tenant findings are the category external testing practice identifies as needing human review. **A passing CI run is not E2 evidence.** | Level | Evidence | Kind | |---|---|---| | **I2** | Identifiers validated against the vocabulary; rejection test for a malformed id; binding shown to come from a verified token | Mechanical | | **I3** | Live re-query demonstrated on an `aal2`-class path; cached-claim path shown unused there | Mechanical | | **A2** | Choke point identified; test that an unbound request is refused | Mechanical | | **A3** | Live decision with a denial observed at the endpoint, not only at the decision surface | Mechanical | | **A4** | Decision served over the standard interface; a second PDP substituted without PEP change | Mechanical | | **E1** | Every tenant-owned table carries the tenant key | Mechanical | | **E2** | Choke point identified; identity bound to tenant A demonstrably cannot read tenant B | **Adversarial**, with a review date | | **E3** | `FORCE ROW LEVEL SECURITY` on every tenant table; no `BYPASSRLS` on leased roles; probe that a session without the GUC reads nothing; probe that a wrong GUC reads nothing; `EXPLAIN` comparison | Mechanical | | **E4** | Per-tenant credential demonstrated unable to connect to another tenant's substrate | Mechanical | | **P1–P4** | Provisioning declaration plus the platform's isolation probes | Mechanical | | **P1–P2 (noisy neighbour)** | One consumer saturating its connection or CPU allowance demonstrably does not breach another's | **Adversarial**, load-generated | | **R2** | Declared retention rendered; erasure horizon published and reported in the operator surface | Mechanical | | **R3** | Sweep evidence records: timestamp, dataset, identifiers removed, authorising policy reference | Mechanical | | **R4** | Erasure demonstrated across live data, backups and derived copies within the horizon | **Adversarial** | **Decision 13.3:** the E2, E3 and noisy-neighbour artifacts do not exist anywhere in the estate today. `rapp-postgres` runs 15 adversarial probes, all against the *consumer* boundary, none against the tenant boundary inside a consumer. Externally, what this framework calls a tenant boundary failure is **Broken Object Level Authorization** — OWASP API1, top of the API Security Top 10 since that list launched, and the most commonly exploited API vulnerability in published assessments. We have no coverage for the highest-ranked risk in our class of system. §19.3 seeks an owner. ## 14. Adoption stance — structure, not tooling **Decision 14.1:** external research is design input. This estate adopts published standards and structural patterns; it does not adopt tooling unless that tooling is an established industry standard with broad application. Everything else is built ground-up, so it can be optimised and refactored as the estate sees fit. | Class | Stance | |---|---| | Security baselines (OWASP Multi-Tenant Security Cheat Sheet, API Security Top 10) | Adopt as the external reference our ladders answer to | | Standards bodies (OpenID AuthZEN 1.0) | Adopt — this is what A4 is | | Reference taxonomies (Azure tenancy models, AWS SaaS Lens, cell architecture) | Adopt as structure | | Engine behaviour (PostgreSQL RLS mechanics) | Facts, not tooling | | Third-party analyzers and test frameworks | **Do not adopt.** Take their rule taxonomies as checklists for probes we write ourselves | The practical effect is small and good: `rapp-postgres` already owns a ground-up probe harness — bash and psql, no dependency tree — that found four real defects in its own provisioning SQL. The evidence artifacts in §13 become new probes in a tool we control. One idea worth reimplementing from the external survey is **policy-diff classification**: labelling a change to an enforcement policy as safe or breaking *before* it lands. ## 15. Alternatives considered **One fixed model with a single set of characteristics** (draft-1). *Rejected:* cannot describe a repo that is not there yet, forcing absorbed repos to misrepresent their posture or stay outside. A framework that can only describe its own end state is not a framework. **A maturity model with a single overall level.** *Rejected:* collapses the plane separation. A service strong on identity and weak on enforcement has a specific, actionable gap; one composite score hides it and invites averaging. **Prohibiting row-level security** (draft-2's inherited position). *Rejected in draft-2, refined in draft-3:* RLS is a real rung against the common threat. The error was never RLS — it was describing E3 in E4's language. **Schema-per-consumer in one database.** *Rejected:* `pg_catalog` is readable per-database, so every co-resident enumerates every other's table and column names regardless of grants. Retained as a describable state, never a target. **Mandating E4 for everyone.** *Rejected:* the tenant taxonomy includes `consumer` (private individuals) and `family`. A cluster per private individual is economically impossible; the taxonomy is itself evidence pooling is required. **Per-consumer physical backup retention.** *Rejected:* CNPG retention is a property of the instance's WAL archive. There is no mechanism, and claiming it would be a fabricated guarantee. Hence the derived maximum in §4.5. **Platform-scheduled row expiry.** *Rejected:* requires the platform to hold DML authority over consumer schemas and interpret consumer data semantics, both forbidden by ADR-0001. The consumer's migration lease is the correct instrument. **Leaving each repo to its own model.** *Rejected:* the status quo, which produced two contradictory ratified defaults and an unowned placement question. ## 16. Held against outside practice **The graduated reframe is corroborated, not invented here.** Microsoft's tenancy-model guidance states it almost verbatim: *"Instead of viewing isolation as a discrete property, consider it a spectrum. You can deploy components of your architecture that are more isolated or less isolated than other components in the same architecture."* The same guidance derives our E↔P coupling independently — shared deployment means enforcement lives in application code; dedicated deployment means it is structural. **Stronger than typical.** Most multi-tenancy literature models one boundary, tenant-to-tenant. This estate has **two stacked boundaries**: platform-service to platform-service, and tenant to tenant inside a consumer. Naming them separately and refusing to enforce both with one mechanism is uncommon and correct. Graduated per-plane levels also beat the silo/pool/bridge trichotomy, which is approximately our P plane with the other four missing — which is why it cannot express "pooled infrastructure, structurally enforced boundary". **Weaker than typical.** The pool model's standard mitigation is a *verified* enforcement layer every service is demonstrably routed through. We have the concept and none of the verification (§13.3). **Adopted without naming it.** Short-lived leased credentials re-read at checkout beat the long-lived-secret norm. §9 promotes it to a tenancy control. **Still unexplored.** Neither P nor R describes a **cell** — a slice of infrastructure with a *fixed maximum size*, sized so one cell's failure is survivable and cell count scales linearly. `platform-pg` is, in these terms, an uncapped cell: §17 computes a ceiling and nothing enforces it (§19.8). Sources: the four research digests in `research/2026-08-17-adr008-*`, which carry full citations for every claim in this section. ## 17. Scaling demands Measured against the live `platform-pg` specification, not estimated. ``` instances: 1 (no HA; single-node rail) max_connections: 100 memory limit: 1Gi per consumer: 14 connections (12 runtime + 2 migration) ``` **Connection ceiling: roughly six consumers.** Seven request 98 of 100 before CNPG's instance manager, metrics exporter and reserved slots. Two are provisioned. We are at roughly a third of capacity and the third request will not feel like a scaling event. **Memory likely binds first.** 100 backends against 1Gi is ~10MB per backend. Connection exhaustion errors clearly; memory pressure OOM-kills and degrades every co-resident at once. **E3 and pooling.** *Corrected from draft-2, which had this backwards.* Transaction-scoped context (`SET LOCAL` inside an explicit transaction) is what makes E3 **safe** under a pooler. Statement-level pooling is what breaks it, serving other tenants' rows under concurrency with no error. E3 constrains which pooling mode is available, not whether pooling is available. **Retention consumes the volume.** WAL accumulates with the window, and §4.5 makes the window the maximum across consumers. A consumer declaring a long retention extends everyone's horizon *and* everyone's storage draw against a 20Gi volume. **Restore time couples all consumers.** Physical backup is instance-wide, so a consumer's RTO is a function of *total* instance size, not its own. **No P1 tenant has HA.** `instances: 1` means a tier promising uptime cannot be satisfied at P1 as built — an availability floor belongs in §11's minimum-level vocabulary alongside isolation. ## 18. Consequences - The estate gains one vocabulary and a way to be honest about partial adoption. - Absorbed repos get a described state and a path instead of a failing grade. - `tenantIsolation` in `PostgresConsumer` is revealed as a mislabelled field. - The verification problem becomes tractable: guard against declaration. - Draft-2's RLS prohibition is reversed and its E3 description corrected; `rapp-postgres` acquires an obligation to define and offer the mechanism. - Adding a consumer with long retention **silently extends everyone's erasure horizon**. This must reach the consumer review checklist, not only this document. - A service selling an isolation tier must maintain a tenant→substrate mapping it does not have today. - Nothing here changes a running system. ## 19. Open questions 1. **`tenantIsolation` field** — `rapp-postgres`: rename to name its plane and carry a level (`tenancy.E: 2`), or move it out of the storage declaration. 2. **Placement ownership** — `railiance-platform` with `adaptive-pricing`: accept the ladder, triggers and the §8.1 monitoring obligation; appoint a recorded placement owner per workload. 3. **E2, E3 and noisy-neighbour evidence** — *owner needed.* Now three artifacts of two kinds: E3 and noisy-neighbour are buildable as probes in the existing harness; E2 is adversarial and needs a review cadence. Both `audit-core` and `tenant-engine` have declined fleet-scope work on correct boundary reasoning, so this needs appointing. Highest-severity gap. 4. **Business app vs platform service** — Custodian canon: a classification rule. Candidate: reuse `repo-classification-standard_v1.0`. 5. **Tier → minimum level mapping** — `adaptive-pricing` and `tenant-engine`: required only for tiers making isolation, availability or retention claims. 6. **The E3 mechanism** — `rapp-postgres`: publish the GUC contract with the `FORCE`/`BYPASSRLS`/`SECURITY INVOKER`/`EXPLAIN` requirements in §4.3. 7. **Identity-provider placement** — owner of `key-cape`: realm-per-tenant or Organizations? Realm-per-tenant's ~5–20 tenant ceiling is below our target. 8. **Cell sizing** — reframed from "should we adopt cells" to **"what is `platform-pg`'s declared maximum size, and what is the overflow target?"** The connection ceiling forces this whether or not we adopt the vocabulary. 9. **Retention floor and ceiling** — should `backupRetentionDays` have a platform minimum (so a consumer asking for 1 day gets a validation error rather than a quiet disappointment) and a maximum (so nobody exhausts the volume)? 10. **Engine neutrality** — the P ladder rests on a PostgreSQL property. State it engine-specifically and say so, or abstract it and risk a non-Postgres implementation that silently differs? **Routed elsewhere, deliberately.** The tenant identifier `tenant::` embeds headcount bands (`small`, `medium`, `large`) that change as a tenant grows, contradicting the consensus that identifiers should not encode mutable attributes. That is a critique of ADR-0013, not of this framework, and belongs to `tenant-engine` and NetKingdom canon. Folding it in here would overreach. ## 20. Ratification path 1. Reviewed by `tenant-engine`, `flex-auth`, `rapp-postgres`, `railiance-platform` and `adaptive-pricing` against §19. 2. Each publishes its own posture vector (§5) as part of review. **The framework is validated by whether it can describe them accurately** — if a repo cannot express itself in these five ladders, the ladders are wrong and this document changes, not the repo. 3. On acceptance, **supersedes** the routing of `rapp-postgres/docs/canon-drafts/shared-platform-relational-storage_v0.1-draft.md`, whose §§3–8 are absorbed here. That draft is withdrawn rather than left pending. 4. On acceptance, `rapp-postgres` ADR-0001 and ADR-0002 move to `accepted` and are annotated as the PostgreSQL implementation of the E, P and R ladders.