the-custodian/canon/architecture/adr-008-multi-tenancy-model.md
codex f8c3c1d8fd ADR-008 draft-2: reframe from fixed model to graduated framework
Draft-1 proposed one model of multi-tenancy with fixed characteristics.
Rejected: the estate needs a framework that can hold several situations at
once, including repos that do not separate tenants rigorously today and must
be developed toward doing so.

What changed:

- Every plane now carries an ordered ladder (I0-I3, A0-A3, E0-E4, P0-P4),
  not just placement. A service is a posture vector, not a verdict.
- D3 reversed. Draft-1 forbade RLS as a control that "looks like a database
  guarantee without being one". The observation was right, the conclusion
  wrong: RLS is E3, materially stronger than E2, and the actual error was
  calling E3 by E4's name. Platform enforcement is now the direction of
  travel and an obligation on the platform, not only on consumers.
- New governing rule: conformance is accuracy, not altitude. Declaring E0 is
  conformant; concealing it is not. Overclaiming is the only violation.
- Fixed a flaw of draft-1's own making: R0 "shared tables, tenant column" was
  an enforcement state mislabelled as placement. Removed; P ladder renumbered.
- Added methodology (analyze/establish/improve/guard) and per-level evidence
  artifacts. Guarding checks a service against its own declaration, which is
  what makes the verification problem tractable at all.
- D7 softened per operator: tiers carry minimum levels internally, marketing
  language stays free, the constraint is on evidence not vocabulary. One hard
  line kept: "another tenant cannot reach your data" requires E4.

Ratification now also tests the framework — each reviewing repo publishes its
posture vector, and if a repo cannot express itself in these ladders the
ladders are wrong, not the repo.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 12:20:49 +02:00

28 KiB
Raw Blame History

id type title status decided_by date revision tags
ADR-008 architecture-decision-record Multi-Tenancy Framework: Four Planes, Graduated Levels, Declared Posture proposed Bernd Worsch 2026-08-17 draft-2
architecture
multi-tenancy
isolation
placement
maturity
tenant-engine
flex-auth
rapp-postgres
scaling

ADR-008: Multi-Tenancy Framework — Four Planes, Graduated Levels, Declared Posture

Status

Proposed, draft-2. Draft-1 (2026-08-17) proposed a single model with fixed characteristics. It was rejected on the grounds that the estate does not need one model of multi-tenancy — it needs a framework able to describe, hold, and improve several situations at once, including repos that do not separate tenants rigorously today and must be developed toward doing so. Draft-2 is that reframe.

This document decides things currently owned by tenant-engine, flex-auth, rapp-postgres, adaptive-pricing, and railiance-platform. It must be reviewed by those owners before ratification; §18 lists what each is being asked to accept.

1. Context

The estate has been building multi-tenancy for months and has never written down what it is building. Five documents each cover a slice:

Document Covers Status
iam-profile_v0.3 (NetKingdom) Tenant identifier shape, tenant_roles claim, staleness rules Ratified
tenant-engine-boundary-contract_v0.1 (NetKingdom) Who owns tenant records, roles, plan assignment Ratified
business-app-service-contract_v0.1 §1 (Custodian) Business apps: instance-per-client, tenant-keyed data Ratified
rapp-postgres/docs/adr/ADR-0001 Consumer + tenant isolation in PostgreSQL Proposed, governs one repo
rapp-postgres/docs/canon-drafts/shared-platform-relational-storage_v0.1 The stacked-boundary gap Routed 2026-08-10, still unratified

Four failures follow.

The gap has been diagnosed once and the fix stalled. The shared-platform-relational-storage draft was written to fill this hole. It has sat unratified for a week and appears in neither canon directory. §19 attaches a ratification path so this one does not join it.

Placement is owned by nobody. user-engine-pg and target-revenue-pg are dedicated; apps-pg, net-kingdom-pg, platform-pg, state-hub-db and forgejo-db are shared. Both patterns are live, neither is written down, and each new service copies whichever neighbour it consulted. tenant-engine raised this with railiance-platform on 2026-08-16; unanswered.

Two contradictory tenancy defaults are already ratified. business-app-service-contract §1.2 makes instance-per-client the default for business apps. Platform services pool all tenants into one database. Nothing states which shape a new service takes, and no definition distinguishes the two categories.

There is no honest way to describe a repo that is not there yet. This is the failure draft-1 missed. The estate absorbs and will keep absorbing repos with weak or absent tenant separation — app-local accounts, no tenant column, ad-hoc filtering. Today such a repo is simply non-conformant, which gives it two bad options: misrepresent its posture, or stay outside the framework entirely. Both are worse than a framework that can say "this repo is at level 1, its target is level 3, here is the gap and the evidence."

2. What this document is

A framework, not a model. It does not specify one correct multi-tenancy implementation with a fixed set of characteristics. It supplies:

  • Terminology — four planes, and graduated levels within each (§3, §4).
  • A declaration — the posture vector, so any repo can state where it actually is (§5).
  • A conformance rule — accuracy, not altitude (§6).
  • Methodology — analyze, establish, improve, guard (§12).
  • Evidence definitions — what proves a claimed level (§13).

A service is conformant when its declared posture is accurate and its trajectory is recorded. A service is non-conformant when it claims a level it cannot evidence — regardless of how high or low that level is.

3. Four orthogonal planes

"Is this multi-tenant?" is treated as one question. It is four, and they are independent:

Plane Question Vocabulary owner
Identity (I) How is a tenant named and validated? tenant-engine / IAM Profile
Authorization (A) How is a request bound to the tenants it may act for? flex-auth
Enforcement (E) Where, mechanically, is the tenant boundary enforced? This framework
Placement (P) Which physical substrate holds a tenant's data? railiance-platform

Treating them as one produces concrete errors today. rapp-postgres's PostgresConsumer carries a field tenantIsolation: consumer-service-boundary — an E-plane fact recorded in a P-plane artifact, reading as though the storage platform enforces something it does not. The "dedicated versus shared" argument routinely mixes P (capacity, blast radius) with E (correctness).

The planes are separated precisely so that each may sit at a different level. A repo can be strong on identity and weak on enforcement, or the reverse, and the framework must state that without flattening it to a single verdict. That is the whole point, and draft-1 lost it by collapsing E into a single fixed rule.

Decision 3.1: every document, declaration, and plan tier that says "isolation" MUST name which plane it means.

Decision 3.2: the planes couple at their top levels — E4 is only reachable at P3 or above — and that coupling MUST be stated where it applies rather than used to argue the planes are one.

4. Graduated levels

Each plane carries an ordered ladder. Higher is stronger, not better: the right level is the one a service can evidence and its risk warrants.

4.1 Identity (I) — how the tenant is named

Level State
I0 No tenant concept. Data is not attributable to a tenant.
I1 A local tenant notion exists but does not conform to tenant:<grouping>:<name>.
I2 Canonical identifiers used, validated against the ratified vocabulary; tenant-engine is the source of existence.
I3 I2 plus capability roles (PLTF/IAM/VEN/CUS) honoured, with live tenant-engine re-query for privileged, destructive, credential-vending or aal2-class decisions rather than trusting the cached tenant_roles claim.

Note business-app-service-contract §2.1 sets app-local accounts as the v1 baseline for business apps — an explicitly sanctioned low level with recorded triggers for moving up. That is the pattern this framework generalises.

4.2 Authorization (A) — how a request is bound to a tenant

Level State
A0 No authorization, or tenant context not carried on the request.
A1 Ad-hoc checks scattered through handlers.
A2 A single local authorization boundary; tenant context bound once, centrally.
A3 Decisions delegated to flex-auth as PDP, with live re-query where the IAM Profile requires it.

4.3 Enforcement (E) — where the tenant boundary actually holds

This is the reframed core. Draft-1 asserted one rule for everyone; this is a ladder whose top is platform-enforced.

Level Mechanism Bypassed by
E0 None. Data not tenant-keyed; separation is incidental or absent. Anything.
E1 Data tenant-keyed, filtering applied per query at call sites. One missing predicate.
E2 Filtering centralised at a single service-side choke point that binds authenticated identity to permitted tenants. Code that bypasses the choke point.
E3 Platform-assisted: row-level security keyed on a session GUC set at pool checkout, or an equivalent data-access layer the application cannot trivially route around. A SECURITY DEFINER function, a missing policy, or a wrong GUC.
E4 Structural: the connection a workload holds cannot address another tenant's data at all. Requires per-tenant credentials and per-tenant substrate. Only a platform-level compromise.

Three properties of this ladder matter.

E3 is real and is not forbidden. Draft-1 rejected RLS on the grounds that a control depending on the consumer setting a session variable "looks like a database guarantee without being one". The observation is correct; the conclusion was wrong. E3 is materially stronger than E2 — it converts "every call site must remember" into "one checkout path must be right" — and it is the standard mitigation in pool-model systems. What was actually wrong was calling E3 by E4's name. The honesty requirement survives; the prohibition does not.

Higher enforcement is the direction of travel, not an optional extra. Services holding data whose cross-tenant exposure would be a reportable breach SHOULD target E3 or above. The framework's default expectation for a new platform service is E2 at first serve, E3 recorded as target.

Platform enforcement is a platform obligation, not only a consumer one. Draft-1 placed the whole burden on consumers permanently. It is now stated that reaching E3 requires the storage platform to offer the mechanism — provisioned RLS policies, a documented GUC contract, and a probe. rapp-postgres ADR-0001 §4 already commits to provisioning it on request; that offer becomes an obligation with a defined contract (§18.6). Where a consumer wants E3 and the platform has not supplied the mechanism, the gap is the platform's, not the consumer's.

4.4 Placement (P) — which substrate holds the data

Draft-1's R0 ("shared tables, tenant column") is removed: that was an E-plane state mislabelled as placement. Corrected ladder:

Level Shape Live occupants
P0 Shares a database with another consumer (schema or tables co-resident). None sanctioned; the state absorbed repos often arrive in.
P1 Database per consumer, shared cluster. audit-core, tenant-engine on platform-pg
P2 Dedicated cluster per consumer. user-engine-pg, target-revenue-pg
P3 Dedicated cluster per tenant. Business apps per business-app-service-contract §1.2
P4 P3 plus separate region or jurisdiction. None

P0 → P1 → P2 does not raise the E level. Those steps buy consumer isolation, capacity predictability, independent backup retention, and a smaller operational blast radius. They leave the tenant boundary exactly where the E level puts it. Only P3 makes E4 reachable. This is the single most misusable fact in the framework and §11 governs how it may be described commercially.

Decision 4.5: P1 is the default for platform services; P3 is the default for client-facing business apps, as already ratified. A service unsure which it is must resolve that before choosing (§18.4).

5. The posture vector

A service states its tenancy posture as one level per plane, plus a target and a date:

tenancy:
  current:  { I: 2, A: 3, E: 2, P: 1 }
  target:   { I: 2, A: 3, E: 3, P: 1 }
  reviewed: "2026-08-17"
  gap:
    E: "Choke point exists and is tested; RLS not yet provisioned. Blocked on
        rapp-postgres offering the GUC contract. Target Q4."

Worked examples, as best assessed today and subject to owner correction:

Service Current Notes
tenant-engine I2 A3 E2 P1 Moving to P1 under TEN-WP-0009; E2 via its own service boundary.
audit-core I2 A3 E2 P1 Same shape; holds audit data, so E3 is the more urgent target.
A newly absorbed repo I1 A1 E1 P0 Conformant if declared, with a recorded path.

The vector is what makes the framework hold both situations honestly. A repo at I1 A1 E1 P0 is not a failure — it is a described state with a direction. A repo claiming E3 with no RLS policy is a failure regardless of how good the rest of it is.

Decision 5.1: the posture vector is declared in the repo, not in the hub, consistent with local-files-are-source-of-truth.

6. Honest absorption — conformance is accuracy, not altitude

The governing rule of this framework:

A service is conformant when its declared posture is accurate, its target is recorded, and it does not claim a level it cannot evidence. It is non-conformant when it overclaims — at any altitude.

Consequences, stated deliberately:

  • Declaring E0 is conformant. Concealing E0 is not.
  • A repo may be absorbed at any posture. It may not be absorbed silently.
  • No service is blocked from the estate for being low on a ladder. Services MAY be blocked from specific work — serving a tenant grouping, holding a data class, carrying a plan tier — by requirements expressed as minimum levels (§11).
  • Downgrading is permitted and must be declared. A regression discovered by guarding (§12) is a defect; a regression declared in advance is a decision.

This is what the plane separation is for. Without it, "not rigorous about tenant separation" is a single verdict that a repo either passes or fails. With it, the same repo is I1 A1 E1 P0 with a path — which is a plan, not an indictment.

7. Portability across placement levels

Movement between P levels must be an operational change, not a rebuild. Adopted from tenant-engine's TEN-WP-0009 design:

  • Connect by injected credential only — no cluster, host, namespace, or database name in source.
  • Own a whole database, never tables inside someone else's. Sharing a cluster is reversible; sharing a database makes relocation a schema merge.
  • Idempotent schema creation.
  • No cross-database joins or co-location assumptions.

A service meeting these moves P1 → P2 by: create target, dump, restore, swap the credential reference, restart.

Decision 7.1: mandatory at P1 and above. At P3, SHOULD rather than MUST — the consolidation business-app-service-contract §1.3 preserves is a real benefit, but a per-client instance that never moves is not misconformant for naming its own database.

8. Placement triggers

A level change decided under pressure is decided badly. Each service records at provisioning time what would move it:

  • Noisy neighbour on a latency-critical path
  • A compliance or residency requirement
  • A plan tier requiring a higher minimum (§11)
  • Backup retention or RTO no longer fitting
  • Connection or memory ceiling reached (§16)

Decision 8.1: triggers MUST be monitored, not merely recorded. A trigger in a YAML comment that nobody re-reads is documentation, not control. The platform-pg ceiling in §16 is the first that requires an actual alert.

Decision 8.2: placement policy ownership is proposed to railiance-platform. This framework supplies the ladder, the triggers, and the declaration; it does not appoint per-workload placement. See §18.2 — this remains their call and their open request.

9. Credentials as a tenancy control

Short-lived leased credentials from OpenBao, re-read at connection checkout, with overlap-first rotation, bound the residual risk at every E level below E4: a leaked credential expires rather than persisting. This is stronger than the industry norm of a long-lived per-service secret and belongs in the tenancy model, not only in credential-management canon.

Decision 9.1: static long-lived database credentials are not a sanctioned path for any service above E0.

10. Blast radius must be published

Decision 10.1: every platform holding consumer data MUST publish, in concrete terms, what a leaked runtime credential can and cannot reach at the levels it operates. rapp-postgres ADR-0001 §5 is the reference implementation. Where the model cannot provide a guarantee, the platform says so and names the escalation, rather than describing a weaker control in stronger language.

11. Commercial expression — room to be successful

Isolation will be sold. The framework must let commercial and presentation work rather than obstruct them, while preventing a claim that cannot be evidenced.

Decision 11.1: plan tiers are expressed internally as minimum levels, not as implementation labels. A tier may require E3 P2; it need not print that string anywhere customer-facing.

Decision 11.2: marketing and product language is free. There is no requirement to expose level labels, ladder names, or this document to customers. "Dedicated infrastructure", "isolated tenancy", "private instance" all remain available.

Decision 11.3: the constraint is on evidence, not vocabulary. Any customer-facing isolation or availability claim must map to a minimum level the delivering service actually holds, recorded once when the tier is defined. The review is internal and happens at tier definition — not per campaign, per page, or per conversation.

Decision 11.4: the one hard line — a claim that another tenant cannot reach the customer's data requires E4. Everything softer than that absolute has room. This is the claim that reaches compliance questionnaires and contracts, and it is the one that must be true.

adaptive-pricing owns plan terms; tenant-engine owns which plan a tenant holds. Neither currently maps a tier to minimum levels; §18.5 asks them to, for tiers making isolation or availability claims only.

12. Methodology — analyze, establish, improve, guard

The framework is operated as a cycle, not consulted as a reference.

Analyze. Assess a repo against the four ladders and produce its current vector. Output: the tenancy.current block with the reasoning recorded. Applies to new services and absorbed ones alike.

Establish. Declare the target vector and the gap. The target is set by the data class held, the tenant groupings served, and any plan tier carried — not by ambition. Output: tenancy.target plus a recorded gap per plane.

Improve. Move one plane at a time. Cross-plane moves are where mistakes happen — raising P while leaving E untouched is exactly the misstep §4.4 warns about. Each step is ordinary repo work with ordinary evidence.

Guard. Verify continuously that the declared posture still holds. Guarding checks a service against its own declaration, not against a universal maximum. This is what makes the verification problem tractable: nobody has to prove every service is at E4; the check is that nobody is below what they declared, and nobody claims what they cannot evidence.

Regression found by guarding is a defect; regression declared in advance is a decision. The estate has been bitten twice by silent pin rollbacks that produced ordinary-looking 403s and 404s rather than errors. Posture regression will look the same, so guarding must be designed for invisible failure rather than for crashes.

13. Evidence per level

Guarding requires each level to have a defined evidence artifact. Proposed:

Level Evidence
I2 Identifiers validated against the ratified vocabulary; rejection test for a malformed tenant id.
I3 Live re-query demonstrated on an aal2-class path; cached-claim path shown not to be used there.
A2 Single choke point identified; test that an unbound request is refused.
A3 Live decision against flex-auth with a denial observed at the endpoint, not only at the decision surface.
E1 Every tenant-owned table carries the tenant key. Schema-level check, mechanical.
E2 Choke point identified; adversarial test that an identity bound to tenant A cannot read tenant B.
E3 RLS policies present on every tenant-owned table; probe that a session without the GUC set reads nothing; probe that a wrong GUC reads nothing.
E4 Per-tenant credential demonstrated unable to connect to another tenant's substrate.
P1P4 Provisioning declaration plus the platform's own isolation probes.

Decision 13.1: a level is claimed only with its evidence artifact present. This turns §6's accuracy rule from an honour system into a check.

Decision 13.2: the E2 and E3 evidence artifacts do not exist anywhere in the estate today. rapp-postgres runs 15 adversarial probes, all against the consumer boundary, none against the tenant boundary inside a consumer. This is the framework's largest live gap and §18.3 seeks an owner.

14. Alternatives considered

One fixed model with a single set of characteristics (draft-1). Rejected by the operator: it cannot describe a repo that is not there yet, forcing absorbed repos either to misrepresent their posture or to stay outside the framework. A framework that can only describe its own end state is not a framework.

A maturity model with a single overall level (bronze/silver/gold). Rejected: collapses the plane separation that makes the framework useful. A service strong on identity and weak on enforcement has a specific, actionable gap; one composite score hides it and invites averaging.

Schema-per-consumer in one database (sanctioning P0). Rejected: pg_catalog is readable per-database, so every co-resident enumerates every other's table and column names regardless of GRANTs. Retained in the ladder as a describable state, never as a sanctioned target.

Row-level security as the consumer boundary. Rejected: puts a correctness-critical boundary between two services inside query predicates. Retained as E3 for the tenant boundary, where the realistic alternative is application code rather than a database.

Mandating E4 for everyone. Rejected: the tenant grouping taxonomy includes consumer (private individuals) and family. A cluster per private individual is economically impossible; the taxonomy is itself evidence that pooling is required.

Leaving each repo to its own model. Rejected: this is the status quo, and it produced two contradictory ratified defaults plus an unowned placement question.

15. Challenge against outside practice

Where the framework is stronger than typical. Most SaaS multi-tenancy literature (AWS's silo/pool/bridge framing among them) models exactly one boundary — tenant-to-tenant. This estate has two stacked boundaries: platform-service-to-platform-service, and tenant-to-tenant inside a consumer. Naming them separately and refusing to enforce both with one mechanism is uncommon and correct.

Graduated per-plane levels are also stronger than the usual silo/pool/bridge trichotomy, which describes deployment shape only. Silo/pool/bridge is approximately this framework's P plane with the other three missing — which is why it cannot express "pooled infrastructure, structurally enforced boundary", or the reverse.

Where it is weaker, and this is the finding that matters. The pool model's standard mitigation is a verified enforcement layer that every service is demonstrably routed through. The estate has the concept (E2/E3) and none of the verification (§13.2). A single missing WHERE tenant_id = ? in any consumer is a cross-tenant breach that no platform control catches, no probe fails, and no log shows as an error.

Where the estate has quietly adopted a good pattern without naming it. Short-lived leased credentials re-read at checkout, with overlap-first rotation, beat the industry norm of a long-lived per-service secret. §9 promotes it to a tenancy control.

Cell-based architecture remains unexplored. Neither P nor E describes a cell — a fixed-size slice of infrastructure serving a bounded set of tenants, sized so one cell's failure is survivable and cell count scales linearly. Given the estate spans railiance, NetKingdom, HelixForge and Coulomb, cells may fit better than per-tenant clusters at the point P1 runs out (§18.8).

16. Scaling demands

Measured against the live platform-pg specification, not estimated.

instances:        1              (no HA; single-node rail)
max_connections:  100
memory limit:     1Gi
per consumer:     14 connections (12 runtime + 2 migration)

Connection ceiling: roughly six consumers. Seven consumers request 98 of 100 before CNPG's instance manager, metrics exporter, and reserved slots. Two are provisioned today. We are at roughly a third of capacity and the third request will not feel like a scaling event.

Memory likely binds first. 100 backends against 1Gi is ~10MB per backend for shared buffers, work_mem and overhead. Connection exhaustion errors clearly; memory pressure OOM-kills and degrades every co-resident at once.

Raising E is not free on a shared substrate. RLS with a per-checkout GUC does not itself increase connection count, but it forecloses aggressive transaction-level pooling for consumers that adopt it. E3 on P1 interacts with the ceiling above and should be planned against it.

Restore time couples all consumers. Physical backup is instance-wide, so a consumer's RTO is a function of total instance size, not its own. One large consumer silently degrades every co-resident's recovery promise.

instances: 1 means no P1 tenant has HA. A tier promising uptime cannot be satisfied at P1 as currently built — an availability floor belongs in §11's minimum-level vocabulary alongside isolation.

17. Consequences

  • The estate gains one vocabulary and, more importantly, a way to be honest about partial adoption.
  • Absorbed repos get a described state and a path instead of a failing grade.
  • tenantIsolation in PostgresConsumer is revealed as a mislabelled field (§18.1).
  • The two ratified defaults are reconciled by scope, at the cost of needing a classification rule.
  • The verification problem becomes tractable: guard against declaration, not against a universal maximum.
  • Draft-1's prohibition on RLS is reversed; rapp-postgres acquires an obligation to define and offer the E3 mechanism.
  • Nothing here changes a running system.

18. Open questions

  1. tenantIsolation fieldrapp-postgres: rename to name its plane and carry a level (tenancy.E: 2), or move it out of the storage declaration.
  2. Placement ownershiprailiance-platform: accept the ladder, the triggers, and the monitoring obligation in §8.1, and appoint a recorded placement owner per workload.
  3. E2/E3 evidence artifactsowner needed. Who builds the adversarial cross-tenant probes and the RLS conformance check? Both audit-core and tenant-engine have declined fleet-scope work on correct boundary reasoning, so this needs appointing. Highest-severity gap (§13.2).
  4. Business app vs platform service — Custodian canon: a classification rule, so a new repo knows whether P3 or P1 is its default. Candidate: reuse repo-classification-standard_v1.0 rather than invent a second taxonomy.
  5. Tier → minimum level mappingadaptive-pricing and tenant-engine: required only for tiers making isolation or availability claims (§11).
  6. E3 mechanismrapp-postgres: what exactly is the GUC contract, and is it provisioned per consumer on request or offered as a standard option?
  7. Engine neutrality — the P ladder rests on a PostgreSQL property (per-database catalogs). State it engine-specifically and say so, or abstract it and risk a non-Postgres implementation that silently differs? Inherited unresolved from the v0.1 draft.
  8. Cell-based partitioning — evaluate at consumer #4 rather than now? Neither P nor E describes a cell.

19. Ratification path

The predecessor draft stalled because it was routed without one.

  1. Reviewed by tenant-engine, flex-auth, rapp-postgres, railiance-platform, and adaptive-pricing against §18.
  2. Each of those repos publishes its own posture vector (§5) as part of review. The framework is validated by whether it can describe them accurately — if a repo cannot express itself in these four ladders, the ladders are wrong and this document changes, not the repo.
  3. On acceptance, supersedes the routing of rapp-postgres/docs/canon-drafts/shared-platform-relational-storage_v0.1-draft.md, whose §§38 are absorbed here. That draft is then withdrawn rather than left pending.
  4. On acceptance, rapp-postgres/docs/adr/ADR-0001 moves to accepted and is annotated as the PostgreSQL implementation of the E and P ladders.