tenant-engine assessed itself against the ladders and came back with corrections. All five are adopted, because the ratification test says that if a repo cannot express itself the ladders are wrong - and it could not, in three places. I2 conflated authority with verification. Its text named tenant-engine as the source of existence, which made the level describing canonical identity unclaimable by the service that provides it. An axis is assessed on a service's own inbound surface, never on its authority over the concept. tenant-engine is the source of tenant records and is at I1, because the acting identity arrives in the request body rather than in a verified token. A level now reports the weakest surface. tenant-engine has PDP-authorized mutations and three unauthorized read routes - including the one flex-auth calls for aal2-class decisions - and reported A2 rather than A3. Publishing the stronger surface would be accurate about that surface and misleading about the service. A per-surface vector was considered and rejected as premature. The E ladder assumed all data is tenant-keyed. A registry whose rows ARE the tenants has no predicate to scope a policy by, and enforcing one would break the service rather than secure it. Mixed-shape services now declare an E level plus a named registry exception; an unnamed exception is an overclaim. Without this they overclaim or sit at E2 forever, which is what tenant-engine was facing. Retention and erasure are two dimensions and one level cannot carry both. tenant-engine is R1 on backup and R0 on erasure - its lifecycle contract deliberately defines no hard-delete, so a tenant record cannot be deleted ever, by design, while carrying display_name and contact_email. Declared R1/R0 now. And the compounding - personal data, no erasure path, a backup window set by the longest-retaining co-resident - is the substrate owner's to surface, because each part looks locally reasonable alone. My own error, corrected: I listed tenant-engine as a live P1 occupant in both the ladder and the E/P matrix. They are on SQLite. P1 is TEN-WP-0009's target and the provisioning is my own unapplied intake. Asserting a placement that a workplan exists to create is exactly the kind of claim this document forbids. Worth recording: they found an unfiltered cross-tenant read in their own event accessor while assessing against the ladder, before publishing anything. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
49 KiB
| id | type | title | domain | status | version | created | updated | scope | revision | adr | related | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| netkingdom-tenancy-posture | standard | NetKingdom Tenancy Posture v0.1 | netkingdom | proposed | 0.1 | 2026-08-17 | 2026-08-17 | multi-tenancy-security-framework | draft-6 |
|
|
NetKingdom Tenancy Posture v0.1 — Five Axes, Graduated Levels, Declared Conformance
Status
Proposed, draft-6. Relocated from the-custodian/canon/architecture on
2026-08-17: multi-tenancy is part of the IT-security framework NetKingdom
provides, so this framework belongs in NetKingdom canon beside the IAM Profile
and the tenant-engine boundary contract, not in the work-factory canon.
- draft-1 proposed a single model with fixed characteristics. Rejected: it could not describe a repo that is not there yet.
- draft-2 reframed to graduated levels per axis. Externally corroborated (§16), but four of its statements were wrong and one thing it needed was missing.
- draft-3 applied those corrections, added the retention axis, and recorded an adoption stance.
- draft-4 closed the two gaps draft-3 left open:
R4had no mechanism beyond waiting, and the noisy-neighbour evidence artifact asserted something shared infrastructure cannot provide. - draft-5 relocated to NetKingdom and renamed the dimensions from planes to axes, because the word was already taken (§0).
- draft-6 applies the first review.
tenant-engineassessed itself, corrected a guess downward on two axes, found an axis that did not fit its data shape, and found a real cross-tenant defect in its own code while reading the ladder. Five changes followed (§4.1, §4.4, §4.5, §5.2, and the E registry exception). The ratification test worked: the document changed, not the repo.
Every correction so far was found by research or by relocation, not by review.
Informed by five external research digests in research/2026-08-17-adr008-*,
which carry full citations for every external claim made here.
Reviewed by nobody yet. §19 lists what each owner is being asked to accept.
0. Terminology: axes, not planes
docs/platform-identity-security-architecture.md — accepted, 2026-07-23 —
already uses plane for a trust and deployment layer: the bootstrap plane,
the platform control plane, and tenant planes. That meaning is established,
ratified, and owned by this repo.
Drafts 1–4 of this document, written elsewhere, used plane for something different: an independent dimension of concern. Two incompatible senses of one word inside one canon is exactly the concept-ownership collision the estate has been careful about elsewhere, and the newcomer yields.
This framework therefore describes five axes. They are orthogonal to NetKingdom's planes, not a subdivision of them:
- A plane is where something runs and what trust it carries — bootstrap, platform control, tenant.
- An axis is which property of tenancy is being described — identity, authorization, enforcement, placement, retention.
A workload in the tenant plane has a position on all five axes. A platform control plane service does too. The two vocabularies compose and neither replaces the other.
The rename is also an improvement. A posture vector is literally a point in five-dimensional space, and "axis" says that where "plane" did not.
1. Context
Drafts 1–4 opened by claiming the estate "has never written down what it is
building". Relocation proved that wrong, and the correction is worth keeping
visible: docs/platform-identity-security-architecture.md has described the
trust model, the tenant model and a capability progression since 2026-07-23.
The accurate claim is narrower — what was missing is a way to say how far a
given service has got, and to hold several answers at once. Seven documents
cover slices of the subject and none of them does that:
| Document | Covers | Status |
|---|---|---|
iam-profile_v0.3 (NetKingdom) |
Tenant identifier shape, tenant_roles claim, staleness rules |
Ratified |
tenant-engine-boundary-contract_v0.1 (NetKingdom) |
Who owns tenant records, roles, plan assignment | Ratified |
business-app-service-contract_v0.1 §1 (Custodian) |
Business apps: instance-per-client, tenant-keyed data | Ratified |
rapp-postgres ADR-0001 |
Consumer + tenant isolation in PostgreSQL | Proposed, governs one repo |
rapp-postgres ADR-0002 |
Per-consumer retention and the erasure horizon | Proposed, governs one repo |
shared-platform-relational-storage_v0.1 |
The stacked-boundary gap | Routed 2026-08-10, still unratified |
platform-identity-security-architecture (NetKingdom) |
Trust model, planes, tenant model, capability progression | Accepted 2026-07-23 |
This document is downstream of that architecture and must not restate it. It answers one question the architecture leaves open: given the model, where is this particular service today, and how would anyone know?
Four failures follow.
The gap was diagnosed once and the fix stalled. The v0.1 draft was written to fill this hole and has sat unratified in neither canon directory. §20 attaches a ratification path so this one does not join it.
Placement is owned by nobody. user-engine-pg and target-revenue-pg are
dedicated; apps-pg, net-kingdom-pg, platform-pg, state-hub-db and
forgejo-db are shared. Both live, neither written down. tenant-engine
raised this with railiance-platform on 2026-08-16; unanswered.
Two contradictory defaults are already ratified. Business apps get instance-per-client; platform services pool. Nothing says which shape a new service takes, and no definition separates the categories.
There is no honest way to describe a repo that is not there yet. The estate absorbs repos with weak or absent tenant separation. Today such a repo is simply non-conformant, leaving it two bad options: misrepresent its posture, or stay outside the framework.
2. What this document is
A framework, not a model. It specifies no single correct implementation. It supplies terminology (§3, §4), a declaration (§5), a conformance rule (§6), methodology (§12), and evidence definitions (§13).
A service is conformant when its declared posture is accurate and its trajectory recorded. A service is non-conformant when it claims a level it cannot evidence — regardless of how high or low that level is.
3. Five orthogonal axes
"Is this multi-tenant?" is treated as one question. It is five, and they are independent:
| Axis | Question | Vocabulary owner |
|---|---|---|
| Identity (I) | How is a tenant named and validated? | tenant-engine / IAM Profile |
| Authorization (A) | How is a request bound to the tenants it may act for? | flex-auth |
| Enforcement (E) | Where, mechanically, is the tenant boundary enforced? | This framework |
| Placement (P) | Which substrate holds a tenant's data? | railiance-platform |
| Retention (R) | How long does data persist, and how is it erased? | The storage platform; policy by the consumer |
Conflation produces errors today. rapp-postgres's PostgresConsumer carries
tenantIsolation: consumer-service-boundary — an E-axis fact in a
P-axis artifact, reading as though storage enforces something it does not.
The "dedicated versus shared" argument mixes P (capacity, blast radius) with E
(correctness).
The axes are separated precisely so each may sit at a different level.
Decision 3.1: every document, declaration and plan tier that says "isolation" MUST name which axis it means.
Decision 3.2: the axes couple at their tops and the couplings MUST be stated where they apply, not used to argue the axes are one:
E4is reachable only atP3or above.R's erasure horizon is bounded below byP— on shared substrate, a consumer's horizon is the instance maximum (§4.5).R4by key destruction is bounded by the key boundary, which is an E-axis property. Shredding a single tenant's data requires the application to encrypt under a per-tenant key before writing; the storage platform cannot supply it. Reaching the top of the retention ladder is not a retention project.
Decision 3.3 — scope. The P and R ladders describe a service's primary datastore. Caches, search indices, message queues and background jobs are named leak surfaces in the external baselines and are assessed separately, not covered by a posture vector. Saying so is honest; implying the vector covers them would not be.
4. Graduated levels
Each axis carries an ordered ladder. Higher is stronger, not better: the right level is the one a service can evidence and its risk warrants.
4.1 Identity (I)
| Level | State |
|---|---|
| I0 | No tenant concept. Data not attributable to a tenant. |
| I1 | A local tenant notion exists but is not canonical, or the tenant is taken from the request rather than from a verified token. |
| I2 | Canonical identifiers, bound at the identity provider and carried as a verified claim, and verified by this service on its own inbound calls. |
| I3 | I2 plus capability roles honoured, with live tenant-engine re-query for privileged, destructive, credential-vending or aal2-class decisions. |
I1 now explicitly absorbs request-supplied tenant identifiers. "Never trust client-supplied tenant IDs without validation" is a named anti-pattern; a service reading the tenant from a header is at I1 however canonical the string.
An axis is assessed on a service's own inbound surface, never on its
authority over the concept. tenant-engine is the source of existence for
tenant records and is nonetheless at I1, because it takes the acting identity
from the request body rather than from a verified token. Draft-5 conflated
these by naming the authority inside the I2 definition, which made the level
describing canonical identity unclaimable by the service that provides it.
Corrected on tenant-engine's review — a reader would otherwise assume the
authority must be at I2 by definition.
business-app-service-contract §2.1 sets app-local accounts as the v1 baseline
for business apps — a sanctioned low level with recorded triggers for moving
up. That is the pattern this framework generalises.
4.2 Authorization (A)
| Level | State |
|---|---|
| A0 | No authorization, or tenant context not carried. |
| A1 | Ad-hoc checks scattered through handlers. |
| A2 | A single local authorization boundary; tenant context bound once, centrally. |
| A3 | Decisions delegated to flex-auth as PDP, with live re-query where the IAM Profile requires it. |
| A4 | A3 over a standard PDP interface (OpenID AuthZEN Authorization API 1.0), so the decision point is swappable and the enforcement point is not coupled to one engine's request shape. |
A4 is new. flex-auth uses a bespoke CheckRequest and a bespoke action
vocabulary, with action strings copied verbatim between repos to avoid
re-derivation — exactly the coupling AuthZEN removes. The specification reached
Final in January 2026 and Keycloak shipped experimental support in May. We are
not wrong, we are pre-standard, and the ladder should have somewhere to go.
Internal service-to-service calls are in scope for this axis. "Skipping
tenant validation for internal services" is a named anti-pattern, and our
estate is mostly internal calls — flex-auth calls tenant-engine
synchronously on the authorization path. A service identity acting on behalf of
a tenant must carry and revalidate tenant context to claim A2 or above.
4.3 Enforcement (E)
| Level | Mechanism |
|---|---|
| E0 | None. Data not tenant-keyed; separation incidental or absent. |
| E1 | Data tenant-keyed, filtering applied per query at call sites. |
| E2 | Filtering centralised at a single service-side choke point binding authenticated identity to permitted tenants. |
| E3 | E2 plus platform-assisted filtering: row-level security keyed on a tenant GUC set transaction-locally, or an equivalent enforced data-access layer. |
| E4 | Structural: the credential a workload holds cannot address another tenant's data at all. Requires per-tenant credentials and per-tenant substrate. |
Correction from draft-2. Draft-2 described E3 as something "the application
cannot trivially route around". That is false and it was this document
overclaiming in exactly the way §6 prohibits. Any session can re-issue SET on
a custom GUC, so an attacker with SQL execution can reset the tenant and read
across the boundary. What E3 buys is precise, and the ladder must say so:
| Threat | E1 | E2 | E3 | E4 |
|---|---|---|---|---|
| A developer forgets a tenant predicate | ✗ | ✓ | ✓ | ✓ |
| A new code path bypasses the choke point | ✗ | ✗ | ✓ | ✓ |
| SQL injection reaching the connection | ✗ | ✗ | ✗ | ✓ |
| The application process is compromised | ✗ | ✗ | ✗ | ✓ |
E3 is a strong control against accident — the common case, and the one that causes real breaches — and no control at all against compromise. Only E4 holds against both, because the credential itself cannot address another tenant's data.
Correction: E3 layers on E2, it does not replace it. External practice treats application-layer and database-layer filtering as complementary. A service that dropped its choke point on reaching E3 would be worse off, since E3 fails open under injection. Claiming E3 therefore requires the E2 evidence artifact as well.
Correction: the GUC is set transaction-locally. Draft-2 said "at pool
checkout", which is session scope and the wrong instrument. Under a pooler in
statement mode, SET leaks between clients and returns other tenants' rows —
a failure that appears only under production concurrency and produces no error.
Use SET LOCAL inside an explicit transaction.
Platform enforcement is a platform obligation. Reaching E3 requires the
storage platform to offer the mechanism: provisioned policies, a documented
GUC contract, and a probe. Where a consumer wants E3 and the platform has not
supplied it, the gap is the platform's. §19.6 asks rapp-postgres to define
that contract, which must carry FORCE ROW LEVEL SECURITY on every tenant
table (without it the table owner bypasses policies silently, and ADR-0001
already established that our migration role owns the tables it creates), no
BYPASSRLS on leased roles, SECURITY INVOKER for ordinary logic, and an
EXPLAIN comparison because RLS disables functional indexes built on
non-leakproof functions.
Not all data is tenant-keyed, and the ladder must not pretend otherwise. A
registry whose rows are the tenants has no per-tenant predicate to scope a
policy by; enforcing one would break the service's function rather than secure
it. tenant-engine's tenants table is the worked example — key-cape
enumerates it at token issuance and flex-auth queries it live, both of which
are cross-tenant reads by design.
A service with mixed data shapes declares E-level plus a registry
exception: the level its tenant-keyed tables hold, and a named list of tables
excluded because they are registries rather than tenant data. The exception is
part of the claim and is reviewable; an unnamed exception is an overclaim.
Without this, mixed-shape services either overclaim or stay at E2 permanently,
and tenant-engine declined to claim E3 on precisely that reasoning.
Default expectation for a new platform service: E2 at first serve, E3 recorded as target. Services whose cross-tenant exposure would be a reportable breach SHOULD target E3 or above.
4.4 Placement (P)
| Level | Shape | Live occupants |
|---|---|---|
| P0 | Shares a database with another consumer. | None sanctioned; the state absorbed repos arrive in. |
| P1 | Database per consumer, shared cluster. | audit-core on platform-pg; tenant-engine (target, TEN-WP-0009 — still on SQLite) |
| P2 | Dedicated cluster per consumer. | user-engine-pg, target-revenue-pg |
| P3 | Dedicated cluster per tenant. | Business apps per business-app-service-contract §1.2 |
| P4 | P3 plus separate region or jurisdiction. | None |
Enforcement and placement are independent axes. Plotted together, with where
each service actually sits — parenthesised entries are targets or defaults
rather than current positions, and — marks a cell the coupling in §3.2 makes
unreachable:
| E \ P | P0 | P1 | P2 | P3 | P4 |
|---|---|---|---|---|---|
| E4 | — | — | — | (business app) | |
| E3 | (target) | ||||
| E2 | (tenant-engine) | audit-core | |||
| E1 | (absorbed repo) | ||||
| E0 |
P0 → P1 → P2 is movement along the horizontal axis only. Those steps buy consumer isolation, capacity predictability, independent retention and a smaller operational blast radius. They do not raise the tenant boundary by one step. Only P3 makes E4 reachable. This is the most misusable fact in the framework and §11 governs how it may be described.
Decision 4.4.1: P1 is the default for platform services; P3 for client-facing business apps, as already ratified. A service unsure which it is must resolve that first (§19.4).
Decision 4.4.2 — placement scopes to data substrate. Identity-provider placement (realm-per-tenant versus Organizations) is the same silo/pool decision on a different substrate, is live in our estate, and is undecided. Realm-per-tenant carries a stated ceiling around 5–20 tenants, far below our target. Recorded here as a parallel question (§19.7), not folded into P.
4.5 Retention and erasure (R)
New in draft-3. Implemented abstractly by the storage platform for any dataset;
policy is built on top of that interface by the consumer or its governance
layer. Reference implementation: rapp-postgres ADR-0002.
| Level | State |
|---|---|
| R0 | No retention or deletion position. Data kept indefinitely by default; no deletion path exists. |
| R1 | Platform default retention applies (N=30 days). The consumer has declared no requirement. |
| R2 | Retention declared as N days per dataset; the erasure horizon is published, and the consumer makes no promise shorter than it. |
| R3 | Policy-driven deletion: the consumer or its governance layer declares what is due, the platform sweeps whole datasets on that instruction and evidences each run. |
| R4 | Verified erasure: data proven unrecoverable across live storage, backups and derived copies, by one of the two routes below. |
R4 has two routes and a service MUST name which one it uses.
| Route | Mechanism | Cost |
|---|---|---|
| Horizon-elapsed | Wait out the published erasure horizon; the data ages out of every retained copy. | Available to everyone, proves little, and the wait is set by a co-resident's retention requirement rather than your own. |
| Key-destroyed | Encrypt per entity, then destroy the key. Retained copies survive but are unreadable. | Requires per-entity keys, strong encryption, and an auditable destruction record. Immediate. |
Regulatory standing of the key-destroyed route, stated carefully because overclaiming here is worse than anywhere else in this framework. Data protection authorities have accepted key destruction as erasure where physical deletion would be manifestly disproportionate, and the practice is recognised under conditions — strong encryption, irreversible destruction, and an auditable record of it. The EDPB has not formally endorsed it as Article 17 erasure. A service reaching R4 by key destruction is making a defensible claim, not a settled one, and must say so rather than reporting a clean "deleted".
Three further properties.
The erasure horizon is the interval between deleting data and it ceasing to be recoverable from anything the platform holds. Deleting a row does not remove it from yesterday's backup. With an N-day window, deleted data remains recoverable for N days. That is the difference between "deleted" and "erased" and the estate had never written it down.
On shared substrate, retention is not per-consumer. Physical backup is instance-wide — one WAL stream, one window — so the instance retention is derived as the maximum across co-resident consumers, and every consumer's horizon is that maximum. A consumer declaring 7 days beside one declaring 90 gets 90. This is the retention analogue of ADR-0001's blast-radius disclosure: state the coupling rather than imply an isolation that is not there.
Retention is therefore a placement trigger. A consumer needing a horizon shorter than the instance floor cannot have one at P1. It moves to P2 for a reason with nothing to do with performance — which is exactly why it needs recording, since nobody looks for a retention argument when reviewing placement.
Deletion splits mechanism from policy. The platform deletes whole datasets on instruction and records an opaque policy reference it never interprets, so every deletion traces to what authorised it. Rows are not a dataset: row expiry is the consumer's own DML under its migration lease. Dropping a consumer's whole database is an operator-gated offboarding step, never a scheduled one.
5. The posture vector
A service states one level per axis, plus a target, a date, and any placement exceptions:
tenancy:
current: { I: 2, A: 3, E: 2, P: 1, R: 1 }
target: { I: 2, A: 3, E: 3, P: 1, R: 2 }
reviewed: "2026-08-17"
gap:
E: "Choke point exists and is tested; RLS not provisioned. Blocked on
rapp-postgres publishing the GUC contract. Target Q4."
R: "Retention declared; erasure horizon not yet published to consumers."
Placement exceptions. Draft-2 assigned one P level per service, which
cannot express the vertically partitioned model — most tenants pooled, some
dedicated — that §11's isolation tiers require. A tier requiring P2 bought by
three tenants would put the service at two levels at once, forcing an over- or
under-claim. Placement is therefore declared as a default plus exceptions:
placement_exceptions:
- tenants: ["tenant:enterprise:*"]
P: 3
reason: "isolation tier; see adaptive-pricing tier definition"
A service with exceptions must be able to say which tenants are on which substrate. That mapping is a first-class artifact, not archaeology.
Worked examples, best-effort and subject to owner correction:
| Service | Current | Notes |
|---|---|---|
tenant-engine |
I1 A2 E2 P— R0/R1 |
Self-reported on review, correcting a more generous guess. I1: acting identity comes from the request body, not a verified token. A2: three read routes unauthorized. P—: still on SQLite, P1 is TEN-WP-0009's target. R: see §4.5 on the erasure/retention split. |
audit-core |
I2 A3 E2 P1 R1 |
Guessed, not yet self-reported. Holds audit evidence, so E3 is urgent. |
| A newly absorbed repo | I1 A1 E1 P0 R0 |
Conformant if declared, with a recorded path. |
Decision 5.1: the posture vector is declared in the repo, not in the hub, consistent with local-files-are-source-of-truth.
Decision 5.2 — a level reports the weakest surface, not the best one. A service whose mutations are authorized by a PDP and whose read routes are unauthenticated is at the read routes' level, not the mutations'. Publishing the stronger surface would be accurate about that surface and misleading about the service, which §6 forbids.
Raised by tenant-engine, which found exactly this shape in itself during
review — A3 on writes, no authorization on three read routes including the one
flex-auth calls for aal2-class decisions — and reported A2. A per-surface
vector was considered and rejected as premature: it multiplies the declaration
before anyone has shown the single weakest number is insufficient. Services
with a materially split surface should record the split in the gap field.
6. Conformance is accuracy, not altitude
A service is conformant when its declared posture is accurate, its target is recorded, and it does not claim a level it cannot evidence. It is non-conformant when it overclaims — at any altitude.
- Declaring
E0is conformant. ConcealingE0is not. - A repo may be absorbed at any posture. It may not be absorbed silently.
- No service is blocked from the estate for being low on a ladder. Services MAY be blocked from specific work — serving a tenant grouping, holding a data class, carrying a plan tier — by requirements expressed as minimum levels.
- Downgrading is permitted and must be declared. A regression found by guarding is a defect; a regression declared in advance is a decision.
Without the axis separation, "not rigorous about tenant separation" is one
verdict a repo passes or fails. With it, the same repo is I1 A1 E1 P0 R0 with
a path — a plan, not an indictment.
7. Portability across placement levels
Movement between P levels must be operational, not a rebuild:
- Connect by injected credential only — no cluster, host, namespace or database name in source.
- Own a whole database, never tables inside someone else's.
- Idempotent schema creation.
- No cross-database joins or co-location assumptions.
Decision 7.1: mandatory at P1 and above. At P3, SHOULD rather than MUST — a per-client instance that never moves is not misconformant for naming its own database.
8. Placement triggers
Recorded at provisioning time: noisy neighbour on a latency-critical path; a compliance or residency requirement; a plan tier requiring a higher minimum; an erasure horizon that no longer fits (§4.5); connection or memory ceiling reached.
Decision 8.1: triggers MUST be monitored, not merely recorded. A trigger in a YAML comment nobody re-reads is documentation, not control.
Decision 8.2: placement policy ownership is proposed to
railiance-platform, co-signed by adaptive-pricing. Tenancy model
selection is a commercial decision as much as a technical one; an
operations-shaped repo should not hold it alone.
8.3 Service class — a placement input, never a priority
A latency-critical consumer and a batch consumer can share an instance today
with nothing distinguishing them. tenant-engine sits on flex-auth's
synchronous authorization path and chose a 5s statement timeout for that
reason; audit-core, co-resident, is not latency-critical. Nothing prioritises
between them.
The framework does not add a QoS axis, because the platform cannot enforce
one. Community PostgreSQL has no resource governor: no per-role CPU or I/O
priority, no resource queues, no workload classes. Those exist in EDB's
enterprise variant, in Greenplum, and in SQL Server — not in what we run. A
declared priority level would therefore be an unenforced claim sitting in a
declaration, which is precisely what retiring tenantIsolation was about. An
axis implies graduation and enforcement; this has neither.
Decision 8.3.1 — co-residents are equal. On shared substrate no consumer's
query yields to another's. A consumer whose latency requirement cannot survive
an unprioritised neighbour must escalate to P2. That is the honest mechanism
and it is the only one we have.
Decision 8.3.2 — service class is declared anyway, as a category rather
than a level: latency-critical, interactive, or batch. It buys three
things, none of which is priority:
- A placement input. Mixing
latency-criticalwithbatchon one instance is a recognised mismatch. It may still be the right call — it is right today — but it should be a decision, not an accident of who was provisioned when. - A trigger. A
latency-criticalconsumer acquiring abatchco-resident is a recorded placement trigger under §8, on the same footing as noisy neighbour. - An acceptance criterion for evidence. The noisy-neighbour artifact in §13
asks whether measured degradation is acceptable; without a declared class
that word has no referent. Degradation tolerable for
batchmay be an outage forlatency-critical.
Decision 8.3.3 — class mixture must be visible. The platform reports which classes are co-resident. An unenforceable risk that nobody can see is strictly worse than one that is stated.
The known escalation short of P2 is gateway-level prioritisation — ordering submissions in a connection proxy by the requesting tenant's current consumption. It is real, it is where the industry puts this when it must, and it is new infrastructure we do not run. Recorded as the option, not adopted.
9. Credentials as a tenancy control
Short-lived leased credentials re-read at connection checkout, with overlap-first rotation, bound the residual risk at every E level below E4: a leaked credential expires rather than persisting. Stronger than the industry norm of a long-lived per-service secret.
Decision 9.1: static long-lived database credentials are not a sanctioned path for any service above E0.
10. Blast radius must be published
Decision 10.1: every platform holding consumer data MUST publish, in
concrete terms, what a leaked runtime credential can and cannot reach at the
levels it operates. rapp-postgres ADR-0001 §5 is the reference. Where the
model cannot provide a guarantee, the platform says so and names the
escalation.
Decision 10.2 — quotas are disclosed, not discovered. The same obligation
extends from what a leaked credential can reach to what the platform will
refuse to do for you. Every consumer MUST be told, at provisioning, the
throttles and quotas enforced against it — connection limits, statement
timeouts, idle-transaction timeouts — and told again when they change. A
consumer learning its statement timeout by hitting it in production is a
disclosure failure, not a consumer bug. This is how tenant-engine was
provisioned, by good practice rather than by rule; the rule now exists.
11. Commercial expression
- 11.1 Plan tiers are expressed internally as minimum levels. A tier may
require
E3 P2 R2; it need not print that anywhere customer-facing. - 11.2 Marketing and product language is free. No requirement to expose level labels or this document. "Dedicated infrastructure", "isolated tenancy", "private instance" all remain available.
- 11.3 The constraint is on evidence, not vocabulary. A customer-facing isolation, availability or retention claim must map to a minimum level the delivering service actually holds, recorded once when the tier is defined. The review is internal and happens at tier definition — not per campaign.
- 11.4 Two hard lines, because these reach contracts and compliance
questionnaires:
- A claim that another tenant cannot reach the customer's data requires E4.
- A claim that deleted data is gone requires R4, or an erasure horizon disclosed alongside it. Where R4 is reached by key destruction, the claim is defensible but not settled law (§4.5) — it may be made, and it may not be made in language that implies a regulator has blessed it.
12. Methodology — analyze, establish, improve, guard
Analyze. Assess a repo against the ladders; produce tenancy.current with
reasoning recorded. Applies to new and absorbed services alike.
Establish. Declare the target and gap. The target is set by data class, tenant groupings served and plan tiers carried — not by ambition.
Improve. Move one axis at a time. Raising P while leaving E untouched is the characteristic misstep.
Guard. Verify continuously that the declared posture holds — against the service's own declaration, not a universal maximum. Nobody must prove every service is at E4; the check is that none is below what it declared.
Regression found by guarding is a defect; regression declared in advance is a decision. The estate has been bitten twice by silent pin rollbacks producing ordinary-looking 403s and 404s rather than errors. Posture regression looks the same — an RLS context leak returns correct-looking rows for the wrong tenant. Guarding must be designed for invisible failure, not for crashes.
13. Evidence per level
Decision 13.1: a level is claimed only with its evidence artifact present. This turns §6's accuracy rule from an honour system into a check.
Decision 13.4 — an artifact must assert something achievable. Draft-3's noisy-neighbour evidence required proof that a saturating consumer "does not breach" another's allowance. Shared infrastructure cannot provide that; the risk is inherent and cannot be wholly removed. An artifact that can only fail, or that passes by being run gently enough, is an overclaim wearing the costume of evidence. Where a property cannot be guaranteed, the artifact measures and records it instead.
Decision 13.2 — evidence is of two kinds, and conflating them is an overclaim. Mechanical evidence is a structural assertion a machine can make and belongs in CI. Adversarial evidence is semantic, requires setting up separate tenant contexts and comparing responses, and carries a review date rather than a green build. Cross-tenant findings are the category external testing practice identifies as needing human review. A passing CI run is not E2 evidence.
| Level | Evidence | Kind |
|---|---|---|
| I2 | Identifiers validated against the vocabulary; rejection test for a malformed id; binding shown to come from a verified token | Mechanical |
| I3 | Live re-query demonstrated on an aal2-class path; cached-claim path shown unused there |
Mechanical |
| A2 | Choke point identified; test that an unbound request is refused | Mechanical |
| A3 | Live decision with a denial observed at the endpoint, not only at the decision surface | Mechanical |
| A4 | Decision served over the standard interface; a second PDP substituted without PEP change | Mechanical |
| E1 | Every tenant-owned table carries the tenant key | Mechanical |
| E2 | Choke point identified; identity bound to tenant A demonstrably cannot read tenant B | Adversarial, with a review date |
| E3 | FORCE ROW LEVEL SECURITY on every tenant table; no BYPASSRLS on leased roles; probe that a session without the GUC reads nothing; probe that a wrong GUC reads nothing; EXPLAIN comparison |
Mechanical |
| E4 | Per-tenant credential demonstrated unable to connect to another tenant's substrate | Mechanical |
| P1–P4 | Provisioning declaration plus the platform's isolation probes | Mechanical |
| P1–P2 (noisy neighbour) | A recorded baseline of per-consumer resource usage; a run in which one consumer saturates its declared allowance; evidence that the governance controls bind (the greedy consumer is held at its limits) and that the degradation co-residents experience is measured, recorded and judged acceptable against each one's declared service class (§8.3); the aggregate headroom at time of measurement | Adversarial, load-generated, with a review date |
| R2 | Declared retention rendered; erasure horizon published and reported in the operator surface | Mechanical |
| R3 | Sweep evidence records: timestamp, dataset, identifiers removed, authorising policy reference | Mechanical |
| R4 | Erasure demonstrated across live data, backups and derived copies within the horizon | Adversarial |
Decision 13.3: the E2, E3 and noisy-neighbour artifacts do not exist
anywhere in the estate today. rapp-postgres runs 15 adversarial probes, all
against the consumer boundary, none against the tenant boundary inside a
consumer. Externally, what this framework calls a tenant boundary failure is
Broken Object Level Authorization — OWASP API1, top of the API Security Top
10 since that list launched, and the most commonly exploited API vulnerability
in published assessments. We have no coverage for the highest-ranked risk in
our class of system. §19.3 seeks an owner.
14. Adoption stance — structure, not tooling
Decision 14.1: external research is design input. This estate adopts published standards and structural patterns; it does not adopt tooling unless that tooling is an established industry standard with broad application. Everything else is built ground-up, so it can be optimised and refactored as the estate sees fit.
| Class | Stance |
|---|---|
| Security baselines (OWASP Multi-Tenant Security Cheat Sheet, API Security Top 10) | Adopt as the external reference our ladders answer to |
| Standards bodies (OpenID AuthZEN 1.0) | Adopt — this is what A4 is |
| Reference taxonomies (Azure tenancy models, AWS SaaS Lens, cell architecture) | Adopt as structure |
| Engine behaviour (PostgreSQL RLS mechanics) | Facts, not tooling |
| Third-party analyzers and test frameworks | Do not adopt. Take their rule taxonomies as checklists for probes we write ourselves |
The practical effect is small and good: rapp-postgres already owns a
ground-up probe harness — bash and psql, no dependency tree — that found four
real defects in its own provisioning SQL. The evidence artifacts in §13 become
new probes in a tool we control. One idea worth reimplementing from the
external survey is policy-diff classification: labelling a change to an
enforcement policy as safe or breaking before it lands.
15. Alternatives considered
One fixed model with a single set of characteristics (draft-1). Rejected: cannot describe a repo that is not there yet, forcing absorbed repos to misrepresent their posture or stay outside. A framework that can only describe its own end state is not a framework.
A maturity model with a single overall level. Rejected: collapses the axis separation. A service strong on identity and weak on enforcement has a specific, actionable gap; one composite score hides it and invites averaging.
Prohibiting row-level security (draft-2's inherited position). Rejected in draft-2, refined in draft-3: RLS is a real rung against the common threat. The error was never RLS — it was describing E3 in E4's language.
Schema-per-consumer in one database. Rejected: pg_catalog is readable
per-database, so every co-resident enumerates every other's table and column
names regardless of grants. Retained as a describable state, never a target.
Mandating E4 for everyone. Rejected: the tenant taxonomy includes
consumer (private individuals) and family. A cluster per private individual
is economically impossible; the taxonomy is itself evidence pooling is
required.
Per-consumer physical backup retention. Rejected: CNPG retention is a property of the instance's WAL archive. There is no mechanism, and claiming it would be a fabricated guarantee. Hence the derived maximum in §4.5.
Platform-scheduled row expiry. Rejected: requires the platform to hold DML authority over consumer schemas and interpret consumer data semantics, both forbidden by ADR-0001. The consumer's migration lease is the correct instrument.
Leaving each repo to its own model. Rejected: the status quo, which produced two contradictory ratified defaults and an unowned placement question.
16. Held against outside practice
The graduated reframe is corroborated, not invented here. Microsoft's tenancy-model guidance states it almost verbatim: "Instead of viewing isolation as a discrete property, consider it a spectrum. You can deploy components of your architecture that are more isolated or less isolated than other components in the same architecture." The same guidance derives our E↔P coupling independently — shared deployment means enforcement lives in application code; dedicated deployment means it is structural.
Stronger than typical. Most multi-tenancy literature models one boundary, tenant-to-tenant. This estate has two stacked boundaries: platform-service to platform-service, and tenant to tenant inside a consumer. Naming them separately and refusing to enforce both with one mechanism is uncommon and correct. Graduated per-axis levels also beat the silo/pool/bridge trichotomy, which is approximately our P axis with the other four missing — which is why it cannot express "pooled infrastructure, structurally enforced boundary".
Weaker than typical. The pool model's standard mitigation is a verified enforcement layer every service is demonstrably routed through. We have the concept and none of the verification (§13.3).
Adopted without naming it. Short-lived leased credentials re-read at checkout beat the long-lived-secret norm. §9 promotes it to a tenancy control.
Still unexplored. Neither P nor R describes a cell — a slice of
infrastructure with a fixed maximum size, sized so one cell's failure is
survivable and cell count scales linearly. platform-pg is, in these terms, an
uncapped cell: §17 computes a ceiling and nothing enforces it (§19.8).
Sources: the four research digests in research/2026-08-17-adr008-*, which
carry full citations for every claim in this section.
17. Scaling demands
Measured against the live platform-pg specification, not estimated.
instances: 1 (no HA; single-node rail)
max_connections: 100
memory limit: 1Gi
per consumer: 14 connections (12 runtime + 2 migration)
Connection ceiling: roughly six consumers — and this is the aggregate noisy-neighbour bound, not a capacity statistic. Seven consumers request 98 of 100 before CNPG's instance manager, metrics exporter and reserved slots. Every one of them is politely inside its declared 14-connection allowance; the instance still fails.
That distinction matters because our governance addresses the wrong shape.
Per-consumer connection_limit, statement_timeout and
idle_in_transaction_session_timeout guard well against one greedy
consumer. They do nothing about the aggregate of many modest ones, which
is the second and less intuitive noisy-neighbour failure and the one this
number describes. Two consumers are provisioned. We are at roughly a third of
the bound, and the third request will not feel like a scaling event.
Memory likely binds first. 100 backends against 1Gi is ~10MB per backend. Connection exhaustion errors clearly; memory pressure OOM-kills and degrades every co-resident at once.
E3 and pooling. Corrected from draft-2, which had this backwards.
Transaction-scoped context (SET LOCAL inside an explicit transaction) is what
makes E3 safe under a pooler. Statement-level pooling is what breaks it,
serving other tenants' rows under concurrency with no error. E3 constrains
which pooling mode is available, not whether pooling is available.
Retention consumes the volume. WAL accumulates with the window, and §4.5 makes the window the maximum across consumers. A consumer declaring a long retention extends everyone's horizon and everyone's storage draw against a 20Gi volume.
Restore time couples all consumers. Physical backup is instance-wide, so a consumer's RTO is a function of total instance size, not its own.
No P1 tenant has HA. instances: 1 means a tier promising uptime cannot be
satisfied at P1 as built — an availability floor belongs in §11's
minimum-level vocabulary alongside isolation.
18. Consequences
- The estate gains one vocabulary and a way to be honest about partial adoption.
- Absorbed repos get a described state and a path instead of a failing grade.
tenantIsolationinPostgresConsumeris revealed as a mislabelled field.- The verification problem becomes tractable: guard against declaration.
- Draft-2's RLS prohibition is reversed and its E3 description corrected;
rapp-postgresacquires an obligation to define and offer the mechanism. - Adding a consumer with long retention silently extends everyone's erasure horizon. This must reach the consumer review checklist, not only this document.
- A service selling an isolation tier must maintain a tenant→substrate mapping it does not have today.
- Nothing here changes a running system.
19. Open questions
-
tenantIsolationfield —rapp-postgres: rename to name its axis and carry a level (tenancy.E: 2), or move it out of the storage declaration. -
Placement ownership —
railiance-platformwithadaptive-pricing: accept the ladder, triggers and the §8.1 monitoring obligation; appoint a recorded placement owner per workload. -
E2, E3 and noisy-neighbour evidence — owned as of 2026-08-17 by
whitehat-security(WHITEHAT-WP-0001), an independent adversarial evidence facility seeded for this purpose.audit-coreandtenant-enginewere right to decline it as fleet-scope work; the answer was a home of its own rather than a volunteer.Owned by NetKingdom — corrected 2026-08-17; an earlier revision of this section proposed otherwise on independence grounds and was overruled. Offensive security is security work and belongs with the repo that owns security. The facility is framed offensively rather than as a conformance checker: it is pointed at infrastructure we choose, our own estate among them, and conformance testing is one use of a general capability.
The residual tension is recorded rather than resolved: NetKingdom owns this framework and the facility that tests conformance to it, so those findings are NetKingdom assessing NetKingdom. The mitigation is that findings leave for
risk-nexus, underthe-custodian, rather than being closed in place. Proportionate, not perfect. Revisit if conformance findings start getting quietly closed.Two consequences land back here. Cadence is now a security parameter, not a schedule — for any control whose guarantee is detection rather than prevention, the interval between probe runs is the exposure window, and
rapp-postgresADR-0003 leaves that number to the facility. And a passing suite is not proof of isolation; it is proof that the attacks attempted did not work. §13's evidence artifacts should be read with that distinction, because a green run recorded as "E2 verified" would be exactly the overclaim §6 prohibits. -
Business app vs platform service — Custodian canon: a classification rule. Candidate: reuse
repo-classification-standard_v1.0. -
Tier → minimum level mapping —
adaptive-pricingandtenant-engine: required only for tiers making isolation, availability or retention claims. -
The E3 mechanism —
rapp-postgres: publish the GUC contract with theFORCE/BYPASSRLS/SECURITY INVOKER/EXPLAINrequirements in §4.3. -
Identity-provider placement — owner of
key-cape: realm-per-tenant or Organizations? Realm-per-tenant's ~5–20 tenant ceiling is below our target. -
Cell sizing — reframed from "should we adopt cells" to "what is
platform-pg's declared maximum size, and what is the overflow target?" The connection ceiling forces this whether or not we adopt the vocabulary. -
Retention floor and ceiling — should
backupRetentionDayshave a platform minimum (so a consumer asking for 1 day gets a validation error rather than a quiet disappointment) and a maximum (so nobody exhausts the volume)? -
Engine neutrality — the P ladder rests on a PostgreSQL property. State it engine-specifically and say so, or abstract it and risk a non-Postgres implementation that silently differs?
-
Erasure versus audit —
audit-core: crypto-shredding a tenant's audit records destroys the evidence the service exists to hold, and ADR-0001 §2 deliberately built the role model so history could not be rewritten. The usual resolution separates the fact of an event, retained, from its personal payload, encrypted per subject and shreddable. Raised because a naive "R4 everywhere" target would instruct the audit service to destroy its own evidence. The answer isaudit-core's, not this framework's. -
Quality of service — resolved 2026-08-17. Co-residents are equal; a declared service class informs placement but never grants priority. See §8.3. The question asked whether to add a QoS dimension; the answer is no, and the reason is that we could not enforce one.
Routed elsewhere, deliberately. The tenant identifier
tenant:<grouping>:<name> embeds headcount bands (small, medium, large)
that change as a tenant grows, contradicting the consensus that identifiers
should not encode mutable attributes. That is a critique of ADR-0013, not of
this framework, and belongs to tenant-engine and NetKingdom canon. Folding it
in here would overreach.
20. Ratification path
- Reviewed by
tenant-engine,flex-auth,rapp-postgres,railiance-platformandadaptive-pricingagainst §19. - Each publishes its own posture vector (§5) as part of review. The framework is validated by whether it can describe them accurately — if a repo cannot express itself in these five ladders, the ladders are wrong and this document changes, not the repo.
- On acceptance, supersedes the routing of
rapp-postgres/docs/canon-drafts/shared-platform-relational-storage_v0.1-draft.md, whose §§3–8 are absorbed here. That draft is withdrawn rather than left pending. - On acceptance,
rapp-postgresADR-0001 and ADR-0002 move toacceptedand are annotated as the PostgreSQL implementation of the E, P and R ladders.