research: external survey of multi-tenancy practice, one digest per ADR-008 plane
Four search-and-read passes against AWS Well-Architected SaaS Lens, Azure Architecture Center, OWASP, the OpenID Foundation, and PostgreSQL practitioner literature. Each digest holds one of our ladders against outside practice and lists its sources. Headline: the draft-1 -> draft-2 reframe is corroborated. Microsoft states "isolation is a spectrum, not a discrete property" and derives our E/P coupling independently. The shape is sound; the details are not. Five corrections draft-2 needs: - E3 overclaims. Any session can re-SET the GUC, so RLS is strong against a forgotten predicate and useless against SQL injection or a compromised process. The ADR says "the application cannot trivially route around" - the exact overclaim it prohibits elsewhere. - The pooling claim in section 16 is backwards. Transaction-scoped SET LOCAL is what makes E3 safe under pooling; statement-level pooling is what serves other tenants' rows, only under production concurrency. - E3 layers on E2 rather than superseding it. - Placement cannot be expressed per tenant, but the commercial section requires exactly that - a tier bought by some tenants puts a service at two P levels at once, which the posture vector cannot represent. - FORCE ROW LEVEL SECURITY is mandatory: ADR-0001 already established that our migration role owns the tables it creates, so it would silently bypass every policy. Also: pgTAP and pgrls make the E3 evidence artifact far cheaper than assumed; OWASP publishes a Multi-Tenant Security Cheat Sheet we reference nowhere; our tenant-boundary gap is BOLA, OWASP API1, top of that list since it launched; AuthZEN 1.0 went Final in January and flex-auth is pre-standard; and no plane covers tenant offboarding, which OWASP makes a MUST. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
9a6e14e733
commit
04f888399f
5 changed files with 1052 additions and 0 deletions
234
research/2026-08-17-adr008-plane-i-identity.md
Normal file
234
research/2026-08-17-adr008-plane-i-identity.md
Normal file
|
|
@ -0,0 +1,234 @@
|
|||
# Plane I — Identity: prior knowledge and best practices
|
||||
|
||||
> Research artefact — 2026-08-17. Supports `canon/architecture/adr-008-multi-tenancy-model.md`
|
||||
> (draft-2), plane I of four. Surveys external practice for how multi-tenant
|
||||
> systems name, carry, and validate tenant identity, and holds our `I0–I3`
|
||||
> ladder against it.
|
||||
> Companions: plane A (authorization), plane E (enforcement), plane P (placement).
|
||||
> Index: `2026-08-17-adr008-multi-tenancy-research-index.md`.
|
||||
|
||||
## 1. What the plane covers
|
||||
|
||||
How a tenant is named, where that name is bound to a principal, how it travels
|
||||
with a request, and how fresh the binding must be. Our ladder:
|
||||
|
||||
| Level | State |
|
||||
|---|---|
|
||||
| I0 | No tenant concept |
|
||||
| I1 | Local tenant notion, non-canonical |
|
||||
| I2 | Canonical `tenant:<grouping>:<name>`, validated; `tenant-engine` owns existence |
|
||||
| I3 | I2 + capability roles honoured, live re-query for `aal2`-class decisions |
|
||||
|
||||
## 2. Consensus practice: bind at registration, carry in the token
|
||||
|
||||
The dominant pattern across vendor guidance is that tenant identity is bound
|
||||
to the principal **at the identity provider**, not derived per-application.
|
||||
|
||||
- AWS calls the unpacked token claims the **tenant context**, and frames it as
|
||||
environment-like: it "influences how the SaaS application processes the
|
||||
request" rather than being one parameter among many
|
||||
([SaaS Lens — preventing cross-tenant access](https://docs.aws.amazon.com/wellarchitected/latest/saas-lens/preventing-cross-tenant-access.html)).
|
||||
- Binding is explicitly a **registration-time** responsibility: storing the
|
||||
tenant identifier as a claim or attribute on the user in the IdP "shifts the
|
||||
responsibility of mapping identities to tenants from each application to the
|
||||
user registration process"
|
||||
([AWS Prescriptive Guidance — tenant onboarding and registration](https://docs.aws.amazon.com/prescriptive-guidance/latest/saas-multitenant-api-access-authorization/avp-design-onboarding-registration.html)).
|
||||
- The recommendation is to define tenant context **at the IAM layer** rather
|
||||
than relying on the application layer
|
||||
([AWS — 5 multi-tenant SaaS architecture best practices](https://aws.amazon.com/isv/resources/5-multi-tenant-saas-architecture/)).
|
||||
|
||||
**Against our model:** this is what `I2` describes and what the IAM Profile
|
||||
already mandates (`tenant` claim required on every token accepted by profile
|
||||
consumers). We are aligned with consensus here, and `key-cape`/Keycloak
|
||||
stamping the claim at issuance is the sanctioned shape.
|
||||
|
||||
**Gap:** the ladder's I2 says "canonical identifiers used, validated". It does
|
||||
not say *where the binding happens*. A service could satisfy I2 by parsing a
|
||||
tenant id out of a request body. Consider tightening I2 to require the binding
|
||||
be IdP-sourced, with request-supplied tenant ids explicitly at I1.
|
||||
|
||||
## 3. Propagation is a separate discipline from naming
|
||||
|
||||
Practice treats "the tenant id must reach every layer" as its own requirement,
|
||||
not a consequence of having one:
|
||||
|
||||
- A stable identifier should follow every request and stay consistent across
|
||||
databases, services, logs, billing records, and background processes; without
|
||||
a reliable tenant registry, other isolation controls are hard to enforce
|
||||
consistently ([WorkOS — developer's guide to SaaS multi-tenant architecture](https://workos.com/blog/developers-guide-saas-multi-tenant-architecture)).
|
||||
- Requests should execute in a tenant context established by middleware or
|
||||
context variables, tagging each request with its owner
|
||||
([SSOJet — multi-tenant identity management](https://ssojet.com/blog/multi-tenant-identity-management)).
|
||||
- Two leak surfaces get called out repeatedly and are **not** database
|
||||
concerns: search indices (Elasticsearch queries that omit the tenant filter)
|
||||
and cache keys that do not include the tenant id
|
||||
([WorkOS](https://workos.com/blog/developers-guide-saas-multi-tenant-architecture)).
|
||||
|
||||
**Against our model:** ADR-008 has nothing to say about background jobs,
|
||||
caches, or search indices. Our E ladder is described almost entirely in
|
||||
database terms (tenant-keyed tables, RLS, per-tenant credentials). A service
|
||||
could sit honestly at `E3` and still leak through a Redis key or an unfiltered
|
||||
index.
|
||||
|
||||
**Proposed amendment:** either widen the E-ladder evidence artifacts to cover
|
||||
non-database stores, or add an explicit scope note that the E ladder describes
|
||||
the *primary datastore* only and other stores are assessed separately. The
|
||||
second is cheaper and more honest.
|
||||
|
||||
## 4. Identity-plane placement is its own decision — and it has a known ceiling
|
||||
|
||||
Identity providers face the same silo/pool question the P plane describes, and
|
||||
the industry has converged on an answer with a number attached.
|
||||
|
||||
- Keycloak offers **realm-per-tenant** (strong isolation, heavyweight) versus
|
||||
**Organizations** (multiple tenants inside one realm, stable since Keycloak
|
||||
26, designed for the SaaS case)
|
||||
([Phase Two — multi-tenancy options in Keycloak](https://phasetwo.io/blog/multi-tenancy-options-keycloak/),
|
||||
[Skycloak — multitenancy using Organizations](https://skycloak.io/blog/multitenancy-in-keycloak-using-the-organizations-feature/)).
|
||||
- The operational verdict is blunt: "realm-per-tenant works for 5–20 tenants;
|
||||
beyond that, you are managing an IAM platform, not building your product"
|
||||
([Izylife — Keycloak multi-tenancy: realm per tenant or Organizations?](https://izylifesolutions.com/articles/keycloak-multi-tenancy-b2b-iam/)).
|
||||
- Auth0's equivalent is Organizations within a single Auth0 tenant, with
|
||||
per-organization connections
|
||||
([KeycloakPro — Keycloak multi-tenancy with Organizations](https://keycloakpro.com/blog/keycloak-multi-tenancy-organizations-guide)).
|
||||
|
||||
**Against our model:** ADR-008 treats placement as a property of *data
|
||||
substrate* only. Identity has a placement axis too, it is live in our estate
|
||||
(`key-cape`/Keycloak), and nobody has decided realm-per-tenant versus
|
||||
Organizations. Given the tenant taxonomy includes `consumer` (private
|
||||
individuals) and `family`, realm-per-tenant is as economically impossible here
|
||||
as cluster-per-tenant — the 5–20 ceiling is well below our target tenant count.
|
||||
|
||||
**Proposed amendment:** note in §3 of the ADR that the P plane, as written,
|
||||
scopes to data substrate; identity-provider placement is a parallel decision
|
||||
with the same silo/pool shape and should be recorded, probably by whoever owns
|
||||
`key-cape`.
|
||||
|
||||
## 5. Freshness: the tradeoff we already made, confirmed
|
||||
|
||||
Our I3 requires live `tenant-engine` re-query for privileged, destructive,
|
||||
credential-vending or `aal2`-class decisions, with the cached `tenant_roles`
|
||||
claim acceptable for ordinary ones. External practice supports the split and
|
||||
names the cost precisely.
|
||||
|
||||
- The tradeoff is explicit: "tighter revocation controls often increase latency
|
||||
and operational overhead, requiring organisations to balance immediate cutoff
|
||||
against service reliability"
|
||||
([NHI Mgmt Group — JWT vs OAuth revocation](https://nhimg.org/faq/why-do-jwt-revocation-and-oauth-revocation-need-different-handling/)).
|
||||
- For high-stakes decisions the sanctioned answer is opaque tokens plus
|
||||
introspection (RFC 7662) — "a real kill switch at the cost of a lookup on the
|
||||
hot path" ([oneuptime — how to handle JWT revocation](https://oneuptime.com/blog/post/2026-02-02-jwt-revocation/view)).
|
||||
- Short lifetimes are the mainstream mitigation: "a 5-minute access token caps
|
||||
how stale any resource server's view can be, which is the quiet reason the
|
||||
default is 5 minutes and not an hour" (ibid.). Our IAM Profile's 5–30 minute
|
||||
service/agent token lifetime sits inside this norm.
|
||||
- Known failure mode directly relevant to us: controls "break down when
|
||||
services cache JWT validation results" (ibid.). Our hybrid explicitly permits
|
||||
caching for ordinary decisions, so the boundary between ordinary and
|
||||
high-stakes is load-bearing and must be enforced, not documented.
|
||||
- **Token versioning** is an alternative we do not use: a `token_version` per
|
||||
subject, bumped on change, rejecting stale tokens service-side
|
||||
([TechNet Experts — JWT token versioning](https://www.technetexperts.com/jwt-token-versioning-stale-data/)).
|
||||
It gives cheap global invalidation without a per-request lookup, at the cost
|
||||
of a version read. Worth considering as a middle rung between our cached
|
||||
claim and full live re-query.
|
||||
|
||||
**Against our model:** we match best practice and, unusually, we wrote the
|
||||
boundary down. The gap is evidential rather than architectural — our I3
|
||||
evidence artifact ("live re-query demonstrated on an `aal2`-class path; cached
|
||||
path shown not to be used there") is exactly right and does not exist yet for
|
||||
any service.
|
||||
|
||||
## 6. Finding: our tenant identifier encodes a mutable attribute
|
||||
|
||||
This is the significant new finding, and it is a critique of ratified canon
|
||||
rather than of ADR-008.
|
||||
|
||||
Our identifier is `tenant:<grouping>:<name>`, where `<grouping>` is an
|
||||
onboarding-risk / entity-shape classification from ADR-0013: `trial`,
|
||||
`friendly`, `single`, `small` (≤10 employees), `medium` (≤100), `large`
|
||||
(≤1000), `enterprise` (1001+), `consumer`, `family`, `community`,
|
||||
`association`, `agentic`.
|
||||
|
||||
Several of those are **explicitly headcount bands**, and headcount changes. A
|
||||
tenant onboarded as `tenant:small:acme` that grows past ten employees is
|
||||
either misclassified forever or must be renamed — and renaming a tenant
|
||||
identifier is a estate-wide data migration, since the identifier is the key
|
||||
shared with `user-engine`, appears in OpenBao paths (`tenants/<name>/...`),
|
||||
and is carried in every token.
|
||||
|
||||
External practice is unambiguous:
|
||||
|
||||
- "Avoid encoding mutable data or business logic in the identifier" — the
|
||||
**ID Entanglement Effect**, where composite ids encode business relationships
|
||||
or mutable state that is later parsed
|
||||
([Cogna — the unreasonable power of identifiers](https://cogna.co/blog/the-unreasonable-power-of-identifiers/)).
|
||||
- Opaque identifiers "decouple identification from descriptive attributes,
|
||||
thereby enhancing stability against entity changes"
|
||||
([Grokipedia — unique identifier](https://grokipedia.com/page/Unique_identifier)).
|
||||
- The FAIR/persistent-identifier literature reaches the same conclusion: make
|
||||
identifiers opaque from the start, so that when a name changes "just the
|
||||
`rdfs:label` and some other properties may change — but not the URI"
|
||||
([Cool URIs for FAIR Knowledge Graphs](https://arxiv.org/pdf/2407.09237)).
|
||||
- A structured id is acceptable *only if treated as opaque* — "the key is that
|
||||
we treat it as an unchanging reference and don't try to parse meaning from
|
||||
its structure" ([Cogna](https://cogna.co/blog/the-unreasonable-power-of-identifiers/)).
|
||||
|
||||
The IAM Profile half-anticipates this. It states that capability roles were
|
||||
deliberately separated from the identifier so "a tenant's roles may change
|
||||
without renaming its identifier", and that grouping "does not gate which
|
||||
capability roles a tenant may hold". The same reasoning applies to grouping
|
||||
itself and was not carried through.
|
||||
|
||||
**Two honest readings**, and this needs an owner decision rather than an
|
||||
assertion from this document:
|
||||
|
||||
1. **Grouping is onboarding-time and frozen by design.** It records the shape
|
||||
of the entity *when it was onboarded*, not its current headcount, and is
|
||||
never expected to change. If so, this should be stated explicitly in
|
||||
ADR-0013 — because every reader will otherwise assume `small` means
|
||||
currently-small, and some system will eventually make a decision on it.
|
||||
2. **Grouping is current classification.** Then the identifier encodes a
|
||||
mutable attribute, the estate has a rename problem it has not hit yet
|
||||
because tenant counts are tiny, and the fix (opaque id + grouping as a
|
||||
mutable property on the tenant record) gets more expensive every month.
|
||||
|
||||
Either way, `rapp-postgres` is already correct to treat these as opaque
|
||||
strings — ADR-0001 §4 prohibits the storage platform from parsing them. The
|
||||
risk sits with any consumer that *does* parse the grouping segment.
|
||||
|
||||
**Proposed action:** raise as a question to `tenant-engine` and NetKingdom
|
||||
canon rather than folding into ADR-008. It is out of ADR-008's scope but
|
||||
materially affects plane I's top rung.
|
||||
|
||||
## 7. Summary against our ladder
|
||||
|
||||
| Finding | Verdict |
|
||||
|---|---|
|
||||
| Bind tenant at IdP, carry in token, treat as context not parameter | **Aligned.** Matches I2 + IAM Profile. |
|
||||
| I2 does not require the binding be IdP-sourced | **Gap.** Tighten wording; request-supplied ids belong at I1. |
|
||||
| Propagation must cover caches, search indices, background jobs | **Gap.** ADR-008's ladders are datastore-shaped. Add a scope note. |
|
||||
| Identity-provider placement is its own silo/pool decision, ceiling ~5–20 tenants for realm-per-tenant | **Missing.** P plane scopes to data only; identity placement is undecided and unowned. |
|
||||
| Hybrid cached-claim / live-re-query for high-stakes | **Aligned and unusually well specified.** Evidence artifact does not exist yet. |
|
||||
| Token versioning as a middle option | **Unconsidered.** Cheap global invalidation without hot-path lookup. |
|
||||
| Identifier encodes mutable headcount bands | **Finding.** Contradicts identifier-design consensus; needs an ADR-0013 clarification. |
|
||||
|
||||
## Sources
|
||||
|
||||
- [AWS SaaS Lens — Preventing cross-tenant access](https://docs.aws.amazon.com/wellarchitected/latest/saas-lens/preventing-cross-tenant-access.html)
|
||||
- [AWS Well-Architected SaaS Lens (PDF)](https://docs.aws.amazon.com/pdfs/wellarchitected/latest/saas-lens/wellarchitected-saas-lens.pdf)
|
||||
- [AWS Prescriptive Guidance — Tenant onboarding and user tenant registration](https://docs.aws.amazon.com/prescriptive-guidance/latest/saas-multitenant-api-access-authorization/avp-design-onboarding-registration.html)
|
||||
- [AWS — 5 multi-tenant SaaS architecture best practices](https://aws.amazon.com/isv/resources/5-multi-tenant-saas-architecture/)
|
||||
- [AWS Security Blog — SaaS tenant isolation with ABAC using STS support for tags in JWT](https://aws.amazon.com/blogs/security/saas-tenant-isolation-with-abac-using-aws-sts-support-for-tags-in-jwt/)
|
||||
- [WorkOS — The developer's guide to SaaS multi-tenant architecture](https://workos.com/blog/developers-guide-saas-multi-tenant-architecture)
|
||||
- [SSOJet — Multi-tenant identity management for SaaS](https://ssojet.com/blog/multi-tenant-identity-management)
|
||||
- [Phase Two — Understanding multi-tenancy options in Keycloak](https://phasetwo.io/blog/multi-tenancy-options-keycloak/)
|
||||
- [Skycloak — Multitenancy in Keycloak using the Organizations feature](https://skycloak.io/blog/multitenancy-in-keycloak-using-the-organizations-feature/)
|
||||
- [Izylife — Keycloak multi-tenancy: realm per tenant or Organizations?](https://izylifesolutions.com/articles/keycloak-multi-tenancy-b2b-iam/)
|
||||
- [KeycloakPro — Keycloak multi-tenancy with Organizations](https://keycloakpro.com/blog/keycloak-multi-tenancy-organizations-guide)
|
||||
- [oneuptime — How to handle JWT revocation](https://oneuptime.com/blog/post/2026-02-02-jwt-revocation/view)
|
||||
- [NHI Management Group — Why JWT revocation and OAuth revocation need different handling](https://nhimg.org/faq/why-do-jwt-revocation-and-oauth-revocation-need-different-handling/)
|
||||
- [TechNet Experts — How to fix stale JWTs via token versioning](https://www.technetexperts.com/jwt-token-versioning-stale-data/)
|
||||
- [Cogna — The unreasonable power of identifiers](https://cogna.co/blog/the-unreasonable-power-of-identifiers/)
|
||||
- [Grokipedia — Unique identifier](https://grokipedia.com/page/Unique_identifier)
|
||||
- [Cool URIs for FAIR Knowledge Graphs (arXiv 2407.09237)](https://arxiv.org/pdf/2407.09237)
|
||||
Loading…
Add table
Add a link
Reference in a new issue