the-custodian/research/2026-08-17-adr008-plane-i-identity.md
codex 04f888399f research: external survey of multi-tenancy practice, one digest per ADR-008 plane
Four search-and-read passes against AWS Well-Architected SaaS Lens, Azure
Architecture Center, OWASP, the OpenID Foundation, and PostgreSQL practitioner
literature. Each digest holds one of our ladders against outside practice and
lists its sources.

Headline: the draft-1 -> draft-2 reframe is corroborated. Microsoft states
"isolation is a spectrum, not a discrete property" and derives our E/P coupling
independently. The shape is sound; the details are not.

Five corrections draft-2 needs:

- E3 overclaims. Any session can re-SET the GUC, so RLS is strong against a
  forgotten predicate and useless against SQL injection or a compromised
  process. The ADR says "the application cannot trivially route around" - the
  exact overclaim it prohibits elsewhere.
- The pooling claim in section 16 is backwards. Transaction-scoped SET LOCAL is
  what makes E3 safe under pooling; statement-level pooling is what serves
  other tenants' rows, only under production concurrency.
- E3 layers on E2 rather than superseding it.
- Placement cannot be expressed per tenant, but the commercial section requires
  exactly that - a tier bought by some tenants puts a service at two P levels
  at once, which the posture vector cannot represent.
- FORCE ROW LEVEL SECURITY is mandatory: ADR-0001 already established that our
  migration role owns the tables it creates, so it would silently bypass every
  policy.

Also: pgTAP and pgrls make the E3 evidence artifact far cheaper than assumed;
OWASP publishes a Multi-Tenant Security Cheat Sheet we reference nowhere; our
tenant-boundary gap is BOLA, OWASP API1, top of that list since it launched;
AuthZEN 1.0 went Final in January and flex-auth is pre-standard; and no plane
covers tenant offboarding, which OWASP makes a MUST.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 13:00:02 +02:00

234 lines
15 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Plane I — Identity: prior knowledge and best practices
> Research artefact — 2026-08-17. Supports `canon/architecture/adr-008-multi-tenancy-model.md`
> (draft-2), plane I of four. Surveys external practice for how multi-tenant
> systems name, carry, and validate tenant identity, and holds our `I0I3`
> ladder against it.
> Companions: plane A (authorization), plane E (enforcement), plane P (placement).
> Index: `2026-08-17-adr008-multi-tenancy-research-index.md`.
## 1. What the plane covers
How a tenant is named, where that name is bound to a principal, how it travels
with a request, and how fresh the binding must be. Our ladder:
| Level | State |
|---|---|
| I0 | No tenant concept |
| I1 | Local tenant notion, non-canonical |
| I2 | Canonical `tenant:<grouping>:<name>`, validated; `tenant-engine` owns existence |
| I3 | I2 + capability roles honoured, live re-query for `aal2`-class decisions |
## 2. Consensus practice: bind at registration, carry in the token
The dominant pattern across vendor guidance is that tenant identity is bound
to the principal **at the identity provider**, not derived per-application.
- AWS calls the unpacked token claims the **tenant context**, and frames it as
environment-like: it "influences how the SaaS application processes the
request" rather than being one parameter among many
([SaaS Lens — preventing cross-tenant access](https://docs.aws.amazon.com/wellarchitected/latest/saas-lens/preventing-cross-tenant-access.html)).
- Binding is explicitly a **registration-time** responsibility: storing the
tenant identifier as a claim or attribute on the user in the IdP "shifts the
responsibility of mapping identities to tenants from each application to the
user registration process"
([AWS Prescriptive Guidance — tenant onboarding and registration](https://docs.aws.amazon.com/prescriptive-guidance/latest/saas-multitenant-api-access-authorization/avp-design-onboarding-registration.html)).
- The recommendation is to define tenant context **at the IAM layer** rather
than relying on the application layer
([AWS — 5 multi-tenant SaaS architecture best practices](https://aws.amazon.com/isv/resources/5-multi-tenant-saas-architecture/)).
**Against our model:** this is what `I2` describes and what the IAM Profile
already mandates (`tenant` claim required on every token accepted by profile
consumers). We are aligned with consensus here, and `key-cape`/Keycloak
stamping the claim at issuance is the sanctioned shape.
**Gap:** the ladder's I2 says "canonical identifiers used, validated". It does
not say *where the binding happens*. A service could satisfy I2 by parsing a
tenant id out of a request body. Consider tightening I2 to require the binding
be IdP-sourced, with request-supplied tenant ids explicitly at I1.
## 3. Propagation is a separate discipline from naming
Practice treats "the tenant id must reach every layer" as its own requirement,
not a consequence of having one:
- A stable identifier should follow every request and stay consistent across
databases, services, logs, billing records, and background processes; without
a reliable tenant registry, other isolation controls are hard to enforce
consistently ([WorkOS — developer's guide to SaaS multi-tenant architecture](https://workos.com/blog/developers-guide-saas-multi-tenant-architecture)).
- Requests should execute in a tenant context established by middleware or
context variables, tagging each request with its owner
([SSOJet — multi-tenant identity management](https://ssojet.com/blog/multi-tenant-identity-management)).
- Two leak surfaces get called out repeatedly and are **not** database
concerns: search indices (Elasticsearch queries that omit the tenant filter)
and cache keys that do not include the tenant id
([WorkOS](https://workos.com/blog/developers-guide-saas-multi-tenant-architecture)).
**Against our model:** ADR-008 has nothing to say about background jobs,
caches, or search indices. Our E ladder is described almost entirely in
database terms (tenant-keyed tables, RLS, per-tenant credentials). A service
could sit honestly at `E3` and still leak through a Redis key or an unfiltered
index.
**Proposed amendment:** either widen the E-ladder evidence artifacts to cover
non-database stores, or add an explicit scope note that the E ladder describes
the *primary datastore* only and other stores are assessed separately. The
second is cheaper and more honest.
## 4. Identity-plane placement is its own decision — and it has a known ceiling
Identity providers face the same silo/pool question the P plane describes, and
the industry has converged on an answer with a number attached.
- Keycloak offers **realm-per-tenant** (strong isolation, heavyweight) versus
**Organizations** (multiple tenants inside one realm, stable since Keycloak
26, designed for the SaaS case)
([Phase Two — multi-tenancy options in Keycloak](https://phasetwo.io/blog/multi-tenancy-options-keycloak/),
[Skycloak — multitenancy using Organizations](https://skycloak.io/blog/multitenancy-in-keycloak-using-the-organizations-feature/)).
- The operational verdict is blunt: "realm-per-tenant works for 520 tenants;
beyond that, you are managing an IAM platform, not building your product"
([Izylife — Keycloak multi-tenancy: realm per tenant or Organizations?](https://izylifesolutions.com/articles/keycloak-multi-tenancy-b2b-iam/)).
- Auth0's equivalent is Organizations within a single Auth0 tenant, with
per-organization connections
([KeycloakPro — Keycloak multi-tenancy with Organizations](https://keycloakpro.com/blog/keycloak-multi-tenancy-organizations-guide)).
**Against our model:** ADR-008 treats placement as a property of *data
substrate* only. Identity has a placement axis too, it is live in our estate
(`key-cape`/Keycloak), and nobody has decided realm-per-tenant versus
Organizations. Given the tenant taxonomy includes `consumer` (private
individuals) and `family`, realm-per-tenant is as economically impossible here
as cluster-per-tenant — the 520 ceiling is well below our target tenant count.
**Proposed amendment:** note in §3 of the ADR that the P plane, as written,
scopes to data substrate; identity-provider placement is a parallel decision
with the same silo/pool shape and should be recorded, probably by whoever owns
`key-cape`.
## 5. Freshness: the tradeoff we already made, confirmed
Our I3 requires live `tenant-engine` re-query for privileged, destructive,
credential-vending or `aal2`-class decisions, with the cached `tenant_roles`
claim acceptable for ordinary ones. External practice supports the split and
names the cost precisely.
- The tradeoff is explicit: "tighter revocation controls often increase latency
and operational overhead, requiring organisations to balance immediate cutoff
against service reliability"
([NHI Mgmt Group — JWT vs OAuth revocation](https://nhimg.org/faq/why-do-jwt-revocation-and-oauth-revocation-need-different-handling/)).
- For high-stakes decisions the sanctioned answer is opaque tokens plus
introspection (RFC 7662) — "a real kill switch at the cost of a lookup on the
hot path" ([oneuptime — how to handle JWT revocation](https://oneuptime.com/blog/post/2026-02-02-jwt-revocation/view)).
- Short lifetimes are the mainstream mitigation: "a 5-minute access token caps
how stale any resource server's view can be, which is the quiet reason the
default is 5 minutes and not an hour" (ibid.). Our IAM Profile's 530 minute
service/agent token lifetime sits inside this norm.
- Known failure mode directly relevant to us: controls "break down when
services cache JWT validation results" (ibid.). Our hybrid explicitly permits
caching for ordinary decisions, so the boundary between ordinary and
high-stakes is load-bearing and must be enforced, not documented.
- **Token versioning** is an alternative we do not use: a `token_version` per
subject, bumped on change, rejecting stale tokens service-side
([TechNet Experts — JWT token versioning](https://www.technetexperts.com/jwt-token-versioning-stale-data/)).
It gives cheap global invalidation without a per-request lookup, at the cost
of a version read. Worth considering as a middle rung between our cached
claim and full live re-query.
**Against our model:** we match best practice and, unusually, we wrote the
boundary down. The gap is evidential rather than architectural — our I3
evidence artifact ("live re-query demonstrated on an `aal2`-class path; cached
path shown not to be used there") is exactly right and does not exist yet for
any service.
## 6. Finding: our tenant identifier encodes a mutable attribute
This is the significant new finding, and it is a critique of ratified canon
rather than of ADR-008.
Our identifier is `tenant:<grouping>:<name>`, where `<grouping>` is an
onboarding-risk / entity-shape classification from ADR-0013: `trial`,
`friendly`, `single`, `small` (≤10 employees), `medium` (≤100), `large`
(≤1000), `enterprise` (1001+), `consumer`, `family`, `community`,
`association`, `agentic`.
Several of those are **explicitly headcount bands**, and headcount changes. A
tenant onboarded as `tenant:small:acme` that grows past ten employees is
either misclassified forever or must be renamed — and renaming a tenant
identifier is a estate-wide data migration, since the identifier is the key
shared with `user-engine`, appears in OpenBao paths (`tenants/<name>/...`),
and is carried in every token.
External practice is unambiguous:
- "Avoid encoding mutable data or business logic in the identifier" — the
**ID Entanglement Effect**, where composite ids encode business relationships
or mutable state that is later parsed
([Cogna — the unreasonable power of identifiers](https://cogna.co/blog/the-unreasonable-power-of-identifiers/)).
- Opaque identifiers "decouple identification from descriptive attributes,
thereby enhancing stability against entity changes"
([Grokipedia — unique identifier](https://grokipedia.com/page/Unique_identifier)).
- The FAIR/persistent-identifier literature reaches the same conclusion: make
identifiers opaque from the start, so that when a name changes "just the
`rdfs:label` and some other properties may change — but not the URI"
([Cool URIs for FAIR Knowledge Graphs](https://arxiv.org/pdf/2407.09237)).
- A structured id is acceptable *only if treated as opaque* — "the key is that
we treat it as an unchanging reference and don't try to parse meaning from
its structure" ([Cogna](https://cogna.co/blog/the-unreasonable-power-of-identifiers/)).
The IAM Profile half-anticipates this. It states that capability roles were
deliberately separated from the identifier so "a tenant's roles may change
without renaming its identifier", and that grouping "does not gate which
capability roles a tenant may hold". The same reasoning applies to grouping
itself and was not carried through.
**Two honest readings**, and this needs an owner decision rather than an
assertion from this document:
1. **Grouping is onboarding-time and frozen by design.** It records the shape
of the entity *when it was onboarded*, not its current headcount, and is
never expected to change. If so, this should be stated explicitly in
ADR-0013 — because every reader will otherwise assume `small` means
currently-small, and some system will eventually make a decision on it.
2. **Grouping is current classification.** Then the identifier encodes a
mutable attribute, the estate has a rename problem it has not hit yet
because tenant counts are tiny, and the fix (opaque id + grouping as a
mutable property on the tenant record) gets more expensive every month.
Either way, `rapp-postgres` is already correct to treat these as opaque
strings — ADR-0001 §4 prohibits the storage platform from parsing them. The
risk sits with any consumer that *does* parse the grouping segment.
**Proposed action:** raise as a question to `tenant-engine` and NetKingdom
canon rather than folding into ADR-008. It is out of ADR-008's scope but
materially affects plane I's top rung.
## 7. Summary against our ladder
| Finding | Verdict |
|---|---|
| Bind tenant at IdP, carry in token, treat as context not parameter | **Aligned.** Matches I2 + IAM Profile. |
| I2 does not require the binding be IdP-sourced | **Gap.** Tighten wording; request-supplied ids belong at I1. |
| Propagation must cover caches, search indices, background jobs | **Gap.** ADR-008's ladders are datastore-shaped. Add a scope note. |
| Identity-provider placement is its own silo/pool decision, ceiling ~520 tenants for realm-per-tenant | **Missing.** P plane scopes to data only; identity placement is undecided and unowned. |
| Hybrid cached-claim / live-re-query for high-stakes | **Aligned and unusually well specified.** Evidence artifact does not exist yet. |
| Token versioning as a middle option | **Unconsidered.** Cheap global invalidation without hot-path lookup. |
| Identifier encodes mutable headcount bands | **Finding.** Contradicts identifier-design consensus; needs an ADR-0013 clarification. |
## Sources
- [AWS SaaS Lens — Preventing cross-tenant access](https://docs.aws.amazon.com/wellarchitected/latest/saas-lens/preventing-cross-tenant-access.html)
- [AWS Well-Architected SaaS Lens (PDF)](https://docs.aws.amazon.com/pdfs/wellarchitected/latest/saas-lens/wellarchitected-saas-lens.pdf)
- [AWS Prescriptive Guidance — Tenant onboarding and user tenant registration](https://docs.aws.amazon.com/prescriptive-guidance/latest/saas-multitenant-api-access-authorization/avp-design-onboarding-registration.html)
- [AWS — 5 multi-tenant SaaS architecture best practices](https://aws.amazon.com/isv/resources/5-multi-tenant-saas-architecture/)
- [AWS Security Blog — SaaS tenant isolation with ABAC using STS support for tags in JWT](https://aws.amazon.com/blogs/security/saas-tenant-isolation-with-abac-using-aws-sts-support-for-tags-in-jwt/)
- [WorkOS — The developer's guide to SaaS multi-tenant architecture](https://workos.com/blog/developers-guide-saas-multi-tenant-architecture)
- [SSOJet — Multi-tenant identity management for SaaS](https://ssojet.com/blog/multi-tenant-identity-management)
- [Phase Two — Understanding multi-tenancy options in Keycloak](https://phasetwo.io/blog/multi-tenancy-options-keycloak/)
- [Skycloak — Multitenancy in Keycloak using the Organizations feature](https://skycloak.io/blog/multitenancy-in-keycloak-using-the-organizations-feature/)
- [Izylife — Keycloak multi-tenancy: realm per tenant or Organizations?](https://izylifesolutions.com/articles/keycloak-multi-tenancy-b2b-iam/)
- [KeycloakPro — Keycloak multi-tenancy with Organizations](https://keycloakpro.com/blog/keycloak-multi-tenancy-organizations-guide)
- [oneuptime — How to handle JWT revocation](https://oneuptime.com/blog/post/2026-02-02-jwt-revocation/view)
- [NHI Management Group — Why JWT revocation and OAuth revocation need different handling](https://nhimg.org/faq/why-do-jwt-revocation-and-oauth-revocation-need-different-handling/)
- [TechNet Experts — How to fix stale JWTs via token versioning](https://www.technetexperts.com/jwt-token-versioning-stale-data/)
- [Cogna — The unreasonable power of identifiers](https://cogna.co/blog/the-unreasonable-power-of-identifiers/)
- [Grokipedia — Unique identifier](https://grokipedia.com/page/Unique_identifier)
- [Cool URIs for FAIR Knowledge Graphs (arXiv 2407.09237)](https://arxiv.org/pdf/2407.09237)