Tenancy Posture: question 3 has an owner

whitehat-security takes the adversarial evidence artifacts - the framework's
highest-severity gap, unowned since it was drafted. audit-core and tenant-engine
were right to decline it as fleet-scope work; the answer was a home of its own
rather than a volunteer.

Recorded here with the part that bears on this document: the facility is
deliberately not owned by NetKingdom, which owns this framework. Verifying
conformance to a standard while reporting to the standard's owner is
self-grading one level up.

Two consequences land back on the framework. Cadence becomes a security
parameter rather than a schedule, since for a detection-based control the
interval between runs is the exposure window. And a passing suite is proof that
the attacks attempted did not work, not proof of isolation - recording a green
run as "E2 verified" would be exactly the overclaim section 6 prohibits.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
tegwick 2026-08-17 16:41:56 +02:00
parent c31a5e3f8b
commit 101e659725

View file

@ -0,0 +1,809 @@
---
id: netkingdom-tenancy-posture
type: standard
title: "NetKingdom Tenancy Posture v0.1"
domain: netkingdom
status: proposed
version: "0.1"
created: "2026-08-17"
updated: "2026-08-17"
scope: multi-tenancy-security-framework
revision: "draft-5"
adr:
- docs/adr/ADR-0006-recursive-multi-tenant-identity-authorization.md
- docs/adr/ADR-0013-tenant-onboarding-grouping-taxonomy.md
- docs/adr/ADR-0014-tenant-capability-roles-and-tenant-engine-ownership.md
related:
- canon/standards/iam-profile_v0.3.md
- canon/standards/tenant-engine-boundary-contract_v0.1.md
- canon/standards/credential-management_v0.2.md
- docs/platform-identity-security-architecture.md
---
# NetKingdom Tenancy Posture v0.1 — Five Axes, Graduated Levels, Declared Conformance
## Status
**Proposed, draft-5.** Relocated from `the-custodian/canon/architecture` on
2026-08-17: multi-tenancy is part of the IT-security framework NetKingdom
provides, so this framework belongs in NetKingdom canon beside the IAM Profile
and the tenant-engine boundary contract, not in the work-factory canon.
- **draft-1** proposed a single model with fixed characteristics. Rejected: it
could not describe a repo that is not there yet.
- **draft-2** reframed to graduated levels per axis. Externally corroborated
(§16), but four of its statements were wrong and one thing it needed was
missing.
- **draft-3** applied those corrections, added the retention axis, and
recorded an adoption stance.
- **draft-4** closed the two gaps draft-3 left open: `R4` had no mechanism
beyond waiting, and the noisy-neighbour evidence artifact asserted something
shared infrastructure cannot provide.
- **draft-5** relocates to NetKingdom and renames the dimensions from *planes*
to *axes*, because the word was already taken (§0).
Every correction so far was found by research or by relocation, not by review.
Informed by five external research digests in `research/2026-08-17-adr008-*`,
which carry full citations for every external claim made here.
Reviewed by nobody yet. §19 lists what each owner is being asked to accept.
## 0. Terminology: axes, not planes
`docs/platform-identity-security-architecture.md` — accepted, 2026-07-23 —
already uses **plane** for a trust and deployment layer: the *bootstrap plane*,
the *platform control plane*, and *tenant planes*. That meaning is established,
ratified, and owned by this repo.
Drafts 14 of this document, written elsewhere, used **plane** for something
different: an independent dimension of concern. Two incompatible senses of one
word inside one canon is exactly the concept-ownership collision the estate has
been careful about elsewhere, and the newcomer yields.
This framework therefore describes five **axes**. They are orthogonal to
NetKingdom's planes, not a subdivision of them:
- A **plane** is *where* something runs and what trust it carries — bootstrap,
platform control, tenant.
- An **axis** is *which property* of tenancy is being described — identity,
authorization, enforcement, placement, retention.
A workload in the tenant plane has a position on all five axes. A platform
control plane service does too. The two vocabularies compose and neither
replaces the other.
The rename is also an improvement. A posture vector is literally a point in
five-dimensional space, and "axis" says that where "plane" did not.
## 1. Context
Drafts 14 opened by claiming the estate "has never written down what it is
building". Relocation proved that wrong, and the correction is worth keeping
visible: `docs/platform-identity-security-architecture.md` has described the
trust model, the tenant model and a capability progression since 2026-07-23.
The accurate claim is narrower — **what was missing is a way to say how far a
given service has got, and to hold several answers at once.** Seven documents
cover slices of the subject and none of them does that:
| Document | Covers | Status |
|---|---|---|
| `iam-profile_v0.3` (NetKingdom) | Tenant identifier shape, `tenant_roles` claim, staleness rules | Ratified |
| `tenant-engine-boundary-contract_v0.1` (NetKingdom) | Who owns tenant records, roles, plan assignment | Ratified |
| `business-app-service-contract_v0.1` §1 (Custodian) | Business apps: instance-per-client, tenant-keyed data | Ratified |
| `rapp-postgres` ADR-0001 | Consumer + tenant isolation in PostgreSQL | Proposed, governs one repo |
| `rapp-postgres` ADR-0002 | Per-consumer retention and the erasure horizon | Proposed, governs one repo |
| `shared-platform-relational-storage_v0.1` | The stacked-boundary gap | Routed 2026-08-10, **still unratified** |
| `platform-identity-security-architecture` (NetKingdom) | Trust model, planes, tenant model, capability progression | Accepted 2026-07-23 |
This document is downstream of that architecture and must not restate it. It
answers one question the architecture leaves open: given the model, **where is
this particular service today, and how would anyone know?**
Four failures follow.
**The gap was diagnosed once and the fix stalled.** The v0.1 draft was written
to fill this hole and has sat unratified in neither canon directory. §20
attaches a ratification path so this one does not join it.
**Placement is owned by nobody.** `user-engine-pg` and `target-revenue-pg` are
dedicated; `apps-pg`, `net-kingdom-pg`, `platform-pg`, `state-hub-db` and
`forgejo-db` are shared. Both live, neither written down. `tenant-engine`
raised this with `railiance-platform` on 2026-08-16; unanswered.
**Two contradictory defaults are already ratified.** Business apps get
instance-per-client; platform services pool. Nothing says which shape a new
service takes, and no definition separates the categories.
**There is no honest way to describe a repo that is not there yet.** The estate
absorbs repos with weak or absent tenant separation. Today such a repo is
simply non-conformant, leaving it two bad options: misrepresent its posture, or
stay outside the framework.
## 2. What this document is
**A framework, not a model.** It specifies no single correct implementation. It
supplies terminology (§3, §4), a declaration (§5), a conformance rule (§6),
methodology (§12), and evidence definitions (§13).
A service is conformant when its declared posture is accurate and its
trajectory recorded. A service is non-conformant when it claims a level it
cannot evidence — regardless of how high or low that level is.
## 3. Five orthogonal axes
"Is this multi-tenant?" is treated as one question. It is five, and they are
independent:
| Axis | Question | Vocabulary owner |
|---|---|---|
| **Identity (I)** | How is a tenant named and validated? | `tenant-engine` / IAM Profile |
| **Authorization (A)** | How is a request bound to the tenants it may act for? | `flex-auth` |
| **Enforcement (E)** | Where, mechanically, is the tenant boundary enforced? | This framework |
| **Placement (P)** | Which substrate holds a tenant's data? | `railiance-platform` |
| **Retention (R)** | How long does data persist, and how is it erased? | The storage platform; policy by the consumer |
Conflation produces errors today. `rapp-postgres`'s `PostgresConsumer` carries
`tenantIsolation: consumer-service-boundary` — an **E**-axis fact in a
**P**-axis artifact, reading as though storage enforces something it does not.
The "dedicated versus shared" argument mixes P (capacity, blast radius) with E
(correctness).
The axes are separated *precisely so each may sit at a different level*.
**Decision 3.1:** every document, declaration and plan tier that says
"isolation" MUST name which axis it means.
**Decision 3.2:** the axes couple at their tops and the couplings MUST be
stated where they apply, not used to argue the axes are one:
- `E4` is reachable only at `P3` or above.
- `R`'s erasure horizon is bounded below by `P` — on shared substrate, a
consumer's horizon is the instance maximum (§4.5).
- `R4` by key destruction is bounded by the **key boundary**, which is an
E-axis property. Shredding a single tenant's data requires the application
to encrypt under a per-tenant key before writing; the storage platform cannot
supply it. **Reaching the top of the retention ladder is not a retention
project.**
**Decision 3.3 — scope.** The P and R ladders describe a service's **primary
datastore**. Caches, search indices, message queues and background jobs are
named leak surfaces in the external baselines and are assessed separately, not
covered by a posture vector. Saying so is honest; implying the vector covers
them would not be.
## 4. Graduated levels
Each axis carries an ordered ladder. Higher is stronger, not better: the right
level is the one a service can evidence and its risk warrants.
### 4.1 Identity (I)
| Level | State |
|---|---|
| **I0** | No tenant concept. Data not attributable to a tenant. |
| **I1** | A local tenant notion exists but is not canonical, **or** the tenant is taken from the request rather than from a verified token. |
| **I2** | Canonical identifiers, bound at the identity provider and carried as a verified claim; `tenant-engine` is the source of existence. |
| **I3** | I2 plus capability roles honoured, with live `tenant-engine` re-query for privileged, destructive, credential-vending or `aal2`-class decisions. |
I1 now explicitly absorbs request-supplied tenant identifiers. "Never trust
client-supplied tenant IDs without validation" is a named anti-pattern; a
service reading the tenant from a header is at I1 however canonical the string.
`business-app-service-contract` §2.1 sets app-local accounts as the v1 baseline
for business apps — a sanctioned low level with recorded triggers for moving
up. That is the pattern this framework generalises.
### 4.2 Authorization (A)
| Level | State |
|---|---|
| **A0** | No authorization, or tenant context not carried. |
| **A1** | Ad-hoc checks scattered through handlers. |
| **A2** | A single local authorization boundary; tenant context bound once, centrally. |
| **A3** | Decisions delegated to `flex-auth` as PDP, with live re-query where the IAM Profile requires it. |
| **A4** | A3 over a **standard** PDP interface (OpenID AuthZEN Authorization API 1.0), so the decision point is swappable and the enforcement point is not coupled to one engine's request shape. |
A4 is new. `flex-auth` uses a bespoke `CheckRequest` and a bespoke action
vocabulary, with action strings copied verbatim between repos to avoid
re-derivation — exactly the coupling AuthZEN removes. The specification reached
Final in January 2026 and Keycloak shipped experimental support in May. We are
not wrong, we are pre-standard, and the ladder should have somewhere to go.
**Internal service-to-service calls are in scope for this axis.** "Skipping
tenant validation for internal services" is a named anti-pattern, and our
estate is mostly internal calls — `flex-auth` calls `tenant-engine`
synchronously on the authorization path. A service identity acting on behalf of
a tenant must carry and revalidate tenant context to claim A2 or above.
### 4.3 Enforcement (E)
| Level | Mechanism |
|---|---|
| **E0** | None. Data not tenant-keyed; separation incidental or absent. |
| **E1** | Data tenant-keyed, filtering applied per query at call sites. |
| **E2** | Filtering centralised at a single service-side choke point binding authenticated identity to permitted tenants. |
| **E3** | E2 **plus** platform-assisted filtering: row-level security keyed on a tenant GUC set transaction-locally, or an equivalent enforced data-access layer. |
| **E4** | Structural: the credential a workload holds cannot address another tenant's data at all. Requires per-tenant credentials and per-tenant substrate. |
**Correction from draft-2.** Draft-2 described E3 as something "the application
cannot trivially route around". That is false and it was this document
overclaiming in exactly the way §6 prohibits. Any session can re-issue `SET` on
a custom GUC, so an attacker with SQL execution can reset the tenant and read
across the boundary. What E3 buys is precise, and the ladder must say so:
| Threat | E1 | E2 | E3 | E4 |
|---|:--:|:--:|:--:|:--:|
| A developer forgets a tenant predicate | ✗ | ✓ | ✓ | ✓ |
| A new code path bypasses the choke point | ✗ | ✗ | ✓ | ✓ |
| SQL injection reaching the connection | ✗ | ✗ | ✗ | ✓ |
| The application process is compromised | ✗ | ✗ | ✗ | ✓ |
E3 is a strong control against **accident** — the common case, and the one that
causes real breaches — and no control at all against **compromise**. Only E4
holds against both, because the credential itself cannot address another
tenant's data.
**Correction: E3 layers on E2, it does not replace it.** External practice
treats application-layer and database-layer filtering as complementary. A
service that dropped its choke point on reaching E3 would be *worse* off, since
E3 fails open under injection. Claiming E3 therefore requires the E2 evidence
artifact as well.
**Correction: the GUC is set transaction-locally.** Draft-2 said "at pool
checkout", which is session scope and the wrong instrument. Under a pooler in
statement mode, `SET` leaks between clients and returns other tenants' rows —
a failure that appears only under production concurrency and produces no error.
Use `SET LOCAL` inside an explicit transaction.
**Platform enforcement is a platform obligation.** Reaching E3 requires the
storage platform to *offer* the mechanism: provisioned policies, a documented
GUC contract, and a probe. Where a consumer wants E3 and the platform has not
supplied it, the gap is the platform's. §19.6 asks `rapp-postgres` to define
that contract, which must carry `FORCE ROW LEVEL SECURITY` on every tenant
table (without it the table owner bypasses policies silently, and ADR-0001
already established that our migration role owns the tables it creates), no
`BYPASSRLS` on leased roles, `SECURITY INVOKER` for ordinary logic, and an
`EXPLAIN` comparison because RLS disables functional indexes built on
non-leakproof functions.
**Default expectation** for a new platform service: E2 at first serve, E3
recorded as target. Services whose cross-tenant exposure would be a reportable
breach SHOULD target E3 or above.
### 4.4 Placement (P)
| Level | Shape | Live occupants |
|---|---|---|
| **P0** | Shares a database with another consumer. | None sanctioned; the state absorbed repos arrive in. |
| **P1** | Database per consumer, shared cluster. | `audit-core`, `tenant-engine` on `platform-pg` |
| **P2** | Dedicated cluster per consumer. | `user-engine-pg`, `target-revenue-pg` |
| **P3** | Dedicated cluster per tenant. | Business apps per `business-app-service-contract` §1.2 |
| **P4** | P3 plus separate region or jurisdiction. | None |
Enforcement and placement are independent axes. Plotted together, with where
each service actually sits — parenthesised entries are targets or defaults
rather than current positions, and `—` marks a cell the coupling in §3.2 makes
unreachable:
| E \ P | P0 | P1 | P2 | P3 | P4 |
|---|---|---|---|---|---|
| **E4** | — | — | — | (business app) | |
| **E3** | | (target) | | | |
| **E2** | | tenant-engine<br>audit-core | | | |
| **E1** | (absorbed repo) | | | | |
| **E0** | | | | | |
**P0 → P1 → P2 is movement along the horizontal axis only.** Those steps buy
consumer isolation, capacity predictability, independent retention and a
smaller operational blast radius. They do not raise the tenant boundary by one
step. Only P3 makes E4 reachable. This is the most misusable fact in the
framework and §11 governs how it may be described.
**Decision 4.4.1:** P1 is the default for platform services; P3 for
client-facing business apps, as already ratified. A service unsure which it is
must resolve that first (§19.4).
**Decision 4.4.2 — placement scopes to data substrate.** Identity-provider
placement (realm-per-tenant versus Organizations) is the same silo/pool
decision on a different substrate, is live in our estate, and is undecided.
Realm-per-tenant carries a stated ceiling around 520 tenants, far below our
target. Recorded here as a parallel question (§19.7), not folded into P.
### 4.5 Retention and erasure (R)
New in draft-3. Implemented abstractly by the storage platform for any dataset;
policy is built on top of that interface by the consumer or its governance
layer. Reference implementation: `rapp-postgres` ADR-0002.
| Level | State |
|---|---|
| **R0** | No retention or deletion position. Data kept indefinitely by default; no deletion path exists. |
| **R1** | Platform default retention applies (N=30 days). The consumer has declared no requirement. |
| **R2** | Retention declared as N days per dataset; the **erasure horizon** is published, and the consumer makes no promise shorter than it. |
| **R3** | Policy-driven deletion: the consumer or its governance layer declares what is due, the platform sweeps whole datasets on that instruction and evidences each run. |
| **R4** | Verified erasure: data proven unrecoverable across live storage, backups and derived copies, by one of the two routes below. |
**R4 has two routes and a service MUST name which one it uses.**
| Route | Mechanism | Cost |
|---|---|---|
| **Horizon-elapsed** | Wait out the published erasure horizon; the data ages out of every retained copy. | Available to everyone, proves little, and the wait is set by a co-resident's retention requirement rather than your own. |
| **Key-destroyed** | Encrypt per entity, then destroy the key. Retained copies survive but are unreadable. | Requires per-entity keys, strong encryption, and an auditable destruction record. Immediate. |
**Regulatory standing of the key-destroyed route, stated carefully because
overclaiming here is worse than anywhere else in this framework.** Data
protection authorities have accepted key destruction as erasure where physical
deletion would be manifestly disproportionate, and the practice is recognised
under conditions — strong encryption, irreversible destruction, and an auditable
record of it. **The EDPB has not formally endorsed it as Article 17 erasure.** A
service reaching R4 by key destruction is making a defensible claim, not a
settled one, and must say so rather than reporting a clean "deleted".
Three further properties.
**The erasure horizon is the interval between deleting data and it ceasing to
be recoverable from anything the platform holds.** Deleting a row does not
remove it from yesterday's backup. With an N-day window, deleted data remains
recoverable for N days. That is the difference between "deleted" and "erased"
and the estate had never written it down.
**On shared substrate, retention is not per-consumer.** Physical backup is
instance-wide — one WAL stream, one window — so the instance retention is
*derived* as the maximum across co-resident consumers, and every consumer's
horizon is that maximum. A consumer declaring 7 days beside one declaring 90
gets 90. This is the retention analogue of ADR-0001's blast-radius disclosure:
state the coupling rather than imply an isolation that is not there.
**Retention is therefore a placement trigger.** A consumer needing a horizon
shorter than the instance floor cannot have one at P1. It moves to P2 for a
reason with nothing to do with performance — which is exactly why it needs
recording, since nobody looks for a retention argument when reviewing
placement.
Deletion splits mechanism from policy. The platform deletes whole **datasets**
on instruction and records an opaque policy reference it never interprets, so
every deletion traces to what authorised it. Rows are not a dataset: row expiry
is the consumer's own DML under its migration lease. Dropping a consumer's
whole database is an operator-gated offboarding step, never a scheduled one.
## 5. The posture vector
A service states one level per axis, plus a target, a date, and any placement
exceptions:
```yaml
tenancy:
current: { I: 2, A: 3, E: 2, P: 1, R: 1 }
target: { I: 2, A: 3, E: 3, P: 1, R: 2 }
reviewed: "2026-08-17"
gap:
E: "Choke point exists and is tested; RLS not provisioned. Blocked on
rapp-postgres publishing the GUC contract. Target Q4."
R: "Retention declared; erasure horizon not yet published to consumers."
```
**Placement exceptions.** Draft-2 assigned one P level per service, which
cannot express the vertically partitioned model — most tenants pooled, some
dedicated — that §11's isolation tiers require. A tier requiring `P2` bought by
three tenants would put the service at two levels at once, forcing an over- or
under-claim. Placement is therefore declared as a default plus exceptions:
```yaml
placement_exceptions:
- tenants: ["tenant:enterprise:*"]
P: 3
reason: "isolation tier; see adaptive-pricing tier definition"
```
A service with exceptions must be able to say which tenants are on which
substrate. That mapping is a first-class artifact, not archaeology.
Worked examples, best-effort and subject to owner correction:
| Service | Current | Notes |
|---|---|---|
| `tenant-engine` | `I2 A3 E2 P1 R1` | Moving to P1 under TEN-WP-0009; retention declared, horizon not yet published. |
| `audit-core` | `I2 A3 E2 P1 R1` | Holds audit evidence, so both E3 and R2 are urgent targets. |
| A newly absorbed repo | `I1 A1 E1 P0 R0` | Conformant **if declared**, with a recorded path. |
**Decision 5.1:** the posture vector is declared in the repo, not in the hub,
consistent with local-files-are-source-of-truth.
## 6. Conformance is accuracy, not altitude
> **A service is conformant when its declared posture is accurate, its target
> is recorded, and it does not claim a level it cannot evidence. It is
> non-conformant when it overclaims — at any altitude.**
- Declaring `E0` is conformant. Concealing `E0` is not.
- A repo may be absorbed at any posture. It may not be absorbed silently.
- No service is blocked from the estate for being low on a ladder. Services MAY
be blocked from *specific work* — serving a tenant grouping, holding a data
class, carrying a plan tier — by requirements expressed as minimum levels.
- Downgrading is permitted and must be declared. A regression found by guarding
is a defect; a regression declared in advance is a decision.
Without the axis separation, "not rigorous about tenant separation" is one
verdict a repo passes or fails. With it, the same repo is `I1 A1 E1 P0 R0` with
a path — a plan, not an indictment.
## 7. Portability across placement levels
Movement between P levels must be operational, not a rebuild:
- Connect by injected credential only — no cluster, host, namespace or database
name in source.
- Own a whole database, never tables inside someone else's.
- Idempotent schema creation.
- No cross-database joins or co-location assumptions.
**Decision 7.1:** mandatory at P1 and above. At P3, SHOULD rather than MUST — a
per-client instance that never moves is not misconformant for naming its own
database.
## 8. Placement triggers
Recorded at provisioning time: noisy neighbour on a latency-critical path; a
compliance or residency requirement; a plan tier requiring a higher minimum; an
erasure horizon that no longer fits (§4.5); connection or memory ceiling
reached.
**Decision 8.1:** triggers MUST be *monitored*, not merely recorded. A trigger
in a YAML comment nobody re-reads is documentation, not control.
**Decision 8.2:** placement policy ownership is proposed to
`railiance-platform`, **co-signed by `adaptive-pricing`**. Tenancy model
selection is a commercial decision as much as a technical one; an
operations-shaped repo should not hold it alone.
## 9. Credentials as a tenancy control
Short-lived leased credentials re-read at connection checkout, with
overlap-first rotation, bound the residual risk at every E level below E4: a
leaked credential expires rather than persisting. Stronger than the industry
norm of a long-lived per-service secret.
**Decision 9.1:** static long-lived database credentials are not a sanctioned
path for any service above E0.
## 10. Blast radius must be published
**Decision 10.1:** every platform holding consumer data MUST publish, in
concrete terms, what a leaked runtime credential can and cannot reach at the
levels it operates. `rapp-postgres` ADR-0001 §5 is the reference. Where the
model cannot provide a guarantee, the platform says so and names the
escalation.
**Decision 10.2 — quotas are disclosed, not discovered.** The same obligation
extends from what a leaked credential can reach to what the platform will
refuse to do for you. Every consumer MUST be told, at provisioning, the
throttles and quotas enforced against it — connection limits, statement
timeouts, idle-transaction timeouts — and told again when they change. A
consumer learning its statement timeout by hitting it in production is a
disclosure failure, not a consumer bug. This is how `tenant-engine` was
provisioned, by good practice rather than by rule; the rule now exists.
## 11. Commercial expression
- **11.1** Plan tiers are expressed *internally* as minimum levels. A tier may
require `E3 P2 R2`; it need not print that anywhere customer-facing.
- **11.2** Marketing and product language is free. No requirement to expose
level labels or this document. "Dedicated infrastructure", "isolated
tenancy", "private instance" all remain available.
- **11.3** The constraint is on **evidence, not vocabulary**. A customer-facing
isolation, availability or retention claim must map to a minimum level the
delivering service actually holds, recorded once when the tier is defined.
The review is internal and happens at tier definition — not per campaign.
- **11.4** Two hard lines, because these reach contracts and compliance
questionnaires:
- A claim that another tenant **cannot** reach the customer's data requires
**E4**.
- A claim that deleted data **is gone** requires **R4**, or an erasure
horizon disclosed alongside it. Where R4 is reached by key destruction, the
claim is defensible but not settled law (§4.5) — it may be made, and it may
not be made in language that implies a regulator has blessed it.
## 12. Methodology — analyze, establish, improve, guard
**Analyze.** Assess a repo against the ladders; produce `tenancy.current` with
reasoning recorded. Applies to new and absorbed services alike.
**Establish.** Declare the target and gap. The target is set by data class,
tenant groupings served and plan tiers carried — not by ambition.
**Improve.** Move one axis at a time. Raising P while leaving E untouched is
the characteristic misstep.
**Guard.** Verify continuously that the declared posture holds — **against the
service's own declaration**, not a universal maximum. Nobody must prove every
service is at E4; the check is that none is below what it declared.
Regression found by guarding is a defect; regression declared in advance is a
decision. The estate has been bitten twice by silent pin rollbacks producing
ordinary-looking 403s and 404s rather than errors. Posture regression looks the
same — an RLS context leak returns correct-looking rows for the wrong tenant.
Guarding must be designed for invisible failure, not for crashes.
## 13. Evidence per level
**Decision 13.1:** a level is claimed only with its evidence artifact present.
This turns §6's accuracy rule from an honour system into a check.
**Decision 13.4 — an artifact must assert something achievable.** Draft-3's
noisy-neighbour evidence required proof that a saturating consumer "does not
breach" another's allowance. Shared infrastructure cannot provide that; the
risk is inherent and cannot be wholly removed. An artifact that can only fail,
or that passes by being run gently enough, is an overclaim wearing the costume
of evidence. Where a property cannot be guaranteed, the artifact measures and
records it instead.
**Decision 13.2 — evidence is of two kinds, and conflating them is an
overclaim.** *Mechanical* evidence is a structural assertion a machine can make
and belongs in CI. *Adversarial* evidence is semantic, requires setting up
separate tenant contexts and comparing responses, and carries a review date
rather than a green build. Cross-tenant findings are the category external
testing practice identifies as needing human review. **A passing CI run is not
E2 evidence.**
| Level | Evidence | Kind |
|---|---|---|
| **I2** | Identifiers validated against the vocabulary; rejection test for a malformed id; binding shown to come from a verified token | Mechanical |
| **I3** | Live re-query demonstrated on an `aal2`-class path; cached-claim path shown unused there | Mechanical |
| **A2** | Choke point identified; test that an unbound request is refused | Mechanical |
| **A3** | Live decision with a denial observed at the endpoint, not only at the decision surface | Mechanical |
| **A4** | Decision served over the standard interface; a second PDP substituted without PEP change | Mechanical |
| **E1** | Every tenant-owned table carries the tenant key | Mechanical |
| **E2** | Choke point identified; identity bound to tenant A demonstrably cannot read tenant B | **Adversarial**, with a review date |
| **E3** | `FORCE ROW LEVEL SECURITY` on every tenant table; no `BYPASSRLS` on leased roles; probe that a session without the GUC reads nothing; probe that a wrong GUC reads nothing; `EXPLAIN` comparison | Mechanical |
| **E4** | Per-tenant credential demonstrated unable to connect to another tenant's substrate | Mechanical |
| **P1P4** | Provisioning declaration plus the platform's isolation probes | Mechanical |
| **P1P2 (noisy neighbour)** | A recorded baseline of per-consumer resource usage; a run in which one consumer saturates its declared allowance; evidence that the governance controls **bind** (the greedy consumer is held at its limits) and that the degradation co-residents experience is **measured, recorded and judged acceptable**; the aggregate headroom at time of measurement | **Adversarial**, load-generated, with a review date |
| **R2** | Declared retention rendered; erasure horizon published and reported in the operator surface | Mechanical |
| **R3** | Sweep evidence records: timestamp, dataset, identifiers removed, authorising policy reference | Mechanical |
| **R4** | Erasure demonstrated across live data, backups and derived copies within the horizon | **Adversarial** |
**Decision 13.3:** the E2, E3 and noisy-neighbour artifacts do not exist
anywhere in the estate today. `rapp-postgres` runs 15 adversarial probes, all
against the *consumer* boundary, none against the tenant boundary inside a
consumer. Externally, what this framework calls a tenant boundary failure is
**Broken Object Level Authorization** — OWASP API1, top of the API Security Top
10 since that list launched, and the most commonly exploited API vulnerability
in published assessments. We have no coverage for the highest-ranked risk in
our class of system. §19.3 seeks an owner.
## 14. Adoption stance — structure, not tooling
**Decision 14.1:** external research is design input. This estate adopts
published standards and structural patterns; it does not adopt tooling unless
that tooling is an established industry standard with broad application.
Everything else is built ground-up, so it can be optimised and refactored as
the estate sees fit.
| Class | Stance |
|---|---|
| Security baselines (OWASP Multi-Tenant Security Cheat Sheet, API Security Top 10) | Adopt as the external reference our ladders answer to |
| Standards bodies (OpenID AuthZEN 1.0) | Adopt — this is what A4 is |
| Reference taxonomies (Azure tenancy models, AWS SaaS Lens, cell architecture) | Adopt as structure |
| Engine behaviour (PostgreSQL RLS mechanics) | Facts, not tooling |
| Third-party analyzers and test frameworks | **Do not adopt.** Take their rule taxonomies as checklists for probes we write ourselves |
The practical effect is small and good: `rapp-postgres` already owns a
ground-up probe harness — bash and psql, no dependency tree — that found four
real defects in its own provisioning SQL. The evidence artifacts in §13 become
new probes in a tool we control. One idea worth reimplementing from the
external survey is **policy-diff classification**: labelling a change to an
enforcement policy as safe or breaking *before* it lands.
## 15. Alternatives considered
**One fixed model with a single set of characteristics** (draft-1). *Rejected:*
cannot describe a repo that is not there yet, forcing absorbed repos to
misrepresent their posture or stay outside. A framework that can only describe
its own end state is not a framework.
**A maturity model with a single overall level.** *Rejected:* collapses the
axis separation. A service strong on identity and weak on enforcement has a
specific, actionable gap; one composite score hides it and invites averaging.
**Prohibiting row-level security** (draft-2's inherited position). *Rejected in
draft-2, refined in draft-3:* RLS is a real rung against the common threat. The
error was never RLS — it was describing E3 in E4's language.
**Schema-per-consumer in one database.** *Rejected:* `pg_catalog` is readable
per-database, so every co-resident enumerates every other's table and column
names regardless of grants. Retained as a describable state, never a target.
**Mandating E4 for everyone.** *Rejected:* the tenant taxonomy includes
`consumer` (private individuals) and `family`. A cluster per private individual
is economically impossible; the taxonomy is itself evidence pooling is
required.
**Per-consumer physical backup retention.** *Rejected:* CNPG retention is a
property of the instance's WAL archive. There is no mechanism, and claiming it
would be a fabricated guarantee. Hence the derived maximum in §4.5.
**Platform-scheduled row expiry.** *Rejected:* requires the platform to hold
DML authority over consumer schemas and interpret consumer data semantics, both
forbidden by ADR-0001. The consumer's migration lease is the correct
instrument.
**Leaving each repo to its own model.** *Rejected:* the status quo, which
produced two contradictory ratified defaults and an unowned placement question.
## 16. Held against outside practice
**The graduated reframe is corroborated, not invented here.** Microsoft's
tenancy-model guidance states it almost verbatim: *"Instead of viewing
isolation as a discrete property, consider it a spectrum. You can deploy
components of your architecture that are more isolated or less isolated than
other components in the same architecture."* The same guidance derives our E↔P
coupling independently — shared deployment means enforcement lives in
application code; dedicated deployment means it is structural.
**Stronger than typical.** Most multi-tenancy literature models one boundary,
tenant-to-tenant. This estate has **two stacked boundaries**: platform-service
to platform-service, and tenant to tenant inside a consumer. Naming them
separately and refusing to enforce both with one mechanism is uncommon and
correct. Graduated per-axis levels also beat the silo/pool/bridge trichotomy,
which is approximately our P axis with the other four missing — which is why
it cannot express "pooled infrastructure, structurally enforced boundary".
**Weaker than typical.** The pool model's standard mitigation is a *verified*
enforcement layer every service is demonstrably routed through. We have the
concept and none of the verification (§13.3).
**Adopted without naming it.** Short-lived leased credentials re-read at
checkout beat the long-lived-secret norm. §9 promotes it to a tenancy control.
**Still unexplored.** Neither P nor R describes a **cell** — a slice of
infrastructure with a *fixed maximum size*, sized so one cell's failure is
survivable and cell count scales linearly. `platform-pg` is, in these terms, an
uncapped cell: §17 computes a ceiling and nothing enforces it (§19.8).
Sources: the four research digests in `research/2026-08-17-adr008-*`, which
carry full citations for every claim in this section.
## 17. Scaling demands
Measured against the live `platform-pg` specification, not estimated.
```
instances: 1 (no HA; single-node rail)
max_connections: 100
memory limit: 1Gi
per consumer: 14 connections (12 runtime + 2 migration)
```
**Connection ceiling: roughly six consumers — and this is the aggregate
noisy-neighbour bound, not a capacity statistic.** Seven consumers request 98
of 100 before CNPG's instance manager, metrics exporter and reserved slots.
Every one of them is politely inside its declared 14-connection allowance; the
instance still fails.
That distinction matters because our governance addresses the wrong shape.
Per-consumer `connection_limit`, `statement_timeout` and
`idle_in_transaction_session_timeout` guard well against **one greedy
consumer**. They do nothing about **the aggregate of many modest ones**, which
is the second and less intuitive noisy-neighbour failure and the one this
number describes. Two consumers are provisioned. We are at roughly a third of
the bound, and the third request will not feel like a scaling event.
**Memory likely binds first.** 100 backends against 1Gi is ~10MB per backend.
Connection exhaustion errors clearly; memory pressure OOM-kills and degrades
every co-resident at once.
**E3 and pooling.** *Corrected from draft-2, which had this backwards.*
Transaction-scoped context (`SET LOCAL` inside an explicit transaction) is what
makes E3 **safe** under a pooler. Statement-level pooling is what breaks it,
serving other tenants' rows under concurrency with no error. E3 constrains
which pooling mode is available, not whether pooling is available.
**Retention consumes the volume.** WAL accumulates with the window, and §4.5
makes the window the maximum across consumers. A consumer declaring a long
retention extends everyone's horizon *and* everyone's storage draw against a
20Gi volume.
**Restore time couples all consumers.** Physical backup is instance-wide, so a
consumer's RTO is a function of *total* instance size, not its own.
**No P1 tenant has HA.** `instances: 1` means a tier promising uptime cannot be
satisfied at P1 as built — an availability floor belongs in §11's
minimum-level vocabulary alongside isolation.
## 18. Consequences
- The estate gains one vocabulary and a way to be honest about partial
adoption.
- Absorbed repos get a described state and a path instead of a failing grade.
- `tenantIsolation` in `PostgresConsumer` is revealed as a mislabelled field.
- The verification problem becomes tractable: guard against declaration.
- Draft-2's RLS prohibition is reversed and its E3 description corrected;
`rapp-postgres` acquires an obligation to define and offer the mechanism.
- Adding a consumer with long retention **silently extends everyone's erasure
horizon**. This must reach the consumer review checklist, not only this
document.
- A service selling an isolation tier must maintain a tenant→substrate mapping
it does not have today.
- Nothing here changes a running system.
## 19. Open questions
1. **`tenantIsolation` field** — `rapp-postgres`: rename to name its axis and
carry a level (`tenancy.E: 2`), or move it out of the storage declaration.
2. **Placement ownership**`railiance-platform` with `adaptive-pricing`:
accept the ladder, triggers and the §8.1 monitoring obligation; appoint a
recorded placement owner per workload.
3. **E2, E3 and noisy-neighbour evidence****owned as of 2026-08-17** by
`whitehat-security` (WHITEHAT-WP-0001), an independent adversarial evidence
facility seeded for this purpose. `audit-core` and `tenant-engine` were
right to decline it as fleet-scope work; the answer was a home of its own
rather than a volunteer.
Independence is the design point, and it bears on this document: a facility
verifying conformance to this framework is deliberately **not** owned by
NetKingdom, which owns the framework. Self-grading one level up is still
self-grading.
Two consequences land back here. **Cadence is now a security parameter, not
a schedule** — for any control whose guarantee is detection rather than
prevention, the interval between probe runs *is* the exposure window, and
`rapp-postgres` ADR-0003 leaves that number to the facility. And **a passing
suite is not proof of isolation**; it is proof that the attacks attempted
did not work. §13's evidence artifacts should be read with that distinction,
because a green run recorded as "E2 verified" would be exactly the overclaim
§6 prohibits.
4. **Business app vs platform service** — Custodian canon: a classification
rule. Candidate: reuse `repo-classification-standard_v1.0`.
5. **Tier → minimum level mapping**`adaptive-pricing` and `tenant-engine`:
required only for tiers making isolation, availability or retention claims.
6. **The E3 mechanism**`rapp-postgres`: publish the GUC contract with the
`FORCE`/`BYPASSRLS`/`SECURITY INVOKER`/`EXPLAIN` requirements in §4.3.
7. **Identity-provider placement** — owner of `key-cape`: realm-per-tenant or
Organizations? Realm-per-tenant's ~520 tenant ceiling is below our target.
8. **Cell sizing** — reframed from "should we adopt cells" to **"what is
`platform-pg`'s declared maximum size, and what is the overflow target?"**
The connection ceiling forces this whether or not we adopt the vocabulary.
9. **Retention floor and ceiling** — should `backupRetentionDays` have a
platform minimum (so a consumer asking for 1 day gets a validation error
rather than a quiet disappointment) and a maximum (so nobody exhausts the
volume)?
10. **Engine neutrality** — the P ladder rests on a PostgreSQL property.
State it engine-specifically and say so, or abstract it and risk a
non-Postgres implementation that silently differs?
11. **Erasure versus audit**`audit-core`: crypto-shredding a tenant's audit
records destroys the evidence the service exists to hold, and ADR-0001 §2
deliberately built the role model so history could not be rewritten. The
usual resolution separates the *fact* of an event, retained, from its
*personal payload*, encrypted per subject and shreddable. Raised because a
naive "R4 everywhere" target would instruct the audit service to destroy
its own evidence. The answer is `audit-core`'s, not this framework's.
12. **Quality of service***owner needed.* The framework has no vocabulary
for saying one consumer's latency matters more than another's.
`tenant-engine` sits on `flex-auth`'s synchronous authorization path and
chose a 5s statement timeout for that reason; it shares an instance with
`audit-core`, which is not latency-critical. Nothing prioritises between
them. Either add a QoS dimension or state that all co-residents are equal
and latency-critical consumers must escalate to P2.
**Routed elsewhere, deliberately.** The tenant identifier
`tenant:<grouping>:<name>` embeds headcount bands (`small`, `medium`, `large`)
that change as a tenant grows, contradicting the consensus that identifiers
should not encode mutable attributes. That is a critique of ADR-0013, not of
this framework, and belongs to `tenant-engine` and NetKingdom canon. Folding it
in here would overreach.
## 20. Ratification path
1. Reviewed by `tenant-engine`, `flex-auth`, `rapp-postgres`,
`railiance-platform` and `adaptive-pricing` against §19.
2. Each publishes its own posture vector (§5) as part of review. **The
framework is validated by whether it can describe them accurately** — if a
repo cannot express itself in these five ladders, the ladders are wrong and
this document changes, not the repo.
3. On acceptance, **supersedes** the routing of
`rapp-postgres/docs/canon-drafts/shared-platform-relational-storage_v0.1-draft.md`,
whose §§38 are absorbed here. That draft is withdrawn rather than left
pending.
4. On acceptance, `rapp-postgres` ADR-0001 and ADR-0002 move to `accepted` and
are annotated as the PostgreSQL implementation of the E, P and R ladders.