From 101e65972521dd08b4b2c6e7809f658dd776d741 Mon Sep 17 00:00:00 2001 From: tegwick Date: Mon, 17 Aug 2026 16:41:56 +0200 Subject: [PATCH] Tenancy Posture: question 3 has an owner whitehat-security takes the adversarial evidence artifacts - the framework's highest-severity gap, unowned since it was drafted. audit-core and tenant-engine were right to decline it as fleet-scope work; the answer was a home of its own rather than a volunteer. Recorded here with the part that bears on this document: the facility is deliberately not owned by NetKingdom, which owns this framework. Verifying conformance to a standard while reporting to the standard's owner is self-grading one level up. Two consequences land back on the framework. Cadence becomes a security parameter rather than a schedule, since for a detection-based control the interval between runs is the exposure window. And a passing suite is proof that the attacks attempted did not work, not proof of isolation - recording a green run as "E2 verified" would be exactly the overclaim section 6 prohibits. Co-Authored-By: Claude Opus 5 --- canon/standards/tenancy-posture_v0.1.md | 809 ++++++++++++++++++++++++ 1 file changed, 809 insertions(+) create mode 100644 canon/standards/tenancy-posture_v0.1.md diff --git a/canon/standards/tenancy-posture_v0.1.md b/canon/standards/tenancy-posture_v0.1.md new file mode 100644 index 0000000..5469291 --- /dev/null +++ b/canon/standards/tenancy-posture_v0.1.md @@ -0,0 +1,809 @@ +--- +id: netkingdom-tenancy-posture +type: standard +title: "NetKingdom Tenancy Posture v0.1" +domain: netkingdom +status: proposed +version: "0.1" +created: "2026-08-17" +updated: "2026-08-17" +scope: multi-tenancy-security-framework +revision: "draft-5" +adr: + - docs/adr/ADR-0006-recursive-multi-tenant-identity-authorization.md + - docs/adr/ADR-0013-tenant-onboarding-grouping-taxonomy.md + - docs/adr/ADR-0014-tenant-capability-roles-and-tenant-engine-ownership.md +related: + - canon/standards/iam-profile_v0.3.md + - canon/standards/tenant-engine-boundary-contract_v0.1.md + - canon/standards/credential-management_v0.2.md + - docs/platform-identity-security-architecture.md +--- + +# NetKingdom Tenancy Posture v0.1 — Five Axes, Graduated Levels, Declared Conformance + +## Status + +**Proposed, draft-5.** Relocated from `the-custodian/canon/architecture` on +2026-08-17: multi-tenancy is part of the IT-security framework NetKingdom +provides, so this framework belongs in NetKingdom canon beside the IAM Profile +and the tenant-engine boundary contract, not in the work-factory canon. + +- **draft-1** proposed a single model with fixed characteristics. Rejected: it + could not describe a repo that is not there yet. +- **draft-2** reframed to graduated levels per axis. Externally corroborated + (§16), but four of its statements were wrong and one thing it needed was + missing. +- **draft-3** applied those corrections, added the retention axis, and + recorded an adoption stance. +- **draft-4** closed the two gaps draft-3 left open: `R4` had no mechanism + beyond waiting, and the noisy-neighbour evidence artifact asserted something + shared infrastructure cannot provide. +- **draft-5** relocates to NetKingdom and renames the dimensions from *planes* + to *axes*, because the word was already taken (§0). + +Every correction so far was found by research or by relocation, not by review. + +Informed by five external research digests in `research/2026-08-17-adr008-*`, +which carry full citations for every external claim made here. + +Reviewed by nobody yet. §19 lists what each owner is being asked to accept. + +## 0. Terminology: axes, not planes + +`docs/platform-identity-security-architecture.md` — accepted, 2026-07-23 — +already uses **plane** for a trust and deployment layer: the *bootstrap plane*, +the *platform control plane*, and *tenant planes*. That meaning is established, +ratified, and owned by this repo. + +Drafts 1–4 of this document, written elsewhere, used **plane** for something +different: an independent dimension of concern. Two incompatible senses of one +word inside one canon is exactly the concept-ownership collision the estate has +been careful about elsewhere, and the newcomer yields. + +This framework therefore describes five **axes**. They are orthogonal to +NetKingdom's planes, not a subdivision of them: + +- A **plane** is *where* something runs and what trust it carries — bootstrap, + platform control, tenant. +- An **axis** is *which property* of tenancy is being described — identity, + authorization, enforcement, placement, retention. + +A workload in the tenant plane has a position on all five axes. A platform +control plane service does too. The two vocabularies compose and neither +replaces the other. + +The rename is also an improvement. A posture vector is literally a point in +five-dimensional space, and "axis" says that where "plane" did not. + +## 1. Context + +Drafts 1–4 opened by claiming the estate "has never written down what it is +building". Relocation proved that wrong, and the correction is worth keeping +visible: `docs/platform-identity-security-architecture.md` has described the +trust model, the tenant model and a capability progression since 2026-07-23. +The accurate claim is narrower — **what was missing is a way to say how far a +given service has got, and to hold several answers at once.** Seven documents +cover slices of the subject and none of them does that: + +| Document | Covers | Status | +|---|---|---| +| `iam-profile_v0.3` (NetKingdom) | Tenant identifier shape, `tenant_roles` claim, staleness rules | Ratified | +| `tenant-engine-boundary-contract_v0.1` (NetKingdom) | Who owns tenant records, roles, plan assignment | Ratified | +| `business-app-service-contract_v0.1` §1 (Custodian) | Business apps: instance-per-client, tenant-keyed data | Ratified | +| `rapp-postgres` ADR-0001 | Consumer + tenant isolation in PostgreSQL | Proposed, governs one repo | +| `rapp-postgres` ADR-0002 | Per-consumer retention and the erasure horizon | Proposed, governs one repo | +| `shared-platform-relational-storage_v0.1` | The stacked-boundary gap | Routed 2026-08-10, **still unratified** | +| `platform-identity-security-architecture` (NetKingdom) | Trust model, planes, tenant model, capability progression | Accepted 2026-07-23 | + +This document is downstream of that architecture and must not restate it. It +answers one question the architecture leaves open: given the model, **where is +this particular service today, and how would anyone know?** + +Four failures follow. + +**The gap was diagnosed once and the fix stalled.** The v0.1 draft was written +to fill this hole and has sat unratified in neither canon directory. §20 +attaches a ratification path so this one does not join it. + +**Placement is owned by nobody.** `user-engine-pg` and `target-revenue-pg` are +dedicated; `apps-pg`, `net-kingdom-pg`, `platform-pg`, `state-hub-db` and +`forgejo-db` are shared. Both live, neither written down. `tenant-engine` +raised this with `railiance-platform` on 2026-08-16; unanswered. + +**Two contradictory defaults are already ratified.** Business apps get +instance-per-client; platform services pool. Nothing says which shape a new +service takes, and no definition separates the categories. + +**There is no honest way to describe a repo that is not there yet.** The estate +absorbs repos with weak or absent tenant separation. Today such a repo is +simply non-conformant, leaving it two bad options: misrepresent its posture, or +stay outside the framework. + +## 2. What this document is + +**A framework, not a model.** It specifies no single correct implementation. It +supplies terminology (§3, §4), a declaration (§5), a conformance rule (§6), +methodology (§12), and evidence definitions (§13). + +A service is conformant when its declared posture is accurate and its +trajectory recorded. A service is non-conformant when it claims a level it +cannot evidence — regardless of how high or low that level is. + +## 3. Five orthogonal axes + +"Is this multi-tenant?" is treated as one question. It is five, and they are +independent: + +| Axis | Question | Vocabulary owner | +|---|---|---| +| **Identity (I)** | How is a tenant named and validated? | `tenant-engine` / IAM Profile | +| **Authorization (A)** | How is a request bound to the tenants it may act for? | `flex-auth` | +| **Enforcement (E)** | Where, mechanically, is the tenant boundary enforced? | This framework | +| **Placement (P)** | Which substrate holds a tenant's data? | `railiance-platform` | +| **Retention (R)** | How long does data persist, and how is it erased? | The storage platform; policy by the consumer | + +Conflation produces errors today. `rapp-postgres`'s `PostgresConsumer` carries +`tenantIsolation: consumer-service-boundary` — an **E**-axis fact in a +**P**-axis artifact, reading as though storage enforces something it does not. +The "dedicated versus shared" argument mixes P (capacity, blast radius) with E +(correctness). + +The axes are separated *precisely so each may sit at a different level*. + +**Decision 3.1:** every document, declaration and plan tier that says +"isolation" MUST name which axis it means. + +**Decision 3.2:** the axes couple at their tops and the couplings MUST be +stated where they apply, not used to argue the axes are one: + +- `E4` is reachable only at `P3` or above. +- `R`'s erasure horizon is bounded below by `P` — on shared substrate, a + consumer's horizon is the instance maximum (§4.5). +- `R4` by key destruction is bounded by the **key boundary**, which is an + E-axis property. Shredding a single tenant's data requires the application + to encrypt under a per-tenant key before writing; the storage platform cannot + supply it. **Reaching the top of the retention ladder is not a retention + project.** + +**Decision 3.3 — scope.** The P and R ladders describe a service's **primary +datastore**. Caches, search indices, message queues and background jobs are +named leak surfaces in the external baselines and are assessed separately, not +covered by a posture vector. Saying so is honest; implying the vector covers +them would not be. + +## 4. Graduated levels + +Each axis carries an ordered ladder. Higher is stronger, not better: the right +level is the one a service can evidence and its risk warrants. + +### 4.1 Identity (I) + +| Level | State | +|---|---| +| **I0** | No tenant concept. Data not attributable to a tenant. | +| **I1** | A local tenant notion exists but is not canonical, **or** the tenant is taken from the request rather than from a verified token. | +| **I2** | Canonical identifiers, bound at the identity provider and carried as a verified claim; `tenant-engine` is the source of existence. | +| **I3** | I2 plus capability roles honoured, with live `tenant-engine` re-query for privileged, destructive, credential-vending or `aal2`-class decisions. | + +I1 now explicitly absorbs request-supplied tenant identifiers. "Never trust +client-supplied tenant IDs without validation" is a named anti-pattern; a +service reading the tenant from a header is at I1 however canonical the string. + +`business-app-service-contract` §2.1 sets app-local accounts as the v1 baseline +for business apps — a sanctioned low level with recorded triggers for moving +up. That is the pattern this framework generalises. + +### 4.2 Authorization (A) + +| Level | State | +|---|---| +| **A0** | No authorization, or tenant context not carried. | +| **A1** | Ad-hoc checks scattered through handlers. | +| **A2** | A single local authorization boundary; tenant context bound once, centrally. | +| **A3** | Decisions delegated to `flex-auth` as PDP, with live re-query where the IAM Profile requires it. | +| **A4** | A3 over a **standard** PDP interface (OpenID AuthZEN Authorization API 1.0), so the decision point is swappable and the enforcement point is not coupled to one engine's request shape. | + +A4 is new. `flex-auth` uses a bespoke `CheckRequest` and a bespoke action +vocabulary, with action strings copied verbatim between repos to avoid +re-derivation — exactly the coupling AuthZEN removes. The specification reached +Final in January 2026 and Keycloak shipped experimental support in May. We are +not wrong, we are pre-standard, and the ladder should have somewhere to go. + +**Internal service-to-service calls are in scope for this axis.** "Skipping +tenant validation for internal services" is a named anti-pattern, and our +estate is mostly internal calls — `flex-auth` calls `tenant-engine` +synchronously on the authorization path. A service identity acting on behalf of +a tenant must carry and revalidate tenant context to claim A2 or above. + +### 4.3 Enforcement (E) + +| Level | Mechanism | +|---|---| +| **E0** | None. Data not tenant-keyed; separation incidental or absent. | +| **E1** | Data tenant-keyed, filtering applied per query at call sites. | +| **E2** | Filtering centralised at a single service-side choke point binding authenticated identity to permitted tenants. | +| **E3** | E2 **plus** platform-assisted filtering: row-level security keyed on a tenant GUC set transaction-locally, or an equivalent enforced data-access layer. | +| **E4** | Structural: the credential a workload holds cannot address another tenant's data at all. Requires per-tenant credentials and per-tenant substrate. | + +**Correction from draft-2.** Draft-2 described E3 as something "the application +cannot trivially route around". That is false and it was this document +overclaiming in exactly the way §6 prohibits. Any session can re-issue `SET` on +a custom GUC, so an attacker with SQL execution can reset the tenant and read +across the boundary. What E3 buys is precise, and the ladder must say so: + +| Threat | E1 | E2 | E3 | E4 | +|---|:--:|:--:|:--:|:--:| +| A developer forgets a tenant predicate | ✗ | ✓ | ✓ | ✓ | +| A new code path bypasses the choke point | ✗ | ✗ | ✓ | ✓ | +| SQL injection reaching the connection | ✗ | ✗ | ✗ | ✓ | +| The application process is compromised | ✗ | ✗ | ✗ | ✓ | + +E3 is a strong control against **accident** — the common case, and the one that +causes real breaches — and no control at all against **compromise**. Only E4 +holds against both, because the credential itself cannot address another +tenant's data. + +**Correction: E3 layers on E2, it does not replace it.** External practice +treats application-layer and database-layer filtering as complementary. A +service that dropped its choke point on reaching E3 would be *worse* off, since +E3 fails open under injection. Claiming E3 therefore requires the E2 evidence +artifact as well. + +**Correction: the GUC is set transaction-locally.** Draft-2 said "at pool +checkout", which is session scope and the wrong instrument. Under a pooler in +statement mode, `SET` leaks between clients and returns other tenants' rows — +a failure that appears only under production concurrency and produces no error. +Use `SET LOCAL` inside an explicit transaction. + +**Platform enforcement is a platform obligation.** Reaching E3 requires the +storage platform to *offer* the mechanism: provisioned policies, a documented +GUC contract, and a probe. Where a consumer wants E3 and the platform has not +supplied it, the gap is the platform's. §19.6 asks `rapp-postgres` to define +that contract, which must carry `FORCE ROW LEVEL SECURITY` on every tenant +table (without it the table owner bypasses policies silently, and ADR-0001 +already established that our migration role owns the tables it creates), no +`BYPASSRLS` on leased roles, `SECURITY INVOKER` for ordinary logic, and an +`EXPLAIN` comparison because RLS disables functional indexes built on +non-leakproof functions. + +**Default expectation** for a new platform service: E2 at first serve, E3 +recorded as target. Services whose cross-tenant exposure would be a reportable +breach SHOULD target E3 or above. + +### 4.4 Placement (P) + +| Level | Shape | Live occupants | +|---|---|---| +| **P0** | Shares a database with another consumer. | None sanctioned; the state absorbed repos arrive in. | +| **P1** | Database per consumer, shared cluster. | `audit-core`, `tenant-engine` on `platform-pg` | +| **P2** | Dedicated cluster per consumer. | `user-engine-pg`, `target-revenue-pg` | +| **P3** | Dedicated cluster per tenant. | Business apps per `business-app-service-contract` §1.2 | +| **P4** | P3 plus separate region or jurisdiction. | None | + +Enforcement and placement are independent axes. Plotted together, with where +each service actually sits — parenthesised entries are targets or defaults +rather than current positions, and `—` marks a cell the coupling in §3.2 makes +unreachable: + +| E \ P | P0 | P1 | P2 | P3 | P4 | +|---|---|---|---|---|---| +| **E4** | — | — | — | (business app) | | +| **E3** | | (target) | | | | +| **E2** | | tenant-engine
audit-core | | | | +| **E1** | (absorbed repo) | | | | | +| **E0** | | | | | | + +**P0 → P1 → P2 is movement along the horizontal axis only.** Those steps buy +consumer isolation, capacity predictability, independent retention and a +smaller operational blast radius. They do not raise the tenant boundary by one +step. Only P3 makes E4 reachable. This is the most misusable fact in the +framework and §11 governs how it may be described. + +**Decision 4.4.1:** P1 is the default for platform services; P3 for +client-facing business apps, as already ratified. A service unsure which it is +must resolve that first (§19.4). + +**Decision 4.4.2 — placement scopes to data substrate.** Identity-provider +placement (realm-per-tenant versus Organizations) is the same silo/pool +decision on a different substrate, is live in our estate, and is undecided. +Realm-per-tenant carries a stated ceiling around 5–20 tenants, far below our +target. Recorded here as a parallel question (§19.7), not folded into P. + +### 4.5 Retention and erasure (R) + +New in draft-3. Implemented abstractly by the storage platform for any dataset; +policy is built on top of that interface by the consumer or its governance +layer. Reference implementation: `rapp-postgres` ADR-0002. + +| Level | State | +|---|---| +| **R0** | No retention or deletion position. Data kept indefinitely by default; no deletion path exists. | +| **R1** | Platform default retention applies (N=30 days). The consumer has declared no requirement. | +| **R2** | Retention declared as N days per dataset; the **erasure horizon** is published, and the consumer makes no promise shorter than it. | +| **R3** | Policy-driven deletion: the consumer or its governance layer declares what is due, the platform sweeps whole datasets on that instruction and evidences each run. | +| **R4** | Verified erasure: data proven unrecoverable across live storage, backups and derived copies, by one of the two routes below. | + +**R4 has two routes and a service MUST name which one it uses.** + +| Route | Mechanism | Cost | +|---|---|---| +| **Horizon-elapsed** | Wait out the published erasure horizon; the data ages out of every retained copy. | Available to everyone, proves little, and the wait is set by a co-resident's retention requirement rather than your own. | +| **Key-destroyed** | Encrypt per entity, then destroy the key. Retained copies survive but are unreadable. | Requires per-entity keys, strong encryption, and an auditable destruction record. Immediate. | + +**Regulatory standing of the key-destroyed route, stated carefully because +overclaiming here is worse than anywhere else in this framework.** Data +protection authorities have accepted key destruction as erasure where physical +deletion would be manifestly disproportionate, and the practice is recognised +under conditions — strong encryption, irreversible destruction, and an auditable +record of it. **The EDPB has not formally endorsed it as Article 17 erasure.** A +service reaching R4 by key destruction is making a defensible claim, not a +settled one, and must say so rather than reporting a clean "deleted". + +Three further properties. + +**The erasure horizon is the interval between deleting data and it ceasing to +be recoverable from anything the platform holds.** Deleting a row does not +remove it from yesterday's backup. With an N-day window, deleted data remains +recoverable for N days. That is the difference between "deleted" and "erased" +and the estate had never written it down. + +**On shared substrate, retention is not per-consumer.** Physical backup is +instance-wide — one WAL stream, one window — so the instance retention is +*derived* as the maximum across co-resident consumers, and every consumer's +horizon is that maximum. A consumer declaring 7 days beside one declaring 90 +gets 90. This is the retention analogue of ADR-0001's blast-radius disclosure: +state the coupling rather than imply an isolation that is not there. + +**Retention is therefore a placement trigger.** A consumer needing a horizon +shorter than the instance floor cannot have one at P1. It moves to P2 for a +reason with nothing to do with performance — which is exactly why it needs +recording, since nobody looks for a retention argument when reviewing +placement. + +Deletion splits mechanism from policy. The platform deletes whole **datasets** +on instruction and records an opaque policy reference it never interprets, so +every deletion traces to what authorised it. Rows are not a dataset: row expiry +is the consumer's own DML under its migration lease. Dropping a consumer's +whole database is an operator-gated offboarding step, never a scheduled one. + +## 5. The posture vector + +A service states one level per axis, plus a target, a date, and any placement +exceptions: + +```yaml +tenancy: + current: { I: 2, A: 3, E: 2, P: 1, R: 1 } + target: { I: 2, A: 3, E: 3, P: 1, R: 2 } + reviewed: "2026-08-17" + gap: + E: "Choke point exists and is tested; RLS not provisioned. Blocked on + rapp-postgres publishing the GUC contract. Target Q4." + R: "Retention declared; erasure horizon not yet published to consumers." +``` + +**Placement exceptions.** Draft-2 assigned one P level per service, which +cannot express the vertically partitioned model — most tenants pooled, some +dedicated — that §11's isolation tiers require. A tier requiring `P2` bought by +three tenants would put the service at two levels at once, forcing an over- or +under-claim. Placement is therefore declared as a default plus exceptions: + +```yaml + placement_exceptions: + - tenants: ["tenant:enterprise:*"] + P: 3 + reason: "isolation tier; see adaptive-pricing tier definition" +``` + +A service with exceptions must be able to say which tenants are on which +substrate. That mapping is a first-class artifact, not archaeology. + +Worked examples, best-effort and subject to owner correction: + +| Service | Current | Notes | +|---|---|---| +| `tenant-engine` | `I2 A3 E2 P1 R1` | Moving to P1 under TEN-WP-0009; retention declared, horizon not yet published. | +| `audit-core` | `I2 A3 E2 P1 R1` | Holds audit evidence, so both E3 and R2 are urgent targets. | +| A newly absorbed repo | `I1 A1 E1 P0 R0` | Conformant **if declared**, with a recorded path. | + +**Decision 5.1:** the posture vector is declared in the repo, not in the hub, +consistent with local-files-are-source-of-truth. + +## 6. Conformance is accuracy, not altitude + +> **A service is conformant when its declared posture is accurate, its target +> is recorded, and it does not claim a level it cannot evidence. It is +> non-conformant when it overclaims — at any altitude.** + +- Declaring `E0` is conformant. Concealing `E0` is not. +- A repo may be absorbed at any posture. It may not be absorbed silently. +- No service is blocked from the estate for being low on a ladder. Services MAY + be blocked from *specific work* — serving a tenant grouping, holding a data + class, carrying a plan tier — by requirements expressed as minimum levels. +- Downgrading is permitted and must be declared. A regression found by guarding + is a defect; a regression declared in advance is a decision. + +Without the axis separation, "not rigorous about tenant separation" is one +verdict a repo passes or fails. With it, the same repo is `I1 A1 E1 P0 R0` with +a path — a plan, not an indictment. + +## 7. Portability across placement levels + +Movement between P levels must be operational, not a rebuild: + +- Connect by injected credential only — no cluster, host, namespace or database + name in source. +- Own a whole database, never tables inside someone else's. +- Idempotent schema creation. +- No cross-database joins or co-location assumptions. + +**Decision 7.1:** mandatory at P1 and above. At P3, SHOULD rather than MUST — a +per-client instance that never moves is not misconformant for naming its own +database. + +## 8. Placement triggers + +Recorded at provisioning time: noisy neighbour on a latency-critical path; a +compliance or residency requirement; a plan tier requiring a higher minimum; an +erasure horizon that no longer fits (§4.5); connection or memory ceiling +reached. + +**Decision 8.1:** triggers MUST be *monitored*, not merely recorded. A trigger +in a YAML comment nobody re-reads is documentation, not control. + +**Decision 8.2:** placement policy ownership is proposed to +`railiance-platform`, **co-signed by `adaptive-pricing`**. Tenancy model +selection is a commercial decision as much as a technical one; an +operations-shaped repo should not hold it alone. + +## 9. Credentials as a tenancy control + +Short-lived leased credentials re-read at connection checkout, with +overlap-first rotation, bound the residual risk at every E level below E4: a +leaked credential expires rather than persisting. Stronger than the industry +norm of a long-lived per-service secret. + +**Decision 9.1:** static long-lived database credentials are not a sanctioned +path for any service above E0. + +## 10. Blast radius must be published + +**Decision 10.1:** every platform holding consumer data MUST publish, in +concrete terms, what a leaked runtime credential can and cannot reach at the +levels it operates. `rapp-postgres` ADR-0001 §5 is the reference. Where the +model cannot provide a guarantee, the platform says so and names the +escalation. + +**Decision 10.2 — quotas are disclosed, not discovered.** The same obligation +extends from what a leaked credential can reach to what the platform will +refuse to do for you. Every consumer MUST be told, at provisioning, the +throttles and quotas enforced against it — connection limits, statement +timeouts, idle-transaction timeouts — and told again when they change. A +consumer learning its statement timeout by hitting it in production is a +disclosure failure, not a consumer bug. This is how `tenant-engine` was +provisioned, by good practice rather than by rule; the rule now exists. + +## 11. Commercial expression + +- **11.1** Plan tiers are expressed *internally* as minimum levels. A tier may + require `E3 P2 R2`; it need not print that anywhere customer-facing. +- **11.2** Marketing and product language is free. No requirement to expose + level labels or this document. "Dedicated infrastructure", "isolated + tenancy", "private instance" all remain available. +- **11.3** The constraint is on **evidence, not vocabulary**. A customer-facing + isolation, availability or retention claim must map to a minimum level the + delivering service actually holds, recorded once when the tier is defined. + The review is internal and happens at tier definition — not per campaign. +- **11.4** Two hard lines, because these reach contracts and compliance + questionnaires: + - A claim that another tenant **cannot** reach the customer's data requires + **E4**. + - A claim that deleted data **is gone** requires **R4**, or an erasure + horizon disclosed alongside it. Where R4 is reached by key destruction, the + claim is defensible but not settled law (§4.5) — it may be made, and it may + not be made in language that implies a regulator has blessed it. + +## 12. Methodology — analyze, establish, improve, guard + +**Analyze.** Assess a repo against the ladders; produce `tenancy.current` with +reasoning recorded. Applies to new and absorbed services alike. + +**Establish.** Declare the target and gap. The target is set by data class, +tenant groupings served and plan tiers carried — not by ambition. + +**Improve.** Move one axis at a time. Raising P while leaving E untouched is +the characteristic misstep. + +**Guard.** Verify continuously that the declared posture holds — **against the +service's own declaration**, not a universal maximum. Nobody must prove every +service is at E4; the check is that none is below what it declared. + +Regression found by guarding is a defect; regression declared in advance is a +decision. The estate has been bitten twice by silent pin rollbacks producing +ordinary-looking 403s and 404s rather than errors. Posture regression looks the +same — an RLS context leak returns correct-looking rows for the wrong tenant. +Guarding must be designed for invisible failure, not for crashes. + +## 13. Evidence per level + +**Decision 13.1:** a level is claimed only with its evidence artifact present. +This turns §6's accuracy rule from an honour system into a check. + +**Decision 13.4 — an artifact must assert something achievable.** Draft-3's +noisy-neighbour evidence required proof that a saturating consumer "does not +breach" another's allowance. Shared infrastructure cannot provide that; the +risk is inherent and cannot be wholly removed. An artifact that can only fail, +or that passes by being run gently enough, is an overclaim wearing the costume +of evidence. Where a property cannot be guaranteed, the artifact measures and +records it instead. + +**Decision 13.2 — evidence is of two kinds, and conflating them is an +overclaim.** *Mechanical* evidence is a structural assertion a machine can make +and belongs in CI. *Adversarial* evidence is semantic, requires setting up +separate tenant contexts and comparing responses, and carries a review date +rather than a green build. Cross-tenant findings are the category external +testing practice identifies as needing human review. **A passing CI run is not +E2 evidence.** + +| Level | Evidence | Kind | +|---|---|---| +| **I2** | Identifiers validated against the vocabulary; rejection test for a malformed id; binding shown to come from a verified token | Mechanical | +| **I3** | Live re-query demonstrated on an `aal2`-class path; cached-claim path shown unused there | Mechanical | +| **A2** | Choke point identified; test that an unbound request is refused | Mechanical | +| **A3** | Live decision with a denial observed at the endpoint, not only at the decision surface | Mechanical | +| **A4** | Decision served over the standard interface; a second PDP substituted without PEP change | Mechanical | +| **E1** | Every tenant-owned table carries the tenant key | Mechanical | +| **E2** | Choke point identified; identity bound to tenant A demonstrably cannot read tenant B | **Adversarial**, with a review date | +| **E3** | `FORCE ROW LEVEL SECURITY` on every tenant table; no `BYPASSRLS` on leased roles; probe that a session without the GUC reads nothing; probe that a wrong GUC reads nothing; `EXPLAIN` comparison | Mechanical | +| **E4** | Per-tenant credential demonstrated unable to connect to another tenant's substrate | Mechanical | +| **P1–P4** | Provisioning declaration plus the platform's isolation probes | Mechanical | +| **P1–P2 (noisy neighbour)** | A recorded baseline of per-consumer resource usage; a run in which one consumer saturates its declared allowance; evidence that the governance controls **bind** (the greedy consumer is held at its limits) and that the degradation co-residents experience is **measured, recorded and judged acceptable**; the aggregate headroom at time of measurement | **Adversarial**, load-generated, with a review date | +| **R2** | Declared retention rendered; erasure horizon published and reported in the operator surface | Mechanical | +| **R3** | Sweep evidence records: timestamp, dataset, identifiers removed, authorising policy reference | Mechanical | +| **R4** | Erasure demonstrated across live data, backups and derived copies within the horizon | **Adversarial** | + +**Decision 13.3:** the E2, E3 and noisy-neighbour artifacts do not exist +anywhere in the estate today. `rapp-postgres` runs 15 adversarial probes, all +against the *consumer* boundary, none against the tenant boundary inside a +consumer. Externally, what this framework calls a tenant boundary failure is +**Broken Object Level Authorization** — OWASP API1, top of the API Security Top +10 since that list launched, and the most commonly exploited API vulnerability +in published assessments. We have no coverage for the highest-ranked risk in +our class of system. §19.3 seeks an owner. + +## 14. Adoption stance — structure, not tooling + +**Decision 14.1:** external research is design input. This estate adopts +published standards and structural patterns; it does not adopt tooling unless +that tooling is an established industry standard with broad application. +Everything else is built ground-up, so it can be optimised and refactored as +the estate sees fit. + +| Class | Stance | +|---|---| +| Security baselines (OWASP Multi-Tenant Security Cheat Sheet, API Security Top 10) | Adopt as the external reference our ladders answer to | +| Standards bodies (OpenID AuthZEN 1.0) | Adopt — this is what A4 is | +| Reference taxonomies (Azure tenancy models, AWS SaaS Lens, cell architecture) | Adopt as structure | +| Engine behaviour (PostgreSQL RLS mechanics) | Facts, not tooling | +| Third-party analyzers and test frameworks | **Do not adopt.** Take their rule taxonomies as checklists for probes we write ourselves | + +The practical effect is small and good: `rapp-postgres` already owns a +ground-up probe harness — bash and psql, no dependency tree — that found four +real defects in its own provisioning SQL. The evidence artifacts in §13 become +new probes in a tool we control. One idea worth reimplementing from the +external survey is **policy-diff classification**: labelling a change to an +enforcement policy as safe or breaking *before* it lands. + +## 15. Alternatives considered + +**One fixed model with a single set of characteristics** (draft-1). *Rejected:* +cannot describe a repo that is not there yet, forcing absorbed repos to +misrepresent their posture or stay outside. A framework that can only describe +its own end state is not a framework. + +**A maturity model with a single overall level.** *Rejected:* collapses the +axis separation. A service strong on identity and weak on enforcement has a +specific, actionable gap; one composite score hides it and invites averaging. + +**Prohibiting row-level security** (draft-2's inherited position). *Rejected in +draft-2, refined in draft-3:* RLS is a real rung against the common threat. The +error was never RLS — it was describing E3 in E4's language. + +**Schema-per-consumer in one database.** *Rejected:* `pg_catalog` is readable +per-database, so every co-resident enumerates every other's table and column +names regardless of grants. Retained as a describable state, never a target. + +**Mandating E4 for everyone.** *Rejected:* the tenant taxonomy includes +`consumer` (private individuals) and `family`. A cluster per private individual +is economically impossible; the taxonomy is itself evidence pooling is +required. + +**Per-consumer physical backup retention.** *Rejected:* CNPG retention is a +property of the instance's WAL archive. There is no mechanism, and claiming it +would be a fabricated guarantee. Hence the derived maximum in §4.5. + +**Platform-scheduled row expiry.** *Rejected:* requires the platform to hold +DML authority over consumer schemas and interpret consumer data semantics, both +forbidden by ADR-0001. The consumer's migration lease is the correct +instrument. + +**Leaving each repo to its own model.** *Rejected:* the status quo, which +produced two contradictory ratified defaults and an unowned placement question. + +## 16. Held against outside practice + +**The graduated reframe is corroborated, not invented here.** Microsoft's +tenancy-model guidance states it almost verbatim: *"Instead of viewing +isolation as a discrete property, consider it a spectrum. You can deploy +components of your architecture that are more isolated or less isolated than +other components in the same architecture."* The same guidance derives our E↔P +coupling independently — shared deployment means enforcement lives in +application code; dedicated deployment means it is structural. + +**Stronger than typical.** Most multi-tenancy literature models one boundary, +tenant-to-tenant. This estate has **two stacked boundaries**: platform-service +to platform-service, and tenant to tenant inside a consumer. Naming them +separately and refusing to enforce both with one mechanism is uncommon and +correct. Graduated per-axis levels also beat the silo/pool/bridge trichotomy, +which is approximately our P axis with the other four missing — which is why +it cannot express "pooled infrastructure, structurally enforced boundary". + +**Weaker than typical.** The pool model's standard mitigation is a *verified* +enforcement layer every service is demonstrably routed through. We have the +concept and none of the verification (§13.3). + +**Adopted without naming it.** Short-lived leased credentials re-read at +checkout beat the long-lived-secret norm. §9 promotes it to a tenancy control. + +**Still unexplored.** Neither P nor R describes a **cell** — a slice of +infrastructure with a *fixed maximum size*, sized so one cell's failure is +survivable and cell count scales linearly. `platform-pg` is, in these terms, an +uncapped cell: §17 computes a ceiling and nothing enforces it (§19.8). + +Sources: the four research digests in `research/2026-08-17-adr008-*`, which +carry full citations for every claim in this section. + +## 17. Scaling demands + +Measured against the live `platform-pg` specification, not estimated. + +``` +instances: 1 (no HA; single-node rail) +max_connections: 100 +memory limit: 1Gi +per consumer: 14 connections (12 runtime + 2 migration) +``` + +**Connection ceiling: roughly six consumers — and this is the aggregate +noisy-neighbour bound, not a capacity statistic.** Seven consumers request 98 +of 100 before CNPG's instance manager, metrics exporter and reserved slots. +Every one of them is politely inside its declared 14-connection allowance; the +instance still fails. + +That distinction matters because our governance addresses the wrong shape. +Per-consumer `connection_limit`, `statement_timeout` and +`idle_in_transaction_session_timeout` guard well against **one greedy +consumer**. They do nothing about **the aggregate of many modest ones**, which +is the second and less intuitive noisy-neighbour failure and the one this +number describes. Two consumers are provisioned. We are at roughly a third of +the bound, and the third request will not feel like a scaling event. + +**Memory likely binds first.** 100 backends against 1Gi is ~10MB per backend. +Connection exhaustion errors clearly; memory pressure OOM-kills and degrades +every co-resident at once. + +**E3 and pooling.** *Corrected from draft-2, which had this backwards.* +Transaction-scoped context (`SET LOCAL` inside an explicit transaction) is what +makes E3 **safe** under a pooler. Statement-level pooling is what breaks it, +serving other tenants' rows under concurrency with no error. E3 constrains +which pooling mode is available, not whether pooling is available. + +**Retention consumes the volume.** WAL accumulates with the window, and §4.5 +makes the window the maximum across consumers. A consumer declaring a long +retention extends everyone's horizon *and* everyone's storage draw against a +20Gi volume. + +**Restore time couples all consumers.** Physical backup is instance-wide, so a +consumer's RTO is a function of *total* instance size, not its own. + +**No P1 tenant has HA.** `instances: 1` means a tier promising uptime cannot be +satisfied at P1 as built — an availability floor belongs in §11's +minimum-level vocabulary alongside isolation. + +## 18. Consequences + +- The estate gains one vocabulary and a way to be honest about partial + adoption. +- Absorbed repos get a described state and a path instead of a failing grade. +- `tenantIsolation` in `PostgresConsumer` is revealed as a mislabelled field. +- The verification problem becomes tractable: guard against declaration. +- Draft-2's RLS prohibition is reversed and its E3 description corrected; + `rapp-postgres` acquires an obligation to define and offer the mechanism. +- Adding a consumer with long retention **silently extends everyone's erasure + horizon**. This must reach the consumer review checklist, not only this + document. +- A service selling an isolation tier must maintain a tenant→substrate mapping + it does not have today. +- Nothing here changes a running system. + +## 19. Open questions + +1. **`tenantIsolation` field** — `rapp-postgres`: rename to name its axis and + carry a level (`tenancy.E: 2`), or move it out of the storage declaration. +2. **Placement ownership** — `railiance-platform` with `adaptive-pricing`: + accept the ladder, triggers and the §8.1 monitoring obligation; appoint a + recorded placement owner per workload. +3. **E2, E3 and noisy-neighbour evidence** — **owned as of 2026-08-17** by + `whitehat-security` (WHITEHAT-WP-0001), an independent adversarial evidence + facility seeded for this purpose. `audit-core` and `tenant-engine` were + right to decline it as fleet-scope work; the answer was a home of its own + rather than a volunteer. + + Independence is the design point, and it bears on this document: a facility + verifying conformance to this framework is deliberately **not** owned by + NetKingdom, which owns the framework. Self-grading one level up is still + self-grading. + + Two consequences land back here. **Cadence is now a security parameter, not + a schedule** — for any control whose guarantee is detection rather than + prevention, the interval between probe runs *is* the exposure window, and + `rapp-postgres` ADR-0003 leaves that number to the facility. And **a passing + suite is not proof of isolation**; it is proof that the attacks attempted + did not work. §13's evidence artifacts should be read with that distinction, + because a green run recorded as "E2 verified" would be exactly the overclaim + §6 prohibits. +4. **Business app vs platform service** — Custodian canon: a classification + rule. Candidate: reuse `repo-classification-standard_v1.0`. +5. **Tier → minimum level mapping** — `adaptive-pricing` and `tenant-engine`: + required only for tiers making isolation, availability or retention claims. +6. **The E3 mechanism** — `rapp-postgres`: publish the GUC contract with the + `FORCE`/`BYPASSRLS`/`SECURITY INVOKER`/`EXPLAIN` requirements in §4.3. +7. **Identity-provider placement** — owner of `key-cape`: realm-per-tenant or + Organizations? Realm-per-tenant's ~5–20 tenant ceiling is below our target. +8. **Cell sizing** — reframed from "should we adopt cells" to **"what is + `platform-pg`'s declared maximum size, and what is the overflow target?"** + The connection ceiling forces this whether or not we adopt the vocabulary. +9. **Retention floor and ceiling** — should `backupRetentionDays` have a + platform minimum (so a consumer asking for 1 day gets a validation error + rather than a quiet disappointment) and a maximum (so nobody exhausts the + volume)? +10. **Engine neutrality** — the P ladder rests on a PostgreSQL property. + State it engine-specifically and say so, or abstract it and risk a + non-Postgres implementation that silently differs? +11. **Erasure versus audit** — `audit-core`: crypto-shredding a tenant's audit + records destroys the evidence the service exists to hold, and ADR-0001 §2 + deliberately built the role model so history could not be rewritten. The + usual resolution separates the *fact* of an event, retained, from its + *personal payload*, encrypted per subject and shreddable. Raised because a + naive "R4 everywhere" target would instruct the audit service to destroy + its own evidence. The answer is `audit-core`'s, not this framework's. +12. **Quality of service** — *owner needed.* The framework has no vocabulary + for saying one consumer's latency matters more than another's. + `tenant-engine` sits on `flex-auth`'s synchronous authorization path and + chose a 5s statement timeout for that reason; it shares an instance with + `audit-core`, which is not latency-critical. Nothing prioritises between + them. Either add a QoS dimension or state that all co-residents are equal + and latency-critical consumers must escalate to P2. + +**Routed elsewhere, deliberately.** The tenant identifier +`tenant::` embeds headcount bands (`small`, `medium`, `large`) +that change as a tenant grows, contradicting the consensus that identifiers +should not encode mutable attributes. That is a critique of ADR-0013, not of +this framework, and belongs to `tenant-engine` and NetKingdom canon. Folding it +in here would overreach. + +## 20. Ratification path + +1. Reviewed by `tenant-engine`, `flex-auth`, `rapp-postgres`, + `railiance-platform` and `adaptive-pricing` against §19. +2. Each publishes its own posture vector (§5) as part of review. **The + framework is validated by whether it can describe them accurately** — if a + repo cannot express itself in these five ladders, the ladders are wrong and + this document changes, not the repo. +3. On acceptance, **supersedes** the routing of + `rapp-postgres/docs/canon-drafts/shared-platform-relational-storage_v0.1-draft.md`, + whose §§3–8 are absorbed here. That draft is withdrawn rather than left + pending. +4. On acceptance, `rapp-postgres` ADR-0001 and ADR-0002 move to `accepted` and + are annotated as the PostgreSQL implementation of the E, P and R ladders.