190 lines
8.5 KiB
Markdown
190 lines
8.5 KiB
Markdown
|
|
---
|
||
|
|
id: WHITEHAT-WP-0001
|
||
|
|
type: workplan
|
||
|
|
title: "Produce the adversarial evidence the Tenancy Posture ladders require"
|
||
|
|
domain: infotech
|
||
|
|
repo: whitehat-security
|
||
|
|
status: proposed
|
||
|
|
owner: the-custodian
|
||
|
|
topic_slug: whitehat-security
|
||
|
|
created: "2026-08-17"
|
||
|
|
updated: "2026-08-17"
|
||
|
|
---
|
||
|
|
|
||
|
|
# WHITEHAT-WP-0001 — cross-tenant evidence
|
||
|
|
|
||
|
|
## Goal
|
||
|
|
|
||
|
|
Produce, on a schedule, the adversarial evidence artifacts that NetKingdom's
|
||
|
|
*Tenancy Posture* standard requires and no repo can honestly produce about
|
||
|
|
itself — starting with the one the framework calls its highest-severity gap.
|
||
|
|
|
||
|
|
Done means: a service claiming `E2` has been attacked as an adversary holding
|
||
|
|
its runtime credential, the attempt is recorded with a date and an attacker
|
||
|
|
model, and a failure routes to `risk-nexus` rather than to a log nobody reads.
|
||
|
|
|
||
|
|
## The forcing case
|
||
|
|
|
||
|
|
*Tenancy Posture* open question 3 has been unowned since the framework was
|
||
|
|
drafted. What it calls a tenant-boundary failure the security industry calls
|
||
|
|
**Broken Object Level Authorization** — OWASP API1, top of the API Security Top
|
||
|
|
10 since that list launched, and the most commonly exploited API vulnerability
|
||
|
|
in published assessments.
|
||
|
|
|
||
|
|
The estate has no coverage for it. `rapp-postgres` runs fifteen adversarial
|
||
|
|
probes and every one targets the *consumer* boundary — service versus service.
|
||
|
|
None targets the tenant boundary *inside* a consumer, which is where the
|
||
|
|
framework says the residual risk actually lives.
|
||
|
|
|
||
|
|
## Rules of engagement — T01, and nothing else starts first
|
||
|
|
|
||
|
|
An automated facility that probes systems without written scope is
|
||
|
|
indistinguishable from the threat it models. This is the gating task and it is
|
||
|
|
not paperwork.
|
||
|
|
|
||
|
|
- **Scope.** Which systems, which namespaces, which credentials. Explicitly:
|
||
|
|
build-mode environments now; production requires a separate, recorded
|
||
|
|
authorization.
|
||
|
|
- **Prohibited actions**, stated as hard rules rather than intentions:
|
||
|
|
no destructive operations against data the estate did not create for the
|
||
|
|
test; no exfiltration of real tenant data even as proof of a finding —
|
||
|
|
a count and a schema shape are proof enough; no denial-of-service against a
|
||
|
|
shared substrate outside a declared window, because the connection ceiling
|
||
|
|
means a saturation probe is an outage for every co-resident.
|
||
|
|
- **Credentials.** The facility holds leased credentials like any workload,
|
||
|
|
through the sanctioned OpenBao lane. It gets no standing privilege, and
|
||
|
|
notably **no `BYPASSRLS` and no superuser** — an attacker would not have them
|
||
|
|
and a probe holding them proves nothing.
|
||
|
|
- **Attribution.** Every probe connection is identifiable as a probe in
|
||
|
|
`pg_stat_activity` and in logs, so an operator investigating an anomaly can
|
||
|
|
tell us from a real adversary in seconds.
|
||
|
|
- **Abort.** How a run is stopped, by whom, and what state it leaves behind.
|
||
|
|
|
||
|
|
**Output:** `docs/rules-of-engagement.md`, reviewed by the operator personally.
|
||
|
|
This is exactly the class of thing `risk-nexus`'s escalation duty exists for.
|
||
|
|
|
||
|
|
## Tasks
|
||
|
|
|
||
|
|
### T01 — Rules of engagement
|
||
|
|
As above. Gates everything.
|
||
|
|
|
||
|
|
### T02 — The attacker model per axis
|
||
|
|
|
||
|
|
What the adversary is assumed to hold, so a probe is judged against a threat
|
||
|
|
rather than against taste. Drawn from *Tenancy Posture* §4.3, which already
|
||
|
|
distinguishes them:
|
||
|
|
|
||
|
|
| Axis | Adversary holds | Probe answers |
|
||
|
|
|---|---|---|
|
||
|
|
| E1/E2 | A legitimate runtime credential and the ability to make ordinary requests as tenant A | Can it reach tenant B's rows? |
|
||
|
|
| E3 | The above, plus SQL execution on the connection | Can it re-`SET` the tenant GUC and read across? |
|
||
|
|
| E4 | A leaked per-tenant credential | Can it address another tenant's substrate at all? |
|
||
|
|
| P1/P2 | A co-resident consumer behaving badly within its own allowance | What degradation do neighbours experience? |
|
||
|
|
| R | A copy of a backup taken before an erasure | Is the erased data still readable? |
|
||
|
|
|
||
|
|
**Output:** `docs/attacker-model.md`. Note that the E3 row exists because the
|
||
|
|
framework corrected itself: E3 stops accident, not compromise, and a probe that
|
||
|
|
only tested accident would report a strength E3 does not have.
|
||
|
|
|
||
|
|
### T03 — Differential cross-tenant harness (the E2 artifact)
|
||
|
|
|
||
|
|
The core technique: run the same request as two tenants and compare.
|
||
|
|
|
||
|
|
- Provision two disposable tenants against a target service.
|
||
|
|
- Exercise its surface as tenant A; attempt every object identifier observed
|
||
|
|
from tenant B's context.
|
||
|
|
- Assert: B receives a denial or an empty result. **A well-formed 200
|
||
|
|
containing A's data is the finding**, and the harness must be built to notice
|
||
|
|
that rather than to notice errors — both of this estate's real regressions
|
||
|
|
produced ordinary-looking responses, a 403 and a 404, and nothing alerted.
|
||
|
|
- Capture evidence as a count and a schema shape, never as tenant data (T01).
|
||
|
|
|
||
|
|
**Acceptance:** run against `tenant-engine` and `audit-core`, both of which
|
||
|
|
currently claim `E2`. The artifact is the run record, not a green tick.
|
||
|
|
|
||
|
|
### T04 — Prove the probes fail
|
||
|
|
|
||
|
|
A probe that has only ever passed is not evidence.
|
||
|
|
|
||
|
|
Build known-bad fixtures — a service with a deliberately missing tenant
|
||
|
|
predicate — and confirm each probe fails against them. `rapp-postgres` verified
|
||
|
|
its drift check this way, by re-pinning to a bad digest and confirming exit 1;
|
||
|
|
the same discipline applies here and is not optional.
|
||
|
|
|
||
|
|
**Acceptance:** every probe in T03 demonstrated failing before any of them is
|
||
|
|
trusted passing.
|
||
|
|
|
||
|
|
### T05 — RLS conformance under attack (the E3 artifact)
|
||
|
|
|
||
|
|
`rapp-postgres` ADR-0003 supplies an `rls_conformance` view and a template, and
|
||
|
|
states plainly that the platform's guarantee is **detection, not prevention** —
|
||
|
|
a table created by a later migration ships without a policy until something
|
||
|
|
notices. This repo is that something.
|
||
|
|
|
||
|
|
- Query the conformance view on a cadence; any row is a finding.
|
||
|
|
- Attack what the view cannot see: a session that sets no GUC must read
|
||
|
|
nothing; a session with another tenant's value must see nothing; an insert
|
||
|
|
attributed to another tenant must be refused.
|
||
|
|
- Attempt the documented bypasses: a `SECURITY DEFINER` function owned by the
|
||
|
|
table owner, and a role holding `BYPASSRLS`.
|
||
|
|
|
||
|
|
**Cadence is the deliverable here, not a detail.** For a detection-based
|
||
|
|
control the interval between runs *is* the exposure window, and ADR-0003 leaves
|
||
|
|
the number to this repo. Set it, and state the resulting window in the record.
|
||
|
|
|
||
|
|
### T06 — Noisy-neighbour characterisation (the P1/P2 artifact)
|
||
|
|
|
||
|
|
The framework had to reword this artifact once already: its first draft
|
||
|
|
required proof that a saturating consumer "does not breach" another's
|
||
|
|
allowance, which shared infrastructure cannot provide.
|
||
|
|
|
||
|
|
What is achievable and therefore what this produces: a recorded baseline of
|
||
|
|
per-consumer resource usage; a run in which one consumer saturates its declared
|
||
|
|
allowance; evidence that the governance controls **bind**; and the degradation
|
||
|
|
co-residents experience, **measured and written down** rather than asserted
|
||
|
|
acceptable.
|
||
|
|
|
||
|
|
Runs inside a declared window per T01 — on a single-node rail with a six-
|
||
|
|
consumer connection ceiling, a saturation probe is an outage if run carelessly.
|
||
|
|
|
||
|
|
### T07 — Reporting into risk-nexus
|
||
|
|
|
||
|
|
Findings leave this repo in one direction. A run produces: what was attempted,
|
||
|
|
under which attacker model, when, against which posture claim, and the outcome.
|
||
|
|
It carries no severity — that is `risk-nexus`'s.
|
||
|
|
|
||
|
|
A **passing** run is also reported. "The attacks we thought of did not work" is
|
||
|
|
the honest claim, and recording it dated is what lets anyone see how stale the
|
||
|
|
assurance has become.
|
||
|
|
|
||
|
|
## Sequencing
|
||
|
|
|
||
|
|
T01 gates all. T02 shapes T03/T05/T06. T04 gates trusting any of them. T07 can
|
||
|
|
follow T03.
|
||
|
|
|
||
|
|
## Risks
|
||
|
|
|
||
|
|
**The facility becomes the threat.** Mitigated by T01, and by holding no
|
||
|
|
standing privilege.
|
||
|
|
|
||
|
|
**Probes weaken silently.** A probe that starts passing after a change to
|
||
|
|
itself rather than to the system is a finding, not a fix. Probe changes are
|
||
|
|
reviewed as security changes.
|
||
|
|
|
||
|
|
**Green is mistaken for safe.** Every report states that a pass means the
|
||
|
|
attacks attempted did not work, not that the boundary holds.
|
||
|
|
|
||
|
|
**It drifts into fixing things.** The boundary in INTENT is load-bearing:
|
||
|
|
findings route out, work does not come in.
|
||
|
|
|
||
|
|
## Open questions
|
||
|
|
|
||
|
|
1. **Owner confirmation.** Proposed `the-custodian`, deliberately not
|
||
|
|
NetKingdom, whose framework this verifies. Needs the operator's yes.
|
||
|
|
2. **Where probes run from.** In-cluster gives realistic network position;
|
||
|
|
outside gives independence from the substrate under test. Probably both,
|
||
|
|
eventually; pick one to start.
|
||
|
|
3. **Does a service get told it is being probed?** Announced runs are easier to
|
||
|
|
operate; unannounced ones test the alerting too. Build mode probably
|
||
|
|
announced, production probably not — which is itself a T01 decision.
|