whitehat-security/workplans/WHITEHAT-WP-0001-cross-tenant-evidence.md
tegwick 75df0207c3 Declare Staff in INTENT.md and record WP-0001 DoD-Ok
Replace the transcribed gate-house review note with this repository's
own v0.7 §11 declaration. Close WHITEHAT-IN-0001: Staff, blocked-clean,
probe authorship retained here. Record DoD-Ok on WHITEHAT-WP-0001 so
the finished plan is not quality-debt-open.

Assistant: grok
Assistant-Session: 01a05e32-c776-72a3-86ec-c490e027aca9
2026-09-01 20:49:53 +02:00

373 lines
17 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
id: WHITEHAT-WP-0001
type: workplan
title: "Produce the adversarial evidence the Tenancy Posture ladders require"
domain: infotech
repo: whitehat-security
status: finished
owner: net-kingdom
topic_slug: whitehat-security
created: "2026-08-17"
updated: "2026-09-01"
state_hub_workstream_id: "3049aa1e-b188-514f-9ad7-bf3026094fb9"
quality_dod: DoD-Ok
quality_dod_at: "2026-09-01"
quality_dod_by: grok
quality_dod_note: "T01T08 done for every applicable target; live residuals owned by WHITEHAT-WP-0006; make check passed; no live traffic."
---
# WHITEHAT-WP-0001 — cross-tenant evidence
## Goal
Produce, on a schedule, the adversarial evidence artifacts that NetKingdom's
*Tenancy Posture* standard requires and no repo can honestly produce about
itself — starting with the one the framework calls its highest-severity gap.
Done means: a service claiming `E2` has been attacked as an adversary holding
its runtime credential, the attempt is recorded with a date and an attacker
model, and a failure routes to `risk-nexus` rather than to a log nobody reads.
## The forcing case
*Tenancy Posture* open question 3 has been unowned since the framework was
drafted. What it calls a tenant-boundary failure the security industry calls
**Broken Object Level Authorization** — OWASP API1, top of the API Security Top
10 since that list launched, and the most commonly exploited API vulnerability
in published assessments.
The estate has no coverage for it. `rapp-postgres` runs fifteen adversarial
probes and every one targets the *consumer* boundary — service versus service.
None targets the tenant boundary *inside* a consumer, which is where the
framework says the residual risk actually lives.
## Rules of engagement — T01, and nothing else starts first
An automated facility that probes systems without written scope is
indistinguishable from the threat it models. This is the gating task and it is
not paperwork.
- **Target authorization**, the control that matters most now the facility is
scoped to any surface we choose rather than only our own. No target without a
recorded authorization from its responsible party. Our estate in build mode
has standing authorization only as a prerequisite: every live run still
needs the dated engagement record and target-owner acknowledgement defined in
the rules of engagement. Production needs its own authorization; anything we
do not own needs written per-engagement authorization recorded here before a
packet is sent. A commercial relationship, public reachability, and "they
would obviously be fine with it" are each explicitly not authorization.
- **Scope.** Which systems, which namespaces, which credentials — and a hard
stop at the engagement boundary. A probe that discovers an adjacent system
reports what it saw and does not follow it.
- **Prohibited actions**, stated as hard rules rather than intentions:
no destructive operations against data the estate did not create for the
test; no exfiltration of real tenant data even as proof of a finding —
a count and a schema shape are proof enough; no denial-of-service against a
shared substrate outside a declared window, because the connection ceiling
means a saturation probe is an outage for every co-resident.
- **Credentials.** The facility holds leased credentials like any workload,
through the sanctioned OpenBao lane. It gets no standing privilege, and
notably **no `BYPASSRLS` and no superuser** — an attacker would not have them
and a probe holding them proves nothing.
- **Attribution.** Every probe connection is identifiable as a probe in
`pg_stat_activity` and in logs, so an operator investigating an anomaly can
tell us from a real adversary in seconds.
- **Abort.** How a run is stopped, by whom, and what state it leaves behind.
**Output:** `docs/rules-of-engagement.md`, reviewed by the operator personally.
This is exactly the class of thing `risk-nexus`'s escalation duty exists for.
## Tasks
### T01 — Rules of engagement
As above. Gates everything.
```task
id: WHITEHAT-WP-0001-T01
status: done
priority: high
state_hub_task_id: "13c2feb1-e9ce-5ec5-8887-729541351429"
```
Drafted in `docs/rules-of-engagement.md` on 2026-08-18 with authorization
classes, per-run records, initial target envelope, hard prohibitions,
credential/attribution rules, rate defaults, abort/cleanup and evidence
schema. The operator accepted v0.1 on 2026-08-21 with the scope recorded in
what is now §11. v0.2 (2026-08-22) adds §10, the governed test plane, as a
stricter admission control. It does not expand authorization. The acceptance
approves the operating rules and offline fixture work; it does not
pre-authorize any live target.
### T02 — The attacker model per axis
```task
id: WHITEHAT-WP-0001-T02
status: done
priority: high
state_hub_task_id: "b467057d-3c1e-5580-b5c8-1b821a1322ef"
```
What the adversary is assumed to hold, so a probe is judged against a threat
rather than against taste. Drawn from *Tenancy Posture* §4.3, which already
distinguishes them:
| Axis | Adversary holds | Probe answers |
|---|---|---|
| E1/E2 | A legitimate runtime credential and the ability to make ordinary requests as tenant A | Can it reach tenant B's rows? |
| E3 | The above, plus SQL execution on the connection | Can it re-`SET` the tenant GUC and read across? |
| E4 | A leaked per-tenant credential | Can it address another tenant's substrate at all? |
| P1/P2 | A co-resident consumer behaving badly within its own allowance | What degradation do neighbours experience? |
| R | A copy of a backup taken before an erasure | Is the erased data still readable? |
**Output:** `docs/attacker-model.md`, completed 2026-08-21. It separates
credential-bearing tenant attacks, omitted-predicate accidents, SQL-capable
compromise, structural credential confinement, bounded co-resident saturation,
and both R4 erasure routes. It preserves the framework's correction that E3
stops accident, not compromise: resetting the tenant GUC after SQL execution
is recorded as E3's documented limit, not misreported as an E3 conformance
failure.
### T03 — Differential cross-tenant harness (the E2 artifact)
```task
id: WHITEHAT-WP-0001-T03
status: done
priority: high
state_hub_task_id: "e03edc76-a314-50dc-8625-3666c6351818"
```
The core technique: run the same request as two tenants and compare.
- Provision two disposable tenants against a target service.
- Exercise its surface as tenant A; attempt every object identifier observed
from tenant B's context.
- Assert: B receives a denial or an empty result. **A well-formed 200
containing A's data is the finding**, and the harness must be built to notice
that rather than to notice errors — both of this estate's real regressions
produced ordinary-looking responses, a 403 and a 404, and nothing alerted.
- Capture evidence as a count and a schema shape, never as tenant data (T01).
**Acceptance:** a dated run record against every *applicable* E2 target. The
artifact is the run record, not a green tick. `tenant-engine` is registered
`not_applicable` for E2; that record is the artifact for that target.
`flex-auth` remains `pending` and is not an applicable E2 target.
Done 2026-08-22: `WH-ENG-20260822-AUDIT-E2-03` is a dated target pass against
the only applicable live E2 target, `audit-core`. Ten operations, three
calibrated probes, cleanup before expiry, sanitized report in
`evidence/WH-ENG-20260822-AUDIT-E2-03.json`. `tenant-engine` stays
`not_applicable`; that record is the artifact, not a deferral. `-01` expired
unused and `-02` aborted with zero packets; those identifiers remain terminal.
Whitehat will not relabel pending or not-applicable targets to finish this
task. A later audit-core run needs a new engagement ID; this pass is due for
review or replacement at 2026-08-23T22:10:25Z.
### T04 — Prove the probes fail
```task
id: WHITEHAT-WP-0001-T04
status: done
priority: high
state_hub_task_id: "3956bdad-14cd-5b60-91b5-9a27ac4ab745"
```
A probe that has only ever passed is not evidence.
Build known-bad fixtures — a service with a deliberately missing tenant
predicate — and confirm each probe fails against them. `rapp-postgres` verified
its drift check this way, by re-pinning to a bad digest and confirming exit 1;
the same discipline applies here and is not optional.
**Acceptance:** every probe in T03 demonstrated failing before any of them is
trusted passing.
Completed 2026-08-22. Five generic read/list/create/update/delete probes and
the three audit-core shaped probes pass the enforcing in-process fixture and
all produce findings when the tenant predicate is removed. Tenant-engine has
no applicable E2 identity, so its pack is not calibrated as if it were E2.
`make fixture-evidence` refreshes `evidence/offline-calibration.json`.
Target probes are still not trusted passing against a live service until a
new admitted engagement runs.
### T05 — RLS conformance under attack (the E3 artifact)
```task
id: WHITEHAT-WP-0001-T05
status: done
priority: medium
state_hub_task_id: "7c2c3dab-a135-543b-86b1-d543ec6af0fc"
```
`rapp-postgres` ADR-0003 supplies an `rls_conformance` view and a template, and
states plainly that the platform's guarantee is **detection, not prevention**
a table created by a later migration ships without a policy until something
notices. This repo is that something.
- Query the conformance view on a cadence; any row is a finding.
- Attack what the view cannot see: a session that sets no GUC must read
nothing; a session with another tenant's value must see nothing; an insert
attributed to another tenant must be refused.
- Attempt the documented bypasses: a `SECURITY DEFINER` function owned by the
table owner, and a role holding `BYPASSRLS`.
**Cadence is the deliverable here, not a detail.** For a detection-based
control the interval between runs *is* the exposure window, and ADR-0003 leaves
the number to this repo. Set it, and state the resulting window in the record.
**Acceptance:** a cadence and calibration record against every *applicable* E3
target. The artifact is the record, not a live SQL run. `platform-pg` is
registered `not_applicable` for an ordinary runtime conformance-view identity;
that record is the artifact for that target.
Done 2026-09-01: cadence is 24 hours plus run and reporting latency, with
event-triggered pre-promotion runs after schema, role, RLS or security-definer
changes. `src/whitehat_security/e3.py` encodes the seven expected outcomes,
keeps the SQL-compromise GUC reset labeled as E3's documented limit, and
calibrates known-good/known-bad in-process.
`evidence/offline-e3-calibration.json` is the fixture artifact. `platform-pg`
stays `not_applicable`. Whitehat will not relabel it to finish this task. A
live database run is owned by `WHITEHAT-WP-0006` and still requires a separately
reviewed runtime-safe surface, named database, and window.
### T06 — Noisy-neighbour characterisation (the P1/P2 artifact)
```task
id: WHITEHAT-WP-0001-T06
status: done
priority: medium
state_hub_task_id: "3a2d7a31-86b9-5c82-ab88-606b64381b2a"
```
The framework had to reword this artifact once already: its first draft
required proof that a saturating consumer "does not breach" another's
allowance, which shared infrastructure cannot provide.
What is achievable and therefore what this produces: a recorded baseline of
per-consumer resource usage; a run in which one consumer saturates its declared
allowance; evidence that the governance controls **bind**; and the degradation
co-residents experience, **measured and written down** rather than asserted
acceptable.
Runs inside a declared window per T01 — on a single-node rail with a six-
consumer connection ceiling, a saturation probe is an outage if run carelessly.
**Acceptance:** an in-process characterization evaluator that records governor
binding and neighbour degradation, proven against known-good and known-bad
samples. `shared-substrate` remains `pending` and is not an applicable live
target.
Done 2026-09-01: `src/whitehat_security/capacity.py` records baseline/loaded
latency, errors and throughput per consumer, governor binding, aggressor
peak/ceiling and neighbour degradation. Known-good binds and stays within
ceiling; known-bad detects an unbound governor, an exceeded ceiling, and a
missing neighbour sample. `evidence/offline-capacity-calibration.json` is the
fixture artifact. The in-process fixture is `fixture-capacity`.
`shared-substrate` stays `pending`. Whitehat will not relabel it to finish this
task. No live load has been generated; that residual is owned by
`WHITEHAT-WP-0006`.
### T07 — Reporting into risk-nexus
```task
id: WHITEHAT-WP-0001-T07
status: done
priority: medium
state_hub_task_id: "602d0f23-8406-5882-8f93-ec90151dd475"
```
Findings leave this repo in one direction. A run produces: what was attempted,
under which attacker model, when, against which posture claim, and the outcome.
It carries no severity — that is `risk-nexus`'s.
A **passing** run is also reported. "The attacks we thought of did not work" is
the honest claim, and recording it dated is what lets anyone see how stale the
assurance has become.
Done 2026-08-22: `schemas/run-report.schema.json` defines the minimized
evidence contract, `whitehat risk-message` renders pass and finding
deliveries without severity, and `whitehat deliver` queues target reports.
The first authorized target report is
`evidence/WH-ENG-20260822-AUDIT-E2-03.json`, delivered to `risk-nexus` as
State Hub message `40e3f825-fc70-4091-96d2-9ab01d42184a`. Fixture calibration
remains refused as target assurance.
### T08 — Governed test plane
The 2026-08-22 cutoff's missing infrastructure, encoded here so live work can
resume later without assembling authority during the run.
```task
id: WHITEHAT-WP-0001-T08
status: done
priority: high
state_hub_task_id: "98b818b2-30d5-584f-8c04-ead58136dcb5"
```
Completed 2026-08-22 as a repository contract, not a cluster provision:
- Target registration schema and catalog, including an honest
`not_applicable` state.
- Fail-closed admission: retired IDs, kill switch, approval class, namespace,
pinned digest, known-bad calibration, two identity handles.
- Credential broker interface that never returns secret values. The live
broker is unconnected and raises before any custody call.
- Rate watcher, lease cleanup, default-deny plane manifests, runner identity.
- Automatic outbox delivery of target reports only.
`ops-mason` was asked on 2026-08-22 to provision only the namespace, default
deny policy and runner service account from `plane/`. That message does not
authorize a pod, a credential, or traffic. A real custody projection still
waits on a new engagement ID.
## Sequencing
T01 gates all. T02 shapes T03/T05/T06. T04 gates trusting any of them. T08
gates live T03. T07 can follow T03.
## Session cutoff — 2026-08-22
The coordinating session ended with the workplan deliberately **active**. T01,
T02, T03, T04, T07 and T08 were done. T05 and T06 remained in progress.
`WH-ENG-20260822-AUDIT-E2-01` expired unused, `-02` aborted with zero packets,
and `-03` completed as a bounded target pass. Those identifiers are terminal
and must never be reused.
The exact earlier cutoff scope is recorded in
`docs/session-cutoff-2026-08-22.md` and `docs/test-plane.md`.
## Closeout — 2026-09-01
This workplan is **finished**. T01T08 are done for every applicable target.
`tenant-engine` E2 and `platform-pg` E3 stay `not_applicable`. `flex-auth` E2
and `shared-substrate` P1/P2 stay `pending`. Live E3, live P1/P2, flex-auth E2,
and a later audit-core E2 run are owned by `WHITEHAT-WP-0006`. That plan
authorizes no packet.
## Risks
**The facility becomes the threat.** Mitigated by T01, and by holding no
standing privilege.
**Probes weaken silently.** A probe that starts passing after a change to
itself rather than to the system is a finding, not a fix. Probe changes are
reviewed as security changes.
**Green is mistaken for safe.** Every report states that a pass means the
attacks attempted did not work, not that the boundary holds.
**It drifts into fixing things.** The boundary in INTENT is load-bearing:
findings route out, work does not come in.
## Open questions
1. **Owner: NetKingdom** — settled 2026-08-17. Offensive security is security
work. The residual tension (NetKingdom owning both the Tenancy Posture
framework and the facility that tests conformance to it) is mitigated by
findings routing out to `risk-nexus` under separate ownership, and is
recorded in INTENT rather than argued away.
2. **Where probes run from.** In-cluster gives realistic network position;
outside gives independence from the substrate under test. Probably both,
eventually; pick one to start.
3. **Does a service get told it is being probed?** Announced runs are easier to
operate; unannounced ones test the alerting too. Build mode probably
announced, production probably not — which is itself a T01 decision.