--- id: RISK-F-0001 type: finding title: "flex-auth /v1/check authenticates no caller" status: open reported_by: flex-auth reported_via: rapp-postgres routed_by: rapp-postgres date_reported: "2026-08-17" system: flex-auth environment: production fix_owner: flex-auth fix_tracking: FLEX-WP-0015-T02 # The three fields below are risk-nexus's, not the reporter's. Left unset # deliberately: the reporter says what is true, this repo says how bad it is # and who hears about it (INTENT, "What it does not own"). severity: unset disclosure: unset escalation: unset --- # RISK-F-0001 — flex-auth authenticates no caller on the decision surface ## What is true `POST /v1/check` and `POST /v1/batch_check` authenticate no caller. Any workload with network reach to the ClusterIP Service can assert any subject and any tenant and receive an authoritative **allow**. `flex-auth` is the estate's authorization oracle. Every service that delegates a decision to it is relying on an answer that anyone able to reach the pod can obtain for any identity they care to name. Self-reported by `flex-auth` as `A0` on their own inbound surface, in their Tenancy Posture review. Their words: "flex-auth is the estate's authorization oracle and it trusts its callers completely." ## How it was found Not by a probe, an incident, or an alert. By `flex-auth` assessing themselves against the Tenancy Posture A ladder during a review they were asked to do — and their own note says they did not know they were carrying it. That provenance matters for triage: nothing was watching for this, and nothing would have found it. It has presumably been true for as long as the endpoint has existed. ## Exposure, as far as the reporter stated it - The Service is `ClusterIP`, so reach requires a workload inside the cluster. - No claim was made that network policy restricts which workloads can reach it, and this record does not assume one. **If a default-deny NetworkPolicy fronts the service, that materially changes the exposure and should be verified rather than inferred** — `flex-auth` did not state it either way, and I have not checked, because doing so would be reporting on a system I do not own. ## What makes it worse than a single service's defect A false allow from this endpoint is not confined to `flex-auth`. It is the answer other services act on. `tenant-engine` separately reports that its own mutations are authorized by `flex-auth` and that direct authority over its rows would mean "privilege escalation across NetKingdom rather than data tampering confined to one store". The same reasoning applies to a forged allow. ## Owner and state `flex-auth` owns the fix and has tracked it as `FLEX-WP-0015-T02`, to ship through the staged-promotion path rather than a direct apply. They classify it as the only urgent item of their five follow-ups. Nothing is asked of them by this record beyond what they have already committed to. ## What this repo is asked to decide 1. **Severity.** Not the reporter's to set. 2. **Disclosure.** Build mode is currently public-by-default, and this is precisely the class of finding where that stops being obviously right — a live authorization bypass in the service every other service trusts. The controlled-disclosure scheme this repo anticipates does not exist yet, so the choice today is publish or hold, with no mechanism between them. 3. **Escalation.** Whether this reaches the operator personally. The candidate triggers in INTENT include "anything exposing real tenant data" — this exposes the decision that governs access to it, which may or may not be the same thing, and that judgement is this repo's. ## Related, reported at the same time and not yet filed Three further defects surfaced from the same review round. They are recorded here so they are visible, not filed as findings, because filing them was not asked for: - `tenant-engine` — `events()` returns the entire event log unfiltered. A live cross-tenant read at `E2`. - `audit-core` — read path applies no tenant filter; a credential with `may_read` can read any tenant's events. Bounded by deployment (`may_read: false` on the production sender) and not by code. Tracked `AUDIT-WP-0008-T04`. - `apps-pg` (`railiance-platform`) — no backup configured at all: no `barmanObjectStore`, no retention policy, `BestEffort` QoS. `R0` there means no recovery, not merely no erasure policy. All four were found the same way, by repos reading their own code against a ladder, within a day of each other. That is a fact about the estate's observability worth carrying into triage: **four live defects, none found by monitoring.**