risk-nexus/findings/RISK-F-0004-tenant-engine-unfiltered-event-read.md

147 lines
6.5 KiB
Markdown
Raw Normal View History

---
id: RISK-F-0004
type: finding
title: "tenant-engine events() returns the entire event log unfiltered"
status: fixed
owner: risk-nexus
reported_by: tenant-engine
reported_via: flex-auth
routed_by: risk-nexus
date_reported: "2026-08-17"
date_filed: "2026-08-19"
system: tenant-engine
environment: production
fix_owner: tenant-engine
fix_tracking: TEN-WP-0011-T05 (finished 2026-08-29)
related: [RISK-F-0001]
# Graded by risk-nexus 2026-08-19 — docs/rulings/2026-08-19-second-grading.md
Five owner replies worked through: one fix closed, two grades corrected, one control retired RISK-F-0006 fixed and public — railiance-platform's restore evidence was read, not taken: 56s restore, all 13 coulomb_social row counts matching, plus the BestEffort QoS this register had graded on, plus a failed first WAL attempt recorded alongside the successful one. RISK-F-0004 high -> medium. tenant-engine corrected in both directions: payloads are returned (worse than graded) but there is no HTTP event-read route, so the live network-reachable read this register wrote down does not exist. L3 was a reachability claim inherited from a summary and never tested. RISK-F-0002: reading (c) confirmed — nothing blocks policy.enabled, it is off by decision. ADR-0006 retires it in favour of zone-scoped enforcement. Ruled: the framing is superseded, the risk is not. A control retired before its replacement exists is still an absent control. The successor's blocker is 26 of 27 lanes having no identifiable workload, which is RISK-N-0004 with a number on it. RISK-F-0009: uncovered count 8 -> 6, corrected by the reporter against themselves; the token was never expired; and the deployed policy differs from the file, which moves 'a file is not a safe proxy for the server' from suspicion to evidence and amends verification.md — including the admission that fix_tracker.py reads records, and a record can be stale. RISK-V-0001 reconciled: ops-warden reaches the pin from the node through a tunnel, so a podSelector ingress rule does not constrain it. The observation was right and the inference was not. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 09:32:32 +02:00
severity: medium
severity_at_production: high
severity_superseded: "high (2026-08-19), graded on a live network-reachable read that does not exist"
impact: I3
Five owner replies worked through: one fix closed, two grades corrected, one control retired RISK-F-0006 fixed and public — railiance-platform's restore evidence was read, not taken: 56s restore, all 13 coulomb_social row counts matching, plus the BestEffort QoS this register had graded on, plus a failed first WAL attempt recorded alongside the successful one. RISK-F-0004 high -> medium. tenant-engine corrected in both directions: payloads are returned (worse than graded) but there is no HTTP event-read route, so the live network-reachable read this register wrote down does not exist. L3 was a reachability claim inherited from a summary and never tested. RISK-F-0002: reading (c) confirmed — nothing blocks policy.enabled, it is off by decision. ADR-0006 retires it in favour of zone-scoped enforcement. Ruled: the framing is superseded, the risk is not. A control retired before its replacement exists is still an absent control. The successor's blocker is 26 of 27 lanes having no identifiable workload, which is RISK-N-0004 with a number on it. RISK-F-0009: uncovered count 8 -> 6, corrected by the reporter against themselves; the token was never expired; and the deployed policy differs from the file, which moves 'a file is not a safe proxy for the server' from suspicion to evidence and amends verification.md — including the admission that fix_tracker.py reads records, and a record can be stale. RISK-V-0001 reconciled: ops-warden reaches the pin from the node through a tunnel, so a podSelector ingress rule does not constrain it. The observation was right and the inference was not. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 09:32:32 +02:00
likelihood: L1
fidelity_modifier: false
production_rescore: true
disclosure: public
publication: pending-handover
publication_id: risk-f-0004-tenant-engine-unfiltered-event-read
publication_path: "findings/tenant-engine-unfiltered-event-read/v1/index.html"
publication_subtitle: "Tenant Engine's production store protocol exposed an unfiltered event-list method; it now exposes tenant-scoped reads only."
revision: "fixed-1"
last_reviewed: "2026-09-01"
review_interval: 6m
embargo_lifted: "2026-08-29 — unfiltered events() removed from the production protocol and tenant-scoped negative tests landed"
embargo_was_since: "2026-08-19"
escalation: none
date_fixed: "2026-08-29"
last_checked: "2026-09-01T00:32:44Z"
next_check: "2026-09-01T00:32:44Z"
Five owner replies worked through: one fix closed, two grades corrected, one control retired RISK-F-0006 fixed and public — railiance-platform's restore evidence was read, not taken: 56s restore, all 13 coulomb_social row counts matching, plus the BestEffort QoS this register had graded on, plus a failed first WAL attempt recorded alongside the successful one. RISK-F-0004 high -> medium. tenant-engine corrected in both directions: payloads are returned (worse than graded) but there is no HTTP event-read route, so the live network-reachable read this register wrote down does not exist. L3 was a reachability claim inherited from a summary and never tested. RISK-F-0002: reading (c) confirmed — nothing blocks policy.enabled, it is off by decision. ADR-0006 retires it in favour of zone-scoped enforcement. Ruled: the framing is superseded, the risk is not. A control retired before its replacement exists is still an absent control. The successor's blocker is 26 of 27 lanes having no identifiable workload, which is RISK-N-0004 with a number on it. RISK-F-0009: uncovered count 8 -> 6, corrected by the reporter against themselves; the token was never expired; and the deployed policy differs from the file, which moves 'a file is not a safe proxy for the server' from suspicion to evidence and amends verification.md — including the admission that fix_tracker.py reads records, and a record can be stale. RISK-V-0001 reconciled: ops-warden reaches the pin from the node through a tunnel, so a podSelector ingress rule does not constrain it. The observation was right and the inference was not. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 09:32:32 +02:00
cadence: instant
clean_streak: 0
graded_by: risk-nexus
ruling: RISK-RULING-2026-08-19-B
checked_by: "codex/risk-nexus"
---
# RISK-F-0004 — the tenant event log is readable across tenants
## What is true, as reported
`tenant-engine`'s `events()` returns the entire event log with no tenant
filter. Reported by `tenant-engine` as a live cross-tenant read at `E2` on the
Tenancy Posture E ladder, surfaced in the same review round as `RISK-F-0001`
and recorded inside that finding as visible-but-unfiled.
**This repo has not verified it and does not own the code.** It is filed on the
owning repo's own self-report. `tenant-engine` confirms or corrects it.
## Why it is being filed now
It was held back pending the precedent this register set for what warrants a
record of its own. That precedent now exists (`docs/method/severity.md`), and
"still waiting on the precedent" has stopped being an available answer. This
finding clears the floor on both tests: `tenant-engine` can act, and recording
it changes when they act.
## What is not established
- Whether any caller other than `tenant-engine` itself currently reaches
`events()`.
- What the log contains — whether the rows are metadata or carry tenant
payload. The grade assumes cross-tenant *visibility*, not payload
disclosure, and would rise if it is the latter.
- Whether a fix is tracked anywhere. `fix_tracking` is `unset` and this repo
has asked.
## Register ruling — 2026-08-19
`high` today (`I3` × `L3`), `critical` at production, embargoed until the read
path filters in code, no escalation.
`I3`: a cross-tenant read crosses a tenant boundary inside one system. `L3`:
no additional step is needed by anything that can already call the method, and
the authorization that would otherwise constrain the caller is `RISK-F-0001`,
which authenticates nobody.
`production_rescore: true`. Today the log holds no real counterparty's events,
which lowers what one occurrence costs; it does not lower what the defect is.
At production the same read is `I4` — real tenant data crossing a boundary
that nothing verifies — and the re-score is owed before any production
declaration completes.
No escalation: no real tenant data yet (trigger 1 reads *real*), owned, and
not yet stalled. It becomes an escalation on the day the estate takes real
tenant data, and that is what `production_rescore` is for.
## Reviews
- **2026-08-19** — filed and graded from `RISK-F-0001`'s unfiled list.
Open at review: does `tenant-engine` confirm; is a fix tracked; what does
the log actually contain.
- **2026-08-20** — clean check: checked against the inbox and the owner's record; nothing moved. Cadence instant → 1h (1 clean in a row); next check 2026-08-20 11:02Z.
Five owner replies worked through: one fix closed, two grades corrected, one control retired RISK-F-0006 fixed and public — railiance-platform's restore evidence was read, not taken: 56s restore, all 13 coulomb_social row counts matching, plus the BestEffort QoS this register had graded on, plus a failed first WAL attempt recorded alongside the successful one. RISK-F-0004 high -> medium. tenant-engine corrected in both directions: payloads are returned (worse than graded) but there is no HTTP event-read route, so the live network-reachable read this register wrote down does not exist. L3 was a reachability claim inherited from a summary and never tested. RISK-F-0002: reading (c) confirmed — nothing blocks policy.enabled, it is off by decision. ADR-0006 retires it in favour of zone-scoped enforcement. Ruled: the framing is superseded, the risk is not. A control retired before its replacement exists is still an absent control. The successor's blocker is 26 of 27 lanes having no identifiable workload, which is RISK-N-0004 with a number on it. RISK-F-0009: uncovered count 8 -> 6, corrected by the reporter against themselves; the token was never expired; and the deployed policy differs from the file, which moves 'a file is not a safe proxy for the server' from suspicion to evidence and amends verification.md — including the admission that fix_tracker.py reads records, and a record can be stale. RISK-V-0001 reconciled: ops-warden reaches the pin from the node through a tunnel, so a podSelector ingress rule does not constrain it. The observation was right and the inference was not. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 09:32:32 +02:00
## Check — 2026-08-21: confirmed, with a correction that cuts both ways
`tenant-engine` answered, and the answer changes the grade in **both**
directions.
**Worse than recorded:** `TenantStore.events()` returns the complete event list
**and payloads** without a tenant argument. The original grading assumed
cross-tenant *visibility* and said explicitly that it would rise if payloads
were involved. They are — payloads carry mutation evidence.
**Much less reachable than recorded:** it is an in-process store protocol
method used by repository tests. **`tenant-engine` exposes no HTTP event-read
route**, so there is no live network-reachable cross-tenant read, which is what
this register wrote down and what `L3` was scored on. That was a reachability
claim the register inherited from a summary and never tested.
`L3 → L1`. `I3` holds — the payload correction and the boundary crossing offset
each other rather than compounding. **`high → medium`**, `high` at production,
because the latent interface is still too broad and any future export route
inherits it.
Fix tracking is now `TEN-IN-0002`: remove it from the production protocol, or
replace it with an authorized, deliberately scoped export interface plus
cross-tenant negatives. Their separate `TEN-IN-0001` covers a tamper-evident
audit emission gap and is not this finding.
The correction is the useful part of the exchange, and it is the direction
reporters are usually reluctant to push: this register had overstated their
exposure for two days, in public-facing language, and they said so plainly.
- **2026-08-21** — not clean: owner replied; see the dated check section Cadence 1h → instant; checked again immediately.
## Closure — 2026-09-01: the broad method is gone
Tenant Engine reports `TEN-WP-0011-T05` complete: `TenantStore.events()` was
removed from the production protocol and replaced by `events_for(tenant_id)`.
Risk Nexus read the current in-memory, SQLite, and PostgreSQL implementations,
confirmed that all resolve and query by tenant, and ran the dedicated negative
suite: **3 passed**, including absence of the broad method and cross-tenant
isolation.
That is concrete source and regression evidence for the precise defect. The
finding is `fixed`; the embargo condition is met and publication is handed over.
- **2026-09-01** — not clean: TEN-WP-0011-T05 removed the broad protocol method and the focused tenant-scope suite passes; status fixed and embargo lifted. Cadence instant → instant; checked again immediately.