diff --git a/REGISTER.md b/REGISTER.md index b950082..c492898 100644 --- a/REGISTER.md +++ b/REGISTER.md @@ -2,18 +2,18 @@ Generated by `tools/register_index.py` from `findings/`. Do not edit by hand. Last built 2026-08-21. -8 live of 9 findings; 3 notes below the floor. +7 live of 9 findings; 3 notes below the floor. ## Findings | ID | Finding | System | Severity | Disclosure | Escalation | Fix owner | Status | Cadence | Next check | | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | -| [RISK-F-0009](findings/RISK-F-0009-openbao-deny-set-covers-a-third-of-high-risk-lanes.md) | agent-high-risk-boundary denies 6 of 17 high-risk lanes; the direct bao path is unprotected for the rest | railiance-platform | **high** | embargoed | none | railiance-platform | open | 1h (1) | **due** | +| [RISK-F-0009](findings/RISK-F-0009-openbao-deny-set-covers-a-third-of-high-risk-lanes.md) | agent-high-risk-boundary denies 6 of 17 high-risk lanes; the direct bao path is unprotected for the rest | railiance-platform | **high** | embargoed | none | railiance-platform | open | instant (0) | **due** | | [RISK-F-0008](findings/RISK-F-0008-audit-retention-legal-basis-assumed.md) | The legal basis for retaining audit facts against an erasure request has been assumed, never established | audit-core | medium | public | **answered** (t2, answered) | risk-nexus | accepted | instant (0) | **due** | -| [RISK-F-0007](findings/RISK-F-0007-unverified-tenant-boundary.md) | No consumer's tenant boundary is verified anywhere | estate | **high** | embargoed | **answered** (t4, assigned) | per-consumer, on request | accepted | instant (0) | **due** | -| [RISK-F-0006](findings/RISK-F-0006-apps-pg-no-backup-configured.md) | apps-pg has no backup configured at all: R0 means no recovery | railiance-platform | **high** | embargoed | **answered** (t3, approved) | railiance-platform | open | 1h (1) | **due** | +| [RISK-F-0007](findings/RISK-F-0007-unverified-tenant-boundary.md) | No consumer's tenant boundary is verified anywhere | estate | **high** | embargoed | **answered** (t4, assigned) | per-consumer, on request | accepted | 1h (1) | 2026-08-21 08:32Z | +| [RISK-F-0006](findings/RISK-F-0006-apps-pg-no-backup-configured.md) | apps-pg has no backup configured at all: R0 means no recovery | railiance-platform | **high** | public | **answered** (t3, answered) | railiance-platform | fixed | instant (0) | **due** | | [RISK-F-0005](findings/RISK-F-0005-audit-core-unfiltered-read-path.md) | audit-core read path applies no tenant filter; the bound is deployment, not code | audit-core | medium | public | none | audit-core | mitigated | instant (0) | **due** | -| [RISK-F-0004](findings/RISK-F-0004-tenant-engine-unfiltered-event-read.md) | tenant-engine events() returns the entire event log unfiltered | tenant-engine | **high** | embargoed | none | tenant-engine | open | 1h (1) | **due** | +| [RISK-F-0004](findings/RISK-F-0004-tenant-engine-unfiltered-event-read.md) | tenant-engine events() returns the entire event log unfiltered | tenant-engine | medium | embargoed | none | tenant-engine | open | instant (0) | **due** | | [RISK-F-0003](findings/RISK-F-0003-ops-warden-read-boundary-ungraded-lanes.md) | ops-warden agent read-boundary does not fire on ungraded catalog lanes | ops-warden | medium | embargoed | none | ops-warden | mitigated | 8h (2) | 2026-08-21 14:32Z | | [RISK-F-0002](findings/RISK-F-0002-ops-warden-sign-ungated.md) | ops-warden signs SSH certificates with no authorization decision, and its unblock is now unsafe | ops-warden | medium | embargoed | **withdrawn** (t6, withdrawn-hazard-window-closed) | ops-warden | open | instant (0) | **due** | | [RISK-F-0001](findings/RISK-F-0001-flex-auth-unauthenticated-check.md) | flex-auth /v1/check authenticates no caller | flex-auth | **high** | public | **withdrawn** (t1, withdrawn-before-sending) | flex-auth | fixed | instant (0) | **due** | @@ -36,10 +36,7 @@ Silence never buys a softer grade — see `docs/method/dependencies.md`. | RISK-F-0009 | railiance-platform | embargo lifts on coverage; live verification would refine the grade but is not required for it | the eight uncovered paths stand as recorded and the finding is re-raised | 2026-09-03 | | RISK-F-0008 | audit-core | a working keyed commitment narrows RISK-REG-0001 to retained-by-obligation categories only | encrypt-then-hash recorded as the only known route, and the retention period recorded as unstateable | 2026-11-17 | | RISK-F-0007 | user-engine | likelihood falls for user-engine if a verification exists; a defect becomes its own finding if not | the on-request path is recorded as having produced no answer, which makes the acceptance itself unsupported and is escalated | 2026-09-03 | -| RISK-F-0006 | railiance-platform | embargo lifts on a demonstrated restore; the approved spend becomes a real figure | recorded as stalled with approval already granted, which is the worst kind of stall | 2026-09-18 | | RISK-F-0005 | audit-core | likelihood rises to L3 if any other production credential carries may_read | graded on the sender alone, as stated; the wider question is recorded as unanswered | 2026-09-19 | -| RISK-F-0004 | tenant-engine | grade rises if the log carries tenant payload rather than metadata | grade stands as recorded; absent fix tracking is recorded as a stall | 2026-09-03 | -| RISK-F-0002 | ops-warden | if it admits no ingress, enabling the gate stops all signing — an availability blocker, not a risk one | the register records the ordering as unverified and re-raises it; the finding stands at medium | 2026-08-27 | | RISK-F-0001 | policy-nexus | publication: published plus the permanent URL comes back onto each record | the findings stay disclosure: public with no address, which the register records as a claim rather than a publication | 2026-09-17 | ## Embargoes @@ -50,7 +47,6 @@ Held from publication with a stated condition. A hold with no moving condition i | --- | --- | --- | --- | | RISK-F-0009 | 2026-08-20 | railiance-platform reports the deny set covers every high-risk lane with a KV path (live verification refines the grade, it is not the condition) | — | | RISK-F-0007 | 2026-08-19 | a verification exists for at least one consumer boundary | 2026-09-18 | -| RISK-F-0006 | 2026-08-19 | a backup exists and a restore has been demonstrated once | 2026-09-18 | | RISK-F-0004 | 2026-08-19 | the read path filters by tenant in code | 2026-09-18 | | RISK-F-0003 | 2026-08-19 | RISK-F-0009 resolved — the OpenBao deny set covers every high-risk lane with a KV path | 2026-09-18 | | RISK-F-0002 | 2026-08-19 | FLEX-WP-0015-T02 shipped and ops-warden policy.enabled true in production | 2026-11-17 | diff --git a/docs/method/verification.md b/docs/method/verification.md index 312b8bb..e841cbe 100644 --- a/docs/method/verification.md +++ b/docs/method/verification.md @@ -66,6 +66,31 @@ worse one. If OpenBao claims need verifying, the right answer is for the owner to run the read and report it, which is what `ops-warden` did — the obstacle there was an expired token, not the arrangement. +## A file is not a safe proxy for the server + +Established 2026-08-21, by evidence rather than by caution. `ops-warden` ran +the live OpenBao read that `RISK-F-0009` had been graded without, and **the +deployed policy differs from the committed file.** No lane maps to the drifted +path, so nothing was exposed — but the general claim is now proven rather than +suspected. + +Consequences this register accepts: + +- **Every grade made off a checkout is a grade on a document**, and says so on + its face. `RISK-F-0009` does. +- A verification that reads a file is a **weaker artifact** than one that reads + a server, and the record must not blur them. `RISK-V-0001` reads a server; + the OpenBao comparison in `RISK-F-0009` reads a file. +- Where only the owner can reach the server, the owner's probe is the evidence + and the register says whose it is. That is not a lesser standard — it is the + correct one, given `verification.md`'s own limits. + +The corollary is uncomfortable and worth stating: **drift between file and +server is invisible to anyone reading files**, which is most of this estate's +tooling, including `fix_tracker.py`. What that tool reads is what a repo +*recorded*, and a record can be as stale as any other claim — as four +self-reported stale blockers in twelve hours demonstrated on 2026-08-21. + ## Recording a verification One file per verification in `docs/verifications/`, `RISK-V-NNNN`, stating the diff --git a/docs/verifications/2026-08-20-flex-auth-networkpolicy.md b/docs/verifications/2026-08-20-flex-auth-networkpolicy.md index 0d9d3d2..95bbe8a 100644 --- a/docs/verifications/2026-08-20-flex-auth-networkpolicy.md +++ b/docs/verifications/2026-08-20-flex-auth-networkpolicy.md @@ -98,3 +98,53 @@ what `RISK-F-0002` exists to prevent happening blind. `bao policy read agent-high-risk-boundary` both return **403 permission denied** from this host, the same wall `ops-warden` hit. `RISK-F-0009` therefore still rests on file comparison. + +--- + +## Reconciled — 2026-08-21: ingress does reach that pin, and both facts hold + +`ops-warden` answered ahead of their default date, and the answer is **no** — +enabling the gate would not have stopped signing. + +**Their evidence.** `GET http://127.0.0.1:19090/healthz` returns 200 through +the `flex-auth-ops-warden-railiance01` ops-bridge tunnel — a plain `-L` forward +from the railiance01 node to ClusterIP `10.43.1.165:8080` — answering +continuously since 2026-08-19, *after* the deny-all policy was created at +12:47:18Z that day. A full authenticated `/v1/check` through it returned 200, +`effect=allow`, `decision:f3f7c88f9585582a`. + +**The reconciliation, which keeps both observations true.** NetworkPolicy +governs pod-to-pod traffic. `ops-warden` is not a pod: `warden sign` runs on the +operator workstation and reaches the pin through an SSH tunnel terminating on +the node, so the traffic originates from the **node**, not from a pod. A +`podSelector`-scoped ingress rule does not constrain that path. + +A supporting datapoint neither they nor this register gathered: `net-kingdom` +reported on 2026-08-20 that a `tenant-engine` pod could not reach the +`user-engine` pin at all, `Errno 111`, attributed to NetworkPolicy. If that +holds, pod-sourced ingress *is* enforced and node-sourced is not — the expected +shape rather than an anomaly. + +**So the deny-all was deliberate and correctly scoped**, not a rollout artifact: +`flex-auth` set `consumer.isolated: true` on that pin, isolated until there is +an in-cluster PEP namespace to admit. + +### What this verification got right, and what it got wrong + +Right: reading the live policy, reporting the discrepancy, and routing it to +both owners **before either flipped a switch**. The observation was accurate. + +Wrong: the inference. "Ingress policyTypes with no rules means nothing reaches +this pin" is true of pod traffic and this register presented it without that +qualifier. `docs/method/verification.md` says a verification is evidence and +not a ruling; this one drifted toward a ruling and was corrected by the owner +inside a day, which is the system working. + +### The caveat `ops-warden` refused to paper over, carried here + +Their evidence proves ingress reaches the pin. It does **not** prove +NetworkPolicy is enforced everywhere it should be. **Anyone relying on a +NetworkPolicy to isolate something reachable from a node should probe that +assumption.** Nobody has. It is not a finding — no defect is established — and +it is exactly the shape of thing that becomes one the first time somebody +checks. diff --git a/findings/RISK-F-0002-ops-warden-sign-ungated.md b/findings/RISK-F-0002-ops-warden-sign-ungated.md index 5d89f38..8290b02 100644 --- a/findings/RISK-F-0002-ops-warden-sign-ungated.md +++ b/findings/RISK-F-0002-ops-warden-sign-ungated.md @@ -29,19 +29,13 @@ embargo_review: "2026-11-17" escalation: withdrawn escalation_trigger: 6 escalation_status: withdrawn-hazard-window-closed -last_checked: "2026-08-21T06:29:10Z" -next_check: "2026-08-21T06:29:10Z" +last_checked: "2026-08-21T07:32:09Z" +next_check: "2026-08-21T07:32:09Z" cadence: instant clean_streak: 0 -waiting_on: - - who: ops-warden - what: "probe whether the flex-auth pin you call admits ingress, before enabling policy.enabled" - since: "2026-08-20" - would_change: "if it admits no ingress, enabling the gate stops all signing — an availability blocker, not a risk one" - default: "the register records the ordering as unverified and re-raises it; the finding stands at medium" - default_at: "2026-08-27" graded_by: risk-nexus ruling: RISK-RULING-2026-08-19 +checked_by: "risk-nexus" --- # RISK-F-0002 — the SSH signing gate is off, and turning it on is now the more dangerous move @@ -324,3 +318,65 @@ Previously: probe whether the flex-auth pin admits ingress. Now, additionally: except the ingress question", that is a much shorter path than the one both repos have been describing. - **2026-08-21** — not clean: Fix tracking read for the first time: WARDEN-WP-0007 archived 2026-07-08, FLEX-WP-0007 finished 2026-06-29 — both closed before the finding was filed. Cadence instant → instant; checked again immediately. + +## Check — 2026-08-21: reading (c) confirmed — the blocker was stale, and the control is being retired + +`ops-warden` answered: **nothing blocks `policy.enabled`. It is off by +decision, not by blocker.** + +The gate is ready and verified — `flex-auth`'s `flex-auth-ops-warden` pin runs +`callerAuth.mode: enforce`, confirmed by both sides +(`decision:f3f7c88f9585582a`, anonymous `/v1/check` returns 401), and +`ops-warden` shipped the calling identity in `WARDEN-WP-0031`. `FLEX-WP-0007` +is finished, and so is `FLEX-WP-0016`, which `flex-auth` closed on 2026-08-19 +recording explicitly that the flip it was named after is not coming. + +**What replaced the blocker is a decision, not an obstacle.** `ops-warden` +`ADR-0006` (accepted): enforcement is zone-scoped, never a global flag. +`policy.enabled` is a single estate-wide boolean, and flipping it would enforce +uniformly across an estate under deep refactor. So it is not waiting to be +turned on — **it is being retired and replaced** by a per-zone control +(`ZONE-WP-0001`, with `WARDEN-WP-0032` consuming). + +### The disposition: the framing is superseded, the risk is not + +`ops-warden` proposed closed-by-supersession and left the grade here. Ruled: + +- **The framing is superseded.** "Gate shipped, disabled, blocked on + `FLEX-WP-0007`" describes a world that ended in June. Recording it as blocked + remediation misdescribes it, and they are right about that. +- **The risk stands, unchanged at `medium`.** Every production `warden sign` + still proceeds with no per-request authorization decision. That is what the + finding is about, and no part of it improved — a control retired before its + replacement exists is still an absent control. + +So the finding stays `open`, with a corrected fix path. `fix_tracking` becomes +`ZONE-WP-0001` / `WARDEN-WP-0032`. Nothing about the estate is worse today than +yesterday; what changed is that the register now knows why it is not better. + +### The successor's blocker is real, and it is the same missing join + +`ZONE-WP-0001-T03` cannot model stance because **26 of 27 credential lanes have +no identifiable workload** to attach a maturity ladder to. Escalated by +`ops-warden` to `repo-manager` and `net-kingdom` on 2026-08-20, unanswered. + +That is `RISK-N-0004` — the zone lookup this register routed as a note — with a +number on it. Three findings already wanted that facility; now the estate's +per-zone enforcement is blocked on it too. **The note is one instance short of +being a finding**, and the missing instance is somebody stating that the join +cannot be built. + +### Four stale blockers in twelve hours, self-reported + +`ops-warden` volunteered three more, all found the same day: an OpenBao token +recorded as expired that was valid; a verification script recorded as "ready" +that had never been written; and a ten-day blocker against `secrets-engine` +answerable from that repo's source. + +Their diagnosis is the one this register has been circling since `RISK-F-0001`: +**a blocker is written once, as prose, and then read as fact forever, because +nothing re-derives it and nothing expires it.** They have asked to adopt +whatever staleness convention this register settles rather than inventing a +second one. That is worth answering properly and is recorded as an open item +for the next round. +- **2026-08-21** — not clean: owner replied; see the dated check section Cadence instant → instant; checked again immediately. diff --git a/findings/RISK-F-0004-tenant-engine-unfiltered-event-read.md b/findings/RISK-F-0004-tenant-engine-unfiltered-event-read.md index cc37445..6143e73 100644 --- a/findings/RISK-F-0004-tenant-engine-unfiltered-event-read.md +++ b/findings/RISK-F-0004-tenant-engine-unfiltered-event-read.md @@ -11,13 +11,14 @@ date_filed: "2026-08-19" system: tenant-engine environment: production fix_owner: tenant-engine -fix_tracking: unset +fix_tracking: TEN-IN-0002 related: [RISK-F-0001] # Graded by risk-nexus 2026-08-19 — docs/rulings/2026-08-19-second-grading.md -severity: high -severity_at_production: critical +severity: medium +severity_at_production: high +severity_superseded: "high (2026-08-19), graded on a live network-reachable read that does not exist" impact: I3 -likelihood: L3 +likelihood: L1 fidelity_modifier: false production_rescore: true disclosure: embargoed @@ -25,19 +26,13 @@ embargo_condition: "the read path filters by tenant in code" embargo_since: "2026-08-19" embargo_review: "2026-09-18" escalation: none -last_checked: "2026-08-20T10:02:42Z" -next_check: "2026-08-20T11:02:42Z" -cadence: 1h -clean_streak: 1 -waiting_on: - - who: tenant-engine - what: "confirm or correct the unfiltered events() read; open fix tracking" - since: "2026-08-19" - would_change: "grade rises if the log carries tenant payload rather than metadata" - default: "grade stands as recorded; absent fix tracking is recorded as a stall" - default_at: "2026-09-03" +last_checked: "2026-08-21T07:32:09Z" +next_check: "2026-08-21T07:32:09Z" +cadence: instant +clean_streak: 0 graded_by: risk-nexus ruling: RISK-RULING-2026-08-19-B +checked_by: "risk-nexus" --- # RISK-F-0004 — the tenant event log is readable across tenants @@ -96,3 +91,34 @@ tenant data, and that is what `production_rescore` is for. Open at review: does `tenant-engine` confirm; is a fix tracked; what does the log actually contain. - **2026-08-20** — clean check: checked against the inbox and the owner's record; nothing moved. Cadence instant → 1h (1 clean in a row); next check 2026-08-20 11:02Z. + +## Check — 2026-08-21: confirmed, with a correction that cuts both ways + +`tenant-engine` answered, and the answer changes the grade in **both** +directions. + +**Worse than recorded:** `TenantStore.events()` returns the complete event list +**and payloads** without a tenant argument. The original grading assumed +cross-tenant *visibility* and said explicitly that it would rise if payloads +were involved. They are — payloads carry mutation evidence. + +**Much less reachable than recorded:** it is an in-process store protocol +method used by repository tests. **`tenant-engine` exposes no HTTP event-read +route**, so there is no live network-reachable cross-tenant read, which is what +this register wrote down and what `L3` was scored on. That was a reachability +claim the register inherited from a summary and never tested. + +`L3 → L1`. `I3` holds — the payload correction and the boundary crossing offset +each other rather than compounding. **`high → medium`**, `high` at production, +because the latent interface is still too broad and any future export route +inherits it. + +Fix tracking is now `TEN-IN-0002`: remove it from the production protocol, or +replace it with an authorized, deliberately scoped export interface plus +cross-tenant negatives. Their separate `TEN-IN-0001` covers a tamper-evident +audit emission gap and is not this finding. + +The correction is the useful part of the exchange, and it is the direction +reporters are usually reluctant to push: this register had overstated their +exposure for two days, in public-facing language, and they said so plainly. +- **2026-08-21** — not clean: owner replied; see the dated check section Cadence 1h → instant; checked again immediately. diff --git a/findings/RISK-F-0006-apps-pg-no-backup-configured.md b/findings/RISK-F-0006-apps-pg-no-backup-configured.md index 3b2b757..10e03ad 100644 --- a/findings/RISK-F-0006-apps-pg-no-backup-configured.md +++ b/findings/RISK-F-0006-apps-pg-no-backup-configured.md @@ -2,7 +2,7 @@ id: RISK-F-0006 type: finding title: "apps-pg has no backup configured at all: R0 means no recovery" -status: open +status: fixed reported_by: railiance-platform reported_via: flex-auth routed_by: risk-nexus @@ -11,7 +11,7 @@ date_filed: "2026-08-19" system: railiance-platform environment: production fix_owner: railiance-platform -fix_tracking: unset (railiance-platform to open) +fix_tracking: RPF-WP-0019 (finished) related: [RISK-F-0001] # Graded by risk-nexus 2026-08-19 — docs/rulings/2026-08-19-second-grading.md severity: high @@ -20,30 +20,26 @@ impact: I4 likelihood: L2 fidelity_modifier: false production_rescore: true -disclosure: embargoed -embargo_condition: "a backup exists and a restore has been demonstrated once" +disclosure: public +publication: pending-handover +embargo_lifted: "2026-08-20 — backup live, restore demonstrated with matching row counts" embargo_since: "2026-08-19" embargo_review: "2026-09-18" escalation: answered escalation_trigger: 3 -escalation_status: approved +escalation_status: answered +date_fixed: "2026-08-20" escalation_answered: "2026-08-19" escalation_answered_by: the-custodian escalation_act: approve decision: "spend for apps-pg backup storage approved; no ceiling stated" -last_checked: "2026-08-20T10:02:42Z" -next_check: "2026-08-20T11:02:42Z" -cadence: 1h -clean_streak: 1 -waiting_on: - - who: railiance-platform - what: "the backup target chosen, its monthly cost, and a demonstrated restore" - since: "2026-08-19" - would_change: "embargo lifts on a demonstrated restore; the approved spend becomes a real figure" - default: "recorded as stalled with approval already granted, which is the worst kind of stall" - default_at: "2026-09-18" +last_checked: "2026-08-21T07:32:09Z" +next_check: "2026-08-21T07:32:09Z" +cadence: instant +clean_streak: 0 graded_by: risk-nexus ruling: RISK-RULING-2026-08-19-B +checked_by: "risk-nexus" --- # RISK-F-0006 — the platform database cannot be recovered @@ -123,3 +119,31 @@ claim, not a control. - **2026-08-19** — escalation answered, spend approved. Open at review: is a backup configured; has a restore been demonstrated; what does it cost. - **2026-08-20** — clean check: checked against the inbox and the owner's record; nothing moved. Cadence instant → 1h (1 clean in a row); next check 2026-08-20 11:02Z. + +## Check — 2026-08-21: fixed, and the evidence was read rather than taken + +`railiance-platform` finished `RPF-WP-0019` on 2026-08-20. This register read +their evidence file rather than accepting the report: +`railiance-platform/docs/evidence/RPF-WP-0019-backup-restore-2026-08-20.md`. + +- Governed target live at `s3://railiance-platform-pg-backup/platform-pg/apps-pg/`, + 30-day retention, continuous WAL. +- `apps-pg-daily-20260820204148` completed in 8s, `LastBackupSucceeded=True`. +- **A separately named scratch cluster restored both consumer databases in 56s, + with exact database sizes and all 13 `coulomb_social` user-table row counts + matching.** The scratch namespace was deleted; production stayed `Ready`. + +**The embargo condition required a demonstrated restore and got one.** That +condition existed because a backup nobody has restored from is a claim, not a +control — and the demonstration is the difference between this finding closing +and merely appearing to. + +They also fixed something the grading had cited but not asked for: `apps-pg-1` +moved from `BestEffort` to `Burstable` QoS, with requests and limits, which was +part of why `L2` rather than `L1`. And their evidence records a failed first WAL +attempt against the wrong prefix rather than only the successful one, which is +the standard this register keeps asking of reporters. + +`fixed`, `public`, escalation closed. The spend approved on 2026-08-19 has a +real target behind it. +- **2026-08-21** — not clean: owner replied; see the dated check section Cadence 1h → instant; checked again immediately. diff --git a/findings/RISK-F-0007-unverified-tenant-boundary.md b/findings/RISK-F-0007-unverified-tenant-boundary.md index 680fc2c..c3c2b18 100644 --- a/findings/RISK-F-0007-unverified-tenant-boundary.md +++ b/findings/RISK-F-0007-unverified-tenant-boundary.md @@ -34,10 +34,10 @@ accepted_by: the-custodian accepted_on: "2026-08-19" accepted_until: "production transition (hard expiry, not a date)" decision: "pragmatic default before production — carried unverified; verification of a named consumer boundary on request" -last_checked: "2026-08-20T10:02:41Z" -next_check: "2026-08-20T10:02:41Z" -cadence: instant -clean_streak: 0 +last_checked: "2026-08-21T07:32:10Z" +next_check: "2026-08-21T08:32:10Z" +cadence: 1h +clean_streak: 1 waiting_on: - who: user-engine what: "does anything verify that a caller for tenant A cannot reach tenant B (RISK-V-0002)" @@ -47,6 +47,7 @@ waiting_on: default_at: "2026-09-03" graded_by: risk-nexus ruling: RISK-RULING-2026-08-19-B +checked_by: "risk-nexus" --- # RISK-F-0007 — nothing checks that tenants stay apart @@ -182,3 +183,4 @@ assumption that asking works. Grade unchanged. Nothing about the boundary itself has moved. - **2026-08-20** — not clean: On-request verification walked for the first time: RISK-V-0002 asks user-engine. Cadence instant → instant; checked again immediately. +- **2026-08-21** — clean check: no answer yet from user-engine; nothing about the boundary moved. Cadence instant → 1h (1 clean in a row); next check 2026-08-21 08:32Z. diff --git a/findings/RISK-F-0009-openbao-deny-set-covers-a-third-of-high-risk-lanes.md b/findings/RISK-F-0009-openbao-deny-set-covers-a-third-of-high-risk-lanes.md index 8e9c74f..4dd62e3 100644 --- a/findings/RISK-F-0009-openbao-deny-set-covers-a-third-of-high-risk-lanes.md +++ b/findings/RISK-F-0009-openbao-deny-set-covers-a-third-of-high-risk-lanes.md @@ -10,7 +10,7 @@ date_reported: "2026-08-20" system: railiance-platform environment: production fix_owner: railiance-platform -fix_tracking: unset +fix_tracking: unset (railiance-platform) filed_as: "RISK-F-0004 by ops-warden; renumbered by risk-nexus 2026-08-20 (id collision)" answers: RISK-F-0003 related: [RISK-F-0003] @@ -25,10 +25,10 @@ disclosure: embargoed embargo_condition: "railiance-platform reports the deny set covers every high-risk lane with a KV path (live verification refines the grade, it is not the condition)" embargo_since: "2026-08-20" escalation: none -last_checked: "2026-08-20T10:02:42Z" -next_check: "2026-08-20T11:02:42Z" -cadence: 1h -clean_streak: 1 +last_checked: "2026-08-21T07:32:09Z" +next_check: "2026-08-21T07:32:09Z" +cadence: instant +clean_streak: 0 waiting_on: - who: railiance-platform what: "report whether the deny set covers every high-risk lane with a KV path" @@ -38,6 +38,7 @@ waiting_on: default_at: "2026-09-03" graded_by: risk-nexus ruling: RISK-RULING-2026-08-20 +checked_by: "risk-nexus" --- # RISK-F-0009 — the OpenBao half of the agent read-boundary covers a third of the lanes @@ -239,3 +240,45 @@ otherwise. deployed policy match the file; do any agent tokens carry both policies; has `railiance-platform` taken the catalog-generated deny set. - **2026-08-20** — clean check: checked against the inbox and the owner's record; nothing moved. Cadence instant → 1h (1 clean in a row); next check 2026-08-20 11:02Z. + +## Check — 2026-08-21: measured rather than inferred, and one number corrected down + +`ops-warden` ran the live verification. Three changes, one of which matters +beyond this finding. + +**The headline stands and is now measured:** coverage confirmed at **6 of 17** +high-risk lanes. + +**The uncovered count was wrong in the direction that overstated the finding.** +Eight became **six**: the original eight included `openbao-api-key`, which is a +path *pattern* rather than an address, and `ops-warden-warden-sign-token`, which +is a broker grant — neither is deniable by a policy. The finding's own prose +had already flagged the first as "probably cannot be expressed as a policy +deny" while the table counted it anyway, and `ops-warden` corrected against +themselves rather than leaving it. + +The `I4` grade is unaffected: it was scored on the worst uncovered lane, and +`scaleway-bootstrap` and `agent-harness-forgejo-deploy` are both still in the +six. + +**The token was never expired.** The finding recorded `bao policy read` as +returning 403 and the comparison as static. `ops-warden` reports the token was +valid the whole time and the read succeeded on the first attempt. A stale claim +about the world sat in the record for a day — the fourth such instance they +self-reported in twelve hours. + +### The part that reaches past this finding + +**The deployed policy differs from the file.** `RISK-F-0009` named that as an +unconfirmed risk; it is now established. No `ops-warden` lane maps to the +drifted path, so exposure is unchanged — but *"the file is not a safe proxy for +the server"* has moved from suspicion to evidence, and that bears on **every +grade this register makes off a checkout**, including its own method. + +`docs/method/verification.md` is amended accordingly. + +**The embargo does not lift.** Its condition is a coverage report from +`railiance-platform`, and coverage did not improve — it stopped being an +inference, which is a different thing. `ops-warden` put the correction to them +directly and said so. +- **2026-08-21** — not clean: owner replied; see the dated check section Cadence 1h → instant; checked again immediately.