Five owner replies worked through: one fix closed, two grades corrected, one control retired

RISK-F-0006 fixed and public — railiance-platform's restore evidence was
read, not taken: 56s restore, all 13 coulomb_social row counts matching,
plus the BestEffort QoS this register had graded on, plus a failed first
WAL attempt recorded alongside the successful one.

RISK-F-0004 high -> medium. tenant-engine corrected in both directions:
payloads are returned (worse than graded) but there is no HTTP event-read
route, so the live network-reachable read this register wrote down does
not exist. L3 was a reachability claim inherited from a summary and never
tested.

RISK-F-0002: reading (c) confirmed — nothing blocks policy.enabled, it is
off by decision. ADR-0006 retires it in favour of zone-scoped
enforcement. Ruled: the framing is superseded, the risk is not. A control
retired before its replacement exists is still an absent control. The
successor's blocker is 26 of 27 lanes having no identifiable workload,
which is RISK-N-0004 with a number on it.

RISK-F-0009: uncovered count 8 -> 6, corrected by the reporter against
themselves; the token was never expired; and the deployed policy differs
from the file, which moves 'a file is not a safe proxy for the server'
from suspicion to evidence and amends verification.md — including the
admission that fix_tracker.py reads records, and a record can be stale.

RISK-V-0001 reconciled: ops-warden reaches the pin from the node through
a tunnel, so a podSelector ingress rule does not constrain it. The
observation was right and the inference was not.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
tegwick 2026-08-21 09:32:32 +02:00
parent 59e3e0a23a
commit e482423523
8 changed files with 280 additions and 58 deletions

View file

@ -29,19 +29,13 @@ embargo_review: "2026-11-17"
escalation: withdrawn
escalation_trigger: 6
escalation_status: withdrawn-hazard-window-closed
last_checked: "2026-08-21T06:29:10Z"
next_check: "2026-08-21T06:29:10Z"
last_checked: "2026-08-21T07:32:09Z"
next_check: "2026-08-21T07:32:09Z"
cadence: instant
clean_streak: 0
waiting_on:
- who: ops-warden
what: "probe whether the flex-auth pin you call admits ingress, before enabling policy.enabled"
since: "2026-08-20"
would_change: "if it admits no ingress, enabling the gate stops all signing — an availability blocker, not a risk one"
default: "the register records the ordering as unverified and re-raises it; the finding stands at medium"
default_at: "2026-08-27"
graded_by: risk-nexus
ruling: RISK-RULING-2026-08-19
checked_by: "risk-nexus"
---
# RISK-F-0002 — the SSH signing gate is off, and turning it on is now the more dangerous move
@ -324,3 +318,65 @@ Previously: probe whether the flex-auth pin admits ingress. Now, additionally:
except the ingress question", that is a much shorter path than the one both
repos have been describing.
- **2026-08-21** — not clean: Fix tracking read for the first time: WARDEN-WP-0007 archived 2026-07-08, FLEX-WP-0007 finished 2026-06-29 — both closed before the finding was filed. Cadence instant → instant; checked again immediately.
## Check — 2026-08-21: reading (c) confirmed — the blocker was stale, and the control is being retired
`ops-warden` answered: **nothing blocks `policy.enabled`. It is off by
decision, not by blocker.**
The gate is ready and verified — `flex-auth`'s `flex-auth-ops-warden` pin runs
`callerAuth.mode: enforce`, confirmed by both sides
(`decision:f3f7c88f9585582a`, anonymous `/v1/check` returns 401), and
`ops-warden` shipped the calling identity in `WARDEN-WP-0031`. `FLEX-WP-0007`
is finished, and so is `FLEX-WP-0016`, which `flex-auth` closed on 2026-08-19
recording explicitly that the flip it was named after is not coming.
**What replaced the blocker is a decision, not an obstacle.** `ops-warden`
`ADR-0006` (accepted): enforcement is zone-scoped, never a global flag.
`policy.enabled` is a single estate-wide boolean, and flipping it would enforce
uniformly across an estate under deep refactor. So it is not waiting to be
turned on — **it is being retired and replaced** by a per-zone control
(`ZONE-WP-0001`, with `WARDEN-WP-0032` consuming).
### The disposition: the framing is superseded, the risk is not
`ops-warden` proposed closed-by-supersession and left the grade here. Ruled:
- **The framing is superseded.** "Gate shipped, disabled, blocked on
`FLEX-WP-0007`" describes a world that ended in June. Recording it as blocked
remediation misdescribes it, and they are right about that.
- **The risk stands, unchanged at `medium`.** Every production `warden sign`
still proceeds with no per-request authorization decision. That is what the
finding is about, and no part of it improved — a control retired before its
replacement exists is still an absent control.
So the finding stays `open`, with a corrected fix path. `fix_tracking` becomes
`ZONE-WP-0001` / `WARDEN-WP-0032`. Nothing about the estate is worse today than
yesterday; what changed is that the register now knows why it is not better.
### The successor's blocker is real, and it is the same missing join
`ZONE-WP-0001-T03` cannot model stance because **26 of 27 credential lanes have
no identifiable workload** to attach a maturity ladder to. Escalated by
`ops-warden` to `repo-manager` and `net-kingdom` on 2026-08-20, unanswered.
That is `RISK-N-0004` — the zone lookup this register routed as a note — with a
number on it. Three findings already wanted that facility; now the estate's
per-zone enforcement is blocked on it too. **The note is one instance short of
being a finding**, and the missing instance is somebody stating that the join
cannot be built.
### Four stale blockers in twelve hours, self-reported
`ops-warden` volunteered three more, all found the same day: an OpenBao token
recorded as expired that was valid; a verification script recorded as "ready"
that had never been written; and a ten-day blocker against `secrets-engine`
answerable from that repo's source.
Their diagnosis is the one this register has been circling since `RISK-F-0001`:
**a blocker is written once, as prose, and then read as fact forever, because
nothing re-derives it and nothing expires it.** They have asked to adopt
whatever staleness convention this register settles rather than inventing a
second one. That is worth answering properly and is recorded as an open item
for the next round.
- **2026-08-21** — not clean: owner replied; see the dated check section Cadence instant → instant; checked again immediately.

View file

@ -11,13 +11,14 @@ date_filed: "2026-08-19"
system: tenant-engine
environment: production
fix_owner: tenant-engine
fix_tracking: unset
fix_tracking: TEN-IN-0002
related: [RISK-F-0001]
# Graded by risk-nexus 2026-08-19 — docs/rulings/2026-08-19-second-grading.md
severity: high
severity_at_production: critical
severity: medium
severity_at_production: high
severity_superseded: "high (2026-08-19), graded on a live network-reachable read that does not exist"
impact: I3
likelihood: L3
likelihood: L1
fidelity_modifier: false
production_rescore: true
disclosure: embargoed
@ -25,19 +26,13 @@ embargo_condition: "the read path filters by tenant in code"
embargo_since: "2026-08-19"
embargo_review: "2026-09-18"
escalation: none
last_checked: "2026-08-20T10:02:42Z"
next_check: "2026-08-20T11:02:42Z"
cadence: 1h
clean_streak: 1
waiting_on:
- who: tenant-engine
what: "confirm or correct the unfiltered events() read; open fix tracking"
since: "2026-08-19"
would_change: "grade rises if the log carries tenant payload rather than metadata"
default: "grade stands as recorded; absent fix tracking is recorded as a stall"
default_at: "2026-09-03"
last_checked: "2026-08-21T07:32:09Z"
next_check: "2026-08-21T07:32:09Z"
cadence: instant
clean_streak: 0
graded_by: risk-nexus
ruling: RISK-RULING-2026-08-19-B
checked_by: "risk-nexus"
---
# RISK-F-0004 — the tenant event log is readable across tenants
@ -96,3 +91,34 @@ tenant data, and that is what `production_rescore` is for.
Open at review: does `tenant-engine` confirm; is a fix tracked; what does
the log actually contain.
- **2026-08-20** — clean check: checked against the inbox and the owner's record; nothing moved. Cadence instant → 1h (1 clean in a row); next check 2026-08-20 11:02Z.
## Check — 2026-08-21: confirmed, with a correction that cuts both ways
`tenant-engine` answered, and the answer changes the grade in **both**
directions.
**Worse than recorded:** `TenantStore.events()` returns the complete event list
**and payloads** without a tenant argument. The original grading assumed
cross-tenant *visibility* and said explicitly that it would rise if payloads
were involved. They are — payloads carry mutation evidence.
**Much less reachable than recorded:** it is an in-process store protocol
method used by repository tests. **`tenant-engine` exposes no HTTP event-read
route**, so there is no live network-reachable cross-tenant read, which is what
this register wrote down and what `L3` was scored on. That was a reachability
claim the register inherited from a summary and never tested.
`L3 → L1`. `I3` holds — the payload correction and the boundary crossing offset
each other rather than compounding. **`high → medium`**, `high` at production,
because the latent interface is still too broad and any future export route
inherits it.
Fix tracking is now `TEN-IN-0002`: remove it from the production protocol, or
replace it with an authorized, deliberately scoped export interface plus
cross-tenant negatives. Their separate `TEN-IN-0001` covers a tamper-evident
audit emission gap and is not this finding.
The correction is the useful part of the exchange, and it is the direction
reporters are usually reluctant to push: this register had overstated their
exposure for two days, in public-facing language, and they said so plainly.
- **2026-08-21** — not clean: owner replied; see the dated check section Cadence 1h → instant; checked again immediately.

View file

@ -2,7 +2,7 @@
id: RISK-F-0006
type: finding
title: "apps-pg has no backup configured at all: R0 means no recovery"
status: open
status: fixed
reported_by: railiance-platform
reported_via: flex-auth
routed_by: risk-nexus
@ -11,7 +11,7 @@ date_filed: "2026-08-19"
system: railiance-platform
environment: production
fix_owner: railiance-platform
fix_tracking: unset (railiance-platform to open)
fix_tracking: RPF-WP-0019 (finished)
related: [RISK-F-0001]
# Graded by risk-nexus 2026-08-19 — docs/rulings/2026-08-19-second-grading.md
severity: high
@ -20,30 +20,26 @@ impact: I4
likelihood: L2
fidelity_modifier: false
production_rescore: true
disclosure: embargoed
embargo_condition: "a backup exists and a restore has been demonstrated once"
disclosure: public
publication: pending-handover
embargo_lifted: "2026-08-20 — backup live, restore demonstrated with matching row counts"
embargo_since: "2026-08-19"
embargo_review: "2026-09-18"
escalation: answered
escalation_trigger: 3
escalation_status: approved
escalation_status: answered
date_fixed: "2026-08-20"
escalation_answered: "2026-08-19"
escalation_answered_by: the-custodian
escalation_act: approve
decision: "spend for apps-pg backup storage approved; no ceiling stated"
last_checked: "2026-08-20T10:02:42Z"
next_check: "2026-08-20T11:02:42Z"
cadence: 1h
clean_streak: 1
waiting_on:
- who: railiance-platform
what: "the backup target chosen, its monthly cost, and a demonstrated restore"
since: "2026-08-19"
would_change: "embargo lifts on a demonstrated restore; the approved spend becomes a real figure"
default: "recorded as stalled with approval already granted, which is the worst kind of stall"
default_at: "2026-09-18"
last_checked: "2026-08-21T07:32:09Z"
next_check: "2026-08-21T07:32:09Z"
cadence: instant
clean_streak: 0
graded_by: risk-nexus
ruling: RISK-RULING-2026-08-19-B
checked_by: "risk-nexus"
---
# RISK-F-0006 — the platform database cannot be recovered
@ -123,3 +119,31 @@ claim, not a control.
- **2026-08-19** — escalation answered, spend approved. Open at review: is a
backup configured; has a restore been demonstrated; what does it cost.
- **2026-08-20** — clean check: checked against the inbox and the owner's record; nothing moved. Cadence instant → 1h (1 clean in a row); next check 2026-08-20 11:02Z.
## Check — 2026-08-21: fixed, and the evidence was read rather than taken
`railiance-platform` finished `RPF-WP-0019` on 2026-08-20. This register read
their evidence file rather than accepting the report:
`railiance-platform/docs/evidence/RPF-WP-0019-backup-restore-2026-08-20.md`.
- Governed target live at `s3://railiance-platform-pg-backup/platform-pg/apps-pg/`,
30-day retention, continuous WAL.
- `apps-pg-daily-20260820204148` completed in 8s, `LastBackupSucceeded=True`.
- **A separately named scratch cluster restored both consumer databases in 56s,
with exact database sizes and all 13 `coulomb_social` user-table row counts
matching.** The scratch namespace was deleted; production stayed `Ready`.
**The embargo condition required a demonstrated restore and got one.** That
condition existed because a backup nobody has restored from is a claim, not a
control — and the demonstration is the difference between this finding closing
and merely appearing to.
They also fixed something the grading had cited but not asked for: `apps-pg-1`
moved from `BestEffort` to `Burstable` QoS, with requests and limits, which was
part of why `L2` rather than `L1`. And their evidence records a failed first WAL
attempt against the wrong prefix rather than only the successful one, which is
the standard this register keeps asking of reporters.
`fixed`, `public`, escalation closed. The spend approved on 2026-08-19 has a
real target behind it.
- **2026-08-21** — not clean: owner replied; see the dated check section Cadence 1h → instant; checked again immediately.

View file

@ -34,10 +34,10 @@ accepted_by: the-custodian
accepted_on: "2026-08-19"
accepted_until: "production transition (hard expiry, not a date)"
decision: "pragmatic default before production — carried unverified; verification of a named consumer boundary on request"
last_checked: "2026-08-20T10:02:41Z"
next_check: "2026-08-20T10:02:41Z"
cadence: instant
clean_streak: 0
last_checked: "2026-08-21T07:32:10Z"
next_check: "2026-08-21T08:32:10Z"
cadence: 1h
clean_streak: 1
waiting_on:
- who: user-engine
what: "does anything verify that a caller for tenant A cannot reach tenant B (RISK-V-0002)"
@ -47,6 +47,7 @@ waiting_on:
default_at: "2026-09-03"
graded_by: risk-nexus
ruling: RISK-RULING-2026-08-19-B
checked_by: "risk-nexus"
---
# RISK-F-0007 — nothing checks that tenants stay apart
@ -182,3 +183,4 @@ assumption that asking works.
Grade unchanged. Nothing about the boundary itself has moved.
- **2026-08-20** — not clean: On-request verification walked for the first time: RISK-V-0002 asks user-engine. Cadence instant → instant; checked again immediately.
- **2026-08-21** — clean check: no answer yet from user-engine; nothing about the boundary moved. Cadence instant → 1h (1 clean in a row); next check 2026-08-21 08:32Z.

View file

@ -10,7 +10,7 @@ date_reported: "2026-08-20"
system: railiance-platform
environment: production
fix_owner: railiance-platform
fix_tracking: unset
fix_tracking: unset (railiance-platform)
filed_as: "RISK-F-0004 by ops-warden; renumbered by risk-nexus 2026-08-20 (id collision)"
answers: RISK-F-0003
related: [RISK-F-0003]
@ -25,10 +25,10 @@ disclosure: embargoed
embargo_condition: "railiance-platform reports the deny set covers every high-risk lane with a KV path (live verification refines the grade, it is not the condition)"
embargo_since: "2026-08-20"
escalation: none
last_checked: "2026-08-20T10:02:42Z"
next_check: "2026-08-20T11:02:42Z"
cadence: 1h
clean_streak: 1
last_checked: "2026-08-21T07:32:09Z"
next_check: "2026-08-21T07:32:09Z"
cadence: instant
clean_streak: 0
waiting_on:
- who: railiance-platform
what: "report whether the deny set covers every high-risk lane with a KV path"
@ -38,6 +38,7 @@ waiting_on:
default_at: "2026-09-03"
graded_by: risk-nexus
ruling: RISK-RULING-2026-08-20
checked_by: "risk-nexus"
---
# RISK-F-0009 — the OpenBao half of the agent read-boundary covers a third of the lanes
@ -239,3 +240,45 @@ otherwise.
deployed policy match the file; do any agent tokens carry both policies; has
`railiance-platform` taken the catalog-generated deny set.
- **2026-08-20** — clean check: checked against the inbox and the owner's record; nothing moved. Cadence instant → 1h (1 clean in a row); next check 2026-08-20 11:02Z.
## Check — 2026-08-21: measured rather than inferred, and one number corrected down
`ops-warden` ran the live verification. Three changes, one of which matters
beyond this finding.
**The headline stands and is now measured:** coverage confirmed at **6 of 17**
high-risk lanes.
**The uncovered count was wrong in the direction that overstated the finding.**
Eight became **six**: the original eight included `openbao-api-key`, which is a
path *pattern* rather than an address, and `ops-warden-warden-sign-token`, which
is a broker grant — neither is deniable by a policy. The finding's own prose
had already flagged the first as "probably cannot be expressed as a policy
deny" while the table counted it anyway, and `ops-warden` corrected against
themselves rather than leaving it.
The `I4` grade is unaffected: it was scored on the worst uncovered lane, and
`scaleway-bootstrap` and `agent-harness-forgejo-deploy` are both still in the
six.
**The token was never expired.** The finding recorded `bao policy read` as
returning 403 and the comparison as static. `ops-warden` reports the token was
valid the whole time and the read succeeded on the first attempt. A stale claim
about the world sat in the record for a day — the fourth such instance they
self-reported in twelve hours.
### The part that reaches past this finding
**The deployed policy differs from the file.** `RISK-F-0009` named that as an
unconfirmed risk; it is now established. No `ops-warden` lane maps to the
drifted path, so exposure is unchanged — but *"the file is not a safe proxy for
the server"* has moved from suspicion to evidence, and that bears on **every
grade this register makes off a checkout**, including its own method.
`docs/method/verification.md` is amended accordingly.
**The embargo does not lift.** Its condition is a coverage report from
`railiance-platform`, and coverage did not improve — it stopped being an
inference, which is a different thing. `ops-warden` put the correction to them
directly and said so.
- **2026-08-21** — not clean: owner replied; see the dated check section Cadence 1h → instant; checked again immediately.