Five owner replies worked through: one fix closed, two grades corrected, one control retired

RISK-F-0006 fixed and public — railiance-platform's restore evidence was
read, not taken: 56s restore, all 13 coulomb_social row counts matching,
plus the BestEffort QoS this register had graded on, plus a failed first
WAL attempt recorded alongside the successful one.

RISK-F-0004 high -> medium. tenant-engine corrected in both directions:
payloads are returned (worse than graded) but there is no HTTP event-read
route, so the live network-reachable read this register wrote down does
not exist. L3 was a reachability claim inherited from a summary and never
tested.

RISK-F-0002: reading (c) confirmed — nothing blocks policy.enabled, it is
off by decision. ADR-0006 retires it in favour of zone-scoped
enforcement. Ruled: the framing is superseded, the risk is not. A control
retired before its replacement exists is still an absent control. The
successor's blocker is 26 of 27 lanes having no identifiable workload,
which is RISK-N-0004 with a number on it.

RISK-F-0009: uncovered count 8 -> 6, corrected by the reporter against
themselves; the token was never expired; and the deployed policy differs
from the file, which moves 'a file is not a safe proxy for the server'
from suspicion to evidence and amends verification.md — including the
admission that fix_tracker.py reads records, and a record can be stale.

RISK-V-0001 reconciled: ops-warden reaches the pin from the node through
a tunnel, so a podSelector ingress rule does not constrain it. The
observation was right and the inference was not.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
tegwick 2026-08-21 09:32:32 +02:00
parent 59e3e0a23a
commit e482423523
8 changed files with 280 additions and 58 deletions

View file

@ -2,7 +2,7 @@
id: RISK-F-0006
type: finding
title: "apps-pg has no backup configured at all: R0 means no recovery"
status: open
status: fixed
reported_by: railiance-platform
reported_via: flex-auth
routed_by: risk-nexus
@ -11,7 +11,7 @@ date_filed: "2026-08-19"
system: railiance-platform
environment: production
fix_owner: railiance-platform
fix_tracking: unset (railiance-platform to open)
fix_tracking: RPF-WP-0019 (finished)
related: [RISK-F-0001]
# Graded by risk-nexus 2026-08-19 — docs/rulings/2026-08-19-second-grading.md
severity: high
@ -20,30 +20,26 @@ impact: I4
likelihood: L2
fidelity_modifier: false
production_rescore: true
disclosure: embargoed
embargo_condition: "a backup exists and a restore has been demonstrated once"
disclosure: public
publication: pending-handover
embargo_lifted: "2026-08-20 — backup live, restore demonstrated with matching row counts"
embargo_since: "2026-08-19"
embargo_review: "2026-09-18"
escalation: answered
escalation_trigger: 3
escalation_status: approved
escalation_status: answered
date_fixed: "2026-08-20"
escalation_answered: "2026-08-19"
escalation_answered_by: the-custodian
escalation_act: approve
decision: "spend for apps-pg backup storage approved; no ceiling stated"
last_checked: "2026-08-20T10:02:42Z"
next_check: "2026-08-20T11:02:42Z"
cadence: 1h
clean_streak: 1
waiting_on:
- who: railiance-platform
what: "the backup target chosen, its monthly cost, and a demonstrated restore"
since: "2026-08-19"
would_change: "embargo lifts on a demonstrated restore; the approved spend becomes a real figure"
default: "recorded as stalled with approval already granted, which is the worst kind of stall"
default_at: "2026-09-18"
last_checked: "2026-08-21T07:32:09Z"
next_check: "2026-08-21T07:32:09Z"
cadence: instant
clean_streak: 0
graded_by: risk-nexus
ruling: RISK-RULING-2026-08-19-B
checked_by: "risk-nexus"
---
# RISK-F-0006 — the platform database cannot be recovered
@ -123,3 +119,31 @@ claim, not a control.
- **2026-08-19** — escalation answered, spend approved. Open at review: is a
backup configured; has a restore been demonstrated; what does it cost.
- **2026-08-20** — clean check: checked against the inbox and the owner's record; nothing moved. Cadence instant → 1h (1 clean in a row); next check 2026-08-20 11:02Z.
## Check — 2026-08-21: fixed, and the evidence was read rather than taken
`railiance-platform` finished `RPF-WP-0019` on 2026-08-20. This register read
their evidence file rather than accepting the report:
`railiance-platform/docs/evidence/RPF-WP-0019-backup-restore-2026-08-20.md`.
- Governed target live at `s3://railiance-platform-pg-backup/platform-pg/apps-pg/`,
30-day retention, continuous WAL.
- `apps-pg-daily-20260820204148` completed in 8s, `LastBackupSucceeded=True`.
- **A separately named scratch cluster restored both consumer databases in 56s,
with exact database sizes and all 13 `coulomb_social` user-table row counts
matching.** The scratch namespace was deleted; production stayed `Ready`.
**The embargo condition required a demonstrated restore and got one.** That
condition existed because a backup nobody has restored from is a claim, not a
control — and the demonstration is the difference between this finding closing
and merely appearing to.
They also fixed something the grading had cited but not asked for: `apps-pg-1`
moved from `BestEffort` to `Burstable` QoS, with requests and limits, which was
part of why `L2` rather than `L1`. And their evidence records a failed first WAL
attempt against the wrong prefix rather than only the successful one, which is
the standard this register keeps asking of reporters.
`fixed`, `public`, escalation closed. The spend approved on 2026-08-19 has a
real target behind it.
- **2026-08-21** — not clean: owner replied; see the dated check section Cadence 1h → instant; checked again immediately.