The register had nine waits in four days, one four hops deep: F-0003's embargo waited on F-0009, which waited on railiance-platform, which waited on live OpenBao verification, which waited on a credential nobody has. No single link was wrong, which is why it needed a rule. docs/method/dependencies.md: the register never waits to decide, it decides and revises. Every wait carries who, what, since, what it would change, what happens if nobody answers, and the date that default applies. Depth one — a record never waits on a record that is itself waiting. Defaults are dates and are pessimistic: silence costs the grade the evidence supports rather than buying a softer one, and owners are told the default in advance because a default nobody was warned about is an ambush. Applied: F-0009's embargo now lifts on railiance-platform reporting coverage, with live verification as a refinement rather than a condition, cutting the F-0003 chain from four hops to two. All eight open waits are typed with defaults. make check reports them with age, owner and default date, flags defaults come due, and catches depth-two violations. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
125 lines
5.1 KiB
Markdown
125 lines
5.1 KiB
Markdown
---
|
||
id: RISK-F-0006
|
||
type: finding
|
||
title: "apps-pg has no backup configured at all: R0 means no recovery"
|
||
status: open
|
||
reported_by: railiance-platform
|
||
reported_via: flex-auth
|
||
routed_by: risk-nexus
|
||
date_reported: "2026-08-17"
|
||
date_filed: "2026-08-19"
|
||
system: railiance-platform
|
||
environment: production
|
||
fix_owner: railiance-platform
|
||
fix_tracking: unset (railiance-platform to open)
|
||
related: [RISK-F-0001]
|
||
# Graded by risk-nexus 2026-08-19 — docs/rulings/2026-08-19-second-grading.md
|
||
severity: high
|
||
severity_at_production: critical
|
||
impact: I4
|
||
likelihood: L2
|
||
fidelity_modifier: false
|
||
production_rescore: true
|
||
disclosure: embargoed
|
||
embargo_condition: "a backup exists and a restore has been demonstrated once"
|
||
embargo_since: "2026-08-19"
|
||
embargo_review: "2026-09-18"
|
||
escalation: answered
|
||
escalation_trigger: 3
|
||
escalation_status: approved
|
||
escalation_answered: "2026-08-19"
|
||
escalation_answered_by: the-custodian
|
||
escalation_act: approve
|
||
decision: "spend for apps-pg backup storage approved; no ceiling stated"
|
||
last_checked: "2026-08-20T10:02:42Z"
|
||
next_check: "2026-08-20T11:02:42Z"
|
||
cadence: 1h
|
||
clean_streak: 1
|
||
waiting_on:
|
||
- who: railiance-platform
|
||
what: "the backup target chosen, its monthly cost, and a demonstrated restore"
|
||
since: "2026-08-19"
|
||
would_change: "embargo lifts on a demonstrated restore; the approved spend becomes a real figure"
|
||
default: "recorded as stalled with approval already granted, which is the worst kind of stall"
|
||
default_at: "2026-09-18"
|
||
graded_by: risk-nexus
|
||
ruling: RISK-RULING-2026-08-19-B
|
||
---
|
||
|
||
# RISK-F-0006 — the platform database cannot be recovered
|
||
|
||
## What is true, as reported
|
||
|
||
`apps-pg` (`railiance-platform`) has no backup configured at all: no
|
||
`barmanObjectStore`, no retention policy, and `BestEffort` QoS on the pod.
|
||
Reported in the same round as `RISK-F-0001`, whose author noted that `R0` here
|
||
means **no recovery**, not merely no erasure policy.
|
||
|
||
Not verified by this repo. `railiance-platform` owns it.
|
||
|
||
## Why it is graded highest of the round despite being the least exciting
|
||
|
||
The other findings in this round are about someone reaching something. This one
|
||
is about nothing coming back. There is no attacker in it and no clever step —
|
||
an ordinary node failure, a lost volume, a mistaken `DROP`, and the data that
|
||
was in `apps-pg` is gone.
|
||
|
||
`BestEffort` QoS makes the pod the first candidate for eviction under memory
|
||
pressure. Eviction alone is an availability event and recoverable. What is not
|
||
recoverable is anything that damages the volume, and that is the case with no
|
||
answer today.
|
||
|
||
## Register ruling — 2026-08-19
|
||
|
||
`high` today (`I4` × `L2`, non-adversarial reading), `critical` at production,
|
||
embargoed, and **escalated on trigger 3 (spend)**.
|
||
|
||
`I4`: `apps-pg` is shared platform infrastructure, so the loss is not confined
|
||
to one system, and it is unrecoverable rather than degraded. `L2` on the
|
||
non-adversarial reading the scale gained for this finding: irrecoverable loss
|
||
needs a volume or corruption event, which is an ordinary failure the estate
|
||
has seen the shape of but not an expected weekly occurrence.
|
||
|
||
`production_rescore: true`: at production this is `L3`, because the surface
|
||
area of ordinary operational events grows with real traffic and real operators.
|
||
|
||
**Escalation, trigger 3.** Object storage for backups is recurring spend that
|
||
is not currently committed, and no repo can authorise it for itself. The ask
|
||
is narrow: approve a backup target and its cost, or state that the estate
|
||
accepts running `apps-pg` with no recovery and for how long. The second is a
|
||
legitimate answer in build mode and needs saying out loud rather than
|
||
happening by default.
|
||
|
||
The disclosure condition deliberately includes a demonstrated restore. A
|
||
backup nobody has restored from is a claim, not a control — the same class of
|
||
error as the gate in `RISK-F-0002`.
|
||
|
||
## Reviews
|
||
|
||
- **2026-08-19** — filed and graded from `RISK-F-0001`'s unfiled list.
|
||
Open at review: has the spend been ruled on; is a backup configured; has a
|
||
restore been demonstrated; does `railiance-platform` track it anywhere.
|
||
|
||
## Operator decision — 2026-08-19: approved
|
||
|
||
The spend is approved. `railiance-platform` may provision backup storage for
|
||
`apps-pg` without returning for authorisation.
|
||
|
||
No ceiling was stated, so none is recorded. The register asks
|
||
`railiance-platform` to report the actual target and its monthly cost once
|
||
chosen; a figure materially above the trigger-3 band (recurring €50/month)
|
||
comes back for confirmation rather than being assumed covered. That is the
|
||
register being careful with an open approval, not a condition on it.
|
||
|
||
**What is now blocking is work, not permission.** The finding stays `open` at
|
||
`high`, and the embargo condition is unchanged: a backup exists **and** a
|
||
restore has been demonstrated once. A backup nobody has restored from is a
|
||
claim, not a control.
|
||
|
||
`fix_tracking` is still `unset` and is now `railiance-platform`'s to open.
|
||
|
||
## Reviews
|
||
|
||
- **2026-08-19** — escalation answered, spend approved. Open at review: is a
|
||
backup configured; has a restore been demonstrated; what does it cost.
|
||
- **2026-08-20** — clean check: checked against the inbox and the owner's record; nothing moved. Cadence instant → 1h (1 clean in a row); next check 2026-08-20 11:02Z.
|