railiance-platform/docs/adr/ADR-0004-service-assurance-disclosure.md
codex 234b1b559f
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 1s
Implement S3 service assurance and admission checks
Assistant: codex
Assistant-Model: gpt-6-astra
Assistant-Session: 01a06ecb-456a-71c2-b41e-0755d336e883
2026-09-05 11:43:55 +02:00

3 KiB

id title status revision owner date last_reviewed review_interval
RPLAT-ADR-0004 ADR-0004 — Disclose unsupported service guarantees and evidence freshness accepted 1.0 railiance-platform 2026-09-05 2026-09-05 6m

ADR-0004 — Service assurance disclosure

Context

The platform has dated backup/restore evidence and three single-instance CNPG cells, but no accepted numeric end-to-end availability/RPO/RTO commitments. One successful scratch restore is not a recovery guarantee. A successful source validator is not current live health. RPF-WP-0036 makes that distinction executable and visible to consumers.

Decision

Publish assurance/service-records.json as the S3-owned disclosure of supported and unsupported guarantees. Availability, RPO and RTO remain explicitly unsupported until the named service/package/consumer owners accept measurable targets backed by evidence. This is a platform disclosure decision, not an external owner's approval. A consumer requesting a stronger guarantee needs a recorded placement/recovery decision before admission. In particular, CNPG on one host does not provide host-loss HA; the dependency is S1 host failure domains, S2 runtime placement and package-level failover proof.

Use assurance/service-contract.json only for local diagnostics: 15-minute capture/health age, 36-hour successful backup/upload age, 30-day restore age, one-hour ESO refresh age, and 80% connection/memory occupancy for headroom. These intentionally conservative defaults expose missing evidence; they are neither SLAs nor approvals to install schedules. Service-specific cadence and receiver acceptance remain RPF-WP-0036-T03/T04. A different legitimate refresh interval needs an explicit service contract, not automatic acceptance of an indefinitely stale observation.

Consequences

The checker distinguishes healthy, stale, missing, unavailable and failed samples. It rejects wrong clusters, unknown payload fields and future times. Its receipt states transport: unmonitored and guarantees: unsupported. Receipt syntax and freshness validation do not prove the truth of a supplied observation or replace a signed native restore receipt. The live collector queries only fixed metadata/status fields, a pg_stat_activity count and unauthenticated seal status; it cannot fetch Secrets or run backup, restore or seal/unseal mutations.

Sources of restore proof remain the owning packages and their existing validators. The platform accepts those artifacts after an approved isolated exercise; it does not convert old Markdown or example JSON into fresh proof. No unattended jobs or Q2 transport are deployed under this decision.

Alternatives considered

Invent numeric targets from elapsed restore time: rejected because it creates an untested promise. Keep all checks waiting for telemetry: rejected because safe local diagnostics and unsupported-guarantee disclosure are useful now. Treat every source declaration as live: rejected because desired placement and actual deployment can diverge.