Assistant: codex Assistant-Model: gpt-6-astra Assistant-Session: 01a06ecb-456a-71c2-b41e-0755d336e883
8.4 KiB
S3 service assurance — v1
Workplan: RPF-WP-0036. Disclosure authority: ADR-0004.
Operator entry points
make assurance-admission
make assurance-capture
make assurance-check
Admission requires the adjacent reviewed rapp-postgres checkout and PyYAML,
which existing platform tooling already uses. Capture requires the existing
noninteractive SSH route to railiance01, with kubectl on that host. It pins
the kube-system UID from existing platform evidence and stops on mismatch.
It does not trust the default workstation kube context. Failure prints only a
fixed error code, never raw command output. No Secret, application row/log or credential query is run. A read-only
pg_stat_activity count and kubectl metrics supply aggregate headroom;
unauthenticated bao status supplies seal state. The observation is written to
/tmp/railiance-platform-observation.json by default; override with
ASSURANCE_OBSERVATION. Run the check only after capture succeeds.
evaluate exits 0 when every required sample is healthy, 1 for incomplete,
stale, failed or unavailable samples, and 2 for invalid/unavailable input.
There is intentionally no green overall service claim while evidence is
missing. To replay a dated observation:
python3 scripts/service_assurance.py evaluate observation.json --now 2026-09-05T12:00:00Z
--now is for replay/testing, not a way to freshen old evidence. The output
retains evaluation time, unsupported guarantees and unmonitored transport.
A pass sample is an attestation from its producer, not cryptographic proof.
Only approved owner receipts can close the workplan's live acceptance gates.
Supported source contract and missing adapters
The closed observation envelope carries schema, pinned cluster UID, capture
time and a configured set of signals. Each sample has only observed_at and
result (pass, fail, unavailable). Free text, values, paths and unknown
signals are rejected. For backup/restore/upload signals the observed time is
the completion time of that action; it is not the time someone read the file.
A capture older than 15 minutes never passes even when a backup is recent.
The collector supplies database readiness, last completed backup and continuous archiving for apps-pg/platform-pg/platform-pg-2, plus fleet ESO Ready/oldest-refresh status, OpenBao seal state, and aggregate connection/memory headroom. Headroom uses a local diagnostic ceiling of 80% of source connection/memory limits; it is not a workload latency guarantee. Metrics timestamps retain their age. ContinuousArchiving is a controller condition, not a measured WAL lag guarantee. The fleet ESO aggregate is deliberately a conservative inventory check; an obsolete resource must be explicitly classified by its owner before it is excluded. One ESO failure cannot disappear inside an average.
The following remain missing until a native value-safe adapter and acceptance exist: validated isolated restore receipts, OpenBao snapshot/restore proof, and offsite upload/restore receipts. Missing adapters are not inferred healthy from pod readiness. The local producer is not the Q2 standard; railiance-telemetry has no implemented receiving contract in the reviewed checkout. Integration, routing and scheduled delivery remain T04, and no notification was sent during implementation.
Service records and evidence inventory
assurance/service-records.json covers all three database cells, OpenBao,
Forgejo backup and multi-host CNPG logical offsite delivery. It distinguishes
unsupported objectives, declared retention, recovery custody and decision
owners. The old Forgejo cron example mentions an RPO 24h target; that is not a
newly accepted or verified current service guarantee. RPF-WP-0029 still gates
provider invalidation/replacement recovery for the shared offsite lane.
| Evidence | Existing authority | Current acceptance limit |
|---|---|---|
| CNPG completed backup/WAL | Cluster controller status; live capture | Current status only; does not exercise restore |
| apps-pg restore | RPF-WP-0019 evidence, 2026-08-20 | Dated 56-second scratch restore; no numeric RTO commitment |
| platform-pg restore/PITR | rapp-postgres restore procedure and evidence | Recurring current proof and independent recovery access remain T03 |
| platform-pg-2 restore | RAPP-POSTGRES-WP-0005-T04, 2026-08-22 | 51-second isolated restore after SBOM migration; not a continuous guarantee |
| OpenBao snapshot | WARDEN-WP-0027 preparation receipt, 2026-08-23 | Snapshot/encrypted off-host preparation, not isolated restore |
| OpenBao restore | Existing openbao-validate-restore-evidence.sh and package procedure |
Example receipt cannot pass as a fresh execution |
| Recovery exercise | RPF-WP-0015-T02/T03 | Separate windows, synthetic driver/quorum and abort operator required |
| Logical/Forgejo offsite | activity-core backup definitions, existing helper/runbooks | No new upload/restore performed; RPF-WP-0029 remains open |
Admission and disclosure drift
assurance/admission-baseline.json is generated by the admission command from
local CNPG CRs, placement-owners.json, and the native rapp-postgres renderer.
The checker reuses both native capacity validators. It also rejects missing
owners, changed cell placement, unhonoured inherited retention and mismatch
between effective package retention and the actual CR. Hashes pin every
consumed source artifact; changing package code or declarations requires review
before refreshing the baseline. Do not blindly regenerate it in CI.
Source occupancy: apps-pg 2/3, apps-pg-2 0/3 (not deployed), platform-pg 4/4 including the probe, platform-pg-2 1/4 for sbom-nexus. Tenant-engine's PostgreSQL cutover completed 2026-08-21; its old SQLite-current claim was stale. Platform-pg-2 was accepted and restored on 2026-08-22; it is no longer merely named overflow. Core Hub is latency-critical alongside audit-core batch work; no service-class priority is implied. See the generated baseline for per-cell details.
Source admission checks happen before deployment in the existing apps-pg and package owner paths; this additional check detects disclosure drift. It does not apply resources or become a second provisioning engine. Any new consumer needs an owner entry and a reviewed baseline update, even within free capacity.
Native recovery receipt adapter
capture_service_observation.py loads assurance/recovery-evidence.json through
recovery_evidence.py. The reviewed SHA-256 pins bind apps-pg and forgejo-db
restore samples to their Scaleway receipts. Original finished_at values drive
freshness; recapturing cannot extend their 30-day diagnostic validity. Missing,
changed, invalid or incorrectly scoped receipts yield unavailable samples.
Update pins only after reviewing replacement evidence. Native database recovery
does not attest full application or essentials recovery. Legacy archive receipts
without completion timestamps remain manual evidence; no timestamp is inferred
from file modification time. Automatic cadence and alert delivery remain pending.
New archive runs record UTC started_at and terminal finished_at in transfer,
decryption and isolated restore receipts, including failures. A timestamp alone
is not success: consumers must validate status, provider/profile, hashes and
cleanup. Restore completion is recorded after scratch cleanup, and cleanup
failure sets status to failed. Decryption and restore also record SHA-256 of
the exact input receipt bytes as transfer_receipt_sha256.
Decryption now retains platform.forgejo-primary-decryption.v1 and its own
operation times; older decryption receipts used the transfer schema through an
overwrite bug. The restore tool explicitly accepts both forms, with verified
hash/decryption flags. Existing historical receipts are unchanged. These producer
fixes enable future dated archive evidence; automatic archive adapters, fresh
end-to-end receipts and recurring execution are still pending.
The OpenBao snapshot adapter also accepts the reviewed, hash-pinned receipt in
reviews/. It requires the expected source cluster identity, encrypted off-host
custody, verified hashes and creation time. That time is a conservative age
anchor for snapshot creation, not proof of a later restore or renewed custody.
The August 22 receipt evaluates stale under the 36-hour budget. Updating the
index or reading the file cannot renew it; an isolated OpenBao restore remains
a separate obligation.