Assistant: codex Assistant-Model: gpt-6-astra Assistant-Session: 01a06ecb-456a-71c2-b41e-0755d336e883
6.3 KiB
S3 service assurance — v1
Workplan: RPF-WP-0036. Disclosure authority: ADR-0004.
Operator entry points
make assurance-admission
make assurance-capture
make assurance-check
Admission requires the adjacent reviewed rapp-postgres checkout and PyYAML,
which existing platform tooling already uses. Capture requires the existing
noninteractive SSH route to railiance01, with kubectl on that host. It pins
the kube-system UID from existing platform evidence and stops on mismatch.
It does not trust the default workstation kube context. Failure prints only a
fixed error code, never raw command output. No Secret, application row/log or credential query is run. A read-only
pg_stat_activity count and kubectl metrics supply aggregate headroom;
unauthenticated bao status supplies seal state. The observation is written to
/tmp/railiance-platform-observation.json by default; override with
ASSURANCE_OBSERVATION. Run the check only after capture succeeds.
evaluate exits 0 when every required sample is healthy, 1 for incomplete,
stale, failed or unavailable samples, and 2 for invalid/unavailable input.
There is intentionally no green overall service claim while evidence is
missing. To replay a dated observation:
python3 scripts/service_assurance.py evaluate observation.json --now 2026-09-05T12:00:00Z
--now is for replay/testing, not a way to freshen old evidence. The output
retains evaluation time, unsupported guarantees and unmonitored transport.
A pass sample is an attestation from its producer, not cryptographic proof.
Only approved owner receipts can close the workplan's live acceptance gates.
Supported source contract and missing adapters
The closed observation envelope carries schema, pinned cluster UID, capture
time and a configured set of signals. Each sample has only observed_at and
result (pass, fail, unavailable). Free text, values, paths and unknown
signals are rejected. For backup/restore/upload signals the observed time is
the completion time of that action; it is not the time someone read the file.
A capture older than 15 minutes never passes even when a backup is recent.
The collector supplies database readiness, last completed backup and continuous archiving for apps-pg/platform-pg/platform-pg-2, plus fleet ESO Ready/oldest-refresh status, OpenBao seal state, and aggregate connection/memory headroom. Headroom uses a local diagnostic ceiling of 80% of source connection/memory limits; it is not a workload latency guarantee. Metrics timestamps retain their age. ContinuousArchiving is a controller condition, not a measured WAL lag guarantee. The fleet ESO aggregate is deliberately a conservative inventory check; an obsolete resource must be explicitly classified by its owner before it is excluded. One ESO failure cannot disappear inside an average.
The following remain missing until a native value-safe adapter and acceptance exist: validated isolated restore receipts, OpenBao snapshot/restore proof, and offsite upload/restore receipts. Missing adapters are not inferred healthy from pod readiness. The local producer is not the Q2 standard; railiance-telemetry has no implemented receiving contract in the reviewed checkout. Integration, routing and scheduled delivery remain T04, and no notification was sent during implementation.
Service records and evidence inventory
assurance/service-records.json covers all three database cells, OpenBao,
Forgejo backup and multi-host CNPG logical offsite delivery. It distinguishes
unsupported objectives, declared retention, recovery custody and decision
owners. The old Forgejo cron example mentions an RPO 24h target; that is not a
newly accepted or verified current service guarantee. RPF-WP-0029 still gates
provider invalidation/replacement recovery for the shared offsite lane.
| Evidence | Existing authority | Current acceptance limit |
|---|---|---|
| CNPG completed backup/WAL | Cluster controller status; live capture | Current status only; does not exercise restore |
| apps-pg restore | RPF-WP-0019 evidence, 2026-08-20 | Dated 56-second scratch restore; no numeric RTO commitment |
| platform-pg restore/PITR | rapp-postgres restore procedure and evidence | Recurring current proof and independent recovery access remain T03 |
| platform-pg-2 restore | RAPP-POSTGRES-WP-0005-T04, 2026-08-22 | 51-second isolated restore after SBOM migration; not a continuous guarantee |
| OpenBao snapshot | WARDEN-WP-0027 preparation receipt, 2026-08-23 | Snapshot/encrypted off-host preparation, not isolated restore |
| OpenBao restore | Existing openbao-validate-restore-evidence.sh and package procedure |
Example receipt cannot pass as a fresh execution |
| Recovery exercise | RPF-WP-0015-T02/T03 | Separate windows, synthetic driver/quorum and abort operator required |
| Logical/Forgejo offsite | activity-core backup definitions, existing helper/runbooks | No new upload/restore performed; RPF-WP-0029 remains open |
Admission and disclosure drift
assurance/admission-baseline.json is generated by the admission command from
local CNPG CRs, placement-owners.json, and the native rapp-postgres renderer.
The checker reuses both native capacity validators. It also rejects missing
owners, changed cell placement, unhonoured inherited retention and mismatch
between effective package retention and the actual CR. Hashes pin every
consumed source artifact; changing package code or declarations requires review
before refreshing the baseline. Do not blindly regenerate it in CI.
Source occupancy: apps-pg 2/3, apps-pg-2 0/3 (not deployed), platform-pg 4/4 including the probe, platform-pg-2 1/4 for sbom-nexus. Tenant-engine's PostgreSQL cutover completed 2026-08-21; its old SQLite-current claim was stale. Platform-pg-2 was accepted and restored on 2026-08-22; it is no longer merely named overflow. Core Hub is latency-critical alongside audit-core batch work; no service-class priority is implied. See the generated baseline for per-cell details.
Source admission checks happen before deployment in the existing apps-pg and package owner paths; this additional check detects disclosure drift. It does not apply resources or become a second provisioning engine. Any new consumer needs an owner entry and a reviewed baseline update, even within free capacity.