railiance-platform/workplans/RPF-WP-0036-platform-service-assurance.md
codex 08a406f2f8
All checks were successful
CI Smoke / host-smoke (push) Successful in 1s
CI Smoke / container-smoke (push) Successful in 2s
Close three-lane ESO recovery with live verification and cleanup evidence
Assistant: codex
Assistant-Model: gpt-6-astra
Assistant-Session: 01a06ecb-456a-71c2-b41e-0755d336e883
2026-09-05 19:00:19 +02:00

11 KiB

id type title domain repo status owner created updated state_hub_workstream_id
RPF-WP-0036 workplan Close S3 service assurance and ownership gaps financials railiance-platform blocked codex 2026-09-05 2026-09-05 ca639c3d-3a87-5fa4-ad13-6f2e014b0c84

S3 service assurance and ownership gaps

Source: history/2026-09-05-platform-intent-workplan-assessment.md. Reviewed against current repository evidence. This plan supplies the missing continuing obligations; it does not reopen completed bootstrap projects or duplicate incident/lane work. Repository design and read-only implementation can progress now. Every live drill, scheduler, credential operation or migration retains its own owner and execution gate.

Record the portfolio assessment and consolidate source work

id: RPF-WP-0036-T01
status: done
priority: high
state_hub_task_id: "9c6d0439-8a14-5921-9d88-95a10b303057"

Completed 2026-09-05. Assessed all 37 existing plans and their 168 task records, corrected SCOPE, grouped the remaining obligations, consolidated the three design follow-ups under RPF-WP-0035, archived completed plans with identities preserved, and recorded owner handoffs and before/after inventory in history. This certifies the source review, not live service health or external acceptance.

Publish achievable service guarantees and recovery ownership

id: RPF-WP-0036-T02
status: done
priority: high
state_hub_task_id: "eebcd5c7-4cfd-5084-9a21-ca1e4e748bdd"

For apps-pg, platform-pg, OpenBao and each supported backup delivery lane, publish a versioned service record: accountable S3/package/operator owners, consumers, failure domain, availability objective, RPO/RTO, retention, recovery key/quorum availability, maintenance/abort path and evidence freshness budget. Separate measured results from accepted targets and unknowns. A 56-second scratch restore is not an RTO commitment; one replica on one host is not HA. Reuse docs/s3-consumer-interfaces.md and existing package declarations.

Done when: every supported service has owner-reviewed numeric targets or an explicit unsupported guarantee and decision owner; consumer requirements are compared to the current substrate; any HA/node-loss gap has an exact S1/S2 and package dependency rather than a blanket new-cluster project here.

Make backup freshness and recurring recovery evidence checkable

id: RPF-WP-0036-T03
status: wait
priority: high
state_hub_task_id: "d64446fb-870b-5ae9-9628-f1fb9d06c5a4"

Inventory authoritative CNPG backup/PITR, OpenBao snapshot/isolated restore, encrypted off-host copy and custody-recovery receipts. Reuse existing validators and package status commands. Define cadence/expiry from T02; return distinct healthy, stale, missing and unavailable states using metadata only. Schedule execution only through the accepted execution owner and separately approved authority. Keep RPF-WP-0015's pending database/reboot experiments as the sole live tasks for those experiments; RPF-WP-0029 retains provider-key recovery.

Done when: a current off-host backup and a current isolated restore receipt exist for each supported data service, the approved cadence is installed and its execution is evidenced, and missing/stale/failed evidence reaches a named operator. A template, dated successful snapshot, or same-PVC reboot does not pass as restore proof. Record independent recovery-key access without values.

Produce S3 signals and prove their delivery to the evidence owner

id: RPF-WP-0036-T04
status: wait
priority: high
state_hub_task_id: "5351e0e4-6263-58f0-afb2-7c78c4cd6f68"

Define service-owned health semantics for backup/WAL age, restore age, seal state, ESO freshness, connection/memory headroom and consumer ceiling. Reuse package emitters and the Q2 owner's standard contract; retain an explicit unmonitored state and named manual checker until transport is accepted. Request a concrete receiving contract from railiance-telemetry when routing is authorized; do not implement a competing monitoring plane in S3.

Done when: bounded metadata-only samples pass contract validation, a controlled stale/failure sample reaches a named recipient through the accepted Q2 route, and missing emission itself is detectable. Local fixture tests may finish before the receiver, but end-to-end acceptance cannot.

Reconcile admission, placement and consumer interface drift

id: RPF-WP-0036-T05
status: done
priority: high
state_hub_task_id: "5cded2e9-7edd-57bd-9f7e-7977b75004dc"

Join actual package declarations and authorized metadata to the S3 interface, tenancy and placement records. Correct stale platform-pg occupancy/co-residency (Core Hub admission versus older tenant-engine descriptions), distinguish desired placement from observed placement, and verify the named overflow targets remain provisionable. Add a bounded check for missing owners, unsupported retention requests, quota/ceiling drift and stale evidence; consume package admission checks instead of reimplementing them.

Done when: every admitted consumer has one authoritative placement/contract, capacity and retention disclosures match package source and dated live proof, and synthetic invalid admissions fail before provisioning. No workload moves under this task without its own owner-reviewed migration.

Obtain acceptance for compatibility assets and derived-record cleanup

id: RPF-WP-0036-T06
status: wait
priority: medium
state_hub_task_id: "3b74f79c-7366-5d00-bf0e-31c0464a8f33"

Prepare exact source/entry-point inventories and owner-ready handoffs for Forgejo backup/pruning/image inventory (railiance-forge, activity-core execution), retained OpenBao package wrappers (rapp-openbao), and ArgoCD bootstrap/application manifests (S2/S4/S5 according to artifact). Keep S3 custody contracts and the RPF-WP-0029 exposure obligation here until accepted closure. No new framework or app-specific helper belongs here by default.

Supply repo-manager/State Hub with the exact legacy alias/source identity map from the assessment. Their apparent duplicate active records and stale brief are derived-state defects, not additional workplans. Use scoped reconciliation; never change managed UUIDs or blanket-acknowledge retirements to clean a view.

Done when: each retained compatibility surface has an accepting owner, canonical replacement and tested callers or a dated retention decision; the repo-filtered projection and generated brief agree with source identities. Unaccepted transfer remains explicitly pending. No requests were sent during the assessment and this task does not assert acceptance for another repo.

Decide demand and reuse for undeployed stateful capabilities

id: RPF-WP-0036-T07
status: done
priority: medium
state_hub_task_id: "b32709d2-581d-5750-a04c-e3494a42d27e"

Review cache, general object storage and messaging separately with potential consumers. Inventory existing providers/contracts (including artifact-store and the external backup bucket) before selecting an engine. Record workload, durability/latency/retention needs, capacity, tenancy, custody, package owner, recovery cost and operating owner for any accepted demand. Ask railiance-master to resolve fleet Q3 ownership through its architecture process; do not assign it to S3 by implication.

Done when: each capability has a dated decision to reuse, defer with a review trigger, or start a bounded consumer-backed delivery plan with explicit acceptance criteria. “No accepted demand; keep deploy gated” is a valid result. No Valkey, MinIO, RabbitMQ or new provider purchase is authorized by this plan.

Implementation and remaining acceptance — 2026-09-05

T02 completed with assurance/service-records.json and ADR-0004: all three CNPG cells, OpenBao and both backup delivery surfaces disclose unsupported availability/RPO/RTO guarantees, existing evidence, retention and named decision owners. Numeric commitments have not been invented or approved for other owners.

T05 completed with the owner-native admission checker, hash-bound baseline and placement registry. It reuses local/package validators and rejects unowned consumers, fifth consumers, unhonoured retention and source disclosure drift. Corrected tenant-engine's completed PostgreSQL cutover and the deployed SBOM overflow cell. Live metadata matches one instance, 1Gi limit, 100 connections and 30-day retention on all three deployed cells. Dated owner evidence remains the authority for database/application cutover; this run moved no workloads.

T07 completed with ADR-0005: reuse approved backup storage only for backup; evaluate artifact-store's existing S3 interface for artifact demand; defer cache, general object store and broker deployment with concrete demand/review triggers. The fleet Q3 question is retained in T06's prepared master handoff; no external architecture assignment or provisioning occurred.

T03/T04 have working read-only assurance-capture, assurance-check and assurance-admission entry points plus adversarial tests. The collector pins the cluster UID and reads selected status, native aggregate connection counts, metrics and unauthenticated seal status. It captures no Secrets, SQL text from sessions, application rows or logs. Healthy/stale/missing/unavailable/failed are distinct; unknown fields, future timestamps and wrong clusters fail closed. The output explicitly says unmonitored and unsupported guarantees.

  • T03 waits: fresh owner-validated isolated restore/snapshot/offsite receipts, independent recovery access, approved cadence and execution evidence. Existing live experiments stay in RPF-WP-0015; offsite exposure stays in RPF-WP-0029. No restore, snapshot creation, upload, scheduler or seal/unseal was run.
  • T04 waits: accepted Q2 receiving contract and controlled failure/absence delivery to a named recipient. The local producer is not a ratified Q2 contract. Live metadata also surfaced failed ESO resources: forgejo-mailer, reuse-surface-runtime and target-revenue-runtime. Their current/obsolete scope and exact repair require consumer/platform acceptance before lane mutation; failures remain visible rather than excluded.
  • T06 waits: ownership acceptance and scoped alias/brief repair. The exact path/hash/caller inventory, dated retention decision through 2026-10-05 and prepared owner requests are in assurance/ownership-handoffs.json and docs/platform-ownership-handoffs.md. The installed brief generator still includes open legacy aliases; rerunning it alone cannot satisfy acceptance. No messages or external handoff acceptance were fabricated.

Evidence: docs/evidence/RPF-WP-0036-assurance-2026-09-05.json. The plan is blocked on these explicit live/owner gates, not finished merely because its repository implementation and tests pass.

2026-09-05 ESO recovery follow-up: RPF-WP-0037 resolved all three delivery failures using exact Kubernetes authentication. Custody values were unchanged, negative checks and repeated refresh passed, obsolete invalid delivery tokens were removed, and all 27 ExternalSecrets now report Ready. The earlier dated assurance snapshot is retained as historical evidence. T04 still waits on the accepted Q2 receiver and controlled failure/absence transport proof.