railiance-platform/workplans/RPF-WP-0036-platform-service-assurance.md
codex 4c320cf053
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 1s
Track Q2 receiver implementation and live acceptance dependency
Assistant: codex
Assistant-Model: gpt-6-astra
Assistant-Session: 01a06ecb-456a-71c2-b41e-0755d336e883
2026-09-06 19:18:14 +02:00

16 KiB

id type title domain repo status owner created updated state_hub_workstream_id
RPF-WP-0036 workplan Close S3 service assurance and ownership gaps financials railiance-platform blocked codex 2026-09-05 2026-09-06 ca639c3d-3a87-5fa4-ad13-6f2e014b0c84

S3 service assurance and ownership gaps

Source: history/2026-09-05-platform-intent-workplan-assessment.md. Reviewed against current repository evidence. This plan supplies the missing continuing obligations; it does not reopen completed bootstrap projects or duplicate incident/lane work. Repository design and read-only implementation can progress now. Every live drill, scheduler, credential operation or migration retains its own owner and execution gate.

Record the portfolio assessment and consolidate source work

id: RPF-WP-0036-T01
status: done
priority: high
state_hub_task_id: "9c6d0439-8a14-5921-9d88-95a10b303057"

Completed 2026-09-05. Assessed all 37 existing plans and their 168 task records, corrected SCOPE, grouped the remaining obligations, consolidated the three design follow-ups under RPF-WP-0035, archived completed plans with identities preserved, and recorded owner handoffs and before/after inventory in history. This certifies the source review, not live service health or external acceptance.

Publish achievable service guarantees and recovery ownership

id: RPF-WP-0036-T02
status: done
priority: high
state_hub_task_id: "eebcd5c7-4cfd-5084-9a21-ca1e4e748bdd"

For apps-pg, platform-pg, OpenBao and each supported backup delivery lane, publish a versioned service record: accountable S3/package/operator owners, consumers, failure domain, availability objective, RPO/RTO, retention, recovery key/quorum availability, maintenance/abort path and evidence freshness budget. Separate measured results from accepted targets and unknowns. A 56-second scratch restore is not an RTO commitment; one replica on one host is not HA. Reuse docs/s3-consumer-interfaces.md and existing package declarations.

Done when: every supported service has owner-reviewed numeric targets or an explicit unsupported guarantee and decision owner; consumer requirements are compared to the current substrate; any HA/node-loss gap has an exact S1/S2 and package dependency rather than a blanket new-cluster project here.

Make backup freshness and recurring recovery evidence checkable

id: RPF-WP-0036-T03
status: wait
priority: high
state_hub_task_id: "d64446fb-870b-5ae9-9628-f1fb9d06c5a4"

Inventory authoritative CNPG backup/PITR, OpenBao snapshot/isolated restore, encrypted off-host copy and custody-recovery receipts. Reuse existing validators and package status commands. Define cadence/expiry from T02; return distinct healthy, stale, missing and unavailable states using metadata only. Schedule execution only through the accepted execution owner and separately approved authority. Keep RPF-WP-0015's pending database/reboot experiments as the sole live tasks for those experiments; RPF-WP-0029 retains provider-key recovery.

Done when: a current off-host backup and a current isolated restore receipt exist for each supported data service, the approved cadence is installed and its execution is evidenced, and missing/stale/failed evidence reaches a named operator. A template, dated successful snapshot, or same-PVC reboot does not pass as restore proof. Record independent recovery-key access without values.

Produce S3 signals and prove their delivery to the evidence owner

id: RPF-WP-0036-T04
status: wait
priority: high
state_hub_task_id: "5351e0e4-6263-58f0-afb2-7c78c4cd6f68"

Define service-owned health semantics for backup/WAL age, restore age, seal state, ESO freshness, connection/memory headroom and consumer ceiling. Reuse package emitters and the Q2 owner's standard contract; retain an explicit unmonitored state and named manual checker until transport is accepted. Request a concrete receiving contract from railiance-telemetry when routing is authorized; do not implement a competing monitoring plane in S3.

Done when: bounded metadata-only samples pass contract validation, a controlled stale/failure sample reaches a named recipient through the accepted Q2 route, and missing emission itself is detectable. Local fixture tests may finish before the receiver, but end-to-end acceptance cannot.

Reconcile admission, placement and consumer interface drift

id: RPF-WP-0036-T05
status: done
priority: high
state_hub_task_id: "5cded2e9-7edd-57bd-9f7e-7977b75004dc"

Join actual package declarations and authorized metadata to the S3 interface, tenancy and placement records. Correct stale platform-pg occupancy/co-residency (Core Hub admission versus older tenant-engine descriptions), distinguish desired placement from observed placement, and verify the named overflow targets remain provisionable. Add a bounded check for missing owners, unsupported retention requests, quota/ceiling drift and stale evidence; consume package admission checks instead of reimplementing them.

Done when: every admitted consumer has one authoritative placement/contract, capacity and retention disclosures match package source and dated live proof, and synthetic invalid admissions fail before provisioning. No workload moves under this task without its own owner-reviewed migration.

Obtain acceptance for compatibility assets and derived-record cleanup

id: RPF-WP-0036-T06
status: wait
priority: medium
state_hub_task_id: "3b74f79c-7366-5d00-bf0e-31c0464a8f33"

Prepare exact source/entry-point inventories and owner-ready handoffs for Forgejo backup/pruning/image inventory (railiance-forge, activity-core execution), retained OpenBao package wrappers (rapp-openbao), and ArgoCD bootstrap/application manifests (S2/S4/S5 according to artifact). Keep S3 custody contracts and the RPF-WP-0029 exposure obligation here until accepted closure. No new framework or app-specific helper belongs here by default.

Supply repo-manager/State Hub with the exact legacy alias/source identity map from the assessment. Their apparent duplicate active records and stale brief are derived-state defects, not additional workplans. Use scoped reconciliation; never change managed UUIDs or blanket-acknowledge retirements to clean a view.

Done when: each retained compatibility surface has an accepting owner, canonical replacement and tested callers or a dated retention decision; the repo-filtered projection and generated brief agree with source identities. Unaccepted transfer remains explicitly pending. No requests were sent during the assessment and this task does not assert acceptance for another repo.

Decide demand and reuse for undeployed stateful capabilities

id: RPF-WP-0036-T07
status: done
priority: medium
state_hub_task_id: "b32709d2-581d-5750-a04c-e3494a42d27e"

Review cache, general object storage and messaging separately with potential consumers. Inventory existing providers/contracts (including artifact-store and the external backup bucket) before selecting an engine. Record workload, durability/latency/retention needs, capacity, tenancy, custody, package owner, recovery cost and operating owner for any accepted demand. Ask railiance-master to resolve fleet Q3 ownership through its architecture process; do not assign it to S3 by implication.

Done when: each capability has a dated decision to reuse, defer with a review trigger, or start a bounded consumer-backed delivery plan with explicit acceptance criteria. “No accepted demand; keep deploy gated” is a valid result. No Valkey, MinIO, RabbitMQ or new provider purchase is authorized by this plan.

Implementation and remaining acceptance — 2026-09-05

T02 completed with assurance/service-records.json and ADR-0004: all three CNPG cells, OpenBao and both backup delivery surfaces disclose unsupported availability/RPO/RTO guarantees, existing evidence, retention and named decision owners. Numeric commitments have not been invented or approved for other owners.

T05 completed with the owner-native admission checker, hash-bound baseline and placement registry. It reuses local/package validators and rejects unowned consumers, fifth consumers, unhonoured retention and source disclosure drift. Corrected tenant-engine's completed PostgreSQL cutover and the deployed SBOM overflow cell. Live metadata matches one instance, 1Gi limit, 100 connections and 30-day retention on all three deployed cells. Dated owner evidence remains the authority for database/application cutover; this run moved no workloads.

T07 completed with ADR-0005: reuse approved backup storage only for backup; evaluate artifact-store's existing S3 interface for artifact demand; defer cache, general object store and broker deployment with concrete demand/review triggers. The fleet Q3 question is retained in T06's prepared master handoff; no external architecture assignment or provisioning occurred.

T03/T04 have working read-only assurance-capture, assurance-check and assurance-admission entry points plus adversarial tests. The collector pins the cluster UID and reads selected status, native aggregate connection counts, metrics and unauthenticated seal status. It captures no Secrets, SQL text from sessions, application rows or logs. Healthy/stale/missing/unavailable/failed are distinct; unknown fields, future timestamps and wrong clusters fail closed. The output explicitly says unmonitored and unsupported guarantees.

  • T03 waits: fresh owner-validated isolated restore/snapshot/offsite receipts, independent recovery access, approved cadence and execution evidence. Existing live experiments stay in RPF-WP-0015; offsite exposure stays in RPF-WP-0029. No restore, snapshot creation, upload, scheduler or seal/unseal was run.
  • T04 waits: accepted Q2 receiving contract and controlled failure/absence delivery to a named recipient. The local producer is not a ratified Q2 contract. Live metadata also surfaced failed ESO resources: forgejo-mailer, reuse-surface-runtime and target-revenue-runtime. Their current/obsolete scope and exact repair require consumer/platform acceptance before lane mutation; failures remain visible rather than excluded.
  • T06 waits: ownership acceptance and scoped alias/brief repair. The exact path/hash/caller inventory, dated retention decision through 2026-10-05 and prepared owner requests are in assurance/ownership-handoffs.json and docs/platform-ownership-handoffs.md. The installed brief generator still includes open legacy aliases; rerunning it alone cannot satisfy acceptance. No messages or external handoff acceptance were fabricated.

Evidence: docs/evidence/RPF-WP-0036-assurance-2026-09-05.json. The plan is blocked on these explicit live/owner gates, not finished merely because its repository implementation and tests pass.

2026-09-05 ESO recovery follow-up: RPF-WP-0037 resolved all three delivery failures using exact Kubernetes authentication. Custody values were unchanged, negative checks and repeated refresh passed, obsolete invalid delivery tokens were removed, and all 27 ExternalSecrets now report Ready. The earlier dated assurance snapshot is retained as historical evidence. T04 still waits on the accepted Q2 receiver and controlled failure/absence transport proof.

Blocker closure review — 2026-09-05

Fresh metadata capture verified all 15 collected signals healthy; seven recovery signals remain absent from the evaluator. RPF-WP-0037 is archived. The completed Backup account fixture proof remains separate from a full backup/restore receipt. T03/T04 remain waiting on recovery evidence/cadence and Q2 delivery respectively. T06's source inventory hashes were refreshed; its dated retention decision is unchanged. A repo-filtered Hub read confirms three retired aliases still appear open. The installed generator would reproduce them; exact UUID mapping and remaining owner requirements are persisted in history/2026-09-05-blocked-workplan-closure-review.md. No duplicate recovery, monitoring or owner-transfer workplan was created.

Primary backup coverage — 2026-09-06

Scaleway is the selected primary; Nextcloud is the independent secondary. Fresh apps-pg recovery from Scaleway passed in 42.64 seconds with expected consumer databases and limits, production Ready and scratch cleanup complete. Evidence: docs/evidence/scaleway-primary-restore-2026-09-06.json. T03 now has this fresh physical recovery receipt; recurring assurance cadence and the remaining recovery surfaces are still incomplete.

The source/live coverage inventory docs/backup-provider-coverage.md exposes remaining native primary configuration gaps on net-kingdom-pg/state-hub-db. Forgejo native database recovery and full Scaleway archive recovery have since passed; independent Nextcloud essentials recovery also passed (WP-0038). Track primary coverage here with forge/package/storage owners; do not silently claim the Nextcloud account cutover filled it or weaken WP-0029's separate incident closure.

Recovery evidence adapter follow-up — 2026-09-06

T03 advanced: capture now consumes hash-pinned apps-pg and forgejo-db native Scaleway restore receipts through scripts/recovery_evidence.py. Completion timestamps survive every capture; wrong destination, drift, missing/naive/future times and unsuccessful cleanup cannot become passing samples. Forgejo database restore is its own signal. Older application archive receipts lack completion timestamps and remain outside the automatic adapter; essentials evidence never substitutes for full recovery. Six previous recovery signals still lack adapters or accepted evidence. Cadence, independent custody and Q2 transport gates remain.

The Forgejo service record now reflects the completed native/full/essentials proofs while retaining unsupported guarantees and the pending scheduled cutover.

Live adapter verification: docs/evidence/RPF-WP-0036-assurance-2026-09-06.json reports 16 healthy, six missing and one stale ESO refresh signal. The aggregate crossed the one-hour diagnostic boundary; subsequent metadata inspection found all 27 ExternalSecrets Ready with newer refreshes. Cadence/grace acceptance is still needed; no outage or successful alert transport is inferred.

Archive receipt producer repair — 2026-09-06

T03 progressed further: primary/secondary transfer, primary decryption and isolated restore now emit their own start/finish timestamps, including failures. Decryption preserves its schema instead of inheriting the transfer's. Restore checks receipt schema and decryption proof, hashes the exact input receipt and marks failed cleanup as failed recovery. This closes the producer timestamp gap for future runs; old evidence remains unchanged. Native adapters remain the only automatic recovery adapters pending reviewed fresh archive receipts and adapter implementation. Validation and residual gates are recorded in history/2026-09-06-archive-receipt-review.md; T03 remains waiting on its wider cadence/custody/acceptance gates.

Closure review — 2026-09-06

User requested completion. Source review still finds T03/T04/T06 acceptance unmet; see history/2026-09-06-WP-0036-closure-gates.md. Telemetry's checkout has no receiver implementation, and the older platform-pg scratch MinIO drill does not establish production off-host recovery. Added the verified August 22 OpenBao snapshot to the hash-pinned evidence adapter: it reports stale using its original creation time, separately from the still-missing isolated restore. No criteria were weakened, external acceptance inferred or live window reused.

Q2 owner implementation — 2026-09-06

The user directed implementation in railiance-telemetry. RTEL-WP-0002 now owns its metadata contract, durable private reference receiver, platform report adapter and missing-emission semantics. Nine tests and a separate-process local CLI failure/absence proof pass. See telemetry docs/signal-contract.md and history/2026-09-06-local-receiver-proof.json. T04's next dependency is explicitly RTEL-WP-0002-T04: accepted package/runtime, authenticated execution, actual operator recipient, cadence/storage/retention and independent receiver watchdog, then acknowledged controlled delivery. Local test-inbox visibility is not that live receipt; S3 remains unmonitored until acceptance.