Assistant: codex Assistant-Model: gpt-5.6-sol Assistant-Session: 01a02669-87ee-7a31-b111-edc95a16e0fa
4.7 KiB
Audit-core restart-free database lease recovery
This is the review and attended-execution contract for
RAILIANCE-WP-0024-T02. It proves that revoking the lease currently used by
audit-core causes retryable unavailability and that External Secrets plus the
mounted credential reader restore service without replacing or restarting the
pod. It does not authorize a live revocation.
Fixed scope and owners
| Boundary | Exact scope | Responsible owner |
|---|---|---|
| Consumer | Deployment audit-core/audit-core; Secret audit-core/audit-core-database; ExternalSecret audit-core/audit-core-database |
audit-core |
| Lease | database/creds/audit-core-runtime; only the lease matching the current mounted username |
railiance-platform custody operator |
| Database | CNPG databases/platform-pg, database audit_core, runtime group audit_core_app |
rapp-postgres |
| Load | One separately approved synthetic sender/tenant contract; no production event or identity | audit-core + load owner |
| Coordination and evidence | RAILIANCE-WP-0024-T02 |
railiance-platform |
The live window requires audit-core and rapp-postgres acknowledgement, a named abort operator, and an approved synthetic-load identifier. Run the read-only gate before requesting that approval:
python3 scripts/audit-core-recovery-preflight.py database-lease
The result must have automated_checks_passed: true. It will remain
ready_for_live_execution: false until the explicit owner/window fields are
provided. The helper reads Secret metadata and key names only; it never reads
Secret data.
Attended sequence
Keep one terminal attached from baseline through recovery. A value-handling
helper used in the approved window must suppress OpenBao/Kubernetes response
bodies and clear all transient JSON from memory or mode-0600 temporary files.
- Record the audit-core pod UID, restart count, database Secret
resourceVersion,/healthzstatus,/readyzstatus, CNPG readiness, ExternalSecret condition, and ClusterSecretStore condition. - Start bounded retrying synthetic traffic. Its evidence records only request identifier, attempt time, response class, and final accepted/duplicate result—never the bearer or payload.
- Inside the attended process, read the current mounted username without
emitting it. List only lease handles below
sys/leases/lookup/database/creds/audit-core-runtime, look each up, and retain the one handle whose lease username matches. Abort unless exactly one live handle matches. Clear every lookup response and username immediately. - Recheck that the Secret
resourceVersion, pod UID, and selected handle have not changed. Revoke only that handle. Do not revoke the role prefix and do not revoke the ESO parent token. - Prove the bounded failure state:
/healthzremains 200,/readyzbecomes 503, at least one synthetic attempt receives retryable 503, pod UID is unchanged, and restart count is unchanged. Abort if liveness fails, the pod restarts, a non-synthetic request is involved, or 503 is not retryable. - Annotate only ExternalSecret
audit-core/audit-core-databaseto force reconciliation. Wait forReady=True/SecretSynced, a changed SecretresourceVersion, and Kubernetes' mounted..datageneration to advance. Do not restart or roll out audit-core. - Prove recovery:
/readyzreturns 200, the retrying synthetic request reaches accepted or duplicate, the pre-exercise event remains readable, pod UID is unchanged, and restart count is unchanged. - Stop the load, remove its separately approved fixture/identity, and confirm the predecessor lease is revoked. The replacement remains managed by ESO and expires through its normal database role TTL.
Time bounds and abort path
- Planned outage observation: at most 120 seconds.
- Full exercise: at most 10 minutes.
- If
/readyzhas not recovered 120 seconds after forced reconciliation, stop load and page both custody and database owners. Do not bounce the pod. - If ESO is not
SecretSynced, check OpenBao seal state and the exact store. Restart External Secrets only for the known post-unseal stale-provider case and only with the platform owner present. - If PostgreSQL is not 1/1 Ready or continuous archiving is false, abort without revocation.
Value-safe evidence
Record timestamps, approved window/load identifiers, owner acknowledgements, lease accessor or one-way handle fingerprint, Secret resource versions, mount generation changes, HTTP status sequence, pod UID/restart counts, CNPG/ESO conditions, and cleanup outcome. Never record usernames, passwords, DSNs, Secret data, bearer values, request payloads, OpenBao response bodies, or unseal material.