audit-core/workplans/AUDIT-WP-0005-postgres-store-and-production-deployment.md
tegwick bd274f6269
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 1s
Add the PostgreSQL audit backend and a shared conformance suite
AUDIT-WP-0005-T01, built and verified against PostgreSQL 16 locally in
Docker; the Railiance cluster was not needed.

tests/test_backend_conformance.py is one suite run against every backend, so
"the Postgres backend is done" means it satisfies the same contract SQLite
already does rather than having its own green tests. It skips cleanly with no
server reachable; make pg-test-up and make test-pg run it. Suite 50 -> 71.

RetentionPolicy declares immutable=True and earns it: migration 0002 installs
a trigger rejecting UPDATE and DELETE on the events table, so a leaked runtime
credential can append but cannot rewrite or erase the trail. That materially
narrows the residual risk ADR-0001 section 5 called out. tamper_evidence stays
False because nothing here would prove a database owner had dropped the
trigger - hash-chaining or external anchoring would be needed and is not
implemented.

Idempotency is one statement (INSERT ... ON CONFLICT DO NOTHING RETURNING),
verified to behave identically to the SQLite backend under 12 concurrent
submissions of the same event. Migrations are ordered, recorded and
idempotent. Replay reconciles rather than duplicating - the piece deferred out
of WP-0004-T05 - and is tested to leave exactly one custody record.

Backend selection is by AUDIT_CORE_DATABASE_URL; the SQLite fallback logs a
warning so a deployment that lost its URL is visible rather than quietly
running on the wrong store.

Also fixed: ingestion had no __main__ guard, so python -m audit_core.ingestion
silently did nothing. Found during end-to-end smoke.

Counting semantics documented: occurrences counts transmissions, not stored
events, so a retry of a secret-shaped field increments it again. That is the
sender behaviour being optimized away.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 17:09:46 +02:00

8.8 KiB

id type title domain repo status owner topic_slug created updated depends_on state_hub_workstream_id
AUDIT-WP-0005 workplan Deploy audit-core on Railiance with durable Postgres custody infotech audit-core proposed codex netkingdom 2026-08-10 2026-08-10
AUDIT-WP-0004
RAPP-POSTGRES-WP-0002
NK-WP-0024
7b24a844-c9f2-4d2d-ac7d-20978bdf6b38

AUDIT-WP-0005 - Postgres store and production deployment

Goal

Put audit-core into production on Railiance with PostgreSQL as the custody store, and prove the delivery path end to end.

This is first on the critical path: email-connect delivery, OpenBao sender credentials, the user-engine runtime integrations, and the live failure matrix all sit downstream of a receiver that actually holds evidence.

Dependencies

  • AUDIT-WP-0004 — the receiver must route through the backend contract and return deliberate status codes before a Postgres backend is worth writing behind it, and before a failure matrix can distinguish a correct refusal from a crash.
  • RAPP-POSTGRES-WP-0002 — the database, the tenancy model, and the credential lane. T01 below can be drafted against the model as soon as RAPP-POSTGRES-WP-0002-T01 settles it; T02 needs a provisioned database.

Boundaries

This workplan owns the Postgres backend implementation, audit-core's deployment, and the live verification. It does not own the database platform, its tenancy model, or its backup machinery — those belong to rapp-postgres. Where audit-core needs a guarantee from the platform, it states the requirement and consumes it; it does not implement it here.

T01 - Implement the Postgres audit backend

id: AUDIT-WP-0005-T01
status: done
priority: high
state_hub_task_id: "b1601d0b-922a-40f7-92c0-ea06af6c4468"

Implement AuditBackend against PostgreSQL, declaring an honest RetentionPolicy — the custody class, retention window, and whether immutability and tamper evidence are genuinely provided rather than aspired to. If the schema does not prevent an operator from silently editing a recorded event, the policy must not claim immutable.

Push idempotency into the database rather than a read-then-write in application code: an insert conflicting on event ID must resolve atomically into duplicate-accepted or conflict, with no window where two concurrent identical events both write. Store the payload hash for conflict detection.

Handle connection lifecycle properly — pooling, reconnection after a database restart, and a bounded statement timeout so a stalled write surfaces as unavailable instead of hanging the request.

Own the migrations. The consuming service owns its schema; rapp-postgres owns the space it runs in.

Done when the backend passes the same contract tests as the existing backends, concurrent duplicate submissions produce exactly one record, and a database restart mid-write does not produce an acknowledged-but-absent event.

Done 2026-08-10, built and verified against PostgreSQL 16 locally in Docker — the Railiance cluster was not needed for any of it.

tests/test_backend_conformance.py is a single suite run against every backend, so "the Postgres backend is done" means it satisfies the same contract SQLite already does rather than having its own tests that happen to be green. It skips cleanly when no server is reachable; make pg-test-up and make test-pg run it. Suite 50 -> 71.

RetentionPolicy declares immutable=True, and that is earned: migration 0002 installs a trigger rejecting UPDATE and DELETE on the events table, so a leaked runtime credential can append but cannot rewrite or erase the trail. This materially narrows the residual risk ADR-0001 §5 called out — a leaked credential could previously forge the audit record. tamper_evidence stays False, because nothing here would prove a database owner had dropped the trigger; hash-chaining or external anchoring would be needed and is not implemented.

Migrations are ordered, recorded in schema_migrations, and idempotent. Replay reconciles rather than duplicating — the piece deferred out of WP-0004-T05 — and is tested to leave exactly one custody record.

Backend selection is by AUDIT_CORE_DATABASE_URL; falling back to SQLite logs a warning, so a deployment that lost its URL is visible rather than quietly running on the wrong store. End-to-end smoke through waitress against Postgres confirmed accept, duplicate, cross-tenant refusal, read/write privilege separation, visible redaction, secret-finding counters, and readiness reporting custody_class=archive.

Also fixed: audit_core.ingestion had no __main__ guard, so python -m audit_core.ingestion silently did nothing.

Not covered here: behaviour across a real database failover, which needs the cluster and belongs to T05.

T02 - Provision storage through the platform lane

id: AUDIT-WP-0005-T02
status: todo
priority: high
state_hub_task_id: "831b2472-0d80-4369-a5e3-eb08ef3526b1"

Declare audit-core's database requirement against rapp-postgres and take delivery of a least-privilege runtime role through the OpenBao lane. The runtime role connects; a separate role runs migrations; neither owns more than it needs.

Verify audit-core's side of the isolation model: the runtime credential reaches audit-core's data and nothing else, and rotation completes without a delivery gap.

Done when audit-core runs against a provisioned database using a credential it never received as a literal, and rotating that credential does not drop events.

T03 - Deploy the receiver

id: AUDIT-WP-0005-T03
status: todo
priority: high
state_hub_task_id: "598af2ac-e772-4a4e-9a65-dde9d4ca167f"

Publish an immutable image — base pinned by digest, not a mutable tag, and labelled with the build commit — and deploy on railiance01 with a Service, health and readiness probes wired to the checks from WP-0004-T01, resource requests and limits, a restricted security context, and a default-deny NetworkPolicy admitting only user-engine as sender and the operator path for reads.

State the rollback position: which image digest and which schema version the deployment can return to, and whether the migration in T01 is reversible. A rollback plan that assumes reversible migrations without checking is not a plan.

The container currently runs as uid 10001 and expects a writable /data; once custody is in Postgres that path should carry no durable state at all. Confirm nothing of value is left on the pod filesystem.

Done when the receiver is reachable only by its declared peers, survives pod restart and rescheduling without loss, and has a tested path back to the previous version.

T04 - Migrate existing SQLite records

id: AUDIT-WP-0005-T04
status: todo
priority: medium
state_hub_task_id: "9010fb4a-a1b8-4ef7-b143-e33ca7cc0619"

Any events accepted by the pre-production SQLite receiver are audit records and cannot simply be dropped. Move them into the Postgres store with their original identifiers, timestamps, and payload hashes intact, or record an explicit decision that they are development artifacts with no custody value.

Whichever holds, it must be written down — silently discarding accepted audit events is the exact failure this service is meant to make impossible.

Done when the disposition of every pre-production record is either migrated and verified, or explicitly and justifiably discarded.

T05 - Run the live failure matrix

id: AUDIT-WP-0005-T05
status: todo
priority: high
state_hub_task_id: "1da30fec-b9f1-4be0-b42c-15797a8c4392"

Exercise the deployed path: successful delivery; receiver timeout and unavailability; bounded user-engine retry against each documented status code; dead-letter visibility; operator replay and duplicate replay; redaction; and correlation lookup. Include a database failover or restart during active ingestion, and a credential rotation during active ingestion.

The controlling assertion is that one source outbox event produces exactly one durable normalized event across retries, replay, and infrastructure disruption — no loss, no duplication.

Hand non-secret evidence back to NK-WP-0024.

Done when the matrix has been run against the deployed system and each outcome recorded, including any case where behaviour differed from the documented contract.

T06 - Operational handover

id: AUDIT-WP-0005-T06
status: todo
priority: medium
state_hub_task_id: "0856c80d-abe1-4bff-ba8d-87295cf76819"

Document what an operator needs: how to look up an event by correlation ID, how to inspect and replay dead-lettered events, how to rotate the sender credential, what the alert conditions mean, and how to restore audit data from a rapp-postgres backup.

Verify audit-core's recovery requirement against what rapp-postgres actually provides — the retention window audit-core declares must not exceed the retention the platform guarantees.

Done when the runbook exists and the restore path has been walked once.