audit-core/workplans/AUDIT-WP-0005-postgres-store-and-production-deployment.md

316 lines
14 KiB
Markdown
Raw Normal View History

---
id: AUDIT-WP-0005
type: workplan
title: "Deploy audit-core on Railiance with durable Postgres custody"
domain: infotech
repo: audit-core
status: proposed
owner: codex
topic_slug: netkingdom
created: "2026-08-10"
updated: "2026-08-10"
depends_on:
- AUDIT-WP-0004
- RAPP-POSTGRES-WP-0002
- NK-WP-0024
Route ingestion through the backend contract; fix error semantics AUDIT-WP-0004 T01, T02, T07. T01 - ingestion wrote to SQLite directly and never called the AuditBackend contract, so a 202 meant a row existed rather than that a backend with a declared retention policy had accepted the event. Adds IdempotentAuditBackend to the contract: duplicate detection lives inside the backend so custody and idempotency state share a transaction and cannot diverge. SQLiteAuditBackend implements it with WAL, synchronous=FULL and a busy timeout. Ingestion now refuses any backend declaring durable=False, so the development file backend cannot silently become the production sink. The atomicity claim was tested rather than asserted, and the first attempt failed: with a single shared connection, 16 racing submissions of one event told two callers they were first. Storage was correct but the response was not. Fixed with per-thread connections and BEGIN IMMEDIATE around the insert/read pair, and locked in by a test. T02 - storage errors previously escaped the handler with start_response never called, and the auth check sat outside the try block so a non-ASCII Authorization header crashed the request. Adds a catch-all, maps conflict to 409, backend unavailability to 503 and unexpected faults to 500, and documents the full response contract with the retry semantics each status implies, since senders key their behaviour off it. T07 - ingestion tests 2 -> 23, suite 15 -> 36. accepted_at is now UTC rather than local time, and naive timestamps are rejected instead of silently assumed. Remaining in WP-0004: T03 tenant/source binding, T04 redaction policy, T05 operator read surface, T06 production serving layer. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 14:30:20 +02:00
state_hub_workstream_id: "7b24a844-c9f2-4d2d-ac7d-20978bdf6b38"
---
# AUDIT-WP-0005 - Postgres store and production deployment
## Goal
Put audit-core into production on Railiance with PostgreSQL as the custody
store, and prove the delivery path end to end.
This is first on the critical path: email-connect delivery, OpenBao sender
credentials, the user-engine runtime integrations, and the live failure matrix
all sit downstream of a receiver that actually holds evidence.
## Dependencies
- **AUDIT-WP-0004** — the receiver must route through the backend contract and
return deliberate status codes before a Postgres backend is worth writing
behind it, and before a failure matrix can distinguish a correct refusal
from a crash.
- **RAPP-POSTGRES-WP-0002** — the database, the tenancy model, and the
credential lane. T01 below can be drafted against the model as soon as
RAPP-POSTGRES-WP-0002-T01 settles it; T02 needs a provisioned database.
## Boundaries
This workplan owns the Postgres backend implementation, audit-core's
deployment, and the live verification. It does not own the database platform,
its tenancy model, or its backup machinery — those belong to rapp-postgres.
Where audit-core needs a guarantee from the platform, it states the
requirement and consumes it; it does not implement it here.
## T01 - Implement the Postgres audit backend
```task
id: AUDIT-WP-0005-T01
Add the PostgreSQL audit backend and a shared conformance suite AUDIT-WP-0005-T01, built and verified against PostgreSQL 16 locally in Docker; the Railiance cluster was not needed. tests/test_backend_conformance.py is one suite run against every backend, so "the Postgres backend is done" means it satisfies the same contract SQLite already does rather than having its own green tests. It skips cleanly with no server reachable; make pg-test-up and make test-pg run it. Suite 50 -> 71. RetentionPolicy declares immutable=True and earns it: migration 0002 installs a trigger rejecting UPDATE and DELETE on the events table, so a leaked runtime credential can append but cannot rewrite or erase the trail. That materially narrows the residual risk ADR-0001 section 5 called out. tamper_evidence stays False because nothing here would prove a database owner had dropped the trigger - hash-chaining or external anchoring would be needed and is not implemented. Idempotency is one statement (INSERT ... ON CONFLICT DO NOTHING RETURNING), verified to behave identically to the SQLite backend under 12 concurrent submissions of the same event. Migrations are ordered, recorded and idempotent. Replay reconciles rather than duplicating - the piece deferred out of WP-0004-T05 - and is tested to leave exactly one custody record. Backend selection is by AUDIT_CORE_DATABASE_URL; the SQLite fallback logs a warning so a deployment that lost its URL is visible rather than quietly running on the wrong store. Also fixed: ingestion had no __main__ guard, so python -m audit_core.ingestion silently did nothing. Found during end-to-end smoke. Counting semantics documented: occurrences counts transmissions, not stored events, so a retry of a secret-shaped field increments it again. That is the sender behaviour being optimized away. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 17:09:46 +02:00
status: done
priority: high
Route ingestion through the backend contract; fix error semantics AUDIT-WP-0004 T01, T02, T07. T01 - ingestion wrote to SQLite directly and never called the AuditBackend contract, so a 202 meant a row existed rather than that a backend with a declared retention policy had accepted the event. Adds IdempotentAuditBackend to the contract: duplicate detection lives inside the backend so custody and idempotency state share a transaction and cannot diverge. SQLiteAuditBackend implements it with WAL, synchronous=FULL and a busy timeout. Ingestion now refuses any backend declaring durable=False, so the development file backend cannot silently become the production sink. The atomicity claim was tested rather than asserted, and the first attempt failed: with a single shared connection, 16 racing submissions of one event told two callers they were first. Storage was correct but the response was not. Fixed with per-thread connections and BEGIN IMMEDIATE around the insert/read pair, and locked in by a test. T02 - storage errors previously escaped the handler with start_response never called, and the auth check sat outside the try block so a non-ASCII Authorization header crashed the request. Adds a catch-all, maps conflict to 409, backend unavailability to 503 and unexpected faults to 500, and documents the full response contract with the retry semantics each status implies, since senders key their behaviour off it. T07 - ingestion tests 2 -> 23, suite 15 -> 36. accepted_at is now UTC rather than local time, and naive timestamps are rejected instead of silently assumed. Remaining in WP-0004: T03 tenant/source binding, T04 redaction policy, T05 operator read surface, T06 production serving layer. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 14:30:20 +02:00
state_hub_task_id: "b1601d0b-922a-40f7-92c0-ea06af6c4468"
```
Implement `AuditBackend` against PostgreSQL, declaring an honest
`RetentionPolicy` — the custody class, retention window, and whether
immutability and tamper evidence are genuinely provided rather than aspired
to. If the schema does not prevent an operator from silently editing a
recorded event, the policy must not claim `immutable`.
Push idempotency into the database rather than a read-then-write in
application code: an insert conflicting on event ID must resolve atomically
into duplicate-accepted or conflict, with no window where two concurrent
identical events both write. Store the payload hash for conflict detection.
Handle connection lifecycle properly — pooling, reconnection after a database
restart, and a bounded statement timeout so a stalled write surfaces as
unavailable instead of hanging the request.
Own the migrations. The consuming service owns its schema; rapp-postgres owns
the space it runs in.
Done when the backend passes the same contract tests as the existing backends,
concurrent duplicate submissions produce exactly one record, and a database
restart mid-write does not produce an acknowledged-but-absent event.
Add the PostgreSQL audit backend and a shared conformance suite AUDIT-WP-0005-T01, built and verified against PostgreSQL 16 locally in Docker; the Railiance cluster was not needed. tests/test_backend_conformance.py is one suite run against every backend, so "the Postgres backend is done" means it satisfies the same contract SQLite already does rather than having its own green tests. It skips cleanly with no server reachable; make pg-test-up and make test-pg run it. Suite 50 -> 71. RetentionPolicy declares immutable=True and earns it: migration 0002 installs a trigger rejecting UPDATE and DELETE on the events table, so a leaked runtime credential can append but cannot rewrite or erase the trail. That materially narrows the residual risk ADR-0001 section 5 called out. tamper_evidence stays False because nothing here would prove a database owner had dropped the trigger - hash-chaining or external anchoring would be needed and is not implemented. Idempotency is one statement (INSERT ... ON CONFLICT DO NOTHING RETURNING), verified to behave identically to the SQLite backend under 12 concurrent submissions of the same event. Migrations are ordered, recorded and idempotent. Replay reconciles rather than duplicating - the piece deferred out of WP-0004-T05 - and is tested to leave exactly one custody record. Backend selection is by AUDIT_CORE_DATABASE_URL; the SQLite fallback logs a warning so a deployment that lost its URL is visible rather than quietly running on the wrong store. Also fixed: ingestion had no __main__ guard, so python -m audit_core.ingestion silently did nothing. Found during end-to-end smoke. Counting semantics documented: occurrences counts transmissions, not stored events, so a retry of a secret-shaped field increments it again. That is the sender behaviour being optimized away. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 17:09:46 +02:00
Done 2026-08-10, built and verified against PostgreSQL 16 locally in Docker —
the Railiance cluster was not needed for any of it.
`tests/test_backend_conformance.py` is a single suite run against every
backend, so "the Postgres backend is done" means it satisfies the same
contract SQLite already does rather than having its own tests that happen to
be green. It skips cleanly when no server is reachable; `make pg-test-up` and
`make test-pg` run it. Suite 50 -> 71.
`RetentionPolicy` declares `immutable=True`, and that is earned: migration
0002 installs a trigger rejecting UPDATE and DELETE on the events table, so a
leaked runtime credential can append but cannot rewrite or erase the trail.
This materially narrows the residual risk ADR-0001 §5 called out — a leaked
credential could previously forge the audit record. `tamper_evidence` stays
False, because nothing here would *prove* a database owner had dropped the
trigger; hash-chaining or external anchoring would be needed and is not
implemented.
Migrations are ordered, recorded in `schema_migrations`, and idempotent.
Replay reconciles rather than duplicating — the piece deferred out of
WP-0004-T05 — and is tested to leave exactly one custody record.
Backend selection is by `AUDIT_CORE_DATABASE_URL`; falling back to SQLite logs
a warning, so a deployment that lost its URL is visible rather than quietly
running on the wrong store. End-to-end smoke through waitress against Postgres
confirmed accept, duplicate, cross-tenant refusal, read/write privilege
separation, visible redaction, secret-finding counters, and readiness
reporting `custody_class=archive`.
Also fixed: `audit_core.ingestion` had no `__main__` guard, so `python -m
audit_core.ingestion` silently did nothing.
Not covered here: behaviour across a real database failover, which needs the
cluster and belongs to T05.
## T02 - Provision storage through the platform lane
```task
id: AUDIT-WP-0005-T02
status: todo
priority: high
Route ingestion through the backend contract; fix error semantics AUDIT-WP-0004 T01, T02, T07. T01 - ingestion wrote to SQLite directly and never called the AuditBackend contract, so a 202 meant a row existed rather than that a backend with a declared retention policy had accepted the event. Adds IdempotentAuditBackend to the contract: duplicate detection lives inside the backend so custody and idempotency state share a transaction and cannot diverge. SQLiteAuditBackend implements it with WAL, synchronous=FULL and a busy timeout. Ingestion now refuses any backend declaring durable=False, so the development file backend cannot silently become the production sink. The atomicity claim was tested rather than asserted, and the first attempt failed: with a single shared connection, 16 racing submissions of one event told two callers they were first. Storage was correct but the response was not. Fixed with per-thread connections and BEGIN IMMEDIATE around the insert/read pair, and locked in by a test. T02 - storage errors previously escaped the handler with start_response never called, and the auth check sat outside the try block so a non-ASCII Authorization header crashed the request. Adds a catch-all, maps conflict to 409, backend unavailability to 503 and unexpected faults to 500, and documents the full response contract with the retry semantics each status implies, since senders key their behaviour off it. T07 - ingestion tests 2 -> 23, suite 15 -> 36. accepted_at is now UTC rather than local time, and naive timestamps are rejected instead of silently assumed. Remaining in WP-0004: T03 tenant/source binding, T04 redaction policy, T05 operator read surface, T06 production serving layer. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 14:30:20 +02:00
state_hub_task_id: "831b2472-0d80-4369-a5e3-eb08ef3526b1"
```
Declare audit-core's database requirement against rapp-postgres and take
delivery of a least-privilege runtime role through the OpenBao lane. The
runtime role connects; a separate role runs migrations; neither owns more than
it needs.
Verify audit-core's side of the isolation model: the runtime credential
reaches audit-core's data and nothing else, and rotation completes without a
delivery gap.
Done when audit-core runs against a provisioned database using a credential it
never received as a literal, and rotating that credential does not drop
events.
## T03 - Deploy the receiver
```task
id: AUDIT-WP-0005-T03
Add deployment manifests, custody-class guard and request counters AUDIT-WP-0005-T03 (progress). Manifests validated --dry-run=server --validate=strict against railiance01; not applied, since deployment is gated on RAPP-POSTGRES-WP-0002 and T02 credentials. Nothing here mutates the cluster. Conventions read off the deployed user-engine workload rather than invented: digest-pinned image from forgejo.coulomb.social, runAsNonRoot with RuntimeDefault seccomp, no privilege escalation, all capabilities dropped, readOnlyRootFilesystem, probes on a named http port, same resource envelope. The namespace carries railiance.io/postgres-client: platform-pg, which is what platform-pg-consumer-ingress in rapp-postgres admits; without that label the pod cannot reach the database at all. NetworkPolicies default-deny both directions, then permit ingress from the user-engine namespace only, a separately labelled operator read path, and egress to PostgreSQL in databases plus DNS. Three decisions worth naming. Liveness is /healthz while readiness is /readyz, so a database outage drops the pod from the Service rather than restarting it in a loop. readOnlyRootFilesystem enforces the empty-filesystem property rather than trusting it, so the SQLite fallback physically cannot accumulate audit records on ephemeral storage. AUDIT_CORE_REQUIRE_CUSTODY_CLASS=archive makes a missing database URL a startup failure instead of a silent downgrade to the development store. Counters deferred from WP-0004-T06 are exposed as JSON at /v1/stats behind the read privilege, not as Prometheus exposition format: the cluster runs no Prometheus, no ServiceMonitor CRD and no other scrape target, so an exposition endpoint would target a scrape path that does not exist. Usable with curl now and a small step from /metrics later. Tests 77 -> 80. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 17:42:43 +02:00
status: progress
priority: high
Route ingestion through the backend contract; fix error semantics AUDIT-WP-0004 T01, T02, T07. T01 - ingestion wrote to SQLite directly and never called the AuditBackend contract, so a 202 meant a row existed rather than that a backend with a declared retention policy had accepted the event. Adds IdempotentAuditBackend to the contract: duplicate detection lives inside the backend so custody and idempotency state share a transaction and cannot diverge. SQLiteAuditBackend implements it with WAL, synchronous=FULL and a busy timeout. Ingestion now refuses any backend declaring durable=False, so the development file backend cannot silently become the production sink. The atomicity claim was tested rather than asserted, and the first attempt failed: with a single shared connection, 16 racing submissions of one event told two callers they were first. Storage was correct but the response was not. Fixed with per-thread connections and BEGIN IMMEDIATE around the insert/read pair, and locked in by a test. T02 - storage errors previously escaped the handler with start_response never called, and the auth check sat outside the try block so a non-ASCII Authorization header crashed the request. Adds a catch-all, maps conflict to 409, backend unavailability to 503 and unexpected faults to 500, and documents the full response contract with the retry semantics each status implies, since senders key their behaviour off it. T07 - ingestion tests 2 -> 23, suite 15 -> 36. accepted_at is now UTC rather than local time, and naive timestamps are rejected instead of silently assumed. Remaining in WP-0004: T03 tenant/source binding, T04 redaction policy, T05 operator read surface, T06 production serving layer. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 14:30:20 +02:00
state_hub_task_id: "598af2ac-e772-4a4e-9a65-dde9d4ca167f"
```
Publish an immutable image — base pinned by digest, not a mutable tag, and
labelled with the build commit — and deploy on railiance01 with a Service,
health and readiness probes wired to the checks from WP-0004-T01, resource
requests and limits, a restricted security context, and a default-deny
NetworkPolicy admitting only user-engine as sender and the operator path for
reads.
State the rollback position: which image digest and which schema version the
deployment can return to, and whether the migration in T01 is reversible. A
rollback plan that assumes reversible migrations without checking is not a
plan.
The container currently runs as uid 10001 and expects a writable `/data`; once
custody is in Postgres that path should carry no durable state at all. Confirm
nothing of value is left on the pod filesystem.
Done when the receiver is reachable only by its declared peers, survives pod
restart and rescheduling without loss, and has a tested path back to the
previous version.
Add deployment manifests, custody-class guard and request counters AUDIT-WP-0005-T03 (progress). Manifests validated --dry-run=server --validate=strict against railiance01; not applied, since deployment is gated on RAPP-POSTGRES-WP-0002 and T02 credentials. Nothing here mutates the cluster. Conventions read off the deployed user-engine workload rather than invented: digest-pinned image from forgejo.coulomb.social, runAsNonRoot with RuntimeDefault seccomp, no privilege escalation, all capabilities dropped, readOnlyRootFilesystem, probes on a named http port, same resource envelope. The namespace carries railiance.io/postgres-client: platform-pg, which is what platform-pg-consumer-ingress in rapp-postgres admits; without that label the pod cannot reach the database at all. NetworkPolicies default-deny both directions, then permit ingress from the user-engine namespace only, a separately labelled operator read path, and egress to PostgreSQL in databases plus DNS. Three decisions worth naming. Liveness is /healthz while readiness is /readyz, so a database outage drops the pod from the Service rather than restarting it in a loop. readOnlyRootFilesystem enforces the empty-filesystem property rather than trusting it, so the SQLite fallback physically cannot accumulate audit records on ephemeral storage. AUDIT_CORE_REQUIRE_CUSTODY_CLASS=archive makes a missing database URL a startup failure instead of a silent downgrade to the development store. Counters deferred from WP-0004-T06 are exposed as JSON at /v1/stats behind the read privilege, not as Prometheus exposition format: the cluster runs no Prometheus, no ServiceMonitor CRD and no other scrape target, so an exposition endpoint would target a scrape path that does not exist. Usable with curl now and a small step from /metrics later. Tests 77 -> 80. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 17:42:43 +02:00
Progress 2026-08-10: manifests written and validated `--dry-run=server
--validate=strict` against railiance01. Not applied — deployment is gated on
RAPP-POSTGRES-WP-0002 and on T02's credentials, and nothing here mutates the
cluster.
Conventions were read off the deployed `user-engine` workload rather than
invented: digest-pinned image from `forgejo.coulomb.social/coulomb/<name>`,
`runAsNonRoot` + `RuntimeDefault` seccomp, `allowPrivilegeEscalation: false`,
all capabilities dropped, `readOnlyRootFilesystem: true`, probes on a named
`http` port, and the same resource envelope.
`deploy/audit-core.yaml` — namespace, Service, Deployment. The namespace
carries `railiance.io/postgres-client: platform-pg`, which is what
`platform-pg-consumer-ingress` in rapp-postgres admits; without that label the
pod cannot reach the database at all.
`deploy/networkpolicies.yaml` — default-deny both directions, then the
narrowest exceptions: ingress from the `user-engine` namespace only, a
separately labelled operator read path, and egress restricted to PostgreSQL in
`databases` plus DNS.
Three decisions worth recording:
- **Liveness is `/healthz`, readiness is `/readyz`.** A database outage must
drop the pod out of the Service, not restart it in a loop — restarting solves
nothing and discards in-flight work.
- **`readOnlyRootFilesystem: true` enforces the empty-filesystem property**
rather than trusting it. Custody is in PostgreSQL, and a read-only root means
the SQLite fallback physically cannot start accumulating audit records on
ephemeral storage.
- **`AUDIT_CORE_REQUIRE_CUSTODY_CLASS=archive`** makes a missing database URL a
startup failure instead of a silent downgrade to the development store. Added
in this task; covered by a test.
Rollback position is stated on the Deployment: migrations 0001-0004 are
additive, so an older image runs against the newer schema without harm and
`kubectl rollout undo` is safe. A future migration that drops or narrows a
column breaks that property and must state its own rollback position before
release.
Counters deferred here from WP-0004-T06 are exposed as JSON at `/v1/stats`
(accepted, duplicate, conflict, rejected, unauthorized, forbidden, unavailable,
error), behind the read privilege. Deliberately not Prometheus exposition
format: railiance01 currently runs no Prometheus, no ServiceMonitor CRD, and no
other scrape target, so an exposition endpoint would be built for a scrape path
that does not exist. This is usable with curl today and a small step from
`/metrics` when a metrics stack lands.
Remaining before done: build and publish the image to get a real digest, apply,
and verify restart/rescheduling and the rollback path on the cluster.
## T04 - Migrate existing SQLite records
```task
id: AUDIT-WP-0005-T04
Establish pre-production record disposition and add the migration tool AUDIT-WP-0005-T04. Disposition: there are no pre-production records. No audit-core SQLite store on this host, no mock-file-backend output, and no audit-core pod, deployment or PVC on railiance01 - the only audit-* PVC there is OpenBao's own audit device. Consistent with the history: WP-0003-T03 was cancelled before the receiver was ever deployed, so every SQLite store that has existed was a test fixture. Nothing is being discarded because nothing was ever accepted outside tests. The tool is built anyway because the SQLite path stays reachable - the entrypoint falls back to it when AUDIT_CORE_DATABASE_URL is unset. If that fallback is ever used in anger the records are audit records, and writing the migration afterwards under pressure is the wrong time. audit_core.migrate_store and `python -m audit_core migrate-store` transfer events, dead letters and secret-finding counters. Records keep their original event_id, payload_hash and accepted_at, which is why this bypasses accept(): that stamps acceptance with the current time, and a migration that rewrote acceptance times would destroy the evidence it exists to preserve. Idempotent, and verification reads back from the destination rather than trusting the write path. A destination record with a differing payload hash is reported as a conflict and left untouched - silently overwriting a stored audit record is the same class of failure as losing it. Conflicts and failed verification exit non-zero; a partial migration is not a success. Tests 71 -> 77. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 17:31:18 +02:00
status: done
priority: medium
Route ingestion through the backend contract; fix error semantics AUDIT-WP-0004 T01, T02, T07. T01 - ingestion wrote to SQLite directly and never called the AuditBackend contract, so a 202 meant a row existed rather than that a backend with a declared retention policy had accepted the event. Adds IdempotentAuditBackend to the contract: duplicate detection lives inside the backend so custody and idempotency state share a transaction and cannot diverge. SQLiteAuditBackend implements it with WAL, synchronous=FULL and a busy timeout. Ingestion now refuses any backend declaring durable=False, so the development file backend cannot silently become the production sink. The atomicity claim was tested rather than asserted, and the first attempt failed: with a single shared connection, 16 racing submissions of one event told two callers they were first. Storage was correct but the response was not. Fixed with per-thread connections and BEGIN IMMEDIATE around the insert/read pair, and locked in by a test. T02 - storage errors previously escaped the handler with start_response never called, and the auth check sat outside the try block so a non-ASCII Authorization header crashed the request. Adds a catch-all, maps conflict to 409, backend unavailability to 503 and unexpected faults to 500, and documents the full response contract with the retry semantics each status implies, since senders key their behaviour off it. T07 - ingestion tests 2 -> 23, suite 15 -> 36. accepted_at is now UTC rather than local time, and naive timestamps are rejected instead of silently assumed. Remaining in WP-0004: T03 tenant/source binding, T04 redaction policy, T05 operator read surface, T06 production serving layer. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 14:30:20 +02:00
state_hub_task_id: "9010fb4a-a1b8-4ef7-b143-e33ca7cc0619"
```
Any events accepted by the pre-production SQLite receiver are audit records
and cannot simply be dropped. Move them into the Postgres store with their
original identifiers, timestamps, and payload hashes intact, or record an
explicit decision that they are development artifacts with no custody value.
Whichever holds, it must be written down — silently discarding accepted audit
events is the exact failure this service is meant to make impossible.
Done when the disposition of every pre-production record is either migrated
and verified, or explicitly and justifiably discarded.
Establish pre-production record disposition and add the migration tool AUDIT-WP-0005-T04. Disposition: there are no pre-production records. No audit-core SQLite store on this host, no mock-file-backend output, and no audit-core pod, deployment or PVC on railiance01 - the only audit-* PVC there is OpenBao's own audit device. Consistent with the history: WP-0003-T03 was cancelled before the receiver was ever deployed, so every SQLite store that has existed was a test fixture. Nothing is being discarded because nothing was ever accepted outside tests. The tool is built anyway because the SQLite path stays reachable - the entrypoint falls back to it when AUDIT_CORE_DATABASE_URL is unset. If that fallback is ever used in anger the records are audit records, and writing the migration afterwards under pressure is the wrong time. audit_core.migrate_store and `python -m audit_core migrate-store` transfer events, dead letters and secret-finding counters. Records keep their original event_id, payload_hash and accepted_at, which is why this bypasses accept(): that stamps acceptance with the current time, and a migration that rewrote acceptance times would destroy the evidence it exists to preserve. Idempotent, and verification reads back from the destination rather than trusting the write path. A destination record with a differing payload hash is reported as a conflict and left untouched - silently overwriting a stored audit record is the same class of failure as losing it. Conflicts and failed verification exit non-zero; a partial migration is not a success. Tests 71 -> 77. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 17:31:18 +02:00
**Disposition, established 2026-08-10: there are no pre-production records.**
Checked and found empty: no `/data/audit-core.db` or any other audit-core
SQLite store on this host, no mock-file-backend output under `/tmp/audit-core`,
and no audit-core pod, deployment, or PVC on railiance01. (The only `audit-*`
PVC there is `audit-openbao-0`, which is OpenBao's own audit device and
unrelated to this service.)
That is consistent with the history rather than surprising: WP-0003-T03 was
cancelled before the receiver was ever deployed, so the only SQLite stores that
have ever existed were test fixtures with no custody value. Nothing is being
discarded, because nothing was ever accepted outside tests.
The migration tool was built regardless, because the SQLite path remains
reachable: the entrypoint falls back to it when `AUDIT_CORE_DATABASE_URL` is
unset. If that fallback is ever used in anger, the records it accumulates are
audit records and cannot simply be dropped — and building the tool afterwards,
under pressure, is the wrong time.
`audit_core.migrate_store` plus `python -m audit_core migrate-store` transfer
events, dead letters, and secret-finding counters. Fidelity is the point:
records keep their original `event_id`, `payload_hash`, and `accepted_at`,
which is why the tool bypasses `accept()` — that stamps acceptance with the
current time, and a migration that rewrote acceptance times would destroy the
evidence it exists to preserve.
Migration is idempotent, and verification reads back from the destination
rather than trusting the write path. Where the destination already holds an
event with a *different* payload hash, the tool reports a conflict and leaves
the destination untouched: silently overwriting a stored audit record with a
differing one is the same class of failure as losing it. A conflict or failed
verification exits non-zero — a partial migration is not a success.
Six tests cover fidelity, dead letters and counters, idempotency, conflict
handling, post-migration readability under the append-only trigger, and the
empty-source case that is today's actual situation.
## T05 - Run the live failure matrix
```task
id: AUDIT-WP-0005-T05
status: todo
priority: high
Route ingestion through the backend contract; fix error semantics AUDIT-WP-0004 T01, T02, T07. T01 - ingestion wrote to SQLite directly and never called the AuditBackend contract, so a 202 meant a row existed rather than that a backend with a declared retention policy had accepted the event. Adds IdempotentAuditBackend to the contract: duplicate detection lives inside the backend so custody and idempotency state share a transaction and cannot diverge. SQLiteAuditBackend implements it with WAL, synchronous=FULL and a busy timeout. Ingestion now refuses any backend declaring durable=False, so the development file backend cannot silently become the production sink. The atomicity claim was tested rather than asserted, and the first attempt failed: with a single shared connection, 16 racing submissions of one event told two callers they were first. Storage was correct but the response was not. Fixed with per-thread connections and BEGIN IMMEDIATE around the insert/read pair, and locked in by a test. T02 - storage errors previously escaped the handler with start_response never called, and the auth check sat outside the try block so a non-ASCII Authorization header crashed the request. Adds a catch-all, maps conflict to 409, backend unavailability to 503 and unexpected faults to 500, and documents the full response contract with the retry semantics each status implies, since senders key their behaviour off it. T07 - ingestion tests 2 -> 23, suite 15 -> 36. accepted_at is now UTC rather than local time, and naive timestamps are rejected instead of silently assumed. Remaining in WP-0004: T03 tenant/source binding, T04 redaction policy, T05 operator read surface, T06 production serving layer. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 14:30:20 +02:00
state_hub_task_id: "1da30fec-b9f1-4be0-b42c-15797a8c4392"
```
Exercise the deployed path: successful delivery; receiver timeout and
unavailability; bounded user-engine retry against each documented status code;
dead-letter visibility; operator replay and duplicate replay; redaction; and
correlation lookup. Include a database failover or restart during active
ingestion, and a credential rotation during active ingestion.
The controlling assertion is that one source outbox event produces exactly one
durable normalized event across retries, replay, and infrastructure
disruption — no loss, no duplication.
Hand non-secret evidence back to NK-WP-0024.
Done when the matrix has been run against the deployed system and each
outcome recorded, including any case where behaviour differed from the
documented contract.
## T06 - Operational handover
```task
id: AUDIT-WP-0005-T06
status: todo
priority: medium
Route ingestion through the backend contract; fix error semantics AUDIT-WP-0004 T01, T02, T07. T01 - ingestion wrote to SQLite directly and never called the AuditBackend contract, so a 202 meant a row existed rather than that a backend with a declared retention policy had accepted the event. Adds IdempotentAuditBackend to the contract: duplicate detection lives inside the backend so custody and idempotency state share a transaction and cannot diverge. SQLiteAuditBackend implements it with WAL, synchronous=FULL and a busy timeout. Ingestion now refuses any backend declaring durable=False, so the development file backend cannot silently become the production sink. The atomicity claim was tested rather than asserted, and the first attempt failed: with a single shared connection, 16 racing submissions of one event told two callers they were first. Storage was correct but the response was not. Fixed with per-thread connections and BEGIN IMMEDIATE around the insert/read pair, and locked in by a test. T02 - storage errors previously escaped the handler with start_response never called, and the auth check sat outside the try block so a non-ASCII Authorization header crashed the request. Adds a catch-all, maps conflict to 409, backend unavailability to 503 and unexpected faults to 500, and documents the full response contract with the retry semantics each status implies, since senders key their behaviour off it. T07 - ingestion tests 2 -> 23, suite 15 -> 36. accepted_at is now UTC rather than local time, and naive timestamps are rejected instead of silently assumed. Remaining in WP-0004: T03 tenant/source binding, T04 redaction policy, T05 operator read surface, T06 production serving layer. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 14:30:20 +02:00
state_hub_task_id: "0856c80d-abe1-4bff-ba8d-87295cf76819"
```
Document what an operator needs: how to look up an event by correlation ID,
how to inspect and replay dead-lettered events, how to rotate the sender
credential, what the alert conditions mean, and how to restore audit data from
a rapp-postgres backup.
Verify audit-core's recovery requirement against what rapp-postgres actually
provides — the retention window audit-core declares must not exceed the
retention the platform guarantees.
Done when the runbook exists and the restore path has been walked once.