Rescope WP-0003 and open WP-0004/WP-0005 after pre-deploy review
A pre-deploy review found the receiver is not deployable as built. The
ingestion path never calls AuditBackend.emit(), so a 202 means a SQLite row
exists, not that a backend with a declared retention policy accepted the
event. Storage exceptions escape the handler with start_response never
called. Tenant isolation and source binding are recorded as done but are not
implemented. There is no read, replay, or correlation-lookup surface, so the
failure matrix cannot produce the evidence NK-WP-0024 needs.
WP-0003 closes as finished on its narrowed scope (contract + reference
implementation, T01/T02). T03 and T04 are cancelled with rationale.
WP-0004 covers receiver correctness and hardening, storage-agnostic so it
runs in parallel with the database platform work.
WP-0005 covers the Postgres backend, deployment, SQLite record migration,
and the live failure matrix. Production storage moves from SQLite-on-a-volume
to the Railiance PostgreSQL platform (RAPP-POSTGRES-WP-0002).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 12:44:23 +02:00
|
|
|
---
|
|
|
|
|
id: AUDIT-WP-0005
|
|
|
|
|
type: workplan
|
|
|
|
|
title: "Deploy audit-core on Railiance with durable Postgres custody"
|
|
|
|
|
domain: infotech
|
|
|
|
|
repo: audit-core
|
|
|
|
|
status: proposed
|
|
|
|
|
owner: codex
|
|
|
|
|
topic_slug: netkingdom
|
|
|
|
|
created: "2026-08-10"
|
|
|
|
|
updated: "2026-08-10"
|
|
|
|
|
depends_on:
|
|
|
|
|
- AUDIT-WP-0004
|
|
|
|
|
- RAPP-POSTGRES-WP-0002
|
|
|
|
|
- NK-WP-0024
|
Route ingestion through the backend contract; fix error semantics
AUDIT-WP-0004 T01, T02, T07.
T01 - ingestion wrote to SQLite directly and never called the AuditBackend
contract, so a 202 meant a row existed rather than that a backend with a
declared retention policy had accepted the event. Adds IdempotentAuditBackend
to the contract: duplicate detection lives inside the backend so custody and
idempotency state share a transaction and cannot diverge. SQLiteAuditBackend
implements it with WAL, synchronous=FULL and a busy timeout. Ingestion now
refuses any backend declaring durable=False, so the development file backend
cannot silently become the production sink.
The atomicity claim was tested rather than asserted, and the first attempt
failed: with a single shared connection, 16 racing submissions of one event
told two callers they were first. Storage was correct but the response was
not. Fixed with per-thread connections and BEGIN IMMEDIATE around the
insert/read pair, and locked in by a test.
T02 - storage errors previously escaped the handler with start_response never
called, and the auth check sat outside the try block so a non-ASCII
Authorization header crashed the request. Adds a catch-all, maps conflict to
409, backend unavailability to 503 and unexpected faults to 500, and
documents the full response contract with the retry semantics each status
implies, since senders key their behaviour off it.
T07 - ingestion tests 2 -> 23, suite 15 -> 36. accepted_at is now UTC rather
than local time, and naive timestamps are rejected instead of silently
assumed.
Remaining in WP-0004: T03 tenant/source binding, T04 redaction policy,
T05 operator read surface, T06 production serving layer.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 14:30:20 +02:00
|
|
|
state_hub_workstream_id: "7b24a844-c9f2-4d2d-ac7d-20978bdf6b38"
|
Rescope WP-0003 and open WP-0004/WP-0005 after pre-deploy review
A pre-deploy review found the receiver is not deployable as built. The
ingestion path never calls AuditBackend.emit(), so a 202 means a SQLite row
exists, not that a backend with a declared retention policy accepted the
event. Storage exceptions escape the handler with start_response never
called. Tenant isolation and source binding are recorded as done but are not
implemented. There is no read, replay, or correlation-lookup surface, so the
failure matrix cannot produce the evidence NK-WP-0024 needs.
WP-0003 closes as finished on its narrowed scope (contract + reference
implementation, T01/T02). T03 and T04 are cancelled with rationale.
WP-0004 covers receiver correctness and hardening, storage-agnostic so it
runs in parallel with the database platform work.
WP-0005 covers the Postgres backend, deployment, SQLite record migration,
and the live failure matrix. Production storage moves from SQLite-on-a-volume
to the Railiance PostgreSQL platform (RAPP-POSTGRES-WP-0002).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 12:44:23 +02:00
|
|
|
---
|
|
|
|
|
|
|
|
|
|
# AUDIT-WP-0005 - Postgres store and production deployment
|
|
|
|
|
|
|
|
|
|
## Goal
|
|
|
|
|
|
|
|
|
|
Put audit-core into production on Railiance with PostgreSQL as the custody
|
|
|
|
|
store, and prove the delivery path end to end.
|
|
|
|
|
|
|
|
|
|
This is first on the critical path: email-connect delivery, OpenBao sender
|
|
|
|
|
credentials, the user-engine runtime integrations, and the live failure matrix
|
|
|
|
|
all sit downstream of a receiver that actually holds evidence.
|
|
|
|
|
|
|
|
|
|
## Dependencies
|
|
|
|
|
|
|
|
|
|
- **AUDIT-WP-0004** — the receiver must route through the backend contract and
|
|
|
|
|
return deliberate status codes before a Postgres backend is worth writing
|
|
|
|
|
behind it, and before a failure matrix can distinguish a correct refusal
|
|
|
|
|
from a crash.
|
|
|
|
|
- **RAPP-POSTGRES-WP-0002** — the database, the tenancy model, and the
|
|
|
|
|
credential lane. T01 below can be drafted against the model as soon as
|
|
|
|
|
RAPP-POSTGRES-WP-0002-T01 settles it; T02 needs a provisioned database.
|
|
|
|
|
|
|
|
|
|
## Boundaries
|
|
|
|
|
|
|
|
|
|
This workplan owns the Postgres backend implementation, audit-core's
|
|
|
|
|
deployment, and the live verification. It does not own the database platform,
|
|
|
|
|
its tenancy model, or its backup machinery — those belong to rapp-postgres.
|
|
|
|
|
Where audit-core needs a guarantee from the platform, it states the
|
|
|
|
|
requirement and consumes it; it does not implement it here.
|
|
|
|
|
|
|
|
|
|
## T01 - Implement the Postgres audit backend
|
|
|
|
|
|
|
|
|
|
```task
|
|
|
|
|
id: AUDIT-WP-0005-T01
|
|
|
|
|
status: todo
|
|
|
|
|
priority: high
|
Route ingestion through the backend contract; fix error semantics
AUDIT-WP-0004 T01, T02, T07.
T01 - ingestion wrote to SQLite directly and never called the AuditBackend
contract, so a 202 meant a row existed rather than that a backend with a
declared retention policy had accepted the event. Adds IdempotentAuditBackend
to the contract: duplicate detection lives inside the backend so custody and
idempotency state share a transaction and cannot diverge. SQLiteAuditBackend
implements it with WAL, synchronous=FULL and a busy timeout. Ingestion now
refuses any backend declaring durable=False, so the development file backend
cannot silently become the production sink.
The atomicity claim was tested rather than asserted, and the first attempt
failed: with a single shared connection, 16 racing submissions of one event
told two callers they were first. Storage was correct but the response was
not. Fixed with per-thread connections and BEGIN IMMEDIATE around the
insert/read pair, and locked in by a test.
T02 - storage errors previously escaped the handler with start_response never
called, and the auth check sat outside the try block so a non-ASCII
Authorization header crashed the request. Adds a catch-all, maps conflict to
409, backend unavailability to 503 and unexpected faults to 500, and
documents the full response contract with the retry semantics each status
implies, since senders key their behaviour off it.
T07 - ingestion tests 2 -> 23, suite 15 -> 36. accepted_at is now UTC rather
than local time, and naive timestamps are rejected instead of silently
assumed.
Remaining in WP-0004: T03 tenant/source binding, T04 redaction policy,
T05 operator read surface, T06 production serving layer.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 14:30:20 +02:00
|
|
|
state_hub_task_id: "b1601d0b-922a-40f7-92c0-ea06af6c4468"
|
Rescope WP-0003 and open WP-0004/WP-0005 after pre-deploy review
A pre-deploy review found the receiver is not deployable as built. The
ingestion path never calls AuditBackend.emit(), so a 202 means a SQLite row
exists, not that a backend with a declared retention policy accepted the
event. Storage exceptions escape the handler with start_response never
called. Tenant isolation and source binding are recorded as done but are not
implemented. There is no read, replay, or correlation-lookup surface, so the
failure matrix cannot produce the evidence NK-WP-0024 needs.
WP-0003 closes as finished on its narrowed scope (contract + reference
implementation, T01/T02). T03 and T04 are cancelled with rationale.
WP-0004 covers receiver correctness and hardening, storage-agnostic so it
runs in parallel with the database platform work.
WP-0005 covers the Postgres backend, deployment, SQLite record migration,
and the live failure matrix. Production storage moves from SQLite-on-a-volume
to the Railiance PostgreSQL platform (RAPP-POSTGRES-WP-0002).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 12:44:23 +02:00
|
|
|
```
|
|
|
|
|
|
|
|
|
|
Implement `AuditBackend` against PostgreSQL, declaring an honest
|
|
|
|
|
`RetentionPolicy` — the custody class, retention window, and whether
|
|
|
|
|
immutability and tamper evidence are genuinely provided rather than aspired
|
|
|
|
|
to. If the schema does not prevent an operator from silently editing a
|
|
|
|
|
recorded event, the policy must not claim `immutable`.
|
|
|
|
|
|
|
|
|
|
Push idempotency into the database rather than a read-then-write in
|
|
|
|
|
application code: an insert conflicting on event ID must resolve atomically
|
|
|
|
|
into duplicate-accepted or conflict, with no window where two concurrent
|
|
|
|
|
identical events both write. Store the payload hash for conflict detection.
|
|
|
|
|
|
|
|
|
|
Handle connection lifecycle properly — pooling, reconnection after a database
|
|
|
|
|
restart, and a bounded statement timeout so a stalled write surfaces as
|
|
|
|
|
unavailable instead of hanging the request.
|
|
|
|
|
|
|
|
|
|
Own the migrations. The consuming service owns its schema; rapp-postgres owns
|
|
|
|
|
the space it runs in.
|
|
|
|
|
|
|
|
|
|
Done when the backend passes the same contract tests as the existing backends,
|
|
|
|
|
concurrent duplicate submissions produce exactly one record, and a database
|
|
|
|
|
restart mid-write does not produce an acknowledged-but-absent event.
|
|
|
|
|
|
|
|
|
|
## T02 - Provision storage through the platform lane
|
|
|
|
|
|
|
|
|
|
```task
|
|
|
|
|
id: AUDIT-WP-0005-T02
|
|
|
|
|
status: todo
|
|
|
|
|
priority: high
|
Route ingestion through the backend contract; fix error semantics
AUDIT-WP-0004 T01, T02, T07.
T01 - ingestion wrote to SQLite directly and never called the AuditBackend
contract, so a 202 meant a row existed rather than that a backend with a
declared retention policy had accepted the event. Adds IdempotentAuditBackend
to the contract: duplicate detection lives inside the backend so custody and
idempotency state share a transaction and cannot diverge. SQLiteAuditBackend
implements it with WAL, synchronous=FULL and a busy timeout. Ingestion now
refuses any backend declaring durable=False, so the development file backend
cannot silently become the production sink.
The atomicity claim was tested rather than asserted, and the first attempt
failed: with a single shared connection, 16 racing submissions of one event
told two callers they were first. Storage was correct but the response was
not. Fixed with per-thread connections and BEGIN IMMEDIATE around the
insert/read pair, and locked in by a test.
T02 - storage errors previously escaped the handler with start_response never
called, and the auth check sat outside the try block so a non-ASCII
Authorization header crashed the request. Adds a catch-all, maps conflict to
409, backend unavailability to 503 and unexpected faults to 500, and
documents the full response contract with the retry semantics each status
implies, since senders key their behaviour off it.
T07 - ingestion tests 2 -> 23, suite 15 -> 36. accepted_at is now UTC rather
than local time, and naive timestamps are rejected instead of silently
assumed.
Remaining in WP-0004: T03 tenant/source binding, T04 redaction policy,
T05 operator read surface, T06 production serving layer.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 14:30:20 +02:00
|
|
|
state_hub_task_id: "831b2472-0d80-4369-a5e3-eb08ef3526b1"
|
Rescope WP-0003 and open WP-0004/WP-0005 after pre-deploy review
A pre-deploy review found the receiver is not deployable as built. The
ingestion path never calls AuditBackend.emit(), so a 202 means a SQLite row
exists, not that a backend with a declared retention policy accepted the
event. Storage exceptions escape the handler with start_response never
called. Tenant isolation and source binding are recorded as done but are not
implemented. There is no read, replay, or correlation-lookup surface, so the
failure matrix cannot produce the evidence NK-WP-0024 needs.
WP-0003 closes as finished on its narrowed scope (contract + reference
implementation, T01/T02). T03 and T04 are cancelled with rationale.
WP-0004 covers receiver correctness and hardening, storage-agnostic so it
runs in parallel with the database platform work.
WP-0005 covers the Postgres backend, deployment, SQLite record migration,
and the live failure matrix. Production storage moves from SQLite-on-a-volume
to the Railiance PostgreSQL platform (RAPP-POSTGRES-WP-0002).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 12:44:23 +02:00
|
|
|
```
|
|
|
|
|
|
|
|
|
|
Declare audit-core's database requirement against rapp-postgres and take
|
|
|
|
|
delivery of a least-privilege runtime role through the OpenBao lane. The
|
|
|
|
|
runtime role connects; a separate role runs migrations; neither owns more than
|
|
|
|
|
it needs.
|
|
|
|
|
|
|
|
|
|
Verify audit-core's side of the isolation model: the runtime credential
|
|
|
|
|
reaches audit-core's data and nothing else, and rotation completes without a
|
|
|
|
|
delivery gap.
|
|
|
|
|
|
|
|
|
|
Done when audit-core runs against a provisioned database using a credential it
|
|
|
|
|
never received as a literal, and rotating that credential does not drop
|
|
|
|
|
events.
|
|
|
|
|
|
|
|
|
|
## T03 - Deploy the receiver
|
|
|
|
|
|
|
|
|
|
```task
|
|
|
|
|
id: AUDIT-WP-0005-T03
|
|
|
|
|
status: todo
|
|
|
|
|
priority: high
|
Route ingestion through the backend contract; fix error semantics
AUDIT-WP-0004 T01, T02, T07.
T01 - ingestion wrote to SQLite directly and never called the AuditBackend
contract, so a 202 meant a row existed rather than that a backend with a
declared retention policy had accepted the event. Adds IdempotentAuditBackend
to the contract: duplicate detection lives inside the backend so custody and
idempotency state share a transaction and cannot diverge. SQLiteAuditBackend
implements it with WAL, synchronous=FULL and a busy timeout. Ingestion now
refuses any backend declaring durable=False, so the development file backend
cannot silently become the production sink.
The atomicity claim was tested rather than asserted, and the first attempt
failed: with a single shared connection, 16 racing submissions of one event
told two callers they were first. Storage was correct but the response was
not. Fixed with per-thread connections and BEGIN IMMEDIATE around the
insert/read pair, and locked in by a test.
T02 - storage errors previously escaped the handler with start_response never
called, and the auth check sat outside the try block so a non-ASCII
Authorization header crashed the request. Adds a catch-all, maps conflict to
409, backend unavailability to 503 and unexpected faults to 500, and
documents the full response contract with the retry semantics each status
implies, since senders key their behaviour off it.
T07 - ingestion tests 2 -> 23, suite 15 -> 36. accepted_at is now UTC rather
than local time, and naive timestamps are rejected instead of silently
assumed.
Remaining in WP-0004: T03 tenant/source binding, T04 redaction policy,
T05 operator read surface, T06 production serving layer.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 14:30:20 +02:00
|
|
|
state_hub_task_id: "598af2ac-e772-4a4e-9a65-dde9d4ca167f"
|
Rescope WP-0003 and open WP-0004/WP-0005 after pre-deploy review
A pre-deploy review found the receiver is not deployable as built. The
ingestion path never calls AuditBackend.emit(), so a 202 means a SQLite row
exists, not that a backend with a declared retention policy accepted the
event. Storage exceptions escape the handler with start_response never
called. Tenant isolation and source binding are recorded as done but are not
implemented. There is no read, replay, or correlation-lookup surface, so the
failure matrix cannot produce the evidence NK-WP-0024 needs.
WP-0003 closes as finished on its narrowed scope (contract + reference
implementation, T01/T02). T03 and T04 are cancelled with rationale.
WP-0004 covers receiver correctness and hardening, storage-agnostic so it
runs in parallel with the database platform work.
WP-0005 covers the Postgres backend, deployment, SQLite record migration,
and the live failure matrix. Production storage moves from SQLite-on-a-volume
to the Railiance PostgreSQL platform (RAPP-POSTGRES-WP-0002).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 12:44:23 +02:00
|
|
|
```
|
|
|
|
|
|
|
|
|
|
Publish an immutable image — base pinned by digest, not a mutable tag, and
|
|
|
|
|
labelled with the build commit — and deploy on railiance01 with a Service,
|
|
|
|
|
health and readiness probes wired to the checks from WP-0004-T01, resource
|
|
|
|
|
requests and limits, a restricted security context, and a default-deny
|
|
|
|
|
NetworkPolicy admitting only user-engine as sender and the operator path for
|
|
|
|
|
reads.
|
|
|
|
|
|
|
|
|
|
State the rollback position: which image digest and which schema version the
|
|
|
|
|
deployment can return to, and whether the migration in T01 is reversible. A
|
|
|
|
|
rollback plan that assumes reversible migrations without checking is not a
|
|
|
|
|
plan.
|
|
|
|
|
|
|
|
|
|
The container currently runs as uid 10001 and expects a writable `/data`; once
|
|
|
|
|
custody is in Postgres that path should carry no durable state at all. Confirm
|
|
|
|
|
nothing of value is left on the pod filesystem.
|
|
|
|
|
|
|
|
|
|
Done when the receiver is reachable only by its declared peers, survives pod
|
|
|
|
|
restart and rescheduling without loss, and has a tested path back to the
|
|
|
|
|
previous version.
|
|
|
|
|
|
|
|
|
|
## T04 - Migrate existing SQLite records
|
|
|
|
|
|
|
|
|
|
```task
|
|
|
|
|
id: AUDIT-WP-0005-T04
|
|
|
|
|
status: todo
|
|
|
|
|
priority: medium
|
Route ingestion through the backend contract; fix error semantics
AUDIT-WP-0004 T01, T02, T07.
T01 - ingestion wrote to SQLite directly and never called the AuditBackend
contract, so a 202 meant a row existed rather than that a backend with a
declared retention policy had accepted the event. Adds IdempotentAuditBackend
to the contract: duplicate detection lives inside the backend so custody and
idempotency state share a transaction and cannot diverge. SQLiteAuditBackend
implements it with WAL, synchronous=FULL and a busy timeout. Ingestion now
refuses any backend declaring durable=False, so the development file backend
cannot silently become the production sink.
The atomicity claim was tested rather than asserted, and the first attempt
failed: with a single shared connection, 16 racing submissions of one event
told two callers they were first. Storage was correct but the response was
not. Fixed with per-thread connections and BEGIN IMMEDIATE around the
insert/read pair, and locked in by a test.
T02 - storage errors previously escaped the handler with start_response never
called, and the auth check sat outside the try block so a non-ASCII
Authorization header crashed the request. Adds a catch-all, maps conflict to
409, backend unavailability to 503 and unexpected faults to 500, and
documents the full response contract with the retry semantics each status
implies, since senders key their behaviour off it.
T07 - ingestion tests 2 -> 23, suite 15 -> 36. accepted_at is now UTC rather
than local time, and naive timestamps are rejected instead of silently
assumed.
Remaining in WP-0004: T03 tenant/source binding, T04 redaction policy,
T05 operator read surface, T06 production serving layer.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 14:30:20 +02:00
|
|
|
state_hub_task_id: "9010fb4a-a1b8-4ef7-b143-e33ca7cc0619"
|
Rescope WP-0003 and open WP-0004/WP-0005 after pre-deploy review
A pre-deploy review found the receiver is not deployable as built. The
ingestion path never calls AuditBackend.emit(), so a 202 means a SQLite row
exists, not that a backend with a declared retention policy accepted the
event. Storage exceptions escape the handler with start_response never
called. Tenant isolation and source binding are recorded as done but are not
implemented. There is no read, replay, or correlation-lookup surface, so the
failure matrix cannot produce the evidence NK-WP-0024 needs.
WP-0003 closes as finished on its narrowed scope (contract + reference
implementation, T01/T02). T03 and T04 are cancelled with rationale.
WP-0004 covers receiver correctness and hardening, storage-agnostic so it
runs in parallel with the database platform work.
WP-0005 covers the Postgres backend, deployment, SQLite record migration,
and the live failure matrix. Production storage moves from SQLite-on-a-volume
to the Railiance PostgreSQL platform (RAPP-POSTGRES-WP-0002).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 12:44:23 +02:00
|
|
|
```
|
|
|
|
|
|
|
|
|
|
Any events accepted by the pre-production SQLite receiver are audit records
|
|
|
|
|
and cannot simply be dropped. Move them into the Postgres store with their
|
|
|
|
|
original identifiers, timestamps, and payload hashes intact, or record an
|
|
|
|
|
explicit decision that they are development artifacts with no custody value.
|
|
|
|
|
|
|
|
|
|
Whichever holds, it must be written down — silently discarding accepted audit
|
|
|
|
|
events is the exact failure this service is meant to make impossible.
|
|
|
|
|
|
|
|
|
|
Done when the disposition of every pre-production record is either migrated
|
|
|
|
|
and verified, or explicitly and justifiably discarded.
|
|
|
|
|
|
|
|
|
|
## T05 - Run the live failure matrix
|
|
|
|
|
|
|
|
|
|
```task
|
|
|
|
|
id: AUDIT-WP-0005-T05
|
|
|
|
|
status: todo
|
|
|
|
|
priority: high
|
Route ingestion through the backend contract; fix error semantics
AUDIT-WP-0004 T01, T02, T07.
T01 - ingestion wrote to SQLite directly and never called the AuditBackend
contract, so a 202 meant a row existed rather than that a backend with a
declared retention policy had accepted the event. Adds IdempotentAuditBackend
to the contract: duplicate detection lives inside the backend so custody and
idempotency state share a transaction and cannot diverge. SQLiteAuditBackend
implements it with WAL, synchronous=FULL and a busy timeout. Ingestion now
refuses any backend declaring durable=False, so the development file backend
cannot silently become the production sink.
The atomicity claim was tested rather than asserted, and the first attempt
failed: with a single shared connection, 16 racing submissions of one event
told two callers they were first. Storage was correct but the response was
not. Fixed with per-thread connections and BEGIN IMMEDIATE around the
insert/read pair, and locked in by a test.
T02 - storage errors previously escaped the handler with start_response never
called, and the auth check sat outside the try block so a non-ASCII
Authorization header crashed the request. Adds a catch-all, maps conflict to
409, backend unavailability to 503 and unexpected faults to 500, and
documents the full response contract with the retry semantics each status
implies, since senders key their behaviour off it.
T07 - ingestion tests 2 -> 23, suite 15 -> 36. accepted_at is now UTC rather
than local time, and naive timestamps are rejected instead of silently
assumed.
Remaining in WP-0004: T03 tenant/source binding, T04 redaction policy,
T05 operator read surface, T06 production serving layer.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 14:30:20 +02:00
|
|
|
state_hub_task_id: "1da30fec-b9f1-4be0-b42c-15797a8c4392"
|
Rescope WP-0003 and open WP-0004/WP-0005 after pre-deploy review
A pre-deploy review found the receiver is not deployable as built. The
ingestion path never calls AuditBackend.emit(), so a 202 means a SQLite row
exists, not that a backend with a declared retention policy accepted the
event. Storage exceptions escape the handler with start_response never
called. Tenant isolation and source binding are recorded as done but are not
implemented. There is no read, replay, or correlation-lookup surface, so the
failure matrix cannot produce the evidence NK-WP-0024 needs.
WP-0003 closes as finished on its narrowed scope (contract + reference
implementation, T01/T02). T03 and T04 are cancelled with rationale.
WP-0004 covers receiver correctness and hardening, storage-agnostic so it
runs in parallel with the database platform work.
WP-0005 covers the Postgres backend, deployment, SQLite record migration,
and the live failure matrix. Production storage moves from SQLite-on-a-volume
to the Railiance PostgreSQL platform (RAPP-POSTGRES-WP-0002).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 12:44:23 +02:00
|
|
|
```
|
|
|
|
|
|
|
|
|
|
Exercise the deployed path: successful delivery; receiver timeout and
|
|
|
|
|
unavailability; bounded user-engine retry against each documented status code;
|
|
|
|
|
dead-letter visibility; operator replay and duplicate replay; redaction; and
|
|
|
|
|
correlation lookup. Include a database failover or restart during active
|
|
|
|
|
ingestion, and a credential rotation during active ingestion.
|
|
|
|
|
|
|
|
|
|
The controlling assertion is that one source outbox event produces exactly one
|
|
|
|
|
durable normalized event across retries, replay, and infrastructure
|
|
|
|
|
disruption — no loss, no duplication.
|
|
|
|
|
|
|
|
|
|
Hand non-secret evidence back to NK-WP-0024.
|
|
|
|
|
|
|
|
|
|
Done when the matrix has been run against the deployed system and each
|
|
|
|
|
outcome recorded, including any case where behaviour differed from the
|
|
|
|
|
documented contract.
|
|
|
|
|
|
|
|
|
|
## T06 - Operational handover
|
|
|
|
|
|
|
|
|
|
```task
|
|
|
|
|
id: AUDIT-WP-0005-T06
|
|
|
|
|
status: todo
|
|
|
|
|
priority: medium
|
Route ingestion through the backend contract; fix error semantics
AUDIT-WP-0004 T01, T02, T07.
T01 - ingestion wrote to SQLite directly and never called the AuditBackend
contract, so a 202 meant a row existed rather than that a backend with a
declared retention policy had accepted the event. Adds IdempotentAuditBackend
to the contract: duplicate detection lives inside the backend so custody and
idempotency state share a transaction and cannot diverge. SQLiteAuditBackend
implements it with WAL, synchronous=FULL and a busy timeout. Ingestion now
refuses any backend declaring durable=False, so the development file backend
cannot silently become the production sink.
The atomicity claim was tested rather than asserted, and the first attempt
failed: with a single shared connection, 16 racing submissions of one event
told two callers they were first. Storage was correct but the response was
not. Fixed with per-thread connections and BEGIN IMMEDIATE around the
insert/read pair, and locked in by a test.
T02 - storage errors previously escaped the handler with start_response never
called, and the auth check sat outside the try block so a non-ASCII
Authorization header crashed the request. Adds a catch-all, maps conflict to
409, backend unavailability to 503 and unexpected faults to 500, and
documents the full response contract with the retry semantics each status
implies, since senders key their behaviour off it.
T07 - ingestion tests 2 -> 23, suite 15 -> 36. accepted_at is now UTC rather
than local time, and naive timestamps are rejected instead of silently
assumed.
Remaining in WP-0004: T03 tenant/source binding, T04 redaction policy,
T05 operator read surface, T06 production serving layer.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 14:30:20 +02:00
|
|
|
state_hub_task_id: "0856c80d-abe1-4bff-ba8d-87295cf76819"
|
Rescope WP-0003 and open WP-0004/WP-0005 after pre-deploy review
A pre-deploy review found the receiver is not deployable as built. The
ingestion path never calls AuditBackend.emit(), so a 202 means a SQLite row
exists, not that a backend with a declared retention policy accepted the
event. Storage exceptions escape the handler with start_response never
called. Tenant isolation and source binding are recorded as done but are not
implemented. There is no read, replay, or correlation-lookup surface, so the
failure matrix cannot produce the evidence NK-WP-0024 needs.
WP-0003 closes as finished on its narrowed scope (contract + reference
implementation, T01/T02). T03 and T04 are cancelled with rationale.
WP-0004 covers receiver correctness and hardening, storage-agnostic so it
runs in parallel with the database platform work.
WP-0005 covers the Postgres backend, deployment, SQLite record migration,
and the live failure matrix. Production storage moves from SQLite-on-a-volume
to the Railiance PostgreSQL platform (RAPP-POSTGRES-WP-0002).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 12:44:23 +02:00
|
|
|
```
|
|
|
|
|
|
|
|
|
|
Document what an operator needs: how to look up an event by correlation ID,
|
|
|
|
|
how to inspect and replay dead-lettered events, how to rotate the sender
|
|
|
|
|
credential, what the alert conditions mean, and how to restore audit data from
|
|
|
|
|
a rapp-postgres backup.
|
|
|
|
|
|
|
|
|
|
Verify audit-core's recovery requirement against what rapp-postgres actually
|
|
|
|
|
provides — the retention window audit-core declares must not exceed the
|
|
|
|
|
retention the platform guarantees.
|
|
|
|
|
|
|
|
|
|
Done when the runbook exists and the restore path has been walked once.
|