Add deployment manifests, custody-class guard and request counters
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 1s

AUDIT-WP-0005-T03 (progress). Manifests validated --dry-run=server
--validate=strict against railiance01; not applied, since deployment is gated
on RAPP-POSTGRES-WP-0002 and T02 credentials. Nothing here mutates the cluster.

Conventions read off the deployed user-engine workload rather than invented:
digest-pinned image from forgejo.coulomb.social, runAsNonRoot with
RuntimeDefault seccomp, no privilege escalation, all capabilities dropped,
readOnlyRootFilesystem, probes on a named http port, same resource envelope.

The namespace carries railiance.io/postgres-client: platform-pg, which is what
platform-pg-consumer-ingress in rapp-postgres admits; without that label the
pod cannot reach the database at all.

NetworkPolicies default-deny both directions, then permit ingress from the
user-engine namespace only, a separately labelled operator read path, and
egress to PostgreSQL in databases plus DNS.

Three decisions worth naming. Liveness is /healthz while readiness is /readyz,
so a database outage drops the pod from the Service rather than restarting it
in a loop. readOnlyRootFilesystem enforces the empty-filesystem property rather
than trusting it, so the SQLite fallback physically cannot accumulate audit
records on ephemeral storage. AUDIT_CORE_REQUIRE_CUSTODY_CLASS=archive makes a
missing database URL a startup failure instead of a silent downgrade to the
development store.

Counters deferred from WP-0004-T06 are exposed as JSON at /v1/stats behind the
read privilege, not as Prometheus exposition format: the cluster runs no
Prometheus, no ServiceMonitor CRD and no other scrape target, so an exposition
endpoint would target a scrape path that does not exist. Usable with curl now
and a small step from /metrics later.

Tests 77 -> 80.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
tegwick 2026-08-10 17:42:43 +02:00
parent 88d16847ff
commit 2f4e1adf66
6 changed files with 393 additions and 5 deletions

View file

@ -137,7 +137,7 @@ events.
```task
id: AUDIT-WP-0005-T03
status: todo
status: progress
priority: high
state_hub_task_id: "598af2ac-e772-4a4e-9a65-dde9d4ca167f"
```
@ -162,6 +162,57 @@ Done when the receiver is reachable only by its declared peers, survives pod
restart and rescheduling without loss, and has a tested path back to the
previous version.
Progress 2026-08-10: manifests written and validated `--dry-run=server
--validate=strict` against railiance01. Not applied — deployment is gated on
RAPP-POSTGRES-WP-0002 and on T02's credentials, and nothing here mutates the
cluster.
Conventions were read off the deployed `user-engine` workload rather than
invented: digest-pinned image from `forgejo.coulomb.social/coulomb/<name>`,
`runAsNonRoot` + `RuntimeDefault` seccomp, `allowPrivilegeEscalation: false`,
all capabilities dropped, `readOnlyRootFilesystem: true`, probes on a named
`http` port, and the same resource envelope.
`deploy/audit-core.yaml` — namespace, Service, Deployment. The namespace
carries `railiance.io/postgres-client: platform-pg`, which is what
`platform-pg-consumer-ingress` in rapp-postgres admits; without that label the
pod cannot reach the database at all.
`deploy/networkpolicies.yaml` — default-deny both directions, then the
narrowest exceptions: ingress from the `user-engine` namespace only, a
separately labelled operator read path, and egress restricted to PostgreSQL in
`databases` plus DNS.
Three decisions worth recording:
- **Liveness is `/healthz`, readiness is `/readyz`.** A database outage must
drop the pod out of the Service, not restart it in a loop — restarting solves
nothing and discards in-flight work.
- **`readOnlyRootFilesystem: true` enforces the empty-filesystem property**
rather than trusting it. Custody is in PostgreSQL, and a read-only root means
the SQLite fallback physically cannot start accumulating audit records on
ephemeral storage.
- **`AUDIT_CORE_REQUIRE_CUSTODY_CLASS=archive`** makes a missing database URL a
startup failure instead of a silent downgrade to the development store. Added
in this task; covered by a test.
Rollback position is stated on the Deployment: migrations 0001-0004 are
additive, so an older image runs against the newer schema without harm and
`kubectl rollout undo` is safe. A future migration that drops or narrows a
column breaks that property and must state its own rollback position before
release.
Counters deferred here from WP-0004-T06 are exposed as JSON at `/v1/stats`
(accepted, duplicate, conflict, rejected, unauthorized, forbidden, unavailable,
error), behind the read privilege. Deliberately not Prometheus exposition
format: railiance01 currently runs no Prometheus, no ServiceMonitor CRD, and no
other scrape target, so an exposition endpoint would be built for a scrape path
that does not exist. This is usable with curl today and a small step from
`/metrics` when a metrics stack lands.
Remaining before done: build and publish the image to get a real digest, apply,
and verify restart/rescheduling and the rollback path on the cluster.
## T04 - Migrate existing SQLite records
```task