audit-core/deploy/audit-core.yaml
tegwick 2f4e1adf66
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 1s
Add deployment manifests, custody-class guard and request counters
AUDIT-WP-0005-T03 (progress). Manifests validated --dry-run=server
--validate=strict against railiance01; not applied, since deployment is gated
on RAPP-POSTGRES-WP-0002 and T02 credentials. Nothing here mutates the cluster.

Conventions read off the deployed user-engine workload rather than invented:
digest-pinned image from forgejo.coulomb.social, runAsNonRoot with
RuntimeDefault seccomp, no privilege escalation, all capabilities dropped,
readOnlyRootFilesystem, probes on a named http port, same resource envelope.

The namespace carries railiance.io/postgres-client: platform-pg, which is what
platform-pg-consumer-ingress in rapp-postgres admits; without that label the
pod cannot reach the database at all.

NetworkPolicies default-deny both directions, then permit ingress from the
user-engine namespace only, a separately labelled operator read path, and
egress to PostgreSQL in databases plus DNS.

Three decisions worth naming. Liveness is /healthz while readiness is /readyz,
so a database outage drops the pod from the Service rather than restarting it
in a loop. readOnlyRootFilesystem enforces the empty-filesystem property rather
than trusting it, so the SQLite fallback physically cannot accumulate audit
records on ephemeral storage. AUDIT_CORE_REQUIRE_CUSTODY_CLASS=archive makes a
missing database URL a startup failure instead of a silent downgrade to the
development store.

Counters deferred from WP-0004-T06 are exposed as JSON at /v1/stats behind the
read privilege, not as Prometheus exposition format: the cluster runs no
Prometheus, no ServiceMonitor CRD and no other scrape target, so an exposition
endpoint would target a scrape path that does not exist. Usable with curl now
and a small step from /metrics later.

Tests 77 -> 80.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 17:42:43 +02:00

157 lines
5.4 KiB
YAML

# audit-core receiver — railiance01 (AUDIT-WP-0005-T03).
#
# Conventions match the deployed user-engine workload: digest-pinned image from
# forgejo.coulomb.social, non-root with a read-only root filesystem, and probes
# on a named http port.
#
# Apply order matters: the namespace label railiance.io/postgres-client is what
# platform-pg-consumer-ingress in the databases namespace admits, so without it
# the pod cannot reach the database at all.
---
apiVersion: v1
kind: Namespace
metadata:
name: audit-core
labels:
kubernetes.io/metadata.name: audit-core
railiance.io/workload-class: platform
# Admitted by NetworkPolicy platform-pg-consumer-ingress (rapp-postgres).
railiance.io/postgres-client: platform-pg
---
apiVersion: v1
kind: Service
metadata:
name: audit-core
namespace: audit-core
labels:
app.kubernetes.io/name: audit-core
spec:
type: ClusterIP
selector:
app.kubernetes.io/name: audit-core
ports:
- name: http
port: 8080
targetPort: http
protocol: TCP
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: audit-core
namespace: audit-core
labels:
app.kubernetes.io/name: audit-core
annotations:
# Rollback position. Update both together; `kubectl rollout undo` returns to
# the previous digest, and the schema note records whether that is safe.
audit-core.railiance.io/rollback-note: >-
Migrations 0001-0004 are additive (CREATE TABLE/INDEX/TRIGGER IF NOT
EXISTS) and are not reversed by a rollback. An older image runs against
the newer schema without harm. A future migration that drops or narrows a
column breaks that property and must state its own rollback position
before it is released.
spec:
replicas: 1
revisionHistoryLimit: 5
strategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 0
maxSurge: 1
selector:
matchLabels:
app.kubernetes.io/name: audit-core
template:
metadata:
labels:
app.kubernetes.io/name: audit-core
spec:
securityContext:
runAsNonRoot: true
seccompProfile:
type: RuntimeDefault
# In-flight events must not be lost on rollout; the app shuts down
# gracefully on SIGTERM (AUDIT-WP-0004-T06).
terminationGracePeriodSeconds: 30
containers:
- name: audit-core
# REPLACE at release time with the built digest. A mutable tag is not
# an immutable image, and `:latest` must never be the only reference.
image: forgejo.coulomb.social/coulomb/audit-core@sha256:REPLACE_AT_RELEASE
imagePullPolicy: IfNotPresent
ports:
- name: http
containerPort: 8080
env:
- name: AUDIT_CORE_HOST
value: "0.0.0.0"
- name: AUDIT_CORE_HTTP_PORT
value: "8080"
# Refuses to start on anything but archive-class custody, so a
# missing database URL fails loudly instead of silently downgrading
# to the development store.
- name: AUDIT_CORE_REQUIRE_CUSTODY_CLASS
value: archive
- name: AUDIT_CORE_DATABASE_SCHEMA
value: audit_core
- name: AUDIT_CORE_THREADS
value: "8"
- name: AUDIT_CORE_REQUEST_TIMEOUT
value: "30"
- name: AUDIT_CORE_DB_STATEMENT_TIMEOUT_MS
value: "30000"
# Both secrets are delivered through the OpenBao lane
# (AUDIT-WP-0005-T02) — never committed, never set by hand.
- name: AUDIT_CORE_DATABASE_URL
valueFrom:
secretKeyRef:
name: audit-core-database
key: url
- name: AUDIT_CORE_SENDERS
valueFrom:
secretKeyRef:
name: audit-core-senders
key: senders.json
resources:
requests:
cpu: 50m
memory: 64Mi
limits:
cpu: 500m
memory: 256Mi
securityContext:
allowPrivilegeEscalation: false
capabilities:
drop: ["ALL"]
# Custody lives in PostgreSQL. Nothing of value is written to the
# pod filesystem, and a read-only root enforces that rather than
# trusting it — the SQLite fallback cannot silently start
# accumulating audit records on ephemeral storage.
readOnlyRootFilesystem: true
volumeMounts:
- name: tmp
mountPath: /tmp
startupProbe:
httpGet: {path: /healthz, port: http}
periodSeconds: 3
failureThreshold: 20
readinessProbe:
# /readyz checks the backend is reachable and durable, so the pod
# leaves the Service when custody is unavailable rather than
# accepting events it cannot store.
httpGet: {path: /readyz, port: http}
periodSeconds: 10
timeoutSeconds: 2
failureThreshold: 3
livenessProbe:
# Deliberately /healthz, not /readyz: a database outage must not
# restart the pod in a loop. Losing readiness is the correct
# response; restarting solves nothing and loses in-flight work.
httpGet: {path: /healthz, port: http}
periodSeconds: 20
timeoutSeconds: 2
failureThreshold: 3
volumes:
- name: tmp
emptyDir: {}