AUDIT-WP-0009 T02/T10 — schedule attestation, and make the §5 check total

T02. deploy/attest-cronjob.yaml: daily at 03:17 UTC against the 168h window,
its own ServiceAccount, and a Role reaching exactly one named ConfigMap —
get/update/patch, no create, no list. audit_core/attest_publish.py does the
publish in stdlib; the image carries no kubectl, and adding one to an audit
receiver's image to write a single file is the worse trade.

Three refusals, all deliberate:

  The producer is not the receiver. A receiver that could rewrite its own
  attestation could forge it. audit-core-egress is now scoped to
  component: receiver and a separate audit-core-attest-egress carries the 6443
  rule, so the receiver never gains API-server reach. Asserted by test.

  It refuses to publish over a broken chain. A fresh head written over a break
  replaces an honest chain_break with a fresh-looking attestation. Stale
  degrades the claim visibly; false does not.

  Mounted as a directory, not subPath. Found while writing the manifest: a
  subPath ConfigMap mount is resolved once at pod start and never updates, so
  the daily attestation would land in the ConfigMap and never reach the running
  receiver — tamper_evidence would age out to false while the job reported
  success every night, silent in both directions.

The offsite copy stays an operator step. audit-core holds no Nextcloud
credential and should not acquire one to publish a hash, so docs/integrity.md
states the bound plainly: until that copy exists the delivered control defends
against a database owner, not a cluster owner, and no stronger claim may be
made from it.

T10. layer.yaml lists four infrastructure contacts — platform-pg, state-hub,
kube-apiserver, the container registry — each with its role and whether another
layer reads it. tooling_contacts stays [], which is true under §5 as written;
the companion's totality request is met by the uncatalogued list rather than by
inventing a Tooling row. tests/test_layer_conformance.py derives the egress
destinations from the manifests and the registry from the pinned digests, so a
new contact appearing in deploy/ without a row fails the test rather than
waiting for a reviewer to notice.

Applying the manifests remains an operator action; nothing here was applied.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nb7Q6ZmXppNDkTWytfYqfv

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 2069992@bnt-lap001
Assistant-Session: 167dd7f8-2a25-4be1-aa46-3b6f1a5f94c6
This commit is contained in:
tegwick 2026-09-10 16:39:20 +02:00
parent 3c2cdcdf79
commit de9e3abe5f
9 changed files with 666 additions and 8 deletions

163
deploy/attest-cronjob.yaml Normal file
View file

@ -0,0 +1,163 @@
# Chain-head attestation on a schedule (AUDIT-WP-0009-T02).
#
# T01 made `tamper_evidence` conditional on a fresh attestation. Until this
# runs, the honest answer in production is `false` — the precondition is simply
# not met. This is the job that meets it.
#
# WHERE THE ATTESTATION GOES, AND WHY IT MATTERS MORE THAN THE SCHEDULE.
# An attestation is worth exactly as much as its independence from the thing it
# attests. Two copies, with different properties, and neither is optional:
#
# 1. ConfigMap `audit-core-chain-head`, in-cluster, written here and mounted
# read-only by the Deployment. Outside platform-pg, so a database owner
# who rewrites a suffix cannot also rewrite the attestation without
# separate cluster access. This is the copy the receiver reads and the one
# that makes /readyz truthful.
#
# 2. The logical-offsite copy (rapp-postgres / Nextcloud + age, the path
# RESOURCE-WP-0002-T06 already uses). Survives loss of the cluster. This
# job does NOT write it — audit-core holds no Nextcloud credential and
# should not — so it remains an operator step, recorded in
# docs/integrity.md.
#
# NOT the Barman prefix. A copy restored alongside the table proves nothing:
# whoever rewrote the table restores the attestation that matches it.
#
# So copy 1 alone is a real but bounded control: it defends against a database
# owner, not against a cluster owner. docs/integrity.md states that bound; do
# not let this file be read as delivering more.
#
# Image digest must match deploy/audit-core.yaml.
---
apiVersion: v1
kind: ServiceAccount
metadata:
name: audit-core-attest
namespace: audit-core
labels:
app.kubernetes.io/name: audit-core
app.kubernetes.io/component: attest
---
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: audit-core-attest
namespace: audit-core
rules:
# One named ConfigMap, in this namespace, and no `create` on the collection.
# The ConfigMap is created once by the operator; this job may only replace
# its contents. Deliberately not `list` — a job that can enumerate the
# namespace's ConfigMaps has more reach than writing one head needs.
- apiGroups: [""]
resources: ["configmaps"]
resourceNames: ["audit-core-chain-head"]
verbs: ["get", "update", "patch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: audit-core-attest
namespace: audit-core
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: Role
name: audit-core-attest
subjects:
- kind: ServiceAccount
name: audit-core-attest
namespace: audit-core
---
apiVersion: batch/v1
kind: CronJob
metadata:
name: audit-core-attest-chain
namespace: audit-core
labels:
app.kubernetes.io/name: audit-core
app.kubernetes.io/component: attest
spec:
# Daily, against the 168h freshness window declared in docs/integrity.md.
# Seven cadences of headroom on purpose: a missed run degrades the claim
# gradually rather than flapping tamper_evidence false on one bad night.
schedule: "17 3 * * *"
timeZone: Etc/UTC
concurrencyPolicy: Forbid
startingDeadlineSeconds: 3600
successfulJobsHistoryLimit: 3
failedJobsHistoryLimit: 7
jobTemplate:
spec:
backoffLimit: 2
ttlSecondsAfterFinished: 172800
template:
metadata:
labels:
app.kubernetes.io/name: audit-core
app.kubernetes.io/component: attest
spec:
restartPolicy: OnFailure
serviceAccountName: audit-core-attest
securityContext:
runAsNonRoot: true
runAsUser: 10001
fsGroup: 10001
seccompProfile:
type: RuntimeDefault
containers:
- name: attest
image: forgejo.coulomb.social/coulomb/audit-core@sha256:c2fe39a0185b99be3fc0cb14d2de69772b8e66e20490097c9d11d90cc39719a6
imagePullPolicy: IfNotPresent
command:
- python
- -m
- audit_core.attest_publish
# attest_publish runs `attest-chain`, refuses to publish over a
# broken chain, and PATCHes the ConfigMap through the API with
# the projected ServiceAccount token. Written in stdlib rather
# than shelling out because this image carries no kubectl, and
# adding one to an audit receiver's image to write one file is a
# worse trade than twenty lines of urllib.
env:
- name: AUDIT_CORE_CREDENTIAL_DIR
value: /etc/audit-core/db
- name: AUDIT_CORE_DATABASE_SCHEMA
value: audit_core
# Read-only work: never migrate from the attestation job.
- name: AUDIT_CORE_AUTO_MIGRATE
value: "0"
resources:
requests: {cpu: 50m, memory: 64Mi}
limits: {cpu: 500m, memory: 256Mi}
securityContext:
allowPrivilegeEscalation: false
capabilities:
drop: ["ALL"]
readOnlyRootFilesystem: true
volumeMounts:
- name: tmp
mountPath: /tmp
- name: database-credential
mountPath: /etc/audit-core/db
readOnly: true
volumes:
- name: tmp
emptyDir: {}
- name: database-credential
secret:
secretName: audit-core-database
defaultMode: 0440
---
# Created empty by the operator so the CronJob's Role needs no `create`, and
# so the Deployment can mount it before the first run. An absent or undated
# attestation degrades the claim rather than breaking the receiver — that is
# T01's `no_attestation` path, and it is the correct behaviour on day one.
apiVersion: v1
kind: ConfigMap
metadata:
name: audit-core-chain-head
namespace: audit-core
labels:
app.kubernetes.io/name: audit-core
app.kubernetes.io/component: attest
data:
chain-head.json: "{}"

View file

@ -66,6 +66,12 @@ spec:
metadata:
labels:
app.kubernetes.io/name: audit-core
# Distinguishes receiver pods from the attestation CronJob's pods,
# which share the name label. NetworkPolicy audit-core-egress is
# scoped to this value so the receiver never gains the attest job's
# API-server reach (AUDIT-WP-0009-T02). Not in the Deployment's
# selector, which is immutable and does not need it.
app.kubernetes.io/component: receiver
spec:
securityContext:
runAsNonRoot: true
@ -127,6 +133,17 @@ spec:
# below deploy/senders-scope.json.
- name: AUDIT_CORE_SENDERS_SCOPE_PATH
value: /etc/audit-core/senders-scope.json
# Chain-head attestation, written by the audit-core-attest CronJob
# and mounted read-only here (AUDIT-WP-0009-T02). The receiver
# reads it and never writes it: a receiver that could rewrite its
# own attestation could forge it, which is why the job is a
# separate workload with a separate identity.
#
# Absent, empty or stale degrades tamper_evidence rather than
# failing the pod — that is T01's intended behaviour and the
# correct state before the first run.
- name: AUDIT_CORE_ATTESTATION_PATH
value: /etc/audit-core/attestation/chain-head.json
resources:
requests:
cpu: 50m
@ -153,6 +170,16 @@ spec:
mountPath: /etc/audit-core/senders-scope.json
subPath: senders-scope.json
readOnly: true
# Mounted as a DIRECTORY, deliberately, and not with subPath.
# A subPath ConfigMap mount is resolved once at pod start and
# never updates — the daily attestation would land in the
# ConfigMap and never reach the running receiver, so
# tamper_evidence would age out to false while the job reported
# success every night. A directory mount is updated in place by
# the kubelet, and the backend re-reads the file per check.
- name: chain-head
mountPath: /etc/audit-core/attestation
readOnly: true
startupProbe:
httpGet: {path: /healthz, port: http}
periodSeconds: 3
@ -188,3 +215,9 @@ spec:
configMap:
name: audit-core-senders-scope
defaultMode: 0444
- name: chain-head
configMap:
name: audit-core-chain-head
defaultMode: 0444
# optional: the pod must start before the first attestation exists.
optional: true

View file

@ -171,6 +171,42 @@ spec:
---
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: audit-core-attest-egress
namespace: audit-core
spec:
# Scoped to the attestation job by component label, so the receiver itself
# gains nothing from this rule. The receiver must not be able to reach the
# API server: a compromised receiver that could rewrite the chain-head
# ConfigMap could forge its own attestation, which is the one thing the
# separation of these two workloads exists to prevent.
podSelector:
matchLabels:
app.kubernetes.io/name: audit-core
app.kubernetes.io/component: attest
policyTypes: [Egress]
egress:
- to:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: databases
ports:
- {protocol: TCP, port: 5432}
- to:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: kube-system
ports:
- {protocol: UDP, port: 53}
- {protocol: TCP, port: 53}
# kube-apiserver. On this single-node k3s cluster the API server is the
# host itself, so this is a host-network destination rather than a pod
# selector; narrow it to the API port.
- ports:
- {protocol: TCP, port: 6443}
---
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: audit-core-egress
namespace: audit-core
@ -178,6 +214,7 @@ spec:
podSelector:
matchLabels:
app.kubernetes.io/name: audit-core
app.kubernetes.io/component: receiver
policyTypes: [Egress]
egress:
# PostgreSQL custody store. This is the only destination the receiver needs;