audit-core/docs/audit-backend-contract.md
tegwick ded432a63f
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 2s
Implement AUDIT-WP-0006 honest operational custody.
Postgres now reports custody_class=operational with a cited 30-day
recoverable window. Join ITC-CAP operations.audit at D4, publish the
interface card, and overlay user-engine tenants [*] from Git so an
ExternalSecret refresh cannot shrink it.
2026-08-16 00:24:33 +02:00

13 KiB

Audit Backend Contract

Audit Core separates event producers (integrations, CLI, future HTTP API) from audit backends (sinks that persist normalized events). This document defines the replaceable backend interface, the current event schema, retention guarantees, and how to migrate from the development mock file backend to durable custody.

AuditBackend protocol

Every backend implements audit_core.interface.AuditBackend:

class AuditBackend(Protocol):
    @property
    def retention_policy(self) -> RetentionPolicy: ...

    def emit(self, event: AuditEvent) -> str: ...

emit(event) -> str

Persist one normalized AuditEvent and return a backend-specific reference. Callers use the reference for debugging and correlation; it is not a stable cross-backend identifier.

Backend kind Typical return value
Mock file Absolute path to the hourly JSONL file
Archive (planned) Batch URI or object key
Hot search (planned) Stream offset or index document id

Requirements:

  • Idempotent references: Re-emitting the same logical event (same event_id) must not corrupt prior records. Backends may append duplicates unless deduplication is documented.
  • No secret dumping: Backends must not log or persist plaintext secrets, tokens, keys, or passwords from details or future payload fields.
  • Redaction is recorded, not silent: where ingestion masks a field, the stored record says so (see Secret-shaped fields).
  • Visible failure: Raise on persistence failure; do not silently drop events.
  • Normalization only: Backends receive AuditEvent instances. Source-specific adapters run upstream.

retention_policy

Each backend exposes a frozen RetentionPolicy describing what it guarantees. Integrators and readiness checks use this to decide whether a sink satisfies a scope's policy. See Retention policy.

Event schema (audit-core.event.v1alpha1)

The first implementation uses a flat JSON record produced by AuditEvent.as_record(). This is a deliberate simplification for local wiring; the long-term envelope in spec/ProductRequirementsDefinition.md (audit-core.event.v1) nests source, tenant, scope, actor, and result objects.

Required fields

Field Type Description
schema_version string Must be audit-core.event.v1alpha1 for this contract
event_id string UUID or source-stable identifier
observed_at string UTC ISO-8601 timestamp (no subsecond precision)
tenant string Tenant id; default platform for control-plane events
scope string Scope id; default platform-control-plane
source string Emitter id (e.g. openbao, audit-core)
action string Namespaced action (e.g. openbao.audit.list)
resource string Affected resource path or id
outcome string Result label (e.g. success, failure, denied)

Optional fields

Field Type Description
actor string or null Subject performing the action
reason string or null Human-readable result explanation
details object Source-specific extension map; must not contain secrets. May carry a redaction entry — see below

Example record

{
  "action": "openbao.authenticated_readiness_proof",
  "actor": null,
  "details": {"backend": "mock-file", "file_audit_visible": true},
  "event_id": "6f3e2b1a-4c5d-6e7f-8a9b-0c1d2e3f4a5b",
  "observed_at": "2026-06-01T20:30:00+00:00",
  "outcome": "success",
  "reason": null,
  "resource": "openbao/openbao-0",
  "schema_version": "audit-core.event.v1alpha1",
  "scope": "platform-control-plane",
  "source": "openbao",
  "tenant": "platform"
}

Validation

Use audit_core.interface.validate_event() before emit. It checks required fields, schema_version, and rejects empty strings on required identifiers.

Evolution to audit-core.event.v1

Future backends will accept the nested v1 envelope at the HTTP ingestion layer. Module-level backends may continue using AuditEvent with adapter translation. Compatibility rules (not yet implemented):

  • v1alpha1 records remain readable in archive export.
  • New required v1 fields get sensible defaults during adapter migration.
  • schema_version gates parser selection.

Retention policy

RetentionPolicy is declarative metadata; enforcement is backend-specific.

Field Meaning
custody_class development, operational, archive, or hot_search
retention_days Maximum age before eligible deletion; None is a lifecycle statement (no expiry), not a recovery guarantee
immutable Whether stored records are protected from in-place alteration
tamper_evidence Whether manifests, hash chains, or signatures exist
durable Whether survival is expected across process restarts and host reboots
recoverable_days Cited platform backup window; None if not declared
recoverable_source Where the recoverable window is cited from
recoverable_basis ITC-GOV EvidenceBasis of that citation (measured, quoted, …)

Custody classes

Class Purpose Guarantees
development Local integration and bootstrap wiring Ephemeral local files; best-effort cleanup; not audit custody
operational Durable production custody Append-only Postgres; recoverable through the platform data.backup provision; not ITC-CAP data.archive
archive Long-term evidence (future sink) Reserved for a backend that can satisfy data.archive hooks (retention policy, integrity verification, retrieval test)
hot_search Operational investigation (planned) Shorter retention; searchable; not the evidence record

AUDIT_CORE_REQUIRE_CUSTODY_CLASS=operational is the production fail-closed gate. archive is accepted as an alias of operational for one mixed rollout so an old manifest cannot refuse a new image. A development backend satisfies neither.

Mock file backend policy

MockFileAuditBackend.retention_policy:

  • custody_class: development
  • retention_days: 7 (override via AUDIT_CORE_MOCK_RETENTION_DAYS or constructor)
  • immutable: false
  • tamper_evidence: false
  • durable: false

Enforcement:

  • Events append to hourly JSONL files under AUDIT_CORE_MOCK_DIR (default /tmp/audit-core).
  • cleanup_old_files() deletes audit-*.jsonl whose mtime is older than retention_days. Cleanup runs on each emit and via python3 -m audit_core cleanup.
  • Negative retention_days disables automatic deletion.

Not guaranteed: crash-safe writes, replication, encryption, tenant isolation, integrity proofs, or survival of /tmp across reboots.

Production operational policy (Postgres, live)

PostgresAuditBackend.retention_policy:

  • custody_class: operational
  • retention_days: unset in production (the service does not expire rows)
  • immutable: true (trigger events_append_only; not a claim against the database owner)
  • tamper_evidence: false (a superuser can drop the trigger; no hash-chain)
  • durable: true
  • recoverable_days: 30, cited from the platform data.backup provision
  • recoverable_source: resource-control/data/capability/platform-audit-storage.json#provisions[capability=data.backup]
  • recoverable_basis: measured

None retention is a lifecycle statement, not unbounded archive. Rows older than the recoverable window are not promised after a restore. A future archive backend that satisfies ITC-CAP data.archive is not implemented.

Migration path: mock file → durable backend

Phase 0 — today (mock file)

Use MockFileAuditBackend or the CLI:

python3 -m audit_core emit \
  --source my-service \
  --action my_service.audit.smoke \
  --resource my-service/instance-0 \
  --outcome success

Integrations import the protocol, not the mock:

from audit_core import AuditBackend, AuditEvent, MockFileAuditBackend

backend: AuditBackend = MockFileAuditBackend()
backend.emit(AuditEvent(source="...", action="...", resource="...", outcome="success"))

Phase 1 — dual-write readiness (planned)

  1. Register a durable archive backend implementing AuditBackend.
  2. Configure routing: development scopes may keep mock; production scopes require custody_class=operational (the live Postgres backend).
  3. Readiness checks compare backend.retention_policy against scope policy and fail closed when custody is insufficient.

Phase 2 — operational primary (live)

  1. HTTP ingestion writes through PostgresAuditBackend (custody_class=operational).
  2. Retain mock only for local make mock-audit-smoke and unit tests.
  3. A future data.archive sink is a separate backend, not a rename of Postgres.

Phase 3 — hot search adjunct (planned)

Add a second AuditBackend with custody_class=hot_search for investigation. Archive remains the evidence record; hot search may use shorter retention_days.

Code migration checklist

Step Action
1 Depend on AuditBackend, not MockFileAuditBackend, in integration code
2 Build AuditEvent with explicit tenant, scope, and source
3 Call validate_event() before emit
4 Inspect retention_policy in readiness gates
5 Replace mock construction with injected backend from configuration
6 Verify export/manifest workflow before decommissioning mock files

Reference implementations

Backend Module Custody class
Mock file JSONL audit_core.mock_file_backend.MockFileAuditBackend development
SQLite (local / test) audit_core.sqlite_backend.SQLiteAuditBackend development
PostgreSQL (production) audit_core.postgres_backend.PostgresAuditBackend operational
Archive (planned data.archive sink) TBD archive
Hot search (planned) TBD hot_search
  • INTENT.md — product purpose and principles
  • spec/ProductRequirementsDefinition.md — full v1 envelope and API requirements
  • registry/capabilities/capability.audit.event-retain.md — capability registry entry

Secret-shaped fields

Ingestion detects fields whose key name contains password, secret, token, credential, or private_key, at any depth in the payload. Values are not inspected: a value-shape heuristic produces false positives on legitimate identifiers, and a false positive here silently mangles an audit record.

Policy

The default is redact and accept. Rejecting an otherwise valid event because of one field loses the audit record entirely, which is a worse outcome than storing it with that field masked.

Policy is set per sender identity via secret_policy in AUDIT_CORE_SENDERS, so a higher-assurance channel can be switched to reject without changing the posture for every other sender:

[{"name": "user-engine", "tokens": ["..."], "sources": ["user-engine"],
  "secret_policy": "redact"},
 {"name": "payments-engine", "tokens": ["..."], "sources": ["payments-engine"],
  "secret_policy": "reject"}]
Policy Response Effect
redact (default) 202 / 200 Value replaced with [redacted]; key preserved; event stored
reject 400 secret_shaped_field Event not stored; dead-lettered with its payload withheld

Keys are preserved under redaction. Dropping them would hide the fact that the sender transmitted the field at all — which is exactly what an operator needs in order to stop it.

Recorded redaction

A redacted record carries the fact in details.redaction, so a reader never has to infer whether what they are looking at is what the sender sent:

"details": {
  "correlation_id": "corr-1",
  "data": {"membership_id": "m-1", "auth_token": "[redacted]"},
  "redaction": {"policy": "redact", "paths": ["data.auth_token"]}
}

Idempotency is unaffected: the payload hash is taken over the original request body, so a resubmission of the same original is still recognised as a duplicate and redaction is deterministic.

Counting

Both outcomes are counted durably, aggregated by sender, source, action, and field path — because the actionable unit is "stop emitting data.auth.token on membership.added", not "there were 47 redactions". Counters survive restart, since the fix they drive lives in another service.

Read them at GET /v1/secret-findings (requires the read privilege):

{"secret_findings": [
  {"sender": "user-engine", "source": "user-engine",
   "action": "membership.added", "field_path": "data.auth_token",
   "outcome": "redacted", "persisted": true, "occurrences": 3,
   "first_seen": "...", "last_seen": "..."}]}

occurrences counts transmissions, not stored events: a retry resubmitting the same secret-shaped field increments it again, even though the event reconciles as a duplicate and produces no second custody record. That is deliberate — the number measures how often the sender emitted the field, which is the behaviour being optimized away.

persisted distinguishes a field that reached the stored record from one that sat elsewhere in the envelope and was dropped by normalization anyway. A healthy sender trends to zero occurrences; a non-empty list is a backlog item for the sending service, not a steady state.