TEN-WP-0009-T01: record rapp-postgres acceptance, add postgres-client label

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
tegwick 2026-08-17 10:59:32 +02:00
parent 05d12fe39f
commit 5a9e78f657
2 changed files with 134 additions and 3 deletions

View file

@ -4,6 +4,11 @@ metadata:
name: tenant-engine
labels:
net-kingdom/component: tenant-engine
# TEN-WP-0009: rapp-postgres' consumer-ingress NetworkPolicy selects on
# this label, not on the declaration's spec.clientNamespaces -- that field
# is review metadata, this label is enforcement. Without it the pod cannot
# reach 5432 and the failure looks like DNS or a bad credential.
railiance.io/postgres-client: platform-pg
---
apiVersion: v1
kind: PersistentVolumeClaim

View file

@ -205,9 +205,130 @@ Done when `rapp-postgres` has provisioned the consumer, the connection reaches
the app only through an injected, refreshable credential, and the
move-to-dedicated trigger list and runbook are recorded.
Status 2026-08-17: declaration requested from `rapp-postgres`; awaiting their
provisioning and an answer on how a workload consumes the broker lease in
practice. Nothing to apply on our side until that returns.
Status 2026-08-17: **accepted.** `consumers/tenant-engine.yaml` merged at
`deda11e`. Provisioning is operator-gated as RAPP-IN-0004; nothing applied yet.
**They caught a real defect in our proposal.** The declaration had a single
`statementTimeout` field that the renderer applied to the migration role, the
runtime role, *and* the database — so our 5s would have given us 5s DDL and
failed the first non-trivial migration. Rather than reject it they made it
expressible, adding `limits.migrationStatementTimeout` (defaulting to
`statementTimeout`, so audit-core's rendered SQL is byte-identical). Ours is
now `statementTimeout: 5s` / `migrationStatementTimeout: 30s`. The 5s runtime
reasoning was accepted as-is.
Mechanical detail worth keeping: the *database*-level `statement_timeout` is
what a leased login actually gets, because group-role settings do not pass
through membership and OpenBao vends a fresh login role that merely inherits
`tenant_engine_migrate`. The 30s therefore has to be set on the leased login in
the migration role's creation statement. That lands in their OpenBao config —
but if a migration ever dies at 5s, that is where to look.
`tenantKeyingRequired: true` confirmed correct. The field asserts our data is
tenant-keyed from the first migration; nothing inspects the key's shape, and
the isolation probes run against a synthetic consumer, so there was no
audit-core-shaped assumption to trip. They queued a probe run with our consumer
present anyway, on the grounds that "we do not think it will false-positive" is
not evidence — which is the right standard.
### Answers to the three open questions
**Q1 — lease delivery: the credential *directory* is sanctioned.** A
`ClusterSecretStore` over OpenBao's database secrets engine projects the lease
to a mounted path. **Read the credential at connection checkout, not once at
startup.** Their rotation is overlap-first — acquire and verify the
replacement, switch the pool, drain the old, revoke the predecessor — so a
consumer that re-reads per connection rotates with no restart, while one that
takes the lease as process environment needs a restart every rotation. They
require that to be an explicit per-consumer choice; we take the directory.
There is also a broker `exec` path injecting `PG*` vars into a child process
(`railiance-platform/scripts/credential.py`) — for reviewed migrations and
operator sessions, not the runtime pod.
Theirs to create: `ClusterSecretStore openbao-tenant-engine`, and OpenBao roles
`rapp-postgres/tenant-engine-{runtime,migration}` — bundled with audit-core's
outstanding RAPP-IN-0003.
**Ours to do:** label the namespace
`railiance.io/postgres-client=platform-pg`. Their consumer-ingress
NetworkPolicy selects on that label, *not* on `spec.clientNamespaces` — the
declaration field is review metadata, the label is enforcement. Done
declaratively in `deploy/tenant-engine.yaml`'s Namespace document rather than
by an imperative `kubectl label`, so it cannot drift out of the repo.
**Q2 — out-of-band migration is the contract, not a preference.** The runtime
role gets DML only and owns nothing; migrate-on-startup would require DDL on
the runtime credential and collapse the separation. So our idempotent startup
migration becomes a step run under a `tenant-engine-migration` lease *before*
the new image serves traffic. Keep it idempotent regardless — that property is
what makes a move to dedicated capacity a restore rather than a rebuild.
**Q3 — backup is inherited and now real:** continuous WAL archiving to Scaleway
S3 (gzip), daily base backup at 02:15, 30-day retention, plus an encrypted
logical dump lane to Nextcloud without PITR.
No certified RPO. Restore and PITR are proven against a scratch cluster and
scratch object store, not the production `serverName`, so the honest status is
"designed for minutes, not yet measured". They volunteered that rather than
quoting a number they had not earned.
### Two things to push back on
**1. Dedicated capacity does not mitigate §5 for us.** They offer a dedicated
instance as the escalation for a consumer that cannot accept ADR-0001 §5's
residual risk — a leaked `tenant_engine_app` credential reading and writing
every row we own. For us that escalation does not work: a dedicated instance
changes co-residency, not what our own runtime credential can reach. The blast
radius is identical either way. So we should *not* argue for dedicated on §5
grounds; it would buy a different guarantee than the one at risk.
What §5 actually means here is worse than for `audit-core`, and worth stating
plainly: direct DML bypasses every control this service exists to enforce —
flex-auth authorization, version CAS, the guardrail loosening guard, and the
audit trail. Someone with the runtime credential could grant `PLTF` on any
tenant, or raise any ceiling, with no authorization check and no audit event.
Since `tenant-engine` is the source of the `tenant_roles` claim `key-cape`
mints and `flex-auth` reads live, that is privilege escalation across
NetKingdom, not just data tampering. §4's opt-in RLS does not help: our rows
*are* the tenants, so there is no per-tenant predicate to scope by.
The real mitigations are lease lifetime and network reach, both of which they
provide — plus one we do not have, below.
**2. Our audit trail is only as trustworthy as our own database.** `events`
lives in the same database as the rows it attests to, so the §5 scenario
forges the evidence along with the data. That is a genuine gap, and the fleet
already has the answer: `audit-core` exists as an append-only sink, and
`AUDIT_CORE_SENDERS` implies a registration path. Emitting there would make our
trail tamper-evident independently of our store.
**Out of scope for this workplan** — it is a new cross-service dependency, not
a store migration — but it should not be discovered later. Raised with
`rapp-postgres`; needs its own workplan.
### RTO answer they asked for
They asked us to say now if our RTO on the authorization path cannot absorb a
single-consumer restore (a logical dump taken from a scratch-restored copy of
the whole instance, then a controlled import).
Split the cases, because the answer differs:
- **Whole-instance loss:** we are restored alongside everyone else. Acceptable,
and dedicated capacity would make it *worse* by adding a second instance to
recover.
- **tenant-engine-only corruption** — a bad migration, or the §5 credential
scenario — needs the slow single-consumer path. This is the case that matters,
and it is sharpened by our being the authority: there is no upstream to replay
from. If our data is gone or forged, nothing else in the fleet can reconstruct
it.
So the ask is not dedicated capacity but a **per-consumer logical dump
cadence**. They already run an encrypted logical dump lane to Nextcloud; if it
covers per-database dumps, our single-consumer restore comes from there rather
than from a scratch-restored instance, and the slow path stops being the only
path. Asked.
## T02 - Implement `PostgresTenantStore`
@ -377,6 +498,11 @@ Done when the notes are sent and TEN-WP-0008-T02 is updated.
- **Fleet placement policy.** Which element or tenant gets dedicated capacity,
and how isolation maps to plan tiers, is a platform decision. This workplan
makes `tenant-engine` movable and stops there.
- **Emitting the audit trail to `audit-core`.** Identified under T01 as a real
gap — our `events` table shares a database with the rows it attests to, so
ADR-0001 §5 forges both together. It is a new cross-service dependency rather
than part of a store migration, and it needs its own workplan. Recorded here
so it is not rediscovered by accident.
- **Per-tenant isolation.** A future plan tier may sell dedicated
infrastructure. `tenant-engine` already records plan assignment by id, and
`adaptive-pricing` owns what a plan *means*, so neither the tier definition