Corrects the credential assumption: shared-cluster access is a broker lease, not a CNPG-minted secretKeyRef. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
384 lines
17 KiB
Markdown
384 lines
17 KiB
Markdown
---
|
|
id: TEN-WP-0009
|
|
type: workplan
|
|
title: "PostgreSQL as the production store, SQLite for dev and test"
|
|
domain: infotech
|
|
repo: tenant-engine
|
|
status: ready
|
|
owner: claude
|
|
topic_slug: tenant-lifecycle
|
|
created: "2026-08-16"
|
|
updated: "2026-08-16"
|
|
depends_on:
|
|
- TEN-WP-0007
|
|
unblocks:
|
|
- TEN-WP-0008
|
|
state_hub_workstream_id: "3bf3b4ab-a116-4914-b2d8-d3c430754956"
|
|
---
|
|
|
|
# TEN-WP-0009 - PostgreSQL production store
|
|
|
|
Move the production store from SQLite-on-a-PVC to the fleet's CloudNativePG
|
|
infrastructure. SQLite stays, deliberately, as the dev and test backend.
|
|
|
|
## Why this reverses an earlier decision
|
|
|
|
TEN-WP-0005-T02 chose SQLite over PostgreSQL on purpose, and the reasoning was
|
|
sound at the time: TEN-WP-0004 had shipped SQLite on a PVC, so *"adding an
|
|
unused Postgres path would have been dead code"*. The `TenantStore` Protocol
|
|
was kept precisely as the seam for this change.
|
|
|
|
Three things have changed since:
|
|
|
|
1. **The PVC is now blocking work.** The volume is `ReadWriteOnce`, which forces
|
|
`strategy: Recreate` and makes a side-by-side canary impossible — the
|
|
constraint TEN-WP-0008-T02 ran into when onboarding to staged promotion. The
|
|
store choice is no longer invisible to anything outside this repo.
|
|
2. **The infrastructure is already there and already proven.** The
|
|
CloudNativePG operator runs on railiance01 with seven clusters, and
|
|
`user-engine` — our sibling service, sharing the `tenant_id` key — has run
|
|
its own `user-engine-pg` for 19 days.
|
|
3. **A Postgres path is no longer unused.** It becomes the production path, so
|
|
the dead-code objection no longer applies.
|
|
|
|
SQLite is *not* being removed. It stays as the dev and test backend, where a
|
|
zero-setup file-backed store is genuinely the right tool, and it keeps the
|
|
conformance suite honest by forcing the Protocol to stay a real seam rather
|
|
than a formality wrapped around one implementation.
|
|
|
|
## Credential handling
|
|
|
|
**Corrected 2026-08-17.** This section originally said CloudNativePG mints an
|
|
`X-app` secret we reference with `secretKeyRef`, as `user-engine` does. That is
|
|
true for a cluster's *own* app database — and wrong for the path we are
|
|
actually taking.
|
|
|
|
`railiance-platform/docs/rapp-postgres-boundary.md` is explicit: `rapp-postgres`
|
|
owns the role and database provisioning surface for the shared cluster, and
|
|
*"the workload receives a short-lived lease through the platform broker"* — the
|
|
OpenBao database secrets engine — rather than a static secret. `railiance-platform`
|
|
retains cluster-wide governance and the credential-broker grant catalog.
|
|
|
|
Consequences, and the second one is real work:
|
|
|
|
- **We do not provision our own database.** We *declare a consumer request* and
|
|
`rapp-postgres` provisions it. Writing a CNPG `Database` manifest into this
|
|
repo's `deploy/`, or adding a managed role to a shared cluster's spec, would
|
|
both be reaching across that boundary — the latter also mutating shared
|
|
infrastructure other consumers depend on.
|
|
- **Credentials are leased, not mounted.** A short-lived lease means the
|
|
connection string can change under a running pod, so the store cannot read a
|
|
DSN once at startup and hold it forever. T03 and T05 must account for
|
|
credential refresh; a long-lived pool built on an expired lease fails at the
|
|
worst possible moment.
|
|
|
|
Either way **no secret enters this repo, this workplan, or any evidence dump**.
|
|
Runtime secret custody stays OpenBao's per `SCOPE.md`, which is exactly what
|
|
the broker lease implements.
|
|
|
|
## Placement: shared by default, movable by design
|
|
|
|
Dedicated-vs-shared is **not a decision this workplan should settle once**, and
|
|
`tenant-engine` is the wrong place to settle it for the platform. The estate
|
|
has distinct elements — railiance, NetKingdom, HelixForge, Coulomb — plus
|
|
tenants layered on top, and which repo belongs to which grouping is not yet
|
|
organised. On top of that, isolation level is heading toward being a
|
|
*product* property: plans will differ in how much isolation they buy.
|
|
|
|
So the requirement here is **portability, not placement**:
|
|
|
|
- start on a **shared** cluster, because it is cheaper on a single node and
|
|
nothing yet justifies dedicated capacity;
|
|
- make moving to a dedicated cluster — or to a different shared one — an
|
|
**operational change, not a code change**;
|
|
- never let the decision leak into the application.
|
|
|
|
What that demands of this repo, and all of it is cheap if done now and
|
|
expensive to retrofit:
|
|
|
|
1. **Connect by injected URL only.** No cluster name, namespace, host, or
|
|
database name anywhere in `src/`. The application must be unable to tell
|
|
whether it is on a shared or dedicated cluster.
|
|
2. **Own a whole database, not a set of tables in someone else's.** Sharing a
|
|
*cluster* is a capacity decision; sharing a *database* would entangle
|
|
schemas and make relocation a merge rather than a move. A dedicated
|
|
database inside a shared cluster relocates with a dump and a restore.
|
|
3. **No cross-database joins or co-location assumptions**, which the repo
|
|
boundary already forbids — `tenant_id` is the only key shared with
|
|
`user-engine`. Placement must not quietly become a dependency.
|
|
4. **Idempotent schema creation**, so a fresh target comes up correct without a
|
|
hand-built database.
|
|
|
|
Then "move `tenant-engine` to dedicated" is: create the target cluster, dump,
|
|
restore, swap the `secretKeyRef`, restart. No rebuild, no release, no code
|
|
review.
|
|
|
|
**What this repo does not own:** the placement *policy* — which element or
|
|
tenant gets dedicated capacity and when. That is a platform concern, and
|
|
`tenant-engine` implementing its own would be the same boundary error as
|
|
building a fleet drift-detector because we got bitten. Raised with
|
|
`railiance-platform` separately; see T01.
|
|
|
|
## T01 - Make placement portable and provision on shared capacity
|
|
|
|
```task
|
|
id: TEN-WP-0009-T01
|
|
status: todo
|
|
priority: high
|
|
state_hub_task_id: "b4701256-f235-4474-9dcf-7cb09f62b873"
|
|
```
|
|
|
|
**Target: `platform-pg`, via a `PostgresConsumer` declaration to
|
|
`rapp-postgres`.** Not `net-kingdom-pg` or `apps-pg`, and not provisioned by us.
|
|
|
|
Reasoning, revised 2026-08-17 after finding the actual contract. The candidate
|
|
list originally weighed product-family fit and would have picked
|
|
`net-kingdom-pg`. But `platform-pg` is the shared cluster that has a
|
|
*governed provisioning path* — `rapp-postgres` owns its consumer declarations,
|
|
role provisioning, isolation tests, and recovery procedure. The other clusters
|
|
carry their roles as hand-added entries in the shared cluster spec
|
|
(`net-kingdom-pg` has `privacyidea`, `apps-pg` has `vergabe` and
|
|
`coulomb_social`). Choosing product-family fit would mean improvising outside
|
|
the one contract that exists, to gain an affinity that portability makes cheap
|
|
to change later anyway. `audit-core` is the precedent: a platform service, on
|
|
`platform-pg`, declared as a consumer.
|
|
|
|
Note the deviation from `user-engine` and why: they run a dedicated
|
|
`user-engine-pg`. We are deliberately *not* copying that. Their choice is fine
|
|
and ours is reversible; picking dedicated now would spend capacity on an
|
|
isolation guarantee nothing has asked for yet.
|
|
|
|
The declaration to request, modelled on `consumers/audit-core.yaml`:
|
|
|
|
```yaml
|
|
apiVersion: rapp-postgres.railiance.io/v1alpha1
|
|
kind: PostgresConsumer
|
|
metadata:
|
|
name: tenant-engine
|
|
spec:
|
|
database: tenant_engine
|
|
schema: tenant_engine
|
|
costAttributionKey: platform:tenant-engine
|
|
clientNamespaces: [tenant-engine]
|
|
roles:
|
|
owner: tenant_engine_owner
|
|
migration: tenant_engine_migrate
|
|
runtime: tenant_engine_app
|
|
limits:
|
|
migrationConnections: 2
|
|
runtimeConnections: 12
|
|
statementTimeout: 30s
|
|
idleInTransactionSessionTimeout: 15s
|
|
tenantKeyingRequired: true
|
|
tenantIsolation: consumer-service-boundary
|
|
```
|
|
|
|
Two things to raise with them rather than assume:
|
|
|
|
- **`tenantKeyingRequired: true`** is right for `audit-core`, whose rows are
|
|
per-tenant. Ours are *about* tenants — `tenant_id` is the primary key of the
|
|
`tenants` table, not a partition key on someone else's data. Confirm that
|
|
their isolation tests read our shape correctly rather than flagging a false
|
|
positive.
|
|
- **Statement timeout.** 30s is generous for this workload; every query here is
|
|
a single-row lookup or a small transaction. A tighter timeout is a better
|
|
failure mode for a service `flex-auth` calls synchronously on the
|
|
authorization path — a slow query should fail closed fast, not hold the PDP.
|
|
|
|
Record what would trigger a move to dedicated, so the reversal is a judgement
|
|
already made rather than one improvised under pressure. Candidate triggers:
|
|
a noisy neighbour affecting the authorization path, a compliance or residency
|
|
requirement, a plan tier that sells isolation, or the shared cluster's backup
|
|
policy no longer fitting.
|
|
|
|
Also decide **backups**: whether the shared cluster's existing
|
|
`scheduledbackups` policy covers us, or we need our own. SQLite-on-a-PVC had no
|
|
backup story at all, so this is a gain to claim explicitly rather than inherit
|
|
by accident — and a shared cluster means inheriting *someone else's* retention
|
|
choice, which is worth checking rather than assuming.
|
|
|
|
**Do not write a CNPG `Database` or `Cluster` manifest into this repo's
|
|
`deploy/`.** That surface belongs to `rapp-postgres`. Our `deploy/` gains only
|
|
the consumption side — how the workload reaches the leased credential.
|
|
|
|
Done when `rapp-postgres` has provisioned the consumer, the connection reaches
|
|
the app only through an injected, refreshable credential, and the
|
|
move-to-dedicated trigger list and runbook are recorded.
|
|
|
|
Status 2026-08-17: declaration requested from `rapp-postgres`; awaiting their
|
|
provisioning and an answer on how a workload consumes the broker lease in
|
|
practice. Nothing to apply on our side until that returns.
|
|
|
|
## T02 - Implement `PostgresTenantStore`
|
|
|
|
```task
|
|
id: TEN-WP-0009-T02
|
|
status: todo
|
|
priority: high
|
|
state_hub_task_id: "d7428bc1-2e5f-4a4c-93f5-9368ebb691d5"
|
|
```
|
|
|
|
Implement the full `TenantStore` Protocol — including
|
|
`guardrail_overrides` / `guardrail_changes` / `set_guardrail_override` from
|
|
TEN-WP-0006 — and add PostgreSQL to the conformance suites, which are already
|
|
parametrised over in-memory and SQLite. A third backend should be a fixture
|
|
parameter, not a new test file.
|
|
|
|
Invariants that must hold identically, not approximately:
|
|
|
|
- `mutate_tenant()` commits idempotency replay, version CAS, mutation, and
|
|
audit event **in one transaction**;
|
|
- `set_guardrail_override()` does the same for the override, its audit record,
|
|
and its receipt;
|
|
- a replayed `Idempotency-Key` short-circuits *before* the version check;
|
|
- reads and writes never observe a row mid-transaction.
|
|
|
|
That last one is where PostgreSQL should simplify rather than complicate:
|
|
SQLite needed an explicit `RLock` around reads because one connection is shared
|
|
across request threads, and the concurrent-writer test caught a real defect
|
|
there. Postgres gives proper per-connection transaction isolation, so the lock
|
|
should disappear rather than be ported. **Do not transliterate the SQLite
|
|
implementation** — port the semantics and let the backend do what it is good
|
|
at. Use a connection pool and appropriate isolation, and prove the guarantees
|
|
with the same concurrent-writer test that caught the SQLite bug.
|
|
|
|
Adds a runtime dependency (`psycopg`); keep it out of the base install path if
|
|
SQLite-only dev is meant to stay dependency-light.
|
|
|
|
Done when all three backends pass the same conformance suites unchanged.
|
|
|
|
## T03 - Backend selection that fails closed
|
|
|
|
```task
|
|
id: TEN-WP-0009-T03
|
|
status: todo
|
|
priority: high
|
|
state_hub_task_id: "89c121c6-8bc5-4abe-80f0-d11406aeffc2"
|
|
```
|
|
|
|
`Settings` currently carries only `database_path`. Add a database URL and make
|
|
selection explicit:
|
|
|
|
- URL set → PostgreSQL;
|
|
- path set → SQLite;
|
|
- **both set → refuse to start.** Do not silently prefer one. An ambiguous
|
|
store configuration in production is exactly the class of silent
|
|
misconfiguration this repo fails closed on everywhere else, and picking a
|
|
winner would let a stale `TENANT_ENGINE_DATABASE_PATH` quietly shadow the
|
|
real database;
|
|
- neither set → in-memory, as today, which is correct for tests and wrong for
|
|
production.
|
|
|
|
Consider surfacing the active backend on `/health`. The last two production
|
|
incidents were invisible because the service looked fine from outside; "which
|
|
store am I actually using" is cheap to answer and expensive to guess.
|
|
|
|
Done when misconfiguration is a startup failure with a clear message, covered
|
|
by tests.
|
|
|
|
## T04 - Migrate production data
|
|
|
|
```task
|
|
id: TEN-WP-0009-T04
|
|
status: todo
|
|
priority: high
|
|
state_hub_task_id: "51a83eae-9e9f-4676-bc1d-9c2d884bac0d"
|
|
```
|
|
|
|
Move the live SQLite database to PostgreSQL. Currently small — two tenants
|
|
(`tenant:trial:portalcheck` active v1, `tenant:trial:ten-wp-0005-t05` retired
|
|
v5) plus the T04 disposables from TEN-WP-0007 — but small is not the same as
|
|
trivial.
|
|
|
|
Everything must move, not just the tenants:
|
|
|
|
- `tenants` with versions and lifecycle timestamps **preserved exactly** — a
|
|
reset version silently breaks every consumer holding an ETag;
|
|
- `grants` including revoked ones, because the trail is append-only and its
|
|
history is the audit record;
|
|
- `plans`;
|
|
- `events` — the audit correlation contract;
|
|
- `idempotency_receipts` — dropping these lets an in-flight retry double-apply
|
|
a mutation that already happened;
|
|
- `guardrail_overrides` and `guardrail_changes`.
|
|
|
|
SQLite is single-writer and the app writes to it, so the cutover must stop
|
|
writes rather than race them: scale to zero, export, import, verify, then start
|
|
against PostgreSQL. Verify by comparing row counts *and* spot-checking the two
|
|
known tenants' versions and lifecycle state, not by trusting the exporter.
|
|
|
|
Keep the SQLite file and its PVC until T05's soak passes. It is the rollback.
|
|
|
|
Done when a verified copy is live in PostgreSQL and the comparison evidence is
|
|
recorded here.
|
|
|
|
## T05 - Cut over the deployment
|
|
|
|
```task
|
|
id: TEN-WP-0009-T05
|
|
status: todo
|
|
priority: high
|
|
state_hub_task_id: "03aeae3b-5e0f-4c77-a763-787ae08078f5"
|
|
```
|
|
|
|
Update `deploy/tenant-engine.yaml`:
|
|
|
|
- drop the PVC and its volume mount;
|
|
- add `TENANT_ENGINE_DATABASE_URL` from the CNPG-minted `secretKeyRef`;
|
|
- switch `strategy: Recreate` → `RollingUpdate`, now that no `ReadWriteOnce`
|
|
volume forces serialization — **this is what unblocks TEN-WP-0008-T02's
|
|
canary**;
|
|
- extend the `NetworkPolicy` egress to the Postgres service, and no wider.
|
|
Ingress stays restricted to `user-engine`.
|
|
|
|
Then decide, deliberately, whether to raise `replicas` above 1. It becomes
|
|
possible here for the first time; possible is not the same as warranted, and a
|
|
second replica changes failure modes for a service on the authorization path.
|
|
|
|
Update `make verify-pin` if the expected shape changes, and re-run it.
|
|
|
|
Rollback: re-pin the previous digest and re-attach the PVC. State plainly that
|
|
this becomes lossy once writes have landed in PostgreSQL — after that point the
|
|
rollback is "restore from Postgres", not "swap back to the file".
|
|
|
|
Done when production runs on PostgreSQL, `make verify-pin` passes, and the
|
|
TEN-WP-0007 T04 verification list still passes end to end against the live
|
|
service.
|
|
|
|
## T06 - Tell the fleet, and close the loop on TEN-WP-0008
|
|
|
|
```task
|
|
id: TEN-WP-0009-T06
|
|
status: todo
|
|
priority: medium
|
|
state_hub_task_id: "e7adbc64-ad1b-4f28-bfd0-0b07034629db"
|
|
```
|
|
|
|
- `flex-auth`: availability characteristics of their data source changed. No
|
|
contract or action change.
|
|
- `user-engine`: no contract change either, but they run the same CNPG pattern
|
|
and may have operational lessons worth having before we repeat their mistakes.
|
|
- **TEN-WP-0008-T02**: record that the `ReadWriteOnce` canary constraint is
|
|
gone, so staged-promotion onboarding no longer has to describe a canary that
|
|
cannot run.
|
|
|
|
Done when the notes are sent and TEN-WP-0008-T02 is updated.
|
|
|
|
## Out of scope
|
|
|
|
- **Removing SQLite.** It stays as the dev and test backend, and as the thing
|
|
that keeps `TenantStore` an honest seam.
|
|
- **Schema redesign.** Port the existing schema; a store migration and a model
|
|
change at the same time makes any regression ambiguous.
|
|
- **Sharing a database with `user-engine`.** `tenant_id` is the only key the
|
|
two services share, and it stays that way — separate stores are the boundary,
|
|
not an implementation detail. Sharing a *cluster* is fine; sharing a
|
|
*database* is not.
|
|
- **Fleet placement policy.** Which element or tenant gets dedicated capacity,
|
|
and how isolation maps to plan tiers, is a platform decision. This workplan
|
|
makes `tenant-engine` movable and stops there.
|
|
- **Per-tenant isolation.** A future plan tier may sell dedicated
|
|
infrastructure. `tenant-engine` already records plan assignment by id, and
|
|
`adaptive-pricing` owns what a plan *means*, so neither the tier definition
|
|
nor the placement that implements it belongs here. Worth designing for —
|
|
which the portability principles above do — not worth building for.
|