580 lines
28 KiB
Markdown
580 lines
28 KiB
Markdown
---
|
|
id: TEN-WP-0009
|
|
type: workplan
|
|
title: "PostgreSQL as the production store, SQLite for dev and test"
|
|
domain: infotech
|
|
repo: tenant-engine
|
|
status: finished
|
|
owner: claude
|
|
topic_slug: tenant-lifecycle
|
|
created: "2026-08-16"
|
|
updated: "2026-08-21"
|
|
depends_on:
|
|
- TEN-WP-0007
|
|
unblocks:
|
|
- TEN-WP-0008
|
|
state_hub_workstream_id: "3bf3b4ab-a116-4914-b2d8-d3c430754956"
|
|
---
|
|
|
|
# TEN-WP-0009 - PostgreSQL production store
|
|
|
|
Move the production store from SQLite-on-a-PVC to the fleet's CloudNativePG
|
|
infrastructure. SQLite stays, deliberately, as the dev and test backend.
|
|
|
|
## Why this reverses an earlier decision
|
|
|
|
TEN-WP-0005-T02 chose SQLite over PostgreSQL on purpose, and the reasoning was
|
|
sound at the time: TEN-WP-0004 had shipped SQLite on a PVC, so *"adding an
|
|
unused Postgres path would have been dead code"*. The `TenantStore` Protocol
|
|
was kept precisely as the seam for this change.
|
|
|
|
Three things have changed since:
|
|
|
|
1. **The PVC is now blocking work.** The volume is `ReadWriteOnce`, which forces
|
|
`strategy: Recreate` and makes a side-by-side canary impossible — the
|
|
constraint TEN-WP-0008-T02 ran into when onboarding to staged promotion. The
|
|
store choice is no longer invisible to anything outside this repo.
|
|
2. **The infrastructure is already there and already proven.** The
|
|
CloudNativePG operator runs on railiance01 with seven clusters, and
|
|
`user-engine` — our sibling service, sharing the `tenant_id` key — has run
|
|
its own `user-engine-pg` for 19 days.
|
|
3. **A Postgres path is no longer unused.** It becomes the production path, so
|
|
the dead-code objection no longer applies.
|
|
|
|
SQLite is *not* being removed. It stays as the dev and test backend, where a
|
|
zero-setup file-backed store is genuinely the right tool, and it keeps the
|
|
conformance suite honest by forcing the Protocol to stay a real seam rather
|
|
than a formality wrapped around one implementation.
|
|
|
|
## Credential handling
|
|
|
|
**Corrected 2026-08-17.** This section originally said CloudNativePG mints an
|
|
`X-app` secret we reference with `secretKeyRef`, as `user-engine` does. That is
|
|
true for a cluster's *own* app database — and wrong for the path we are
|
|
actually taking.
|
|
|
|
`railiance-platform/docs/rapp-postgres-boundary.md` is explicit: `rapp-postgres`
|
|
owns the role and database provisioning surface for the shared cluster, and
|
|
*"the workload receives a short-lived lease through the platform broker"* — the
|
|
OpenBao database secrets engine — rather than a static secret. `railiance-platform`
|
|
retains cluster-wide governance and the credential-broker grant catalog.
|
|
|
|
Consequences, and the second one is real work:
|
|
|
|
- **We do not provision our own database.** We *declare a consumer request* and
|
|
`rapp-postgres` provisions it. Writing a CNPG `Database` manifest into this
|
|
repo's `deploy/`, or adding a managed role to a shared cluster's spec, would
|
|
both be reaching across that boundary — the latter also mutating shared
|
|
infrastructure other consumers depend on.
|
|
- **Credentials are leased, not mounted.** A short-lived lease means the
|
|
connection string can change under a running pod, so the store cannot read a
|
|
DSN once at startup and hold it forever. T03 and T05 must account for
|
|
credential refresh; a long-lived pool built on an expired lease fails at the
|
|
worst possible moment.
|
|
|
|
Either way **no secret enters this repo, this workplan, or any evidence dump**.
|
|
Runtime secret custody stays OpenBao's per `SCOPE.md`, which is exactly what
|
|
the broker lease implements.
|
|
|
|
## Placement: shared by default, movable by design
|
|
|
|
Dedicated-vs-shared is **not a decision this workplan should settle once**, and
|
|
`tenant-engine` is the wrong place to settle it for the platform. The estate
|
|
has distinct elements — railiance, NetKingdom, HelixForge, Coulomb — plus
|
|
tenants layered on top, and which repo belongs to which grouping is not yet
|
|
organised. On top of that, isolation level is heading toward being a
|
|
*product* property: plans will differ in how much isolation they buy.
|
|
|
|
So the requirement here is **portability, not placement**:
|
|
|
|
- start on a **shared** cluster, because it is cheaper on a single node and
|
|
nothing yet justifies dedicated capacity;
|
|
- make moving to a dedicated cluster — or to a different shared one — an
|
|
**operational change, not a code change**;
|
|
- never let the decision leak into the application.
|
|
|
|
What that demands of this repo, and all of it is cheap if done now and
|
|
expensive to retrofit:
|
|
|
|
1. **Connect by injected URL only.** No cluster name, namespace, host, or
|
|
database name anywhere in `src/`. The application must be unable to tell
|
|
whether it is on a shared or dedicated cluster.
|
|
2. **Own a whole database, not a set of tables in someone else's.** Sharing a
|
|
*cluster* is a capacity decision; sharing a *database* would entangle
|
|
schemas and make relocation a merge rather than a move. A dedicated
|
|
database inside a shared cluster relocates with a dump and a restore.
|
|
3. **No cross-database joins or co-location assumptions**, which the repo
|
|
boundary already forbids — `tenant_id` is the only key shared with
|
|
`user-engine`. Placement must not quietly become a dependency.
|
|
4. **Idempotent schema creation**, so a fresh target comes up correct without a
|
|
hand-built database.
|
|
|
|
Then "move `tenant-engine` to dedicated" is: create the target cluster, dump,
|
|
restore, swap the `secretKeyRef`, restart. No rebuild, no release, no code
|
|
review.
|
|
|
|
**What this repo does not own:** the placement *policy* — which element or
|
|
tenant gets dedicated capacity and when. That is a platform concern, and
|
|
`tenant-engine` implementing its own would be the same boundary error as
|
|
building a fleet drift-detector because we got bitten. Raised with
|
|
`railiance-platform` separately; see T01.
|
|
|
|
## T01 - Make placement portable and provision on shared capacity
|
|
|
|
```task
|
|
id: TEN-WP-0009-T01
|
|
status: done
|
|
priority: high
|
|
state_hub_task_id: "b4701256-f235-4474-9dcf-7cb09f62b873"
|
|
```
|
|
|
|
**Target: `platform-pg`, via a `PostgresConsumer` declaration to
|
|
`rapp-postgres`.** Not `net-kingdom-pg` or `apps-pg`, and not provisioned by us.
|
|
|
|
Reasoning, revised 2026-08-17 after finding the actual contract. The candidate
|
|
list originally weighed product-family fit and would have picked
|
|
`net-kingdom-pg`. But `platform-pg` is the shared cluster that has a
|
|
*governed provisioning path* — `rapp-postgres` owns its consumer declarations,
|
|
role provisioning, isolation tests, and recovery procedure. The other clusters
|
|
carry their roles as hand-added entries in the shared cluster spec
|
|
(`net-kingdom-pg` has `privacyidea`, `apps-pg` has `vergabe` and
|
|
`coulomb_social`). Choosing product-family fit would mean improvising outside
|
|
the one contract that exists, to gain an affinity that portability makes cheap
|
|
to change later anyway. `audit-core` is the precedent: a platform service, on
|
|
`platform-pg`, declared as a consumer.
|
|
|
|
Note the deviation from `user-engine` and why: they run a dedicated
|
|
`user-engine-pg`. We are deliberately *not* copying that. Their choice is fine
|
|
and ours is reversible; picking dedicated now would spend capacity on an
|
|
isolation guarantee nothing has asked for yet.
|
|
|
|
The declaration to request, modelled on `consumers/audit-core.yaml`:
|
|
|
|
```yaml
|
|
apiVersion: rapp-postgres.railiance.io/v1alpha1
|
|
kind: PostgresConsumer
|
|
metadata:
|
|
name: tenant-engine
|
|
spec:
|
|
database: tenant_engine
|
|
schema: tenant_engine
|
|
costAttributionKey: platform:tenant-engine
|
|
clientNamespaces: [tenant-engine]
|
|
roles:
|
|
owner: tenant_engine_owner
|
|
migration: tenant_engine_migrate
|
|
runtime: tenant_engine_app
|
|
limits:
|
|
migrationConnections: 2
|
|
runtimeConnections: 12
|
|
statementTimeout: 30s
|
|
idleInTransactionSessionTimeout: 15s
|
|
tenantKeyingRequired: true
|
|
tenantIsolation: consumer-service-boundary
|
|
```
|
|
|
|
Two things to raise with them rather than assume:
|
|
|
|
- **`tenantKeyingRequired: true`** is right for `audit-core`, whose rows are
|
|
per-tenant. Ours are *about* tenants — `tenant_id` is the primary key of the
|
|
`tenants` table, not a partition key on someone else's data. Confirm that
|
|
their isolation tests read our shape correctly rather than flagging a false
|
|
positive.
|
|
- **Statement timeout.** 30s is generous for this workload; every query here is
|
|
a single-row lookup or a small transaction. A tighter timeout is a better
|
|
failure mode for a service `flex-auth` calls synchronously on the
|
|
authorization path — a slow query should fail closed fast, not hold the PDP.
|
|
|
|
Record what would trigger a move to dedicated, so the reversal is a judgement
|
|
already made rather than one improvised under pressure. Candidate triggers:
|
|
a noisy neighbour affecting the authorization path, a compliance or residency
|
|
requirement, a plan tier that sells isolation, or the shared cluster's backup
|
|
policy no longer fitting.
|
|
|
|
Also decide **backups**: whether the shared cluster's existing
|
|
`scheduledbackups` policy covers us, or we need our own. SQLite-on-a-PVC had no
|
|
backup story at all, so this is a gain to claim explicitly rather than inherit
|
|
by accident — and a shared cluster means inheriting *someone else's* retention
|
|
choice, which is worth checking rather than assuming.
|
|
|
|
**Do not write a CNPG `Database` or `Cluster` manifest into this repo's
|
|
`deploy/`.** That surface belongs to `rapp-postgres`. Our `deploy/` gains only
|
|
the consumption side — how the workload reaches the leased credential.
|
|
|
|
Done when `rapp-postgres` has provisioned the consumer, the connection reaches
|
|
the app only through an injected, refreshable credential, and the
|
|
move-to-dedicated trigger list and runbook are recorded.
|
|
|
|
Status 2026-08-17: **accepted.** `consumers/tenant-engine.yaml` merged at
|
|
`deda11e`. Provisioning is operator-gated as RAPP-IN-0004; nothing applied yet.
|
|
|
|
2026-08-18 repository readiness: desired runtime and migration manifests now
|
|
consume separate file-projected leases, label the client namespace, and admit
|
|
egress only to the real `databases/platform-pg` destination (an earlier draft
|
|
incorrectly selected namespace `railiance-platform`). The published retention
|
|
policy replaces the declaration's placeholder. T01 remains progress because
|
|
consumer/lease provisioning and backup evidence are live operator gates.
|
|
|
|
**They caught a real defect in our proposal.** The declaration had a single
|
|
`statementTimeout` field that the renderer applied to the migration role, the
|
|
runtime role, *and* the database — so our 5s would have given us 5s DDL and
|
|
failed the first non-trivial migration. Rather than reject it they made it
|
|
expressible, adding `limits.migrationStatementTimeout` (defaulting to
|
|
`statementTimeout`, so audit-core's rendered SQL is byte-identical). Ours is
|
|
now `statementTimeout: 5s` / `migrationStatementTimeout: 30s`. The 5s runtime
|
|
reasoning was accepted as-is.
|
|
|
|
Mechanical detail worth keeping: the *database*-level `statement_timeout` is
|
|
what a leased login actually gets, because group-role settings do not pass
|
|
through membership and OpenBao vends a fresh login role that merely inherits
|
|
`tenant_engine_migrate`. The 30s therefore has to be set on the leased login in
|
|
the migration role's creation statement. That lands in their OpenBao config —
|
|
but if a migration ever dies at 5s, that is where to look.
|
|
|
|
`tenantKeyingRequired: true` confirmed correct. The field asserts our data is
|
|
tenant-keyed from the first migration; nothing inspects the key's shape, and
|
|
the isolation probes run against a synthetic consumer, so there was no
|
|
audit-core-shaped assumption to trip. They queued a probe run with our consumer
|
|
present anyway, on the grounds that "we do not think it will false-positive" is
|
|
not evidence — which is the right standard.
|
|
|
|
### Answers to the three open questions
|
|
|
|
**Q1 — lease delivery: the credential *directory* is sanctioned.** A
|
|
`ClusterSecretStore` over OpenBao's database secrets engine projects the lease
|
|
to a mounted path. **Read the credential at connection checkout, not once at
|
|
startup.** Their rotation is overlap-first — acquire and verify the
|
|
replacement, switch the pool, drain the old, revoke the predecessor — so a
|
|
consumer that re-reads per connection rotates with no restart, while one that
|
|
takes the lease as process environment needs a restart every rotation. They
|
|
require that to be an explicit per-consumer choice; we take the directory.
|
|
|
|
There is also a broker `exec` path injecting `PG*` vars into a child process
|
|
(`railiance-platform/scripts/credential.py`) — for reviewed migrations and
|
|
operator sessions, not the runtime pod.
|
|
|
|
Theirs to create: `ClusterSecretStore openbao-tenant-engine`, and OpenBao roles
|
|
`rapp-postgres/tenant-engine-{runtime,migration}` — bundled with audit-core's
|
|
outstanding RAPP-IN-0003.
|
|
|
|
**Ours to do:** label the namespace
|
|
`railiance.io/postgres-client=platform-pg`. Their consumer-ingress
|
|
NetworkPolicy selects on that label, *not* on `spec.clientNamespaces` — the
|
|
declaration field is review metadata, the label is enforcement. Done
|
|
declaratively in `deploy/tenant-engine.yaml`'s Namespace document rather than
|
|
by an imperative `kubectl label`, so it cannot drift out of the repo.
|
|
|
|
**Q2 — out-of-band migration is the contract, not a preference.** The runtime
|
|
role gets DML only and owns nothing; migrate-on-startup would require DDL on
|
|
the runtime credential and collapse the separation. So our idempotent startup
|
|
migration becomes a step run under a `tenant-engine-migration` lease *before*
|
|
the new image serves traffic. Keep it idempotent regardless — that property is
|
|
what makes a move to dedicated capacity a restore rather than a rebuild.
|
|
|
|
**Q3 — backup is inherited and now real:** continuous WAL archiving to Scaleway
|
|
S3 (gzip), daily base backup at 02:15, 30-day retention, plus an encrypted
|
|
logical dump lane to Nextcloud without PITR.
|
|
|
|
No certified RPO. Restore and PITR are proven against a scratch cluster and
|
|
scratch object store, not the production `serverName`, so the honest status is
|
|
"designed for minutes, not yet measured". They volunteered that rather than
|
|
quoting a number they had not earned.
|
|
|
|
### Two things to push back on
|
|
|
|
**1. Dedicated capacity does not mitigate §5 for us.** They offer a dedicated
|
|
instance as the escalation for a consumer that cannot accept ADR-0001 §5's
|
|
residual risk — a leaked `tenant_engine_app` credential reading and writing
|
|
every row we own. For us that escalation does not work: a dedicated instance
|
|
changes co-residency, not what our own runtime credential can reach. The blast
|
|
radius is identical either way. So we should *not* argue for dedicated on §5
|
|
grounds; it would buy a different guarantee than the one at risk.
|
|
|
|
What §5 actually means here is worse than for `audit-core`, and worth stating
|
|
plainly: direct DML bypasses every control this service exists to enforce —
|
|
flex-auth authorization, version CAS, the guardrail loosening guard, and the
|
|
audit trail. Someone with the runtime credential could grant `PLTF` on any
|
|
tenant, or raise any ceiling, with no authorization check and no audit event.
|
|
Since `tenant-engine` is the source of the `tenant_roles` claim `key-cape`
|
|
mints and `flex-auth` reads live, that is privilege escalation across
|
|
NetKingdom, not just data tampering. §4's opt-in RLS does not help: our rows
|
|
*are* the tenants, so there is no per-tenant predicate to scope by.
|
|
|
|
The real mitigations are lease lifetime and network reach, both of which they
|
|
provide — plus one we do not have, below.
|
|
|
|
**2. Our audit trail is only as trustworthy as our own database.** `events`
|
|
lives in the same database as the rows it attests to, so the §5 scenario
|
|
forges the evidence along with the data. That is a genuine gap, and the fleet
|
|
already has the answer: `audit-core` exists as an append-only sink, and
|
|
`AUDIT_CORE_SENDERS` implies a registration path. Emitting there would make our
|
|
trail tamper-evident independently of our store.
|
|
|
|
**Out of scope for this workplan** — it is a new cross-service dependency, not
|
|
a store migration — but it should not be discovered later. Raised with
|
|
`rapp-postgres`; needs its own workplan.
|
|
|
|
### RTO answer they asked for
|
|
|
|
They asked us to say now if our RTO on the authorization path cannot absorb a
|
|
single-consumer restore (a logical dump taken from a scratch-restored copy of
|
|
the whole instance, then a controlled import).
|
|
|
|
Split the cases, because the answer differs:
|
|
|
|
- **Whole-instance loss:** we are restored alongside everyone else. Acceptable,
|
|
and dedicated capacity would make it *worse* by adding a second instance to
|
|
recover.
|
|
- **tenant-engine-only corruption** — a bad migration, or the §5 credential
|
|
scenario — needs the slow single-consumer path. This is the case that matters,
|
|
and it is sharpened by our being the authority: there is no upstream to replay
|
|
from. If our data is gone or forged, nothing else in the fleet can reconstruct
|
|
it.
|
|
|
|
So the ask is not dedicated capacity but a **per-consumer logical dump
|
|
cadence**. They already run an encrypted logical dump lane to Nextcloud; if it
|
|
covers per-database dumps, our single-consumer restore comes from there rather
|
|
than from a scratch-restored instance, and the slow path stops being the only
|
|
path. Asked.
|
|
|
|
Completed 2026-08-21. `Database/tenant-engine`, the namespace admission label,
|
|
exact-scope OpenBao runtime/migration roles, exact-namespace SecretStore and
|
|
split ExternalSecrets are live. The runtime lease was forcibly refreshed
|
|
without a pod restart, both local and live isolation suites passed 19/19, and
|
|
the shared placement remains within the reviewed four-consumer ceiling. The
|
|
move-to-dedicated triggers and honest backup/restore boundary remain recorded
|
|
above. Live evidence is in
|
|
`rapp-postgres/docs/evidence/tenant-engine-postgres-cutover-2026-08-21.md`.
|
|
|
|
## T02 - Implement `PostgresTenantStore`
|
|
|
|
```task
|
|
id: TEN-WP-0009-T02
|
|
status: done
|
|
priority: high
|
|
state_hub_task_id: "d7428bc1-2e5f-4a4c-93f5-9368ebb691d5"
|
|
```
|
|
|
|
Implement the full `TenantStore` Protocol — including
|
|
`guardrail_overrides` / `guardrail_changes` / `set_guardrail_override` from
|
|
TEN-WP-0006 — and add PostgreSQL to the conformance suites, which are already
|
|
parametrised over in-memory and SQLite. A third backend should be a fixture
|
|
parameter, not a new test file.
|
|
|
|
Invariants that must hold identically, not approximately:
|
|
|
|
- `mutate_tenant()` commits idempotency replay, version CAS, mutation, and
|
|
audit event **in one transaction**;
|
|
- `set_guardrail_override()` does the same for the override, its audit record,
|
|
and its receipt;
|
|
- a replayed `Idempotency-Key` short-circuits *before* the version check;
|
|
- reads and writes never observe a row mid-transaction.
|
|
|
|
That last one is where PostgreSQL should simplify rather than complicate:
|
|
SQLite needed an explicit `RLock` around reads because one connection is shared
|
|
across request threads, and the concurrent-writer test caught a real defect
|
|
there. Postgres gives proper per-connection transaction isolation, so the lock
|
|
should disappear rather than be ported. **Do not transliterate the SQLite
|
|
implementation** — port the semantics and let the backend do what it is good
|
|
at. Use a connection pool and appropriate isolation, and prove the guarantees
|
|
with the same concurrent-writer test that caught the SQLite bug.
|
|
|
|
Adds a runtime dependency (`psycopg`); keep it out of the base install path if
|
|
SQLite-only dev is meant to stay dependency-light.
|
|
|
|
Done when all three backends pass the same conformance suites unchanged.
|
|
|
|
Completed 2026-08-18. `PostgresTenantStore` implements the full protocol with
|
|
bounded pooling, per-checkout credential-file rotation, row locking and atomic
|
|
receipt/version/event transactions. The existing lifecycle and guardrail
|
|
conformance suites run against PostgreSQL when
|
|
`TENANT_ENGINE_TEST_DATABASE_URL` is present. A disposable PostgreSQL 16 run
|
|
passed 88 shared conformance cases; the full source suite passes 260 with 28
|
|
environment-gated PostgreSQL cases skipped when no database is supplied.
|
|
|
|
## T03 - Backend selection that fails closed
|
|
|
|
```task
|
|
id: TEN-WP-0009-T03
|
|
status: done
|
|
priority: high
|
|
state_hub_task_id: "89c121c6-8bc5-4abe-80f0-d11406aeffc2"
|
|
```
|
|
|
|
`Settings` currently carries only `database_path`. Add a database URL and make
|
|
selection explicit:
|
|
|
|
- URL set → PostgreSQL;
|
|
- path set → SQLite;
|
|
- **both set → refuse to start.** Do not silently prefer one. An ambiguous
|
|
store configuration in production is exactly the class of silent
|
|
misconfiguration this repo fails closed on everywhere else, and picking a
|
|
winner would let a stale `TENANT_ENGINE_DATABASE_PATH` quietly shadow the
|
|
real database;
|
|
- neither set → in-memory, as today, which is correct for tests and wrong for
|
|
production.
|
|
|
|
Consider surfacing the active backend on `/health`. The last two production
|
|
incidents were invisible because the service looked fine from outside; "which
|
|
store am I actually using" is cheap to answer and expensive to guess.
|
|
|
|
Done when misconfiguration is a startup failure with a clear message, covered
|
|
by tests.
|
|
|
|
Completed 2026-08-18. URL-file selects PostgreSQL, path selects SQLite, both
|
|
refuse startup, and neither retains the in-memory test default. `/health`
|
|
reports and pings the active backend so store selection and credential failure
|
|
are externally visible without disclosing a DSN.
|
|
|
|
## T04 - Migrate production data
|
|
|
|
```task
|
|
id: TEN-WP-0009-T04
|
|
status: done
|
|
priority: high
|
|
state_hub_task_id: "51a83eae-9e9f-4676-bc1d-9c2d884bac0d"
|
|
```
|
|
|
|
Move the live SQLite database to PostgreSQL. Currently small — two tenants
|
|
(`tenant:trial:portalcheck` active v1, `tenant:trial:ten-wp-0005-t05` retired
|
|
v5) plus the T04 disposables from TEN-WP-0007 — but small is not the same as
|
|
trivial.
|
|
|
|
Everything must move, not just the tenants:
|
|
|
|
- `tenants` with versions and lifecycle timestamps **preserved exactly** — a
|
|
reset version silently breaks every consumer holding an ETag;
|
|
- `grants` including revoked ones, because the trail is append-only and its
|
|
history is the audit record;
|
|
- `plans`;
|
|
- `events` — the audit correlation contract;
|
|
- `idempotency_receipts` — dropping these lets an in-flight retry double-apply
|
|
a mutation that already happened;
|
|
- `guardrail_overrides` and `guardrail_changes`.
|
|
|
|
SQLite is single-writer and the app writes to it, so the cutover must stop
|
|
writes rather than race them: scale to zero, export, import, verify, then start
|
|
against PostgreSQL. Verify by comparing row counts *and* spot-checking the two
|
|
known tenants' versions and lifecycle state, not by trusting the exporter.
|
|
|
|
Keep the SQLite file and its PVC until T05's soak passes. It is the rollback.
|
|
|
|
Done when a verified copy is live in PostgreSQL and the comparison evidence is
|
|
recorded here.
|
|
|
|
2026-08-18 repository readiness: `tenant-engine-transfer` now refuses a
|
|
non-empty target, copies all seven tables in one transaction, preserves
|
|
identity sequence values and nullable legacy timestamps, compares every source
|
|
and target row, and emits non-secret counts/SHA-256 digests plus requested
|
|
tenant lifecycle/version checks. Its PostgreSQL 16 integration test exercises
|
|
all tables and the next event sequence. The live scale-to-zero/export/import is
|
|
not performed and T04 remains progress.
|
|
|
|
Completed 2026-08-21. Writes were stopped before the source snapshot. The
|
|
transfer verified all seven physical tables by exact row comparison and
|
|
deterministic digest: 5 tenants, 13 events, 8 receipts, 3 guardrail changes, 1
|
|
override, and empty grants/plans. The known active v1 and retired v5 tenant
|
|
records were unchanged. Two failed attempts rolled back to an empty target;
|
|
the successful retry also proved the next event sequence. The stopped-write
|
|
SQLite snapshot hash and full per-table digests are recorded in the linked
|
|
rapp-postgres evidence.
|
|
|
|
## T05 - Cut over the deployment
|
|
|
|
```task
|
|
id: TEN-WP-0009-T05
|
|
status: done
|
|
priority: high
|
|
state_hub_task_id: "03aeae3b-5e0f-4c77-a763-787ae08078f5"
|
|
```
|
|
|
|
Update `deploy/tenant-engine.yaml`:
|
|
|
|
- drop the PVC and its volume mount;
|
|
- add `TENANT_ENGINE_DATABASE_URL` from the CNPG-minted `secretKeyRef`;
|
|
- switch `strategy: Recreate` → `RollingUpdate`, now that no `ReadWriteOnce`
|
|
volume forces serialization — **this is what unblocks TEN-WP-0008-T02's
|
|
canary**;
|
|
- extend the `NetworkPolicy` egress to the Postgres service, and no wider.
|
|
Ingress stays restricted to `user-engine`.
|
|
|
|
Then decide, deliberately, whether to raise `replicas` above 1. It becomes
|
|
possible here for the first time; possible is not the same as warranted, and a
|
|
second replica changes failure modes for a service on the authorization path.
|
|
|
|
Update `make verify-pin` if the expected shape changes, and re-run it.
|
|
|
|
Rollback: re-pin the previous digest and re-attach the PVC. State plainly that
|
|
this becomes lossy once writes have landed in PostgreSQL — after that point the
|
|
rollback is "restore from Postgres", not "swap back to the file".
|
|
|
|
Done when production runs on PostgreSQL, `make verify-pin` passes, and the
|
|
TEN-WP-0007 T04 verification list still passes end to end against the live
|
|
service.
|
|
|
|
2026-08-18 desired state: PVC/mount removed, `RollingUpdate` selected,
|
|
file-projected runtime and caller credentials added, and egress narrowed to
|
|
`databases/platform-pg`, flex-auth and DNS. A separate migration Job has its
|
|
own lease and NetworkPolicy. Replicas deliberately remain one. The manifests
|
|
are not applied and the pinned old digest does not yet contain this code, so
|
|
T05 remains progress and no live PostgreSQL claim is made.
|
|
|
|
Completed 2026-08-21. Production is healthy with backend `postgresql` at
|
|
immutable digest
|
|
`sha256:a8e8086ffc5b772c1391b166f5e1884b90f7d327b152c205eceae129df555c24`.
|
|
The repo pin, Deployment spec and running image ID agree; `make verify-pin` and
|
|
the package `verify-live` target passed. The Deployment is `RollingUpdate`,
|
|
does not mount the SQLite PVC, keeps one deliberate replica, and retained all
|
|
five required routes plus the TEN-WP-0007 guardrail/lifecycle checks. The old
|
|
PVC is a time-bounded rollback artifact, not a current authority; its post-soak
|
|
retirement is `RAPP-IN-0005`.
|
|
|
|
## T06 - Tell the fleet, and close the loop on TEN-WP-0008
|
|
|
|
```task
|
|
id: TEN-WP-0009-T06
|
|
status: done
|
|
priority: medium
|
|
state_hub_task_id: "e7adbc64-ad1b-4f28-bfd0-0b07034629db"
|
|
```
|
|
|
|
- `flex-auth`: availability characteristics of their data source changed. No
|
|
contract or action change.
|
|
- `user-engine`: no contract change either, but they run the same CNPG pattern
|
|
and may have operational lessons worth having before we repeat their mistakes.
|
|
- **TEN-WP-0008-T02**: record that the `ReadWriteOnce` canary constraint is
|
|
gone, so staged-promotion onboarding no longer has to describe a canary that
|
|
cannot run.
|
|
|
|
Done when the notes are sent and TEN-WP-0008-T02 is updated.
|
|
|
|
Completed 2026-08-21. Non-secret completion notices were sent to flex-auth
|
|
(State Hub message `6bb6f923-8909-45d0-9c57-2b6d7c3e41ad`) and user-engine
|
|
(`9512cf0c-f2f0-486b-a710-f737b9c79510`), stating the availability/storage
|
|
change and unchanged service contracts. TEN-WP-0008-T02 now records that the
|
|
RWO constraint is gone and that side-by-side canary execution is
|
|
storage-feasible.
|
|
|
|
## Out of scope
|
|
|
|
- **Removing SQLite.** It stays as the dev and test backend, and as the thing
|
|
that keeps `TenantStore` an honest seam.
|
|
- **Schema redesign.** Port the existing schema; a store migration and a model
|
|
change at the same time makes any regression ambiguous.
|
|
- **Sharing a database with `user-engine`.** `tenant_id` is the only key the
|
|
two services share, and it stays that way — separate stores are the boundary,
|
|
not an implementation detail. Sharing a *cluster* is fine; sharing a
|
|
*database* is not.
|
|
- **Fleet placement policy.** Which element or tenant gets dedicated capacity,
|
|
and how isolation maps to plan tiers, is a platform decision. This workplan
|
|
makes `tenant-engine` movable and stops there.
|
|
- **Emitting the audit trail to `audit-core`.** Identified under T01 as a real
|
|
gap — our `events` table shares a database with the rows it attests to, so
|
|
ADR-0001 §5 forges both together. It is a new cross-service dependency rather
|
|
than part of a store migration. Live residual `TEN-IN-0001` owns promotion
|
|
into a reviewed workplan; it is no longer parked only in this prose.
|
|
- **Per-tenant isolation.** A future plan tier may sell dedicated
|
|
infrastructure. `tenant-engine` already records plan assignment by id, and
|
|
`adaptive-pricing` owns what a plan *means*, so neither the tier definition
|
|
nor the placement that implements it belongs here. Worth designing for —
|
|
which the portability principles above do — not worth building for.
|