railiance-platform/docs/apps-pg.md
codex dc4245361d
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 1s
Finish RPF-WP-0018; RPF-WP-0019 repository-complete
RPF-WP-0018 closed: all seven tasks done. The provider-declaration finding
was adopted upstream and its canonical form is the provider: block in
tenancy.yaml; adaptive-pricing declined the standing co-signature and
supplied typed tier minima instead, recorded in ADR-0002. Three corrections
against our own output are recorded in the documents rather than edited
away.

RPF-WP-0019 T03 done (ceiling of three, memory binding, apps-pg-2 named as
overflow, enforced by make apps-pg-verify-capacity). T01/T02 are
repository-complete: backup target, retention, per-consumer connection
limits, role timeouts and Burstable resources are declared in source and
published in s3-consumer-interfaces 1.1.0 before rollout. They stay in
progress because no live application, backup success or restore proof
exists, and declared configuration is not a section 13 artifact. T04 waits
on that window.

apps-pg R reason corrected to say the target is declared-not-applied rather
than absent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 13:35:04 +02:00

126 lines
5.2 KiB
Markdown

# apps-pg Shared PostgreSQL Cluster
`apps-pg` is the shared CloudNativePG cluster for S5 application
databases. It lives in the `databases` namespace and is owned by
`railiance-platform` as an S3 platform service.
## Cluster Identity
- CNPG Cluster: `apps-pg`
- Namespace: `databases`
- PostgreSQL major version: `16`
- Primary endpoint: `apps-pg-rw.databases.svc.cluster.local:5432`
- Read-only endpoint: `apps-pg-ro.databases.svc.cluster.local:5432`
- Bootstrap database: `apps_meta`
- Bootstrap role: `apps_admin`
`apps_admin` is a platform bootstrap and smoke-test role. Do not copy it
into application namespaces, use it in S5 runtime configuration, or treat
it as a consumer credential.
## Consumer Onboarding
Each S5 application gets its own role, database, and runtime Secret. The
current flow is platform-administered until OpenBao or a dedicated
database onboarding automation owns the lifecycle end to end.
1. Request an app database in the consuming repo workplan. Include the
app name, namespace, database name, role name, intended owners, and
any required extensions.
2. Platform reviews and approves the database/role names. Names should
be app-scoped, for example `vergabe` and `vergabe_db`.
3. Platform provisions the backing PostgreSQL role and credential for
the approved app. Until automation exists, this is a controlled
operator SQL action against `apps-pg`, not a self-service repo apply.
4. Add a CNPG `Database` manifest in the platform-managed database
manifests, with `spec.cluster.name: apps-pg` and `spec.owner` set to
the approved role.
5. Label the consuming namespace so NetworkPolicy allows access:
```bash
kubectl label namespace <app-namespace> \
railiance.io/postgres-client=apps-pg
```
6. Publish or mirror the runtime Secret into the consumer namespace. The
Secret should contain only that app's role credential and DSN fields.
7. Wire the DSN into the application Helm values or runtime
configuration. Prefer the RW service for migrations and writes:
`postgresql://<role>:<password>@apps-pg-rw.databases.svc.cluster.local:5432/<database>`.
Example CNPG `Database` resource:
```yaml
apiVersion: postgresql.cnpg.io/v1
kind: Database
metadata:
name: vergabe-db
namespace: databases
spec:
cluster:
name: apps-pg
name: vergabe_db
owner: vergabe
```
## CNPG Boundary
CNPG 1.28 provides a standalone `Database` CRD. It does not provide a
standalone `Role` CRD in this cluster. Role lifecycle is cluster-scoped,
through `Cluster.spec.managed.roles` or a controlled SQL workflow.
Consumer repos must therefore not assume they can create PostgreSQL
roles themselves. They can request a database and consume a runtime
Secret after the platform role has been provisioned.
## Network Policy
The `databases` namespace has a default-deny posture. `apps-pg` accepts
client traffic on TCP/5432 only from namespaces labeled:
```text
railiance.io/postgres-client=apps-pg
```
The CNPG operator in `cnpg-system` is allowed to manage the cluster on
the standard PostgreSQL, instance manager, and metrics ports.
## Consumers (live)
| App | Role | Database | Password secret (databases ns) | Consumer ns label |
| --- | --- | --- | --- | --- |
| vergabe-teilnahme | `vergabe` | `vergabe_db` | `vergabe-app-credentials` | `vergabe-teilnahme` |
| coulomb-social | `coulomb_social` | `coulomb_social_db` | `coulomb-social-app-credentials` | `coulomb-social` |
Bootstrapped 2026-08-09 on railiance01: cluster healthy; both Database CRs
applied; coulomb-social connectivity smoke from labeled consumer ns OK.
## Backup And Roadmap
`apps-pg` remains a conservative single-instance, 10Gi cluster. The desired
state now includes continuous WAL archival, a daily 02:15 UTC base backup and
30-day retention under the distinct `apps-pg/` object-store prefix. The
credential lane is shared with `platform-pg`; the backup data path is not.
The declared ceiling is three consumers. Each gets at most 20 connections;
40 of the explicit 100-connection aggregate remains for CNPG and operator
headroom. Memory (1Gi limit), not the clean connection refusal, is treated as
the binding safety constraint. A fourth consumer goes to the named
`apps-pg-2` overflow substrate, which must be provisioned before onboarding.
Reviewed, unapplied source for that cell is
`helm/apps-pg-2-{cluster,backup,networkpolicies}.yaml`; it uses a distinct
bootstrap Secret and backup prefix. `make apps-pg-verify-capacity` rejects a
fourth role on either cell, missing connection limits, duplicate roles, missing
resource envelopes, or a reused backup path.
`statement_timeout` and `idle_in_transaction_session_timeout` are 15 seconds
per consumer role. CNPG 1.28 cannot express those role settings, so
`helm/apps-pg-consumer-controls.sql` is the idempotent controlled-operator
step; `connectionLimit` remains declaratively reconciled by CNPG.
Resource evidence for `resource:railiance:apps-pg` (capacity, recovery,
labor, allocation drivers) is published under
`docs/evidence/RAILIANCE-WP-0016-apps-pg-resource-evidence.md`. The
2026-08-14 observation remains historically correct. The 2026-08-18 desired
state is not called verified recovery until the first backup succeeds and a
scratch restore artifact is recorded.