Two reachable clusters each carry a CNPG Cluster named apps-pg in a namespace named databases. KUBECONFIG is an environment variable, so the Makefile ?= default never applied, and RAILIANCE01_KUBECONFIG pointed at config-hosteurope - a different cluster. Had the environment pointed at the other reachable cluster instead of an unauthorized one, make apps-pg-deploy would have applied RPF-WP-0019 connection limits, role timeouts and backup config to the wrong cluster and reported success. The Unauthorized error was the only thing that prevented it. Filename selection cannot protect against this: both kubeconfigs resolve to a 127.0.0.1 tunnel port and the environment wins either way. railiance01-guard pins identity instead, comparing the live kube-system namespace UID against RAILIANCE01_CLUSTER_UID, and fails closed on mismatch or unreachability. It gates apps-pg deploy, backup-deploy, overflow-dry-run, status and shell. Verified refusing on the wrong cluster, refusing when unreachable, and passing on railiance01. Not global: db-status legitimately targets the other cluster for gitea-db. RPF-WP-0019 blocker note corrected - the cluster was never unreachable, our wiring was wrong. RPF-WP-0020 seeded for the pre-existing CCR test failure, which is two unrelated problems: CCR-2026-0010 is an active lane missing its whole openbao.auth block, and CCR-2026-0011 is an honest in-flight draft the suite has no way to express. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
150 lines
6.3 KiB
Markdown
150 lines
6.3 KiB
Markdown
# apps-pg Shared PostgreSQL Cluster
|
|
|
|
`apps-pg` is the shared CloudNativePG cluster for S5 application
|
|
databases. It lives in the `databases` namespace and is owned by
|
|
`railiance-platform` as an S3 platform service.
|
|
|
|
## Cluster Identity
|
|
|
|
- CNPG Cluster: `apps-pg`
|
|
- Namespace: `databases`
|
|
- PostgreSQL major version: `16`
|
|
- Primary endpoint: `apps-pg-rw.databases.svc.cluster.local:5432`
|
|
- Read-only endpoint: `apps-pg-ro.databases.svc.cluster.local:5432`
|
|
- Bootstrap database: `apps_meta`
|
|
- Bootstrap role: `apps_admin`
|
|
|
|
`apps_admin` is a platform bootstrap and smoke-test role. Do not copy it
|
|
into application namespaces, use it in S5 runtime configuration, or treat
|
|
it as a consumer credential.
|
|
|
|
## Which cluster this is
|
|
|
|
**Two reachable clusters each carry a CNPG `Cluster` named `apps-pg` in a
|
|
namespace named `databases`.** The one this document describes is
|
|
**railiance01** (k3s v1.35.1) — it also carries `platform-pg` and `forgejo-db`,
|
|
and holds both apps-pg consumers (`vergabe_db`, `coulomb_social_db`). The other
|
|
cluster carries `gitea-db` and only one apps-pg consumer.
|
|
|
|
Selecting the right one by kubeconfig filename is not safe: `KUBECONFIG` is an
|
|
environment variable that overrides the Makefile default, and both files
|
|
resolve to a `127.0.0.1` tunnel port. Every `apps-pg-*` target therefore runs
|
|
`railiance01-guard` first, which compares the live `kube-system` namespace UID
|
|
against `RAILIANCE01_CLUSTER_UID` and fails closed on mismatch or
|
|
unreachability.
|
|
|
|
```bash
|
|
make cluster-id # what KUBECONFIG currently selects
|
|
make railiance01-guard # assert it is railiance01, or refuse
|
|
```
|
|
|
|
If a target refuses, the fix is `KUBECONFIG=~/.kube/config-railiance01`, not
|
|
`--force` and not editing the pinned UID. `db-status` deliberately targets the
|
|
other cluster for `gitea-db`, which is why the guard is per-target.
|
|
|
|
## Consumer Onboarding
|
|
|
|
Each S5 application gets its own role, database, and runtime Secret. The
|
|
current flow is platform-administered until OpenBao or a dedicated
|
|
database onboarding automation owns the lifecycle end to end.
|
|
|
|
1. Request an app database in the consuming repo workplan. Include the
|
|
app name, namespace, database name, role name, intended owners, and
|
|
any required extensions.
|
|
2. Platform reviews and approves the database/role names. Names should
|
|
be app-scoped, for example `vergabe` and `vergabe_db`.
|
|
3. Platform provisions the backing PostgreSQL role and credential for
|
|
the approved app. Until automation exists, this is a controlled
|
|
operator SQL action against `apps-pg`, not a self-service repo apply.
|
|
4. Add a CNPG `Database` manifest in the platform-managed database
|
|
manifests, with `spec.cluster.name: apps-pg` and `spec.owner` set to
|
|
the approved role.
|
|
5. Label the consuming namespace so NetworkPolicy allows access:
|
|
|
|
```bash
|
|
kubectl label namespace <app-namespace> \
|
|
railiance.io/postgres-client=apps-pg
|
|
```
|
|
|
|
6. Publish or mirror the runtime Secret into the consumer namespace. The
|
|
Secret should contain only that app's role credential and DSN fields.
|
|
7. Wire the DSN into the application Helm values or runtime
|
|
configuration. Prefer the RW service for migrations and writes:
|
|
`postgresql://<role>:<password>@apps-pg-rw.databases.svc.cluster.local:5432/<database>`.
|
|
|
|
Example CNPG `Database` resource:
|
|
|
|
```yaml
|
|
apiVersion: postgresql.cnpg.io/v1
|
|
kind: Database
|
|
metadata:
|
|
name: vergabe-db
|
|
namespace: databases
|
|
spec:
|
|
cluster:
|
|
name: apps-pg
|
|
name: vergabe_db
|
|
owner: vergabe
|
|
```
|
|
|
|
## CNPG Boundary
|
|
|
|
CNPG 1.28 provides a standalone `Database` CRD. It does not provide a
|
|
standalone `Role` CRD in this cluster. Role lifecycle is cluster-scoped,
|
|
through `Cluster.spec.managed.roles` or a controlled SQL workflow.
|
|
|
|
Consumer repos must therefore not assume they can create PostgreSQL
|
|
roles themselves. They can request a database and consume a runtime
|
|
Secret after the platform role has been provisioned.
|
|
|
|
## Network Policy
|
|
|
|
The `databases` namespace has a default-deny posture. `apps-pg` accepts
|
|
client traffic on TCP/5432 only from namespaces labeled:
|
|
|
|
```text
|
|
railiance.io/postgres-client=apps-pg
|
|
```
|
|
|
|
The CNPG operator in `cnpg-system` is allowed to manage the cluster on
|
|
the standard PostgreSQL, instance manager, and metrics ports.
|
|
|
|
## Consumers (live)
|
|
|
|
| App | Role | Database | Password secret (databases ns) | Consumer ns label |
|
|
| --- | --- | --- | --- | --- |
|
|
| vergabe-teilnahme | `vergabe` | `vergabe_db` | `vergabe-app-credentials` | `vergabe-teilnahme` |
|
|
| coulomb-social | `coulomb_social` | `coulomb_social_db` | `coulomb-social-app-credentials` | `coulomb-social` |
|
|
|
|
Bootstrapped 2026-08-09 on railiance01: cluster healthy; both Database CRs
|
|
applied; coulomb-social connectivity smoke from labeled consumer ns OK.
|
|
|
|
## Backup And Roadmap
|
|
|
|
`apps-pg` remains a conservative single-instance, 10Gi cluster. The desired
|
|
state now includes continuous WAL archival, a daily 02:15 UTC base backup and
|
|
30-day retention under the distinct `apps-pg/` object-store prefix. The
|
|
credential lane is shared with `platform-pg`; the backup data path is not.
|
|
|
|
The declared ceiling is three consumers. Each gets at most 20 connections;
|
|
40 of the explicit 100-connection aggregate remains for CNPG and operator
|
|
headroom. Memory (1Gi limit), not the clean connection refusal, is treated as
|
|
the binding safety constraint. A fourth consumer goes to the named
|
|
`apps-pg-2` overflow substrate, which must be provisioned before onboarding.
|
|
Reviewed, unapplied source for that cell is
|
|
`helm/apps-pg-2-{cluster,backup,networkpolicies}.yaml`; it uses a distinct
|
|
bootstrap Secret and backup prefix. `make apps-pg-verify-capacity` rejects a
|
|
fourth role on either cell, missing connection limits, duplicate roles, missing
|
|
resource envelopes, or a reused backup path.
|
|
|
|
`statement_timeout` and `idle_in_transaction_session_timeout` are 15 seconds
|
|
per consumer role. CNPG 1.28 cannot express those role settings, so
|
|
`helm/apps-pg-consumer-controls.sql` is the idempotent controlled-operator
|
|
step; `connectionLimit` remains declaratively reconciled by CNPG.
|
|
|
|
Resource evidence for `resource:railiance:apps-pg` (capacity, recovery,
|
|
labor, allocation drivers) is published under
|
|
`docs/evidence/RAILIANCE-WP-0016-apps-pg-resource-evidence.md`. The
|
|
2026-08-14 observation remains historically correct. The 2026-08-18 desired
|
|
state is not called verified recovery until the first backup succeeds and a
|
|
scratch restore artifact is recorded.
|