--- id: RCP-WP-0002 type: workplan title: "First deployment of canned-prompts on Railiance" domain: agents repo: rapp-canned-prompts status: finished owner: codex topic_slug: practice created: "2026-09-06" updated: "2026-09-06" state_hub_workstream_id: "11874f32-ac36-5bb9-a0a5-e7a259f5972c" --- # First deployment of canned-prompts on Railiance The package is written and validated; what remains needs credentials and a published image, which are operator actions rather than authoring ones. `readiness_state` is `draft` and stays there until T04 produces evidence. Topology is not readiness (ADR-0006): binding this rapp to a reef does not make it deployed. ## Publish the image and pin its digest ```task id: RCP-WP-0002-T01 status: done priority: high state_hub_task_id: "3e7faf50-8fe9-5ef3-8a75-1d9e586d6c0f" ``` `declarations/rapp.yaml` carries `version: pending-publication`, and both manifests carry a matching tag. This is deliberate: a placeholder shaped like a real digest could be mistaken for something deployable, and a rapp that pins nothing is a stub that tells the fleet tooling a deployment exists when none does. The image builds and was verified locally in `canned-prompts` (`CANP-WP-0006-T06`): it starts non-root, all six service-level smoke checks pass against the running container, and the reference CLI installs a package from it over HTTP. **Done, 2026-09-07.** Published as `forgejo.coulomb.social/coulomb/canned-prompts:0.1.0` and pinned by **digest** in the declaration and both manifests: ``` sha256:e0ded3c7fe2548c910445deaa31ec42e84f129a92aa324841e231a53fee1f123 ``` Pinned by digest rather than tag deliberately: a tag can be moved, and `live-image-digest-match` would then pass against something that is no longer what this repo reviewed. Verified before pushing — the image resolves a file-mounted credential, carries the reference validator, and ships the migration scripts — and verified after, by fetching the manifest back **by digest** rather than trusting the push output. `readiness_state` moved `draft` → `declared`. ## Provision database roles and credentials ```task id: RCP-WP-0002-T02 status: done priority: high state_hub_task_id: "61f5cbd7-8fcf-5156-87da-57239ae55d8f" ``` Two roles, deliberately separate: a runtime role that reads and writes rows, and a migration role that owns the schema. A service able to `ALTER` its own tables at runtime turns any code defect into a schema defect. Needed from `rapp-postgres` and the credential broker: - database `canned_prompts` on `platform-pg-2`; - `creds/canned-prompts-runtime` and `creds/canned-prompts-migration` in OpenBao under the `database` path; - the `openbao-canned-prompts-eso-token` secret in `external-secrets`. `creds/canned-prompts-publish` is **optional**. Without it the service is read-only, which is the correct posture until per-publisher identity exists — not a misconfiguration to be worked around. **Requested, 2026-09-07.** `rapp-postgres` now carries `consumers/canned-prompts.yaml`, authored against its documented `PostgresConsumer` shape, and its agent has the request in its inbox. The declaration only. `make provision-consumers` — the step that actually mints OpenBao credentials — was deliberately **not** run: credential issuance is `rapp-postgres`' to perform, and running it from the consuming side would take a decision that is not this repo's, however available the script happens to be. **Received 2026-09-08.** `rapp-postgres` provisioned and sent a database-owner receipt with 12 checks proven live: runtime DDL denied (SQLSTATE 42501), statement timeouts and `search_path` as declared, and both logins refused CONNECT on `sbom_nexus` — the cell is shared, so that last one matters. `creds/canned-prompts-publish` was **not** created, as asked. The service runs read-only, which is the intended posture rather than a gap. ## Apply and migrate ```task id: RCP-WP-0002-T03 status: done priority: high state_hub_task_id: "0b8d206e-bd28-5950-abf7-d824015c09a4" ``` Order matters, and the ordering is the point rather than a convenience: Verified against the live cluster before writing: `platform-pg-2` exists and is healthy, and the `railiance.io/postgres-client: platform-pg-2` label matches what `sbom-nexus` actually carries in the cluster rather than only what its repo says. 1. `manifests/00-namespace.yaml` — the `railiance.io/postgres-client` label is what lets the database namespace accept traffic; 2. `manifests/database-secrets.yaml`, then wait for the secrets to materialize; 3. `manifests/migration.yaml` — the Job runs `alembic upgrade head` under the migration role and must complete before any replica serves; 4. `manifests/runtime.yaml`. The runtime deliberately does not migrate at start-up. Migrations as a Job keep a schema rollback separate from a code rollback and stop replicas racing each other. **Done 2026-09-08, after four defects that only a real rollout could expose.** Each is recorded in `docs/evidence/RCP-WP-0002-T04-first-deployment-2026-09-08.md`; the two worth repeating here failed *silently*: - `SET ROLE` opened an implicit transaction that Alembic then nested inside rather than owning, so every revision logged as applied and was rolled back. Alembic reported success against an empty database. - The egress NetworkPolicy selected `app.kubernetes.io/name`, which the migration Job does not carry. The Job matched only the default-deny and succeeded exactly once — because it ran before the policies existed. The next migration would have failed with a DNS error. Now selects `part-of`, and ingress is a separate policy so the Job is never reachable. `rapp-postgres` asked to be told when the first revision landed so they can re-run ownership reconciliation. Notified. ## Verify and record evidence ```task id: RCP-WP-0002-T04 status: done priority: high state_hub_task_id: "6d9eb97c-57e2-5b74-b6eb-455076713417" ``` Run `tools/smoke.sh`. It checks what only the cluster can answer — secrets materialized, Service is ClusterIP with no Ingress, NetworkPolicies present, live image digest matches the pin — and calls `canned-prompts`' `service/tools/smoke.py` for health and migration head rather than holding a second opinion about whether the service is healthy. Record the output as evidence, then move `readiness_state` to `deployed`, and to `verified` only with that evidence attached. **Done 2026-09-08.** All six deployment checks and all six service-level checks pass; evidence at `docs/evidence/RCP-WP-0002-T04-first-deployment-2026-09-08.md`. **Then corrected.** I recorded `verified` after a passing smoke run and found the deployment `0/1` eight hours later: platform credentials are 30-minute leases, and the service read the mounted URL once at start-up. External Secrets kept the file current; the engine held the URL it booted with. Fixed in 0.1.5 by re-reading the credential for every new connection, with `pool_recycle` inside the lease TTL. `readiness_state` is **`verified`** on evidence: continuously ready across a full lease rotation, zero readiness failures, no restarts. Log at `docs/evidence/RCP-WP-0002-T04-lease-rotation-2026-09-08.log`. The bar for `verified` is observation across a rotation, because a smoke run inside the first window cannot tell a service that works from one that works *once* — an availability property is not provable by a single sample. It took three attempts to build an instrument that could return that verdict honestly; all three failures are recorded in the evidence file. **The check that could never have passed.** `live-image-digest-match` read the pin with a line-offset `grep`, which returned empty once comments were added above `version:` — and the check then degraded to reporting "not pinned yet" instead of failing. It reported that while a digest *was* pinned. Now parsed as YAML. A verification step that cannot fail is worth less than none, because it is trusted. ## Decide per-publisher identity ```task id: RCP-WP-0002-T05 status: done priority: medium state_hub_task_id: "1252f3a1-12ce-5bc1-b003-3e8999d03459" ``` Blocked on a decision in `canned-prompts`, recorded here because it is the thing that decides what this deployment is *for*. **Done 2026-09-08**, upstream in `canned-prompts` (image 0.2.0, migration `0003`). The decision was already made and research found it rather than my judgement supplying it: **DR-3, resolved 2026-07-10** — app-local accounts, platform OIDC demand-gated on client SSO requests, instance consolidation, or local-account toil across more than two apps. None has fired here and no Keycloak is deployed, so OIDC was ruled out by fleet decision. App-local publisher tokens, because a registry is consumed by CLIs and agents: no browser, no session, and a login surface nothing uses is a liability. The entire authentication boundary stays in `auth.py`, so contract § 2.3 holds and a later OIDC switch is bounded. § 20.1 namespace ownership is now a real access decision — `owner` names a publisher. A closed namespace with **no** owner admits nobody, including the operator, because reading a missing owner as "anyone" would invert the point of closing it. **Still open, deliberately:** the NetworkPolicy ingress rule admits any namespace. Publisher identity now gates *writes*, so this is no longer the only control, but it should narrow once the set of legitimate callers is known. Recorded rather than tightened on a guess. **Reversal on record.** I asked `rapp-postgres` not to issue `creds/canned-prompts-publish`, then asked for it. Both were right in their moment: with one indistinguishable identity there was nothing worth authenticating; with real publishers the operator token becomes the bootstrap path that mints the first one. Requested as a reversal rather than quietly asking for the opposite of the earlier argument.