rapp-canned-prompts/workplans/RCP-WP-0002-first-deployment.md
tegwick 5a0f4cb5c8 Package canned-prompts for Railiance
Registers the repo with State Hub (agents / practice, prefix RCP-WP) and fills
in the rapp shape.

declarations/rapp.yaml declares a manifest-managed platform service owned by
canned-prompts, bound to rail-kubernetes and reef-railiance, with rollout,
smoke and rollback contracts.

The image pin says `pending-publication` rather than carrying a placeholder
digest. The image builds and was verified locally (canned-prompts
CANP-WP-0006-T06) but has never been pushed, so no registry digest exists. A
placeholder shaped like a real digest would be worse than a sentinel: it could
be mistaken for something deployable.

manifests/ follows the rapp-sbom-nexus shape: namespace labelled for the
postgres client, external secrets from OpenBao, a migration Job, and the
runtime Deployment with a ClusterIP-only Service, dedicated ServiceAccount and
default-deny plus runtime NetworkPolicies.

Three choices worth stating. Credentials arrive as mounted files, never env
vars — an env var holding a password is visible in kubectl describe, in crash
dumps, and to anything that can read /proc. Migrations run as a Job rather than
at start-up, so a schema rollback stays separate from a code rollback and
replicas do not race. Liveness points at /healthz, which checks only that the
process is up: pointing it at a database-dependent path would restart every
replica during a database blip.

Egress is PostgreSQL and DNS only. A package arrives by publish; the registry
never reaches out, so it is given no path to.

tools/smoke.sh checks what only the cluster can answer and calls
canned-prompts' service/tools/smoke.py for health and migration head, rather
than holding a second opinion about whether the service is healthy.

RCP-WP-0002 carries the two operator actions that block a first rollout —
publishing the image and provisioning database roles — and records
per-publisher identity as a decision belonging upstream, which this repo must
not paper over with cluster configuration implying finer control than exists.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 21:44:17 +02:00

134 lines
4.7 KiB
Markdown

---
id: RCP-WP-0002
type: workplan
title: "First deployment of canned-prompts on Railiance"
domain: agents
repo: rapp-canned-prompts
status: proposed
owner: codex
topic_slug: practice
created: "2026-09-06"
updated: "2026-09-06"
state_hub_workstream_id: "11874f32-ac36-5bb9-a0a5-e7a259f5972c"
---
# First deployment of canned-prompts on Railiance
The package is written and validated; what remains needs credentials and a
published image, which are operator actions rather than authoring ones.
`readiness_state` is `draft` and stays there until T04 produces evidence.
Topology is not readiness (ADR-0006): binding this rapp to a reef does not make
it deployed.
## Publish the image and pin its digest
```task
id: RCP-WP-0002-T01
status: todo
priority: high
state_hub_task_id: "3e7faf50-8fe9-5ef3-8a75-1d9e586d6c0f"
```
`declarations/rapp.yaml` carries `version: pending-publication`, and both
manifests carry a matching tag. This is deliberate: a placeholder shaped like a
real digest could be mistaken for something deployable, and a rapp that pins
nothing is a stub that tells the fleet tooling a deployment exists when none
does.
The image builds and was verified locally in `canned-prompts`
(`CANP-WP-0006-T06`): it starts non-root, all six service-level smoke checks
pass against the running container, and the reference CLI installs a package
from it over HTTP.
What remains is publication to `forgejo.coulomb.social/coulomb/canned-prompts`,
which needs registry credentials. Then replace `pending-publication` in
`declarations/rapp.yaml`, `manifests/runtime.yaml` and `manifests/migration.yaml`
with the `@sha256:` digest — the same digest in all three, since
`live-image-digest-match` compares them.
## Provision database roles and credentials
```task
id: RCP-WP-0002-T02
status: todo
priority: high
state_hub_task_id: "61f5cbd7-8fcf-5156-87da-57239ae55d8f"
```
Two roles, deliberately separate: a runtime role that reads and writes rows, and
a migration role that owns the schema. A service able to `ALTER` its own tables
at runtime turns any code defect into a schema defect.
Needed from `rapp-postgres` and the credential broker:
- database `canned_prompts` on `platform-pg-2`;
- `creds/canned-prompts-runtime` and `creds/canned-prompts-migration` in OpenBao
under the `database` path;
- the `openbao-canned-prompts-eso-token` secret in `external-secrets`.
`creds/canned-prompts-publish` is **optional**. Without it the service is
read-only, which is the correct posture until per-publisher identity exists —
not a misconfiguration to be worked around.
## Apply and migrate
```task
id: RCP-WP-0002-T03
status: todo
priority: high
state_hub_task_id: "0b8d206e-bd28-5950-abf7-d824015c09a4"
```
Order matters, and the ordering is the point rather than a convenience:
1. `manifests/00-namespace.yaml` — the `railiance.io/postgres-client` label is
what lets the database namespace accept traffic;
2. `manifests/database-secrets.yaml`, then wait for the secrets to materialize;
3. `manifests/migration.yaml` — the Job runs `alembic upgrade head` under the
migration role and must complete before any replica serves;
4. `manifests/runtime.yaml`.
The runtime deliberately does not migrate at start-up. Migrations as a Job keep
a schema rollback separate from a code rollback and stop replicas racing each
other.
## Verify and record evidence
```task
id: RCP-WP-0002-T04
status: todo
priority: high
state_hub_task_id: "6d9eb97c-57e2-5b74-b6eb-455076713417"
```
Run `tools/smoke.sh`. It checks what only the cluster can answer — secrets
materialized, Service is ClusterIP with no Ingress, NetworkPolicies present,
live image digest matches the pin — and calls `canned-prompts`'
`service/tools/smoke.py` for health and migration head rather than holding a
second opinion about whether the service is healthy.
Record the output as evidence, then move `readiness_state` to `deployed`, and to
`verified` only with that evidence attached.
## Decide per-publisher identity
```task
id: RCP-WP-0002-T05
status: wait
priority: medium
state_hub_task_id: "1252f3a1-12ce-5bc1-b003-3e8999d03459"
```
Blocked on a decision in `canned-prompts`, recorded here because it is the
thing that decides what this deployment is *for*.
Today the service authenticates a single shared bearer token proving "the
operator". That is adequate for a private in-cluster registry and inadequate for
the collaborative prompting platform the operator described: every token holder
is indistinguishable, so § 20.1 namespace ownership can be enforced against
anonymous callers but not attributed among publishers.
This repo must not paper over that with cluster configuration implying finer
control than exists. When identity lands upstream, revisit the publish-token
secret and the NetworkPolicy ingress rule, which currently admits any namespace.