rapp-canned-prompts/workplans/RCP-WP-0002-first-deployment.md
tegwick 4d5af7698d First deployment: verified on railiance01
RCP-WP-0002 T02-T04 done, readiness_state verified with evidence attached
rather than ahead of it.

rapp-postgres provisioned canned_prompts on platform-pg-2 and sent a
database-owner receipt with 12 checks proven live. creds/canned-prompts-publish
was deliberately not issued, so the service runs read-only and POST /packages
returns 503 explaining why — the intended posture, not a gap.

Four defects surfaced that only a real rollout could expose, two of them
silent:

- SET ROLE opened an implicit transaction that Alembic nested inside rather
  than owning, so every revision logged as applied and was rolled back.
  Alembic reported success against an empty database.
- The egress NetworkPolicy selected app.kubernetes.io/name, which the
  migration Job does not carry. The Job matched only the default-deny and
  succeeded exactly once, because it ran before the policies existed; the next
  migration would have failed on DNS. Now selects part-of, with ingress split
  into its own policy so the Job is never reachable.
- env.py read database_url rather than resolved_database_url, so the migration
  could never run where the credential is a mounted file.
- live-image-digest-match extracted the pin with a line-offset grep, which
  returned empty once comments were added above `version:`. The check degraded
  to reporting "not pinned yet" while a digest was pinned — it could not have
  passed for any pin. Now parsed as YAML. A verification step that cannot fail
  is worth less than none, because it is trusted.

Evidence: docs/evidence/RCP-WP-0002-T04-first-deployment-2026-09-08.md

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-08 08:56:02 +02:00

197 lines
7.8 KiB
Markdown

---
id: RCP-WP-0002
type: workplan
title: "First deployment of canned-prompts on Railiance"
domain: agents
repo: rapp-canned-prompts
status: finished
owner: codex
topic_slug: practice
created: "2026-09-06"
updated: "2026-09-06"
state_hub_workstream_id: "11874f32-ac36-5bb9-a0a5-e7a259f5972c"
---
# First deployment of canned-prompts on Railiance
The package is written and validated; what remains needs credentials and a
published image, which are operator actions rather than authoring ones.
`readiness_state` is `draft` and stays there until T04 produces evidence.
Topology is not readiness (ADR-0006): binding this rapp to a reef does not make
it deployed.
## Publish the image and pin its digest
```task
id: RCP-WP-0002-T01
status: done
priority: high
state_hub_task_id: "3e7faf50-8fe9-5ef3-8a75-1d9e586d6c0f"
```
`declarations/rapp.yaml` carries `version: pending-publication`, and both
manifests carry a matching tag. This is deliberate: a placeholder shaped like a
real digest could be mistaken for something deployable, and a rapp that pins
nothing is a stub that tells the fleet tooling a deployment exists when none
does.
The image builds and was verified locally in `canned-prompts`
(`CANP-WP-0006-T06`): it starts non-root, all six service-level smoke checks
pass against the running container, and the reference CLI installs a package
from it over HTTP.
**Done, 2026-09-07.** Published as
`forgejo.coulomb.social/coulomb/canned-prompts:0.1.0` and pinned by **digest**
in the declaration and both manifests:
```
sha256:e0ded3c7fe2548c910445deaa31ec42e84f129a92aa324841e231a53fee1f123
```
Pinned by digest rather than tag deliberately: a tag can be moved, and
`live-image-digest-match` would then pass against something that is no longer
what this repo reviewed.
Verified before pushing — the image resolves a file-mounted credential, carries
the reference validator, and ships the migration scripts — and verified after,
by fetching the manifest back **by digest** rather than trusting the push
output. `readiness_state` moved `draft``declared`.
## Provision database roles and credentials
```task
id: RCP-WP-0002-T02
status: done
priority: high
state_hub_task_id: "61f5cbd7-8fcf-5156-87da-57239ae55d8f"
```
Two roles, deliberately separate: a runtime role that reads and writes rows, and
a migration role that owns the schema. A service able to `ALTER` its own tables
at runtime turns any code defect into a schema defect.
Needed from `rapp-postgres` and the credential broker:
- database `canned_prompts` on `platform-pg-2`;
- `creds/canned-prompts-runtime` and `creds/canned-prompts-migration` in OpenBao
under the `database` path;
- the `openbao-canned-prompts-eso-token` secret in `external-secrets`.
`creds/canned-prompts-publish` is **optional**. Without it the service is
read-only, which is the correct posture until per-publisher identity exists —
not a misconfiguration to be worked around.
**Requested, 2026-09-07.** `rapp-postgres` now carries
`consumers/canned-prompts.yaml`, authored against its documented
`PostgresConsumer` shape, and its agent has the request in its inbox.
The declaration only. `make provision-consumers` — the step that actually mints
OpenBao credentials — was deliberately **not** run: credential issuance is
`rapp-postgres`' to perform, and running it from the consuming side would take a
decision that is not this repo's, however available the script happens to be.
**Received 2026-09-08.** `rapp-postgres` provisioned and sent a database-owner
receipt with 12 checks proven live: runtime DDL denied (SQLSTATE 42501),
statement timeouts and `search_path` as declared, and both logins refused
CONNECT on `sbom_nexus` — the cell is shared, so that last one matters.
`creds/canned-prompts-publish` was **not** created, as asked. The service runs
read-only, which is the intended posture rather than a gap.
## Apply and migrate
```task
id: RCP-WP-0002-T03
status: done
priority: high
state_hub_task_id: "0b8d206e-bd28-5950-abf7-d824015c09a4"
```
Order matters, and the ordering is the point rather than a convenience:
Verified against the live cluster before writing: `platform-pg-2` exists and is
healthy, and the `railiance.io/postgres-client: platform-pg-2` label matches
what `sbom-nexus` actually carries in the cluster rather than only what its repo
says.
1. `manifests/00-namespace.yaml` — the `railiance.io/postgres-client` label is
what lets the database namespace accept traffic;
2. `manifests/database-secrets.yaml`, then wait for the secrets to materialize;
3. `manifests/migration.yaml` — the Job runs `alembic upgrade head` under the
migration role and must complete before any replica serves;
4. `manifests/runtime.yaml`.
The runtime deliberately does not migrate at start-up. Migrations as a Job keep
a schema rollback separate from a code rollback and stop replicas racing each
other.
**Done 2026-09-08, after four defects that only a real rollout could expose.**
Each is recorded in
`docs/evidence/RCP-WP-0002-T04-first-deployment-2026-09-08.md`; the two worth
repeating here failed *silently*:
- `SET ROLE` opened an implicit transaction that Alembic then nested inside
rather than owning, so every revision logged as applied and was rolled back.
Alembic reported success against an empty database.
- The egress NetworkPolicy selected `app.kubernetes.io/name`, which the
migration Job does not carry. The Job matched only the default-deny and
succeeded exactly once — because it ran before the policies existed. The next
migration would have failed with a DNS error. Now selects `part-of`, and
ingress is a separate policy so the Job is never reachable.
`rapp-postgres` asked to be told when the first revision landed so they can
re-run ownership reconciliation. Notified.
## Verify and record evidence
```task
id: RCP-WP-0002-T04
status: done
priority: high
state_hub_task_id: "6d9eb97c-57e2-5b74-b6eb-455076713417"
```
Run `tools/smoke.sh`. It checks what only the cluster can answer — secrets
materialized, Service is ClusterIP with no Ingress, NetworkPolicies present,
live image digest matches the pin — and calls `canned-prompts`'
`service/tools/smoke.py` for health and migration head rather than holding a
second opinion about whether the service is healthy.
Record the output as evidence, then move `readiness_state` to `deployed`, and to
`verified` only with that evidence attached.
**Done 2026-09-08.** All six deployment checks and all six service-level checks
pass; evidence at
`docs/evidence/RCP-WP-0002-T04-first-deployment-2026-09-08.md`.
`readiness_state` is `verified`, with the evidence attached rather than ahead of
it.
**The check that could never have passed.** `live-image-digest-match` read the
pin with a line-offset `grep`, which returned empty once comments were added
above `version:` — and the check then degraded to reporting "not pinned yet"
instead of failing. It reported that while a digest *was* pinned. Now parsed as
YAML. A verification step that cannot fail is worth less than none, because it
is trusted.
## Decide per-publisher identity
```task
id: RCP-WP-0002-T05
status: wait
priority: medium
state_hub_task_id: "1252f3a1-12ce-5bc1-b003-3e8999d03459"
```
Blocked on a decision in `canned-prompts`, recorded here because it is the
thing that decides what this deployment is *for*.
Today the service authenticates a single shared bearer token proving "the
operator". That is adequate for a private in-cluster registry and inadequate for
the collaborative prompting platform the operator described: every token holder
is indistinguishable, so § 20.1 namespace ownership can be enforced against
anonymous callers but not attributed among publishers.
This repo must not paper over that with cluster configuration implying finer
control than exists. When identity lands upstream, revisit the publish-token
secret and the NetworkPolicy ingress rule, which currently admits any namespace.