Continuously ready across a full credential lease rotation: 11 samples, zero not-ready, zero query failures, zero kubelet readiness 503s with a probe every 5 seconds, no restarts, pod uptime well past the 30-minute lease TTL. Recorded from a working instrument on the third attempt. The first two verdicts were FAILED and both were the watcher's own defects — an empty kubectl result counted as an outage, then grep -c's exit status poisoning a clean count. The verdict was not talked around; the instrument was fixed and the measurement repeated. Evidence: docs/evidence/RCP-WP-0002-T04-lease-rotation-2026-09-08.log Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
212 lines
8.7 KiB
Markdown
212 lines
8.7 KiB
Markdown
---
|
|
id: RCP-WP-0002
|
|
type: workplan
|
|
title: "First deployment of canned-prompts on Railiance"
|
|
domain: agents
|
|
repo: rapp-canned-prompts
|
|
status: active
|
|
owner: codex
|
|
topic_slug: practice
|
|
created: "2026-09-06"
|
|
updated: "2026-09-06"
|
|
state_hub_workstream_id: "11874f32-ac36-5bb9-a0a5-e7a259f5972c"
|
|
---
|
|
|
|
# First deployment of canned-prompts on Railiance
|
|
|
|
The package is written and validated; what remains needs credentials and a
|
|
published image, which are operator actions rather than authoring ones.
|
|
|
|
`readiness_state` is `draft` and stays there until T04 produces evidence.
|
|
Topology is not readiness (ADR-0006): binding this rapp to a reef does not make
|
|
it deployed.
|
|
|
|
## Publish the image and pin its digest
|
|
|
|
```task
|
|
id: RCP-WP-0002-T01
|
|
status: done
|
|
priority: high
|
|
state_hub_task_id: "3e7faf50-8fe9-5ef3-8a75-1d9e586d6c0f"
|
|
```
|
|
|
|
`declarations/rapp.yaml` carries `version: pending-publication`, and both
|
|
manifests carry a matching tag. This is deliberate: a placeholder shaped like a
|
|
real digest could be mistaken for something deployable, and a rapp that pins
|
|
nothing is a stub that tells the fleet tooling a deployment exists when none
|
|
does.
|
|
|
|
The image builds and was verified locally in `canned-prompts`
|
|
(`CANP-WP-0006-T06`): it starts non-root, all six service-level smoke checks
|
|
pass against the running container, and the reference CLI installs a package
|
|
from it over HTTP.
|
|
|
|
**Done, 2026-09-07.** Published as
|
|
`forgejo.coulomb.social/coulomb/canned-prompts:0.1.0` and pinned by **digest**
|
|
in the declaration and both manifests:
|
|
|
|
```
|
|
sha256:e0ded3c7fe2548c910445deaa31ec42e84f129a92aa324841e231a53fee1f123
|
|
```
|
|
|
|
Pinned by digest rather than tag deliberately: a tag can be moved, and
|
|
`live-image-digest-match` would then pass against something that is no longer
|
|
what this repo reviewed.
|
|
|
|
Verified before pushing — the image resolves a file-mounted credential, carries
|
|
the reference validator, and ships the migration scripts — and verified after,
|
|
by fetching the manifest back **by digest** rather than trusting the push
|
|
output. `readiness_state` moved `draft` → `declared`.
|
|
|
|
## Provision database roles and credentials
|
|
|
|
```task
|
|
id: RCP-WP-0002-T02
|
|
status: done
|
|
priority: high
|
|
state_hub_task_id: "61f5cbd7-8fcf-5156-87da-57239ae55d8f"
|
|
```
|
|
|
|
Two roles, deliberately separate: a runtime role that reads and writes rows, and
|
|
a migration role that owns the schema. A service able to `ALTER` its own tables
|
|
at runtime turns any code defect into a schema defect.
|
|
|
|
Needed from `rapp-postgres` and the credential broker:
|
|
|
|
- database `canned_prompts` on `platform-pg-2`;
|
|
- `creds/canned-prompts-runtime` and `creds/canned-prompts-migration` in OpenBao
|
|
under the `database` path;
|
|
- the `openbao-canned-prompts-eso-token` secret in `external-secrets`.
|
|
|
|
`creds/canned-prompts-publish` is **optional**. Without it the service is
|
|
read-only, which is the correct posture until per-publisher identity exists —
|
|
not a misconfiguration to be worked around.
|
|
|
|
**Requested, 2026-09-07.** `rapp-postgres` now carries
|
|
`consumers/canned-prompts.yaml`, authored against its documented
|
|
`PostgresConsumer` shape, and its agent has the request in its inbox.
|
|
|
|
The declaration only. `make provision-consumers` — the step that actually mints
|
|
OpenBao credentials — was deliberately **not** run: credential issuance is
|
|
`rapp-postgres`' to perform, and running it from the consuming side would take a
|
|
decision that is not this repo's, however available the script happens to be.
|
|
|
|
**Received 2026-09-08.** `rapp-postgres` provisioned and sent a database-owner
|
|
receipt with 12 checks proven live: runtime DDL denied (SQLSTATE 42501),
|
|
statement timeouts and `search_path` as declared, and both logins refused
|
|
CONNECT on `sbom_nexus` — the cell is shared, so that last one matters.
|
|
|
|
`creds/canned-prompts-publish` was **not** created, as asked. The service runs
|
|
read-only, which is the intended posture rather than a gap.
|
|
|
|
## Apply and migrate
|
|
|
|
```task
|
|
id: RCP-WP-0002-T03
|
|
status: done
|
|
priority: high
|
|
state_hub_task_id: "0b8d206e-bd28-5950-abf7-d824015c09a4"
|
|
```
|
|
|
|
Order matters, and the ordering is the point rather than a convenience:
|
|
|
|
Verified against the live cluster before writing: `platform-pg-2` exists and is
|
|
healthy, and the `railiance.io/postgres-client: platform-pg-2` label matches
|
|
what `sbom-nexus` actually carries in the cluster rather than only what its repo
|
|
says.
|
|
|
|
1. `manifests/00-namespace.yaml` — the `railiance.io/postgres-client` label is
|
|
what lets the database namespace accept traffic;
|
|
2. `manifests/database-secrets.yaml`, then wait for the secrets to materialize;
|
|
3. `manifests/migration.yaml` — the Job runs `alembic upgrade head` under the
|
|
migration role and must complete before any replica serves;
|
|
4. `manifests/runtime.yaml`.
|
|
|
|
The runtime deliberately does not migrate at start-up. Migrations as a Job keep
|
|
a schema rollback separate from a code rollback and stop replicas racing each
|
|
other.
|
|
|
|
**Done 2026-09-08, after four defects that only a real rollout could expose.**
|
|
Each is recorded in
|
|
`docs/evidence/RCP-WP-0002-T04-first-deployment-2026-09-08.md`; the two worth
|
|
repeating here failed *silently*:
|
|
|
|
- `SET ROLE` opened an implicit transaction that Alembic then nested inside
|
|
rather than owning, so every revision logged as applied and was rolled back.
|
|
Alembic reported success against an empty database.
|
|
- The egress NetworkPolicy selected `app.kubernetes.io/name`, which the
|
|
migration Job does not carry. The Job matched only the default-deny and
|
|
succeeded exactly once — because it ran before the policies existed. The next
|
|
migration would have failed with a DNS error. Now selects `part-of`, and
|
|
ingress is a separate policy so the Job is never reachable.
|
|
|
|
`rapp-postgres` asked to be told when the first revision landed so they can
|
|
re-run ownership reconciliation. Notified.
|
|
|
|
## Verify and record evidence
|
|
|
|
```task
|
|
id: RCP-WP-0002-T04
|
|
status: done
|
|
priority: high
|
|
state_hub_task_id: "6d9eb97c-57e2-5b74-b6eb-455076713417"
|
|
```
|
|
|
|
Run `tools/smoke.sh`. It checks what only the cluster can answer — secrets
|
|
materialized, Service is ClusterIP with no Ingress, NetworkPolicies present,
|
|
live image digest matches the pin — and calls `canned-prompts`'
|
|
`service/tools/smoke.py` for health and migration head rather than holding a
|
|
second opinion about whether the service is healthy.
|
|
|
|
Record the output as evidence, then move `readiness_state` to `deployed`, and to
|
|
`verified` only with that evidence attached.
|
|
|
|
**Done 2026-09-08.** All six deployment checks and all six service-level checks
|
|
pass; evidence at
|
|
`docs/evidence/RCP-WP-0002-T04-first-deployment-2026-09-08.md`.
|
|
**Then corrected.** I recorded `verified` after a passing smoke run and found
|
|
the deployment `0/1` eight hours later: platform credentials are 30-minute
|
|
leases, and the service read the mounted URL once at start-up. External Secrets
|
|
kept the file current; the engine held the URL it booted with. Fixed in 0.1.5 by
|
|
re-reading the credential for every new connection, with `pool_recycle` inside
|
|
the lease TTL.
|
|
|
|
`readiness_state` was returned to `deployed`, and is now **`verified`** on
|
|
evidence: continuously ready across a full lease rotation, zero readiness
|
|
failures, no restarts. Log at
|
|
`docs/evidence/RCP-WP-0002-T04-lease-rotation-2026-09-08.log`.
|
|
|
|
The bar for `verified` is observation across a rotation, because a smoke run
|
|
inside the first window cannot tell a service that works from one that works
|
|
*once* — an availability property is not provable by a single sample. It took
|
|
three attempts to build an instrument that could return that verdict honestly;
|
|
all three failures are recorded in the evidence file.
|
|
|
|
**The check that could never have passed.** `live-image-digest-match` read the
|
|
pin with a line-offset `grep`, which returned empty once comments were added
|
|
above `version:` — and the check then degraded to reporting "not pinned yet"
|
|
instead of failing. It reported that while a digest *was* pinned. Now parsed as
|
|
YAML. A verification step that cannot fail is worth less than none, because it
|
|
is trusted.
|
|
|
|
## Decide per-publisher identity
|
|
|
|
```task
|
|
id: RCP-WP-0002-T05
|
|
status: wait
|
|
priority: medium
|
|
state_hub_task_id: "1252f3a1-12ce-5bc1-b003-3e8999d03459"
|
|
```
|
|
|
|
Blocked on a decision in `canned-prompts`, recorded here because it is the
|
|
thing that decides what this deployment is *for*.
|
|
|
|
Today the service authenticates a single shared bearer token proving "the
|
|
operator". That is adequate for a private in-cluster registry and inadequate for
|
|
the collaborative prompting platform the operator described: every token holder
|
|
is indistinguishable, so § 20.1 namespace ownership can be enforced against
|
|
anonymous callers but not attributed among publishers.
|
|
|
|
This repo must not paper over that with cluster configuration implying finer
|
|
control than exists. When identity lands upstream, revisit the publish-token
|
|
secret and the NetworkPolicy ingress rule, which currently admits any namespace.
|