Continuously ready across a full credential lease rotation: 11 samples, zero not-ready, zero query failures, zero kubelet readiness 503s with a probe every 5 seconds, no restarts, pod uptime well past the 30-minute lease TTL. Recorded from a working instrument on the third attempt. The first two verdicts were FAILED and both were the watcher's own defects — an empty kubectl result counted as an outage, then grep -c's exit status poisoning a clean count. The verdict was not talked around; the instrument was fixed and the measurement repeated. Evidence: docs/evidence/RCP-WP-0002-T04-lease-rotation-2026-09-08.log Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
8.7 KiB
| id | type | title | domain | repo | status | owner | topic_slug | created | updated | state_hub_workstream_id |
|---|---|---|---|---|---|---|---|---|---|---|
| RCP-WP-0002 | workplan | First deployment of canned-prompts on Railiance | agents | rapp-canned-prompts | active | codex | practice | 2026-09-06 | 2026-09-06 | 11874f32-ac36-5bb9-a0a5-e7a259f5972c |
First deployment of canned-prompts on Railiance
The package is written and validated; what remains needs credentials and a published image, which are operator actions rather than authoring ones.
readiness_state is draft and stays there until T04 produces evidence.
Topology is not readiness (ADR-0006): binding this rapp to a reef does not make
it deployed.
Publish the image and pin its digest
id: RCP-WP-0002-T01
status: done
priority: high
state_hub_task_id: "3e7faf50-8fe9-5ef3-8a75-1d9e586d6c0f"
declarations/rapp.yaml carries version: pending-publication, and both
manifests carry a matching tag. This is deliberate: a placeholder shaped like a
real digest could be mistaken for something deployable, and a rapp that pins
nothing is a stub that tells the fleet tooling a deployment exists when none
does.
The image builds and was verified locally in canned-prompts
(CANP-WP-0006-T06): it starts non-root, all six service-level smoke checks
pass against the running container, and the reference CLI installs a package
from it over HTTP.
Done, 2026-09-07. Published as
forgejo.coulomb.social/coulomb/canned-prompts:0.1.0 and pinned by digest
in the declaration and both manifests:
sha256:e0ded3c7fe2548c910445deaa31ec42e84f129a92aa324841e231a53fee1f123
Pinned by digest rather than tag deliberately: a tag can be moved, and
live-image-digest-match would then pass against something that is no longer
what this repo reviewed.
Verified before pushing — the image resolves a file-mounted credential, carries
the reference validator, and ships the migration scripts — and verified after,
by fetching the manifest back by digest rather than trusting the push
output. readiness_state moved draft → declared.
Provision database roles and credentials
id: RCP-WP-0002-T02
status: done
priority: high
state_hub_task_id: "61f5cbd7-8fcf-5156-87da-57239ae55d8f"
Two roles, deliberately separate: a runtime role that reads and writes rows, and
a migration role that owns the schema. A service able to ALTER its own tables
at runtime turns any code defect into a schema defect.
Needed from rapp-postgres and the credential broker:
- database
canned_promptsonplatform-pg-2; creds/canned-prompts-runtimeandcreds/canned-prompts-migrationin OpenBao under thedatabasepath;- the
openbao-canned-prompts-eso-tokensecret inexternal-secrets.
creds/canned-prompts-publish is optional. Without it the service is
read-only, which is the correct posture until per-publisher identity exists —
not a misconfiguration to be worked around.
Requested, 2026-09-07. rapp-postgres now carries
consumers/canned-prompts.yaml, authored against its documented
PostgresConsumer shape, and its agent has the request in its inbox.
The declaration only. make provision-consumers — the step that actually mints
OpenBao credentials — was deliberately not run: credential issuance is
rapp-postgres' to perform, and running it from the consuming side would take a
decision that is not this repo's, however available the script happens to be.
Received 2026-09-08. rapp-postgres provisioned and sent a database-owner
receipt with 12 checks proven live: runtime DDL denied (SQLSTATE 42501),
statement timeouts and search_path as declared, and both logins refused
CONNECT on sbom_nexus — the cell is shared, so that last one matters.
creds/canned-prompts-publish was not created, as asked. The service runs
read-only, which is the intended posture rather than a gap.
Apply and migrate
id: RCP-WP-0002-T03
status: done
priority: high
state_hub_task_id: "0b8d206e-bd28-5950-abf7-d824015c09a4"
Order matters, and the ordering is the point rather than a convenience:
Verified against the live cluster before writing: platform-pg-2 exists and is
healthy, and the railiance.io/postgres-client: platform-pg-2 label matches
what sbom-nexus actually carries in the cluster rather than only what its repo
says.
manifests/00-namespace.yaml— therailiance.io/postgres-clientlabel is what lets the database namespace accept traffic;manifests/database-secrets.yaml, then wait for the secrets to materialize;manifests/migration.yaml— the Job runsalembic upgrade headunder the migration role and must complete before any replica serves;manifests/runtime.yaml.
The runtime deliberately does not migrate at start-up. Migrations as a Job keep a schema rollback separate from a code rollback and stop replicas racing each other.
Done 2026-09-08, after four defects that only a real rollout could expose.
Each is recorded in
docs/evidence/RCP-WP-0002-T04-first-deployment-2026-09-08.md; the two worth
repeating here failed silently:
SET ROLEopened an implicit transaction that Alembic then nested inside rather than owning, so every revision logged as applied and was rolled back. Alembic reported success against an empty database.- The egress NetworkPolicy selected
app.kubernetes.io/name, which the migration Job does not carry. The Job matched only the default-deny and succeeded exactly once — because it ran before the policies existed. The next migration would have failed with a DNS error. Now selectspart-of, and ingress is a separate policy so the Job is never reachable.
rapp-postgres asked to be told when the first revision landed so they can
re-run ownership reconciliation. Notified.
Verify and record evidence
id: RCP-WP-0002-T04
status: done
priority: high
state_hub_task_id: "6d9eb97c-57e2-5b74-b6eb-455076713417"
Run tools/smoke.sh. It checks what only the cluster can answer — secrets
materialized, Service is ClusterIP with no Ingress, NetworkPolicies present,
live image digest matches the pin — and calls canned-prompts'
service/tools/smoke.py for health and migration head rather than holding a
second opinion about whether the service is healthy.
Record the output as evidence, then move readiness_state to deployed, and to
verified only with that evidence attached.
Done 2026-09-08. All six deployment checks and all six service-level checks
pass; evidence at
docs/evidence/RCP-WP-0002-T04-first-deployment-2026-09-08.md.
Then corrected. I recorded verified after a passing smoke run and found
the deployment 0/1 eight hours later: platform credentials are 30-minute
leases, and the service read the mounted URL once at start-up. External Secrets
kept the file current; the engine held the URL it booted with. Fixed in 0.1.5 by
re-reading the credential for every new connection, with pool_recycle inside
the lease TTL.
readiness_state was returned to deployed, and is now verified on
evidence: continuously ready across a full lease rotation, zero readiness
failures, no restarts. Log at
docs/evidence/RCP-WP-0002-T04-lease-rotation-2026-09-08.log.
The bar for verified is observation across a rotation, because a smoke run
inside the first window cannot tell a service that works from one that works
once — an availability property is not provable by a single sample. It took
three attempts to build an instrument that could return that verdict honestly;
all three failures are recorded in the evidence file.
The check that could never have passed. live-image-digest-match read the
pin with a line-offset grep, which returned empty once comments were added
above version: — and the check then degraded to reporting "not pinned yet"
instead of failing. It reported that while a digest was pinned. Now parsed as
YAML. A verification step that cannot fail is worth less than none, because it
is trusted.
Decide per-publisher identity
id: RCP-WP-0002-T05
status: wait
priority: medium
state_hub_task_id: "1252f3a1-12ce-5bc1-b003-3e8999d03459"
Blocked on a decision in canned-prompts, recorded here because it is the
thing that decides what this deployment is for.
Today the service authenticates a single shared bearer token proving "the operator". That is adequate for a private in-cluster registry and inadequate for the collaborative prompting platform the operator described: every token holder is indistinguishable, so § 20.1 namespace ownership can be enforced against anonymous callers but not attributed among publishers.
This repo must not paper over that with cluster configuration implying finer control than exists. When identity lands upstream, revisit the publish-token secret and the NetworkPolicy ingress rule, which currently admits any namespace.