rapp-canned-prompts/declarations/rapp.yaml

65 lines
2.1 KiB
YAML
Raw Normal View History

Package canned-prompts for Railiance Registers the repo with State Hub (agents / practice, prefix RCP-WP) and fills in the rapp shape. declarations/rapp.yaml declares a manifest-managed platform service owned by canned-prompts, bound to rail-kubernetes and reef-railiance, with rollout, smoke and rollback contracts. The image pin says `pending-publication` rather than carrying a placeholder digest. The image builds and was verified locally (canned-prompts CANP-WP-0006-T06) but has never been pushed, so no registry digest exists. A placeholder shaped like a real digest would be worse than a sentinel: it could be mistaken for something deployable. manifests/ follows the rapp-sbom-nexus shape: namespace labelled for the postgres client, external secrets from OpenBao, a migration Job, and the runtime Deployment with a ClusterIP-only Service, dedicated ServiceAccount and default-deny plus runtime NetworkPolicies. Three choices worth stating. Credentials arrive as mounted files, never env vars — an env var holding a password is visible in kubectl describe, in crash dumps, and to anything that can read /proc. Migrations run as a Job rather than at start-up, so a schema rollback stays separate from a code rollback and replicas do not race. Liveness points at /healthz, which checks only that the process is up: pointing it at a database-dependent path would restart every replica during a database blip. Egress is PostgreSQL and DNS only. A package arrives by publish; the registry never reaches out, so it is given no path to. tools/smoke.sh checks what only the cluster can answer and calls canned-prompts' service/tools/smoke.py for health and migration head, rather than holding a second opinion about whether the service is healthy. RCP-WP-0002 carries the two operator actions that block a first rollout — publishing the image and provisioning database roles — and records per-publisher identity as a decision belonging upstream, which this repo must not paper over with cluster configuration implying finer control than exists. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 21:44:17 +02:00
kind: managed-workload-package
repo_family: rapp
rapp_id: rapp-canned-prompts
repo: rapp-canned-prompts
ownership_repo: canned-prompts
contract_version: 1.0.0
readiness_state: verified
Package canned-prompts for Railiance Registers the repo with State Hub (agents / practice, prefix RCP-WP) and fills in the rapp shape. declarations/rapp.yaml declares a manifest-managed platform service owned by canned-prompts, bound to rail-kubernetes and reef-railiance, with rollout, smoke and rollback contracts. The image pin says `pending-publication` rather than carrying a placeholder digest. The image builds and was verified locally (canned-prompts CANP-WP-0006-T06) but has never been pushed, so no registry digest exists. A placeholder shaped like a real digest would be worse than a sentinel: it could be mistaken for something deployable. manifests/ follows the rapp-sbom-nexus shape: namespace labelled for the postgres client, external secrets from OpenBao, a migration Job, and the runtime Deployment with a ClusterIP-only Service, dedicated ServiceAccount and default-deny plus runtime NetworkPolicies. Three choices worth stating. Credentials arrive as mounted files, never env vars — an env var holding a password is visible in kubectl describe, in crash dumps, and to anything that can read /proc. Migrations run as a Job rather than at start-up, so a schema rollback stays separate from a code rollback and replicas do not race. Liveness points at /healthz, which checks only that the process is up: pointing it at a database-dependent path would restart every replica during a database blip. Egress is PostgreSQL and DNS only. A package arrives by publish; the registry never reaches out, so it is given no path to. tools/smoke.sh checks what only the cluster can answer and calls canned-prompts' service/tools/smoke.py for health and migration head, rather than holding a second opinion about whether the service is healthy. RCP-WP-0002 carries the two operator actions that block a first rollout — publishing the image and provisioning database roles — and records per-publisher identity as a decision belonging upstream, which this repo must not paper over with cluster configuration implying finer control than exists. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 21:44:17 +02:00
workload_identity:
name: canned-prompts
package_type: manifest-managed-platform-service
data_classification: internal
criticality: medium
primary_rail: rail-kubernetes
supported_rails:
- rail-kubernetes
bound_reefs:
- reef-railiance
runtime_dependencies:
- kubernetes-api
- openbao-database-secrets-engine
composition:
purpose: >-
Package and operate the canned-prompts hosted registry and index on
Railiance. Format semantics, package validation, API compatibility, the
schema and its migrations, and image publication remain owned by
canned-prompts. PostgreSQL topology, database isolation, backups and
credential issuance remain owned by rapp-postgres and the platform
credential broker.
member_repos:
- repo: rapp-canned-prompts
role: managed runtime package
deployables:
- canned-prompts
upstream_components:
- name: canned-prompts
source: forgejo.coulomb.social/coulomb/canned-prompts
Correct a premature `verified`, and survive credential rotation I recorded readiness_state: verified after a passing smoke run and found the deployment 0/1 eight hours later. Platform credentials are 30-minute leases, not passwords. The service read the mounted URL once at start-up and never again, so External Secrets kept the file current while the engine held the URL it booted with, and every reconnection after the first lease expiry used a credential the database had already revoked. /readyz reported it accurately — "database unreachable: OperationalError" — and the pod sat unready for eight hours without being restarted, because liveness is deliberately independent of the database. That separation behaved exactly as designed: the process was alive and could not serve, and it said so. Fixed in 0.1.5: the engine re-reads the credential for every new connection and pool_recycle is 900s, well inside the lease. Only username and password come from the refreshed URL; host, port and database come from the engine, so a malformed refresh cannot silently redirect the service. readiness_state back to deployed. The bar for verified is now observation across a full lease rotation, because a smoke run inside the first window cannot distinguish a service that works from one that works once — an availability property is not provable by a single sample. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-08 09:01:10 +02:00
# Published 2026-09-07 from canned-prompts service/Dockerfile, tag 0.1.5.
Publish the image, pin it by digest, and request the database RCP-WP-0002-T01 done. Published forgejo.coulomb.social/coulomb/canned-prompts:0.1.0 and pinned sha256:e0ded3c7fe25... in declarations/rapp.yaml and both manifests. Pinned by digest rather than tag: a tag can be moved, and live-image-digest-match would then pass against something that is no longer what this repo reviewed. Verified after the push by fetching the manifest back by digest rather than trusting the push output. readiness_state draft -> declared. Not deployed, so not `deployed`. T02 requested rather than performed. rapp-postgres now carries consumers/canned-prompts.yaml against its documented PostgresConsumer shape, and its agent has the request. `make provision-consumers`, which mints the OpenBao credentials, was deliberately not run: credential issuance belongs to that repo's operator, and running it from the consuming side would take a decision that is not this repo's, however available the script is. T03 and T04 move to wait behind it. Also corrected a check of my own: I briefly read the cluster list as lacking platform-pg-2 and suspected the manifests targeted a host that does not exist. That was my own truncated output. platform-pg-2 is present and healthy, and the postgres-client label matches what sbom-nexus actually carries in the cluster rather than only what its repo says. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-07 08:43:35 +02:00
# Pinned by digest rather than tag: a tag can be moved, and
# live-image-digest-match would then pass against something that is no
# longer what this repo reviewed.
Correct a premature `verified`, and survive credential rotation I recorded readiness_state: verified after a passing smoke run and found the deployment 0/1 eight hours later. Platform credentials are 30-minute leases, not passwords. The service read the mounted URL once at start-up and never again, so External Secrets kept the file current while the engine held the URL it booted with, and every reconnection after the first lease expiry used a credential the database had already revoked. /readyz reported it accurately — "database unreachable: OperationalError" — and the pod sat unready for eight hours without being restarted, because liveness is deliberately independent of the database. That separation behaved exactly as designed: the process was alive and could not serve, and it said so. Fixed in 0.1.5: the engine re-reads the credential for every new connection and pool_recycle is 900s, well inside the lease. Only username and password come from the refreshed URL; host, port and database come from the engine, so a malformed refresh cannot silently redirect the service. readiness_state back to deployed. The bar for verified is now observation across a full lease rotation, because a smoke run inside the first window cannot distinguish a service that works from one that works once — an availability property is not provable by a single sample. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-08 09:01:10 +02:00
version: sha256:14c7b92f20d63f2e70ea17b0b45d3c483521fbbaf14ee5bc277fa541e10452e2
Package canned-prompts for Railiance Registers the repo with State Hub (agents / practice, prefix RCP-WP) and fills in the rapp shape. declarations/rapp.yaml declares a manifest-managed platform service owned by canned-prompts, bound to rail-kubernetes and reef-railiance, with rollout, smoke and rollback contracts. The image pin says `pending-publication` rather than carrying a placeholder digest. The image builds and was verified locally (canned-prompts CANP-WP-0006-T06) but has never been pushed, so no registry digest exists. A placeholder shaped like a real digest would be worse than a sentinel: it could be mistaken for something deployable. manifests/ follows the rapp-sbom-nexus shape: namespace labelled for the postgres client, external secrets from OpenBao, a migration Job, and the runtime Deployment with a ClusterIP-only Service, dedicated ServiceAccount and default-deny plus runtime NetworkPolicies. Three choices worth stating. Credentials arrive as mounted files, never env vars — an env var holding a password is visible in kubectl describe, in crash dumps, and to anything that can read /proc. Migrations run as a Job rather than at start-up, so a schema rollback stays separate from a code rollback and replicas do not race. Liveness points at /healthz, which checks only that the process is up: pointing it at a database-dependent path would restart every replica during a database blip. Egress is PostgreSQL and DNS only. A package arrives by publish; the registry never reaches out, so it is given no path to. tools/smoke.sh checks what only the cluster can answer and calls canned-prompts' service/tools/smoke.py for health and migration head, rather than holding a second opinion about whether the service is healthy. RCP-WP-0002 carries the two operator actions that block a first rollout — publishing the image and provisioning database roles — and records per-publisher identity as a decision belonging upstream, which this repo must not paper over with cluster configuration implying finer control than exists. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 21:44:17 +02:00
rollout_contract:
default_mode: kubectl-server-side-apply
smoke_contract:
required:
- state-health-ok
- migration-at-head
- external-secrets-ready
- private-service-only
- networkpolicies-present
- live-image-digest-match
rollback_contract:
order:
- previous-immutable-image-digest
- apply-reviewed-git-revision
source_documents:
- repo: canned-prompts
path: CannedPromptFormat.md
- repo: canned-prompts
path: service/README.md
- repo: canned-prompts
path: workplans/CANP-WP-0006-hosted-registry-service.md
- repo: repo-manager
path: docs/RailianceAppDeploymentGuide.md