Closes the gap that made section 20.1 ownership advisory. The service could
refuse anonymous callers but could not tell two publishers apart, so a closed
namespace could be protected and never attributed.
Follows DR-3, resolved 2026-07-10: app-local accounts, with platform OIDC
demand-gated on client SSO requests, instance consolidation, or local-account
toil across more than two apps. None of those triggers has fired here, so this
is deliberately not OIDC. Tokens rather than accounts because a registry is
consumed by CLIs and agents — no browser, no session, no UI to log into, and a
login surface nothing uses is a liability.
The whole authentication boundary stays in auth.py, so contract section 2.3 is
met and a later OIDC switch is bounded rather than a search.
The properties that matter are the ones about what a credential cannot do:
- tokens are stored hashed, because a registry that can print its own
credentials back is one database read away from impersonating every publisher
it knows, and are shown once at creation;
- an unknown token and a wrong token get the same answer, so a caller cannot
enumerate which tokens exist;
- a publisher cannot mint publishers — that would be an administrator with
extra steps, and revoking one would no longer revoke what it could do;
- the operator token publishes but owns nothing, so it is a bootstrap path
rather than an identity that can hold a namespace;
- a closed namespace with no owner recorded admits nobody, including the
operator: reading a missing owner as "anyone" would invert the point of
closing it;
- revocation is a timestamp, not a delete, so what someone published stays
attributed to them after their credential is withdrawn.
Migration 0003 adds publishers and index_entries.published_by. The attribution
is a name rather than a foreign key, so deleting a publisher cannot erase the
history of what they published.
Service tests 49 -> 61.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM
Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
The deployment ran for one lease window and then sat unready for eight hours.
Platform credentials are 30-minute leases, not passwords: the service read the
mounted URL once at start-up, so External Secrets kept the file current while
the engine held the URL it booted with, and every reconnection after the first
expiry used a credential the database had already revoked.
make_engine now takes an optional refresh callable, invoked by a do_connect
hook each time the pool opens a connection, and pool_recycle is 900s so a
pooled connection is retired well inside the lease. Only username and password
are taken from the refreshed URL — host, port and database come from the engine,
so a malformed refresh cannot silently redirect the service somewhere else.
Two things behaved correctly and are worth keeping. /readyz reported the real
cause, "database unreachable: OperationalError", rather than a generic failure.
And liveness stayed independent of the database, so the pod was never
restart-looped: it was alive, unable to serve, and said so. Pointing liveness at
a database-dependent path would have masked this as a crash loop.
Service tests 47 -> 49, including one asserting pool_recycle stays inside the
shortest lease the platform issues.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM
Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
Three defects the test suite could not have caught, because each needed a real
cluster, a mounted secret, or a live PostgreSQL.
env.py read database_url rather than resolved_database_url. `configured` was
true because a file was set, and the field it then read was empty — so Alembic
received an empty URL and the migration could never run in the cluster. The
value is also now escaped for ConfigParser interpolation, since a `%` in a
generated password would otherwise raise at credential rotation, which is the
worst time to find out.
SET ROLE opened an implicit transaction that Alembic then nested inside rather
than owning, so it never committed and leaving the connection block rolled
everything back. Alembic logged "Running upgrade" for every revision against a
database that stayed empty. SET ROLE is session-scoped, so committing
immediately ends the implicit transaction without discarding the role.
A missing optional publish-token file was treated as a hard failure. The
absence is the documented read-only posture — the secret is mounted optional
and deliberately not issued — so treating it as a fault turned an intended
state into a 500 rather than the 503 that explains it. `required` now separates
the two cases: a missing database URL still fails loudly, because there the
silence would hide a real fault.
Migrations also assume the durable owner role rather than creating objects as
the leased migration login, per the rapp-postgres database-owner boundary. The
role name is validated against an identifier pattern because SET ROLE cannot be
parameterised.
Service tests 36 -> 47.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM
Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
The foundation of the hosted registry, in canned-prompts so rapp.yaml gets
ownership_repo: canned-prompts — the sbom-nexus shape, where product ownership
stays out of the operations repo.
Stack matches state-hub and sbom-nexus: FastAPI, SQLAlchemy, Alembic,
PostgreSQL, in service/ with its own environment. reference/ is deliberately
untouched: it is the format's conformance witness and stays dependency-light,
and the service is a separate consumer of the same package semantics.
CANNED_PROMPTS_DATABASE_URL has no default. A service that silently falls back
to a local database when its real one is misconfigured is worse than one that
refuses to start.
Health surface per RailianceAppDeploymentGuide.md: unauthenticated /healthz and
/readyz, plus /state/health for fleet consistency. /healthz deliberately checks
nothing beyond the process being up, so a database blip does not restart pods;
/readyz asks the database something it can fail to answer.
Migration 0001 creates package_versions, package_files and index_entries, every
one carrying a tenant key per business-app-service-contract section 1.3 — the
service is single-tenant today, and the key is present so a later consolidation
is a data copy rather than a rewrite. A test asserts every table in the metadata
is tenant-keyed, so adding an unkeyed table fails the suite rather than being
discovered at consolidation time. Uniqueness is (tenant, registry, package_id,
version): registry-scoped because identity is, tenant-scoped so two tenants may
hold the same id.
The schema keeps the format's three things distinct — an immutable package
version, its files as content rather than parsed rows, and an index entry
recording how a version arrived here.
Fixes a bug its own test caught: check_readiness first caught every failure in
one except and reported "database unreachable", so an unmigrated but perfectly
reachable database sent an operator to credentials and networking when the fix
was alembic upgrade. Connectivity and schema are now checked separately.
Service tests 11 passing; reference tests unaffected at 99.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM
Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502