reuse-surface/docs/deploy/reuse-kubernetes.md
tegwick b035664890
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 1s
ci / validate-registry (push) Successful in 2m44s
Build and Publish Container Image / build-and-push (push) Successful in 49s
Correct deploy guide for the Gitea registry retirement (REUSE-WP-0020-T05)
The guide described gitea.coulomb.social as the live registry and its manual
build commands still pushed there. CoulombCore is switched off 2026-08-31.

Also record what the registry actually contains: authenticated tags/list shows
latest, main-bca7165, main-f9d957a and no e3ae22e — so the tag pinned in
railiance-apps/helm/reuse-surface-values.yaml cannot be pulled today, making
ImagePullBackOff a present risk on any restart rather than a future one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 00:06:07 +02:00

165 lines
No EOL
6.6 KiB
Markdown

# reuse-surface Service — Kubernetes Deployment
Companion to **RAILIANCE-WP-0007** (`railiance-apps` Helm release).
## Image
**`gitea.coulomb.social` is going away.** CoulombCore, which hosts it, is
switched off **2026-08-31**. Do not push images there and do not deploy
anything that references it. Tracked as **REUSE-WP-0020**.
This repo's canonical remote migrated to Forgejo in REUSE-WP-0019-T03
(`origin``forgejo-remote:coulomb/reuse-surface.git`; the old Gitea copy is
read-only, not deleted). `.forgejo/workflows/image.yaml` builds and pushes
`forgejo.coulomb.social/coulomb/reuse-surface:latest` and `:main-<short-sha>`
on every push that touches `Dockerfile`/`reuse_surface/**`/`schemas/**`/
`pyproject.toml`. Prefer promoting a CI-built tag over building by hand.
`railiance-apps@04be416` already repointed `charts/reuse-surface/values.yaml`
to the Forgejo repository, but `helm/reuse-surface-values.yaml` still pins
`image.tag: "e3ae22e"` — a commit from 2026-07-07 18:25, whereas CI only began
publishing to Forgejo at 21:25 that day. **A Forgejo `:e3ae22e` tag most likely
never existed**, so the deployment as written risks `ImagePullBackOff` on any
restart or reschedule today, not only after the retirement date. Verify the tag
exists before relying on the current manifest. See REUSE-WP-0020-T05.
If you must build by hand:
```bash
docker build -t forgejo.coulomb.social/coulomb/reuse-surface:<tag> .
docker push forgejo.coulomb.social/coulomb/reuse-surface:<tag>
```
## Required environment
| Variable | Purpose |
|---|---|
| `REUSE_SURFACE_TOKEN` | Bearer token for write API |
| `REUSE_SURFACE_FORGEJO_WEBHOOK_SECRET` | HMAC for `POST /v1/webhooks/forgejo` |
| `REUSE_SURFACE_DB` | SQLite path (default `/data/reuse.db`) |
| `REUSE_SURFACE_CACHE_DIR` | Remote index cache (default `/data/cache`) |
Mount a PVC at `/data` for persistence. Production injects secrets via
`ExternalSecret reuse-surface-runtime` → Kubernetes Secret `reuse-surface-env`
(OpenBao path `platform/workloads/reuse/reuse-surface/runtime-secrets`,
CCR-2026-0005). Operator runbook: `railiance-apps/docs/reuse-surface-on-railiance01.md`;
rotation: `railiance-platform/docs/reuse-surface-runtime-secrets-rotation-runbook.md`.
## Probes
- Liveness/readiness: `GET /health` on port `8000`
## Browser landing page
Production ingress routes HTTPS `/` to a static landing Deployment
(`reuse-surface-landing`, **RAILIANCE-WP-0008**). API paths are unchanged:
- `/health` and `/v1/*` → hub service container
- `/` → informational HTML for browser visitors (no login, no secrets)
Agents and CLI clients should target `/health` and `/v1/*` only, not `/`.
## Public URL and DNS
| Item | Value |
|---|---|
| URL | `https://reuse.coulomb.social` |
| DNS A record | **`92.205.62.239`** (Railiance01 production) |
CoulombCore (`92.205.130.254`) held a bootstrap deploy; production release uses
`KUBECONFIG=~/.kube/config-hosteurope`. Verify propagation:
```bash
dig +short reuse.coulomb.social A # must return 92.205.62.239
```
## Client configuration
```bash
export REUSE_SURFACE_URL=https://reuse.coulomb.social
export REUSE_SURFACE_TOKEN=<write-token>
reuse-surface hub status
```
## Operational hardening
The hub runs as a single-replica Deployment with SQLite on a PVC (**A5**
containerized service). **A6** (managed platform) is deferred until multi-replica
or Postgres backing is required.
### Backup and restore (SQLite PVC)
1. Identify the PVC mounted at `/data` (stores `reuse.db` and remote index cache).
2. Snapshot or copy while the pod is running (SQLite WAL-safe copy) or scale to
zero briefly for a cold copy:
```bash
kubectl -n <namespace> exec deploy/reuse-surface -- \
sqlite3 /data/reuse.db '.backup /tmp/reuse-backup.db'
kubectl -n <namespace> cp deploy/reuse-surface:/tmp/reuse-backup.db ./reuse-backup.db
```
3. Restore by replacing `/data/reuse.db` from backup and restarting the pod.
4. Re-register repos if the database is empty (`reuse-surface hub list`).
Verify backup once per environment after deploy changes.
### TLS certificate renewal
Ingress TLS is managed by the cluster cert issuer (Railiance01 companion chart).
Monitor certificate expiry on `reuse.coulomb.social`. Renewal is automatic when
the issuer is healthy; on failure, check ingress secret `reuse-surface-tls` and
cert-manager / companion operator logs.
### Token rotation
1. Generate a new `REUSE_SURFACE_TOKEN` value.
2. Update Kubernetes Secret `reuse-surface-env`.
3. Rolling restart the hub Deployment.
4. Update operator workstations and CI secrets that call write endpoints.
5. Confirm `reuse-surface hub register` fails with the old token and succeeds
with the new token.
### Image promotion checklist
1. Tag image from CI commit. `.forgejo/workflows/image.yaml` already builds
and pushes `forgejo.coulomb.social/coulomb/reuse-surface:main-<short-sha>`
automatically on every push that touches `Dockerfile`/`reuse_surface/**`/
`schemas/**`/`pyproject.toml` — verify it's green rather than building by
hand. Confirm the tag exists in the registry before bumping the manifest;
`GET /v2/coulomb/reuse-surface/tags/list` needs registry credentials.
2. Run `pytest -q` and `reuse-surface validate` on that commit (CI already
does this; re-verify locally if promoting outside CI).
3. Update Helm values image tag in `railiance-apps`
(`helm/reuse-surface-values.yaml`).
4. Deploy to Railiance01 (`make reuse-deploy`); verify `GET /v1/federated`
and `GET /v1/repos` (not `GET /health` — see the known ingress routing
issue below).
5. Smoke `reuse-surface hub list` and `GET /v1/federated` capability count;
check `reuse-surface stats` shows a fresh `composed_at` post-deploy.
6. Record image digest in workplan or progress log.
**Known issue (found 2026-07-07, not fixed):** the public ingress's
exact-path `/health` rule 404s (shadowed by the catch-all `/` rule to the
landing page) — confirmed ingress-layer only via direct port-forward and
the Deployment's own passing readiness/liveness probes. Use
`GET /v1/repos` or `GET /v1/federated` for external verification instead.
Flagged to `railiance-apps`; not this repo's fix to make (shared ingress
template).
### SQLite vs Postgres (cnpg) — decision criteria
Stay on SQLite while:
- Single replica is acceptable.
- RPO of occasional PVC snapshot is sufficient.
- Write volume is low (repo registration changes only).
Consider Postgres (e.g. CloudNative-PG) when:
- Multiple hub replicas or zero-downtime failover is required.
- RPO/RTO targets need point-in-time recovery beyond PVC snapshots.
- Federation cache metadata or audit tables grow beyond comfortable SQLite size.
**Implementation deferred** unless an operator approves migration. Document only
until then.