activity-core/k8s/railiance/README.md
tegwick fee89c4ea1
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 2s
Finish ACTIVITY-WP-0025 after NK-WP-0021 group allowlist.
Close T06: LLDAP activity-core-operators and Authelia domain rules are live
in net-kingdom. Mark the workplan finished, update G10/runbook/SSO design
with membership pointers, and clear residual handoff notes.
2026-07-22 17:47:57 +02:00

146 lines
7.3 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Railiance01 Kubernetes Deployment
This bundle establishes activity-core as an internal production service on the
railiance01 K3s cluster. Services remain ClusterIP; browser access to the ops
console and Temporal UI is via Traefik + Authelia SSO Ingress
(`activity.coulomb.social`, `temporal.coulomb.social` — ACTIVITY-WP-0025).
## Layout
- `00-namespace.yaml`: namespace and shared labels
- `10-infrastructure.yaml`: PostgreSQL for app data, PostgreSQL for Temporal,
NATS JetStream, Temporal, and Temporal UI
- `15-externalsecret-issue-core.yaml`: OpenBao → `ISSUE_CORE_API_KEY` merge into
`actcore-runtime-secret` via External Secrets
- `15-externalsecret-forgejo-admin.yaml`: OpenBao → `FORGEJO_TOKEN` merge for
weekly package prune (ACTIVITY-WP-0023-T05)
- `20-runtime.yaml`: migrate/sync jobs plus API, worker, and event-router
- `bootstrap-secrets.sh`: idempotently creates generated Kubernetes secrets
The runtime image tag is `activity-core:railiance01-prod` and is expected to be
loaded into the railiance01 K3s containerd image store.
`20-runtime.yaml` also projects the disabled Custodian-owned
`ops-service-inventory-probes.md` ActivityDefinition and a non-secret
`actcore-ops-service-inventory` ConfigMap snapshot. The source of truth for the
inventory source of truth remains `custodian://ops/service-inventory.yml`; update
the ConfigMap projection from that file before enabling the probe schedule.
`OPS_HUB_KEY` is created only as an empty Secret placeholder until the operator
provisions the Inter-Hub ops-hub key.
`ISSUE_SINK_TYPE` defaults to **`state-hub`** (ACTIVITY-WP-0022; no silent Forgejo
issues). Set `rest` only for intentional issue-core projection when the backend
is healthy. `ISSUE_CORE_API_KEY` and `FORGEJO_TOKEN` are synced from OpenBao into
`actcore-runtime-secret` by ExternalSecrets:
| ExternalSecret | OpenBao path | Secret key |
| --- | --- | --- |
| `actcore-issue-core-runtime` | `platform/workloads/issue-core/issue-core/issue-core-runtime` | `ISSUE_CORE_API_KEY` |
| `actcore-forgejo-admin` | `platform/workloads/forgejo/forgejo-admin` (`API_TOKEN`) | `FORGEJO_TOKEN` |
Prereqs: `ClusterSecretStore/openbao-activity-core` and ESO token bootstrap with
both read policies (`OPENBAO_TOKEN_FILE=~/.local/openbao/platform-admin.token
./scripts/openbao-eso-token-apply.sh` — attaches
`workload-kv-read-issue-core-runtime` + `workload-kv-read-forgejo-admin`).
Roll back to audit mode by setting `ISSUE_SINK_TYPE=null` and restarting worker
and event-router deployments. See `docs/issue-core-emission-boundary.md`.
The same runtime projection now includes the active
`daily-statehub-wsjf-triage.md` ActivityDefinition plus its JSON output schema
and a persistent working-memory volume mounted at
`/var/custodian/memory/working` (hostPath → `/home/tegwick/the-custodian/memory/working`).
Before trusting the daily 07:20
Europe/Berlin schedule, verify both runtime dependencies:
- `actcore-statehub-edge-relay` is ready and reports upstream reachability at
`GET /edge/health` (upstream is the in-cluster State Hub API at
`state-hub.state-hub.svc.cluster.local:8000`). `STATE_HUB_URL` points at the
relay so allowlisted `GET` reads can be served from cache during brief upstream
outages and queueable writes survive until replay.
- `LLM_CONNECT_URL` points at the verified in-namespace llm-connect Service,
`http://llm-connect.activity-core.svc.cluster.local:8080`, and the
operator-owned provider Secret lets that Service serve the
`custodian-triage-balanced` profile.
If `LLM_CONNECT_URL` is missing or broken, report-sink instructions write a
visible `execution_failed` diagnostic instead of silently producing no report.
## Deploy
```bash
docker build -t activity-core:railiance01-prod .
docker save -o /tmp/activity-core-railiance01-prod.tar activity-core:railiance01-prod
scp /tmp/activity-core-railiance01-prod.tar railiance01:/tmp/
ssh railiance01 sudo k3s ctr images import /tmp/activity-core-railiance01-prod.tar
rsync -a k8s/railiance/ railiance01:activity-core/k8s/railiance/
ssh railiance01
cd ~/activity-core
bash k8s/railiance/bootstrap-secrets.sh
kubectl apply -f k8s/railiance/10-infrastructure.yaml
# Bootstrap OpenBao ESO token + apply ExternalSecrets (once per cluster):
OPENBAO_TOKEN_FILE=~/.local/openbao/platform-admin.token ./scripts/openbao-eso-token-apply.sh
kubectl apply -f ~/railiance-platform/argocd/platform-addons/openbao-secretstore/openbao-activity-core.clustersecretstore.yaml
kubectl apply -f k8s/railiance/15-externalsecret-issue-core.yaml
kubectl apply -f k8s/railiance/15-externalsecret-forgejo-admin.yaml
kubectl -n activity-core wait --for=condition=Ready externalsecret/actcore-issue-core-runtime --timeout=120s
kubectl -n activity-core wait --for=condition=Ready externalsecret/actcore-forgejo-admin --timeout=120s
kubectl -n activity-core wait --for=condition=ready pod -l app.kubernetes.io/name=actcore-app-db --timeout=180s
kubectl -n activity-core wait --for=condition=ready pod -l app.kubernetes.io/name=actcore-temporal-db --timeout=180s
kubectl -n activity-core wait --for=condition=ready pod -l app.kubernetes.io/name=actcore-nats --timeout=180s
kubectl -n activity-core rollout status deploy/actcore-temporal --timeout=300s
kubectl -n activity-core delete job actcore-migrate --ignore-not-found
kubectl apply -f k8s/railiance/20-runtime.yaml
kubectl -n activity-core wait --for=condition=complete job/actcore-migrate --timeout=180s
kubectl -n activity-core rollout status deploy/actcore-api --timeout=180s
kubectl -n activity-core rollout status deploy/actcore-worker --timeout=180s
kubectl -n activity-core rollout status deploy/actcore-event-router --timeout=180s
kubectl -n activity-core delete job actcore-sync --ignore-not-found
kubectl apply -f k8s/railiance/20-runtime.yaml
kubectl -n activity-core wait --for=condition=complete job/actcore-sync --timeout=180s
```
## Verify
```bash
kubectl -n activity-core exec deploy/actcore-api -- \
python -c "import urllib.request; print(urllib.request.urlopen('http://localhost:8010/health').read().decode())"
kubectl -n activity-core get pods
kubectl -n activity-core get svc
```
## Operator automation console (ACTIVITY-WP-0024 / 0025)
### SSO (primary — live)
Manifests `30-``32-*.yaml` are applied; TLS certs Ready; Authelia ForwardAuth
redirects unauthenticated browsers to `auth.coulomb.social`.
```bash
# Re-apply if needed:
kubectl apply -f k8s/railiance/30-authelia-middleware.yaml
kubectl apply -f k8s/railiance/31-ingress-ops-sso.yaml
kubectl apply -f k8s/railiance/32-ingress-temporal-sso.yaml
kubectl -n activity-core set env deploy/actcore-api \
ACTIVITY_CORE_TEMPORAL_UI_URL=https://temporal.coulomb.social
kubectl -n activity-core set env deploy/actcore-temporal-ui \
TEMPORAL_CORS_ORIGINS=https://temporal.coulomb.social,http://localhost:8080,http://127.0.0.1:8080
```
- Ops: https://activity.coulomb.social/ops/ui (Authelia SSO; group `activity-core-operators`)
- Temporal: https://temporal.coulomb.social
- Design: `docs/ops-sso-access.md`
- Membership: `net-kingdom/sso-mfa/k8s/lldap/OPERATOR-GROUPS.md` (NK-WP-0021)
### Break-glass port-forward
```bash
export KUBECONFIG=~/.kube/config-hosteurope
kubectl -n activity-core port-forward svc/actcore-api 8010:8010
# UI: http://127.0.0.1:8010/ops/ui
```
Mutations: SSO headers when behind Authelia, else `X-Operator-Token` from
`actcore-runtime-secret`. Cron edits remain git-owned.