Mark workplan active. Add Traefik ForwardAuth middleware and Ingress manifests for activity.coulomb.social and activity-temporal.coulomb.social. Prefer Authelia SSO identity for ops mutations; document DNS gate and fleet pattern (docs/ops-sso-access.md).
143 lines
6.9 KiB
Markdown
143 lines
6.9 KiB
Markdown
# Railiance01 Kubernetes Deployment
|
|
|
|
This bundle establishes activity-core as an internal production service on the
|
|
railiance01 K3s cluster. It keeps the unauthenticated API as a ClusterIP service;
|
|
publish it through an authenticated ingress only after choosing the final host
|
|
name and access policy.
|
|
|
|
## Layout
|
|
|
|
- `00-namespace.yaml`: namespace and shared labels
|
|
- `10-infrastructure.yaml`: PostgreSQL for app data, PostgreSQL for Temporal,
|
|
NATS JetStream, Temporal, and Temporal UI
|
|
- `15-externalsecret-issue-core.yaml`: OpenBao → `ISSUE_CORE_API_KEY` merge into
|
|
`actcore-runtime-secret` via External Secrets
|
|
- `15-externalsecret-forgejo-admin.yaml`: OpenBao → `FORGEJO_TOKEN` merge for
|
|
weekly package prune (ACTIVITY-WP-0023-T05)
|
|
- `20-runtime.yaml`: migrate/sync jobs plus API, worker, and event-router
|
|
- `bootstrap-secrets.sh`: idempotently creates generated Kubernetes secrets
|
|
|
|
The runtime image tag is `activity-core:railiance01-prod` and is expected to be
|
|
loaded into the railiance01 K3s containerd image store.
|
|
|
|
`20-runtime.yaml` also projects the disabled Custodian-owned
|
|
`ops-service-inventory-probes.md` ActivityDefinition and a non-secret
|
|
`actcore-ops-service-inventory` ConfigMap snapshot. The source of truth for the
|
|
inventory source of truth remains `custodian://ops/service-inventory.yml`; update
|
|
the ConfigMap projection from that file before enabling the probe schedule.
|
|
`OPS_HUB_KEY` is created only as an empty Secret placeholder until the operator
|
|
provisions the Inter-Hub ops-hub key.
|
|
|
|
`ISSUE_SINK_TYPE` defaults to **`state-hub`** (ACTIVITY-WP-0022; no silent Forgejo
|
|
issues). Set `rest` only for intentional issue-core projection when the backend
|
|
is healthy. `ISSUE_CORE_API_KEY` and `FORGEJO_TOKEN` are synced from OpenBao into
|
|
`actcore-runtime-secret` by ExternalSecrets:
|
|
|
|
| ExternalSecret | OpenBao path | Secret key |
|
|
| --- | --- | --- |
|
|
| `actcore-issue-core-runtime` | `platform/workloads/issue-core/issue-core/issue-core-runtime` | `ISSUE_CORE_API_KEY` |
|
|
| `actcore-forgejo-admin` | `platform/workloads/forgejo/forgejo-admin` (`API_TOKEN`) | `FORGEJO_TOKEN` |
|
|
|
|
Prereqs: `ClusterSecretStore/openbao-activity-core` and ESO token bootstrap with
|
|
both read policies (`OPENBAO_TOKEN_FILE=~/.local/openbao/platform-admin.token
|
|
./scripts/openbao-eso-token-apply.sh` — attaches
|
|
`workload-kv-read-issue-core-runtime` + `workload-kv-read-forgejo-admin`).
|
|
Roll back to audit mode by setting `ISSUE_SINK_TYPE=null` and restarting worker
|
|
and event-router deployments. See `docs/issue-core-emission-boundary.md`.
|
|
|
|
The same runtime projection now includes the active
|
|
`daily-statehub-wsjf-triage.md` ActivityDefinition plus its JSON output schema
|
|
and a persistent working-memory volume mounted at
|
|
`/var/custodian/memory/working` (hostPath → `/home/tegwick/the-custodian/memory/working`).
|
|
Before trusting the daily 07:20
|
|
Europe/Berlin schedule, verify both runtime dependencies:
|
|
|
|
- `actcore-statehub-edge-relay` is ready and reports upstream reachability at
|
|
`GET /edge/health` (upstream is the in-cluster State Hub API at
|
|
`state-hub.state-hub.svc.cluster.local:8000`). `STATE_HUB_URL` points at the
|
|
relay so allowlisted `GET` reads can be served from cache during brief upstream
|
|
outages and queueable writes survive until replay.
|
|
- `LLM_CONNECT_URL` points at the verified in-namespace llm-connect Service,
|
|
`http://llm-connect.activity-core.svc.cluster.local:8080`, and the
|
|
operator-owned provider Secret lets that Service serve the
|
|
`custodian-triage-balanced` profile.
|
|
|
|
If `LLM_CONNECT_URL` is missing or broken, report-sink instructions write a
|
|
visible `execution_failed` diagnostic instead of silently producing no report.
|
|
|
|
## Deploy
|
|
|
|
```bash
|
|
docker build -t activity-core:railiance01-prod .
|
|
docker save -o /tmp/activity-core-railiance01-prod.tar activity-core:railiance01-prod
|
|
scp /tmp/activity-core-railiance01-prod.tar railiance01:/tmp/
|
|
ssh railiance01 sudo k3s ctr images import /tmp/activity-core-railiance01-prod.tar
|
|
rsync -a k8s/railiance/ railiance01:activity-core/k8s/railiance/
|
|
|
|
ssh railiance01
|
|
cd ~/activity-core
|
|
bash k8s/railiance/bootstrap-secrets.sh
|
|
kubectl apply -f k8s/railiance/10-infrastructure.yaml
|
|
# Bootstrap OpenBao ESO token + apply ExternalSecrets (once per cluster):
|
|
OPENBAO_TOKEN_FILE=~/.local/openbao/platform-admin.token ./scripts/openbao-eso-token-apply.sh
|
|
kubectl apply -f ~/railiance-platform/argocd/platform-addons/openbao-secretstore/openbao-activity-core.clustersecretstore.yaml
|
|
kubectl apply -f k8s/railiance/15-externalsecret-issue-core.yaml
|
|
kubectl apply -f k8s/railiance/15-externalsecret-forgejo-admin.yaml
|
|
kubectl -n activity-core wait --for=condition=Ready externalsecret/actcore-issue-core-runtime --timeout=120s
|
|
kubectl -n activity-core wait --for=condition=Ready externalsecret/actcore-forgejo-admin --timeout=120s
|
|
kubectl -n activity-core wait --for=condition=ready pod -l app.kubernetes.io/name=actcore-app-db --timeout=180s
|
|
kubectl -n activity-core wait --for=condition=ready pod -l app.kubernetes.io/name=actcore-temporal-db --timeout=180s
|
|
kubectl -n activity-core wait --for=condition=ready pod -l app.kubernetes.io/name=actcore-nats --timeout=180s
|
|
kubectl -n activity-core rollout status deploy/actcore-temporal --timeout=300s
|
|
|
|
kubectl -n activity-core delete job actcore-migrate --ignore-not-found
|
|
kubectl apply -f k8s/railiance/20-runtime.yaml
|
|
kubectl -n activity-core wait --for=condition=complete job/actcore-migrate --timeout=180s
|
|
kubectl -n activity-core rollout status deploy/actcore-api --timeout=180s
|
|
kubectl -n activity-core rollout status deploy/actcore-worker --timeout=180s
|
|
kubectl -n activity-core rollout status deploy/actcore-event-router --timeout=180s
|
|
kubectl -n activity-core delete job actcore-sync --ignore-not-found
|
|
kubectl apply -f k8s/railiance/20-runtime.yaml
|
|
kubectl -n activity-core wait --for=condition=complete job/actcore-sync --timeout=180s
|
|
```
|
|
|
|
## Verify
|
|
|
|
```bash
|
|
kubectl -n activity-core exec deploy/actcore-api -- \
|
|
python -c "import urllib.request; print(urllib.request.urlopen('http://localhost:8010/health').read().decode())"
|
|
|
|
kubectl -n activity-core get pods
|
|
kubectl -n activity-core get svc
|
|
```
|
|
|
|
## Operator automation console (ACTIVITY-WP-0024 / 0025)
|
|
|
|
### SSO (primary, after DNS)
|
|
|
|
```bash
|
|
# DNS A records → 92.205.62.239 (once):
|
|
# activity.coulomb.social
|
|
# activity-temporal.coulomb.social
|
|
|
|
kubectl apply -f k8s/railiance/30-authelia-middleware.yaml
|
|
kubectl apply -f k8s/railiance/31-ingress-ops-sso.yaml
|
|
kubectl apply -f k8s/railiance/32-ingress-temporal-sso.yaml
|
|
kubectl -n activity-core set env deploy/actcore-api \
|
|
ACTIVITY_CORE_TEMPORAL_UI_URL=https://activity-temporal.coulomb.social
|
|
```
|
|
|
|
- Ops: https://activity.coulomb.social/ops/ui (Authelia SSO)
|
|
- Temporal: https://activity-temporal.coulomb.social
|
|
- Design: `docs/ops-sso-access.md`
|
|
|
|
### Break-glass port-forward
|
|
|
|
```bash
|
|
export KUBECONFIG=~/.kube/config-hosteurope
|
|
kubectl -n activity-core port-forward svc/actcore-api 8010:8010
|
|
# UI: http://127.0.0.1:8010/ops/ui
|
|
```
|
|
|
|
Mutations: SSO headers when behind Authelia, else `X-Operator-Token` from
|
|
`actcore-runtime-secret`. Cron edits remain git-owned.
|