From 87f12967141c5a8ecf44877dc37454f09d095bfc Mon Sep 17 00:00:00 2001 From: tegwick Date: Fri, 21 Aug 2026 16:13:38 +0200 Subject: [PATCH] docs: STATE-WP-0081 cluster self-sufficiency and registrar fix The registrar is broken by workstation coupling rather than beside it. Inside the pod /home/tegwick mounts read-only despite the chart setting no readOnly and securityContext being empty, so the sweep dies before it can mint. That is why 12 sync requests queued and the newest custodian-sync commits are July. Eight tasks: restore the write path, give the pod its own clone volume and service identity, run non-root, correct 72 repo records still pointing at /home/worsch with gitea remote_urls, serve the dashboard from the cluster, write the boundary rule down, and close out the registrar. End state: workstation coding agents push to forgejo, cluster infrastructure reads from forgejo, neither reads the other's disk. Co-Authored-By: Claude Opus 5 --- ...-cluster-self-sufficiency-and-registrar.md | 230 ++++++++++++++++++ 1 file changed, 230 insertions(+) create mode 100644 workplans/STATE-WP-0081-cluster-self-sufficiency-and-registrar.md diff --git a/workplans/STATE-WP-0081-cluster-self-sufficiency-and-registrar.md b/workplans/STATE-WP-0081-cluster-self-sufficiency-and-registrar.md new file mode 100644 index 0000000..aa57338 --- /dev/null +++ b/workplans/STATE-WP-0081-cluster-self-sufficiency-and-registrar.md @@ -0,0 +1,230 @@ +--- +id: STATE-WP-0081 +type: workplan +title: "Cluster self-sufficiency: remove workstation coupling and fix the registrar" +domain: infotech +repo: state-hub +status: proposed +owner: codex +topic_slug: infotech +created: "2026-08-21" +updated: "2026-08-21" +parent_project: prj-state-hub-retirement +parent_workplan: SHR-WP-0001 +related: + - STATE-WP-0079 + - RMGR-WP-0005 + - RMGR-WP-0008 + - ADR-007 + - ADR-010 +--- + +# Cluster self-sufficiency: remove workstation coupling and fix the registrar + +## Goal + +Make the railiance01-hosted State Hub stand on its own: its own clones, its own +git identity, no hostPath into an operator's home directory, and no dependence +on workstation paths, processes, or checkouts. Fix the identifier registrar as +part of that, because the registrar is broken *by* this coupling rather than +beside it. + +End state: **workstation coding agents push to forgejo; cluster infrastructure +reads from forgejo. Neither reads the other's disk.** + +## Admissibility under the retirement freeze + +`policies/retirement-freeze.md` allows changes that fix operational risk or +reduce scope. This is both: the sweep is currently broken in production, and the +work removes a coupling rather than adding capability. It establishes no new +permanent ownership here — the deployment chart already lives in this repo at +`deploy/railiance/apps/charts/state-hub/`. + +## The coupling, as measured 2026-08-21 + +| # | Coupling | Evidence | +| --- | --- | --- | +| 1 | Pod mounts the operator home as a hostPath | `sweep.hostPath: /home/tegwick`, mounted rw at the same path | +| 2 | Pod runs as root | `securityContext: {}`, no `runAsUser`; produced **999 root-owned files** under `~/state-hub` | +| 3 | Pod uses the operator's personal SSH key as its service identity | `sweep.sshHostPath: /home/tegwick/.ssh` → `/root/.ssh` | +| 4 | Hub repo registry points at workstation paths | **72 of 75** records have `local_path: /home/worsch/...`, a path railiance01 will never have | +| 5 | Hub repo registry records the retired git host | `remote_url: gitea-remote:...` after the 2026-08-21 forgejo transition | +| 6 | Dashboard is a workstation process | Observable dev server on `127.0.0.1:3000`, not served from the cluster | + +**And the fault that blocks everything:** inside the pod `/home/tegwick` is +mounted **read-only** — `ro,relatime,discard,errors=remount-ro` — although the +chart sets no `readOnly` on that volumeMount and `securityContext` is empty. The +host has the same device mounted `rw` with a healthy disk and no dmesg errors, so +this is imposed by the runtime or an admission path, not by hardware. + +The sweep therefore dies with +`OSError: [Errno 30] Read-only file system: '/home/tegwick/info-tech-canon/.custodian-brief.md'` +before it can mint anything. That is the whole registrar problem: **the registrar +cannot write identifiers into files it cannot write.** It explains the 12 queued +sync requests and why the newest `custodian-sync` commits are from July. + +## Restore the write path + +```task +id: STATE-WP-0081-T01 +status: todo +priority: high +``` + +Find why the hostPath mounts read-only against the chart's own spec, and restore +writes. Candidates in order: an admission controller or PodSecurity policy +forcing hostPath read-only; a k3s/containerd default for hostPath volumes; a +missing explicit `readOnly: false` on the volumeMount. + +This is the minimum to make the registrar function, and the only task here that +unblocks anything today. Do it first even though T02 later deletes the mount — +a working baseline makes every later change verifiable. + +Verification is end-to-end, not a green pod: run the sweep and confirm +`EBIND-WP-0002` gets a `state_hub_workstream_id` written back into its file. + +## Give the pod its own clones + +```task +id: STATE-WP-0081-T02 +status: todo +priority: high +``` + +Replace the `sweep-repos` hostPath with storage the pod owns — a PVC the sweep +clones into and maintains, seeded from forgejo. + +This is the change that actually severs coupling #1 and #2. Mounting a human's +home directory into a production workload is what created 999 root-owned files, +what made `git pull` fail on the host, and what put the operator's checkouts one +`reset --hard` away from a scheduled job. + +Sizing input: the current tree is ~78 repos; `markitect_project` alone is 24 MB. + +Keep the sweep's repo list driven by the hub's repo registry, not by whatever +happens to be on a disk. + +## Replace the operator SSH key with a service identity + +```task +id: STATE-WP-0081-T03 +status: todo +priority: high +``` + +Issue a dedicated forgejo deploy key or service account for the sweep and mount +it from a Secret. Remove `sweep.sshHostPath` and the `/root/.ssh` mount. + +The current arrangement gives a root-running production workload the operator's +personal private key. It is read-only, so this is a blast-radius problem rather +than a live compromise — but the key that can push to every repository in the +fleet should not be the same key a human uses interactively. + +Credential custody routes through OpenBao, not this repo — see +`.claude/rules/credential-routing.md`. Do not put key material in the chart. + +## Run as a non-root user + +```task +id: STATE-WP-0081-T04 +status: todo +priority: medium +``` + +Set `runAsUser`/`runAsGroup` and a `fsGroup` matching the PVC. Depends on T02: +once the pod owns its storage there is no reason for it to be root. + +Closes the recurrence: today's ownership fix will be undone by the next sweep +while the pod still runs as root. + +## Correct the hub repo registry + +```task +id: STATE-WP-0081-T05 +status: todo +priority: high +``` + +72 of 75 repo records carry `local_path: /home/worsch/...` and stale +`remote_url: gitea-remote:...`. Both are wrong for a cluster that reads from +forgejo into its own clone tree. + +Decide first whether `local_path` should be **per-instance rather than global** — +one column cannot describe a workstation checkout and a cluster clone at once, +and its current single value is precisely how the workstation leaked into +cluster configuration. `RMGR-WP-0008` is building the repository-representation +surface that inherits this; settle the shape with it rather than patching values +that will move. + +`remote_url` correction is unambiguous and can proceed immediately. + +## Serve the dashboard from the cluster + +```task +id: STATE-WP-0081-T06 +status: todo +priority: medium +``` + +The dashboard runs as an Observable dev server on the workstation at +`127.0.0.1:3000`. Serve it from the cluster behind the same ingress as the API so +it survives the workstation being off, and so what the operator sees is what the +cluster holds. + +Check first whether this is worth building here at all: `hub-projection-ui` is +dispositioned `replace` → `hub-core` (14 items, slice B5 in +`docs/retirement-cutover-slice-plan.md`). If B5 lands first this task is a +redirect, not a build. Confirm with `HUB-WP-0004` before writing any chart. + +## State and enforce the boundary + +```task +id: STATE-WP-0081-T07 +status: todo +priority: medium +``` + +Write the rule down so it stops being re-derived: **workstation coding agents +push to forgejo; cluster infrastructure reads from forgejo; neither reads the +other's disk.** + +Record it where agents will meet it — `docs/`, the repo `AGENTS.md` templates, and +as an ADR if it constrains other repos, which it does. + +Include the failure this prevents. On 2026-08-21 railiance01's 70 checkouts were +found still pointed at the retired gitea host, six weeks stale, and the +`evidence-binder` workplan the registrar was asked to index **did not exist** in +the copy the cluster could see. A shared-disk assumption made a stale reader look +like a queue. + +## Close out the registrar + +```task +id: STATE-WP-0081-T08 +status: todo +priority: high +``` + +With the write path restored, confirm the registrar end-to-end and drain the +backlog: 12 queued requests from `evidence-binder`, `kaizen-agentic`, +`glas-harness`, and `agentic-resources`, plus this fleet's own unregistered +`RMGR-WP-0008` and `RMGR-WP-0009`. + +Then remove the interim: `RMGR-WP-0005-T03` (UUIDv5 derivation, keyed on +`(namespace, identifier)` and deriving for live records only per the 2026-08-21 +`ADR-007` amendment) makes writeback idempotent and retires the single-writer +rule entirely. Coordinate rather than duplicate — the derivation belongs to +`repo-manager`. + +Reply to the queued agents when it is done; several have been waiting since +2026-08-20. + +## Acceptance + +- [ ] Sweep writes successfully from the pod; `EBIND-WP-0002` registered +- [ ] Pod uses its own clone volume; no hostPath into any home directory +- [ ] Pod authenticates with a dedicated key, runs as non-root +- [ ] No `/home/worsch` path in any cluster-consumed record; `remote_url` values current +- [ ] Dashboard reachable without the workstation, or formally handed to `hub-core` +- [ ] Boundary rule written and discoverable by agents +- [ ] Registrar queue drained; interim single-writer rule retired or explicitly deferred