state-hub/workplans/STATE-WP-0081-cluster-self-sufficiency-and-registrar.md
codex 3cb256ef5c
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 1s
fix(workplans): adopt ADR-007 derived identifiers for unregistered records
These workplans exist only in the retired local hub. Their random pre-ADR-007
identifiers are refused by C-06 as stale references, so they cannot be
registered. Deriving from the canonical record id takes no identity from
anything: central does not hold them and the old ids die with the cache.

Records central already holds were deliberately left untouched.

Refs CUST-WP-0068-T06

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 2583210@bnt-lap001
Assistant-Session: f2bff2d5-e9b2-4338-92ca-10282a927006
2026-08-25 20:22:27 +02:00

9.2 KiB

id type title domain repo status owner topic_slug created updated parent_project parent_workplan related state_hub_workstream_id
STATE-WP-0081 workplan Cluster self-sufficiency: remove workstation coupling and fix the registrar infotech state-hub proposed codex infotech 2026-08-21 2026-08-21 prj-state-hub-retirement SHR-WP-0001
STATE-WP-0079
RMGR-WP-0005
RMGR-WP-0008
ADR-007
ADR-010
bb2798fd-0680-5027-8479-3af3a5b048de

Cluster self-sufficiency: remove workstation coupling and fix the registrar

Goal

Make the railiance01-hosted State Hub stand on its own: its own clones, its own git identity, no hostPath into an operator's home directory, and no dependence on workstation paths, processes, or checkouts. Fix the identifier registrar as part of that, because the registrar is broken by this coupling rather than beside it.

End state: workstation coding agents push to forgejo; cluster infrastructure reads from forgejo. Neither reads the other's disk.

Admissibility under the retirement freeze

policies/retirement-freeze.md allows changes that fix operational risk or reduce scope. This is both: the sweep is currently broken in production, and the work removes a coupling rather than adding capability. It establishes no new permanent ownership here — the deployment chart already lives in this repo at deploy/railiance/apps/charts/state-hub/.

The coupling, as measured 2026-08-21

# Coupling Evidence
1 Pod mounts the operator home as a hostPath sweep.hostPath: /home/tegwick, mounted rw at the same path
2 Pod runs as root securityContext: {}, no runAsUser; produced 999 root-owned files under ~/state-hub
3 Pod uses the operator's personal SSH key as its service identity sweep.sshHostPath: /home/tegwick/.ssh/root/.ssh
4 Hub repo registry points at workstation paths 72 of 75 records have local_path: /home/worsch/..., a path railiance01 will never have
5 Hub repo registry records the retired git host remote_url: gitea-remote:... after the 2026-08-21 forgejo transition
6 Dashboard is a workstation process Observable dev server on 127.0.0.1:3000, not served from the cluster

And the fault that blocks everything: inside the pod /home/tegwick is mounted read-onlyro,relatime,discard,errors=remount-ro — although the chart sets no readOnly on that volumeMount and securityContext is empty. The host has the same device mounted rw with a healthy disk and no dmesg errors, so this is imposed by the runtime or an admission path, not by hardware.

The sweep therefore dies with OSError: [Errno 30] Read-only file system: '/home/tegwick/info-tech-canon/.custodian-brief.md' before it can mint anything. That is the whole registrar problem: the registrar cannot write identifiers into files it cannot write. It explains the 12 queued sync requests and why the newest custodian-sync commits are from July.

Restore the write path

id: STATE-WP-0081-T01
status: todo
priority: high
state_hub_task_id: "7dca680b-d34d-5bfe-838e-b463a4487a9c"

Find why the hostPath mounts read-only against the chart's own spec, and restore writes. Candidates in order: an admission controller or PodSecurity policy forcing hostPath read-only; a k3s/containerd default for hostPath volumes; a missing explicit readOnly: false on the volumeMount.

This is the minimum to make the registrar function, and the only task here that unblocks anything today. Do it first even though T02 later deletes the mount — a working baseline makes every later change verifiable.

Verification is end-to-end, not a green pod: run the sweep and confirm EBIND-WP-0002 gets a state_hub_workstream_id written back into its file.

Give the pod its own clones

id: STATE-WP-0081-T02
status: todo
priority: high
state_hub_task_id: "3fd4d23c-6b67-5b92-b312-25a7295773fd"

Replace the sweep-repos hostPath with storage the pod owns — a PVC the sweep clones into and maintains, seeded from forgejo.

This is the change that actually severs coupling #1 and #2. Mounting a human's home directory into a production workload is what created 999 root-owned files, what made git pull fail on the host, and what put the operator's checkouts one reset --hard away from a scheduled job.

Sizing input: the current tree is ~78 repos; markitect_project alone is 24 MB.

Keep the sweep's repo list driven by the hub's repo registry, not by whatever happens to be on a disk.

Replace the operator SSH key with a service identity

id: STATE-WP-0081-T03
status: todo
priority: high
state_hub_task_id: "8e5a5eb8-f29c-5cab-a19d-3447b6e0c26a"

Issue a dedicated forgejo deploy key or service account for the sweep and mount it from a Secret. Remove sweep.sshHostPath and the /root/.ssh mount.

The current arrangement gives a root-running production workload the operator's personal private key. It is read-only, so this is a blast-radius problem rather than a live compromise — but the key that can push to every repository in the fleet should not be the same key a human uses interactively.

Credential custody routes through OpenBao, not this repo — see .claude/rules/credential-routing.md. Do not put key material in the chart.

Run as a non-root user

id: STATE-WP-0081-T04
status: todo
priority: medium
state_hub_task_id: "57ab0ced-ef3f-5260-b7d6-e2a78ff4e023"

Set runAsUser/runAsGroup and a fsGroup matching the PVC. Depends on T02: once the pod owns its storage there is no reason for it to be root.

Closes the recurrence: today's ownership fix will be undone by the next sweep while the pod still runs as root.

Correct the hub repo registry

id: STATE-WP-0081-T05
status: todo
priority: high
state_hub_task_id: "530fbd27-e463-51d2-b09c-ba9827d057de"

72 of 75 repo records carry local_path: /home/worsch/... and stale remote_url: gitea-remote:.... Both are wrong for a cluster that reads from forgejo into its own clone tree.

Decide first whether local_path should be per-instance rather than global — one column cannot describe a workstation checkout and a cluster clone at once, and its current single value is precisely how the workstation leaked into cluster configuration. RMGR-WP-0008 is building the repository-representation surface that inherits this; settle the shape with it rather than patching values that will move.

remote_url correction is unambiguous and can proceed immediately.

Serve the dashboard from the cluster

id: STATE-WP-0081-T06
status: todo
priority: medium
state_hub_task_id: "1d74e5e2-fb0f-5342-afbb-6ce2986823f4"

The dashboard runs as an Observable dev server on the workstation at 127.0.0.1:3000. Serve it from the cluster behind the same ingress as the API so it survives the workstation being off, and so what the operator sees is what the cluster holds.

Check first whether this is worth building here at all: hub-projection-ui is dispositioned replacehub-core (14 items, slice B5 in docs/retirement-cutover-slice-plan.md). If B5 lands first this task is a redirect, not a build. Confirm with HUB-WP-0004 before writing any chart.

State and enforce the boundary

id: STATE-WP-0081-T07
status: todo
priority: medium
state_hub_task_id: "12a2ce9f-7cdf-501b-9018-4c960d45e855"

Write the rule down so it stops being re-derived: workstation coding agents push to forgejo; cluster infrastructure reads from forgejo; neither reads the other's disk.

Record it where agents will meet it — docs/, the repo AGENTS.md templates, and as an ADR if it constrains other repos, which it does.

Include the failure this prevents. On 2026-08-21 railiance01's 70 checkouts were found still pointed at the retired gitea host, six weeks stale, and the evidence-binder workplan the registrar was asked to index did not exist in the copy the cluster could see. A shared-disk assumption made a stale reader look like a queue.

Close out the registrar

id: STATE-WP-0081-T08
status: todo
priority: high
state_hub_task_id: "782c502a-5ce2-5ec3-8f07-b550c9705013"

With the write path restored, confirm the registrar end-to-end and drain the backlog: 12 queued requests from evidence-binder, kaizen-agentic, glas-harness, and agentic-resources, plus this fleet's own unregistered RMGR-WP-0008 and RMGR-WP-0009.

Then remove the interim: RMGR-WP-0005-T03 (UUIDv5 derivation, keyed on (namespace, identifier) and deriving for live records only per the 2026-08-21 ADR-007 amendment) makes writeback idempotent and retires the single-writer rule entirely. Coordinate rather than duplicate — the derivation belongs to repo-manager.

Reply to the queued agents when it is done; several have been waiting since 2026-08-20.

Acceptance

  • Sweep writes successfully from the pod; EBIND-WP-0002 registered
  • Pod uses its own clone volume; no hostPath into any home directory
  • Pod authenticates with a dedicated key, runs as non-root
  • No /home/worsch path in any cluster-consumed record; remote_url values current
  • Dashboard reachable without the workstation, or formally handed to hub-core
  • Boundary rule written and discoverable by agents
  • Registrar queue drained; interim single-writer rule retired or explicitly deferred