2026-08-17 10:19:03 +02:00
---
id: RMGR-WP-0005
type: workplan
title: "Registrar consolidation and deterministic hub identifiers"
domain: infotech
repo: repo-manager
2026-08-18 12:40:32 +02:00
status: active
2026-08-17 10:19:03 +02:00
owner: codex
topic_slug: infotech
created: "2026-08-17"
2026-08-21 13:50:02 +02:00
updated: "2026-08-21"
2026-08-17 10:19:03 +02:00
parent_project: prj-state-hub-retirement
parent_workplan: SHR-WP-0001
related:
- RMGR-WP-0004
- STATE-WP-0080
- STATE-WP-0068
- CFED-WP-0001
2026-08-18 12:18:45 +02:00
state_hub_workstream_id: "7ddb5421-d960-4a3c-94b1-40b6c96abfab"
2026-08-17 10:19:03 +02:00
---
# Registrar consolidation and deterministic hub identifiers
## Goal
Make hub identifiers stored in repository files **derivable rather than
database-local**, so that any number of hub instances can reconcile the same
repository without overwriting each other.
Implements `ADR-007` decision 2: interim single-writer (A), target deterministic
derivation (C2).
## The defect
`state_hub_workstream_id` and `state_hub_task_id` are database-local primary
keys stored in a shared git artifact. Two hub instances over two databases each
mint their own value for the same workplan, and every sync overwrites the other.
Observed 2026-08-16 on `STATE-WP-0080` : workplan UUID `03f38314` from the
workstation hub, `bbfce36a` from a second instance (404 against the workstation
database), plus two disjoint sets of task UUIDs. Sync commits appear under both
`+0000` and `+0200` timezones, confirming two machines write to one repository.
It also inverts `ADR-001` . Files are meant to originate work with the hub as read
model; a file carrying a hub's private key is the file holding hub state.
**Scope: 758 workplan files** across the fleet currently carry these fields.
## Apply the interim single-writer rule
```task
id: RMGR-WP-0005-T01
2026-08-18 21:49:08 +02:00
status: done
2026-08-17 10:19:03 +02:00
priority: high
2026-08-18 12:18:45 +02:00
state_hub_task_id: "b57a6882-280d-4f0a-9c73-899843dfc3d3"
2026-08-17 10:19:03 +02:00
```
Until derivation ships, exactly one instance may write hub identifiers into
repository files. The interim registrar is the automated production instance;
2026-08-17 13:04:03 +02:00
workstation hubs are rebuildable caches (`ADR-010` decision 2).
2026-08-17 10:19:03 +02:00
- Make the writeback path refuse to mint identifiers when the instance is not
the registrar, rather than relying on operator discipline.
- Provide the configuration that designates the registrar, and make a
non-registrar instance's read/project behaviour unchanged.
- Document the accepted cost: registration requires connectivity to the
registrar, so disconnected work cannot register until T03 lands.
Interim, and deliberately so — it trades availability for correctness, and T03
removes the need for the trade.
2026-08-18 21:49:08 +02:00
**Result (2026-08-18):** `repo_manager.registrar.is_identifier_registrar`
(`STATEHUB_REGISTRAR` env, else hostname prefix `railiance` ).
`statehub fix-consistency` skips C-06 / C-11 / C-32 mint+writeback when
this instance is not the registrar. Read/project checks are unchanged.
Cost documented in `docs/repository-standards_v0.1.md` .
2026-08-21 09:32:47 +02:00
**The registrar was never stood up (found 2026-08-21).** T01 shipped the guard —
non-registrar instances correctly refuse to mint — but no instance satisfies the
registrar condition, so nothing drains the queue. Twelve registrar-sync requests
had accumulated from `evidence-binder` , `kaizen-agentic` , `glas-harness` and
`agentic-resources` before anyone noticed.
Surveyed on `railiance01` , the designated production registrar:
| Prerequisite | State |
| --- | --- |
| Registrar signal | **fails** — hostname is `239.62.205.92.host.secureserver.net` , so the `railiance*` prefix heuristic never matches, and `STATEHUB_REGISTRAR` is unset |
| `repo-manager` clone | was absent; cloned 2026-08-21, `rmgr` installed into `~/.venvs/registrar` |
| `statehub` CLI | **not installable** — `pip install ~/state-hub` fails `ResolutionImpossible` ; `hub-core` installs, the conflict is elsewhere |
| Checkout ownership | **999 files under `~/state-hub` owned by root** , because the `state-hub` pod bind-mounts `/home/tegwick` as a hostPath and runs as root. Blocks `git pull` and editable installs |
| Checkout lineage | **points at `gitea-remote`, not `forgejo-remote`** — the host was never swept by `forgejo-tier3-remote-url-sweep-playbook` . Its `evidence-binder` tracks a different lineage and cannot see `EBIND-WP-0002` at all |
The hostname heuristic is the cheapest fix and the most misleading defect: it
reads as though a registrar exists whenever a host is named `railiance*` , but the
production host is not. Prefer the explicit `STATEHUB_REGISTRAR` signal and treat
the hostname fallback as unreliable, or drop it.
Two consequences beyond this task:
- **Security.** The `state-hub` pod mounts `/home/tegwick` **and
`/home/tegwick/.ssh` ** as hostPath volumes. A workload with the operator's
private keys is a larger exposure than the registrar problem it surfaced.
2026-08-21 09:37:49 +02:00
- **The remote sweep is incomplete, and this gates the registrar fix.** Counted
2026-08-21:
| Host | `forgejo-remote` | `gitea-remote` |
| --- | --- | --- |
| workstation | 126 | 0 |
| `railiance01` | 7 | **70** |
The two aliases are different servers — `forgejo-remote` is `92.205.62.239`
(railiance01), `gitea-remote` is `92.205.130.254` (coulombcore) — and both are
live. The workstation finished migrating; `railiance01` did not.
`forgejo-tier3-remote-url-sweep-playbook` still reads "Gitea remains canonical",
which was true when written on 2026-07-04 and is now stale guidance.
### Why the registrar cannot simply be switched on
The fix looked like one environment variable: the `state-hub` **deployment** has
`STATEHUB_REGISTRAR` unset, and that pod is the real automated instance — it
mounts `/home/tegwick` read-write, mounts `/home/tegwick/.ssh` at `/root/.ssh`
for push, and carries `STATE_HUB_SWEEP_HOSTNAME` . Host-level provisioning was the
wrong target.
**Setting it today would be actively harmful.** The pod would begin minting
identifiers into 70 checkouts that track the superseded server and pushing them
there — writing hub primary keys into a lineage nothing reads, and re-animating
`gitea-remote` as a write target.
Required order:
1. Re-point `railiance01` 's checkouts to `forgejo-remote` , keeping `gitea` as the
rollback mirror the playbook specifies. Content diverges — `evidence-binder`
there is at `a715492` against the workstation's `9e3f4a6` — so this is a
reconcile, not a URL rewrite.
2. Then set `STATEHUB_REGISTRAR=1` on the deployment.
Reversing the order mints into the wrong lineage at fleet scale.
2026-08-21 15:37:07 +02:00
**Answered, then executed (2026-08-21).** The pod had pushed nothing: a
commit-level audit of all 70 repos found **zero gitea-only commits** . Forgejo was
strictly ahead everywhere (by 4– 165 commits). `railiance01` was a stale *reader* ,
not a divergent writer — which is why its `evidence-binder` sat at a July commit
and could not see `EBIND-WP-0002` .
Two repos looked gitea-only under SSH `ls-remote` but were not: `inter-hub` had
been renamed to `inter-hub-haskell` (Forgejo answers the rename over HTTP with a
307, which SSH does not follow) and `markitect_project` to `markitect-main` .
**Transition completed.** Sweep pod scaled to zero, then per repo: `origin`
re-pointed to `forgejo-remote` , `gitea` retained as the rollback mirror the
playbook specifies, root-owned files chowned back, and the checkout brought onto
the forgejo lineage. Final state — **77 repos in sync, 0 ahead, 0 behind, 0 still
on `gitea-remote` .**
Recovered rather than discarded:
- `freedom-intelligence` held **6 unpushed commits** — daily research briefs from
7– 14 August by `rein-aharness` , present on no server. Rebased and pushed
(`846cccd..8832652` ). The first audit missed them because it measured against
`@{u}` and silently skipped repos with no upstream configured; they surfaced
only when the comparison was redone against the forgejo ref.
- 5 working trees stashed as `pre-forgejo-transition-20260821` .
- 10 stale `custodian-sync` status commits preserved on `stale-sync-20260821`
branches before being dropped.
Remaining for the registrar: set `STATEHUB_REGISTRAR=1` on the deployment and
scale back up. Note the pod runs as root and will re-create root-owned files
under `/home/tegwick` , undoing today's ownership fix — it needs a non-root
`runAsUser` or its own service account.
2026-08-21 21:38:23 +02:00
**Interim recovery revised 2026-08-21:** do not restore the host-wide sweep just
to drain UUID requests. Repo Manager now owns a scoped on-demand command,
`rmgr registrar-reconcile` , which preflights a clean synchronized Forgejo
checkout and the authoritative hub, serializes the run, and sets
`STATEHUB_REGISTRAR=1` only for one repository's `fix-consistency` child.
It commits the identifier writeback and pushes only when explicitly requested.
This is the coding-agent recovery path until T03 lands; agents must neither set
the environment variable directly nor retry C-06/C-11 or send duplicate
messages. The production sweep remains disabled behind T12.
2026-08-21 15:37:07 +02:00
`gitea` cannot be decommissioned yet: the rollback remotes still point at it by
design. Drop them once a sweep or two confirms forgejo is healthy.
2026-08-21 09:32:47 +02:00
2026-08-21 13:50:02 +02:00
## Contain the stale-lineage production sweep
```task
id: RMGR-WP-0005-T11
status: done
priority: high
2026-08-21 21:42:37 +02:00
state_hub_task_id: "bddeb759-bab8-43ff-8cfc-a1d173827fac"
2026-08-21 13:50:02 +02:00
```
Audit the active `state-hub` sweep workload on `railiance01` before enabling the
registrar. Establish which repositories it has fetched from or pushed to
`gitea-remote` , preserve non-secret evidence of the workload configuration and
recent Git outcomes, and prevent further stale-lineage writes with the smallest
reversible control.
Do not rewrite remote URLs as containment: the checkouts have diverged and need
the governed reconciliation described above. Do not enable
`STATEHUB_REGISTRAR` while any swept checkout still targets `gitea-remote` .
Done when the production sweep cannot push repository changes to the stale
lineage, normal State Hub serving remains available, the SSH hostPath exposure
is recorded for remediation, and the control and rollback are documented.
**Result (2026-08-21):** contained without interrupting the State Hub API.
- Production evidence showed the 15-minute Temporal schedule had fired 5,501
times. The final runs processed 12– 15 repositories each, but every Git fetch
and push failed with `Resource temporarily unavailable` ; C-16/C-17 prevented
further writes where repositories were behind or had unpushed commits. No
checkout commit or reflog activity was found after 2026-08-19. The latent
stale-lineage write path nevertheless remained live.
- Paused Temporal schedule
`activity-schedule-7c4e9a12-8f3b-4d5e-9c6a-1b2d3e4f5a6b` . Its last run is
fixed at `2026-08-21T10:00:00Z` ; subsequent intervals did not fire.
- Rolled State Hub deployment revision 11 with `/home/tegwick` read-only and
the `/home/tegwick/.ssh` mount removed. The replacement pod is Ready and
`/state/health` remains healthy.
- Made containment durable: State Hub production Helm sweep disabled in commit
`2841bf3` ; activity-core projection disabled in `5793eb3` ; Custodian-owned
definition disabled in `22d9366` . All three commits are on Forgejo `main` .
- Live activity-core ConfigMap and definition row say `enabled: false` ; the
Helm chart passes lint and renders no sweep/SSH mounts; activity-core targeted
tests pass (16 tests).
Rollback is deliberately gated: do not unpause the Temporal schedule or enable
the Helm sweep until the `railiance01` checkouts are reconciled to
`forgejo-remote` , the registrar preflight passes, and T12 provides a scoped
credential path that does not mount an operator home or private-key directory.
## Replace the host-wide sweep credential with a scoped identity
```task
id: RMGR-WP-0005-T12
status: wait
priority: high
2026-08-21 21:42:37 +02:00
state_hub_task_id: "7b4edf5b-c1a1-4ef3-ab17-95487b9b19c8"
2026-08-21 13:50:02 +02:00
```
The retired sweep design mounted all of `/home/tegwick` read-write and mounted
`/home/tegwick/.ssh` into the State Hub container. Before any remote sweep is
re-enabled, replace that host-wide authority with a workload-specific identity
and explicit repository scope. The runtime must not receive an operator private
key, an operator home directory, or implicit write access to every checkout.
Coordinate credential custody with the platform owner and keep the schedule
disabled until positive allowed-repository and negative unrelated-repository
push evidence exist without exposing credential values.
2026-08-17 10:19:03 +02:00
## Re-register identifiers minted outside the registrar
```task
id: RMGR-WP-0005-T02
status: wait
priority: medium
2026-08-18 12:18:45 +02:00
state_hub_task_id: "8e679ddb-9845-457f-8672-1fd4b7455e7b"
2026-08-17 10:19:03 +02:00
```
Records minted by non-registrar instances before T01 need reconciliation. Known
cases, all created 2026-08-16/17 from the workstation hub:
- `RMGR-WP-0004` (`b8b3f1e0` ) and its seven tasks;
- `CFED-WP-0001` (`7a96da54` ) and its thirteen tasks, plus the
`prj-canon-federation` repo record (`3809b0ff` );
- `STATE-WP-0080` — already reconciled by hand to the second instance's IDs
(`bbfce36a` ), retained here as the worked example.
Prefer waiting for T03 where possible: once identifiers are derived, these
converge without manual intervention. Re-register by hand only what blocks work
before then.
## Derive identifiers deterministically
```task
id: RMGR-WP-0005-T03
2026-08-21 22:07:48 +02:00
status: progress
2026-08-17 10:19:03 +02:00
priority: high
2026-08-18 12:18:45 +02:00
state_hub_task_id: "28067729-498d-4f47-89bd-5b9718e999c7"
2026-08-17 10:19:03 +02:00
```
Replace minted UUIDs with UUIDv5 derived from the globally unique
`PREFIX-WP-NNNN` identifier (and `PREFIX-WP-NNNN-TNN` for tasks).
- Fix the namespace UUID and derivation input as a published contract — the
value must be reproducible by any implementation, not just this one.
- Field shape is unchanged, so consumers reading `state_hub_workstream_id`
keep working; only the provenance of the value changes.
- Writeback becomes idempotent: two instances write identical bytes, so the
flip-flop cannot recur regardless of how many hubs run.
**Blocked on `RMGR-WP-0004-T08` .** Deriving from a non-unique identifier
manufactures collisions: two repositories sharing `PRJ-WP-` would compute the
same UUID for different workplans. Uniqueness must be enforced first.
2026-08-21 06:36:44 +02:00
**Unblocked 2026-08-21.** `RMGR-WP-0004-T08` closed 2026-08-18 (prefix registry
plus `rmgr prefix-uniqueness` ), and `ADR-007` was amended the same day with the
derivation scope this task needs:
- **Derive for live records only**; archived records keep frozen minted
identifiers. This is what reconciles `ADR-007` § Migration option 2 with the
uniqueness derivation requires.
- Derivation input is `(namespace, identifier)` per `ADR-011` decision 7.
**Namespace is the fleet branch, not the repository** — the ecosystem is at
`N1` , one implied namespace, so the pair does not disambiguate intra-namespace
collisions and must not be read as if it did.
- Derivation is **not retroactive** : existing live records keep minted UUIDs
until deliberately re-derived.
- **Un-archiving is a collision hazard** — a record returning to live must be
checked against the live namespace and renumbered if it clashes. Build this
check alongside derivation, not after.
Remaining prerequisite is the live-collision remediation below, tracked on
`RMGR-WP-0004-T09` . 11 files, 5 identifiers.
2026-08-21 22:07:48 +02:00
Progress (2026-08-21): the versioned derivation function and collision guard are
implemented in Repo Manager and published as
`docs/work-record-uuid-derivation_v1.md` . UUIDv5 uses fixed namespace UUID
`a4058507-5c4a-5a00-ab06-fffa4fb46009` and exact name bytes
`<fleet-namespace>\n<canonical-id>` . `rmgr identifier derive|preflight` provides
independent reproduction and a hard live-collision refusal, including the
unarchive hazard. Activation remains correctly gated on T09's 11-file
remediation and declaration of the current N1 fleet namespace name; no existing
minted identifier was rewritten implicitly.
2026-08-17 10:19:03 +02:00
## Migrate the fleet
```task
id: RMGR-WP-0005-T04
status: wait
priority: high
2026-08-18 12:18:45 +02:00
state_hub_task_id: "503a23a9-ede1-4cf1-bd32-e9669b84ce58"
2026-08-17 10:19:03 +02:00
```
One-time pass over the 758 files carrying hub identifiers: compute the derived
value, update the database to match, and write the file.
- Must be all-or-nothing per repository — a half-migrated repo has some derived
and some minted identifiers and reconciles unpredictably.
- Records whose current identifier is already referenced externally (dashboards,
saved queries, progress events) need a mapping table from old to derived, kept
as provenance rather than discarded.
- Repositories with unresolved identifier collisions cannot migrate until
`ADR-007` § Migration is ruled on; skip and report them rather than guessing.
## Retire the interim rule
```task
id: RMGR-WP-0005-T05
status: wait
priority: low
2026-08-18 12:18:45 +02:00
state_hub_task_id: "3946d1fc-2137-4d7b-a400-29b447ca83de"
2026-08-17 10:19:03 +02:00
```
Once derivation is live fleet-wide, remove the single-writer restriction from
T01. Multiple hub instances become an availability choice rather than a
correctness constraint, and disconnected registration works again.
Confirm before removal: two instances reconciling the same repository produce
byte-identical writeback, and neither creates a duplicate record.
2026-08-17 12:18:02 +02:00
## Rebuild local instances as caches
```task
id: RMGR-WP-0005-T07
status: wait
priority: high
2026-08-18 12:18:45 +02:00
state_hub_task_id: "70f83359-0b61-4dc0-83b0-33f289b64e83"
2026-08-17 12:18:02 +02:00
```
2026-08-17 13:04:03 +02:00
Implement `ADR-010` decisions 1– 3: the central hub on railiance is authoritative
2026-08-17 12:18:02 +02:00
as a *reading* of the repositories; local instances become rebuildable caches.
- A cache must be discardable and reconstructable from repository files alone,
with no work lost.
- Local work must not require a hub — repository files are self-describing, so
reading them is sufficient for working inside a repo.
2026-08-17 13:04:03 +02:00
- Cache reads are advisory and must carry their staleness (`ADR-010` decision 8).
2026-08-17 12:18:02 +02:00
Measured 2026-08-17: 955 workplans locally against 649 on the primary, 320
local-only, of which **288 are backed by files that all exist on disk** . That
portion of the divergence is redundant and needs no merge — only a rebuild.
## Separate file-derived from hub-native data
```task
id: RMGR-WP-0005-T08
status: wait
priority: high
2026-08-18 12:18:45 +02:00
state_hub_task_id: "241cf058-2f3e-4d49-8cc9-5c714be4a1cf"
2026-08-17 12:18:02 +02:00
```
2026-08-17 13:04:03 +02:00
Implement `ADR-010` decision 4. The two kinds need opposite handling:
2026-08-17 12:18:02 +02:00
- **File-derived** (workplans, tasks, statuses, dependencies) — central derives
it and must not accept pushes of it (decision 5). Offline, the git commit *is*
the write. No conflict model: conflicts are git conflicts.
- **Hub-native** (progress events, decisions, inbox messages, token events) —
central owns it, needs a real write path and a local append-only buffer for
replay. No conflict model either: append-only merges regardless of order.
Deliverable is an explicit classification of every record type the hub holds,
with its truth source and offline behaviour, so neither kind is handled by the
other's rules.
Feeds a rescope of `STATE-WP-0068` (offline write buffer and edge relay): under
this split most of what it buffers does not need buffering, and only the
append-only stream does. Re-examine before building further on it — this likely
reduces its scope.
## Disposition the orphaned hub-first records
```task
id: RMGR-WP-0005-T09
status: wait
priority: high
2026-08-18 12:18:45 +02:00
state_hub_task_id: "d40cc4a8-4280-4940-ac1d-dc1049f1b678"
2026-08-17 12:18:02 +02:00
```
28 records exist in the local instance with no backing file. They are the only
records a cache rebuild would drop, so they must be classified first
2026-08-17 13:04:03 +02:00
(`ADR-010` § Orphan disposition):
2026-08-17 12:18:02 +02:00
1. **Broken links** — a file exists but `backing_filename` was never recorded.
`RMGR-WP-0004` is a confirmed instance. Repair the link; no data at risk.
Likely the largest class, so classify before estimating the rest.
2. **Live hub-first records** — `proposed` /`ready` /`backlog` with no file, in
`activity-core` , `core-hub` , `hub-core` , `issue-core` , `ops-hub` ,
`prj-forgejo-org-refactor` , `railiance-enablement` , `railiance-infra` ,
`reef-railiance` . Write a repository file or drop explicitly. These are
`ADR-001` violations and must not survive as hub-only records.
3. **Closed hub-first records** — `finished` /`archived` with no file. Retain as
provenance where cheap; do not reconstruct completed plans.
**Blocks T07** — rebuilding the cache before this classification would discard
class 2.
Note: one of these records is already labelled `SPURIOUS bootstrap (statehub
register collision)` in ` repo-manager`, independent corroboration of the
`STATE-WP-0080` defect.
## Assign one authoritative hub per record
```task
id: RMGR-WP-0005-T10
status: wait
priority: medium
2026-08-18 12:18:45 +02:00
state_hub_task_id: "15f0f167-a8d0-4d5c-8576-3e93b1e8792f"
2026-08-17 12:18:02 +02:00
```
2026-08-17 13:04:03 +02:00
Implement `ADR-010` decision 7. The retirement splits one hub into several, which
2026-08-17 12:18:02 +02:00
is permitted only if every record has exactly one authoritative hub, determined
by its repository and domain.
Define and enforce that mapping before the split lands. Without it the
peer-database divergence this workplan exists to remove recurs at larger scale.
Coordinate with the hub-extension architecture in
`prj-state-hub-retirement/architecture/` ; `hub-core` owns the hub-native side.
2026-08-17 10:19:03 +02:00
## Protect lifecycle status from automation
```task
id: RMGR-WP-0005-T06
2026-08-18 13:36:26 +02:00
status: done
2026-08-17 10:19:03 +02:00
priority: medium
2026-08-18 12:18:45 +02:00
state_hub_task_id: "d440d59c-f78e-4752-84c7-f3d5fdf7d3c3"
2026-08-17 10:19:03 +02:00
```
Implement `ADR-007` decision 3: an automated normalization pass may report
lifecycle drift but may not promote a workplan from `proposed` to `active` .
`proposed` means awaiting human review; automated promotion destroys the gate.
Observed: commit `ff909e1` ("renormalize lifecycle state [auto]") promoted
`STATE-WP-0080` to `active` hours after it was drafted for review.
Extend the same protection to task status, where the symptom is currently
sharper: `C-15` forces `CFED-WP-0001-T02` back to `wait` on every sync
regardless of file content — reproduced three times, via file edit and via
`update_task_status` , with the task never holding `todo` . Establish which
direction wins for task status and make it consistent with `ADR-001` , where the
file originates work.
2026-08-18 13:36:26 +02:00
**Result (2026-08-18):** In `state-hub` consistency: C-23 does not
auto-promote `proposed` → `active` (report only). C-15 no longer
writebacks wait over progress/todo; file wins via C-10 (ADR-001).
C-15 remains a non-fixable warning when the DB is terminal and the file
is not.