repo-manager/workplans/RMGR-WP-0005-registrar-consolidation-deterministic-ids.md
tegwick 35e86d7b85 feat: advance repository records and provenance
Assistant: codex
Assistant-Model: gpt-5.6-sol
Assistant-Session: 01a023c0-a0a3-7c03-b395-5a0d2757214d
2026-08-21 22:07:48 +02:00

22 KiB
Raw Blame History

id type title domain repo status owner topic_slug created updated parent_project parent_workplan related state_hub_workstream_id
RMGR-WP-0005 workplan Registrar consolidation and deterministic hub identifiers infotech repo-manager active codex infotech 2026-08-17 2026-08-21 prj-state-hub-retirement SHR-WP-0001
RMGR-WP-0004
STATE-WP-0080
STATE-WP-0068
CFED-WP-0001
7ddb5421-d960-4a3c-94b1-40b6c96abfab

Registrar consolidation and deterministic hub identifiers

Goal

Make hub identifiers stored in repository files derivable rather than database-local, so that any number of hub instances can reconcile the same repository without overwriting each other.

Implements ADR-007 decision 2: interim single-writer (A), target deterministic derivation (C2).

The defect

state_hub_workstream_id and state_hub_task_id are database-local primary keys stored in a shared git artifact. Two hub instances over two databases each mint their own value for the same workplan, and every sync overwrites the other.

Observed 2026-08-16 on STATE-WP-0080: workplan UUID 03f38314 from the workstation hub, bbfce36a from a second instance (404 against the workstation database), plus two disjoint sets of task UUIDs. Sync commits appear under both +0000 and +0200 timezones, confirming two machines write to one repository.

It also inverts ADR-001. Files are meant to originate work with the hub as read model; a file carrying a hub's private key is the file holding hub state.

Scope: 758 workplan files across the fleet currently carry these fields.

Apply the interim single-writer rule

id: RMGR-WP-0005-T01
status: done
priority: high
state_hub_task_id: "b57a6882-280d-4f0a-9c73-899843dfc3d3"

Until derivation ships, exactly one instance may write hub identifiers into repository files. The interim registrar is the automated production instance; workstation hubs are rebuildable caches (ADR-010 decision 2).

  • Make the writeback path refuse to mint identifiers when the instance is not the registrar, rather than relying on operator discipline.
  • Provide the configuration that designates the registrar, and make a non-registrar instance's read/project behaviour unchanged.
  • Document the accepted cost: registration requires connectivity to the registrar, so disconnected work cannot register until T03 lands.

Interim, and deliberately so — it trades availability for correctness, and T03 removes the need for the trade.

Result (2026-08-18): repo_manager.registrar.is_identifier_registrar (STATEHUB_REGISTRAR env, else hostname prefix railiance). statehub fix-consistency skips C-06 / C-11 / C-32 mint+writeback when this instance is not the registrar. Read/project checks are unchanged. Cost documented in docs/repository-standards_v0.1.md.

The registrar was never stood up (found 2026-08-21). T01 shipped the guard — non-registrar instances correctly refuse to mint — but no instance satisfies the registrar condition, so nothing drains the queue. Twelve registrar-sync requests had accumulated from evidence-binder, kaizen-agentic, glas-harness and agentic-resources before anyone noticed.

Surveyed on railiance01, the designated production registrar:

Prerequisite State
Registrar signal fails — hostname is 239.62.205.92.host.secureserver.net, so the railiance* prefix heuristic never matches, and STATEHUB_REGISTRAR is unset
repo-manager clone was absent; cloned 2026-08-21, rmgr installed into ~/.venvs/registrar
statehub CLI not installablepip install ~/state-hub fails ResolutionImpossible; hub-core installs, the conflict is elsewhere
Checkout ownership 999 files under ~/state-hub owned by root, because the state-hub pod bind-mounts /home/tegwick as a hostPath and runs as root. Blocks git pull and editable installs
Checkout lineage points at gitea-remote, not forgejo-remote — the host was never swept by forgejo-tier3-remote-url-sweep-playbook. Its evidence-binder tracks a different lineage and cannot see EBIND-WP-0002 at all

The hostname heuristic is the cheapest fix and the most misleading defect: it reads as though a registrar exists whenever a host is named railiance*, but the production host is not. Prefer the explicit STATEHUB_REGISTRAR signal and treat the hostname fallback as unreliable, or drop it.

Two consequences beyond this task:

  • Security. The state-hub pod mounts /home/tegwick and /home/tegwick/.ssh as hostPath volumes. A workload with the operator's private keys is a larger exposure than the registrar problem it surfaced.

  • The remote sweep is incomplete, and this gates the registrar fix. Counted 2026-08-21:

    Host forgejo-remote gitea-remote
    workstation 126 0
    railiance01 7 70

    The two aliases are different servers — forgejo-remote is 92.205.62.239 (railiance01), gitea-remote is 92.205.130.254 (coulombcore) — and both are live. The workstation finished migrating; railiance01 did not. forgejo-tier3-remote-url-sweep-playbook still reads "Gitea remains canonical", which was true when written on 2026-07-04 and is now stale guidance.

Why the registrar cannot simply be switched on

The fix looked like one environment variable: the state-hub deployment has STATEHUB_REGISTRAR unset, and that pod is the real automated instance — it mounts /home/tegwick read-write, mounts /home/tegwick/.ssh at /root/.ssh for push, and carries STATE_HUB_SWEEP_HOSTNAME. Host-level provisioning was the wrong target.

Setting it today would be actively harmful. The pod would begin minting identifiers into 70 checkouts that track the superseded server and pushing them there — writing hub primary keys into a lineage nothing reads, and re-animating gitea-remote as a write target.

Required order:

  1. Re-point railiance01's checkouts to forgejo-remote, keeping gitea as the rollback mirror the playbook specifies. Content diverges — evidence-binder there is at a715492 against the workstation's 9e3f4a6 — so this is a reconcile, not a URL rewrite.
  2. Then set STATEHUB_REGISTRAR=1 on the deployment.

Reversing the order mints into the wrong lineage at fleet scale.

Answered, then executed (2026-08-21). The pod had pushed nothing: a commit-level audit of all 70 repos found zero gitea-only commits. Forgejo was strictly ahead everywhere (by 4165 commits). railiance01 was a stale reader, not a divergent writer — which is why its evidence-binder sat at a July commit and could not see EBIND-WP-0002.

Two repos looked gitea-only under SSH ls-remote but were not: inter-hub had been renamed to inter-hub-haskell (Forgejo answers the rename over HTTP with a 307, which SSH does not follow) and markitect_project to markitect-main.

Transition completed. Sweep pod scaled to zero, then per repo: origin re-pointed to forgejo-remote, gitea retained as the rollback mirror the playbook specifies, root-owned files chowned back, and the checkout brought onto the forgejo lineage. Final state — 77 repos in sync, 0 ahead, 0 behind, 0 still on gitea-remote.

Recovered rather than discarded:

  • freedom-intelligence held 6 unpushed commits — daily research briefs from 714 August by rein-aharness, present on no server. Rebased and pushed (846cccd..8832652). The first audit missed them because it measured against @{u} and silently skipped repos with no upstream configured; they surfaced only when the comparison was redone against the forgejo ref.
  • 5 working trees stashed as pre-forgejo-transition-20260821.
  • 10 stale custodian-sync status commits preserved on stale-sync-20260821 branches before being dropped.

Remaining for the registrar: set STATEHUB_REGISTRAR=1 on the deployment and scale back up. Note the pod runs as root and will re-create root-owned files under /home/tegwick, undoing today's ownership fix — it needs a non-root runAsUser or its own service account.

Interim recovery revised 2026-08-21: do not restore the host-wide sweep just to drain UUID requests. Repo Manager now owns a scoped on-demand command, rmgr registrar-reconcile, which preflights a clean synchronized Forgejo checkout and the authoritative hub, serializes the run, and sets STATEHUB_REGISTRAR=1 only for one repository's fix-consistency child. It commits the identifier writeback and pushes only when explicitly requested. This is the coding-agent recovery path until T03 lands; agents must neither set the environment variable directly nor retry C-06/C-11 or send duplicate messages. The production sweep remains disabled behind T12.

gitea cannot be decommissioned yet: the rollback remotes still point at it by design. Drop them once a sweep or two confirms forgejo is healthy.

Contain the stale-lineage production sweep

id: RMGR-WP-0005-T11
status: done
priority: high
state_hub_task_id: "bddeb759-bab8-43ff-8cfc-a1d173827fac"

Audit the active state-hub sweep workload on railiance01 before enabling the registrar. Establish which repositories it has fetched from or pushed to gitea-remote, preserve non-secret evidence of the workload configuration and recent Git outcomes, and prevent further stale-lineage writes with the smallest reversible control.

Do not rewrite remote URLs as containment: the checkouts have diverged and need the governed reconciliation described above. Do not enable STATEHUB_REGISTRAR while any swept checkout still targets gitea-remote.

Done when the production sweep cannot push repository changes to the stale lineage, normal State Hub serving remains available, the SSH hostPath exposure is recorded for remediation, and the control and rollback are documented.

Result (2026-08-21): contained without interrupting the State Hub API.

  • Production evidence showed the 15-minute Temporal schedule had fired 5,501 times. The final runs processed 1215 repositories each, but every Git fetch and push failed with Resource temporarily unavailable; C-16/C-17 prevented further writes where repositories were behind or had unpushed commits. No checkout commit or reflog activity was found after 2026-08-19. The latent stale-lineage write path nevertheless remained live.
  • Paused Temporal schedule activity-schedule-7c4e9a12-8f3b-4d5e-9c6a-1b2d3e4f5a6b. Its last run is fixed at 2026-08-21T10:00:00Z; subsequent intervals did not fire.
  • Rolled State Hub deployment revision 11 with /home/tegwick read-only and the /home/tegwick/.ssh mount removed. The replacement pod is Ready and /state/health remains healthy.
  • Made containment durable: State Hub production Helm sweep disabled in commit 2841bf3; activity-core projection disabled in 5793eb3; Custodian-owned definition disabled in 22d9366. All three commits are on Forgejo main.
  • Live activity-core ConfigMap and definition row say enabled: false; the Helm chart passes lint and renders no sweep/SSH mounts; activity-core targeted tests pass (16 tests).

Rollback is deliberately gated: do not unpause the Temporal schedule or enable the Helm sweep until the railiance01 checkouts are reconciled to forgejo-remote, the registrar preflight passes, and T12 provides a scoped credential path that does not mount an operator home or private-key directory.

Replace the host-wide sweep credential with a scoped identity

id: RMGR-WP-0005-T12
status: wait
priority: high
state_hub_task_id: "7b4edf5b-c1a1-4ef3-ab17-95487b9b19c8"

The retired sweep design mounted all of /home/tegwick read-write and mounted /home/tegwick/.ssh into the State Hub container. Before any remote sweep is re-enabled, replace that host-wide authority with a workload-specific identity and explicit repository scope. The runtime must not receive an operator private key, an operator home directory, or implicit write access to every checkout.

Coordinate credential custody with the platform owner and keep the schedule disabled until positive allowed-repository and negative unrelated-repository push evidence exist without exposing credential values.

Re-register identifiers minted outside the registrar

id: RMGR-WP-0005-T02
status: wait
priority: medium
state_hub_task_id: "8e679ddb-9845-457f-8672-1fd4b7455e7b"

Records minted by non-registrar instances before T01 need reconciliation. Known cases, all created 2026-08-16/17 from the workstation hub:

  • RMGR-WP-0004 (b8b3f1e0) and its seven tasks;
  • CFED-WP-0001 (7a96da54) and its thirteen tasks, plus the prj-canon-federation repo record (3809b0ff);
  • STATE-WP-0080 — already reconciled by hand to the second instance's IDs (bbfce36a), retained here as the worked example.

Prefer waiting for T03 where possible: once identifiers are derived, these converge without manual intervention. Re-register by hand only what blocks work before then.

Derive identifiers deterministically

id: RMGR-WP-0005-T03
status: progress
priority: high
state_hub_task_id: "28067729-498d-4f47-89bd-5b9718e999c7"

Replace minted UUIDs with UUIDv5 derived from the globally unique PREFIX-WP-NNNN identifier (and PREFIX-WP-NNNN-TNN for tasks).

  • Fix the namespace UUID and derivation input as a published contract — the value must be reproducible by any implementation, not just this one.
  • Field shape is unchanged, so consumers reading state_hub_workstream_id keep working; only the provenance of the value changes.
  • Writeback becomes idempotent: two instances write identical bytes, so the flip-flop cannot recur regardless of how many hubs run.

Blocked on RMGR-WP-0004-T08. Deriving from a non-unique identifier manufactures collisions: two repositories sharing PRJ-WP- would compute the same UUID for different workplans. Uniqueness must be enforced first.

Unblocked 2026-08-21. RMGR-WP-0004-T08 closed 2026-08-18 (prefix registry plus rmgr prefix-uniqueness), and ADR-007 was amended the same day with the derivation scope this task needs:

  • Derive for live records only; archived records keep frozen minted identifiers. This is what reconciles ADR-007 § Migration option 2 with the uniqueness derivation requires.
  • Derivation input is (namespace, identifier) per ADR-011 decision 7. Namespace is the fleet branch, not the repository — the ecosystem is at N1, one implied namespace, so the pair does not disambiguate intra-namespace collisions and must not be read as if it did.
  • Derivation is not retroactive: existing live records keep minted UUIDs until deliberately re-derived.
  • Un-archiving is a collision hazard — a record returning to live must be checked against the live namespace and renumbered if it clashes. Build this check alongside derivation, not after.

Remaining prerequisite is the live-collision remediation below, tracked on RMGR-WP-0004-T09. 11 files, 5 identifiers.

Progress (2026-08-21): the versioned derivation function and collision guard are implemented in Repo Manager and published as docs/work-record-uuid-derivation_v1.md. UUIDv5 uses fixed namespace UUID a4058507-5c4a-5a00-ab06-fffa4fb46009 and exact name bytes <fleet-namespace>\n<canonical-id>. rmgr identifier derive|preflight provides independent reproduction and a hard live-collision refusal, including the unarchive hazard. Activation remains correctly gated on T09's 11-file remediation and declaration of the current N1 fleet namespace name; no existing minted identifier was rewritten implicitly.

Migrate the fleet

id: RMGR-WP-0005-T04
status: wait
priority: high
state_hub_task_id: "503a23a9-ede1-4cf1-bd32-e9669b84ce58"

One-time pass over the 758 files carrying hub identifiers: compute the derived value, update the database to match, and write the file.

  • Must be all-or-nothing per repository — a half-migrated repo has some derived and some minted identifiers and reconciles unpredictably.
  • Records whose current identifier is already referenced externally (dashboards, saved queries, progress events) need a mapping table from old to derived, kept as provenance rather than discarded.
  • Repositories with unresolved identifier collisions cannot migrate until ADR-007 § Migration is ruled on; skip and report them rather than guessing.

Retire the interim rule

id: RMGR-WP-0005-T05
status: wait
priority: low
state_hub_task_id: "3946d1fc-2137-4d7b-a400-29b447ca83de"

Once derivation is live fleet-wide, remove the single-writer restriction from T01. Multiple hub instances become an availability choice rather than a correctness constraint, and disconnected registration works again.

Confirm before removal: two instances reconciling the same repository produce byte-identical writeback, and neither creates a duplicate record.

Rebuild local instances as caches

id: RMGR-WP-0005-T07
status: wait
priority: high
state_hub_task_id: "70f83359-0b61-4dc0-83b0-33f289b64e83"

Implement ADR-010 decisions 13: the central hub on railiance is authoritative as a reading of the repositories; local instances become rebuildable caches.

  • A cache must be discardable and reconstructable from repository files alone, with no work lost.
  • Local work must not require a hub — repository files are self-describing, so reading them is sufficient for working inside a repo.
  • Cache reads are advisory and must carry their staleness (ADR-010 decision 8).

Measured 2026-08-17: 955 workplans locally against 649 on the primary, 320 local-only, of which 288 are backed by files that all exist on disk. That portion of the divergence is redundant and needs no merge — only a rebuild.

Separate file-derived from hub-native data

id: RMGR-WP-0005-T08
status: wait
priority: high
state_hub_task_id: "241cf058-2f3e-4d49-8cc9-5c714be4a1cf"

Implement ADR-010 decision 4. The two kinds need opposite handling:

  • File-derived (workplans, tasks, statuses, dependencies) — central derives it and must not accept pushes of it (decision 5). Offline, the git commit is the write. No conflict model: conflicts are git conflicts.
  • Hub-native (progress events, decisions, inbox messages, token events) — central owns it, needs a real write path and a local append-only buffer for replay. No conflict model either: append-only merges regardless of order.

Deliverable is an explicit classification of every record type the hub holds, with its truth source and offline behaviour, so neither kind is handled by the other's rules.

Feeds a rescope of STATE-WP-0068 (offline write buffer and edge relay): under this split most of what it buffers does not need buffering, and only the append-only stream does. Re-examine before building further on it — this likely reduces its scope.

Disposition the orphaned hub-first records

id: RMGR-WP-0005-T09
status: wait
priority: high
state_hub_task_id: "d40cc4a8-4280-4940-ac1d-dc1049f1b678"

28 records exist in the local instance with no backing file. They are the only records a cache rebuild would drop, so they must be classified first (ADR-010 § Orphan disposition):

  1. Broken links — a file exists but backing_filename was never recorded. RMGR-WP-0004 is a confirmed instance. Repair the link; no data at risk. Likely the largest class, so classify before estimating the rest.
  2. Live hub-first recordsproposed/ready/backlog with no file, in activity-core, core-hub, hub-core, issue-core, ops-hub, prj-forgejo-org-refactor, railiance-enablement, railiance-infra, reef-railiance. Write a repository file or drop explicitly. These are ADR-001 violations and must not survive as hub-only records.
  3. Closed hub-first recordsfinished/archived with no file. Retain as provenance where cheap; do not reconstruct completed plans.

Blocks T07 — rebuilding the cache before this classification would discard class 2.

Note: one of these records is already labelled SPURIOUS bootstrap (statehub register collision) in repo-manager, independent corroboration of the STATE-WP-0080 defect.

Assign one authoritative hub per record

id: RMGR-WP-0005-T10
status: wait
priority: medium
state_hub_task_id: "15f0f167-a8d0-4d5c-8576-3e93b1e8792f"

Implement ADR-010 decision 7. The retirement splits one hub into several, which is permitted only if every record has exactly one authoritative hub, determined by its repository and domain.

Define and enforce that mapping before the split lands. Without it the peer-database divergence this workplan exists to remove recurs at larger scale.

Coordinate with the hub-extension architecture in prj-state-hub-retirement/architecture/; hub-core owns the hub-native side.

Protect lifecycle status from automation

id: RMGR-WP-0005-T06
status: done
priority: medium
state_hub_task_id: "d440d59c-f78e-4752-84c7-f3d5fdf7d3c3"

Implement ADR-007 decision 3: an automated normalization pass may report lifecycle drift but may not promote a workplan from proposed to active. proposed means awaiting human review; automated promotion destroys the gate.

Observed: commit ff909e1 ("renormalize lifecycle state [auto]") promoted STATE-WP-0080 to active hours after it was drafted for review.

Extend the same protection to task status, where the symptom is currently sharper: C-15 forces CFED-WP-0001-T02 back to wait on every sync regardless of file content — reproduced three times, via file edit and via update_task_status, with the task never holding todo. Establish which direction wins for task status and make it consistent with ADR-001, where the file originates work.

Result (2026-08-18): In state-hub consistency: C-23 does not auto-promote proposedactive (report only). C-15 no longer writebacks wait over progress/todo; file wins via C-10 (ADR-001). C-15 remains a non-fixable warning when the DB is terminal and the file is not.