Assistant: codex Assistant-Model: gpt-5.6-sol Assistant-Session: 01a053ff-1d6f-7fe2-ac1c-a6eb40a42a0c
48 KiB
| id | type | title | domain | repo | status | owner | topic_slug | created | updated | parent_project | parent_workplan | related | state_hub_workstream_id | ||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| RMGR-WP-0005 | workplan | Registrar consolidation and deterministic hub identifiers | infotech | repo-manager | active | codex | infotech | 2026-08-17 | 2026-08-31 | prj-state-hub-retirement | SHR-WP-0001 |
|
6dcf854e-9229-569d-90f0-2d9df235a61d |
Registrar consolidation and deterministic hub identifiers
Goal
Make hub identifiers stored in repository files derivable rather than database-local, so that any number of hub instances can reconcile the same repository without overwriting each other.
Implements ADR-007 decision 2: interim single-writer (A), target deterministic
derivation (C2).
The defect
state_hub_workstream_id and state_hub_task_id are database-local primary
keys stored in a shared git artifact. Two hub instances over two databases each
mint their own value for the same workplan, and every sync overwrites the other.
Observed 2026-08-16 on STATE-WP-0080: workplan UUID 03f38314 from the
workstation hub, bbfce36a from a second instance (404 against the workstation
database), plus two disjoint sets of task UUIDs. Sync commits appear under both
+0000 and +0200 timezones, confirming two machines write to one repository.
It also inverts ADR-001. Files are meant to originate work with the hub as read
model; a file carrying a hub's private key is the file holding hub state.
Scope: 758 workplan files across the fleet currently carry these fields.
Apply the interim single-writer rule
id: RMGR-WP-0005-T01
status: done
priority: high
state_hub_task_id: "b57a6882-280d-4f0a-9c73-899843dfc3d3"
Until derivation ships, exactly one instance may write hub identifiers into
repository files. The interim registrar is the automated production instance;
workstation hubs are rebuildable caches (ADR-010 decision 2).
- Make the writeback path refuse to mint identifiers when the instance is not the registrar, rather than relying on operator discipline.
- Provide the configuration that designates the registrar, and make a non-registrar instance's read/project behaviour unchanged.
- Document the accepted cost: registration requires connectivity to the registrar, so disconnected work cannot register until T03 lands.
Interim, and deliberately so — it trades availability for correctness, and T03 removes the need for the trade.
Result (2026-08-18): repo_manager.registrar.is_identifier_registrar
(STATEHUB_REGISTRAR env, else hostname prefix railiance).
statehub fix-consistency skips C-06 / C-11 / C-32 mint+writeback when
this instance is not the registrar. Read/project checks are unchanged.
Cost documented in docs/repository-standards_v0.1.md.
The registrar was never stood up (found 2026-08-21). T01 shipped the guard —
non-registrar instances correctly refuse to mint — but no instance satisfies the
registrar condition, so nothing drains the queue. Twelve registrar-sync requests
had accumulated from evidence-binder, kaizen-agentic, glas-harness and
agentic-resources before anyone noticed.
Surveyed on railiance01, the designated production registrar:
| Prerequisite | State |
|---|---|
| Registrar signal | fails — hostname is 239.62.205.92.host.secureserver.net, so the railiance* prefix heuristic never matches, and STATEHUB_REGISTRAR is unset |
repo-manager clone |
was absent; cloned 2026-08-21, rmgr installed into ~/.venvs/registrar |
statehub CLI |
not installable — pip install ~/state-hub fails ResolutionImpossible; hub-core installs, the conflict is elsewhere |
| Checkout ownership | 999 files under ~/state-hub owned by root, because the state-hub pod bind-mounts /home/tegwick as a hostPath and runs as root. Blocks git pull and editable installs |
| Checkout lineage | points at gitea-remote, not forgejo-remote — the host was never swept by forgejo-tier3-remote-url-sweep-playbook. Its evidence-binder tracks a different lineage and cannot see EBIND-WP-0002 at all |
The hostname heuristic is the cheapest fix and the most misleading defect: it
reads as though a registrar exists whenever a host is named railiance*, but the
production host is not. Prefer the explicit STATEHUB_REGISTRAR signal and treat
the hostname fallback as unreliable, or drop it.
Two consequences beyond this task:
-
Security. The
state-hubpod mounts/home/tegwickand/home/tegwick/.sshas hostPath volumes. A workload with the operator's private keys is a larger exposure than the registrar problem it surfaced. -
The remote sweep is incomplete, and this gates the registrar fix. Counted 2026-08-21:
Host forgejo-remotegitea-remoteworkstation 126 0 railiance017 70 The two aliases are different servers —
forgejo-remoteis92.205.62.239(railiance01),gitea-remoteis92.205.130.254(coulombcore) — and both are live. The workstation finished migrating;railiance01did not.forgejo-tier3-remote-url-sweep-playbookstill reads "Gitea remains canonical", which was true when written on 2026-07-04 and is now stale guidance.
Why the registrar cannot simply be switched on
The fix looked like one environment variable: the state-hub deployment has
STATEHUB_REGISTRAR unset, and that pod is the real automated instance — it
mounts /home/tegwick read-write, mounts /home/tegwick/.ssh at /root/.ssh
for push, and carries STATE_HUB_SWEEP_HOSTNAME. Host-level provisioning was the
wrong target.
Setting it today would be actively harmful. The pod would begin minting
identifiers into 70 checkouts that track the superseded server and pushing them
there — writing hub primary keys into a lineage nothing reads, and re-animating
gitea-remote as a write target.
Required order:
- Re-point
railiance01's checkouts toforgejo-remote, keepinggiteaas the rollback mirror the playbook specifies. Content diverges —evidence-binderthere is ata715492against the workstation's9e3f4a6— so this is a reconcile, not a URL rewrite. - Then set
STATEHUB_REGISTRAR=1on the deployment.
Reversing the order mints into the wrong lineage at fleet scale.
Answered, then executed (2026-08-21). The pod had pushed nothing: a
commit-level audit of all 70 repos found zero gitea-only commits. Forgejo was
strictly ahead everywhere (by 4–165 commits). railiance01 was a stale reader,
not a divergent writer — which is why its evidence-binder sat at a July commit
and could not see EBIND-WP-0002.
Two repos looked gitea-only under SSH ls-remote but were not: inter-hub had
been renamed to inter-hub-haskell (Forgejo answers the rename over HTTP with a
307, which SSH does not follow) and markitect_project to markitect-main.
Transition completed. Sweep pod scaled to zero, then per repo: origin
re-pointed to forgejo-remote, gitea retained as the rollback mirror the
playbook specifies, root-owned files chowned back, and the checkout brought onto
the forgejo lineage. Final state — 77 repos in sync, 0 ahead, 0 behind, 0 still
on gitea-remote.
Recovered rather than discarded:
freedom-intelligenceheld 6 unpushed commits — daily research briefs from 7–14 August byrein-aharness, present on no server. Rebased and pushed (846cccd..8832652). The first audit missed them because it measured against@{u}and silently skipped repos with no upstream configured; they surfaced only when the comparison was redone against the forgejo ref.- 5 working trees stashed as
pre-forgejo-transition-20260821. - 10 stale
custodian-syncstatus commits preserved onstale-sync-20260821branches before being dropped.
Remaining for the registrar: set STATEHUB_REGISTRAR=1 on the deployment and
scale back up. Note the pod runs as root and will re-create root-owned files
under /home/tegwick, undoing today's ownership fix — it needs a non-root
runAsUser or its own service account.
Interim recovery revised 2026-08-21: do not restore the host-wide sweep just
to drain UUID requests. Repo Manager now owns a scoped on-demand command,
rmgr registrar-reconcile, which preflights a clean synchronized Forgejo
checkout and the authoritative hub, serializes the run, and sets
STATEHUB_REGISTRAR=1 only for one repository's fix-consistency child.
It commits the identifier writeback and pushes only when explicitly requested.
This is the coding-agent recovery path until T03 lands; agents must neither set
the environment variable directly nor retry C-06/C-11 or send duplicate
messages. The production sweep remains disabled behind T12.
gitea cannot be decommissioned yet: the rollback remotes still point at it by
design. Drop them once a sweep or two confirms forgejo is healthy.
Contain the stale-lineage production sweep
id: RMGR-WP-0005-T11
status: done
priority: high
state_hub_task_id: "bddeb759-bab8-43ff-8cfc-a1d173827fac"
Audit the active state-hub sweep workload on railiance01 before enabling the
registrar. Establish which repositories it has fetched from or pushed to
gitea-remote, preserve non-secret evidence of the workload configuration and
recent Git outcomes, and prevent further stale-lineage writes with the smallest
reversible control.
Do not rewrite remote URLs as containment: the checkouts have diverged and need
the governed reconciliation described above. Do not enable
STATEHUB_REGISTRAR while any swept checkout still targets gitea-remote.
Done when the production sweep cannot push repository changes to the stale lineage, normal State Hub serving remains available, the SSH hostPath exposure is recorded for remediation, and the control and rollback are documented.
Result (2026-08-21): contained without interrupting the State Hub API.
- Production evidence showed the 15-minute Temporal schedule had fired 5,501
times. The final runs processed 12–15 repositories each, but every Git fetch
and push failed with
Resource temporarily unavailable; C-16/C-17 prevented further writes where repositories were behind or had unpushed commits. No checkout commit or reflog activity was found after 2026-08-19. The latent stale-lineage write path nevertheless remained live. - Paused Temporal schedule
activity-schedule-7c4e9a12-8f3b-4d5e-9c6a-1b2d3e4f5a6b. Its last run is fixed at2026-08-21T10:00:00Z; subsequent intervals did not fire. - Rolled State Hub deployment revision 11 with
/home/tegwickread-only and the/home/tegwick/.sshmount removed. The replacement pod is Ready and/state/healthremains healthy. - Made containment durable: State Hub production Helm sweep disabled in commit
2841bf3; activity-core projection disabled in5793eb3; Custodian-owned definition disabled in22d9366. All three commits are on Forgejomain. - Live activity-core ConfigMap and definition row say
enabled: false; the Helm chart passes lint and renders no sweep/SSH mounts; activity-core targeted tests pass (16 tests).
Rollback is deliberately gated: do not unpause the Temporal schedule or enable
the Helm sweep until the railiance01 checkouts are reconciled to
forgejo-remote, the registrar preflight passes, and T12 provides a scoped
credential path that does not mount an operator home or private-key directory.
Replace the host-wide sweep credential with a scoped identity
id: RMGR-WP-0005-T12
status: cancel
priority: high
state_hub_task_id: "0d69713d-0596-5415-b40c-8f5035641541"
The retired sweep design mounted all of /home/tegwick read-write and mounted
/home/tegwick/.ssh into the State Hub container. Before any remote sweep is
re-enabled, replace that host-wide authority with a workload-specific identity
and explicit repository scope. The runtime must not receive an operator private
key, an operator home directory, or implicit write access to every checkout.
Coordinate credential custody with the platform owner and keep the schedule disabled until positive allowed-repository and negative unrelated-repository push evidence exist without exposing credential values.
Cancelled 2026-08-31. The host-wide sweep is retired, not awaiting
re-enablement. Production continues with sweep.enabled: false, no checkout
hostPath, and no operator SSH mount. A future scoped sweep would be a new
capability with its own workplan and credential route; keeping a conditional
task open here incorrectly implies the retired design is on the critical path.
Re-register identifiers minted outside the registrar
id: RMGR-WP-0005-T02
status: cancel
priority: medium
state_hub_task_id: "7892ca86-bd89-537e-9361-d648dff8e4b2"
Records minted by non-registrar instances before T01 need reconciliation. Known cases, all created 2026-08-16/17 from the workstation hub:
RMGR-WP-0004(b8b3f1e0) and its seven tasks;CFED-WP-0001(7a96da54) and its thirteen tasks, plus theprj-canon-federationrepo record (3809b0ff);STATE-WP-0080— already reconciled by hand to the second instance's IDs (bbfce36a), retained here as the worked example.
Prefer waiting for T03 where possible: once identifiers are derived, these converge without manual intervention. Re-register by hand only what blocks work before then.
Scoped recovery (2026-08-22): a request from rail-kubernetes exposed two
registrar wrapper defects. The missing scan incorrectly excluded a newly
finished workplan, producing a false noop; after that was fixed, unrelated
C-03 assessment failures produced a false failed even though the requested
projection was complete. Commits 69adfff and c90f701 now distinguish a new
closed parent from frozen child-only gaps and verify the exact requested UUIDs
before accepting assessment exit 1.
The production central hub now contains deterministic RAIL-K8S-WP-0003 and
both done tasks, and the file writeback is pushed in rail-kubernetes commit
29305d4. A read-only repeat has no C-06/C-11 finding for that workplan. Its
two older workplans still carry production-absent random UUIDs and remain T04
work; T02 therefore stays wait. Evidence:
docs/evidence/RMGR-WP-0005-rail-kubernetes-registrar-2026-08-22.md.
Standalone-record and duplicate-source hardening (2026-08-22): owner
feedback exposed that the governed top-level intakes.md/decisions.md
filenames were absent from Repo Manager's explicit record-file allowlist.
Consequently an intake-only reconciliation could return noop; it happened to
work only when another missing record caused State Hub's broader pass to run.
The scanner now recognizes both lowercase and uppercase governed filenames,
and a direct scan of the Custodian source returns exactly CUST-IN-0014 as
missing.
The registrar and conformance paths now also implement the Custodian kind
registry's identity-reconciliation ruling. Repeated canonical id plus the same
non-null UUID is one indexed record with all source occurrences retained and a
governed cleanup warning; conflicting or incomplete UUID assignments fail
closed. The grandfathered MASON-0001 and MASON-0001-TNN forms are loaded
from the machine-readable canon mapping instead of being rewritten. Live proof
against ops-mason indexes MASON-0001-T01 once, preserves both task-block
locations, and emits the cleanup diagnostic. T02 remains wait because its
original pre-derivation reconciliation inventory is not yet exhausted.
Scoped identity compatibility (2026-08-23): Custodian owner work exposed
that the on-demand registrar still applied its collision gate repository-wide:
unrelated grandfathered bare T01–T09 collisions prevented assignment of the
canonical CUST-WP-0065-T01–T04 set. The preflight now intersects invalid IDs
and collision findings with the exact missing records requested by the command.
A requested defect still fails closed; unrelated history stays in evidence;
empty-projection bootstrap retains the full-repository gate. Regression coverage
proves all three cases.
The first Custodian proof passed the repaired identity gate but exceeded the old five-minute child timeout without writing an identifier. The wrapper now allows 15 minutes for large legacy repositories, returns structured timeout evidence instead of a traceback, and isolates its unit-test lock from live registrar activity.
The clean retry completed under correlation
76d57c01-5e16-428d-ae15-01a718f4b1a5: it assigned and centrally verified
CUST-WP-0065-T01–T04 plus CUST-IN-0014, committed their file writeback,
and pushed Custodian revision d792318. The full historical identity report
remained in evidence while blocking_invalid_identifiers and
blocking_identity_collisions were both empty, proving the intended boundary.
Superseded 2026-08-31. T03 and T04 replaced hand re-registration as the
governed convergence path. The fresh fleet plan covers every eligible live
record as replace, assign, or unchanged; closed pre-derivation UUIDs stay
frozen historical evidence. No remaining record should be repaired through the
old registrar-minting procedure, so this task is cancelled rather than left in
permanent conditional wait.
Derive identifiers deterministically
id: RMGR-WP-0005-T03
status: done
priority: high
state_hub_task_id: "28067729-498d-4f47-89bd-5b9718e999c7"
Replace minted UUIDs with UUIDv5 derived from the globally unique
PREFIX-WP-NNNN identifier (and PREFIX-WP-NNNN-TNN for tasks).
- Fix the namespace UUID and derivation input as a published contract — the value must be reproducible by any implementation, not just this one.
- Field shape is unchanged, so consumers reading
state_hub_workstream_idkeep working; only the provenance of the value changes. - Writeback becomes idempotent: two instances write identical bytes, so the flip-flop cannot recur regardless of how many hubs run.
Blocked on RMGR-WP-0004-T08. Deriving from a non-unique identifier
manufactures collisions: two repositories sharing PRJ-WP- would compute the
same UUID for different workplans. Uniqueness must be enforced first.
Unblocked 2026-08-21. RMGR-WP-0004-T08 closed 2026-08-18 (prefix registry
plus rmgr prefix-uniqueness), and ADR-007 was amended the same day with the
derivation scope this task needs:
- Derive for live records only; archived records keep frozen minted
identifiers. This is what reconciles
ADR-007§ Migration option 2 with the uniqueness derivation requires. - Derivation input is
(namespace, identifier)perADR-011decision 7. Namespace is the fleet branch, not the repository — the ecosystem is atN1, one implied namespace, so the pair does not disambiguate intra-namespace collisions and must not be read as if it did. - Derivation is not retroactive: existing live records keep minted UUIDs until deliberately re-derived.
- Un-archiving is a collision hazard — a record returning to live must be checked against the live namespace and renumbered if it clashes. Build this check alongside derivation, not after.
The live-collision remediation tracked on RMGR-WP-0004-T09 completed
2026-08-21. All 11 affected files are terminal history, and a fresh fleet
preflight found 263 live identifiers with zero collisions.
Progress (2026-08-21): the versioned derivation function and collision guard are
implemented in Repo Manager and published as
docs/work-record-uuid-derivation_v1.md. UUIDv5 uses fixed namespace UUID
a4058507-5c4a-5a00-ab06-fffa4fb46009 and exact name bytes
<fleet-namespace>\n<canonical-id>. rmgr identifier derive|preflight provides
independent reproduction and a hard live-collision refusal, including the
unarchive hazard. rmgr identifier migration-plan now emits the versioned,
non-mutating old-to-derived mapping needed by T04 and refuses a whole repository
when any live record is unsafe. Activation is gated only on declaration of the
current fleet namespace boundary/name: ADR-011 intentionally leaves that an
open governance question, so Repo Manager does not infer it from a hostname.
No existing minted identifier was rewritten implicitly.
Result (2026-08-21): fleet owner Bernd Worsch declared helixforge as the
authoritative namespace. config/fleet-namespace.yaml records the decision,
root lineage, and accurate N2 U1 R1 S1 posture with N3 U2 R2 S1 target.
Identifier commands now use that declaration by default while retaining an
explicit override for foreign namespaces. The derivation contract, fixed UUID,
collision/unarchive guard, and independent reproduction test are all active;
bulk replacement remains the separately governed T04 operation.
Ad-hoc identity correction (2026-08-23): Activity Core and Net Kingdom both
created the canon-prescribed ADHOC-2026-08-23 identifier. The repository is
not a namespace, so UUIDv5 correctly collided. Repo Manager's canon copy now
recognizes {PREFIX}-WP-ADHOC-YYYY-MM-DD[-TNN] as canonical and retains the
unqualified form only as grandfathered history. The governed Activity Core
retry registered ACTIVITY-WP-ADHOC-2026-08-23 and its task under correlation
f272ab73-35fb-44e7-aabd-23b124e3c135, with no missing identifiers afterward.
Migrate the fleet
id: RMGR-WP-0005-T04
status: progress
priority: high
state_hub_task_id: "3fd25c8f-25ea-566e-91c7-ac5ccba83009"
One-time pass over the 758 files carrying hub identifiers: compute the derived value, update the database to match, and write the file.
- Must be all-or-nothing per repository — a half-migrated repo has some derived and some minted identifiers and reconciles unpredictably.
- Records whose current identifier is already referenced externally (dashboards, saved queries, progress events) need a mapping table from old to derived, kept as provenance rather than discarded.
- Repositories with unresolved identifier collisions cannot migrate until
ADR-007§ Migration is ruled on; skip and report them rather than guessing.
Progress (2026-08-21): rmgr identifier migration-plan --root ... --namespace ... --output ... implements the non-mutating first phase. Its
versioned JSON retains every current-to-derived UUID, labels the repository as
the atomic unit, and skips an entire affected repository on a collision or
malformed live identifier. Applying those plans remains gated on the namespace
declaration and a compatibility write path that updates file and central
projection together.
Progress (2026-08-21, namespace declared): the helixforge fleet preflight
scanned 114 repository roots and found 242 live identifiers with zero
collisions. Fourteen non-canonical NetKingdom task ids were the sole initial
repository blocker and were normalized in commit 25ed928. The regenerated,
SHA-256-sealed mapping is source-ready across 42/42 affected repositories: 209
UUID replacements and 33 assignments, with per-repository HEAD/source
fingerprints verified by rmgr identifier migration-verify.
Apply is still deliberately unauthorized. State Hub's workplan/task primary
keys are referenced by numerous foreign keys without ON UPDATE CASCADE; a
file-first or direct-PK rewrite would split the projection. Evidence and the
required transactional/alias/rollback contract are recorded in
docs/evidence/RMGR-WP-0005-helixforge-migration-readiness-2026-08-21.md.
Central projection gate passed in isolation (2026-08-22): State Hub commit
cb1b028 adds Alembic revision b8d4f0a2c6e1, durable old→new alias
provenance, and a repository-atomic forward/reverse executor. All 20 foreign
keys into workplans/tasks now retain their delete policy while cascading primary
key updates. An isolated PostgreSQL upgrade/downgrade rehearsal verified 20/20
constraints in both directions and alias-table creation/removal; service tests
prove forward/reverse cascades and all-or-nothing failure, and the full State
Hub suite passes (622 passed). No live database or fleet file was migrated.
Repository file gate implemented (2026-08-22): rmgr identifier migration-files validates the plan seal and repository source before touching
bytes, requires the exact --confirm-plan-sha256 plus an explicit --execute,
prepares all mapped files before replacing any, restores original bytes on a
write failure, and supports reverse. Reverse restores replaced IDs and removes
fields introduced by action: assign. Focused forward/dry-run/reverse tests and
the full Repo Manager suite pass (70 passed); Ruff is clean.
The guard correctly rejected the 2026-08-21 fleet plan after this workplan changed, proving source drift cannot slip into apply. Regenerate and reseal that plan after this implementation commit. The remaining gate is a single-repository pilot that couples database apply, authoritative file rewrite, consistency verification, and database rollback if the file phase fails.
Pilot completed (2026-08-22): the refreshed sealed plan covered 43/43
eligible repositories, 245 mappings (212 replacements, 33 assignments), and no
collisions. The authorized repo-manager slice migrated six existing
workplan/task rows and file identifiers atomically. All referenced rows followed,
six aliases are durable, the local cache is fresh, and both old/new API lookups
have the expected 404/200 result. The production registrar was upgraded with
sweeping still disabled and rebuilt the formerly absent projection from the
same deterministic workplan UUID and all 12 authoritative task UUIDs. Exact
evidence, restore points, hashes, and residuals are in
docs/evidence/RMGR-WP-0005-repo-manager-pilot-2026-08-22.md. T04 remains in
progress: widening beyond this single repository requires a new explicit fleet
cutover decision and refreshed source fingerprints.
Batch 0002 prepared (2026-08-22): the original fleet seal now rejects 15
repositories with legitimate source drift. A refreshed zero-collision plan
covers 41 repositories and 230 records. New rmgr identifier migration-batch-plan pins an exact subset only after source verification plus
clean, Forgejo-synchronized Git preflight, and always emits
apply_authorized: false until a decision cites its batch hash.
RMGR-DEC-2026-001 now offers the directly executable second pilot:
whynot-design only, one workplan UUID replacement, no task replacements or
assignments. Workstation and production both resolve the old UUID, reject the
derived UUID, and agree on 9 tasks, 11 progress events, and 1 decision. The
sealed batch hash is
4039224352c6590fdc6b41e739539f1f9c92d098b43ed6e30b3cd9a87b53da65.
No mutation has run; approval, revision, or rejection of that exact package is
the remaining gate. Evidence:
docs/evidence/RMGR-WP-0005-batch-0002-readiness-2026-08-22.md.
Batch 0002 completed (2026-08-22): RMGR-DEC-2026-001 approved the exact
sealed whynot-design package. Fresh workstation and production restore points
were created and verified before both projection transactions replaced
WHYNOT-WP-0003 with derived UUID
d3a6ec16-ac40-5ffb-99e7-07f997e59a4a. The authoritative one-line file
transaction was pushed in whynot-design commit 9ffe187; both projections
now return old/new API lookups 404/200, retain 9 tasks, 11 progress events, and
1 decision, and carry one applied alias for the exact plan. Two consistency
passes ran against each projection. They exposed an existing active-versus-all-
tasks-done status-rule oscillation but no identifier or reference drift. Full
restore hashes, commits, checks, and the bounded warning are in
docs/evidence/RMGR-WP-0005-batch-0002-cutover-2026-08-22.md. T04 remains in
progress because the remaining fleet repositories require separately sealed
and approved batches.
Batch 0003 prepared (2026-08-22): markitect-main is the smallest next
repository whose two mapped rows exist in both current projections. Its clean,
Forgejo-synchronized source still matches the sealed fleet plan. Both
projections return old/new lookups 200/404 for MARKITECT-WP-0002 and
MARKITECT-WP-0002-T01, with 10 tasks and 9 progress events linked to the
workplan and no foreign-key reference to the mapped task. The fail-closed batch
hash is 494db8e101aeae7f8488082565dec3b4dd5c9c6ccb5f8812aaf233f813768f67.
RMGR-DEC-2026-002 is the directly actionable approval interface; no mutation
has run. Evidence:
docs/evidence/RMGR-WP-0005-batch-0003-markitect-readiness-2026-08-22.md.
Batch 0003 completed (2026-08-22): RMGR-DEC-2026-002 approved the exact
sealed markitect-main package. Fresh workstation and production restore
points were created and verified before both projection transactions replaced
the workplan and task UUIDs. The authoritative two-line file transaction was
pushed in markitect-main commit e26bd8b7. Both projections now return
old/new API lookups 404/200 for both records, retain 10 tasks and 9 progress
events under the derived workplan, retain zero references to the mapped task,
and carry two applied aliases. Two consistency passes against each projection
were identifier- and status-idempotent. Full restore hashes, commits, checks,
and pre-existing conformance warnings are in
docs/evidence/RMGR-WP-0005-batch-0003-markitect-cutover-2026-08-22.md. T04
remains in progress for separately sealed and approved fleet batches.
Batch 0004 prepared (2026-08-22): railiance-cluster is the smallest next
replacement batch represented in both current projections. Its clean source is
published and synchronized at 41fdfd9; the refreshed fleet plan and source
fingerprint verify without drift. The batch replaces RCLUSTER-WP-0007 and
five live task UUIDs. Its two cancelled task UUIDs remain outside the live
migration plan and follow the replaced parent through the database cascade.
Both projections contain the old six mapped rows and none of the targets. Each
retains seven child tasks, one decision, and five progress references to mapped
tasks; workstation and production retain their existing 14 and 12 workplan
progress events respectively. Every current foreign key into workplans/tasks
uses ON UPDATE CASCADE. The fail-closed batch hash is
554c6cbe119e13738efc6e2836254ddae00ea870f1cb99f2b5603336160e308d.
RMGR-DEC-2026-003 is the directly actionable approval interface; no
identifier mutation has run. Evidence:
docs/evidence/RMGR-WP-0005-batch-0004-railiance-cluster-readiness-2026-08-22.md.
Batch 0004 completed (2026-08-22): RMGR-DEC-2026-003 approved the exact
sealed railiance-cluster package. Fresh workstation and production restore
points were created and verified before both projection transactions replaced
the workplan and five live task UUIDs. The authoritative six-line file
transaction was pushed in commit 2bcc3f0; the two cancelled task UUIDs stayed
unchanged beneath the derived parent.
Both projections now return old/new API lookups 404/200 for all six mapped
records, retain seven child tasks, two cancelled tasks, one decision, five
mapped-task progress references, and their original 14/12 workplan progress
histories, and carry six applied aliases. Repeated consistency found no
remaining issue in the migrated slice. Production retains six pre-existing
C-03 failures for unrelated repository records that are absent from that
projection. Full restore hashes, commits, checks, and the bounded transport
note are in
docs/evidence/RMGR-WP-0005-batch-0004-railiance-cluster-cutover-2026-08-22.md.
T04 remains in progress for separately sealed and approved fleet batches.
Mixed projection convergence implemented (2026-08-31). Fast forge-derived
reconciliation can legitimately reach a deterministic target before the sealed
identifier cutover reaches its authoritative file. Projection preflight now
classifies every replacement as legacy_source, derived_target,
both_present, or neither_present. The first two may coexist within the same
repository-atomic batch; the latter two remain refusals. State Hub migrates only
legacy rows, verifies already-derived rows against repository and canonical
record identity, and writes the same durable alias provenance for both. Repo
Manager exposes the explicit primary-only database phase as rmgr identifier migration-projection; the existing migration-files command remains the
atomic file phase, and the projection phase is reversible if file writeback
fails. Focused State Hub and Repo Manager suites cover mixed convergence,
idempotent retry, ambiguity refusal, HTTP apply/reverse, primary identity, and
exact plan-seal checks.
Mixed-convergence pilot completed (2026-08-31). The fresh State Hub plan
contained nine replacements and 21 unchanged records. Projection preflight
classified all nine as legacy_source, with zero already-derived or ambiguous
units. The primary transaction migrated all nine rows, the atomic file phase
rewrote the nine authoritative UUID fields in
CUST-WP-0038-state-hub-threephoenix-ha.md, and verification returned old/new
lookups 404/200 for every mapping. The repository was pushed in commit
14e865a and exact-synced to the central projection. Approval evidence and the
sealed plan/batch are in
docs/evidence/RMGR-WP-0005-state-hub-pilot-{plan,batch}-2026-08-31.json and
RMGR-DEC-2026-004. This closes the operational gate for the mixed-state path;
T04 remains open only for the remaining fleet batches.
Adaptive-pricing rechecked (2026-08-31). A fresh plan at synchronized commit
5ff416f contains five unchanged records and no migration action: its
authoritative files already carry the derived UUIDs. Batch planning therefore
refuses with “repository has no migration action,” which is the correct no-op
classification. No repair or mutation is warranted. The next batch must come
from a regenerated fleet plan and select the smallest clean repository that
still has at least one replace action; stale 2026-08-22 batch evidence must not
be replayed.
Batch 0005 preflight blocked safely (2026-08-22): the refreshed zero-collision
fleet plan covers 39 repositories and 215 live records. A projection-aware
adaptive-pricing batch pinned its clean synchronized source and five UUID
replacements, then proved the workstation has all five old rows while production
has none. It therefore records ready_for_approval: false and authorizes no
mutation. A governed repair attempt retained State Hub's conservative C-03
refusal for absent random pre-derivation UUIDs.
That attempt exposed and fixed a Repo Manager fail-open result classification:
an empty ordinary-registration set could previously make an unverified exact
repair report applied. Repair and bootstrap modes now require their own exact
projection proof. Batch planning can also pin repeated
--projection-api-base endpoints and fails readiness unless every replacement
is current=200/derived=404 on every hub; verification repeats those live checks.
The remaining cross-owner gap is directly adoptable as
helixforge.identifiers.state-hub-sealed-projection-repair.v1. It asks State Hub
for a registrar-only, sealed, atomic restoration path without weakening normal
random stale-reference refusal. Evidence:
docs/evidence/RMGR-WP-0005-batch-0005-adaptive-pricing-preflight-2026-08-22.md.
Batch 0006 prepared (2026-08-31): a new non-mutating fleet plan supersedes
the 2026-08-22 source fingerprints. It covers 42 eligible repositories and 315
live records: 87 replacements, 6 assignments, 222 unchanged, zero skipped. The
plan seal is
af3c8e3fe6f423845a97258a10322eeded4b7814d13bf9541b5d550c285cc0b9.
reef-railiance is the smallest clean repository with a real replacement
action. Its synchronized source at 3e98b10 contains exactly two replacements:
REEF-RAILIANCE-WP-0003 and REEF-RAILIANCE-WP-0003-T04. Central preflight
classifies both as legacy_source, with zero derived, both-present,
neither-present, or unexpected states. The approval package is ready but
unauthorized; exact batch seal:
102061f3ed93eac091e9ceb05d9d3130241a23134b9f7d5a624d91c4ef824759.
RMGR-DEC-2026-005 is the approval interface. Evidence:
docs/evidence/RMGR-WP-0005-fleet-plan-2026-08-31.json and
docs/evidence/RMGR-WP-0005-batch-0006-reef-railiance-preflight-2026-08-31.json.
The inevitable Repo Manager evidence commit invalidates whole-fleet source
verification for repo-manager itself, but repository-filtered verification of
the sealed source plan for reef-railiance remains clean.
Batch 0006 applied (2026-08-31): Bernd Worsch approved exactly batch seal
102061f3ed93eac091e9ceb05d9d3130241a23134b9f7d5a624d91c4ef824759
through RMGR-DEC-2026-005. The final batch verification passed with no drift.
The railiance01 primary then migrated both legacy rows transactionally, with
zero assignments or already-derived rows, and the atomic file phase changed
only the two sealed UUID fields in
REEF-RAILIANCE-WP-0003-rapp-qonto-production-gates.md. Post-cutover checks
confirmed both derived records, the task-to-workplan relationship, unchanged
lifecycle state, a clean consistency audit, and the pushed reef commit
dcf51b3. The initial projection attempt refused safely before mutation because
the CLI default misspelled the primary label as railliance01; the corrected
railiance01 default is covered by the same change set. Evidence:
docs/evidence/RMGR-WP-0005-batch-0006-reef-railiance-cutover-2026-08-31.md.
Retire the interim rule
id: RMGR-WP-0005-T05
status: wait
priority: low
state_hub_task_id: "f2634e3a-b32e-5862-85bf-78437fc976dc"
Once derivation is live fleet-wide, remove the single-writer restriction from T01. Multiple hub instances become an availability choice rather than a correctness constraint, and disconnected registration works again.
Confirm before removal: two instances reconciling the same repository produce byte-identical writeback, and neither creates a duplicate record.
Rebuild local instances as caches
id: RMGR-WP-0005-T07
status: done
priority: high
state_hub_task_id: "e3c2a791-f632-59f3-9929-b179e45c77d0"
Implement ADR-010 decisions 1–3: the central hub on railiance is authoritative
as a reading of the repositories; local instances become rebuildable caches.
- A cache must be discardable and reconstructable from repository files alone, with no work lost.
- Local work must not require a hub — repository files are self-describing, so reading them is sufficient for working inside a repo.
- Cache reads are advisory and must carry their staleness (
ADR-010decision 8).
Measured 2026-08-17: 955 workplans locally against 649 on the primary, 320 local-only, of which 288 are backed by files that all exist on disk. That portion of the divergence is redundant and needs no merge — only a rebuild.
Progress (2026-08-21): rmgr cache status|rebuild now makes the local
repository projection explicitly advisory and reports canonical observation
time, age, source fingerprint, current/indexed Git revision, and concrete stale
reasons. The fingerprint covers the authoritative record/classification/intent
bytes, including uncommitted changes. A live drill correctly rejected the
legacy index as stale, rebuilt 73 records from repository files, and immediately
reported a fresh cache.
The 217 closed unbound rows are preserved in a minimized, integrity-sealed
export at
docs/evidence/RMGR-WP-0005-closed-provenance-2026-08-21.json; live unbound
rows remain zero. docs/cache-rebuild_v1.md gates any database replacement on
an isolated projection rebuild and proof that hub-native records reached their
one central owner. Remaining before done: compare against an isolated rebuilt
State Hub database and verify hub-native centralization. The documented remote
tunnel at 127.0.0.1:18000 was not running during this pass, so no central
comparison was inferred.
Production comparison (2026-08-22): the registrar began with no Repo
Manager workplan projection. After the contained rebuild, RMGR-WP-0005 and
all 12 task UUIDs exactly match the authoritative file, and a second read-only
pass reports no issue for that workplan. This proves discard-and-reconstruct for
the pilot record without identifier writeback. T07 remains in progress because
the full isolated projection rebuild and hub-native centralization proof still
cover more than this one workplan.
Full isolated rebuild (2026-08-22): the guarded
--bootstrap-empty-projection path rebuilt a fresh disposable State Hub from
the repository's existing authoritative identifiers: 9 workplans, 62 tasks,
2 intakes, and 0 decisions, with no missing records. A second read-only pass
reported zero issues and all 9 source bindings synchronized. The path refuses a
non-empty target projection, multi-repository scans, unsynchronized Forgejo
checkouts, and any authoritative file mutation. State Hub commits 03c7924
and 059de93; Repo Manager commits 707fb85 and 6e9c93e.
The remaining gate is now measured rather than inferred. The workstation hub
has 35 Repo Manager-linked progress events and 13 messages to/from
repo-manager; production has 2 and 0 respectively. The disposable rebuild has
zero of each by design. These hub-native records cannot be reconstructed from
files, and mutable message state cannot use unordered append replay. T07 stays
progress until the declared central owner migrates or dispositions them and a
repeat comparison passes. Evidence and the non-destructive procedure are in
docs/evidence/RMGR-WP-0005-isolated-rebuild-2026-08-22.md and
docs/cache-rebuild_v1.md.
Done with explicit disposition 2026-08-31. The workstation State Hub API is
gone; its loopback address is now a relay to the sole railiance01 primary.
Repo Manager's advisory cache rebuilt 96 current records at commit e8e2747
and immediately reported fresh with matching source fingerprints. The old
local-only hub-native rows were not migrated before the local database was
replaced: the current local PostgreSQL database contains zero matching progress
events or messages. They are therefore dispositioned as discarded cache
history, not claimed as reconstructed. Central now exclusively owns new
hub-native writes. Exact counts and the limitation are retained in
docs/evidence/RMGR-WP-0005-cache-centralization-2026-08-31.md.
Separate file-derived from hub-native data
id: RMGR-WP-0005-T08
status: done
priority: high
state_hub_task_id: "241cf058-2f3e-4d49-8cc9-5c714be4a1cf"
Implement ADR-010 decision 4. The two kinds need opposite handling:
- File-derived (workplans, tasks, statuses, dependencies) — central derives it and must not accept pushes of it (decision 5). Offline, the git commit is the write. No conflict model: conflicts are git conflicts.
- Hub-native (progress events, decisions, inbox messages, token events) — central owns it, needs a real write path and a local append-only buffer for replay. No conflict model either: append-only merges regardless of order.
Deliverable is an explicit classification of every record type the hub holds, with its truth source and offline behaviour, so neither kind is handled by the other's rules.
Feeds a rescope of STATE-WP-0068 (offline write buffer and edge relay): under
this split most of what it buffers does not need buffering, and only the
append-only stream does. Re-examine before building further on it — this likely
reduces its scope.
Result (2026-08-21): config/hub-record-authority.yaml classifies all 42
tables imported by the current State Hub model package; a regression test keeps
the inventory explicit. docs/hub-record-authority_v1.md defines truth source,
offline behaviour, write rule, and target owner. The audit adapted the original
binary wording: hub-native origin contains both append-only events and mutable
control records. Only progress_events and token_events enter the generic
append/replay path; mutable message, launch, domain/topic, interface, and catalog
state requires an owner-specific ordered command or stays online-only. Decisions
and intakes are file-derived under the ratified work-record standard, superseding
ADR-010's older example list.
Disposition the orphaned hub-first records
id: RMGR-WP-0005-T09
status: done
priority: high
state_hub_task_id: "d40cc4a8-4280-4940-ac1d-dc1049f1b678"
28 records exist in the local instance with no backing file. They are the only
records a cache rebuild would drop, so they must be classified first
(ADR-010 § Orphan disposition):
- Broken links — a file exists but
backing_filenamewas never recorded.RMGR-WP-0004is a confirmed instance. Repair the link; no data at risk. Likely the largest class, so classify before estimating the rest. - Live hub-first records —
proposed/ready/backlogwith no file, inactivity-core,core-hub,hub-core,issue-core,ops-hub,prj-forgejo-org-refactor,railiance-enablement,railiance-infra,reef-railiance. Write a repository file or drop explicitly. These areADR-001violations and must not survive as hub-only records. - Closed hub-first records —
finished/archivedwith no file. Retain as provenance where cheap; do not reconstruct completed plans.
Blocks T07 — rebuilding the cache before this classification would discard class 2.
Note: one of these records is already labelled SPURIOUS bootstrap (statehub register collision) in repo-manager, independent corroboration of the
STATE-WP-0080 defect.
Rebaselined 2026-08-21: the local cache now has 224 rows without a backing
filename, not 28: 217 are closed (finished/archived) provenance and seven are
live. All seven live rows have matching files across net-kingdom, state-hub,
railiance-platform, railiance-enablement, railiance-infra, and ops-hub;
there are zero confirmed live hub-first rows in this snapshot. Exact ids and
paths are captured in
docs/evidence/RMGR-WP-0005-orphan-disposition-2026-08-21.md. Remaining work is
to repair those seven bindings and retain/export the 217 closed rows before T07
replaces the cache.
Result (2026-08-21): the preferred binding-sync API accepted all seven
verified file mappings atomically (updated: 7), without changing repository
files or lifecycle state. The cache now has zero unbound live workplans. The
remaining 217 unbound rows are exclusively closed provenance (201 finished,
16 archived) and are explicitly retained in the untouched current database;
T07 must export or preserve them as its pre-rebuild gate.
Assign one authoritative hub per record
id: RMGR-WP-0005-T10
status: done
priority: medium
state_hub_task_id: "15f0f167-a8d0-4d5c-8576-3e93b1e8792f"
Implement ADR-010 decision 7. The retirement splits one hub into several, which
is permitted only if every record has exactly one authoritative hub, determined
by its repository and domain.
Define and enforce that mapping before the split lands. Without it the peer-database divergence this workplan exists to remove recurs at larger scale.
Coordinate with the hub-extension architecture in
prj-state-hub-retirement/architecture/; hub-core owns the hub-native side.
Result (2026-08-21): the T08 machine-readable inventory is now an
executable, fail-closed routing contract. rmgr authority route resolves every
known record type to exactly one owner; file-derived routes require repository
and domain context, while hub-native/cache routes name their central owner.
Unknown record types and claimed-owner mismatches are rejected. Contract-wide
tests prove all 42 current record types resolve once. Finished HUB-WP-0004
accepts the complementary boundary: repository authority stays with Repo
Manager and hub-core owns the named hub-native projection/event ports.
Protect lifecycle status from automation
id: RMGR-WP-0005-T06
status: done
priority: medium
state_hub_task_id: "d440d59c-f78e-4752-84c7-f3d5fdf7d3c3"
Implement ADR-007 decision 3: an automated normalization pass may report
lifecycle drift but may not promote a workplan from proposed to active.
proposed means awaiting human review; automated promotion destroys the gate.
Observed: commit ff909e1 ("renormalize lifecycle state [auto]") promoted
STATE-WP-0080 to active hours after it was drafted for review.
Extend the same protection to task status, where the symptom is currently
sharper: C-15 forces CFED-WP-0001-T02 back to wait on every sync
regardless of file content — reproduced three times, via file edit and via
update_task_status, with the task never holding todo. Establish which
direction wins for task status and make it consistent with ADR-001, where the
file originates work.
Result (2026-08-18): In state-hub consistency: C-23 does not
auto-promote proposed → active (report only). C-15 no longer
writebacks wait over progress/todo; file wins via C-10 (ADR-001).
C-15 remains a non-fixable warning when the DB is terminal and the file
is not.