Operator decision 2026-08-24: rather than making two hub instances coexist safely, retire the second one. The local postgres+uvicorn instance is what impersonates central and is redundant with ADR-010 decision 3 plus Repo Manager's file-derived index. With it gone, state-hub-primary binds 127.0.0.1:8000 unchanged and every existing default becomes correct with no call-site edits. Adds the cache-only recovery export (44 records) as the T05 source. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
240 lines
10 KiB
Markdown
240 lines
10 KiB
Markdown
---
|
||
id: CUST-WP-0067
|
||
type: workplan
|
||
title: "Resolve hub target ambiguity and restore repo onboarding to the authoritative hub"
|
||
domain: infotech
|
||
repo: the-custodian
|
||
status: active
|
||
owner: codex
|
||
created: "2026-08-24"
|
||
updated: "2026-08-24"
|
||
quality_dor: DoR-Ok
|
||
quality_dor_at: "2026-08-24"
|
||
quality_dor_by: codex
|
||
quality_dor_note: "Live divergence measured (122 cache / 78 central, strict subset, 0 central-only); root cause located in configuration (port collision plus 127.0.0.1 defaults in four call sites); recovery source confirmed on disk and pushed for 43 of 44 records; owner boundaries follow ADR-010 and hub-record-authority.yaml."
|
||
state_hub_workstream_id: "16249302-2767-55df-aec0-d92c2751c225"
|
||
---
|
||
|
||
# Resolve hub target ambiguity and restore repo onboarding to the authoritative hub
|
||
|
||
## Goal
|
||
|
||
Make "the authoritative hub" a resolvable address rather than a policy claim, so
|
||
that repository onboarding lands on central instead of a local cache that
|
||
believes it is primary. Then recover the seven weeks of workstation-originated
|
||
repo registrations that never reached central, and correct the ADR-010 framing
|
||
that treats this class of gap as a refresh problem.
|
||
|
||
## Context
|
||
|
||
`ADR-010` documented two State Hub instances sharing port 8000, separated only
|
||
by IP family, and named the consequence: every tool defaulting to `127.0.0.1`
|
||
reaches the local instance while believing it is the primary. That condition was
|
||
never remediated. As of 2026-08-24 it is still live:
|
||
|
||
```text
|
||
LISTEN 127.0.0.1:8000 uvicorn (pid 3258) local cache
|
||
LISTEN [::1]:8000 ssh -L 8000:10.43.68.154:8000 central
|
||
```
|
||
|
||
The collision is declared in `~/.config/bridge/tunnels.yaml`, where the tunnel
|
||
`state-hub-primary` takes `local_port: 8000` — the port the local uvicorn
|
||
already holds. `ssh -L` binds both address families, the IPv4 bind loses to
|
||
uvicorn, and the IPv6 bind succeeds, so `ExitOnForwardFailure=yes` does not
|
||
fire. `_reject_duplicate_local_ports` in `ops-bridge` guards only bridge tunnel
|
||
against bridge tunnel; it cannot see a non-bridge listener.
|
||
|
||
Measured repository divergence, 2026-08-24:
|
||
|
||
```text
|
||
cache (127.0.0.1:8000) 122
|
||
central ([::1]:8000) 78
|
||
central-only 0
|
||
cache-only 44
|
||
```
|
||
|
||
Central is a strict subset. That is not the shape ADR-010 anticipated. A refresh
|
||
of the cache would delete the 44 rather than reconcile them, and central holds
|
||
no path to re-derive records for repositories it has never been told exist.
|
||
|
||
Onboarding by registration date shows a clean break:
|
||
|
||
| Registered | central | cache-only |
|
||
|---|---|---|
|
||
| 2026-02 – 2026-06 | 74 | 0 |
|
||
| 2026-07 | 1 | 24 |
|
||
| 2026-08 | 3 | 20 |
|
||
|
||
Central is otherwise live — `last_state_synced_at` runs to 2026-08-22, and the
|
||
three August rows (`repo-manager`, `rail-kubernetes`, `fin-hub`) were registered
|
||
from the railiance side. Only workstation-originated onboarding is lost.
|
||
|
||
Nothing is unrecoverable: 43 of the 44 have a working copy on disk, all 43 have
|
||
an `origin` remote and are pushed (`soul-frame` and `rein-openweights` are ahead
|
||
of their remote), and 35 carry `.repo-classification.yaml`. `agent-harness` has
|
||
no working copy and needs separate disposition.
|
||
|
||
`hub-record-authority.yaml` assigns `managed_repos` to `repo-manager` as
|
||
`file-derived`. Repo Manager owns the record type but exposes no command that
|
||
onboards a repository into a hub, and `registrar-reconcile --confirm-primary`
|
||
defaults its "primary" to `127.0.0.1:8000` — confirming projections against the
|
||
cache and reporting success while central holds nothing.
|
||
|
||
## Repair the crashing status printer
|
||
|
||
```task
|
||
id: CUST-WP-0067-T01
|
||
status: done
|
||
priority: low
|
||
state_hub_task_id: "f3608db4-20a5-58fb-a965-885eb14858af"
|
||
```
|
||
|
||
`statehub status` raised `KeyError: 'in_progress'` at `custodian_cli.py:565`.
|
||
The task vocabulary is `wait|todo|progress|done|cancel`; `in_progress` and
|
||
`blocked` are not task statuses in this schema and the totals block never
|
||
carried them. Fix the key names, tolerate the `workplans`/`workstreams` rename,
|
||
and print the resolved API base so the operator can see which instance answered.
|
||
|
||
**Done (2026-08-24):** `custodian_cli.py` `cmd_status` now reads
|
||
`progress`/`todo`/`wait`, accepts either totals key for workplans, uses `.get()`
|
||
defaults throughout, and prints `API base:` as its second line. Verified against
|
||
the live API.
|
||
|
||
## Retire the local hub instance and the reverse-tunnel relay
|
||
|
||
```task
|
||
id: CUST-WP-0067-T02
|
||
status: progress
|
||
priority: high
|
||
state_hub_task_id: "4093e928-d752-5a91-96c3-2e80f0e1dac5"
|
||
```
|
||
|
||
Decision, 2026-08-24: rather than making two hub instances coexist safely,
|
||
retire the second one. The local `postgres:16-alpine` + uvicorn instance is what
|
||
impersonates central, and it is redundant — `ADR-010` decision 3 already states
|
||
that local work requires no hub at all, and Repo Manager already maintains a
|
||
file-derived local index (`index_store.py`, `rmgr cache status` / `cache
|
||
rebuild`). Two caching layers exist and one of them is a database pretending to
|
||
be the primary.
|
||
|
||
With the local instance gone, `state-hub-primary` binds `127.0.0.1:8000`
|
||
unchanged and every existing `http://127.0.0.1:8000` default becomes correct
|
||
without editing a single call site. The port collision cannot recur because only
|
||
one process binds the port.
|
||
|
||
Retire the reverse tunnels too. `state-hub-railiance01` forwards a remote box's
|
||
`:18000` back to this workstation's `:8000`, so a remote agent following the
|
||
documented port map reaches the workstation rather than the primary — which on
|
||
railiance01 is its own machine. That topology assumed the workstation was the
|
||
hub; it has not been since the primary moved.
|
||
|
||
Sequence matters: export the cache-only records first, stop serving second,
|
||
discard the cache data only after T05 proves central holds everything.
|
||
|
||
Acceptance: no local hub process listening; `127.0.0.1:8000` answers from
|
||
central; MCP `dev-hub` resolves to central; reverse `state-hub-*` tunnels removed
|
||
or repointed; the cache-only recovery export is committed.
|
||
|
||
**Done (2026-08-24):** `docs/recovery/cache-only-repos-2026-08-24.json` captures
|
||
all 44 cache-only repository records with working-copy presence, classification
|
||
file presence, and HEAD sha — the recovery source for T05.
|
||
`ops-bridge` now pins local forwards to `127.0.0.1` (commit `2213847`), so a
|
||
contested port fails loudly instead of silently landing on `[::1]`.
|
||
|
||
## Make the hub target explicit and unspoofable
|
||
|
||
```task
|
||
id: CUST-WP-0067-T03
|
||
status: todo
|
||
priority: high
|
||
state_hub_task_id: "29977448-1a74-5a4d-b729-974c15b6bbde"
|
||
```
|
||
|
||
Retiring the local instance removes today's impersonator but not the ability for
|
||
a future one to appear. Give the hub an instance identity it can assert — role
|
||
served from the health or summary endpoint — and make
|
||
`registrar-reconcile --confirm-primary` refuse anything that does not assert
|
||
`primary`.
|
||
|
||
`_check_primary` currently asserts only `status == ok` and `db == connected`.
|
||
Both instances passed it. It is a liveness check wearing an authority check's
|
||
name, and it is what allowed a cache to certify itself as the registrar.
|
||
|
||
Acceptance: `--confirm-primary` fails against a non-primary instance and names
|
||
what it reached; a hub reports its role; `statehub status` shows role alongside
|
||
the API base.
|
||
|
||
## Give Repo Manager a real onboarding write path
|
||
|
||
```task
|
||
id: CUST-WP-0067-T04
|
||
status: todo
|
||
priority: high
|
||
state_hub_task_id: "ffa141d5-331d-543d-8b87-516f27973a22"
|
||
```
|
||
|
||
Repo Manager owns `managed_repos` as `file-derived` in
|
||
`hub-record-authority.yaml` and exposes no command for it. Add onboarding that
|
||
follows `ADR-010` decision 5: write `.repo-classification.yaml` in the target
|
||
repository, commit, push, and have central derive the record. Central must not
|
||
accept a push of derived state, so the command makes the source file correct and
|
||
reachable, then triggers and verifies derivation.
|
||
|
||
Resolve the bootstrap gap explicitly — central derives from repositories it
|
||
already knows about, so a never-registered repository is never scanned. The
|
||
onboarding path must be able to introduce a repository central has not seen.
|
||
|
||
Acceptance: onboarding a fresh repository from the workstation produces a central
|
||
record with no manual step; re-running is idempotent.
|
||
|
||
## Onboard the 44 cache-only repositories to central
|
||
|
||
```task
|
||
id: CUST-WP-0067-T05
|
||
status: todo
|
||
priority: medium
|
||
state_hub_task_id: "078159e6-5ecd-5f9d-b3ec-1fabf60955f7"
|
||
```
|
||
|
||
Runs after T04. With the cache retired this is no longer a convergence of two
|
||
hubs — it is onboarding 44 repositories to central from their files, which is
|
||
T04 applied to the recovery export.
|
||
|
||
Ten of the 44 lack `.repo-classification.yaml` and need one authored with the
|
||
owner rather than guessed; classification is not mechanical, per
|
||
`CUST-WP-0065-T01`. Push `soul-frame` and `rein-openweights` first — both are
|
||
ahead of their remote, and central derives from what it can fetch.
|
||
`agent-harness` has no working copy: decide restore-from-remote or drop and
|
||
record it. Do not preserve it as a hub-only record.
|
||
|
||
Acceptance: every record in the recovery export exists on central or carries a
|
||
written disposition; only then may the local cache database be discarded.
|
||
|
||
## Correct the ADR-010 framing and record the retirement
|
||
|
||
```task
|
||
id: CUST-WP-0067-T06
|
||
status: todo
|
||
priority: medium
|
||
state_hub_task_id: "007bfcf1-3b17-5d4a-b16a-b80ebf273934"
|
||
```
|
||
|
||
`ADR-010` decision 2 says a divergent database is a merge problem and a stale
|
||
cache is a refresh problem. For `managed_repos` neither held: the gap was a
|
||
strict subset in the cache's favour, and refreshing would have destroyed rather
|
||
than reconciled. Record that third shape — cache-only records whose
|
||
authoritative source exists but was never introduced to central — with
|
||
re-derivation from source as its remedy.
|
||
|
||
Record the mechanical cause, which the ADR observed but did not diagnose: an
|
||
unbound `ssh -L` binds every loopback family, so the IPv4 bind losing to a local
|
||
listener still leaves a working `[::1]` forward and `ExitOnForwardFailure` never
|
||
fires. The ADR treated the shared port as the hazard; the missing bind address
|
||
was what made it silent.
|
||
|
||
Note also that ADR-010 reads as though its remediation landed. It did not — the
|
||
condition it measured was still live seven weeks later. Supersede decision 2 for
|
||
this record class and record the local-instance retirement as the outcome.
|
||
|
||
Acceptance: ADR-010 revised with a dated superseding note linked to this
|
||
workplan; `ops-bridge` and the port map documented as the structural fix.
|