feat(workplan): rescope CUST-WP-0067 to retire the local hub instance

Operator decision 2026-08-24: rather than making two hub instances coexist
safely, retire the second one. The local postgres+uvicorn instance is what
impersonates central and is redundant with ADR-010 decision 3 plus Repo
Manager's file-derived index.

With it gone, state-hub-primary binds 127.0.0.1:8000 unchanged and every
existing default becomes correct with no call-site edits.

Adds the cache-only recovery export (44 records) as the T05 source.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
codex 2026-08-24 22:24:01 +02:00
parent 9dbd1f7ff9
commit bd8f02b785
2 changed files with 1157 additions and 58 deletions

File diff suppressed because it is too large Load diff

View file

@ -100,29 +100,46 @@ and print the resolved API base so the operator can see which instance answered.
defaults throughout, and prints `API base:` as its second line. Verified against defaults throughout, and prints `API base:` as its second line. Verified against
the live API. the live API.
## Eliminate the port collision ## Retire the local hub instance and the reverse-tunnel relay
```task ```task
id: CUST-WP-0067-T02 id: CUST-WP-0067-T02
status: todo status: progress
priority: high priority: high
state_hub_task_id: "4093e928-d752-5a91-96c3-2e80f0e1dac5" state_hub_task_id: "4093e928-d752-5a91-96c3-2e80f0e1dac5"
``` ```
Move `state-hub-primary` off `local_port: 8000` to `18000`, matching the port Decision, 2026-08-24: rather than making two hub instances coexist safely,
map already reserved for the primary in the global agent instructions. Leave the retire the second one. The local `postgres:16-alpine` + uvicorn instance is what
local cache on 8000 so nothing that currently resolves to it changes behaviour — impersonates central, and it is redundant — `ADR-010` decision 3 already states
this step is deliberately non-breaking and must stay that way. that local work requires no hub at all, and Repo Manager already maintains a
file-derived local index (`index_store.py`, `rmgr cache status` / `cache
rebuild`). Two caching layers exist and one of them is a database pretending to
be the primary.
Extend the `ops-bridge` duplicate-port guard so it rejects a `direction: local` With the local instance gone, `state-hub-primary` binds `127.0.0.1:8000`
tunnel whose `local_port` is already bound by any listener, not only by another unchanged and every existing `http://127.0.0.1:8000` default becomes correct
bridge tunnel. The current guard could not have caught this. without editing a single call site. The port collision cannot recur because only
one process binds the port.
Acceptance: `state-hub-primary` binds `127.0.0.1:18000`; `[::1]:8000` no longer Retire the reverse tunnels too. `state-hub-railiance01` forwards a remote box's
serves a hub; `curl 127.0.0.1:18000/state/health` reports the central instance `:18000` back to this workstation's `:8000`, so a remote agent following the
and `curl 127.0.0.1:8000/state/health` reports the cache; the two return documented port map reaches the workstation rather than the primary — which on
different repo counts; the guard fails a synthetic config that reintroduces the railiance01 is its own machine. That topology assumed the workstation was the
collision. hub; it has not been since the primary moved.
Sequence matters: export the cache-only records first, stop serving second,
discard the cache data only after T05 proves central holds everything.
Acceptance: no local hub process listening; `127.0.0.1:8000` answers from
central; MCP `dev-hub` resolves to central; reverse `state-hub-*` tunnels removed
or repointed; the cache-only recovery export is committed.
**Done (2026-08-24):** `docs/recovery/cache-only-repos-2026-08-24.json` captures
all 44 cache-only repository records with working-copy presence, classification
file presence, and HEAD sha — the recovery source for T05.
`ops-bridge` now pins local forwards to `127.0.0.1` (commit `2213847`), so a
contested port fails loudly instead of silently landing on `[::1]`.
## Make the hub target explicit and unspoofable ## Make the hub target explicit and unspoofable
@ -133,19 +150,19 @@ priority: high
state_hub_task_id: "29977448-1a74-5a4d-b729-974c15b6bbde" state_hub_task_id: "29977448-1a74-5a4d-b729-974c15b6bbde"
``` ```
Remove the `http://127.0.0.1:8000` default from every call site that claims to Retiring the local instance removes today's impersonator but not the ability for
reach the primary: `custodian_cli.py:28`, `statehub_register.py:22`, a future one to appear. Give the hub an instance identity it can assert — role
`repo_manager/cli.py:56,405`, `repo_manager/commands/registrar_reconcile.py:410`. served from the health or summary endpoint — and make
`registrar-reconcile --confirm-primary` refuse anything that does not assert
`primary`.
A default that silently resolves to a cache is worse than a missing one. Give `_check_primary` currently asserts only `status == ok` and `db == connected`.
the hub an identity it can assert — instance role served from the health or Both instances passed it. It is a liveness check wearing an authority check's
summary endpoint — and make `registrar-reconcile --confirm-primary` refuse to name, and it is what allowed a cache to certify itself as the registrar.
run against anything that does not assert `primary`. Confirming against a cache
is a false green and is the specific failure that let this run for seven weeks.
Acceptance: `--confirm-primary` against the cache exits non-zero with a message Acceptance: `--confirm-primary` fails against a non-primary instance and names
naming the instance it reached; against central it succeeds; no code path what it reached; a hub reports its role; `statehub status` shows role alongside
reaches a hub without an explicitly resolved target. the API base.
## Give Repo Manager a real onboarding write path ## Give Repo Manager a real onboarding write path
@ -156,21 +173,21 @@ priority: high
state_hub_task_id: "ffa141d5-331d-543d-8b87-516f27973a22" state_hub_task_id: "ffa141d5-331d-543d-8b87-516f27973a22"
``` ```
Repo Manager owns `managed_repos` as `file-derived` and has no command for it. Repo Manager owns `managed_repos` as `file-derived` in
Add onboarding that follows ADR-010 decision 5: write `.repo-classification.yaml` `hub-record-authority.yaml` and exposes no command for it. Add onboarding that
in the target repository, commit, push, and have central derive the record. follows `ADR-010` decision 5: write `.repo-classification.yaml` in the target
Central must not accept a push of derived state, so the command's job is to make repository, commit, push, and have central derive the record. Central must not
the source file correct and reachable, then trigger and verify derivation. accept a push of derived state, so the command makes the source file correct and
reachable, then triggers and verifies derivation.
Resolve the bootstrap gap explicitly — central derives from repositories it Resolve the bootstrap gap explicitly — central derives from repositories it
already knows about, so a never-registered repository is never scanned. The already knows about, so a never-registered repository is never scanned. The
onboarding path must be able to introduce a repository central has not seen. onboarding path must be able to introduce a repository central has not seen.
Acceptance: onboarding a fresh repository from the workstation produces a Acceptance: onboarding a fresh repository from the workstation produces a central
central record with no manual step; re-running is idempotent; the cache is not record with no manual step; re-running is idempotent.
written directly.
## Backfill the 44 cache-only registrations ## Onboard the 44 cache-only repositories to central
```task ```task
id: CUST-WP-0067-T05 id: CUST-WP-0067-T05
@ -179,22 +196,21 @@ priority: medium
state_hub_task_id: "078159e6-5ecd-5f9d-b3ec-1fabf60955f7" state_hub_task_id: "078159e6-5ecd-5f9d-b3ec-1fabf60955f7"
``` ```
Runs only after T02T04; backfilling before the target is unambiguous refills Runs after T04. With the cache retired this is no longer a convergence of two
the cache. Drive the 43 on-disk repositories through the T04 path. Nine lack hubs — it is onboarding 44 repositories to central from their files, which is
`.repo-classification.yaml` (`binky-control`, `clay-borg`, T04 applied to the recovery export.
`direkt-vermittlung-de`, `polycode-sim`, `railiance-telemetry`, `ralph-workplan`,
`rein-openweights`, `testdrive-jsui`, `timeline-svg`) and need one authored with
the owner rather than guessed — classification is not mechanical, per
`CUST-WP-0065-T01`. Push `soul-frame` and `rein-openweights` first.
`agent-harness` has no working copy: decide restore-from-remote or drop, and Ten of the 44 lack `.repo-classification.yaml` and need one authored with the
record the decision. Do not preserve it as a hub-only record — that is the owner rather than guessed; classification is not mechanical, per
ADR-001 violation ADR-010 calls out. `CUST-WP-0065-T01`. Push `soul-frame` and `rein-openweights` first — both are
ahead of their remote, and central derives from what it can fetch.
`agent-harness` has no working copy: decide restore-from-remote or drop and
record it. Do not preserve it as a hub-only record.
Acceptance: central and cache repo counts converge; the cache-only set is empty Acceptance: every record in the recovery export exists on central or carries a
or every remainder has a written disposition. written disposition; only then may the local cache database be discarded.
## Correct the ADR-010 repo-record framing ## Correct the ADR-010 framing and record the retirement
```task ```task
id: CUST-WP-0067-T06 id: CUST-WP-0067-T06
@ -203,16 +219,22 @@ priority: medium
state_hub_task_id: "007bfcf1-3b17-5d4a-b16a-b80ebf273934" state_hub_task_id: "007bfcf1-3b17-5d4a-b16a-b80ebf273934"
``` ```
ADR-010 decision 2 says a divergent database is a merge problem and a stale `ADR-010` decision 2 says a divergent database is a merge problem and a stale
cache is a refresh problem. For `managed_repos` neither holds: the gap is a cache is a refresh problem. For `managed_repos` neither held: the gap was a
strict subset in the cache's favour, and refreshing destroys rather than strict subset in the cache's favour, and refreshing would have destroyed rather
reconciles. Record the third shape — cache-only records whose authoritative than reconciled. Record that third shape — cache-only records whose
source exists but was never introduced to central — and state that its remedy is authoritative source exists but was never introduced to central — with
re-derivation from source, not refresh and not merge. re-derivation from source as its remedy.
Also correct the implicit assumption that the ADR's own remediation happened. Record the mechanical cause, which the ADR observed but did not diagnose: an
The measurement stands; the fix did not land, and the ADR reads as though it did. unbound `ssh -L` binds every loopback family, so the IPv4 bind losing to a local
listener still leaves a working `[::1]` forward and `ExitOnForwardFailure` never
fires. The ADR treated the shared port as the hazard; the missing bind address
was what made it silent.
Acceptance: ADR-010 revised with a superseding note dated and linked to this Note also that ADR-010 reads as though its remediation landed. It did not — the
workplan; the port-collision remediation recorded as an outcome rather than an condition it measured was still live seven weeks later. Supersede decision 2 for
observation. this record class and record the local-instance retirement as the outcome.
Acceptance: ADR-010 revised with a dated superseding note linked to this
workplan; `ops-bridge` and the port map documented as the structural fix.