Compare commits

..

No commits in common. "1da94a207fe69fdd39695013be5dce73c9da61a6" and "9dbd1f7ff9af6280c71f394ae73b2cca4c39315d" have entirely different histories.

3 changed files with 58 additions and 1235 deletions

File diff suppressed because it is too large Load diff

View file

@ -1,58 +0,0 @@
# Proposed `~/.config/bridge/tunnels.yaml` changes — CUST-WP-0067-T02
Claude could not edit this file (outside any repo; blocked by the permission
classifier). Apply manually or grant the write. **Back up first:**
```bash
cp ~/.config/bridge/tunnels.yaml ~/.config/bridge/tunnels.yaml.bak-$(date +%Y%m%d%H%M%S)
```
## 1. `state-hub-primary` — health check probes the wrong thing
Its `local_port: 8000` is already correct and needs no change. The health check
does:
```yaml
health_check:
- url: http://127.0.0.1:8000/state/health
+ url: http://127.0.0.1:8000/state/health # correct only now that no local hub binds 8000
```
No edit required today, but note *why* it read healthy for seven weeks: it
probed the local cache, not the tunnel it opened on `[::1]:8000`. A tunnel
health check that can be satisfied by a different process is not a health check.
Prefer probing through the tunnel's own bind address once instance identity
lands (T03).
## 2. `state-hub-mcp-railiance01` — probes the wrong port
Forwards `:8001`, probes `:8000`:
```yaml
state-hub-mcp-railiance01:
health_check:
- url: http://127.0.0.1:8000/state/health
+ url: http://127.0.0.1:8001/state/health
```
## 3. Reverse relay tunnels — retire
```yaml
- state-hub-railiance01: # -R 18000 -> workstation:8000
- state-hub-mcp-railiance01: # -R 18001 -> workstation:8001
```
Both make a remote box dial back into this workstation to reach a hub. That was
correct when the workstation *was* the hub. It is not: the primary runs on
railiance01, so an agent there currently routes
`localhost:18000 -> workstation:8000 -> jump host -> 10.43.68.154:8000` to reach
a service on its own machine.
**Before removing**, repoint remote agents. On railiance01 the primary is
reachable in-cluster with no tunnel at all — this is where "abandon tunneling"
genuinely applies. The global agent instructions' remote port map
(`State Hub API http://127.0.0.1:18000`) must be updated in the same change, or
remote sessions will silently lose the hub.
Sequencing: repoint remote agents and update the port map first, then remove
these two entries. Removing them first breaks every remote session.

View file

@ -100,66 +100,29 @@ and print the resolved API base so the operator can see which instance answered.
defaults throughout, and prints `API base:` as its second line. Verified against defaults throughout, and prints `API base:` as its second line. Verified against
the live API. the live API.
## Retire the local hub instance and the reverse-tunnel relay ## Eliminate the port collision
```task ```task
id: CUST-WP-0067-T02 id: CUST-WP-0067-T02
status: progress status: todo
priority: high priority: high
state_hub_task_id: "4093e928-d752-5a91-96c3-2e80f0e1dac5" state_hub_task_id: "4093e928-d752-5a91-96c3-2e80f0e1dac5"
``` ```
Decision, 2026-08-24: rather than making two hub instances coexist safely, Move `state-hub-primary` off `local_port: 8000` to `18000`, matching the port
retire the second one. The local `postgres:16-alpine` + uvicorn instance is what map already reserved for the primary in the global agent instructions. Leave the
impersonates central, and it is redundant — `ADR-010` decision 3 already states local cache on 8000 so nothing that currently resolves to it changes behaviour —
that local work requires no hub at all, and Repo Manager already maintains a this step is deliberately non-breaking and must stay that way.
file-derived local index (`index_store.py`, `rmgr cache status` / `cache
rebuild`). Two caching layers exist and one of them is a database pretending to
be the primary.
With the local instance gone, `state-hub-primary` binds `127.0.0.1:8000` Extend the `ops-bridge` duplicate-port guard so it rejects a `direction: local`
unchanged and every existing `http://127.0.0.1:8000` default becomes correct tunnel whose `local_port` is already bound by any listener, not only by another
without editing a single call site. The port collision cannot recur because only bridge tunnel. The current guard could not have caught this.
one process binds the port.
Retire the reverse tunnels too. `state-hub-railiance01` forwards a remote box's Acceptance: `state-hub-primary` binds `127.0.0.1:18000`; `[::1]:8000` no longer
`:18000` back to this workstation's `:8000`, so a remote agent following the serves a hub; `curl 127.0.0.1:18000/state/health` reports the central instance
documented port map reaches the workstation rather than the primary — which on and `curl 127.0.0.1:8000/state/health` reports the cache; the two return
railiance01 is its own machine. That topology assumed the workstation was the different repo counts; the guard fails a synthetic config that reintroduces the
hub; it has not been since the primary moved. collision.
Sequence matters: export the cache-only records first, stop serving second,
discard the cache data only after T05 proves central holds everything.
Acceptance: no local hub process listening; `127.0.0.1:8000` answers from
central; MCP `dev-hub` resolves to central; reverse `state-hub-*` tunnels removed
or repointed; the cache-only recovery export is committed.
**Progress (2026-08-24):** `docs/recovery/cache-only-repos-2026-08-24.json`
captures all 44 cache-only repository records with working-copy presence,
classification file presence, and HEAD sha — the recovery source for T05.
`ops-bridge` now pins local forwards to `127.0.0.1` (commit `2213847`), so a
contested port fails loudly instead of silently landing on `[::1]`.
The local hub instance is retired: `make api` (uvicorn on `127.0.0.1:8000`) is
stopped, `state-hub-primary` restarted onto the freed IPv4 address, and the
orphaned `[::1]` forward removed. Exactly one process now binds 8000 and it is
the tunnel to central — `/repos/` returns 78 there, not the cache's 122, and
`statehub status` reports central's 10 active workplans rather than 27. The MCP
server on `:8001` needed no change: it targets `http://127.0.0.1:8000` and now
proxies central. No call site was edited; retiring the impersonator made the
existing defaults correct.
The Postgres container is deliberately left running with its data intact. It is
the recovery source of last resort until T05 proves central holds all 44, per
the sequencing above.
**Remaining:** the reverse relay tunnels `state-hub-railiance01` and
`state-hub-mcp-railiance01` still route remote agents back to this workstation.
Retiring them requires editing `~/.config/bridge/tunnels.yaml`, which is outside
any repo and was blocked in-session; the change and its sequencing constraint
are recorded in `docs/recovery/tunnels-yaml-proposed-changes-CUST-WP-0067.md`.
Remote agents and the documented port map must be repointed *before* removal.
## Make the hub target explicit and unspoofable ## Make the hub target explicit and unspoofable
@ -170,19 +133,19 @@ priority: high
state_hub_task_id: "29977448-1a74-5a4d-b729-974c15b6bbde" state_hub_task_id: "29977448-1a74-5a4d-b729-974c15b6bbde"
``` ```
Retiring the local instance removes today's impersonator but not the ability for Remove the `http://127.0.0.1:8000` default from every call site that claims to
a future one to appear. Give the hub an instance identity it can assert — role reach the primary: `custodian_cli.py:28`, `statehub_register.py:22`,
served from the health or summary endpoint — and make `repo_manager/cli.py:56,405`, `repo_manager/commands/registrar_reconcile.py:410`.
`registrar-reconcile --confirm-primary` refuse anything that does not assert
`primary`.
`_check_primary` currently asserts only `status == ok` and `db == connected`. A default that silently resolves to a cache is worse than a missing one. Give
Both instances passed it. It is a liveness check wearing an authority check's the hub an identity it can assert — instance role served from the health or
name, and it is what allowed a cache to certify itself as the registrar. summary endpoint — and make `registrar-reconcile --confirm-primary` refuse to
run against anything that does not assert `primary`. Confirming against a cache
is a false green and is the specific failure that let this run for seven weeks.
Acceptance: `--confirm-primary` fails against a non-primary instance and names Acceptance: `--confirm-primary` against the cache exits non-zero with a message
what it reached; a hub reports its role; `statehub status` shows role alongside naming the instance it reached; against central it succeeds; no code path
the API base. reaches a hub without an explicitly resolved target.
## Give Repo Manager a real onboarding write path ## Give Repo Manager a real onboarding write path
@ -193,21 +156,21 @@ priority: high
state_hub_task_id: "ffa141d5-331d-543d-8b87-516f27973a22" state_hub_task_id: "ffa141d5-331d-543d-8b87-516f27973a22"
``` ```
Repo Manager owns `managed_repos` as `file-derived` in Repo Manager owns `managed_repos` as `file-derived` and has no command for it.
`hub-record-authority.yaml` and exposes no command for it. Add onboarding that Add onboarding that follows ADR-010 decision 5: write `.repo-classification.yaml`
follows `ADR-010` decision 5: write `.repo-classification.yaml` in the target in the target repository, commit, push, and have central derive the record.
repository, commit, push, and have central derive the record. Central must not Central must not accept a push of derived state, so the command's job is to make
accept a push of derived state, so the command makes the source file correct and the source file correct and reachable, then trigger and verify derivation.
reachable, then triggers and verifies derivation.
Resolve the bootstrap gap explicitly — central derives from repositories it Resolve the bootstrap gap explicitly — central derives from repositories it
already knows about, so a never-registered repository is never scanned. The already knows about, so a never-registered repository is never scanned. The
onboarding path must be able to introduce a repository central has not seen. onboarding path must be able to introduce a repository central has not seen.
Acceptance: onboarding a fresh repository from the workstation produces a central Acceptance: onboarding a fresh repository from the workstation produces a
record with no manual step; re-running is idempotent. central record with no manual step; re-running is idempotent; the cache is not
written directly.
## Onboard the 44 cache-only repositories to central ## Backfill the 44 cache-only registrations
```task ```task
id: CUST-WP-0067-T05 id: CUST-WP-0067-T05
@ -216,21 +179,22 @@ priority: medium
state_hub_task_id: "078159e6-5ecd-5f9d-b3ec-1fabf60955f7" state_hub_task_id: "078159e6-5ecd-5f9d-b3ec-1fabf60955f7"
``` ```
Runs after T04. With the cache retired this is no longer a convergence of two Runs only after T02–T04; backfilling before the target is unambiguous refills
hubs — it is onboarding 44 repositories to central from their files, which is the cache. Drive the 43 on-disk repositories through the T04 path. Nine lack
T04 applied to the recovery export. `.repo-classification.yaml` (`binky-control`, `clay-borg`,
`direkt-vermittlung-de`, `polycode-sim`, `railiance-telemetry`, `ralph-workplan`,
`rein-openweights`, `testdrive-jsui`, `timeline-svg`) and need one authored with
the owner rather than guessed — classification is not mechanical, per
`CUST-WP-0065-T01`. Push `soul-frame` and `rein-openweights` first.
Ten of the 44 lack `.repo-classification.yaml` and need one authored with the `agent-harness` has no working copy: decide restore-from-remote or drop, and
owner rather than guessed; classification is not mechanical, per record the decision. Do not preserve it as a hub-only record — that is the
`CUST-WP-0065-T01`. Push `soul-frame` and `rein-openweights` first — both are ADR-001 violation ADR-010 calls out.
ahead of their remote, and central derives from what it can fetch.
`agent-harness` has no working copy: decide restore-from-remote or drop and
record it. Do not preserve it as a hub-only record.
Acceptance: every record in the recovery export exists on central or carries a Acceptance: central and cache repo counts converge; the cache-only set is empty
written disposition; only then may the local cache database be discarded. or every remainder has a written disposition.
## Correct the ADR-010 framing and record the retirement ## Correct the ADR-010 repo-record framing
```task ```task
id: CUST-WP-0067-T06 id: CUST-WP-0067-T06
@ -239,22 +203,16 @@ priority: medium
state_hub_task_id: "007bfcf1-3b17-5d4a-b16a-b80ebf273934" state_hub_task_id: "007bfcf1-3b17-5d4a-b16a-b80ebf273934"
``` ```
`ADR-010` decision 2 says a divergent database is a merge problem and a stale ADR-010 decision 2 says a divergent database is a merge problem and a stale
cache is a refresh problem. For `managed_repos` neither held: the gap was a cache is a refresh problem. For `managed_repos` neither holds: the gap is a
strict subset in the cache's favour, and refreshing would have destroyed rather strict subset in the cache's favour, and refreshing destroys rather than
than reconciled. Record that third shape — cache-only records whose reconciles. Record the third shape — cache-only records whose authoritative
authoritative source exists but was never introduced to central — with source exists but was never introduced to central — and state that its remedy is
re-derivation from source as its remedy. re-derivation from source, not refresh and not merge.
Record the mechanical cause, which the ADR observed but did not diagnose: an Also correct the implicit assumption that the ADR's own remediation happened.
unbound `ssh -L` binds every loopback family, so the IPv4 bind losing to a local The measurement stands; the fix did not land, and the ADR reads as though it did.
listener still leaves a working `[::1]` forward and `ExitOnForwardFailure` never
fires. The ADR treated the shared port as the hazard; the missing bind address
was what made it silent.
Note also that ADR-010 reads as though its remediation landed. It did not — the Acceptance: ADR-010 revised with a superseding note dated and linked to this
condition it measured was still live seven weeks later. Supersede decision 2 for workplan; the port-collision remediation recorded as an outcome rather than an
this record class and record the local-instance retirement as the outcome. observation.
Acceptance: ADR-010 revised with a dated superseding note linked to this
workplan; `ops-bridge` and the port map documented as the structural fix.