the-custodian/docs/recovery/tunnels-yaml-proposed-changes-CUST-WP-0067.md
codex 2755e8c91f
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 1s
feat(workplan): close CUST-WP-0067-T08 — central MCP deployed and verified
state-hub-mcp serves SSE on ClusterIP 10.43.110.80:8001, verified from the node
rather than through a tunnel, and reads central's data end-to-end (79 repos).
Remote port map and dev-hub registration repointed off the reverse tunnel.

state-hub-mcp-railiance01 is now safe to remove. state-hub-railiance01 is not:
~120 AGENTS.md files still depend on it until T07.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 23:20:11 +02:00

62 lines
2.5 KiB
Markdown

# Proposed `~/.config/bridge/tunnels.yaml` changes — CUST-WP-0067-T02
Claude could not edit this file (outside any repo; blocked by the permission
classifier). Apply manually or grant the write. **Back up first:**
```bash
cp ~/.config/bridge/tunnels.yaml ~/.config/bridge/tunnels.yaml.bak-$(date +%Y%m%d%H%M%S)
```
## 1. `state-hub-primary` — health check probes the wrong thing
Its `local_port: 8000` is already correct and needs no change. The health check
does:
```yaml
health_check:
- url: http://127.0.0.1:8000/state/health
+ url: http://127.0.0.1:8000/state/health # correct only now that no local hub binds 8000
```
No edit required today, but note *why* it read healthy for seven weeks: it
probed the local cache, not the tunnel it opened on `[::1]:8000`. A tunnel
health check that can be satisfied by a different process is not a health check.
Prefer probing through the tunnel's own bind address once instance identity
lands (T03).
## 2. `state-hub-mcp-railiance01` — retire (was: probes the wrong port)
Superseded 2026-08-24. An MCP server now runs on central
(`state-hub-mcp`, ClusterIP `10.43.110.80:8001`, CUST-WP-0067-T08), so this
tunnel has nothing left depending on it — the documented `dev-hub` registration
has been repointed at the ClusterIP. **Remove the entry** rather than fixing its
health check, which probed `:8000` while forwarding `:8001`.
```yaml
- state-hub-mcp-railiance01: # -R 18001 -> workstation:8001
```
## 3. Reverse relay tunnels — retire
```yaml
- state-hub-railiance01: # -R 18000 -> workstation:8000
- state-hub-mcp-railiance01: # -R 18001 -> workstation:8001
```
Both make a remote box dial back into this workstation to reach a hub. That was
correct when the workstation *was* the hub. It is not: the primary runs on
railiance01, so an agent there currently routes
`localhost:18000 -> workstation:8000 -> jump host -> 10.43.68.154:8000` to reach
a service on its own machine.
Status 2026-08-24:
- `state-hub-mcp-railiance01`**safe to remove now.** Central MCP is serving
and the global port map points at it.
- `state-hub-railiance01`**not yet.** The global instructions are repointed,
but roughly 120 `AGENTS.md` files across both machines still tell agents to
use `127.0.0.1:18000`. Removing it before `CUST-WP-0067-T07` repoints those
breaks any session that follows its own repo's instructions.
On railiance01 both services are reachable in-cluster with no tunnel at all —
this is where "abandon tunneling" genuinely applies.