# Edge relay resilience (ACTIVITY-WP-0027-T06) **Audience:** operators and agents debugging high-frequency activity-core jobs **Scope:** railiance01 `actcore-statehub-edge-relay` + worker `STATE_HUB_URL` ## Topology ```text actcore-worker STATE_HUB_URL=http://actcore-statehub-edge-relay:8000 │ ▼ edge relay (cluster) · GET allowlist → upstream or stale read-cache · POST side-effects → upstream (or 503 if busy/unreachable) │ ▼ state-hub service (cluster DNS state-hub.state-hub.svc …) · may depend on tunnel / central hub posture ``` Host rein-aharness should prefer the **same edge ClusterIP** (or `k8s://…edge-relay:8000`) so completion events outbox when central hub is flaky — see `docs/llm-connect-host-access.md`. ## Resolver classes | Class | Examples | Hard-fail? | Notes | | ----- | -------- | ---------- | ----- | | **GET reads** | `fi_brief_status`, `domain_summary`, `state_summary` | Soft by default (`{}` / empty) unless `required: true` | Edge may serve **stale cache** (`X-StateHub-Edge-Cache: stale`) | | **Side-effect POSTs** | `consistency_sweep_remote_all`, `recently_on_scope_hourly` | After retries: **degrade** by default | See env below | | **Required pure reads** | `todo_md_staleness` when enabled | Yes if `required: true` | Prefer not to make hub-only reads required without cache | ## Side-effect POST behaviour (worker) Implemented in `context_resolvers/state_hub.py` (`_post_json`): 1. **Retry** transient `502/503/504` and network/timeout errors (`STATE_HUB_POST_RETRIES`, default **3**; backoff `STATE_HUB_POST_RETRY_BACKOFF_SECONDS`, default **2s** × attempt). 2. If still failing and degrade is on (`STATE_HUB_SIDE_EFFECT_DEGRADE` default **true**, or per-source `params.degrade_on_unavailable`): - **consistency_sweep:** returns `{exit_code: 75, lock_skipped: true, repos_processed: [], degraded: true, …}` so the workflow **completes** without Temporal ApplicationError storms. - **recently_on_scope:** returns `{generated: [], failed: [{reason: edge_unavailable}], degraded: true, …}`. 3. Set `degrade_on_unavailable: false` on a definition (or env `STATE_HUB_SIDE_EFFECT_DEGRADE=false`) to keep hard-fail + Temporal retry after in-resolver attempts. ## Operator checks ```bash # Edge health (from host with ClusterIP or kubectl exec) curl -sS http://:8000/edge/health | python3 -m json.tool # Worker still pointing at edge kubectl -n activity-core exec deploy/actcore-worker -- env | grep STATE_HUB # Recent consistency fires should COMPLETE even when edge was briefly 503 # (look for degraded:true in context_snapshot if all retries failed) ``` ## What edge 503 usually means - Upstream hub slow / lock held / temporary overload on `/consistency/sweep/remote-all` (large fleet). - Upstream unreachable; edge refuses to invent a false success for POSTs. - Outbox backlog (see `edge/health` → `outbox.pending_count`) is **write** backlog for queued progress — separate from sweep POST 503. ## Non-goals - Making the edge invent successful sweep results when nothing ran. - Public exposure of State Hub or edge. - Replacing Temporal schedules with workstation heartbeats.