activity-core/docs/edge-relay-resilience.md

78 lines
3.2 KiB
Markdown
Raw Permalink Normal View History

# Edge relay resilience (ACTIVITY-WP-0027-T06)
**Audience:** operators and agents debugging high-frequency activity-core jobs
**Scope:** railiance01 `actcore-statehub-edge-relay` + worker `STATE_HUB_URL`
## Topology
```text
actcore-worker
STATE_HUB_URL=http://actcore-statehub-edge-relay:8000
edge relay (cluster)
· GET allowlist → upstream or stale read-cache
· POST side-effects → upstream (or 503 if busy/unreachable)
state-hub service (cluster DNS state-hub.state-hub.svc …)
· may depend on tunnel / central hub posture
```
Host rein-aharness should prefer the **same edge ClusterIP** (or `k8s://…edge-relay:8000`)
so completion events outbox when central hub is flaky — see
`docs/llm-connect-host-access.md`.
## Resolver classes
| Class | Examples | Hard-fail? | Notes |
| ----- | -------- | ---------- | ----- |
| **GET reads** | `fi_brief_status`, `domain_summary`, `state_summary` | Soft by default (`{}` / empty) unless `required: true` | Edge may serve **stale cache** (`X-StateHub-Edge-Cache: stale`) |
| **Side-effect POSTs** | `consistency_sweep_remote_all`, `recently_on_scope_hourly` | After retries: **degrade** by default | See env below |
| **Required pure reads** | `todo_md_staleness` when enabled | Yes if `required: true` | Prefer not to make hub-only reads required without cache |
## Side-effect POST behaviour (worker)
Implemented in `context_resolvers/state_hub.py` (`_post_json`):
1. **Retry** transient `502/503/504` and network/timeout errors
(`STATE_HUB_POST_RETRIES`, default **3**; backoff
`STATE_HUB_POST_RETRY_BACKOFF_SECONDS`, default **2s** × attempt).
2. If still failing and degrade is on (`STATE_HUB_SIDE_EFFECT_DEGRADE` default
**true**, or per-source `params.degrade_on_unavailable`):
- **consistency_sweep:** returns
`{exit_code: 75, lock_skipped: true, repos_processed: [], degraded: true, …}`
so the workflow **completes** without Temporal ApplicationError storms.
- **recently_on_scope:** returns
`{generated: [], failed: [{reason: edge_unavailable}], degraded: true, …}`.
3. Set `degrade_on_unavailable: false` on a definition (or env
`STATE_HUB_SIDE_EFFECT_DEGRADE=false`) to keep hard-fail + Temporal retry
after in-resolver attempts.
## Operator checks
```bash
# Edge health (from host with ClusterIP or kubectl exec)
curl -sS http://<edge-cluster-ip>:8000/edge/health | python3 -m json.tool
# Worker still pointing at edge
kubectl -n activity-core exec deploy/actcore-worker -- env | grep STATE_HUB
# Recent consistency fires should COMPLETE even when edge was briefly 503
# (look for degraded:true in context_snapshot if all retries failed)
```
## What edge 503 usually means
- Upstream hub slow / lock held / temporary overload on
`/consistency/sweep/remote-all` (large fleet).
- Upstream unreachable; edge refuses to invent a false success for POSTs.
- Outbox backlog (see `edge/health``outbox.pending_count`) is **write**
backlog for queued progress — separate from sweep POST 503.
## Non-goals
- Making the edge invent successful sweep results when nothing ran.
- Public exposure of State Hub or edge.
- Replacing Temporal schedules with workstation heartbeats.