Consistency sweep and recently-on-scope retry transient 502/503/504 then return a degraded context snapshot instead of thrashing Temporal. Document edge-relay resilience. Retire Binky dual-clock host timers after claim-loop smoke.
3.2 KiB
3.2 KiB
Edge relay resilience (ACTIVITY-WP-0027-T06)
Audience: operators and agents debugging high-frequency activity-core jobs
Scope: railiance01 actcore-statehub-edge-relay + worker STATE_HUB_URL
Topology
actcore-worker
STATE_HUB_URL=http://actcore-statehub-edge-relay:8000
│
▼
edge relay (cluster)
· GET allowlist → upstream or stale read-cache
· POST side-effects → upstream (or 503 if busy/unreachable)
│
▼
state-hub service (cluster DNS state-hub.state-hub.svc …)
· may depend on tunnel / central hub posture
Host rein-aharness should prefer the same edge ClusterIP (or k8s://…edge-relay:8000)
so completion events outbox when central hub is flaky — see
docs/llm-connect-host-access.md.
Resolver classes
| Class | Examples | Hard-fail? | Notes |
|---|---|---|---|
| GET reads | fi_brief_status, domain_summary, state_summary |
Soft by default ({} / empty) unless required: true |
Edge may serve stale cache (X-StateHub-Edge-Cache: stale) |
| Side-effect POSTs | consistency_sweep_remote_all, recently_on_scope_hourly |
After retries: degrade by default | See env below |
| Required pure reads | todo_md_staleness when enabled |
Yes if required: true |
Prefer not to make hub-only reads required without cache |
Side-effect POST behaviour (worker)
Implemented in context_resolvers/state_hub.py (_post_json):
- Retry transient
502/503/504and network/timeout errors
(STATE_HUB_POST_RETRIES, default 3; backoffSTATE_HUB_POST_RETRY_BACKOFF_SECONDS, default 2s × attempt). - If still failing and degrade is on (
STATE_HUB_SIDE_EFFECT_DEGRADEdefault true, or per-sourceparams.degrade_on_unavailable):- consistency_sweep: returns
{exit_code: 75, lock_skipped: true, repos_processed: [], degraded: true, …}so the workflow completes without Temporal ApplicationError storms. - recently_on_scope: returns
{generated: [], failed: [{reason: edge_unavailable}], degraded: true, …}.
- consistency_sweep: returns
- Set
degrade_on_unavailable: falseon a definition (or envSTATE_HUB_SIDE_EFFECT_DEGRADE=false) to keep hard-fail + Temporal retry after in-resolver attempts.
Operator checks
# Edge health (from host with ClusterIP or kubectl exec)
curl -sS http://<edge-cluster-ip>:8000/edge/health | python3 -m json.tool
# Worker still pointing at edge
kubectl -n activity-core exec deploy/actcore-worker -- env | grep STATE_HUB
# Recent consistency fires should COMPLETE even when edge was briefly 503
# (look for degraded:true in context_snapshot if all retries failed)
What edge 503 usually means
- Upstream hub slow / lock held / temporary overload on
/consistency/sweep/remote-all(large fleet). - Upstream unreachable; edge refuses to invent a false success for POSTs.
- Outbox backlog (see
edge/health→outbox.pending_count) is write backlog for queued progress — separate from sweep POST 503.
Non-goals
- Making the edge invent successful sweep results when nothing ran.
- Public exposure of State Hub or edge.
- Replacing Temporal schedules with workstation heartbeats.