From 5afd2f347dbcd2ffcc9149d19fdbe94c4d120432 Mon Sep 17 00:00:00 2001 From: tegwick Date: Tue, 21 Jul 2026 04:09:57 +0200 Subject: [PATCH] Add ACTIVITY-WP-0021 production automation reliability workplan MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Captures the Sunday–Monday railiance01 schedule review: IssueSink 503s, Binky cutover failures, edge-relay/ROS issues, silent daily triage, sync_schedules bug, and DB restart thrash — with tasks to fix and verify. --- ...ITY-WP-0021-prod-automation-reliability.md | 302 ++++++++++++++++++ 1 file changed, 302 insertions(+) create mode 100644 workplans/ACTIVITY-WP-0021-prod-automation-reliability.md diff --git a/workplans/ACTIVITY-WP-0021-prod-automation-reliability.md b/workplans/ACTIVITY-WP-0021-prod-automation-reliability.md new file mode 100644 index 0000000..3b7ac89 --- /dev/null +++ b/workplans/ACTIVITY-WP-0021-prod-automation-reliability.md @@ -0,0 +1,302 @@ +--- +id: ACTIVITY-WP-0021 +type: workplan +title: "Production automation reliability after Sunday–Monday fire review" +domain: infotech +repo: activity-core +status: ready +owner: codex +topic_slug: activity-core +created: "2026-07-21" +updated: "2026-07-21" +--- + +# Production automation reliability after Sunday–Monday fire review + +## Problem + +Operator review of railiance01 activity-core schedules for +**2026-07-19 00:00 → 2026-07-21 ~01:30 Europe/Berlin** found that **Temporal +schedules fire on time**, but several high-value outcomes failed or were +invisible: + +| Finding | Evidence | +| --- | --- | +| **IssueSink emit path broken** | Binky daily, Binky mail intake, SBOM weekly, and TODO-stale workflows **FAILED** at `emit_tasks` with **HTTP 503** from `actcore-issue-core-bridge…:8765/issues/` (retries exhausted). No `activity_runs` row for failed Binky/SBOM fires. | +| **Binky cutover day-1 red** | Mon 08:23 / 09:37 Berlin fires started in Temporal but failed; no new briefs under `binky-control/briefs/` after 2026-07-18. Mon–Wed verification window from cutover is at risk. | +| **Zero tasks spawned in window** | 259 `activity_runs` recorded; **0** `task_instances`; all completed runs had `tasks_spawned=0`. | +| **Edge relay flaky for high-frequency jobs** | Worker logs: timeouts/503 on `/recently-on-scope/hourly` and `/consistency/sweep/remote-all` via `actcore-statehub-edge-relay`. | +| **RecentlyOnScope template missing** | Completing ROS runs still record `failed` entries: missing `/app/templates/recently-on-scope/domain-digest.md` in the worker/image path used by the relay or resolver. | +| **Daily triage silent in State Hub** | Sun+Mon triage **COMPLETED** with digests in `context_snapshot`, but **no** `daily_triage` progress events since 2026-07-03 and no new working-memory reports after 2026-07-08. | +| **`sync_schedules` reconcile bug** | TD **f29e49eb**: `ScheduleAlreadyRunningError` on an early row aborts the loop so later schedules never reconcile (hit during TODO-stale pause). | +| **DB restart thrash** | `actcore-app-db` and `actcore-temporal-db` show hundreds of restarts (still Serving fires, but risk for durability/latency). | +| **Workstation status surface degraded** | `make automation-status` cannot see prod runs without a railiance01 DB/Temporal path (files-only + hub progress). | + +Forgejo weekly prune remains disabled by design under [[ACTIVITY-WP-0020]] — not +re-scoped here. + +## Goal + +Restore **observable, successful scheduled outcomes** on railiance01: + +1. Task emission works for definitions that intentionally emit (or they use a + non-Forgejo sink aligned with current policy). +2. Binky Mon–Wed verification can complete with clean scheduled runs. +3. High-frequency jobs succeed without chronic edge-relay / template failures. +4. Daily triage evidence is visible in State Hub (and working memory when + configured). +5. Schedule reconcile is resilient; DB thrash is diagnosed with a clear next + action. +6. Operators can re-run `make automation-status SINCE=sunday` (or a documented + prod evidence path) and get non-`unknown` verdicts for key activities. + +## Out of scope + +- Finishing Forgejo package prune enablement ([[ACTIVITY-WP-0020]]). +- Re-enabling Daily TODO.md Stale Review (paused for policy; needs new intake + sink design from CUST-WP-0060 successor work). +- New activity definitions unrelated to the reliability gaps above. + +## Dependencies + +- railiance01 SSH + in-cluster kubectl (no k3s tunnel; host path as in WP-0014). +- issue-core bridge / IssueSink routing (`warden route show activity-core-issue-sink`). +- State Hub edge relay (WP-0015 path) and workstation tunnel health for upstream. +- agent-harness / binky-control only for Binky verification evidence. + +## Tasks + +## Task: Root-cause and fix IssueSink / issue-core bridge 503 + +```task +id: ACTIVITY-WP-0021-T01 +status: todo +priority: high +``` + +**Why first:** Shared failure mode for Binky daily, Binky mail, SBOM weekly, and +any definition whose rules call `emit_tasks` → IssueSink. + +1. Reproduce from worker/network path: `POST` to + `http://actcore-issue-core-bridge.activity-core.svc.cluster.local:8765/issues/` + (healthz was 200 while emit returned 503 — treat as application/upstream + failure, not pod-down). +2. Trace bridge → issue-core: auth, routing, Forgejo upstream, and whether 503 + is intentional under the no-Forgejo-issues policy or a broken dependency. +3. Choose and implement a fix path: + - **A.** Restore a healthy emit path if IssueSink remains canonical, or + - **B.** Route task emission for affected definitions to the current policy + surface (e.g. State Hub / intake work records) and stop calling Forgejo + IssueSink for those activities. +4. Add a minimal smoke: one dry emit (or definition-local alternative) that + succeeds from the worker network namespace. +5. Record non-secret evidence (status codes, chosen path A/B, commit SHAs). + +**Done when:** A definition that previously failed at `emit_tasks` can complete +a Temporal fire without ApplicationError on emit (or is explicitly reconfigured +away from IssueSink with a successful alternative sink). + +## Task: Recover Binky cutover verification (daily + mail) + +```task +id: ACTIVITY-WP-0021-T02 +status: todo +priority: high +``` + +Depends on T01 if Binky still uses IssueSink/task emission. + +1. After emit path is fixed (or Binky definitions use harness-safe sinks), + confirm Temporal schedules for: + - `activity-schedule-ecdfadd3-…` (Binky Daily, `23 8 * * 1-5` Europe/Berlin) + - `activity-schedule-878fbb78-…` (Binky mail, `37 9 * * 1` Europe/Berlin) +2. Accept either the next scheduled business-day fires **or** a controlled + one-shot trigger with the same definition versions as prod. +3. Evidence required: + - Temporal workflow **COMPLETED** + - `activity_runs` row present + - Expected Binky side-effect (brief/commit/harness task) **or** explicit + documented no-op with `tasks_spawned` and context snapshot explaining why +4. Update cutover notes / BINKY verification if needed so Mon–Wed gate is + honest about remaining days. + +**Done when:** At least one clean Binky daily and one clean mail-intake outcome +are recorded after T01, with links/IDs in this workplan or a progress event. + +## Task: Fix Weekly SBOM emit path (or sink policy) + +```task +id: ACTIVITY-WP-0021-T03 +status: todo +priority: high +``` + +Mon 2026-07-20 07:00 UTC SBOM fire **FAILED** after resolver work when +`flag-stale-sbom` hit issue-core bridge 503. Last successful task-spawning SBOM +run was 2026-07-06 (71 tasks). + +1. Reuse T01 emit/sink decision. +2. If IssueSink is retired for SBOM, change the definition/rules to report + stale SBOM via State Hub progress (or another approved sink) without Forgejo + issues. +3. Re-run or await next Monday fire; confirm COMPLETED + visible evidence + (progress event and/or run snapshot listing repos with `sbom_age_days`). + +**Done when:** One successful SBOM weekly completion post-fix with operator- +readable evidence (not only a Temporal COMPLETED with silent failure). + +## Task: Stabilize State Hub edge relay + RecentlyOnScope template + +```task +id: ACTIVITY-WP-0021-T04 +status: todo +priority: high +``` + +1. **Template gap:** locate or add + `/app/templates/recently-on-scope/domain-digest.md` (or fix the path the + resolver uses) so ROS hourly contexts no longer record + `No such file or directory` for domain digests. Ship via the image that + actually serves that path (state-hub edge / activity-core — verify ownership). +2. **Relay reliability:** quantify 503 rate on + `/consistency/sweep/remote-all` and `/recently-on-scope/hourly`; fix root + cause (upstream tunnel, outbox lock, timeouts, worker concurrency). Prefer + durable queue/retry over silent workflow failure. +3. Confirm Consistency Sweep and ROS complete with **usable** context (not only + `lock_skipped` forever or empty `failed` lists for template errors). + +**Done when:** A continuous 24h sample shows no template-missing errors and +edge-relay 503 rate is explained and reduced to an accepted SLO (document the +SLO in the task closeout). + +## Task: Restore daily triage evidence visibility + +```task +id: ACTIVITY-WP-0021-T05 +status: todo +priority: medium +``` + +Sun+Mon 2026-07-19/20 triage runs **completed** (`run_id` present, +`tasks_spawned=0`) with digests in `context_snapshot`, but: + +- State Hub has **no** `event_type=daily_triage` since 2026-07-03 +- working memory last triage note is 2026-07-08 + +1. Trace report/sink instructions for Daily State Hub WSJF Triage in prod + (ConfigMap/external defs vs repo files). +2. Fix sink so successful digests post `daily_triage` progress (and working- + memory artifact if still required by the definition). +3. Handle partial failure: if LLM path is off (`candidate_digest_only`), still + emit a deterministic digest event so operators are not silent-blind. +4. Verify with next morning fire or one-shot trigger. + +**Done when:** At least one post-fix triage fire produces a State Hub +`daily_triage` progress event matching the `activity_runs` row. + +## Task: Harden sync_schedules against ScheduleAlreadyRunningError + +```task +id: ACTIVITY-WP-0021-T06 +status: todo +priority: medium +``` + +Tracks TD **f29e49eb** (registered 2026-07-20): reconcile aborts mid-loop when +one schedule is already running, so later definitions never pause/update. + +1. Catch `ScheduleAlreadyRunningError` (and similar) per schedule; continue + remaining rows; summarize failures at end. +2. Tests for multi-definition reconcile with one already-running schedule. +3. Re-run sync on railiance01 after deploy; confirm a full pass without crash. + +**Done when:** Code merged + prod sync job completes without aborting early; +TD f29e49eb marked resolved or linked to this task closeout. + +## Task: Diagnose app-db and temporal-db restart thrash + +```task +id: ACTIVITY-WP-0021-T07 +status: todo +priority: medium +``` + +Observed on 2026-07-21: `actcore-app-db` ~384 restarts, `actcore-temporal-db` +~500 restarts over long age — pods still Ready but thrash is a durability risk. + +1. Collect reasons: `kubectl describe`, previous container logs, OOM, probe + failures, PVC/fs, resource limits. +2. Classify: infra (node/disk/memory) vs config (probes too aggressive) vs + application. +3. Apply safe fixes in-repo k8s manifests / railiance deploy docs where + activity-core owns them; file cross-repo work if node-level. +4. Record baseline restart rate before/after (24h). + +**Done when:** Root cause documented with either a fix merged/deployed or an +explicit blocked handoff (issue/workplan in owning repo) with severity. + +## Task: Production evidence path for automation-status + +```task +id: ACTIVITY-WP-0021-T08 +status: todo +priority: low +``` + +Workstation `make automation-status SINCE=sunday` reported all key activities +`unknown` because prod DB/Temporal are not on the laptop. + +1. Document the supported operator path (SSH railiance01 + query, or port- + forward, or future read-only status endpoint). +2. Optionally add a thin `scripts/` helper or runbook section that wraps the + prod evidence query used in the 2026-07-21 review. +3. Do **not** put secrets in git; use existing host access patterns. + +**Done when:** `docs/runbook.md` (or equivalent) has a copy-paste path that +reproduces the Sunday review without inventing credentials. + +## Task: Acceptance window — clean status after fixes + +```task +id: ACTIVITY-WP-0021-T09 +status: todo +priority: medium +``` + +After T01–T06 (and T07 if unblocked): + +1. Run a bounded window status (prod evidence path from T08). +2. Confirm for a 24–48h window: + - No FAILED Temporal fires for Binky daily (business day), SBOM if due, + triage, legacy-meter 8h/weekly + - Edge ROS/consistency not dominated by template/503 failures + - At least one fresh `daily_triage` progress event if triage fired +3. Paste summary into this workplan closeout and log a State Hub progress + event on this workplan. + +**Done when:** Operator can answer "how did automations go?" with completed/ +failed verdicts backed by prod data, not all-`unknown`. + +## Success criteria + +- [ ] Emit/sink path fixed or intentionally replaced; Binky + SBOM no longer + fail solely on issue-core 503 +- [ ] At least one clean Binky daily + mail-intake post-fix +- [ ] ROS template present; edge-relay failure rate reduced and documented +- [ ] Daily triage visible again in State Hub progress +- [ ] `sync_schedules` continues past already-running schedules +- [ ] DB restart thrash root-caused with fix or handoff +- [ ] Runbook/status path for prod evidence documented +- [ ] Acceptance window green enough to close this workplan + +## Evidence anchor (2026-07-21 review) + +- Window: 2026-07-19T00:00+02:00 → ~2026-07-21T01:30+02:00 +- Prod: `activity-core` ns on railiance01; API health ok; 18 cron defs +- 259 `activity_runs`; 0 `task_instances` in window +- Temporal: Binky daily/mail, SBOM Mon, TODO Mon **status=FAILED**; triage, + legacy weekly, core-hub daily **COMPLETED** +- Worker: `emit_tasks` → issue-core bridge 503; ROS/consistency edge-relay + timeouts/503; ROS missing `domain-digest.md` +- Related TD: f29e49eb (`sync_schedules` ScheduleAlreadyRunningError) +- Related active work: [[ACTIVITY-WP-0020]] (Forgejo prune still disabled)