# Health & load review — railiance01 — 2026-07-16 **Engagement:** eng-coulomb-railiance01-ho-001 **Phase:** ramp_up (first live observe) **Access class:** host_observe **Protocols:** load-workload-review + sys-medic subset --- ## 1. Executive Summary railiance01 is a **2-core / 3.8 GiB** Ubuntu 24.04 single-node k3s host under **severe memory pressure and sustained CPU/load overload**. Root disk capacity is acceptable (~54%). The k3s API was **ServiceUnavailable** during assessment (control plane struggling). No privileged remediation was performed. **Overall health: Critical** (confidence: High for host-level memory/load; Medium for full pod inventory due to API unavailability) ## 2. Health Status | Dimension | Status | |-----------|--------| | Aggregate | **Critical** | | Memory | Critical (no swap, PSI elevated, ~hundreds of MiB available) | | CPU / load | Degraded–Critical (load ≫ 2 cores) | | Disk capacity | Healthy (~54% root) | | Disk / logs | Watch (`/var/log` ~5.7G, journal ~4.1G) | | k3s API | Critical / unavailable at sample time | | Network listeners | Watch (expected k3s + SSH; review exposure of 6443) | ## 3. Findings ### F1 — Memory exhaustion / no swap - **Severity:** Critical - **Evidence:** MemTotal ~3.8Gi; MemAvailable often <500Mi; SwapTotal 0; PSI memory `full avg60≈24% avg300≈29%`; `kswapd0` long-running - **Why it matters:** OOM risk; thrashing; control-plane instability - **Likely cause:** Workload density (k3s + Forgejo + Temporal + activity-core + state-hub edge + platform pods) on undersized RAM - **Next step:** Human-approved capacity plan (add RAM and/or swap emergency cushion; reduce concurrent workloads). **Do not** kill workloads without owner approval ### F2 — Load far above CPU capacity - **Severity:** High - **Evidence:** load average ~12/11/16 with `nproc=2`; vmstat shows high `sy` and `wa`, heavy block-in - **Why it matters:** Latency for all services; scheduler saturation - **Likely cause:** Memory reclaim + I/O + k3s restart recovery - **Next step:** Correlate with k3s restart time (~10:39 UTC); re-sample after memory relief ### F3 — k3s API ServiceUnavailable - **Severity:** Critical - **Evidence:** `sudo k3s kubectl get node` → server unable to handle request; `k3s server` process high CPU shortly after start - **Why it matters:** Cluster ops and many apps depend on API - **Likely cause:** Memory pressure / control-plane restart under load - **Next step:** Re-check API when MemAvailable improves; avoid concurrent heavy kubectl ### F4 — Large journald footprint - **Severity:** Medium - **Evidence:** `/var/log/journal` ~4.1G of ~5.7G under `/var/log` - **Why it matters:** Disk growth; I/O; eventual fill - **Next step:** Propose `journalctl --vacuum-size=` **with approval** (privileged / system change) ### F5 — Pending OS package updates (incl. security) - **Severity:** Medium - **Evidence:** `apt list --upgradable` shows many packages (bind9-*, curl, ca-certificates, dpkg, …) - **needrestart:** kernel current (KSTA 1); `unattended-upgrades.service` flagged - **Next step:** Scheduled OS security pass with human approval for upgrades ### F6 — Reverse tunnel port contention (18765) - **Severity:** Low–Medium - **Evidence:** repeated sshd errors binding 127.0.0.1:18765 address already in use - **Why it matters:** ops-bridge / reverse-forward fragility - **Next step:** Operator cleanup of stale forwards; bridge config audit ### F7 — Public exposure of k3s API port in UFW - **Severity:** Medium (policy) - **Evidence:** UFW allows 6443/tcp and 8472/udp from Anywhere - **Next step:** Confirm intended threat model; prefer restricted sources if possible (**firewall_change** gated) ## 4. Immediate Safe Actions (no approval needed) 1. Continue observe-only monitoring; re-sample load/memory in 15–30 minutes 2. Prefer not running heavy builds / additional agents on this host until memory recovers 3. Document findings in engagement vault (this report + memory sections) 4. Escalate capacity concern to human operator ## 5. Escalation - **Human operator (Bernd / railiance ops):** capacity (RAM/swap), whether k3s restart was intentional, approval for journal vacuum and package upgrades - **App owners:** Forgejo, activity-core, Temporal — if service degradation reported ## 6. Suggested inspect commands ```bash ssh railiance01 'free -h; uptime; cat /proc/pressure/memory' ssh railiance01 'sudo k3s kubectl get node; sudo k3s kubectl get pods -A | head' ssh railiance01 'sudo du -sh /var/log/journal' ``` ## 7. Privileged proposals (NOT executed — see RU-08 dry-run) See `vault/session-log/2026-07-16-privileged-proposal-dry-run.md`.