kaizen-agentic/engagements/pilots/eng-coulomb-railiance01-ho-001/reports/2026-07-16-health-review.md
tegwick 6ca167ce19
Some checks failed
CI Smoke / host-smoke (push) Successful in 1s
CI Smoke / container-smoke (push) Successful in 36s
ci / test (push) Failing after 3m41s
ops: complete host-operator ramp-up on railiance01 (WP-0009 T10)
Live observe session verified SSH access, captured baseline and Critical
health/load findings (memory pressure, k3s API unavailable). RU checklist
closed; engagement phase operating; schedule enabled; no privileged changes.
2026-07-16 12:43:00 +02:00

96 lines
4.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Health & load review — railiance01 — 2026-07-16
**Engagement:** eng-coulomb-railiance01-ho-001
**Phase:** ramp_up (first live observe)
**Access class:** host_observe
**Protocols:** load-workload-review + sys-medic subset
---
## 1. Executive Summary
railiance01 is a **2-core / 3.8 GiB** Ubuntu 24.04 single-node k3s host under **severe memory pressure and sustained CPU/load overload**. Root disk capacity is acceptable (~54%). The k3s API was **ServiceUnavailable** during assessment (control plane struggling). No privileged remediation was performed.
**Overall health: Critical** (confidence: High for host-level memory/load; Medium for full pod inventory due to API unavailability)
## 2. Health Status
| Dimension | Status |
|-----------|--------|
| Aggregate | **Critical** |
| Memory | Critical (no swap, PSI elevated, ~hundreds of MiB available) |
| CPU / load | Degraded–Critical (load ≫ 2 cores) |
| Disk capacity | Healthy (~54% root) |
| Disk / logs | Watch (`/var/log` ~5.7G, journal ~4.1G) |
| k3s API | Critical / unavailable at sample time |
| Network listeners | Watch (expected k3s + SSH; review exposure of 6443) |
## 3. Findings
### F1 — Memory exhaustion / no swap
- **Severity:** Critical
- **Evidence:** MemTotal ~3.8Gi; MemAvailable often <500Mi; SwapTotal 0; PSI memory `full avg60≈24% avg300≈29%`; `kswapd0` long-running
- **Why it matters:** OOM risk; thrashing; control-plane instability
- **Likely cause:** Workload density (k3s + Forgejo + Temporal + activity-core + state-hub edge + platform pods) on undersized RAM
- **Next step:** Human-approved capacity plan (add RAM and/or swap emergency cushion; reduce concurrent workloads). **Do not** kill workloads without owner approval
### F2 — Load far above CPU capacity
- **Severity:** High
- **Evidence:** load average ~12/11/16 with `nproc=2`; vmstat shows high `sy` and `wa`, heavy block-in
- **Why it matters:** Latency for all services; scheduler saturation
- **Likely cause:** Memory reclaim + I/O + k3s restart recovery
- **Next step:** Correlate with k3s restart time (~10:39 UTC); re-sample after memory relief
### F3 — k3s API ServiceUnavailable
- **Severity:** Critical
- **Evidence:** `sudo k3s kubectl get node` → server unable to handle request; `k3s server` process high CPU shortly after start
- **Why it matters:** Cluster ops and many apps depend on API
- **Likely cause:** Memory pressure / control-plane restart under load
- **Next step:** Re-check API when MemAvailable improves; avoid concurrent heavy kubectl
### F4 — Large journald footprint
- **Severity:** Medium
- **Evidence:** `/var/log/journal` ~4.1G of ~5.7G under `/var/log`
- **Why it matters:** Disk growth; I/O; eventual fill
- **Next step:** Propose `journalctl --vacuum-size=` **with approval** (privileged / system change)
### F5 — Pending OS package updates (incl. security)
- **Severity:** Medium
- **Evidence:** `apt list --upgradable` shows many packages (bind9-*, curl, ca-certificates, dpkg, …)
- **needrestart:** kernel current (KSTA 1); `unattended-upgrades.service` flagged
- **Next step:** Scheduled OS security pass with human approval for upgrades
### F6 — Reverse tunnel port contention (18765)
- **Severity:** Low–Medium
- **Evidence:** repeated sshd errors binding 127.0.0.1:18765 address already in use
- **Why it matters:** ops-bridge / reverse-forward fragility
- **Next step:** Operator cleanup of stale forwards; bridge config audit
### F7 — Public exposure of k3s API port in UFW
- **Severity:** Medium (policy)
- **Evidence:** UFW allows 6443/tcp and 8472/udp from Anywhere
- **Next step:** Confirm intended threat model; prefer restricted sources if possible (**firewall_change** gated)
## 4. Immediate Safe Actions (no approval needed)
1. Continue observe-only monitoring; re-sample load/memory in 15–30 minutes
2. Prefer not running heavy builds / additional agents on this host until memory recovers
3. Document findings in engagement vault (this report + memory sections)
4. Escalate capacity concern to human operator
## 5. Escalation
- **Human operator (Bernd / railiance ops):** capacity (RAM/swap), whether k3s restart was intentional, approval for journal vacuum and package upgrades
- **App owners:** Forgejo, activity-core, Temporal — if service degradation reported
## 6. Suggested inspect commands
```bash
ssh railiance01 'free -h; uptime; cat /proc/pressure/memory'
ssh railiance01 'sudo k3s kubectl get node; sudo k3s kubectl get pods -A | head'
ssh railiance01 'sudo du -sh /var/log/journal'
```
## 7. Privileged proposals (NOT executed — see RU-08 dry-run)
See `vault/session-log/2026-07-16-privileged-proposal-dry-run.md`.