kaizen-agentic/engagements/pilots/eng-coulomb-railiance01-ho-001/reports/2026-07-16-health-review.md
tegwick 6ca167ce19
Some checks failed
CI Smoke / host-smoke (push) Successful in 1s
CI Smoke / container-smoke (push) Successful in 36s
ci / test (push) Failing after 3m41s
ops: complete host-operator ramp-up on railiance01 (WP-0009 T10)
Live observe session verified SSH access, captured baseline and Critical
health/load findings (memory pressure, k3s API unavailable). RU checklist
closed; engagement phase operating; schedule enabled; no privileged changes.
2026-07-16 12:43:00 +02:00

4.6 KiB
Raw Permalink Blame History

Health & load review — railiance01 — 2026-07-16

Engagement: eng-coulomb-railiance01-ho-001 Phase: ramp_up (first live observe) Access class: host_observe Protocols: load-workload-review + sys-medic subset


1. Executive Summary

railiance01 is a 2-core / 3.8 GiB Ubuntu 24.04 single-node k3s host under severe memory pressure and sustained CPU/load overload. Root disk capacity is acceptable (~54%). The k3s API was ServiceUnavailable during assessment (control plane struggling). No privileged remediation was performed.

Overall health: Critical (confidence: High for host-level memory/load; Medium for full pod inventory due to API unavailability)

2. Health Status

Dimension Status
Aggregate Critical
Memory Critical (no swap, PSI elevated, ~hundreds of MiB available)
CPU / load Degraded–Critical (load ≫ 2 cores)
Disk capacity Healthy (~54% root)
Disk / logs Watch (/var/log ~5.7G, journal ~4.1G)
k3s API Critical / unavailable at sample time
Network listeners Watch (expected k3s + SSH; review exposure of 6443)

3. Findings

F1 — Memory exhaustion / no swap

  • Severity: Critical
  • Evidence: MemTotal ~3.8Gi; MemAvailable often <500Mi; SwapTotal 0; PSI memory full avg60≈24% avg300≈29%; kswapd0 long-running
  • Why it matters: OOM risk; thrashing; control-plane instability
  • Likely cause: Workload density (k3s + Forgejo + Temporal + activity-core + state-hub edge + platform pods) on undersized RAM
  • Next step: Human-approved capacity plan (add RAM and/or swap emergency cushion; reduce concurrent workloads). Do not kill workloads without owner approval

F2 — Load far above CPU capacity

  • Severity: High
  • Evidence: load average ~12/11/16 with nproc=2; vmstat shows high sy and wa, heavy block-in
  • Why it matters: Latency for all services; scheduler saturation
  • Likely cause: Memory reclaim + I/O + k3s restart recovery
  • Next step: Correlate with k3s restart time (~10:39 UTC); re-sample after memory relief

F3 — k3s API ServiceUnavailable

  • Severity: Critical
  • Evidence: sudo k3s kubectl get node → server unable to handle request; k3s server process high CPU shortly after start
  • Why it matters: Cluster ops and many apps depend on API
  • Likely cause: Memory pressure / control-plane restart under load
  • Next step: Re-check API when MemAvailable improves; avoid concurrent heavy kubectl

F4 — Large journald footprint

  • Severity: Medium
  • Evidence: /var/log/journal ~4.1G of ~5.7G under /var/log
  • Why it matters: Disk growth; I/O; eventual fill
  • Next step: Propose journalctl --vacuum-size= with approval (privileged / system change)

F5 — Pending OS package updates (incl. security)

  • Severity: Medium
  • Evidence: apt list --upgradable shows many packages (bind9-*, curl, ca-certificates, dpkg, …)
  • needrestart: kernel current (KSTA 1); unattended-upgrades.service flagged
  • Next step: Scheduled OS security pass with human approval for upgrades

F6 — Reverse tunnel port contention (18765)

  • Severity: Low–Medium
  • Evidence: repeated sshd errors binding 127.0.0.1:18765 address already in use
  • Why it matters: ops-bridge / reverse-forward fragility
  • Next step: Operator cleanup of stale forwards; bridge config audit

F7 — Public exposure of k3s API port in UFW

  • Severity: Medium (policy)
  • Evidence: UFW allows 6443/tcp and 8472/udp from Anywhere
  • Next step: Confirm intended threat model; prefer restricted sources if possible (firewall_change gated)

4. Immediate Safe Actions (no approval needed)

  1. Continue observe-only monitoring; re-sample load/memory in 15–30 minutes
  2. Prefer not running heavy builds / additional agents on this host until memory recovers
  3. Document findings in engagement vault (this report + memory sections)
  4. Escalate capacity concern to human operator

5. Escalation

  • Human operator (Bernd / railiance ops): capacity (RAM/swap), whether k3s restart was intentional, approval for journal vacuum and package upgrades
  • App owners: Forgejo, activity-core, Temporal — if service degradation reported

6. Suggested inspect commands

ssh railiance01 'free -h; uptime; cat /proc/pressure/memory'
ssh railiance01 'sudo k3s kubectl get node; sudo k3s kubectl get pods -A | head'
ssh railiance01 'sudo du -sh /var/log/journal'

7. Privileged proposals (NOT executed — see RU-08 dry-run)

See vault/session-log/2026-07-16-privileged-proposal-dry-run.md.