Live observe session verified SSH access, captured baseline and Critical health/load findings (memory pressure, k3s API unavailable). RU checklist closed; engagement phase operating; schedule enabled; no privileged changes.
4.6 KiB
4.6 KiB
Health & load review — railiance01 — 2026-07-16
Engagement: eng-coulomb-railiance01-ho-001 Phase: ramp_up (first live observe) Access class: host_observe Protocols: load-workload-review + sys-medic subset
1. Executive Summary
railiance01 is a 2-core / 3.8 GiB Ubuntu 24.04 single-node k3s host under severe memory pressure and sustained CPU/load overload. Root disk capacity is acceptable (~54%). The k3s API was ServiceUnavailable during assessment (control plane struggling). No privileged remediation was performed.
Overall health: Critical (confidence: High for host-level memory/load; Medium for full pod inventory due to API unavailability)
2. Health Status
| Dimension | Status |
|---|---|
| Aggregate | Critical |
| Memory | Critical (no swap, PSI elevated, ~hundreds of MiB available) |
| CPU / load | Degraded–Critical (load ≫ 2 cores) |
| Disk capacity | Healthy (~54% root) |
| Disk / logs | Watch (/var/log ~5.7G, journal ~4.1G) |
| k3s API | Critical / unavailable at sample time |
| Network listeners | Watch (expected k3s + SSH; review exposure of 6443) |
3. Findings
F1 — Memory exhaustion / no swap
- Severity: Critical
- Evidence: MemTotal ~3.8Gi; MemAvailable often <500Mi; SwapTotal 0; PSI memory
full avg60≈24% avg300≈29%;kswapd0long-running - Why it matters: OOM risk; thrashing; control-plane instability
- Likely cause: Workload density (k3s + Forgejo + Temporal + activity-core + state-hub edge + platform pods) on undersized RAM
- Next step: Human-approved capacity plan (add RAM and/or swap emergency cushion; reduce concurrent workloads). Do not kill workloads without owner approval
F2 — Load far above CPU capacity
- Severity: High
- Evidence: load average ~12/11/16 with
nproc=2; vmstat shows highsyandwa, heavy block-in - Why it matters: Latency for all services; scheduler saturation
- Likely cause: Memory reclaim + I/O + k3s restart recovery
- Next step: Correlate with k3s restart time (~10:39 UTC); re-sample after memory relief
F3 — k3s API ServiceUnavailable
- Severity: Critical
- Evidence:
sudo k3s kubectl get node→ server unable to handle request;k3s serverprocess high CPU shortly after start - Why it matters: Cluster ops and many apps depend on API
- Likely cause: Memory pressure / control-plane restart under load
- Next step: Re-check API when MemAvailable improves; avoid concurrent heavy kubectl
F4 — Large journald footprint
- Severity: Medium
- Evidence:
/var/log/journal~4.1G of ~5.7G under/var/log - Why it matters: Disk growth; I/O; eventual fill
- Next step: Propose
journalctl --vacuum-size=with approval (privileged / system change)
F5 — Pending OS package updates (incl. security)
- Severity: Medium
- Evidence:
apt list --upgradableshows many packages (bind9-*, curl, ca-certificates, dpkg, …) - needrestart: kernel current (KSTA 1);
unattended-upgrades.serviceflagged - Next step: Scheduled OS security pass with human approval for upgrades
F6 — Reverse tunnel port contention (18765)
- Severity: Low–Medium
- Evidence: repeated sshd errors binding 127.0.0.1:18765 address already in use
- Why it matters: ops-bridge / reverse-forward fragility
- Next step: Operator cleanup of stale forwards; bridge config audit
F7 — Public exposure of k3s API port in UFW
- Severity: Medium (policy)
- Evidence: UFW allows 6443/tcp and 8472/udp from Anywhere
- Next step: Confirm intended threat model; prefer restricted sources if possible (firewall_change gated)
4. Immediate Safe Actions (no approval needed)
- Continue observe-only monitoring; re-sample load/memory in 15–30 minutes
- Prefer not running heavy builds / additional agents on this host until memory recovers
- Document findings in engagement vault (this report + memory sections)
- Escalate capacity concern to human operator
5. Escalation
- Human operator (Bernd / railiance ops): capacity (RAM/swap), whether k3s restart was intentional, approval for journal vacuum and package upgrades
- App owners: Forgejo, activity-core, Temporal — if service degradation reported
6. Suggested inspect commands
ssh railiance01 'free -h; uptime; cat /proc/pressure/memory'
ssh railiance01 'sudo k3s kubectl get node; sudo k3s kubectl get pods -A | head'
ssh railiance01 'sudo du -sh /var/log/journal'
7. Privileged proposals (NOT executed — see RU-08 dry-run)
See vault/session-log/2026-07-16-privileged-proposal-dry-run.md.