ops: complete host-operator ramp-up on railiance01 (WP-0009 T10)
Live observe session verified SSH access, captured baseline and Critical health/load findings (memory pressure, k3s API unavailable). RU checklist closed; engagement phase operating; schedule enabled; no privileged changes.
This commit is contained in:
parent
7257e62dba
commit
6ca167ce19
15 changed files with 334 additions and 90 deletions
|
|
@ -0,0 +1,13 @@
|
|||
# Session report — eng-coulomb-railiance01-ho-001
|
||||
|
||||
- **Date:** 2026-07-16
|
||||
- **Duty:** deep_assessment
|
||||
- **Targets:** railiance01
|
||||
- **Outcome:** success
|
||||
- **Phase:** operating
|
||||
|
||||
## Summary
|
||||
|
||||
T10 ramp-up complete: Critical memory/load, k3s API unavailable, RU all done, phase operating
|
||||
|
||||
_Billing metadata only in commercial/ledger.jsonl; no secrets in this report._
|
||||
|
|
@ -0,0 +1,96 @@
|
|||
# Health & load review — railiance01 — 2026-07-16
|
||||
|
||||
**Engagement:** eng-coulomb-railiance01-ho-001
|
||||
**Phase:** ramp_up (first live observe)
|
||||
**Access class:** host_observe
|
||||
**Protocols:** load-workload-review + sys-medic subset
|
||||
|
||||
---
|
||||
|
||||
## 1. Executive Summary
|
||||
|
||||
railiance01 is a **2-core / 3.8 GiB** Ubuntu 24.04 single-node k3s host under **severe memory pressure and sustained CPU/load overload**. Root disk capacity is acceptable (~54%). The k3s API was **ServiceUnavailable** during assessment (control plane struggling). No privileged remediation was performed.
|
||||
|
||||
**Overall health: Critical** (confidence: High for host-level memory/load; Medium for full pod inventory due to API unavailability)
|
||||
|
||||
## 2. Health Status
|
||||
|
||||
| Dimension | Status |
|
||||
|-----------|--------|
|
||||
| Aggregate | **Critical** |
|
||||
| Memory | Critical (no swap, PSI elevated, ~hundreds of MiB available) |
|
||||
| CPU / load | Degraded–Critical (load ≫ 2 cores) |
|
||||
| Disk capacity | Healthy (~54% root) |
|
||||
| Disk / logs | Watch (`/var/log` ~5.7G, journal ~4.1G) |
|
||||
| k3s API | Critical / unavailable at sample time |
|
||||
| Network listeners | Watch (expected k3s + SSH; review exposure of 6443) |
|
||||
|
||||
## 3. Findings
|
||||
|
||||
### F1 — Memory exhaustion / no swap
|
||||
- **Severity:** Critical
|
||||
- **Evidence:** MemTotal ~3.8Gi; MemAvailable often <500Mi; SwapTotal 0; PSI memory `full avg60≈24% avg300≈29%`; `kswapd0` long-running
|
||||
- **Why it matters:** OOM risk; thrashing; control-plane instability
|
||||
- **Likely cause:** Workload density (k3s + Forgejo + Temporal + activity-core + state-hub edge + platform pods) on undersized RAM
|
||||
- **Next step:** Human-approved capacity plan (add RAM and/or swap emergency cushion; reduce concurrent workloads). **Do not** kill workloads without owner approval
|
||||
|
||||
### F2 — Load far above CPU capacity
|
||||
- **Severity:** High
|
||||
- **Evidence:** load average ~12/11/16 with `nproc=2`; vmstat shows high `sy` and `wa`, heavy block-in
|
||||
- **Why it matters:** Latency for all services; scheduler saturation
|
||||
- **Likely cause:** Memory reclaim + I/O + k3s restart recovery
|
||||
- **Next step:** Correlate with k3s restart time (~10:39 UTC); re-sample after memory relief
|
||||
|
||||
### F3 — k3s API ServiceUnavailable
|
||||
- **Severity:** Critical
|
||||
- **Evidence:** `sudo k3s kubectl get node` → server unable to handle request; `k3s server` process high CPU shortly after start
|
||||
- **Why it matters:** Cluster ops and many apps depend on API
|
||||
- **Likely cause:** Memory pressure / control-plane restart under load
|
||||
- **Next step:** Re-check API when MemAvailable improves; avoid concurrent heavy kubectl
|
||||
|
||||
### F4 — Large journald footprint
|
||||
- **Severity:** Medium
|
||||
- **Evidence:** `/var/log/journal` ~4.1G of ~5.7G under `/var/log`
|
||||
- **Why it matters:** Disk growth; I/O; eventual fill
|
||||
- **Next step:** Propose `journalctl --vacuum-size=` **with approval** (privileged / system change)
|
||||
|
||||
### F5 — Pending OS package updates (incl. security)
|
||||
- **Severity:** Medium
|
||||
- **Evidence:** `apt list --upgradable` shows many packages (bind9-*, curl, ca-certificates, dpkg, …)
|
||||
- **needrestart:** kernel current (KSTA 1); `unattended-upgrades.service` flagged
|
||||
- **Next step:** Scheduled OS security pass with human approval for upgrades
|
||||
|
||||
### F6 — Reverse tunnel port contention (18765)
|
||||
- **Severity:** Low–Medium
|
||||
- **Evidence:** repeated sshd errors binding 127.0.0.1:18765 address already in use
|
||||
- **Why it matters:** ops-bridge / reverse-forward fragility
|
||||
- **Next step:** Operator cleanup of stale forwards; bridge config audit
|
||||
|
||||
### F7 — Public exposure of k3s API port in UFW
|
||||
- **Severity:** Medium (policy)
|
||||
- **Evidence:** UFW allows 6443/tcp and 8472/udp from Anywhere
|
||||
- **Next step:** Confirm intended threat model; prefer restricted sources if possible (**firewall_change** gated)
|
||||
|
||||
## 4. Immediate Safe Actions (no approval needed)
|
||||
|
||||
1. Continue observe-only monitoring; re-sample load/memory in 15–30 minutes
|
||||
2. Prefer not running heavy builds / additional agents on this host until memory recovers
|
||||
3. Document findings in engagement vault (this report + memory sections)
|
||||
4. Escalate capacity concern to human operator
|
||||
|
||||
## 5. Escalation
|
||||
|
||||
- **Human operator (Bernd / railiance ops):** capacity (RAM/swap), whether k3s restart was intentional, approval for journal vacuum and package upgrades
|
||||
- **App owners:** Forgejo, activity-core, Temporal — if service degradation reported
|
||||
|
||||
## 6. Suggested inspect commands
|
||||
|
||||
```bash
|
||||
ssh railiance01 'free -h; uptime; cat /proc/pressure/memory'
|
||||
ssh railiance01 'sudo k3s kubectl get node; sudo k3s kubectl get pods -A | head'
|
||||
ssh railiance01 'sudo du -sh /var/log/journal'
|
||||
```
|
||||
|
||||
## 7. Privileged proposals (NOT executed — see RU-08 dry-run)
|
||||
|
||||
See `vault/session-log/2026-07-16-privileged-proposal-dry-run.md`.
|
||||
Loading…
Add table
Add a link
Reference in a new issue