ops: complete host-operator ramp-up on railiance01 (WP-0009 T10)
Some checks failed
CI Smoke / host-smoke (push) Successful in 1s
CI Smoke / container-smoke (push) Successful in 36s
ci / test (push) Failing after 3m41s

Live observe session verified SSH access, captured baseline and Critical
health/load findings (memory pressure, k3s API unavailable). RU checklist
closed; engagement phase operating; schedule enabled; no privileged changes.
This commit is contained in:
tegwick 2026-07-16 12:43:00 +02:00
parent 7257e62dba
commit 6ca167ce19
15 changed files with 334 additions and 90 deletions

View file

@ -0,0 +1,13 @@
# Session report — eng-coulomb-railiance01-ho-001
- **Date:** 2026-07-16
- **Duty:** deep_assessment
- **Targets:** railiance01
- **Outcome:** success
- **Phase:** operating
## Summary
T10 ramp-up complete: Critical memory/load, k3s API unavailable, RU all done, phase operating
_Billing metadata only in commercial/ledger.jsonl; no secrets in this report._

View file

@ -0,0 +1,96 @@
# Health & load review — railiance01 — 2026-07-16
**Engagement:** eng-coulomb-railiance01-ho-001
**Phase:** ramp_up (first live observe)
**Access class:** host_observe
**Protocols:** load-workload-review + sys-medic subset
---
## 1. Executive Summary
railiance01 is a **2-core / 3.8 GiB** Ubuntu 24.04 single-node k3s host under **severe memory pressure and sustained CPU/load overload**. Root disk capacity is acceptable (~54%). The k3s API was **ServiceUnavailable** during assessment (control plane struggling). No privileged remediation was performed.
**Overall health: Critical** (confidence: High for host-level memory/load; Medium for full pod inventory due to API unavailability)
## 2. Health Status
| Dimension | Status |
|-----------|--------|
| Aggregate | **Critical** |
| Memory | Critical (no swap, PSI elevated, ~hundreds of MiB available) |
| CPU / load | Degraded–Critical (load ≫ 2 cores) |
| Disk capacity | Healthy (~54% root) |
| Disk / logs | Watch (`/var/log` ~5.7G, journal ~4.1G) |
| k3s API | Critical / unavailable at sample time |
| Network listeners | Watch (expected k3s + SSH; review exposure of 6443) |
## 3. Findings
### F1 — Memory exhaustion / no swap
- **Severity:** Critical
- **Evidence:** MemTotal ~3.8Gi; MemAvailable often <500Mi; SwapTotal 0; PSI memory `full avg60≈24% avg300≈29%`; `kswapd0` long-running
- **Why it matters:** OOM risk; thrashing; control-plane instability
- **Likely cause:** Workload density (k3s + Forgejo + Temporal + activity-core + state-hub edge + platform pods) on undersized RAM
- **Next step:** Human-approved capacity plan (add RAM and/or swap emergency cushion; reduce concurrent workloads). **Do not** kill workloads without owner approval
### F2 — Load far above CPU capacity
- **Severity:** High
- **Evidence:** load average ~12/11/16 with `nproc=2`; vmstat shows high `sy` and `wa`, heavy block-in
- **Why it matters:** Latency for all services; scheduler saturation
- **Likely cause:** Memory reclaim + I/O + k3s restart recovery
- **Next step:** Correlate with k3s restart time (~10:39 UTC); re-sample after memory relief
### F3 — k3s API ServiceUnavailable
- **Severity:** Critical
- **Evidence:** `sudo k3s kubectl get node` → server unable to handle request; `k3s server` process high CPU shortly after start
- **Why it matters:** Cluster ops and many apps depend on API
- **Likely cause:** Memory pressure / control-plane restart under load
- **Next step:** Re-check API when MemAvailable improves; avoid concurrent heavy kubectl
### F4 — Large journald footprint
- **Severity:** Medium
- **Evidence:** `/var/log/journal` ~4.1G of ~5.7G under `/var/log`
- **Why it matters:** Disk growth; I/O; eventual fill
- **Next step:** Propose `journalctl --vacuum-size=` **with approval** (privileged / system change)
### F5 — Pending OS package updates (incl. security)
- **Severity:** Medium
- **Evidence:** `apt list --upgradable` shows many packages (bind9-*, curl, ca-certificates, dpkg, …)
- **needrestart:** kernel current (KSTA 1); `unattended-upgrades.service` flagged
- **Next step:** Scheduled OS security pass with human approval for upgrades
### F6 — Reverse tunnel port contention (18765)
- **Severity:** Low–Medium
- **Evidence:** repeated sshd errors binding 127.0.0.1:18765 address already in use
- **Why it matters:** ops-bridge / reverse-forward fragility
- **Next step:** Operator cleanup of stale forwards; bridge config audit
### F7 — Public exposure of k3s API port in UFW
- **Severity:** Medium (policy)
- **Evidence:** UFW allows 6443/tcp and 8472/udp from Anywhere
- **Next step:** Confirm intended threat model; prefer restricted sources if possible (**firewall_change** gated)
## 4. Immediate Safe Actions (no approval needed)
1. Continue observe-only monitoring; re-sample load/memory in 15–30 minutes
2. Prefer not running heavy builds / additional agents on this host until memory recovers
3. Document findings in engagement vault (this report + memory sections)
4. Escalate capacity concern to human operator
## 5. Escalation
- **Human operator (Bernd / railiance ops):** capacity (RAM/swap), whether k3s restart was intentional, approval for journal vacuum and package upgrades
- **App owners:** Forgejo, activity-core, Temporal — if service degradation reported
## 6. Suggested inspect commands
```bash
ssh railiance01 'free -h; uptime; cat /proc/pressure/memory'
ssh railiance01 'sudo k3s kubectl get node; sudo k3s kubectl get pods -A | head'
ssh railiance01 'sudo du -sh /var/log/journal'
```
## 7. Privileged proposals (NOT executed — see RU-08 dry-run)
See `vault/session-log/2026-07-16-privileged-proposal-dry-run.md`.