kaizen-agentic/engagements/pilots/eng-coulomb-railiance01-ho-001/reports/2026-07-16-privileged-remediation.md
tegwick 48443a9abb
Some checks failed
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 4s
ci / test (push) Failing after 5s
ops: apply approved host-operator remediations on railiance01
Record operator approval and implement P1–P4: 4G swapfile, journald vacuum,
apt upgrade, and UFW lockdown of k3s API/flannel world-open rules. k3s node
Ready post-change; residual risk is still tight RAM.
2026-07-16 13:05:25 +02:00

2 KiB
Raw Blame History

Privileged remediation report — railiance01 — 2026-07-16

Engagement: eng-coulomb-railiance01-ho-001 Approval: vault/session-log/2026-07-16-privileged-approval.md Outcome: success (with residual capacity risk)

Summary

ID Action Result
P1 4 GiB /swapfile, fstab persist Done — swap active (~1.9 GiB used post-change)
P2 journald vacuum → 500 M Done — freed ~3.5 GiB archived journals; journal now ~461 M
P3 apt-get upgrade (noninteractive) Done — packages updated; kernel still current (no reboot required)
P4 UFW: drop world 6443/8472; allow 6443 from operator IPs Done

Before → after (host signals)

Metric Before (ramp-up) After remediation
Swap none 4 GiB file, ~1.9 GiB used
MemAvailable often <150 Mi ~625 Mi
PSI memory full avg60 ~24% ~8%
Load 1m ~12 ~6
/var/log/journal ~4.1 Gi ~461 Mi
k3s node API often unavailable Ready control-plane
UFW 6443 Anywhere 89.244.90.246, 85.132.220.102 only
UFW 8472 Anywhere removed (single-node)

Side effects / notes

  • fwupd.service failed to restart during upgrade (non-blocking for cluster)
  • needrestart: deferred restarts for cloud-final, dbus, getty, logind, unattended-upgrades; user sessions still on older sshd binaries until re-login
  • RAM is still 3.8 GiB — swap masks OOM risk but is not a substitute for more memory under sustained pressure
  • If remote kubectl from another IP fails, add: ufw allow from <ip> to any port 6443 proto tcp
  • UFW backups: /etc/ufw/user.rules.bak.20260716 (and user6)

Explicitly not done

  • Hardware RAM upgrade (provider console)
  • Host reboot (kernel KSTA=1, not required)
  1. Plan RAM upgrade when convenient
  2. Re-login SSH sessions to pick up new binaries
  3. Watch swap usage; if chronically full, capacity still insufficient
  4. Maintain operator IP allowlist when admin IPs change