kaizen-agentic/engagements/pilots/eng-coulomb-railiance01-ho-001/reports/2026-07-16-capacity-upgrade-check.md
tegwick 776b7f8cc4
Some checks failed
CI Smoke / container-smoke (push) Successful in 2s
ci / test (push) Failing after 5s
CI Smoke / host-smoke (push) Successful in 0s
ops: note railiance01 capacity upgrade claim vs live guest size
Operator reported 4 vCPU / 16 GiB; verify shows guest still 2 / 3.8 GiB
after recent reboot. Document check report and keep open thread until
nproc/free match the new flavor.
2026-07-16 14:33:48 +02:00

50 lines
2 KiB
Markdown
Raw Blame History

This file contains invisible Unicode characters

This file contains invisible Unicode characters that are indistinguishable to humans but may be processed differently by a computer. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Capacity upgrade check — railiance01 — 2026-07-16
**Engagement:** eng-coulomb-railiance01-ho-001
**Operator report:** VM upgraded to **4 vCPU / 16 GiB RAM**
**Agent verify:** observe via `ssh railiance01`
## Result: not yet visible in the guest OS
| Resource | Operator target | Live guest (`/proc`, `lscpu`) |
|----------|-----------------|------------------------------|
| vCPU | 4 | **2** (`nproc` / `lscpu`) |
| RAM | 16 GiB | **~3.8 GiB** (`MemTotal` ≈ 4009884 kB) |
| Swap | (our /swapfile) | 4 GiB file, ~300 Mi used |
| Uptime at check | — | ~12 minutes (recent reboot) |
| Hypervisor | — | KVM / OpenStack Nova (ConfigDrive) |
| k3s node | — | **Ready** |
## Guest health snapshot (at check)
- Load ~2 on 2 cores (much better than Critical era load ~12)
- MemAvailable ~0.9 GiB; PSI memory full avg60 ~0.2% (comfortable **for current size**)
- Swap lightly used (~300 Mi)
So the **workload is calmer after reboot**, but the **instance size has not expanded** to 4/16 from the kernel’s point of view.
## Likely causes
1. Resize still pending or applied to a different VM/instance
2. Provider resize requires a further console action / hard reboot after flavor change
3. Resize failed partially (boot on old flavor)
4. Wrong host checked (we used SSH host alias `railiance01` → 92.205.62.239)
## Recommended operator steps (provider / OpenStack)
1. Confirm in HostEurope / OpenStack console that **this** instance (IP 92.205.62.239) shows flavor **4 vCPU / 16 GiB**
2. If resize is `VERIFY_RESIZE` / confirm-pending, **confirm** it
3. Hard reboot if the panel says so after resize
4. Re-check from shell:
```bash
ssh railiance01 'nproc; free -h; lscpu | grep -E "CPU\\(s\\)|Model"'
# expect: CPU(s): 4 and Mem ~15–16Gi
```
5. Optional later: reduce reliance on 4 Gi swapfile once 16 Gi is confirmed stable (keep small swap as cushion, or leave as-is)
## Agent action
- Vault/baseline **not** updated to 4/16 until live numbers match
- Open thread retained: “confirm capacity upgrade applied to guest”