railiance-cluster/workplans/RAIL-BS-WP-0014-state-hub-surge-headroom.md
codex 551a898e4d
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 2s
Use canonical RAIL-BS-WP-0014 id for the STATE-WP-0091 residual.
Assistant: grok
Assistant-Session: 01a09dc1-b21e-77e1-919e-fcad2f82b267
2026-09-14 10:17:39 +02:00

34 lines
1.1 KiB
Markdown

---
id: RAIL-BS-WP-0014
type: workplan
title: "Keep node remaining CPU honest for State Hub surge and pending demand"
domain: financials
repo: railiance-cluster
status: proposed
owner: codex
topic_slug: railiance
origin: residual
origin_ref: STATE-WP-0091
created: "2026-09-14"
updated: "2026-09-14"
related: [RCLUSTER-WP-0014, STATE-WP-0091, CUST-WP-0071]
---
Residual from STATE-WP-0091. State Hub now refuses promotion when remaining
CPU cannot cover its 100m API surge, and when unrelated pending pods would
eat the 5m sliver above 105m. Cluster observation must keep publishing
allocatable, allocated, **and pending-unrelated** demand. 105m remaining is
not capacity admission for factory or other zero-request workloads.
## Include pending pods in capacity observations
```task
id: RAIL-BS-WP-0014-T01
status: todo
priority: high
```
Extend `tools/observe_cluster_resources.py` / `make cluster-observe` so the
published observation names pending unscheduled pods and their requests.
State Hub's preflight already consumes that shape. Do not treat residual
millicores as a scheduling guarantee.