Assistant: codex Assistant-Model: gpt-6-astra Assistant-Session: 01a0e396-d089-7653-b0a1-734cac532913
3.2 KiB
Knative substrate runbook: declared CPU requests
The Knative Serving and Kourier v1.22.0 install on railiance01 is performed by
railiance-cluster/install/knative/install.sh from checksum-pinned upstream
assets. rail-knative declares the CPU requests that install must carry, in
substrate/v1.22.0/cpu-requests.patch.yaml (strategic-merge patches) and
substrate/v1.22.0/kustomization.yaml (overlay for the staged core.yaml and
kourier.yaml; crds.yaml stays a separate first apply).
| Deployment | Container | Upstream v1.22.0 | Declared |
|---|---|---|---|
| knative-serving/activator | activator | 300m | 50m |
| kourier-system/3scale-kourier-gateway | kourier-gateway | 200m | 50m |
| knative-serving/net-kourier-controller | controller | 200m | 30m |
| knative-serving/controller | controller | 100m | 30m |
| knative-serving/webhook | webhook | 100m | 30m |
| knative-serving/autoscaler | autoscaler | 100m | 30m |
Only CPU requests differ from upstream. Memory requests and all limits are
upstream's. These values were set live on 2026-09-21 as
ADMINISTER @ realm:kubernetes/railiance01, activation=APPROVED by the founder,
because the node had 100% of its allocatable CPU requested and backups could
not be scheduled. Record: the-custodian/docs/kubernetes-change-gate-decision.md.
The owner installer now renders these requests into the manifests before apply,
using separate Serving and Kourier overlays. Its verifier checks all six values.
Implementation and read-only live diff evidence are recorded in
RAIL-BS-WP-0015,
implemented by railiance-cluster commit 3a5432270e275e978d6c8a99529fa7f8be6eef57.
A plain re-apply of unpatched upstream manifests restores the upstream column.
Every apply or upgrade must preserve these patches, and an upgrade to a new
version needs a new substrate/<version>/ with container names re-checked
against that release.
Lessons
A request cut cannot roll out on a node whose requests are exhausted. A rolling update creates the new pod before the old one stops. When no CPU is left to request, the new pod stays Pending, so the old pod never stops, even though the change would free capacity. On 2026-09-21 all six rollouts deadlocked until 50m was freed elsewhere. Free capacity before changing a Deployment's requests; then the first rollout frees more for the next.
The HPAs scale on percent of request. The HPAs on the activator, the Kourier gateway and the webhook target 100% CPU utilisation of the request. Lowering a request makes the same load read as a higher percentage, so the HPA scales out sooner (at about 50m or 30m of use, not 300m or 100m). At 2026-09-21 use they read about 2-10%. If the request is raised or lowered again, check the HPA targets in the same change.
Read-only drift check
ssh railiance01 'kubectl get deploy -n knative-serving -o jsonpath="{range .items[*]}{.metadata.name} {.spec.template.spec.containers[0].resources.requests.cpu}{\"\n\"}{end}"; kubectl get deploy -n kourier-system -o jsonpath="{range .items[*]}{.metadata.name} {.spec.template.spec.containers[0].resources.requests.cpu}{\"\n\"}{end}"; kubectl get hpa -n knative-serving; kubectl get hpa -n kourier-system'