the-custodian/workplans/CUST-WP-0071-measured-workload-sizing-and-weekly-review.md
codex 1bba535d1f Classify open workplans with flavor (CUST-WP-0072).
Set flavor on open workplans from origin/prose/status. Copy existing
depends_on aliases only. Do not promote residuals.
2026-09-14 15:50:49 +02:00

11 KiB

id type title domain repo status flavor owner topic_slug created updated related state_hub_workstream_id
CUST-WP-0071 workplan Size Railiance workloads from actual demand and establish a weekly allocation review infotech the-custodian active planning the-custodian custodian 2026-09-11 2026-09-14
STATE-WP-0091
RCLUSTER-WP-0014
RESOURCE-WP-0003
RAPP-TELEMETRY-WP-0001
RAIL-FAB-WP-0028
RAPPS-WP-0014
VERGABE-WP-0019
HFACT-WP-0001
2249bddb-7524-5add-bd5c-c4163a6ca0f3

Measured sizing and recurring allocation review

User instruction, 2026-09-11: persist and register this work for later follow-up, then continue the invited Vergabe pilot at an explicitly accepted 60m CPU request. This ready workplan is not a prerequisite to deploying that prototype. It is not an assertion that 60m or the inherited 100m is a measured requirement.

The objective is an explainable, repeatable allocation process across Railiance workloads, starting with vergabe-teilnahme. Use existing telemetry, cluster evidence, resource-control portfolio and workload deployment paths. The Custodian coordinates; participating repositories retain implementation and operating authority. Fabric supplies ownership/dependency context when its accepted endpoint is available; its hosted rollout is not a prerequisite to collecting measurements or reviewing allocations.

Starting evidence:

  • docs/evidence/2026-09-11-railiance-cpu-overview.md and its JSON snapshot: 4,000m allocatable, 3,905m requested, 95m available; 20 active pods with zero CPU request, Knative/Kourier requesting 1,000m, and two stale/unhealthy pods retaining 100m. These observations must be refreshed before action.
  • Grafana/Prometheus is installed privately. Its CPU request recording omits Forgejo's 100m, and its cluster CPU utilization panel lacks the required series. Resource-control uses August capacity evidence. A finished collector implementation does not prove fresh or scheduler-correct consumer data.
  • Vergabe's 100m CPU request originated in the May 19 initial chart commit 962c5a1b3692fb27f003bda2612ffd5ed38d7ffb; no measured sizing rationale was found. The user now accepts a 60m request for one tenant with very few users, prioritizing deployment and onboarding learning over production-grade sizing.

The accepted prototype keeps its configured 1,000m CPU limit and 256Mi memory request/1Gi limit unless separately changed. The database, shared services, build workers and host/control-plane demand remain separate capacity consumers. Record actual deployed values and image revisions; do not infer them from this plan. Do not treat reservation headroom as measured spare processing capacity.

Reconcile live capacity, demand and ownership evidence

id: CUST-WP-0071-T01
status: done
priority: high
assignee: the-custodian
state_hub_task_id: "4ae8eb8d-f8c5-5256-ae50-c323cfd73d7a"

With railiance-cluster, rapp-telemetry/railiance-telemetry and resource-control, produce a timestamped observation joining cluster/node, namespace/workload, owner/service and tenant where known. Keep unknown allocations explicit. Count host and cluster representations of the same CPUs only once.

Reconcile scheduler-effective requests against the Kubernetes API, including scheduled nonterminal and terminating pods, pending demand, init/sidecar/pod resources and overhead. Explain differences from dashboard/collector totals; repair the Forgejo omission and absent host CPU signal in their owning sources. Account for zero-request containers, host/control-plane demand and data gaps. Capture CPU usage, requests/limits, throttling and contention, memory working set/peaks, OOM/restarts, storage exposure and workload activity where available. Preserve query/collector revisions, units, timestamps and coverage. Reuse RCLUSTER-WP-0014 and return the release-specific contract to STATE-WP-0091.

Done when a repeatable observation reconciles current allocation totals, identifies missing/stale signals without reporting them as zero demand, and updates the existing portfolio evidence with accountable owner references.

Done 2026-09-14. Repeatable command: python3 scripts/reconcile_railiance_allocation.py. Evidence: docs/evidence/2026-09-14-allocation-reconcile.{json,md} plus docs/evidence/namespace-ownership.yaml. Live totals: 4000m capacity, 3985m scheduled, 25m pending, 15m residual (not a guarantee). Host and cluster CPUs counted once. Unknown namespaces listed, then filled for approval-engine and vergabe-demo-company. Zero-request workloads named. STATE-WP-0091 preflight refuses (15m remaining). Forgejo omission remains a Prometheus/dashboard gap, not a Kubernetes API gap. RESOURCE-WP-0007 already updated reef-railiance-k3s owner evidence.

Measure the Vergabe prototype under representative pilot use

id: CUST-WP-0071-T02
status: wait
priority: high
assignee: the-custodian
depends_on: [CUST-WP-0071-T01]
blocking_reason: "Await reliable measurements and the current invited-pilot deployment or an equivalent isolated fixture."
state_hub_task_id: "82192370-2fd7-5363-88d2-3c67889d3d68"

VERGABE-WP-0019/RAPPS-WP-0014 supply exact release, tenant/user assumptions, runtime configuration and deployment. Observe the accepted 60m prototype and exercise an isolated synthetic-data fixture with representative concurrent users: login, list/search/detail, tender/lot/task edits and document upload/ download. Include startup, idle periods, ordinary usage, bursts and restart. Measure API/UI latency percentiles and errors alongside CPU/memory, throttling, database demand and request volume. Distinguish container CPU from incremental database/shared-service demand. No benchmark writes to existing customer data.

Record workload sizes, concurrency, hardware/image/workers, duration, coverage and limitations so results are reproducible. Agree response-time/error targets with the product owner before claiming adequacy. Recommend request/limit and memory values with a stated margin and revisit trigger; label an incomplete pilot sample provisional. A successful smoke test alone is not sizing proof.

Review ecosystem allocations and reserve operating room

id: CUST-WP-0071-T03
status: wait
priority: high
assignee: the-custodian
depends_on: [CUST-WP-0071-T01, CUST-WP-0071-T02]
blocking_reason: "Await reconciled demand evidence and a measured pilot recommendation."
state_hub_task_id: "6dc67558-eb1e-5bb6-a667-986f884dd495"

Resource-control and the owning workload/cluster repositories compare actual demand with reservations. Prioritize Knative/Kourier's large reservations, Forgejo's higher observed demand, zero-request workloads and the stale probe/ database-drill allocations. Review each service's startup, burst, autoscaling, latency, availability and storage requirements before proposing changes.

Publish proposed per-workload requests/limits and a node-level budget including host operations, pilot growth, deployment surges/migration hooks and admitted factory/build execution. Declare margin and concurrency assumptions, current constraints, owners and rollback. If demand plus margin does not fit, compare capacity options through resource-control using current provider evidence. STATE-WP-0091 retains its implementation of release preflight; this task consumes and supplies that contract rather than replacing its workplan.

Done when the recommended allocation reconciles to admitted capacity and every remaining change has a live receiving owner/task. This plan does not grant unrelated deletion, procurement or automatic resource changes.

Apply accepted sizing changes and verify their effects

id: CUST-WP-0071-T04
status: wait
priority: medium
assignee: the-custodian
depends_on: [CUST-WP-0071-T03]
blocking_reason: "Await reviewed allocation proposals and applicable workload-owner authority."
state_hub_task_id: "944f98e2-7e42-5894-925b-752423f3a1cc"

Apply accepted changes in the existing owner manifests and deployment paths. Use staged changes and retain exact prior configuration and rollback receipts. Compare before/after traffic, latency/errors, resource consumption, throttling, OOM/restarts and release headroom over a representative observation window. Verify startup and a subsequent release at the new allocation. Revert or revise changes that violate the agreed product/service targets. Update resource-control and task evidence with deployed revisions, measured outcome and residual risks.

Do not close based only on successful scheduling or smaller requests. Acceptance requires useful operation and a reconciled budget; unimplemented proposals or unresolved measurement gaps remain live work records with owners.

Establish the regular weekly workload-demand and allocation assessment

id: CUST-WP-0071-T05
status: wait
priority: medium
assignee: the-custodian
depends_on: [CUST-WP-0071-T04]
blocking_reason: "Await the repeatable observation and accepted sizing/review procedure."
state_hub_task_id: "20ddd4d7-d85d-5b95-bb0a-7b0ca8af76ca"

As the final step, install a weekly assessment on an existing durable Railiance executor with scoped read access. Record the executor owner, schedule and time zone; proposed cadence is Monday 08:00 Europe/Berlin. Use the existing telemetry and resource-control surfaces, not a workstation reminder or a parallel monitoring service. Persist weekly aggregates so comparisons survive the current seven-day raw-metric retention, and compare the last seven days with prior weeks.

Each run records per-workload demand and activity, CPU/memory p95 and peaks, requests/limits, throttling/latency/errors/restarts, pending demand, node and release margin, coverage/freshness, ownership and week-on-week changes. Produce explicit keep/increase/decrease/investigate recommendations with evidence and confidence. Deduplicate open recommendations into live owner work records and review their results in the following run. Resource allocation adapts through the accepted owner change procedure; no unrestricted automatic resizing.

Name the weekly review owner and persist delivery/acknowledgment through an accepted internal channel. Define and exercise stale/failed/missed-run handling, retry/idempotency and an accountable recovery route. Preserve scheduler and assessment configuration in source, with non-secret execution receipts.

Done when the durable schedule is installed, the first unattended scheduled assessment has produced its retained report and owner receipt, missed-run handling is demonstrated, and ownership for the ongoing weekly operation is registered before this workplan finishes. Hand off residual recommendations as live records; the recurring assessment continues after this plan is finished.

Scheduling evidence / live handoff for T01 — 2026-09-14

During SECRETS-WP-0010-T03, repeated sso/keycape-factor-renewer Jobs requesting 10m CPU failed scheduling at full requested CPU. The provider credential expired and KeyCape lost readiness, blocking operator authentication. Informed Decision used 1m CPU; reducing its request from 20m to 5m (limit remains 500m) freed 15m. Scheduled Job keycape-factor-renewer-29822490 completed and self-revoked, ESO updated the mounted credential, and KeyCape recovered. This is immediate recovery, not a fleet sizing conclusion. T01 must include recurring maintenance-job demand and reliable scheduling headroom, not only resident pod allocations. Evidence: informed-decision/docs/evidence/2026-09-14-keycape-renewal-capacity-recovery.json.