Set flavor on open workplans from origin/prose/status. Copy existing depends_on aliases only. Do not promote residuals.
223 lines
11 KiB
Markdown
223 lines
11 KiB
Markdown
---
|
|
id: CUST-WP-0071
|
|
type: workplan
|
|
title: "Size Railiance workloads from actual demand and establish a weekly allocation review"
|
|
domain: infotech
|
|
repo: the-custodian
|
|
status: active
|
|
flavor: planning
|
|
owner: the-custodian
|
|
topic_slug: custodian
|
|
created: "2026-09-11"
|
|
updated: "2026-09-14"
|
|
related: [STATE-WP-0091, RCLUSTER-WP-0014, RESOURCE-WP-0003, RAPP-TELEMETRY-WP-0001, RAIL-FAB-WP-0028, RAPPS-WP-0014, VERGABE-WP-0019, HFACT-WP-0001]
|
|
state_hub_workstream_id: "2249bddb-7524-5add-bd5c-c4163a6ca0f3"
|
|
---
|
|
|
|
# Measured sizing and recurring allocation review
|
|
|
|
User instruction, 2026-09-11: persist and register this work for later follow-up,
|
|
then continue the invited Vergabe pilot at an explicitly accepted 60m CPU
|
|
request. This ready workplan is not a prerequisite to deploying that prototype.
|
|
It is not an assertion that 60m or the inherited 100m is a measured requirement.
|
|
|
|
The objective is an explainable, repeatable allocation process across Railiance
|
|
workloads, starting with vergabe-teilnahme. Use existing telemetry, cluster
|
|
evidence, resource-control portfolio and workload deployment paths. The
|
|
Custodian coordinates; participating repositories retain implementation and
|
|
operating authority. Fabric supplies ownership/dependency context when its
|
|
accepted endpoint is available; its hosted rollout is not a prerequisite to
|
|
collecting measurements or reviewing allocations.
|
|
|
|
Starting evidence:
|
|
|
|
- `docs/evidence/2026-09-11-railiance-cpu-overview.md` and its JSON snapshot:
|
|
4,000m allocatable, 3,905m requested, 95m available; 20 active pods with zero
|
|
CPU request, Knative/Kourier requesting 1,000m, and two stale/unhealthy pods
|
|
retaining 100m. These observations must be refreshed before action.
|
|
- Grafana/Prometheus is installed privately. Its CPU request recording omits
|
|
Forgejo's 100m, and its cluster CPU utilization panel lacks the required
|
|
series. Resource-control uses August capacity evidence. A finished collector
|
|
implementation does not prove fresh or scheduler-correct consumer data.
|
|
- Vergabe's 100m CPU request originated in the May 19 initial chart commit
|
|
`962c5a1b3692fb27f003bda2612ffd5ed38d7ffb`; no measured sizing rationale was
|
|
found. The user now accepts a 60m request for one tenant with very few users,
|
|
prioritizing deployment and onboarding learning over production-grade sizing.
|
|
|
|
The accepted prototype keeps its configured 1,000m CPU limit and 256Mi memory
|
|
request/1Gi limit unless separately changed. The database, shared services,
|
|
build workers and host/control-plane demand remain separate capacity consumers.
|
|
Record actual deployed values and image revisions; do not infer them from this
|
|
plan. Do not treat reservation headroom as measured spare processing capacity.
|
|
|
|
## Reconcile live capacity, demand and ownership evidence
|
|
|
|
```task
|
|
id: CUST-WP-0071-T01
|
|
status: done
|
|
priority: high
|
|
assignee: the-custodian
|
|
state_hub_task_id: "4ae8eb8d-f8c5-5256-ae50-c323cfd73d7a"
|
|
```
|
|
|
|
With railiance-cluster, rapp-telemetry/railiance-telemetry and resource-control,
|
|
produce a timestamped observation joining cluster/node, namespace/workload,
|
|
owner/service and tenant where known. Keep unknown allocations explicit.
|
|
Count host and cluster representations of the same CPUs only once.
|
|
|
|
Reconcile scheduler-effective requests against the Kubernetes API, including
|
|
scheduled nonterminal and terminating pods, pending demand, init/sidecar/pod
|
|
resources and overhead. Explain differences from dashboard/collector totals;
|
|
repair the Forgejo omission and absent host CPU signal in their owning sources.
|
|
Account for zero-request containers, host/control-plane demand and data gaps.
|
|
Capture CPU usage, requests/limits, throttling and contention, memory working
|
|
set/peaks, OOM/restarts, storage exposure and workload activity where available.
|
|
Preserve query/collector revisions, units, timestamps and coverage. Reuse
|
|
RCLUSTER-WP-0014 and return the release-specific contract to STATE-WP-0091.
|
|
|
|
Done when a repeatable observation reconciles current allocation totals,
|
|
identifies missing/stale signals without reporting them as zero demand, and
|
|
updates the existing portfolio evidence with accountable owner references.
|
|
|
|
**Done 2026-09-14.** Repeatable command:
|
|
`python3 scripts/reconcile_railiance_allocation.py`. Evidence:
|
|
`docs/evidence/2026-09-14-allocation-reconcile.{json,md}` plus
|
|
`docs/evidence/namespace-ownership.yaml`. Live totals: 4000m capacity,
|
|
3985m scheduled, 25m pending, 15m residual (not a guarantee). Host and
|
|
cluster CPUs counted once. Unknown namespaces listed, then filled for
|
|
`approval-engine` and `vergabe-demo-company`. Zero-request workloads named.
|
|
STATE-WP-0091 preflight refuses (15m remaining). Forgejo omission remains a
|
|
Prometheus/dashboard gap, not a Kubernetes API gap. RESOURCE-WP-0007 already
|
|
updated reef-railiance-k3s owner evidence.
|
|
|
|
## Measure the Vergabe prototype under representative pilot use
|
|
|
|
```task
|
|
id: CUST-WP-0071-T02
|
|
status: wait
|
|
priority: high
|
|
assignee: the-custodian
|
|
depends_on: [CUST-WP-0071-T01]
|
|
blocking_reason: "Await reliable measurements and the current invited-pilot deployment or an equivalent isolated fixture."
|
|
state_hub_task_id: "82192370-2fd7-5363-88d2-3c67889d3d68"
|
|
```
|
|
|
|
VERGABE-WP-0019/RAPPS-WP-0014 supply exact release, tenant/user assumptions,
|
|
runtime configuration and deployment. Observe the accepted 60m prototype and
|
|
exercise an isolated synthetic-data fixture with representative concurrent
|
|
users: login, list/search/detail, tender/lot/task edits and document upload/
|
|
download. Include startup, idle periods, ordinary usage, bursts and restart.
|
|
Measure API/UI latency percentiles and errors alongside CPU/memory, throttling,
|
|
database demand and request volume. Distinguish container CPU from incremental
|
|
database/shared-service demand. No benchmark writes to existing customer data.
|
|
|
|
Record workload sizes, concurrency, hardware/image/workers, duration, coverage
|
|
and limitations so results are reproducible. Agree response-time/error targets
|
|
with the product owner before claiming adequacy. Recommend request/limit and
|
|
memory values with a stated margin and revisit trigger; label an incomplete
|
|
pilot sample provisional. A successful smoke test alone is not sizing proof.
|
|
|
|
## Review ecosystem allocations and reserve operating room
|
|
|
|
```task
|
|
id: CUST-WP-0071-T03
|
|
status: wait
|
|
priority: high
|
|
assignee: the-custodian
|
|
depends_on: [CUST-WP-0071-T01, CUST-WP-0071-T02]
|
|
blocking_reason: "Await reconciled demand evidence and a measured pilot recommendation."
|
|
state_hub_task_id: "6dc67558-eb1e-5bb6-a667-986f884dd495"
|
|
```
|
|
|
|
Resource-control and the owning workload/cluster repositories compare actual
|
|
demand with reservations. Prioritize Knative/Kourier's large reservations,
|
|
Forgejo's higher observed demand, zero-request workloads and the stale probe/
|
|
database-drill allocations. Review each service's startup, burst, autoscaling,
|
|
latency, availability and storage requirements before proposing changes.
|
|
|
|
Publish proposed per-workload requests/limits and a node-level budget including
|
|
host operations, pilot growth, deployment surges/migration hooks and admitted
|
|
factory/build execution. Declare margin and concurrency assumptions, current
|
|
constraints, owners and rollback. If demand plus margin does not fit, compare
|
|
capacity options through resource-control using current provider evidence.
|
|
STATE-WP-0091 retains its implementation of release preflight; this task consumes
|
|
and supplies that contract rather than replacing its workplan.
|
|
|
|
Done when the recommended allocation reconciles to admitted capacity and every
|
|
remaining change has a live receiving owner/task. This plan does not grant
|
|
unrelated deletion, procurement or automatic resource changes.
|
|
|
|
## Apply accepted sizing changes and verify their effects
|
|
|
|
```task
|
|
id: CUST-WP-0071-T04
|
|
status: wait
|
|
priority: medium
|
|
assignee: the-custodian
|
|
depends_on: [CUST-WP-0071-T03]
|
|
blocking_reason: "Await reviewed allocation proposals and applicable workload-owner authority."
|
|
state_hub_task_id: "944f98e2-7e42-5894-925b-752423f3a1cc"
|
|
```
|
|
|
|
Apply accepted changes in the existing owner manifests and deployment paths.
|
|
Use staged changes and retain exact prior configuration and rollback receipts.
|
|
Compare before/after traffic, latency/errors, resource consumption, throttling,
|
|
OOM/restarts and release headroom over a representative observation window.
|
|
Verify startup and a subsequent release at the new allocation. Revert or revise
|
|
changes that violate the agreed product/service targets. Update resource-control
|
|
and task evidence with deployed revisions, measured outcome and residual risks.
|
|
|
|
Do not close based only on successful scheduling or smaller requests. Acceptance
|
|
requires useful operation and a reconciled budget; unimplemented proposals or
|
|
unresolved measurement gaps remain live work records with owners.
|
|
|
|
## Establish the regular weekly workload-demand and allocation assessment
|
|
|
|
```task
|
|
id: CUST-WP-0071-T05
|
|
status: wait
|
|
priority: medium
|
|
assignee: the-custodian
|
|
depends_on: [CUST-WP-0071-T04]
|
|
blocking_reason: "Await the repeatable observation and accepted sizing/review procedure."
|
|
state_hub_task_id: "20ddd4d7-d85d-5b95-bb0a-7b0ca8af76ca"
|
|
```
|
|
|
|
As the final step, install a weekly assessment on an existing durable Railiance
|
|
executor with scoped read access. Record the executor owner, schedule and time
|
|
zone; proposed cadence is Monday 08:00 Europe/Berlin. Use the existing telemetry
|
|
and resource-control surfaces, not a workstation reminder or a parallel
|
|
monitoring service. Persist weekly aggregates so comparisons survive the current
|
|
seven-day raw-metric retention, and compare the last seven days with prior weeks.
|
|
|
|
Each run records per-workload demand and activity, CPU/memory p95 and peaks,
|
|
requests/limits, throttling/latency/errors/restarts, pending demand, node and
|
|
release margin, coverage/freshness, ownership and week-on-week changes. Produce
|
|
explicit keep/increase/decrease/investigate recommendations with evidence and
|
|
confidence. Deduplicate open recommendations into live owner work records and
|
|
review their results in the following run. Resource allocation adapts through
|
|
the accepted owner change procedure; no unrestricted automatic resizing.
|
|
|
|
Name the weekly review owner and persist delivery/acknowledgment through an
|
|
accepted internal channel. Define and exercise stale/failed/missed-run handling,
|
|
retry/idempotency and an accountable recovery route. Preserve scheduler and
|
|
assessment configuration in source, with non-secret execution receipts.
|
|
|
|
Done when the durable schedule is installed, the first unattended scheduled
|
|
assessment has produced its retained report and owner receipt, missed-run
|
|
handling is demonstrated, and ownership for the ongoing weekly operation is
|
|
registered before this workplan finishes. Hand off residual recommendations as
|
|
live records; the recurring assessment continues after this plan is finished.
|
|
|
|
|
|
### Scheduling evidence / live handoff for T01 — 2026-09-14
|
|
|
|
During SECRETS-WP-0010-T03, repeated `sso/keycape-factor-renewer` Jobs requesting
|
|
10m CPU failed scheduling at full requested CPU. The provider credential expired
|
|
and KeyCape lost readiness, blocking operator authentication. Informed Decision
|
|
used 1m CPU; reducing its request from 20m to 5m (limit remains 500m) freed 15m.
|
|
Scheduled Job keycape-factor-renewer-29822490 completed and self-revoked, ESO
|
|
updated the mounted credential, and KeyCape recovered. This is immediate recovery,
|
|
not a fleet sizing conclusion. T01 must include recurring maintenance-job demand
|
|
and reliable scheduling headroom, not only resident pod allocations. Evidence:
|
|
informed-decision/docs/evidence/2026-09-14-keycape-renewal-capacity-recovery.json.
|