the-custodian/workplans/CUST-WP-0071-measured-workload-sizing-and-weekly-review.md
codex db91818e84
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 3s
Python Tests / pytest (push) Successful in 25s
Advance supervised agent records and close verified Secret annotation guard
2026-09-28 18:15:27 +02:00

260 lines
14 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
id: CUST-WP-0071
type: workplan
title: "Size Railiance workloads from actual demand and establish a weekly allocation review"
domain: infotech
repo: the-custodian
status: active
flavor: planning
owner: the-custodian
topic_slug: custodian
created: "2026-09-11"
updated: "2026-09-28"
related: [STATE-WP-0091, RCLUSTER-WP-0014, RESOURCE-WP-0003, RAPP-TELEMETRY-WP-0001, RAIL-FAB-WP-0028, RAPPS-WP-0014, VERGABE-WP-0019, HFACT-WP-0001]
state_hub_workstream_id: "2249bddb-7524-5add-bd5c-c4163a6ca0f3"
---
# Measured sizing and recurring allocation review
User instruction, 2026-09-11: persist and register this work for later follow-up,
then continue the invited Vergabe pilot at an explicitly accepted 60m CPU
request. This workplan is not a prerequisite to deploying that prototype.
It is not an assertion that 60m or the inherited 100m is a measured requirement.
The objective is an explainable, repeatable allocation process across Railiance
workloads, starting with vergabe-teilnahme. Use existing telemetry, cluster
evidence, resource-control portfolio and workload deployment paths. The
Custodian coordinates; participating repositories retain implementation and
operating authority. Fabric supplies ownership/dependency context when its
accepted endpoint is available; its hosted rollout is not a prerequisite to
collecting measurements or reviewing allocations.
Starting evidence:
- `docs/evidence/2026-09-11-railiance-cpu-overview.md` and its JSON snapshot:
4,000m allocatable, 3,905m requested, 95m available; 20 active pods with zero
CPU request, Knative/Kourier requesting 1,000m, and two stale/unhealthy pods
retaining 100m. These observations must be refreshed before action.
- Grafana/Prometheus is installed privately. Its CPU request recording omits
Forgejo's 100m, and its cluster CPU utilization panel lacks the required
series. Resource-control uses August capacity evidence. A finished collector
implementation does not prove fresh or scheduler-correct consumer data.
- Vergabe's 100m CPU request originated in the May 19 initial chart commit
`962c5a1b3692fb27f003bda2612ffd5ed38d7ffb`; no measured sizing rationale was
found. The user now accepts a 60m request for one tenant with very few users,
prioritizing deployment and onboarding learning over production-grade sizing.
The accepted prototype keeps its configured 1,000m CPU limit and 256Mi memory
request/1Gi limit unless separately changed. The database, shared services,
build workers and host/control-plane demand remain separate capacity consumers.
Record actual deployed values and image revisions; do not infer them from this
plan. Do not treat reservation headroom as measured spare processing capacity.
## Reconcile live capacity, demand and ownership evidence
```task
id: CUST-WP-0071-T01
status: done
priority: high
assignee: the-custodian
state_hub_task_id: "4ae8eb8d-f8c5-5256-ae50-c323cfd73d7a"
```
With railiance-cluster, rapp-telemetry/railiance-telemetry and resource-control,
produce a timestamped observation joining cluster/node, namespace/workload,
owner/service and tenant where known. Keep unknown allocations explicit.
Count host and cluster representations of the same CPUs only once.
Reconcile scheduler-effective requests against the Kubernetes API, including
scheduled nonterminal and terminating pods, pending demand, init/sidecar/pod
resources and overhead. Explain differences from dashboard/collector totals;
repair the Forgejo omission and absent host CPU signal in their owning sources.
Account for zero-request containers, host/control-plane demand and data gaps.
Capture CPU usage, requests/limits, throttling and contention, memory working
set/peaks, OOM/restarts, storage exposure and workload activity where available.
Preserve query/collector revisions, units, timestamps and coverage. Reuse
RCLUSTER-WP-0014 and return the release-specific contract to STATE-WP-0091.
Done when a repeatable observation reconciles current allocation totals,
identifies missing/stale signals without reporting them as zero demand, and
updates the existing portfolio evidence with accountable owner references.
**Done 2026-09-14.** Repeatable command:
`python3 scripts/reconcile_railiance_allocation.py`. Evidence:
`docs/evidence/2026-09-14-allocation-reconcile.{json,md}` plus
`docs/evidence/namespace-ownership.yaml`. Live totals: 4000m capacity,
3985m scheduled, 25m pending, 15m residual (not a guarantee). Host and
cluster CPUs counted once. Unknown namespaces listed, then filled for
`approval-engine` and `vergabe-demo-company`. Zero-request workloads named.
STATE-WP-0091 preflight refuses (15m remaining). Forgejo omission remains a
Prometheus/dashboard gap, not a Kubernetes API gap. RESOURCE-WP-0007 already
updated reef-railiance-k3s owner evidence.
## Measure the Vergabe prototype under representative pilot use
```task
id: CUST-WP-0071-T02
status: progress
priority: high
assignee: the-custodian
depends_on: [CUST-WP-0071-T01]
state_hub_task_id: "82192370-2fd7-5363-88d2-3c67889d3d68"
```
VERGABE-WP-0019/RAPPS-WP-0014 supply exact release, tenant/user assumptions,
runtime configuration and deployment. Observe the accepted 60m prototype and
exercise an isolated synthetic-data fixture with representative concurrent
users: login, list/search/detail, tender/lot/task edits and document upload/
download. Include startup, idle periods, ordinary usage, bursts and restart.
Measure API/UI latency percentiles and errors alongside CPU/memory, throttling,
database demand and request volume. Distinguish container CPU from incremental
database/shared-service demand. No benchmark writes to existing customer data.
Record workload sizes, concurrency, hardware/image/workers, duration, coverage
and limitations so results are reproducible. Founder acceptance target, 2026-09-28: with two simultaneous users, p95 of
ordinary operations must be at most 2 seconds and there must be no failed
operations. Report document transfer time separately. This resolves the target
choice; obtain representative evidence before claiming adequacy. Recommend request/limit and
memory values with a stated margin and revisit trigger; label an incomplete
pilot sample provisional. A successful smoke test alone is not sizing proof.
## Review ecosystem allocations and reserve operating room
```task
id: CUST-WP-0071-T03
status: progress
priority: high
assignee: the-custodian
depends_on: [CUST-WP-0071-T01, CUST-WP-0071-T02]
state_hub_task_id: "6dc67558-eb1e-5bb6-a667-986f884dd495"
```
Resource-control and the owning workload/cluster repositories compare actual
demand with reservations. Prioritize Knative/Kourier's large reservations,
Forgejo's higher observed demand, zero-request workloads and the stale probe/
database-drill allocations. Review each service's startup, burst, autoscaling,
latency, availability and storage requirements before proposing changes.
Publish proposed per-workload requests/limits and a node-level budget including
host operations, pilot growth, deployment surges/migration hooks and admitted
factory/build execution. Declare margin and concurrency assumptions, current
constraints, owners and rollback. If demand plus margin does not fit, compare
capacity options through resource-control using current provider evidence.
STATE-WP-0091 retains its implementation of release preflight; this task consumes
and supplies that contract rather than replacing its workplan.
Done when the recommended allocation reconciles to admitted capacity and every
remaining change has a live receiving owner/task. This plan does not grant
unrelated deletion, procurement or automatic resource changes.
## Apply accepted sizing changes and verify their effects
```task
id: CUST-WP-0071-T04
status: wait
priority: medium
assignee: the-custodian
depends_on: [CUST-WP-0071-T03]
blocking_reason: "Await reviewed allocation proposals and applicable workload-owner authority."
state_hub_task_id: "944f98e2-7e42-5894-925b-752423f3a1cc"
```
Apply accepted changes in the existing owner manifests and deployment paths.
Use staged changes and retain exact prior configuration and rollback receipts.
Compare before/after traffic, latency/errors, resource consumption, throttling,
OOM/restarts and release headroom over a representative observation window.
Verify startup and a subsequent release at the new allocation. Revert or revise
changes that violate the agreed product/service targets. Update resource-control
and task evidence with deployed revisions, measured outcome and residual risks.
Do not close based only on successful scheduling or smaller requests. Acceptance
requires useful operation and a reconciled budget; unimplemented proposals or
unresolved measurement gaps remain live work records with owners.
## Establish the regular weekly workload-demand and allocation assessment
```task
id: CUST-WP-0071-T05
status: wait
priority: medium
assignee: the-custodian
depends_on: [CUST-WP-0071-T04]
blocking_reason: "Await the repeatable observation and accepted sizing/review procedure."
state_hub_task_id: "20ddd4d7-d85d-5b95-bb0a-7b0ca8af76ca"
```
As the final step, install a weekly assessment on an existing durable Railiance
executor with scoped read access. Record the executor owner, schedule and time
zone; proposed cadence is Monday 08:00 Europe/Berlin. Use the existing telemetry
and resource-control surfaces, not a workstation reminder or a parallel
monitoring service. Persist weekly aggregates so comparisons survive the current
seven-day raw-metric retention, and compare the last seven days with prior weeks.
Each run records per-workload demand and activity, CPU/memory p95 and peaks,
requests/limits, throttling/latency/errors/restarts, pending demand, node and
release margin, coverage/freshness, ownership and week-on-week changes. Produce
explicit keep/increase/decrease/investigate recommendations with evidence and
confidence. Deduplicate open recommendations into live owner work records and
review their results in the following run. Resource allocation adapts through
the accepted owner change procedure; no unrestricted automatic resizing.
Name the weekly review owner and persist delivery/acknowledgment through an
accepted internal channel. Define and exercise stale/failed/missed-run handling,
retry/idempotency and an accountable recovery route. Preserve scheduler and
assessment configuration in source, with non-secret execution receipts.
Done when the durable schedule is installed, the first unattended scheduled
assessment has produced its retained report and owner receipt, missed-run
handling is demonstrated, and ownership for the ongoing weekly operation is
registered before this workplan finishes. Hand off residual recommendations as
live records; the recurring assessment continues after this plan is finished.
### Scheduling evidence / live handoff for T01 — 2026-09-14
During SECRETS-WP-0010-T03, repeated `sso/keycape-factor-renewer` Jobs requesting
10m CPU failed scheduling at full requested CPU. The provider credential expired
and KeyCape lost readiness, blocking operator authentication. Informed Decision
used 1m CPU; reducing its request from 20m to 5m (limit remains 500m) freed 15m.
Scheduled Job keycape-factor-renewer-29822490 completed and self-revoked, ESO
updated the mounted credential, and KeyCape recovered. This is immediate recovery,
not a fleet sizing conclusion. T01 must include recurring maintenance-job demand
and reliable scheduling headroom, not only resident pod allocations. Evidence:
informed-decision/docs/evidence/2026-09-14-keycape-renewal-capacity-recovery.json.
## September 28 bounded completion review
The founder asks to finish with minimal additional tasks, workplans and
functionality. Keep all remaining work in T02–T05; do not spawn a monitoring
service, benchmark framework or replacement coordination plan.
Evidence: `docs/evidence/2026-09-28-sizing-review.md`, allocation reconcile,
retained cluster observation, source revisions and exact seven-day PromQL
responses alongside it. T01 refreshed: 3420m requested / 4000m, no pending
requests, 580m reservation residual; instantaneous node CPU was 3963m, so this
is not spare processing capacity. No unsupported pod accounting features were
present in this snapshot. Namespace ownership refreshed.
T02 is now in progress: the exact deployed pilot and seven days of measurements
are recorded. CPU p95 0.52m, sampled peak 20.10m, memory peak 191.14Mi; retain
60m/256Mi provisionally. Representative two-user activity, response/error
acceptance and incremental database attribution remain unproven. Use existing
RAPPS-WP-0014-T03 fixture/acceptance work; do not duplicate its recovery scope.
T03 is now in progress: retain current pilot/Knative allocations, investigate
Forgejo/runner demand (namespace CPU p95 1454m), and account for twelve
zero-request workloads. RAIL-KNATIVE-WP-0002 and RAIL-EN-WP-0002 are finished;
their declaration/deployment work must not be repeated. No new resource change
is proposed from the incomplete sample. T04 can verify a justified keep decision;
it must not create an unnecessary resize. Its useful-operation acceptance stays.
T05 still requires a durable activity-core schedule, retained report, owner
receipt and missed-run recovery proof. Monday 08:00 Europe/Berlin remains the
proposed cadence. None is represented as installed by this session. Existing
T02–T05 retain the remaining evidence and execution, with no new work records.
September 28 follow-up: the founder selected the two-user p95 ≤ 2 seconds,
zero-failed-operations target (document transfer excluded from that latency
threshold). The current seven-day telemetry remains provisional until the
representative workflow runs. This is an acceptance criterion, not a claim that
it has passed. Keep the run and its evidence under existing T02.