the-custodian/workplans/CUST-WP-0071-measured-workload-sizing-and-weekly-review.md
codex 693f7f1962
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 7s
Block CUST-WP-0071 and CUST-WP-0073 pending cross-repo requirements
Remaining tasks wait on GLAS-WP-0012, RAPPS-WP-0014-T03 and platform/infra/warden
owners; requirement tables added to both workplans.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-09-28 20:25:41 +02:00

15 KiB
Raw Blame History

id type title domain repo status flavor owner topic_slug created updated related state_hub_workstream_id
CUST-WP-0071 workplan Size Railiance workloads from actual demand and establish a weekly allocation review infotech the-custodian blocked planning the-custodian custodian 2026-09-11 2026-09-28
STATE-WP-0091
RCLUSTER-WP-0014
RESOURCE-WP-0003
RAPP-TELEMETRY-WP-0001
RAIL-FAB-WP-0028
RAPPS-WP-0014
VERGABE-WP-0019
HFACT-WP-0001
2249bddb-7524-5add-bd5c-c4163a6ca0f3

Measured sizing and recurring allocation review

User instruction, 2026-09-11: persist and register this work for later follow-up, then continue the invited Vergabe pilot at an explicitly accepted 60m CPU request. This workplan is not a prerequisite to deploying that prototype. It is not an assertion that 60m or the inherited 100m is a measured requirement.

The objective is an explainable, repeatable allocation process across Railiance workloads, starting with vergabe-teilnahme. Use existing telemetry, cluster evidence, resource-control portfolio and workload deployment paths. The Custodian coordinates; participating repositories retain implementation and operating authority. Fabric supplies ownership/dependency context when its accepted endpoint is available; its hosted rollout is not a prerequisite to collecting measurements or reviewing allocations.

Starting evidence:

  • docs/evidence/2026-09-11-railiance-cpu-overview.md and its JSON snapshot: 4,000m allocatable, 3,905m requested, 95m available; 20 active pods with zero CPU request, Knative/Kourier requesting 1,000m, and two stale/unhealthy pods retaining 100m. These observations must be refreshed before action.
  • Grafana/Prometheus is installed privately. Its CPU request recording omits Forgejo's 100m, and its cluster CPU utilization panel lacks the required series. Resource-control uses August capacity evidence. A finished collector implementation does not prove fresh or scheduler-correct consumer data.
  • Vergabe's 100m CPU request originated in the May 19 initial chart commit 962c5a1b3692fb27f003bda2612ffd5ed38d7ffb; no measured sizing rationale was found. The user now accepts a 60m request for one tenant with very few users, prioritizing deployment and onboarding learning over production-grade sizing.

The accepted prototype keeps its configured 1,000m CPU limit and 256Mi memory request/1Gi limit unless separately changed. The database, shared services, build workers and host/control-plane demand remain separate capacity consumers. Record actual deployed values and image revisions; do not infer them from this plan. Do not treat reservation headroom as measured spare processing capacity.

Reconcile live capacity, demand and ownership evidence

id: CUST-WP-0071-T01
status: done
priority: high
assignee: the-custodian
state_hub_task_id: "4ae8eb8d-f8c5-5256-ae50-c323cfd73d7a"

With railiance-cluster, rapp-telemetry/railiance-telemetry and resource-control, produce a timestamped observation joining cluster/node, namespace/workload, owner/service and tenant where known. Keep unknown allocations explicit. Count host and cluster representations of the same CPUs only once.

Reconcile scheduler-effective requests against the Kubernetes API, including scheduled nonterminal and terminating pods, pending demand, init/sidecar/pod resources and overhead. Explain differences from dashboard/collector totals; repair the Forgejo omission and absent host CPU signal in their owning sources. Account for zero-request containers, host/control-plane demand and data gaps. Capture CPU usage, requests/limits, throttling and contention, memory working set/peaks, OOM/restarts, storage exposure and workload activity where available. Preserve query/collector revisions, units, timestamps and coverage. Reuse RCLUSTER-WP-0014 and return the release-specific contract to STATE-WP-0091.

Done when a repeatable observation reconciles current allocation totals, identifies missing/stale signals without reporting them as zero demand, and updates the existing portfolio evidence with accountable owner references.

Done 2026-09-14. Repeatable command: python3 scripts/reconcile_railiance_allocation.py. Evidence: docs/evidence/2026-09-14-allocation-reconcile.{json,md} plus docs/evidence/namespace-ownership.yaml. Live totals: 4000m capacity, 3985m scheduled, 25m pending, 15m residual (not a guarantee). Host and cluster CPUs counted once. Unknown namespaces listed, then filled for approval-engine and vergabe-demo-company. Zero-request workloads named. STATE-WP-0091 preflight refuses (15m remaining). Forgejo omission remains a Prometheus/dashboard gap, not a Kubernetes API gap. RESOURCE-WP-0007 already updated reef-railiance-k3s owner evidence.

Measure the Vergabe prototype under representative pilot use

id: CUST-WP-0071-T02
status: wait
priority: high
assignee: the-custodian
depends_on: [CUST-WP-0071-T01]
blocking_reason: "Await RAPPS-WP-0014-T03 (blocked, needs_human): populated two-user/document workflow run by Bernd and the company contact, giving representative load for p95 <= 2s / zero failures."
state_hub_task_id: "82192370-2fd7-5363-88d2-3c67889d3d68"

VERGABE-WP-0019/RAPPS-WP-0014 supply exact release, tenant/user assumptions, runtime configuration and deployment. Observe the accepted 60m prototype and exercise an isolated synthetic-data fixture with representative concurrent users: login, list/search/detail, tender/lot/task edits and document upload/ download. Include startup, idle periods, ordinary usage, bursts and restart. Measure API/UI latency percentiles and errors alongside CPU/memory, throttling, database demand and request volume. Distinguish container CPU from incremental database/shared-service demand. No benchmark writes to existing customer data.

Record workload sizes, concurrency, hardware/image/workers, duration, coverage and limitations so results are reproducible. Founder acceptance target, 2026-09-28: with two simultaneous users, p95 of ordinary operations must be at most 2 seconds and there must be no failed operations. Report document transfer time separately. This resolves the target choice; obtain representative evidence before claiming adequacy. Recommend request/limit and memory values with a stated margin and revisit trigger; label an incomplete pilot sample provisional. A successful smoke test alone is not sizing proof.

Review ecosystem allocations and reserve operating room

id: CUST-WP-0071-T03
status: wait
priority: high
assignee: the-custodian
depends_on: [CUST-WP-0071-T01, CUST-WP-0071-T02]
blocking_reason: "Await CUST-WP-0071-T02 representative sample; Forgejo/runner demand and zero-request workload review need owner input (railiance-cluster, railiance-forge, resource-control)."
state_hub_task_id: "6dc67558-eb1e-5bb6-a667-986f884dd495"

Resource-control and the owning workload/cluster repositories compare actual demand with reservations. Prioritize Knative/Kourier's large reservations, Forgejo's higher observed demand, zero-request workloads and the stale probe/ database-drill allocations. Review each service's startup, burst, autoscaling, latency, availability and storage requirements before proposing changes.

Publish proposed per-workload requests/limits and a node-level budget including host operations, pilot growth, deployment surges/migration hooks and admitted factory/build execution. Declare margin and concurrency assumptions, current constraints, owners and rollback. If demand plus margin does not fit, compare capacity options through resource-control using current provider evidence. STATE-WP-0091 retains its implementation of release preflight; this task consumes and supplies that contract rather than replacing its workplan.

Done when the recommended allocation reconciles to admitted capacity and every remaining change has a live receiving owner/task. This plan does not grant unrelated deletion, procurement or automatic resource changes.

Apply accepted sizing changes and verify their effects

id: CUST-WP-0071-T04
status: wait
priority: medium
assignee: the-custodian
depends_on: [CUST-WP-0071-T03]
blocking_reason: "Await reviewed allocation proposals and applicable workload-owner authority."
state_hub_task_id: "944f98e2-7e42-5894-925b-752423f3a1cc"

Apply accepted changes in the existing owner manifests and deployment paths. Use staged changes and retain exact prior configuration and rollback receipts. Compare before/after traffic, latency/errors, resource consumption, throttling, OOM/restarts and release headroom over a representative observation window. Verify startup and a subsequent release at the new allocation. Revert or revise changes that violate the agreed product/service targets. Update resource-control and task evidence with deployed revisions, measured outcome and residual risks.

Do not close based only on successful scheduling or smaller requests. Acceptance requires useful operation and a reconciled budget; unimplemented proposals or unresolved measurement gaps remain live work records with owners.

Establish the regular weekly workload-demand and allocation assessment

id: CUST-WP-0071-T05
status: wait
priority: medium
assignee: the-custodian
depends_on: [CUST-WP-0071-T04]
blocking_reason: "Await the repeatable observation and accepted sizing/review procedure."
state_hub_task_id: "20ddd4d7-d85d-5b95-bb0a-7b0ca8af76ca"

As the final step, install a weekly assessment on an existing durable Railiance executor with scoped read access. Record the executor owner, schedule and time zone; proposed cadence is Monday 08:00 Europe/Berlin. Use the existing telemetry and resource-control surfaces, not a workstation reminder or a parallel monitoring service. Persist weekly aggregates so comparisons survive the current seven-day raw-metric retention, and compare the last seven days with prior weeks.

Each run records per-workload demand and activity, CPU/memory p95 and peaks, requests/limits, throttling/latency/errors/restarts, pending demand, node and release margin, coverage/freshness, ownership and week-on-week changes. Produce explicit keep/increase/decrease/investigate recommendations with evidence and confidence. Deduplicate open recommendations into live owner work records and review their results in the following run. Resource allocation adapts through the accepted owner change procedure; no unrestricted automatic resizing.

Name the weekly review owner and persist delivery/acknowledgment through an accepted internal channel. Define and exercise stale/failed/missed-run handling, retry/idempotency and an accountable recovery route. Preserve scheduler and assessment configuration in source, with non-secret execution receipts.

Done when the durable schedule is installed, the first unattended scheduled assessment has produced its retained report and owner receipt, missed-run handling is demonstrated, and ownership for the ongoing weekly operation is registered before this workplan finishes. Hand off residual recommendations as live records; the recurring assessment continues after this plan is finished.

Scheduling evidence / live handoff for T01 — 2026-09-14

During SECRETS-WP-0010-T03, repeated sso/keycape-factor-renewer Jobs requesting 10m CPU failed scheduling at full requested CPU. The provider credential expired and KeyCape lost readiness, blocking operator authentication. Informed Decision used 1m CPU; reducing its request from 20m to 5m (limit remains 500m) freed 15m. Scheduled Job keycape-factor-renewer-29822490 completed and self-revoked, ESO updated the mounted credential, and KeyCape recovered. This is immediate recovery, not a fleet sizing conclusion. T01 must include recurring maintenance-job demand and reliable scheduling headroom, not only resident pod allocations. Evidence: informed-decision/docs/evidence/2026-09-14-keycape-renewal-capacity-recovery.json.

September 28 bounded completion review

The founder asks to finish with minimal additional tasks, workplans and functionality. Keep all remaining work in T02–T05; do not spawn a monitoring service, benchmark framework or replacement coordination plan.

Evidence: docs/evidence/2026-09-28-sizing-review.md, allocation reconcile, retained cluster observation, source revisions and exact seven-day PromQL responses alongside it. T01 refreshed: 3420m requested / 4000m, no pending requests, 580m reservation residual; instantaneous node CPU was 3963m, so this is not spare processing capacity. No unsupported pod accounting features were present in this snapshot. Namespace ownership refreshed.

T02 is now in progress: the exact deployed pilot and seven days of measurements are recorded. CPU p95 0.52m, sampled peak 20.10m, memory peak 191.14Mi; retain 60m/256Mi provisionally. Representative two-user activity, response/error acceptance and incremental database attribution remain unproven. Use existing RAPPS-WP-0014-T03 fixture/acceptance work; do not duplicate its recovery scope.

T03 is now in progress: retain current pilot/Knative allocations, investigate Forgejo/runner demand (namespace CPU p95 1454m), and account for twelve zero-request workloads. RAIL-KNATIVE-WP-0002 and RAIL-EN-WP-0002 are finished; their declaration/deployment work must not be repeated. No new resource change is proposed from the incomplete sample. T04 can verify a justified keep decision; it must not create an unnecessary resize. Its useful-operation acceptance stays.

T05 still requires a durable activity-core schedule, retained report, owner receipt and missed-run recovery proof. Monday 08:00 Europe/Berlin remains the proposed cadence. None is represented as installed by this session. Existing T02–T05 retain the remaining evidence and execution, with no new work records.

September 28 follow-up: the founder selected the two-user p95 ≤ 2 seconds, zero-failed-operations target (document transfer excluded from that latency threshold). The current seven-day telemetry remains provisional until the representative workflow runs. This is an acceptance criterion, not a claim that it has passed. Keep the run and its evidence under existing T02.

Blocked — requirements on other repos (2026-09-28)

The bounded review is complete: T01 done, seven-day telemetry recorded, and the provisional keep decision (60m/256Mi) is evidenced. No further step is implementable here; workplan set to blocked. Requirements:

Owner Requirement Gates
railiance-apps (RAPPS-WP-0014-T03, needs_human) Two-user/document acceptance run on the deployed pilot with Bernd and the company contact T02
vergabe-teilnahme (VERGABE-WP-0019, blocked) Release/tenant assumptions and fixture for the representative run T02
resource-control, railiance-cluster, railiance-forge Owner review of Forgejo/runner demand (namespace CPU p95 1454m) and twelve zero-request workloads T03
activity-core Durable weekly schedule (Mon 08:00 Europe/Berlin), retained report, owner receipt, missed-run handling T05 (after T04)

Resume when RAPPS-WP-0014-T03 records the acceptance run; T03 → T04 → T05 follow.