CUST-WP-0071-T02 depends on RAPPS-WP-0014-T03 as its reason said; CUST-WP-0073-T02/T04 gain depends_on, T05 is a human gate. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
277 lines
15 KiB
Markdown
277 lines
15 KiB
Markdown
---
|
||
id: CUST-WP-0071
|
||
type: workplan
|
||
title: "Size Railiance workloads from actual demand and establish a weekly allocation review"
|
||
domain: infotech
|
||
repo: the-custodian
|
||
status: blocked
|
||
flavor: planning
|
||
owner: the-custodian
|
||
topic_slug: custodian
|
||
created: "2026-09-11"
|
||
updated: "2026-09-28"
|
||
related: [STATE-WP-0091, RCLUSTER-WP-0014, RESOURCE-WP-0003, RAPP-TELEMETRY-WP-0001, RAIL-FAB-WP-0028, RAPPS-WP-0014, VERGABE-WP-0019, HFACT-WP-0001]
|
||
state_hub_workstream_id: "2249bddb-7524-5add-bd5c-c4163a6ca0f3"
|
||
---
|
||
|
||
# Measured sizing and recurring allocation review
|
||
|
||
User instruction, 2026-09-11: persist and register this work for later follow-up,
|
||
then continue the invited Vergabe pilot at an explicitly accepted 60m CPU
|
||
request. This workplan is not a prerequisite to deploying that prototype.
|
||
It is not an assertion that 60m or the inherited 100m is a measured requirement.
|
||
|
||
The objective is an explainable, repeatable allocation process across Railiance
|
||
workloads, starting with vergabe-teilnahme. Use existing telemetry, cluster
|
||
evidence, resource-control portfolio and workload deployment paths. The
|
||
Custodian coordinates; participating repositories retain implementation and
|
||
operating authority. Fabric supplies ownership/dependency context when its
|
||
accepted endpoint is available; its hosted rollout is not a prerequisite to
|
||
collecting measurements or reviewing allocations.
|
||
|
||
Starting evidence:
|
||
|
||
- `docs/evidence/2026-09-11-railiance-cpu-overview.md` and its JSON snapshot:
|
||
4,000m allocatable, 3,905m requested, 95m available; 20 active pods with zero
|
||
CPU request, Knative/Kourier requesting 1,000m, and two stale/unhealthy pods
|
||
retaining 100m. These observations must be refreshed before action.
|
||
- Grafana/Prometheus is installed privately. Its CPU request recording omits
|
||
Forgejo's 100m, and its cluster CPU utilization panel lacks the required
|
||
series. Resource-control uses August capacity evidence. A finished collector
|
||
implementation does not prove fresh or scheduler-correct consumer data.
|
||
- Vergabe's 100m CPU request originated in the May 19 initial chart commit
|
||
`962c5a1b3692fb27f003bda2612ffd5ed38d7ffb`; no measured sizing rationale was
|
||
found. The user now accepts a 60m request for one tenant with very few users,
|
||
prioritizing deployment and onboarding learning over production-grade sizing.
|
||
|
||
The accepted prototype keeps its configured 1,000m CPU limit and 256Mi memory
|
||
request/1Gi limit unless separately changed. The database, shared services,
|
||
build workers and host/control-plane demand remain separate capacity consumers.
|
||
Record actual deployed values and image revisions; do not infer them from this
|
||
plan. Do not treat reservation headroom as measured spare processing capacity.
|
||
|
||
## Reconcile live capacity, demand and ownership evidence
|
||
|
||
```task
|
||
id: CUST-WP-0071-T01
|
||
status: done
|
||
priority: high
|
||
assignee: the-custodian
|
||
state_hub_task_id: "4ae8eb8d-f8c5-5256-ae50-c323cfd73d7a"
|
||
```
|
||
|
||
With railiance-cluster, rapp-telemetry/railiance-telemetry and resource-control,
|
||
produce a timestamped observation joining cluster/node, namespace/workload,
|
||
owner/service and tenant where known. Keep unknown allocations explicit.
|
||
Count host and cluster representations of the same CPUs only once.
|
||
|
||
Reconcile scheduler-effective requests against the Kubernetes API, including
|
||
scheduled nonterminal and terminating pods, pending demand, init/sidecar/pod
|
||
resources and overhead. Explain differences from dashboard/collector totals;
|
||
repair the Forgejo omission and absent host CPU signal in their owning sources.
|
||
Account for zero-request containers, host/control-plane demand and data gaps.
|
||
Capture CPU usage, requests/limits, throttling and contention, memory working
|
||
set/peaks, OOM/restarts, storage exposure and workload activity where available.
|
||
Preserve query/collector revisions, units, timestamps and coverage. Reuse
|
||
RCLUSTER-WP-0014 and return the release-specific contract to STATE-WP-0091.
|
||
|
||
Done when a repeatable observation reconciles current allocation totals,
|
||
identifies missing/stale signals without reporting them as zero demand, and
|
||
updates the existing portfolio evidence with accountable owner references.
|
||
|
||
**Done 2026-09-14.** Repeatable command:
|
||
`python3 scripts/reconcile_railiance_allocation.py`. Evidence:
|
||
`docs/evidence/2026-09-14-allocation-reconcile.{json,md}` plus
|
||
`docs/evidence/namespace-ownership.yaml`. Live totals: 4000m capacity,
|
||
3985m scheduled, 25m pending, 15m residual (not a guarantee). Host and
|
||
cluster CPUs counted once. Unknown namespaces listed, then filled for
|
||
`approval-engine` and `vergabe-demo-company`. Zero-request workloads named.
|
||
STATE-WP-0091 preflight refuses (15m remaining). Forgejo omission remains a
|
||
Prometheus/dashboard gap, not a Kubernetes API gap. RESOURCE-WP-0007 already
|
||
updated reef-railiance-k3s owner evidence.
|
||
|
||
## Measure the Vergabe prototype under representative pilot use
|
||
|
||
```task
|
||
id: CUST-WP-0071-T02
|
||
status: wait
|
||
priority: high
|
||
assignee: the-custodian
|
||
depends_on: [CUST-WP-0071-T01, RAPPS-WP-0014-T03]
|
||
blocking_reason: "Await RAPPS-WP-0014-T03 (blocked, needs_human): populated two-user/document workflow run by Bernd and the company contact, giving representative load for p95 <= 2s / zero failures."
|
||
state_hub_task_id: "82192370-2fd7-5363-88d2-3c67889d3d68"
|
||
```
|
||
|
||
VERGABE-WP-0019/RAPPS-WP-0014 supply exact release, tenant/user assumptions,
|
||
runtime configuration and deployment. Observe the accepted 60m prototype and
|
||
exercise an isolated synthetic-data fixture with representative concurrent
|
||
users: login, list/search/detail, tender/lot/task edits and document upload/
|
||
download. Include startup, idle periods, ordinary usage, bursts and restart.
|
||
Measure API/UI latency percentiles and errors alongside CPU/memory, throttling,
|
||
database demand and request volume. Distinguish container CPU from incremental
|
||
database/shared-service demand. No benchmark writes to existing customer data.
|
||
|
||
Record workload sizes, concurrency, hardware/image/workers, duration, coverage
|
||
and limitations so results are reproducible. Founder acceptance target, 2026-09-28: with two simultaneous users, p95 of
|
||
ordinary operations must be at most 2 seconds and there must be no failed
|
||
operations. Report document transfer time separately. This resolves the target
|
||
choice; obtain representative evidence before claiming adequacy. Recommend request/limit and
|
||
memory values with a stated margin and revisit trigger; label an incomplete
|
||
pilot sample provisional. A successful smoke test alone is not sizing proof.
|
||
|
||
## Review ecosystem allocations and reserve operating room
|
||
|
||
```task
|
||
id: CUST-WP-0071-T03
|
||
status: wait
|
||
priority: high
|
||
assignee: the-custodian
|
||
depends_on: [CUST-WP-0071-T01, CUST-WP-0071-T02]
|
||
blocking_reason: "Await CUST-WP-0071-T02 representative sample; Forgejo/runner demand and zero-request workload review need owner input (railiance-cluster, railiance-forge, resource-control)."
|
||
state_hub_task_id: "6dc67558-eb1e-5bb6-a667-986f884dd495"
|
||
```
|
||
|
||
Resource-control and the owning workload/cluster repositories compare actual
|
||
demand with reservations. Prioritize Knative/Kourier's large reservations,
|
||
Forgejo's higher observed demand, zero-request workloads and the stale probe/
|
||
database-drill allocations. Review each service's startup, burst, autoscaling,
|
||
latency, availability and storage requirements before proposing changes.
|
||
|
||
Publish proposed per-workload requests/limits and a node-level budget including
|
||
host operations, pilot growth, deployment surges/migration hooks and admitted
|
||
factory/build execution. Declare margin and concurrency assumptions, current
|
||
constraints, owners and rollback. If demand plus margin does not fit, compare
|
||
capacity options through resource-control using current provider evidence.
|
||
STATE-WP-0091 retains its implementation of release preflight; this task consumes
|
||
and supplies that contract rather than replacing its workplan.
|
||
|
||
Done when the recommended allocation reconciles to admitted capacity and every
|
||
remaining change has a live receiving owner/task. This plan does not grant
|
||
unrelated deletion, procurement or automatic resource changes.
|
||
|
||
## Apply accepted sizing changes and verify their effects
|
||
|
||
```task
|
||
id: CUST-WP-0071-T04
|
||
status: wait
|
||
priority: medium
|
||
assignee: the-custodian
|
||
depends_on: [CUST-WP-0071-T03]
|
||
blocking_reason: "Await reviewed allocation proposals and applicable workload-owner authority."
|
||
state_hub_task_id: "944f98e2-7e42-5894-925b-752423f3a1cc"
|
||
```
|
||
|
||
Apply accepted changes in the existing owner manifests and deployment paths.
|
||
Use staged changes and retain exact prior configuration and rollback receipts.
|
||
Compare before/after traffic, latency/errors, resource consumption, throttling,
|
||
OOM/restarts and release headroom over a representative observation window.
|
||
Verify startup and a subsequent release at the new allocation. Revert or revise
|
||
changes that violate the agreed product/service targets. Update resource-control
|
||
and task evidence with deployed revisions, measured outcome and residual risks.
|
||
|
||
Do not close based only on successful scheduling or smaller requests. Acceptance
|
||
requires useful operation and a reconciled budget; unimplemented proposals or
|
||
unresolved measurement gaps remain live work records with owners.
|
||
|
||
## Establish the regular weekly workload-demand and allocation assessment
|
||
|
||
```task
|
||
id: CUST-WP-0071-T05
|
||
status: wait
|
||
priority: medium
|
||
assignee: the-custodian
|
||
depends_on: [CUST-WP-0071-T04]
|
||
blocking_reason: "Await the repeatable observation and accepted sizing/review procedure."
|
||
state_hub_task_id: "20ddd4d7-d85d-5b95-bb0a-7b0ca8af76ca"
|
||
```
|
||
|
||
As the final step, install a weekly assessment on an existing durable Railiance
|
||
executor with scoped read access. Record the executor owner, schedule and time
|
||
zone; proposed cadence is Monday 08:00 Europe/Berlin. Use the existing telemetry
|
||
and resource-control surfaces, not a workstation reminder or a parallel
|
||
monitoring service. Persist weekly aggregates so comparisons survive the current
|
||
seven-day raw-metric retention, and compare the last seven days with prior weeks.
|
||
|
||
Each run records per-workload demand and activity, CPU/memory p95 and peaks,
|
||
requests/limits, throttling/latency/errors/restarts, pending demand, node and
|
||
release margin, coverage/freshness, ownership and week-on-week changes. Produce
|
||
explicit keep/increase/decrease/investigate recommendations with evidence and
|
||
confidence. Deduplicate open recommendations into live owner work records and
|
||
review their results in the following run. Resource allocation adapts through
|
||
the accepted owner change procedure; no unrestricted automatic resizing.
|
||
|
||
Name the weekly review owner and persist delivery/acknowledgment through an
|
||
accepted internal channel. Define and exercise stale/failed/missed-run handling,
|
||
retry/idempotency and an accountable recovery route. Preserve scheduler and
|
||
assessment configuration in source, with non-secret execution receipts.
|
||
|
||
Done when the durable schedule is installed, the first unattended scheduled
|
||
assessment has produced its retained report and owner receipt, missed-run
|
||
handling is demonstrated, and ownership for the ongoing weekly operation is
|
||
registered before this workplan finishes. Hand off residual recommendations as
|
||
live records; the recurring assessment continues after this plan is finished.
|
||
|
||
|
||
### Scheduling evidence / live handoff for T01 — 2026-09-14
|
||
|
||
During SECRETS-WP-0010-T03, repeated `sso/keycape-factor-renewer` Jobs requesting
|
||
10m CPU failed scheduling at full requested CPU. The provider credential expired
|
||
and KeyCape lost readiness, blocking operator authentication. Informed Decision
|
||
used 1m CPU; reducing its request from 20m to 5m (limit remains 500m) freed 15m.
|
||
Scheduled Job keycape-factor-renewer-29822490 completed and self-revoked, ESO
|
||
updated the mounted credential, and KeyCape recovered. This is immediate recovery,
|
||
not a fleet sizing conclusion. T01 must include recurring maintenance-job demand
|
||
and reliable scheduling headroom, not only resident pod allocations. Evidence:
|
||
informed-decision/docs/evidence/2026-09-14-keycape-renewal-capacity-recovery.json.
|
||
|
||
## September 28 bounded completion review
|
||
|
||
The founder asks to finish with minimal additional tasks, workplans and
|
||
functionality. Keep all remaining work in T02–T05; do not spawn a monitoring
|
||
service, benchmark framework or replacement coordination plan.
|
||
|
||
Evidence: `docs/evidence/2026-09-28-sizing-review.md`, allocation reconcile,
|
||
retained cluster observation, source revisions and exact seven-day PromQL
|
||
responses alongside it. T01 refreshed: 3420m requested / 4000m, no pending
|
||
requests, 580m reservation residual; instantaneous node CPU was 3963m, so this
|
||
is not spare processing capacity. No unsupported pod accounting features were
|
||
present in this snapshot. Namespace ownership refreshed.
|
||
|
||
T02 is now in progress: the exact deployed pilot and seven days of measurements
|
||
are recorded. CPU p95 0.52m, sampled peak 20.10m, memory peak 191.14Mi; retain
|
||
60m/256Mi provisionally. Representative two-user activity, response/error
|
||
acceptance and incremental database attribution remain unproven. Use existing
|
||
RAPPS-WP-0014-T03 fixture/acceptance work; do not duplicate its recovery scope.
|
||
|
||
T03 is now in progress: retain current pilot/Knative allocations, investigate
|
||
Forgejo/runner demand (namespace CPU p95 1454m), and account for twelve
|
||
zero-request workloads. RAIL-KNATIVE-WP-0002 and RAIL-EN-WP-0002 are finished;
|
||
their declaration/deployment work must not be repeated. No new resource change
|
||
is proposed from the incomplete sample. T04 can verify a justified keep decision;
|
||
it must not create an unnecessary resize. Its useful-operation acceptance stays.
|
||
|
||
T05 still requires a durable activity-core schedule, retained report, owner
|
||
receipt and missed-run recovery proof. Monday 08:00 Europe/Berlin remains the
|
||
proposed cadence. None is represented as installed by this session. Existing
|
||
T02–T05 retain the remaining evidence and execution, with no new work records.
|
||
|
||
September 28 follow-up: the founder selected the two-user p95 ≤ 2 seconds,
|
||
zero-failed-operations target (document transfer excluded from that latency
|
||
threshold). The current seven-day telemetry remains provisional until the
|
||
representative workflow runs. This is an acceptance criterion, not a claim that
|
||
it has passed. Keep the run and its evidence under existing T02.
|
||
|
||
## Blocked — requirements on other repos (2026-09-28)
|
||
|
||
The bounded review is complete: T01 done, seven-day telemetry recorded, and the
|
||
provisional keep decision (60m/256Mi) is evidenced. No further step is
|
||
implementable here; workplan set to `blocked`. Requirements:
|
||
|
||
| Owner | Requirement | Gates |
|
||
|-------|-------------|-------|
|
||
| railiance-apps (RAPPS-WP-0014-T03, needs_human) | Two-user/document acceptance run on the deployed pilot with Bernd and the company contact | T02 |
|
||
| vergabe-teilnahme (VERGABE-WP-0019, blocked) | Release/tenant assumptions and fixture for the representative run | T02 |
|
||
| resource-control, railiance-cluster, railiance-forge | Owner review of Forgejo/runner demand (namespace CPU p95 1454m) and twelve zero-request workloads | T03 |
|
||
| activity-core | Durable weekly schedule (Mon 08:00 Europe/Berlin), retained report, owner receipt, missed-run handling | T05 (after T04) |
|
||
|
||
Resume when RAPPS-WP-0014-T03 records the acceptance run; T03 → T04 → T05 follow.
|