the-custodian/docs/evidence/2026-09-28-sizing-review.md
codex db91818e84
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 3s
Python Tests / pytest (push) Successful in 25s
Advance supervised agent records and close verified Secret annotation guard
2026-09-28 18:15:27 +02:00

94 lines
5.2 KiB
Markdown

# CUST-WP-0071 — allocation review, 2026-09-28
## Recommendation
Keep the accepted Vergabe pilot at 60m CPU / 256Mi memory requests and
1 CPU / 1Gi limits. Keep the already reduced Knative allocations. Propose no
new live allocation change from this sample. This is a provisional pilot
recommendation, not acceptance of representative user performance.
The node has reservation room but little measured processing room at the
snapshot: 3,420m requested of 4,000m, zero pending requests, **3,963m observed
CPU** and 12,075,401,216 bytes observed memory. The 580m reservation residual
must not be called spare processing capacity. Avoid admitting concurrent builds
on the strength of that residual alone.
## Evidence and coverage
- Node verified as `92.205.62.239`, k3s v1.35.1+k3s1.
- `2026-09-28-allocation-reconcile.json` and `.md`: existing collector plus
updated namespace owners; same hardware counted once. State Hub release
preflight passes at this observation, not as a standing admission grant.
- `2026-09-28-cluster-observation.json`: retained source observation, including
instantaneous demand. Collector revisions are recorded in the companion
provenance file. No Secret objects were queried.
- Collector limitation checked against active pods: no pod-level resources,
overhead or restartable init containers were present. Its simplified init
accounting therefore did not encounter these unsupported features today.
This does not certify its accounting for future workloads.
- `2026-09-28-allocation-telemetry.json`: exact PromQL and responses for seven
days, evaluated at five-minute resolution, namespace aggregates. The five
namespaces below each have 2,016 CPU evaluation points. These are evaluation
points, not proof of complete raw scrape coverage or representative traffic.
| Namespace | CPU p95 (m) | CPU peak (m) | Memory peak (MiB) |
|---|---:|---:|---:|
| vergabe-demo-company | 0.52 | 20.10 | 191.14 |
| knative-serving | 7.75 | 8.54 | 273.54 |
| kourier-system | 3.88 | 4.06 | 59.00 |
| forgejo | 1454.35 | 2208.27 | 4511.85 |
| databases | 238.08 | 341.43 | 1155.20 |
Namespace peaks need not be simultaneous and cannot be added into a node peak.
Forgejo includes its runner; the databases row is shared demand, not an estimate
of Vergabe's incremental database cost. Namespace aggregates can hide missing
individual pod series. Peak means the maximum sampled value, not every burst.
Vergabe's throttled-period ratio is 0.000106 (about 0.0106%); the restart counter
query reports zero increase for the namespaces above. Counter evidence is not
proof that deleted or recreated pods never restarted. Forgejo's throttle ratio
is absent; it is not zero. Host CPU history, Vergabe HTTP request/latency series,
and the Forgejo namespace CPU-request recording remain absent in these queries.
Live Vergabe image:
`forgejo.coulomb.social/coulomb/vergabe-teilnahme@sha256:a26444f59c259698159c69ccb96f73dc648a261ece4c86bb2037a9d977870d91`.
One ready replica, 60m/1 CPU and 256Mi/1Gi. No customer data was written.
## Existing work replaces duplicate changes
`RAIL-KNATIVE-WP-0002` finished on September 27. Its T03 verifies that the
railiance-cluster installer now preserves the reduced requests. The stale inbox
warning about the installer reverting them is superseded by that file evidence.
`RAIL-EN-WP-0002` is also finished: ArgoCD resources are already declared and
applied. Neither warrants another task here.
T03 retains review of Forgejo/runner demand, twelve zero-request workloads and
the distinction between release reservations and real CPU contention. There is
no approved new sizing change for T04 to deploy today. T04 must verify the
accepted final recommendation, including an explicit keep decision if supported,
rather than manufacture a resize to satisfy its title.
## Smallest remaining execution
Use the existing T02 for one bounded synthetic pilot session, with product-owner
response/error targets, two concurrent users, document round-trip, edits and
restart evidence. Reuse the invited-pilot fixture; do not create a load-test
service. Link the existing RAPPS-WP-0014-T03 acceptance work rather than duplicate
its data-recovery tasks. The seven-day idle/light-use sample is useful input but
does not replace this session.
T03/T04 then resolve a single allocation recommendation and verify only accepted
changes. Preserve a 160m reservation envelope for the current State Hub
100m API + 10m MCP + 50m migration requests, and account separately for scheduled
maintenance and competing releases. This is a review assumption, not a global
admission policy or evidence that the node has 160m spare execution capacity.
T05 reuses activity-core's durable scheduler: Monday 08:00 Europe/Berlin,
Custodian review ownership, retained reports and State Hub progress delivery.
Retain the exact queries with every report; missing signals stay unknown.
The minimum first report covers keep/investigate decisions above, allocation,
sample coverage and links to these existing tasks. The first scheduled run,
acknowledgment and missed-run/recovery evidence are still required. No weekly
schedule was installed by this review; no unattended receipt is claimed.
No new task, workplan, intake, monitoring service or resource mutation was made.