Advance supervised agent records and close verified Secret annotation guard
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 3s
Python Tests / pytest (push) Successful in 25s

This commit is contained in:
codex 2026-09-28 18:15:27 +02:00
parent b2f6713721
commit db91818e84
44 changed files with 6868 additions and 54 deletions

View file

@ -9,7 +9,7 @@ flavor: planning
owner: the-custodian
topic_slug: custodian
created: "2026-09-11"
updated: "2026-09-14"
updated: "2026-09-28"
related: [STATE-WP-0091, RCLUSTER-WP-0014, RESOURCE-WP-0003, RAPP-TELEMETRY-WP-0001, RAIL-FAB-WP-0028, RAPPS-WP-0014, VERGABE-WP-0019, HFACT-WP-0001]
state_hub_workstream_id: "2249bddb-7524-5add-bd5c-c4163a6ca0f3"
---
@ -18,7 +18,7 @@ state_hub_workstream_id: "2249bddb-7524-5add-bd5c-c4163a6ca0f3"
User instruction, 2026-09-11: persist and register this work for later follow-up,
then continue the invited Vergabe pilot at an explicitly accepted 60m CPU
request. This ready workplan is not a prerequisite to deploying that prototype.
request. This workplan is not a prerequisite to deploying that prototype.
It is not an assertion that 60m or the inherited 100m is a measured requirement.
The objective is an explainable, repeatable allocation process across Railiance
@ -94,11 +94,10 @@ updated reef-railiance-k3s owner evidence.
```task
id: CUST-WP-0071-T02
status: wait
status: progress
priority: high
assignee: the-custodian
depends_on: [CUST-WP-0071-T01]
blocking_reason: "Await reliable measurements and the current invited-pilot deployment or an equivalent isolated fixture."
state_hub_task_id: "82192370-2fd7-5363-88d2-3c67889d3d68"
```
@ -112,8 +111,10 @@ database demand and request volume. Distinguish container CPU from incremental
database/shared-service demand. No benchmark writes to existing customer data.
Record workload sizes, concurrency, hardware/image/workers, duration, coverage
and limitations so results are reproducible. Agree response-time/error targets
with the product owner before claiming adequacy. Recommend request/limit and
and limitations so results are reproducible. Founder acceptance target, 2026-09-28: with two simultaneous users, p95 of
ordinary operations must be at most 2 seconds and there must be no failed
operations. Report document transfer time separately. This resolves the target
choice; obtain representative evidence before claiming adequacy. Recommend request/limit and
memory values with a stated margin and revisit trigger; label an incomplete
pilot sample provisional. A successful smoke test alone is not sizing proof.
@ -121,11 +122,10 @@ pilot sample provisional. A successful smoke test alone is not sizing proof.
```task
id: CUST-WP-0071-T03
status: wait
status: progress
priority: high
assignee: the-custodian
depends_on: [CUST-WP-0071-T01, CUST-WP-0071-T02]
blocking_reason: "Await reconciled demand evidence and a measured pilot recommendation."
state_hub_task_id: "6dc67558-eb1e-5bb6-a667-986f884dd495"
```
@ -221,3 +221,40 @@ updated the mounted credential, and KeyCape recovered. This is immediate recover
not a fleet sizing conclusion. T01 must include recurring maintenance-job demand
and reliable scheduling headroom, not only resident pod allocations. Evidence:
informed-decision/docs/evidence/2026-09-14-keycape-renewal-capacity-recovery.json.
## September 28 bounded completion review
The founder asks to finish with minimal additional tasks, workplans and
functionality. Keep all remaining work in T02–T05; do not spawn a monitoring
service, benchmark framework or replacement coordination plan.
Evidence: `docs/evidence/2026-09-28-sizing-review.md`, allocation reconcile,
retained cluster observation, source revisions and exact seven-day PromQL
responses alongside it. T01 refreshed: 3420m requested / 4000m, no pending
requests, 580m reservation residual; instantaneous node CPU was 3963m, so this
is not spare processing capacity. No unsupported pod accounting features were
present in this snapshot. Namespace ownership refreshed.
T02 is now in progress: the exact deployed pilot and seven days of measurements
are recorded. CPU p95 0.52m, sampled peak 20.10m, memory peak 191.14Mi; retain
60m/256Mi provisionally. Representative two-user activity, response/error
acceptance and incremental database attribution remain unproven. Use existing
RAPPS-WP-0014-T03 fixture/acceptance work; do not duplicate its recovery scope.
T03 is now in progress: retain current pilot/Knative allocations, investigate
Forgejo/runner demand (namespace CPU p95 1454m), and account for twelve
zero-request workloads. RAIL-KNATIVE-WP-0002 and RAIL-EN-WP-0002 are finished;
their declaration/deployment work must not be repeated. No new resource change
is proposed from the incomplete sample. T04 can verify a justified keep decision;
it must not create an unnecessary resize. Its useful-operation acceptance stays.
T05 still requires a durable activity-core schedule, retained report, owner
receipt and missed-run recovery proof. Monday 08:00 Europe/Berlin remains the
proposed cadence. None is represented as installed by this session. Existing
T02–T05 retain the remaining evidence and execution, with no new work records.
September 28 follow-up: the founder selected the two-user p95 ≤ 2 seconds,
zero-failed-operations target (document transfer excluded from that latency
threshold). The current seven-day telemetry remains provisional until the
representative workflow runs. This is an acceptance criterion, not a claim that
it has passed. Keep the run and its evidence under existing T02.