activity-core/workplans/ACTIVITY-WP-0020-weekly-forgejo-package-prune.md
codex 1fa9dc18ab
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 1s
fix(workplans): adopt ADR-007 derived identifiers for unregistered records
These workplans exist only in the retired local hub. Their random pre-ADR-007
identifiers are refused by C-06 as stale references, so they cannot be
registered. Deriving from the canonical record id takes no identity from
anything: central does not hold them and the old ids die with the cache.

Records central already holds were deliberately left untouched.

Refs CUST-WP-0068-T06

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 2583210@bnt-lap001
Assistant-Session: f2bff2d5-e9b2-4338-92ca-10282a927006
2026-08-25 20:05:21 +02:00

15 KiB

id type title domain repo status owner topic_slug created updated state_hub_workstream_id
ACTIVITY-WP-0020 workplan Weekly Forgejo package prune (retain 3 versions) infotech activity-core finished codex activity-core 2026-07-12 2026-07-21 5e74cce7-9ffd-5f9a-89e9-17faf55c1ba9

Weekly Forgejo package prune (retain 3 versions)

Problem

Forgejo forgejo dump backups are dominated by OCI container image layers under data/packages (~80% of blob zip; ~1.3 GB today). Growth is ~686 MiB (Jul 7) → ~1.0 GiB (Jul 9) → ~1.3 GiB (Jul 12) because CI pushes new image tags on every deploy and no automated prune runs yet (railiance-apps/docs/forgejo-package-registry.md).

Manual UI deletion does not scale. Prune must be policy-driven, dry-run first, and emit non-secret evidence for backup/ops gates.

Goal

Establish a retention depth of 3 versions per package (newest three kept; older versions deleted) across Forgejo package types that affect backup size, and run pruning weekly via activity-core on railiance01.

Retention policy (canonical)

Rule Value
Versions kept per package 3 (newest by created_at; tie-break on version string)
Package types in scope (phase 1) container (OCI), pypi, npm, generic
Protected versions Any tag/version referenced by production deploy manifests in railiance-apps Helm values / chart defaults (never delete even if outside top 3)
Org scope coulomb
Default mode dry-run; deletes only with explicit --apply
Evidence Non-secret summary posted to State Hub (forgejo_package_prune)

Phase 2 (out of scope here): Helm/Maven/Go registries when consumer inventory proves need.

Architecture

activity-core (weekly cron)
  → shell context resolver: forgejo_package_prune
    → railiance-platform: tools/cmd/forgejo-package-prune --apply
      → Forgejo API DELETE /api/v1/packages/{owner}/{type}/{name}/{version}
  → State Hub progress evidence (counts, bytes estimate, errors; no tokens)

Credential routing: Forgejo PAT with read:package + write:package from OpenBao via warden route find "forgejo package" — never stored in Git, workplans, or activity-definition files. Runtime token file path only.

Schedule: Sunday 03:30 UTC — after daily make forgejo-backup cron (02:15 UTC) so a fresh backup exists before destructive prune.

Scope boundary

In scope Out of scope
Prune script + Make target in railiance-platform Rewriting forgejo-backup chunk copy
Weekly ActivityDefinition in activity-core Gitea legacy registry cleanup
Retention docs in railiance-apps + railiance-platform SBOM signing, vulnerability scanning
Production-tag protection from Helm inventory Cross-org packages outside coulomb
Dry-run + first apply evidence Auto-prune git repos/ or repo-archive blobs

Dependencies

  • Forgejo API reachable (forgejo.coulomb.social)
  • Valid Forgejo operator/service PAT lane (OpenBao)
  • activity-core worker on railiance01 with ACTIVITY_DEFINITION_DIRS including repo definitions and network egress to Forgejo
  • make sync-schedules / worker startup schedule reconciliation

Tasks

T01 — Document retention policy and operator guardrails

id: ACTIVITY-WP-0020-T01
status: done
priority: high
state_hub_task_id: "d86bdea5-e41a-590c-99e1-1fe84d6760e3"

Add canonical retention prose to:

  • railiance-platform/docs/forgejo-package-prune.md (new)
  • railiance-apps/docs/forgejo-package-registry.md § Retention (replace "no auto-prune yet")

Cover: depth-3 rule, protected production tags, dry-run default, weekly activity-core schedule, evidence event type, and rollback posture (restore from Nextcloud dump if a needed tag was removed).

Acceptance: operator can answer "what gets deleted, what is protected, how often" without reading source.


T02 — Implement forgejo-package-prune command

id: ACTIVITY-WP-0020-T02
status: done
priority: high
state_hub_task_id: "ef1eb392-5338-54eb-b2fe-f68218cf33e8"

In railiance-platform:

  • tools/cmd/forgejo-package-prune — list packages for owner coulomb, group versions, compute deletions beyond newest 3, subtract protected tags from Helm inventory scan (railiance-apps/helm/*-values.yaml, railiance-apps/charts/*/values.yaml image tags).
  • Flags: --dry-run (default), --apply, --max-versions 3, --types container,pypi,npm,generic, --owner coulomb.
  • Auth: FORGEJO_TOKEN or OpenBao field via existing backup-lane pattern; preflight 401 → clear error.
  • Output: JSON summary on stdout for activity-core parsing (deleted, skipped_protected, would_delete, errors[]); human lines via railiance-print.sh.
  • Makefile targets: forgejo-package-prune-dry-run, forgejo-package-prune.

Acceptance: dry-run against production lists candidates; no deletes without --apply; unit or fixture tests for version sorting and protection logic.


T03 — Register shell resolver query forgejo_package_prune

id: ACTIVITY-WP-0020-T03
status: done
priority: high
state_hub_task_id: "9f4f6ffd-ed9e-52b2-a59a-7e148e8b5a96"

In activity-core, extend the shell context resolver (same pattern as reuse_surface_report_gaps in context_resolvers/reuse_surface.py):

  • Query: forgejo_package_prune
  • Params: prune_script (default ~/railiance-platform/tools/cmd/forgejo-package-prune), apply (bool, default false for first ship), max_versions, timeout_seconds
  • Invoke script with bounded timeout; parse JSON stdout into context.prune
  • Never capture or persist token env vars in context/working memory

Acceptance: resolver tests with mocked subprocess JSON; failure surfaces non-zero exit and stderr snippet (truncated, no secrets).


T04 — Add weekly ActivityDefinition

id: ACTIVITY-WP-0020-T04
status: done
priority: high
state_hub_task_id: "b80326c0-3ce0-571e-ab38-29125a75d3c5"

Add activity-definitions/weekly-forgejo-package-prune.md:

trigger:
  type: cron
  cron_expression: "30 3 * * 0"
  timezone: UTC
  misfire_policy: skip
context_sources:
  - type: shell
    query: forgejo_package_prune
    params:
      apply: true
      max_versions: 3
      evidence_sinks:
        - type: state-hub-progress
          event_type: forgejo_package_prune
          author: activity-core
    bind_to: context.prune

Ship enabled: false until dry-run evidence exists. No rules / task emission — observational prune + evidence only (same shape as weekly-legacy-meter-review).

Acceptance: definition parses via sync-definitions; schedule id derivable from make automation-inventory.


T05 — Dry-run verify, enable, and runbook

id: ACTIVITY-WP-0020-T05
status: done
priority: medium
state_hub_task_id: "f3b907de-0900-534f-872f-92950f260726"

2026-07-12: Implementation landed (script, resolver, activity-definition, runbook). enabled: false and apply: false until operator dry-run against production with FORGEJO_TOKEN and approves enable.

  1. Run resolver smoke with apply: false; capture candidate counts vs Jul 12 backup analysis (~800 MiB packages).
  2. Operator review of dry-run output; confirm protected tags (e.g. vergabe-teilnahme:064d295) never appear in would_delete.
  3. Flip enabled: true; make sync-schedules on railiance01 worker.
  4. Document in activity-core/docs/runbook.md alongside weekly-sbom-staleness and weekly-legacy-meter-review.

Acceptance: one successful dry-run workflow in Temporal; schedule visible in automation inventory; runbook lists cron, credential pointer, and rollback.

2026-07-16 dry-run: After fixing a version-enumeration bug in the prune tool (Forgejo has no /versions sub-endpoint; the list endpoint returns one entry per (name, version)railiance-platform@6f7ca31, incl. network-timeout retry), the dry-run is clean: 29 candidate deletions across 5 container packages (activity-core 16, state-hub 7, vergabe-teilnahme 3, reuse-surface 2, forgejo-actions-probe 1), 0 errors. Protected vergabe-teilnahme:064d295 and deployed reuse-surface:e3ae22e are correctly not in would_delete.

Blocked before enable/apply: protected_count was only 1 — the protection scan reads helm/*-values.yaml, but activity-core and state-hub (the two largest deletion sets) are not pinned there; their live tag can't be confirmed safe from this host. Enable/apply gated on T07 (extend protection to live deployed tags).


T07 — Extend protection to live deployed tags (pre-apply gate)

id: ACTIVITY-WP-0020-T07
status: done
priority: high
state_hub_task_id: "9da65260-c785-56db-864e-e8a965dc89ab"

The Helm-values protection scan under-covers: it only protected 1 tag, missing activity-core / state-hub (CI-deployed, not pinned in railiance-apps/helm/). Extend the protection source so every currently-running image tag is protected regardless of retention depth. Authoritative source is the live production cluster(s) where these images run (railiance01/Forgejo-consuming), not just Git values — query running Deployment/StatefulSet/DaemonSet container images across namespaces (or the ArgoCD app inventory) and union with the Helm scan.

Acceptance: a dry-run shows all live tags for activity-core, state-hub, and other production workloads present in skipped_protected (or absent from would_delete), verified against the cluster. Only then are T05 enable and T06 apply unblocked.

Progress 2026-07-16: live-tag protection implemented in the prune tool (railiance-platform@, collect_live_cluster_versions() — enumerates pod images via kubectl, protects forgejo…/coulomb/<name>:<tag>; on by default; unit-tested). Verified against the reachable production cluster (92.205.130.254): it protects state-hub f2e042a, vergabe-teilnahme, issue-core live tags — closing the state-hub gap. Multi-cluster caveat: activity-core (the largest delete set, 16) does not run on that cluster, so its live tag is not yet covered — the prune must aggregate live tags across all production clusters/contexts (or run where it sees them) before enable/apply. Enable/apply remains gated until every production cluster's live tags are protected in one dry-run.

Progress 2026-07-18: multi-cluster mechanism landed (railiance-platform@ca4e152): repeatable --live-images-file merges image lists exported from other clusters (unit-tested; missing file = WARN). Cluster census: railiance01 k3s runs activity-core as a locally-imported image (activity-core:railiance01-prod) — no live workload pulls the forgejo activity-core packages at all, defusing the original fear; its live forgejo tag is state-hub:main-1cf949b. The coulombcore cluster's live forgejo tags: issue-core:0.2.1, state-hub:f2e042a, vergabe-teilnahme:064d295. Both clusters' exports staged (workstation scratchpad; coulombcore list also copied to railiance01:~/.local/share/forgejo-prune/live-images-coulombcore.txt). Remaining (~5 min, agent-executable after one founder OIDC login — no PAT files, per ops-warden INTENT §7 / WARDEN-WP-0029): bao login -method=oidc -path=netkingdom role=forgejo-admin-workload-kv-read then warden access "forgejo admin pat" --field FORGEJO_ADMIN_TOKEN --exec -- tools/cmd/forgejo-package-prune --live-images-file <railiance01 export> — if all live tags land in protected/absent from would_delete, T07 acceptance is met and T05 enable + T06 apply unblock.

Done 2026-07-18: dry-run executed via the forgejo-admin-api-token warden lane (no PAT files): 34 candidates, 0 errors, no live tag in would_delete — state-hub main-1cf949b protected via the railiance01 live-images export; activity-core's 19 candidates confirmed to have no live registry consumer. Evidence: railiance-platform/docs/evidence/forgejo-package-prune-dryrun-20260718.json. T05 remaining: flip weekly-forgejo-package-prune to enabled (rides the next activity-core image deploy or the external-definitions configmap)

  • runbook entry; then T06 first apply.

T06 — First apply run and backup impact evidence

id: ACTIVITY-WP-0020-T06
status: done
priority: medium
state_hub_task_id: "99f11d58-9dee-5ee8-9b62-9d410fa0274a"

After T05 approval, run one manual make forgejo-package-prune with --apply (or trigger workflow once with apply: true). Then:

  1. Post forgejo_package_prune progress to State Hub (deleted counts, errors).
  2. Run make forgejo-backup-dry-run and record dump size delta vs pre-prune baseline in railiance-platform/docs/evidence/forgejo-package-prune-YYYYMMDD.json.
  3. Update railiance-apps/docs/forgejo-package-registry.md retention table with automated prune row.

Acceptance: evidence file exists; next backup blob zip measurably smaller or stable; no production deploy tag removed (smoke: crictl pull or helm image still present).


Registration

After approval, register in State Hub:

statehub fix-consistency --repo activity-core

Cross-repo implementation commits land in railiance-platform and railiance-apps; link them in T06 evidence — no separate hub workplan required unless scope expands.

Success criteria

  • Weekly Temporal schedule runs weekly-forgejo-package-prune on railiance01.
  • Each package under coulomb retains at most 3 versions unless protected.
  • State Hub receives forgejo_package_prune evidence every run.
  • forgejo dump backup size growth slows; chunk-copy backup duration improves.

Closeout 2026-07-21

T05 — done

  • Dry-run clean with multi-cluster live protection (38 candidates, 0 live overlap).
  • Definition enabled with apply: true, cron 30 3 * * 0 UTC.
  • Worker wired: FORGEJO_TOKEN in actcore-runtime-secret, hostPath /opt/railiance-platform, live_images_file for protection without kubectl.
  • Runbook updated.

T06 — done

  • First apply (workstation): 38 deleted, 0 errors, evidence railiance-platform/docs/evidence/forgejo-package-prune-apply-20260721.json.
  • State Hub progress forgejo_package_prune posted.
  • Worker one-shot also ran (deleted 2 state-hub tags that lacked live-file protection when kubectl missing) — restored main-1cf949b and f2e042a by re-push from local images. Hardened resolver + live-images file thereafter.

Residual ops notes

  • Refresh live-images-all.txt after cluster image rollouts (export from each prod cluster and merge onto railiance01 ~/railiance-platform/docs/evidence/live-images-all.txt).
  • Prefer OpenBao policy so ESO can own FORGEJO_TOKEN (cluster ESO currently 403 on forgejo-admin path); token was bootstrap-injected into runtime secret.
  • ESO ExternalSecret for forgejo-admin left deleted (permission denied).