railiance-platform/docs/argocd-coulombcore-retirement.md
codex debf981097
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 3s
Close verified incident task and finish local workplan loose ends
Assistant: codex
Assistant-Model: gpt-6-astra
Assistant-Session: 01a0e3c3-621b-7350-9f77-50a8d3ee7657
2026-09-27 19:04:54 +02:00

3.8 KiB

Coulombcore ArgoCD retirement preparation

Existing owner task: RPF-WP-0044-T08. Prepared September 27, 2026; blocked pending a readable live inventory. No controller, application or workload was changed. This document is the phase C preparation within the existing workplan.

The host SSH lane works. Kubernetes returned Unauthorized for both the normal kubectl context and sudo -n k3s kubectl --kubeconfig /etc/rancher/k3s/k3s.yaml get nodes. The infrastructure/cluster owner must restore accepted read access; do not rotate cluster credentials or restart k3s as an incidental inventory fix. Current workload ownership cannot be inferred from the old manifests.

Inventory and decision inputs

Capture node/cluster identity, ArgoCD deployments/statefulsets and installed version, Applications and AppProjects, each Application's destination, pinned revision, tracked resource set, hooks, automated sync and finalizers. Read only repository Secret metadata, never data. Identify external cluster destinations: an old controller can still manage a remote cluster. Search platform and tenant source for references to argocd/applications/ and argocd/bootstrap/, including Make entry points, and match each live application to its accepting owner.

The railiance01 root uses argocd/railiance01/applications; the legacy root uses argocd/applications. Do not remove the legacy tree while the old controller can reconcile it with prune enabled. Keep evidence of each workload's current replicas, readiness and image before any retirement execution.

Ordered execution after inventory and owner acceptance

  1. Have the cluster owner pin the installed railiance01 ArgoCD version and reviewed resource requests in its canonical source. Inspect current requests; the original phase A BestEffort observation is historical, not a fresh check.
  2. Classify each old tracked resource: already adopted on railiance01, retained on coulombcore with another owner, or separately approved for retirement. An application name match does not prove matching cluster/resource identity.
  3. Freeze the old root and child reconciliation through the cluster owner's reviewed procedure. Confirm no running sync operation and no second writer. Preserve the old Application specs and controller configuration in protected recovery storage; repository credentials stay under existing custody.
  4. Detach only inventory-approved Application objects without cascading workload deletion. Review resource finalizers first; never delete the ArgoCD namespace or CRDs as a shortcut. Verify every retained workload is still healthy and each replacement owner can reconcile its accepted resource set.
  5. Disable the old controller through its actual installation owner. Prove it no longer reconciles and that no unrelated system uses its repository/auth resources. Revoke retired credentials only through their custody owners.
  6. Once the old controller is inert, remove the legacy source directories and retire or repoint their callers in one reviewed platform change. Render the railiance01 bootstrap and children and verify no legacy reference remains.

Stop for missing inventory, ambiguous tracking, active operations, unknown finalizers, missing acceptance or degraded retained workloads. Before detachment, rollback restores the recorded sync configuration. After replacement ownership, keep the old controller stopped until the replacement writer is explicitly suspended; rollback must never run two reconcilers against one workload. Package or data deletion needs its own exact approved disposition.

Completion evidence must include the inventory, accepting owners, source/caller changes, controller shutdown and retained-workload checks. T08's planning closure still requires the live read-only inventory. Execution remains with the existing cluster/infra owners; no new task or workplan is created here.