prj-state-hub-retirement/DECISIONS.md
tegwick d9867bd783 Pin the CoulombCore decommission at 2026-08-31; find the registry dependency
Operator set the date: CoulombCore retires by end of August. Eleven days, which
decides T03 — absorbing Core Hub directly into hub-core depends on HUB-WP-0004,
which still has not decided whether hub-core becomes a runtime at all. The
project recommends the interim move (the pattern issue-core used a week ago);
the choice stays core-hub's.

Resolving the public service names against the two hosts surfaced something no
plan contained: gitea.coulomb.social is a live container registry on CoulombCore,
and two railiance01 workloads pull images from it — reuse-surface, and State
Hub's own alembic-init migration job. Nothing breaks on the day, because running
pods already hold their images; it breaks at the next restart or reschedule as
ImagePullBackOff, at a time chosen by circumstance.

Added as T07 and routed to railiance-platform. forgejo.coulomb.social already
runs on railiance01, so the work is retag, push, update manifest.

The generalisation goes into T06's proposed G-GEN gate: the inventory was built
from tunnels and workplans and both missed this. A host is not free of dependents
because nothing tunnels to it — it is free when nothing pulls, resolves or
authenticates against it either.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 08:47:23 +02:00

103 lines
5.3 KiB
Markdown

# Decision Log
_Auto-generated by the Custodian State Hub._
## Restore the previously-live seven-action tenant-engine image rather than rebuild
**Date:** 2026-08-16
**Decided by:** grok
The four-action regression is a pin problem, not a policy-authoring problem. Image sha256:9320df39 was CI-built from e9911eb, previously live 2026-08-11, and confirmed by tenant-engine on 2026-08-13. The tenant-engine policy has not changed since that commit. A new image would re-bake unrelated later packages and would not be the artifact already verified. TEN-WP-0006 guardrail actions stay out of this restore.
---
## Use the guardrail read/write split now: flex-auth may read, may not set
**Date:** 2026-08-16
**Decided by:** grok
TEN-WP-0006 split the actions so a PDP can read ceilings without changing them, and their tests call the read as actor=flex-auth. Leaving both actions on the single tenant-engine subject would keep the seam unused and leave the intended reader failing unknown_subject. This is a real difference: flex-auth set is denied action_not_granted. An ops write subject is out of scope; writes stay on the existing tenant-engine identity.
---
## Generation 2 was retired correctly; the correction matters more than the finding
**2026-08-20.** `SHR-WP-0002`'s first draft asserted that `inter-hub` had
"retired itself by attrition", inferring it from three true observations: no
`~/inter-hub` repo, a dead CoulombCore endpoint, and a railiance01 Deployment
scaled `0/0`.
The inference was wrong. `core-hub/workplans/archived/260708-CORE-WP-0007-haskell-retirement.md`
records that `CORE-WP-0005` closed the production cutover gates on 2026-07-03 —
`hub.coulomb.social` serving Core Hub, Inter-Hub compatibility, staging import
and dual-run smokes all closed — and that Haskell/IHP retirement followed on
2026-07-08 after a stabilization window and explicit operator approval to retire
the Inter-Hub rollback deployment. The `0/0` Deployment **is** that rollback
standby, at its designed end state.
Recorded because the failure mode generalises: **live documents described a
retirement that had already happened, and the completed evidence was in an
archived workplan.** `core-hub/SCOPE.md` still lists cutover planning as in
scope. A reader checking current files would reach the wrong conclusion, as this
project did. Retirement evidence needs to be discoverable from the live record,
not only from the archive — a requirement that applies directly to the State Hub
retirement this project is planning.
The real finding survived the correction and sharpened: **the gen-3 runtime
serves from the host being decommissioned.**
## CoulombCore decommission date: end of August 2026
**2026-08-20, operator.** CoulombCore is to be retired **by 2026-08-31**. This
project treats that as a hard external constraint, not a target.
Eleven days. That materially decides `SHR-WP-0002-T03`: absorbing Core Hub
directly into `hub-core` depends on `CORE-WP-0010``HUB-WP-0004`, the latter
still holding an open decision on whether hub-core becomes a runtime at all.
Finishing an architecture decision *and* a production migration inside eleven
days is not a plan, it is a hope. **The project's recommendation to core-hub is
therefore the interim move** (`rmgr rapp wrap` onto railiance01, absorb into
hub-core afterwards on a calm schedule) — the pattern `issue-core` used one week
earlier. The choice remains core-hub's; the reasoning is recorded either way.
### What is actually still on CoulombCore
Established by resolving the public service names and inspecting railiance01,
2026-08-20:
| Service | Status | Consequence of shutdown |
| --- | --- | --- |
| `hub.coulomb.social` → Core Hub | Production since 2026-07-03 | Loss of the gen-3 interaction framework. Tracked, `SHR-WP-0002-T03` |
| `gitea.coulomb.social` → container registry | Live (200) | **Not tracked anywhere before today.** See below |
`bao`, `forgejo`, `policy`, `risk` and `reuse` all resolve to railiance01
already.
### The registry dependency nobody had written down
`gitea.coulomb.social` is a container registry on CoulombCore, and **two
workloads running on railiance01 pull their images from it**:
- `reuse/Deployment/reuse-surface``gitea.coulomb.social/coulomb/reuse-surface:e3ae22e`
- `state-hub/Job/state-hub-alembic-init``gitea.coulomb.social/coulomb/state-hub:f2e042a`
Checked across Deployments, StatefulSets, DaemonSets, Jobs and CronJobs; those
two are the whole set.
This fails in the most inconvenient way available. Running pods survive
decommission because their images are already pulled locally — **so nothing
breaks on the day**. The failure arrives at the next restart, reschedule, node
reboot or scale-up, as `ImagePullBackOff`, at a moment chosen by circumstance
rather than by us. And one of the two is **State Hub's own migration job** — the
service this entire project exists to retire in an orderly way, unable to run its
schema migration.
`forgejo.coulomb.social` already runs on railiance01, so the destination exists
and the work is retag, push, update manifest. Routed to `railiance-platform`.
**Generalisation worth keeping:** the decommission inventory was built from
tunnels and workplans, and both missed this. A host is not free of dependents
because nothing *tunnels* to it — it is free when nothing *pulls, resolves or
authenticates* against it either.