Assistant: codex Assistant-Model: gpt-5.6-sol Assistant-Session: 01a053ff-1d6f-7fe2-ac1c-a6eb40a42a0c
20 KiB
| id | type | title | domain | repo | status | owner | topic_slug | created | updated | state_hub_workstream_id |
|---|---|---|---|---|---|---|---|---|---|---|
| SHR-WP-0002 | workplan | Generation 3 serves from the host being decommissioned | infotech | prj-state-hub-retirement | proposed | unassigned | state-hub-retirement | 2026-08-20 | 2026-08-20 | 53f42c0c-3a35-53b0-99c8-8eaee7b17354 |
SHR-WP-0002 — Predecessor generations and deployment reality
The problem in one sentence
Core Hub's production runtime — the generation-3 replacement, cut over
deliberately and correctly in July — serves from hub.coulomb.social, which
resolves to CoulombCore: the host being decommissioned. No plan in any
repository accounts for that.
Correction to this workplan's first draft
The first draft of SHR-WP-0002 (2026-08-20, superseded before any task ran)
claimed generation 2 had "retired itself by attrition". That was wrong, and
the evidence contradicting it is in core-hub's own archive:
CORE-WP-0005finished 2026-07-03:hub.coulomb.socialingress serves Core Hub; Inter-Hub compatibility, staging import, dual-run smokes, and production cutover gates all closed.CORE-WP-0007finished 2026-07-08: Haskell/IHP infrastructure retired, the production Inter-Hub repo renamed/archived,ihp-railiance-probearchived — after "a short post-cutover stabilization window" and explicit operator approval to retire the Inter-Hub rollback deployment.
The inter-hub Deployment scaled 0/0 on railiance01 is that rollback
deployment, held at zero exactly as the workplan describes. It is not a
lapse; it is the designed end state of a careful retirement.
The lesson is the opposite of the one first drafted: generation 2 was retired well. What nobody checked was whether the successor's runtime was standing on durable ground.
What was observed, 2026-08-20
| Generation | Position | Observed |
|---|---|---|
1 — state-hub |
Being retired by this project, under gates | Running; what the estate uses hourly |
2 — inter-hub |
Retired CORE-WP-0005/0007, Jul 2026 |
Correctly gone. Rollback deployment at 0/0 as designed |
3 — core-hub |
The replacement, in production since 2026-07-03 | hub.coulomb.social → 92.205.130.254 = CoulombCore. No core-hub namespace or Deployment on railiance01 |
— hub-core |
The surviving repository (GOAL §1) | A reusable Python package, not a running service |
Three consequences, none recorded anywhere:
- The gen-1 retirement is being planned onto a runtime that has no home.
G-HUB-RUNTIMEandG-CORE-ABSORBpresume somewhere to absorb traffic into. That somewhere is currently a host with a decommission date, and the gate model does not represent that date at all. core-hub/SCOPE.mdis stale in a way that hides this. It still places "retiring production Inter-Hub before migration and smoke evidence exists" out of scope and lists/api/v2compatibility and "cutover planning from Inter-Hub" as in scope — work its own archive shows finished in July. A reader cannot tell from the live documents that the cutover already happened.inventory/capabilities.yamldoes not mentioninter-hub. The migration was real, so this is a smaller gap than first thought — butG-DISPstill has no record of where gen-2 capabilities went, and the estate cannot currently answer "does anything still provide the unified operator surface?" without reading an archived workplan.
Why this belongs to the project and not to a child repo
No single repository can see it. core-hub knows its own migration plan;
llm-connect knows its consumers; ops-warden knows its access lanes. Only this
project holds the cross-repository authority, the disposition inventory, and the
gates — and G-DISP ("full disposition inventory"), G-CORE-ABSORB and
G-HUB-RUNTIME are all currently gated on assumptions that the observations
above contradict.
Per SCOPE.md, this workplan decides and inventories; it does not implement.
Every implementation task it identifies is routed to the owning repository by
identifier.
Tasks
id: SHR-WP-0002-T01
status: todo
priority: high
state_hub_task_id: "8c54ea9b-8a92-55bc-80ce-8425678f3efe"
Done 2026-08-20 — there was a proper cutover. CORE-WP-0005 closed the
production cutover gates on 2026-07-03; CORE-WP-0007 retired the Haskell/IHP
infrastructure by 2026-07-08 with operator approval and a stabilization window.
The 0/0 rollback deployment is the designed end state.
Remaining sub-task: record this in DECISIONS.md with the correction, so
the project's own record shows that the first reading was wrong and why. A
retirement project that misreads a completed retirement as an attrition should
keep the evidence trail for the next reader.
id: SHR-WP-0002-T02
status: todo
priority: high
state_hub_task_id: "fdae9c55-5a85-5b11-91ed-0b724129d887"
Extend the disposition inventory backwards to generation 2. Smaller than
first thought — the migration was real — but still unrecorded here. G-DISP demands
that every capability reach an explicit destination or retirement decision.
Applied only to State Hub, it lets gen-2 capabilities vanish unrecorded.
Produce inventory/inter-hub-disposition.yaml in the shape of the existing
capability rollup: for each capability gen 2 was to provide — domain hubs, shared
manifests, widgets, registries, events, unified operator surface — record whether
it is served today (and by what), planned (and by whom, under which
workplan), or dropped (and by whose decision).
The expected output is uncomfortable and worth having: a list of things the estate intended to have, does not have, and has not decided to do without.
id: SHR-WP-0002-T03
status: todo
priority: high
state_hub_task_id: "06119690-312f-5d14-aee0-2fc4dac1e89e"
This is now the workplan's centre of gravity: get the gen-3 runtime off
CoulombCore. G-HUB-RUNTIME and G-CORE-ABSORB presume a runtime to absorb
traffic into. That runtime is hub.coulomb.social on CoulombCore — production
since 2026-07-03 — plus a core-hub-staging tunnel to the same host. There is
no core-hub presence on railiance01, and hub-core is a package, not a service.
Decide and record: does core-hub's runtime move to railiance01 first, or does
consolidation into hub-core (GOAL §1) happen directly, skipping a migration to
a host that is itself scheduled to disappear? Sequence it explicitly against
the CoulombCore decommission date, which is the real constraint and is not
currently represented in the gate model.
Precedent worth reusing: issue-core was migrated off CoulombCore on 2026-08-19
via ISSUE-WP-0007 and a rapp-issue-core package. That is the pattern, and it
is one week old.
Routed to core-hub 2026-08-20; waiting on their answer. The vehicle is
CORE-WP-0010 (runtime absorption into hub-core, proposed, all tasks todo),
which depends on HUB-WP-0004 — itself proposed, and still carrying an open
decision about whether hub-core stays an importable library or becomes a
library plus a permanent thin host.
So the chain that would move production off CoulombCore bottoms out in an unanswered architecture question. That is a sequencing problem rather than an implementation one, which is why it sits here and not in a child repo — but the answer is core-hub's.
Two shapes were put to them:
- (a) Interim move — package Core Hub as-is onto railiance01, absorb into hub-core later on a calm schedule. Costs a migration that would otherwise not happen; buys independence from the decommission date.
- (b) Absorb directly — skip the interim host and let
CORE-WP-0010/HUB-WP-0004be the migration. Cheaper in total work; couples production continuity to finishing an architecture decision under a hard external deadline.
The project does not choose between them, but records the reasoning either way.
The one position taken: leaving it implicit is not acceptable, because the
decommission date is a real constraint currently represented in no plan —
including this project's own gate model, which T06 addresses.
Also asked: whether the decommission date changes HUB-WP-0004's open
library-vs-thin-host decision. Pinned 2026-08-20 by the operator: CoulombCore retires by 2026-08-31.
Eleven days. See DECISIONS.md. On that constraint the project recommends the
interim move — an architecture decision plus a production migration inside
eleven days is not a plan. The choice remains core-hub's.
id: SHR-WP-0002-T04
status: todo
priority: medium
state_hub_task_id: "3815d764-b484-5342-9c7d-2edacc02be13"
Reconcile stated scope against reality across participating repositories.
The lineage confusion is visible in the documents: core-hub/SCOPE.md guards a
predecessor that is not running; llm-connect/SCOPE.md names a consumer
relationship that cannot exist; ops-hub describes itself as an extension for
Core Hub, a repository this project intends to retire into hub-core.
Produce a table of every participating repository's stated position versus observed reality, and route each correction to its owner. Do not edit their files from here — the project's authority is the ledger, not the prose.
id: SHR-WP-0002-T05
status: todo
priority: medium
state_hub_task_id: "c2957182-4bb3-5aa9-9940-656956cfe52d"
Sweep the downstream lanes that outlived their subject. A retired generation leaves access and credential lanes behind, and they do not expire on their own.
Known instance: ops-warden's routing catalog carries inter-hub-bootstrap-ssh —
status: active, risk: high, last reviewed 2026-06-24 — a bootstrap SSH
envelope for a system that runs nowhere, whose runbook points at "the ops-hub
production activation lane tracked by CUST-WP-0049". ops-warden cannot retire
it alone: the lane may still serve ops-hub activation independently of inter-hub's
runtime, and only this project can see both sides.
Ask each participating repository for lanes, tunnels, credentials and scheduled jobs whose subject is a retired or lapsed generation. Route retirements to owners; record the sweep here so the next generation change has a checklist rather than a memory.
id: SHR-WP-0002-T06
status: todo
priority: medium
state_hub_task_id: "714ce904-bccd-5cd8-ae2e-a756f5c0bccc"
Add a generation-transition gate, so this cannot recur. The project has gates for retiring State Hub deliberately. It has none that would have caught a predecessor lapsing, or a successor running only on a host scheduled for shutdown.
Propose G-GEN for architecture/retirement-gates_v0.1.md: no generation is
considered superseded until its capabilities carry explicit dispositions, its
consumers are reconciled, its downstream lanes are swept, and its successor is
deployed somewhere that will still exist. Take it through the project's normal
decision route rather than adding it unilaterally.
id: SHR-WP-0002-T07
status: todo
priority: high
state_hub_task_id: "a89cf85d-d613-5434-92cb-11f850889b08"
The registry dependency, found 2026-08-20 and not previously tracked.
gitea.coulomb.social is a live container registry on CoulombCore, and two
railiance01 workloads pull images from it:
reuse/Deployment/reuse-surface and state-hub/Job/state-hub-alembic-init
(checked across Deployments, StatefulSets, DaemonSets, Jobs and CronJobs — that
is the complete set).
Nothing breaks on decommission day, because running pods already hold their
images. It breaks at the next restart, reschedule or scale-up, as
ImagePullBackOff.
The two are not equally urgent (corrected 2026-08-20 after checking each reference's liveness):
-
reuse-surface— worse than stale: a live runtime dependency. A running Deployment on the old registry, pinned to a 2026-07-07 commit and 22 commits behind main. Its CI switched to forgejo three hours after that commit and has published there ever since; the deployment was never repointed.Escalated 2026-08-20. At the deployed commit,
registry/federation/sources.yamlpoints atgitea.coulomb.socialfor 50 of its 61 sources (all 61 are forgejo at HEAD — migrated ind1c1313,RAIL-HO-WP-0006). So production is not merely running an old image: it is actively federating from the host being switched off, and that breaks on 2026-08-31 with no restart required. This is a different and worse failure mode than theImagePullBackOffrisk it was first filed under.It also inverts the conservative option. Rebuilding the pinned commit onto forgejo — the apparently low-risk choice — does not fix federation, because the gitea URLs are in the code at that commit. Deploying HEAD is the smaller intervention once the runtime dependency is counted.
Side finding for that repo:
REUSE-WP-0019is archived as finished and records a production deploy of its T01/T02 work, but T04–T06 (telemetry store, aggregation, hub freshness monitoring) landed after the deployed commit and appear never to have shipped. A workplan closed as done contains a freshness-monitoring feature production does not have. -
state-hub— mostly a false alarm. The Deployment is already onforgejo.coulomb.social/coulomb/state-hub:main-d8808bf. The gitea reference is a completed one-shot Job that will not re-run by itself; the residual risk is a chart template recreating it. Worth cleaning, not urgent.
There was no repository rename. This is a registry migration (gitea → forgejo, early July) in which producers moved and some consumers did not.
Full sweep, 2026-08-20 — CoulombCore runs two registry services
The first pass looked at container images. gitea.coulomb.social also serves a
PyPI package index (/api/packages/coulomb/pypi), confirmed by
railiance-fabric/fabric/interfaces/railiance-forge-python-package-index.yaml.
Both die on 2026-08-31.
Sweeping every repo for gitea.coulomb.social in operational files (yaml, sh,
py, Makefile, Dockerfile, toml, json) and classifying by whether anything can
still act on it:
Live, must move before 2026-08-31
| Reference | What breaks |
|---|---|
reuse-surface Deployment image |
Next restart → ImagePullBackOff. Also 22 commits stale |
kaizen-agentic Makefile + README + 5 docs |
DONE 2026-08-20 (KAIZEN-WP-0010, finished). Forgejo serves 1.4.0 anonymously (wheel + sdist); fresh venv installed 1.4.0 with --no-cache-dir from the documented forgejo extra index, CLI and import both reported 1.4.0. Gitea Make variables/target, consumer URLs, CLI links, release metadata and the .gitea issue-template path retired; historical workplans and changelog left truthful. Tests and make release-check pass |
Live but already dual-pathed — retire the old half
| Reference | Note |
|---|---|
issue-core/Makefile:240 publishes to gitea PyPI |
Line 250 already publishes to forgejo. The gitea target is legacy and should go |
Owned by the thing being decommissioned — expected
railiance-forge (manifests/gitea-ingress.yaml, helm/gitea-registry-values.yaml,
tools/gitea-runner-status.sh) owns gitea itself; these retire with it.
Not live — leave alone
Tests asserting historical facts (railiance-fabric, markitect-main,
reuse-surface/tests/test_forge_host.py, railiance-platform/tests/…),
inventory snapshots and asset/data registers (disaster-control,
domain-tree, railiance-fabric snapshots), agent session blobs, and
issue-core/Dockerfile:5 (a comment). These record that gitea existed, which
stays true after it is switched off. Rewriting them would destroy history to
tidy a grep.
One to check, not ours: railiance-platform/argocd/bootstrap/01-railiance-tenants-project.yaml:14
permits sourceRepos: https://gitea.coulomb.social/coulomb/*.git. railiance01
has no ArgoCD applications resource type, so this appears inert — but it is a
bootstrap file and should be confirmed rather than assumed.
Credential-lane sweep, 2026-08-20 — a third registry service
Checking ops-warden's routing catalog found what the file sweep could not: the host serves three package services, not two.
gitea.coulomb.social/api/packages/coulomb/npm/ is an npm registry, and
ops-warden's catalog lane whynot-design-npm-publish — risk: high,
production-exercised (WP-0018 published @whynot/design@0.4.0 through it) —
vends the NPM_AUTH_TOKEN that publishes to it.
So the complete CoulombCore package surface is:
| Service | Endpoint | Known consumers |
|---|---|---|
| OCI container registry | gitea.coulomb.social/coulomb/… |
reuse-surface (live Deployment) |
| PyPI index | …/api/packages/coulomb/pypi |
kaizen-agentic (never migrated), issue-core (legacy half) |
| npm registry | …/api/packages/coulomb/npm/ |
@whynot/design via ops-warden lane whynot-design-npm-publish |
Consequence for the credential lane: after 2026-08-31 that lane routes to a
registry that does not exist. It does not fail safe — an operator following it
gets a token for a dead endpoint and debugs the token. forgejo-admin-api-token
already exists as the forgejo-side equivalent and its keywords include
forgejo-npm, so the destination is plausibly in place; ops-warden must confirm
rather than assume, and cannot repoint the lane before the packages are on
forgejo — the same publish-before-repoint rule as KAIZEN-WP-0010.
Routed to ops-warden (lane owner for the catalog entry; railiance-platform
owns the credential itself).
CI runner sweep, 2026-08-20 — clear, and the sweep is now closed
86 repositories carry .forgejo/workflows, with jobs on self-hosted (83),
container-build (11), ubuntu-latest (87) and docker (1). Those labels
resolve to registered runners, so the question was where the runners live.
The only runner in the estate is on railiance01: forgejo/forgejo-runner,
a Deployment up 48 days, registering against ${FORGEJO_INSTANCE} — forgejo,
not gitea. No runner exists on CoulombCore.
railiance-forge/tools/gitea-runner-status.sh, which prompted this check, is a
legacy artifact: it defaults to RUNNER_HOST=haskelseed and probes
INTER_HUB_IMAGE. Both are already retired — haskelseed's bridge on 2026-08-19,
inter-hub in July. So the gitea-era runner lived on haskelseed and died with it a
day ago, and nothing broke: evidence in itself that gitea-based CI is no longer
in use.
CoulombCore's dependency surface is therefore fully enumerated across four
methods — tunnels, service DNS, workload image references, operational-file
grep, credential-lane catalog, and CI runners. Each method found something the
previous one structurally could not see, which is the finding T06's G-GEN
gate should encode: an inventory is only as complete as the number of
independent ways you looked.
forgejo.coulomb.social is already on railiance01, so the work is retag, push,
update manifest. Routed to railiance-platform; ownership of the reuse-surface
manifest sits with that repo.
Why this is a project task and not a platform ticket: it was missed because
the decommission inventory was assembled from tunnels and workplans. A host is
not free of dependents because nothing tunnels to it. Feed that into T06's
G-GEN gate — a generation or host transition must enumerate what pulls,
resolves and authenticates against the thing being switched off, not only what
connects to it.
Related
SHR-WP-0001— foundation and architecture baseline (finished)architecture/retirement-gates_v0.1.md—G-DISP,G-CORE-ABSORB,G-HUB-RUNTIMEinventory/capabilities.yaml— the gen-1 rollup this extendscore-hub/INTENT.md— the stated lineagecore-hubCORE-WP-0005/ archivedCORE-WP-0007— the July cutover and Haskell retirement, the evidence that corrected this workplan's first draftcore-hubCORE-WP-0010— runtime absorption intohub-core, all taskstodo; the likely vehicle for T03issue-coreISSUE-WP-0007— the CoulombCore migration precedent- ops-warden
inter-hub-bootstrap-ssh— a lane outliving its subject