Helm revision 8. /health returns 200 from uvicorn instead of 404 from the landing nginx, / still serves the landing page, and make reuse-smoke passes end to end for the first time. Also resolves the composed_at observation: it was not a bug, just a build that predated REUSE-WP-0019-T06. The timestamp now advances per recompose. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
16 KiB
| id | type | title | domain | repo | status | owner | topic_slug | created | updated |
|---|---|---|---|---|---|---|---|---|---|
| REUSE-WP-0020 | workplan | CoulombCore retirement cutover: federation sources, image, and compose resilience | infotech | reuse-surface | active | claude | helix-forge | 2026-08-20 | 2026-08-20 |
CoulombCore retirement cutover
CoulombCore is switched off 2026-08-31 (operator decision, 2026-08-20).
gitea.coulomb.social runs on it. Raised by prj-state-hub-retirement
(SHR-WP-0002-T07) as a stale-image report; investigation found a second,
more urgent problem underneath it.
Two independent failures, not one:
- Federation source data — the production hub's registrations pointed
50 of 61 sources at
gitea.coulomb.social. This breaks on 08-31 with nobody restarting anything. It is database state, not code:registry/federation/sources.yamlat HEAD has been 61/61 Forgejo sinced1c1313(RAIL-HO-WP-0006). - Deployed image —
railiance-apps/helm/reuse-surface-values.yamlpinsimage.tag: "e3ae22e". The chart'srepositorywas already migrated to Forgejo inrailiance-apps@04be416, bute3ae22eis commit-dated 2026-07-07 18:25 and CI only began publishing to Forgejo at 21:25 the same day — so a Forgejo:e3ae22etag most likely never existed. Any restart, reschedule, or node reboot risksImagePullBackOffnow, not on 08-31.
Item 1 was resolved on 2026-08-20 (see T01). Doing so exposed a third issue: a single malformed member index takes the whole federated endpoint down.
Baseline vector: D5 / A4 / C5 / R3 — unchanged by this workplan;
this is operational continuity, not capability growth.
Scope boundary
The Deployment manifest is not in this repo. It lives in railiance-apps
(charts/reuse-surface/, helm/reuse-surface-values.yaml, RAILIANCE-WP-0007).
This repo owns the image contents, the deploy guide, and the federation data;
applying the manifest change requires a working KUBECONFIG=~/.kube/config-hosteurope,
which is not reachable from the workstation without an ops-bridge tunnel.
Repoint Production Federation Sources To Forgejo
id: REUSE-WP-0020-T01
status: done
priority: high
Completed 2026-08-20. Updated 50 enabled hub registrations from
gitea.coulomb.social/<repo>/raw/main/... to
forgejo.coulomb.social/<repo>/raw/branch/main/... via
reuse-surface hub update --repo <r> --url <new>.
Notes for the record:
- All 50 new URLs were pre-verified
200before any write. - The canonical Forgejo raw form is
/raw/branch/main/./raw/main/answers303and works only becauseurllib.request.urlopenfollows redirects —registry/federation/sources.yamlstill uses the redirecting form. Harmless, but prefer the canonical form in new writes. - One registration (
inter-hub, alreadydisabled) still carries a Gitea URL. Its Forgejo equivalent answers307, so it was left alone rather than repointed to something unverified. It contributes nothing while disabled.
Verified: GET /v1/federated → 60 sources, all Forgejo, 62 capabilities.
Harden Compose Against A Malformed Member Index
id: REUSE-WP-0020-T02
status: done
priority: high
T01 briefly took /v1/federated to HTTP 500. Root cause: evidence-binder's
live index has two capability entries with no id field, and
compose_federated_index does cap_id = item["id"] unguarded. Its Gitea copy
was a stale snapshot returning a non-mapping, so the entries had never been
parsed before.
One bad member index must not take down the federated endpoint. In
reuse_surface/federation.py::compose_federated_index, guard the per-source
body so a failure degrades to a warning and the remaining sources still
compose:
| Failure | Current | Wanted |
|---|---|---|
item["id"] missing |
KeyError, 500 |
skip entry, warn, keep source |
yaml.safe_load → non-mapping |
AttributeError, 500 |
skip source, warn |
yaml.safe_load → None (empty file) |
AttributeError, 500 |
skip source, warn |
| YAML parse error | YAMLError, 500 |
skip source, warn |
sorted(key=item["id"]) |
KeyError |
unreachable once entries are filtered |
Confirmed present at HEAD, not only in the deployed image — deploying HEAD would not have fixed this. Land T02 before T04.
Pytest in tests/test_federation.py: member index with a missing id, an
empty file, a non-mapping body, and unparseable YAML — each composes the
other sources successfully and emits a warning naming the source.
Done 2026-08-20. Extracted _read_index_entries(); every read failure now
returns an empty row list plus a warning, and rows without an id are skipped
individually. A failed source stays in sources with count: 0 rather than
vanishing, so operators can still see it. Six tests added. Verified against the
real evidence-binder index: composes clean with two skip warnings instead of
raising.
Fix The evidence-binder Capability Index
id: REUSE-WP-0020-T03
status: wait
priority: medium
Cross-repo, in ~/evidence-binder. Routed 2026-08-20, not edited directly —
that repo is active (last commit 2026-07-30) and has its own identity and topic.
The defect is larger than the missing id. Both rows use an entirely
non-conforming shape — type / title / description / keywords instead of
id / name / summary / vector / domain / status / owner / path /
tags / consumption_modes. History shows 597b489 added the
REUSE-WP-0014-T04 scaffold and then a10f080 ("Extract binder package from
citation-evidence") hand-wrote these two rows around it. registry/capabilities/
is empty, so both index rows are orphans with no backing entry Markdown.
Left to that repo's owner because the remaining judgment is theirs: reshaping
the rows is mechanical, but assigning D/A/C/R maturity vectors is an assessment
of their own delivery state, and invented values would be wrong in a way that is
hard to notice later. The message offers capabilities: [] as an honest
alternative — ten other federation members sit at count 0 — and offers to do the
work here if they supply vectors or accept a conservative D1/A0/C0/R0
placeholder marked for review.
Suggested ids: capability.evidence.binding,
capability.evidence.rect-registry.
Re-Enable The evidence-binder Source
id: REUSE-WP-0020-T04
status: wait
priority: medium
Blocked on T02 and T03. evidence-binder was set --no-enabled on the
production hub to restore service. Once its index is valid and compose is
hardened, re-enable and confirm /v1/federated returns 61 sources.
Repoint The Production Deployment To A Forgejo Image
id: REUSE-WP-0020-T05
status: done
priority: high
Deadline-bound: must land before 2026-08-31.
-
Verify which tags actually exist.Done 2026-08-20. Authenticated against the registry with the workstation's stored Forgejo credentials:GET /v2/coulomb/reuse-surface/tags/list → {"tags": ["latest", "main-bca7165", "main-f9d957a"]}There is no
e3ae22etag. The pinned manifest cannot pull today — this is confirmed, not inferred.main-f9d957a(2026-07-08, REUSE-WP-0019-T06) is the newest build; the 10 commits between it and HEAD were docs, workplans, and licence changes that do not match the image workflow's path filters, so no newer image exists yet. -
Bump
image.taginrailiance-apps/helm/reuse-surface-values.yamlfrome3ae22eto a verified Forgejo tag.Done 2026-08-20 —
railiance-apps@029460dpinsmain-b035664, verified to resolve (manifest200). CI built it afterfix-consistencypushed the T02/T08 commits. It carries the compose hardening, whichmain-f9d957adoes not: without it, re-enablingevidence-binder(T04) would take production back to HTTP 500.Committed but not applied.
railiance-appshas no ArgoCD or Flux controller, so a commit there deploys nothing —make reuse-deployis a manualhelm upgraderequiringREUSE_KUBECONFIGfor Railiance01. -
CorrectDone 2026-08-20 (docs/deploy/reuse-kubernetes.md.b035664). -
Applied 2026-08-21 00:44 CEST — Helm revision 7,
Upgrade complete.Access note worth keeping:
RAILIANCE01_KUBECONFIGdefaults to~/.kube/config-hosteurope(127.0.0.1:16443), which has no tunnel. The working path is~/.kube/config-railiance01(127.0.0.1:16444), served by thek3s-api-railiance01ops-bridge tunnel. Deploy withRAILIANCE01_KUBECONFIG=$HOME/.kube/config-railiance01 make reuse-deployuntil the Makefile default is corrected.Pre-flight checks that mattered:
- The live Deployment really was on
gitea.coulomb.social/coulomb/reuse-surface:e3ae22ewithpullPolicy: IfNotPresent— it had survived 18 restarts only because the image was cached on the node. One eviction or image GC would have ended it, retirement date or not. - No
imagePullSecretsanywhere in the chart or the namespace, and the Forgejo/v2/endpoint answers401anonymously — which looked fatal. It is not: that401is Forgejo's token handshake, and peers likeaudit-corepull from the same registry with no secret. Verified properly via the anonymous token flow (/v2/token?scope=repository:coulomb/reuse-surface:pull) → manifest200. No pull secret needed.
Post-deploy verification:
Check Result Pod reuse-surface-7f79d88c-fz7fd, 0 restarts, readyImage forgejo.coulomb.social/coulomb/reuse-surface:main-b035664PVC reuse-surface-databound to the samepvc-98165f95…— data keptRegistrations all 61 survived /v1/repos,/v1/federated200; 60 sources, all Forgejo, 62 capabilities /v1/reuse-events200 (was 404) — REUSE-WP-0019 T04/T05 finally live TLS certificate condition metRollback if ever needed:
helm rollback reuse 6 -n reuse— but note revision 6 pins the Gitea image, which only pulls while the node cache holds it and not at all after 2026-08-31. Rolling forward is the safer direction.Superseded steps, kept for context:
make reuse-dry-run # helm template, no apply make reuse-deploy # helm upgrade --install --wait make reuse-status # pods, svc, ingress, pvc, certificate curl -s https://reuse.coulomb.social/v1/federated | head -c 300 curl -s https://reuse.coulomb.social/v1/reuse-events # 404 today; 200 proves T04/T05 shippedConfirm the PVC at
/datasurvived, sincereuse.dband the remote index cache live there.make reuse-smokestill fails — it runscurl -fsS $(REUSE_URL)/health, and/healthreturns nginx404through the ingress (T06). Everything else in that target passes; the failure is a false negative. - The live Deployment really was on
Ships REUSE-WP-0019 T04/T05/T06 as a side effect — telemetry store,
R-axis aggregation, and hub freshness monitoring were closed as finished but
landed after the deployed commit and have never run in production.
GET /v1/reuse-events returning 404 on production confirms this.
Review the six weeks of change between e3ae22e and HEAD before applying.
Correct Hub Freshness And Health Routing
id: REUSE-WP-0020-T06
status: done
priority: medium
Two smaller production inconsistencies found while verifying T01:
-
GET /healthreturns nginx404through the ingress. Diagnosed 2026-08-21, not yet fixed. The application is fine: in-clusterhttp://reuse-surface:8000/healthreturns200. The ingress rule is also correct on paper —/healthExact→reuse-surface:8000, ahead of/Prefix→reuse-surface-landing:8080. The problem is Traefik: it derives router priority from the length of the generated rule string, andPath(`/health`)andPathPrefix(`/`)are the same length, so the tiebreak is arbitrary and the landing route wins. Confirmed by the response carryingserver: nginx/1.27.5(the landing container) rather thanuvicorn.Fixed 2026-08-21 —
railiance-apps@a365635, Helm revision 8. Therouter.priorityannotation applies per Ingress, not per path, so the landing catch-all moved into its own Ingress (landing-ingress.yaml) withpriority: 1; every API route now outranks it regardless of rule length.Path Before After /health404,server: nginx200,server: uvicorn/200, landing200, landing (unchanged)/v1/*200200(unchanged)make reuse-smokenow passes end to end. -
composed_atstayed at2026-08-20T20:44:12withstale: falseacross a recompose that demonstrably changed output (61 → 62 capabilities). Resolved by T05, not a bug. The build deployed at the time predated REUSE-WP-0019-T06;composed_atnow advances with each recompose (2026-08-21T00:18:02observed after the redeploy).
Refresh SCOPE.md Standard Sections
id: REUSE-WP-0020-T07
status: done
priority: low
Unrelated to the retirement, recorded so it is not lost. The hub's repo scope
check reported C5b/C5c warnings: SCOPE.md was missing the standard H2
sections Relevant When, Not Relevant When, How It Fits, Terminology,
Related / Overlapping, Provided Capabilities, and carried no fenced
capability block.
Done 2026-08-21 (823ce9e). All six sections added plus three capability
blocks.
Writing this section surfaced the root cause of T03. The fenced capability
blocks that belong in SCOPE.md use type / title / description /
keywords — which is exactly the shape evidence-binder put in its
registry/indexes/capabilities.yaml, where the required shape is id / name
/ summary / vector / domain / status / owner / path / tags /
consumption_modes. They did not invent a broken format; they copied a valid
one from the wrong place.
That makes it a systemic trap rather than a one-off mistake, so the Terminology section now documents the distinction explicitly, and the compose warning now names it:
evidence-binder: capability #0 has no id, skipped — looks like a SCOPE.md
capability block (type/title/description/keywords); an index row needs
id/name/summary/vector/domain/status/owner/path
Federation swept 2026-08-21. Composed all 61 sources with --refresh:
evidence-binder is the only affected member. The other ten zero-count
repos are genuine empty scaffolds and produced no warnings, so the confusion has
not spread.
Fix Wall-Clock Rot In The Plan-Check Test Suite
id: REUSE-WP-0020-T08
status: done
priority: medium
Found while verifying T02: tests/test_plan_check.py was already failing at
clean HEAD, independent of any change here. Three tests pinned the federated
index compose date to the literal "2026-07-06", and
test_run_plan_check_reuse_verdict asserts no staleness warning — so the test
passed when written and started failing once wall-clock time drifted past
STALE_DAYS. CI is red today for this reason alone, which matters because
T05 calls for verifying the image workflow is green.
Done 2026-08-20. Replaced the literals with a _recent_date() helper
derived from STALE_DAYS. Full suite: 179 passed.
Verification
python3 -m pytest tests/test_federation.py -q
python3 -m reuse_surface.cli validate
python3 -m reuse_surface.cli federation compose --refresh
REUSE_SURFACE_URL=https://reuse.coulomb.social python3 -m reuse_surface.cli hub list
curl -s https://reuse.coulomb.social/v1/federated | python3 -m json.tool | head -20
No source may reference gitea.coulomb.social after this workplan closes.