activity-core/docs/issue-core-emission-boundary.md
tegwick 9a7ae8b59a
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 6s
Build and Publish Container Image / build-and-push (push) Successful in 13s
Add ExternalSecret for ISSUE_CORE_API_KEY on Railiance
Sync the shared issue-core ingestion key from OpenBao into
actcore-runtime-secret via External Secrets, with an interim coulombcore
ClusterSecretStore bootstrap script and deploy docs. Removes manual key
injection from bootstrap-secrets.sh.
2026-07-08 00:04:38 +02:00

4.5 KiB

Issue-Core Emission Boundary

activity-core owns the decision to spawn a task and the audit trail that says why it spawned. It does not own downstream task lifecycle state after emission.

Current authoritative endpoint

The current authoritative boundary is the issue-core REST API:

POST {ISSUE_CORE_URL}/issues/

IssueCoreRestSink authenticates with the shared ISSUE_CORE_API_KEY env var (same value as the issue-core server) via Authorization: Bearer <key> and sends this payload:

{
  "title": "Run SBOM rescan for activity-core",
  "description": "",
  "target_repo": "activity-core",
  "priority": "medium",
  "labels": ["sbom", "security", "automated"],
  "due_in_days": null,
  "source_type": "rule",
  "source_id": "flag-stale-sbom",
  "triggering_event_id": "event-or-schedule-key",
  "activity_definition_id": "activity-definition-uuid"
}

The expected response contains issue_id and may include issue_url and backend. activity-core stores only the returned task reference in task_spawn_log; issue-core remains authoritative for task status, assignment, comments, closure, and cancellation.

REST versus NATS

Keep REST as the active emission contract until issue-core publishes and owns a durable NATS consumer for task-creation commands. NATS is still appropriate for event intake into activity-core, but task creation needs an acknowledged, idempotent command boundary. A future NATS sink must return or later correlate a task reference before it can replace IssueCoreRestSink.

Safe operating modes

  • ISSUE_SINK_TYPE=null: dry-run/audit mode. Task specs are rendered and the workflow records synthetic null-* references. Use this for contract review and emergency rollback.
  • ISSUE_SINK_TYPE=rest: live task creation. Sink failures raise out of emit_tasks, so Temporal retries and the workflow history make failures visible. Railiance runtime ConfigMap uses this mode once ISSUE_CORE_API_KEY is present in actcore-runtime-secret.

Weekly SBOM staleness is the canonical promotion candidate because the rule contract is deterministic and tested. Promote it only after a null-sink dry-run review and one live IssueCoreRestSink smoke against the target endpoint.

Promotion and rollback

Promote one definition safely

  1. Keep ISSUE_SINK_TYPE=null and run or wait for the target definition.
  2. Review rendered task specs in task_spawn_log (source id, condition, target repo, synthetic null-* reference).
  3. Confirm ISSUE_CORE_URL reachability and a populated ISSUE_CORE_API_KEY on both activity-core and issue-core (same value). Credential custody: warden route show issue-core-ingestion-api-key --json.
  4. Run the repo smoke:
    uv run python scripts/smoke_issue_core_emission.py
    ISSUE_CORE_URL=http://127.0.0.1:8765 ISSUE_CORE_API_KEY=... \
      uv run python scripts/smoke_issue_core_emission.py --live
    
  5. Set ISSUE_SINK_TYPE=rest in actcore-runtime-config, apply k8s/railiance/15-externalsecret-issue-core.yaml so External Secrets merges ISSUE_CORE_API_KEY into actcore-runtime-secret, and restart actcore-worker / actcore-event-router after the ExternalSecret is Ready.
  6. Trigger one known-safe run (weekly SBOM staleness on a stale fixture or manual /activity-definitions/<id>/trigger) and confirm task_spawn_log stores the real issue_id returned by issue-core.

Roll back to null-sink

  1. Set ISSUE_SINK_TYPE=null in actcore-runtime-config.
  2. kubectl -n activity-core rollout restart deploy/actcore-worker deploy/actcore-event-router
  3. Verify the next run records synthetic null-* references again.
  4. Leave issue-core tasks already created in place; activity-core does not own downstream task lifecycle. Close or cancel duplicates in issue-core if a promotion experiment created unexpected tasks.

Duplicate handling today: issue-core REST ingest does not yet dedupe on triggering_event_id; Temporal retry visibility is the current guardrail. Treat promotion as one-definition-at-a-time until server-side idempotency ships.

Verification

Local contract tests cover the rendered weekly SBOM task path and the REST payload shape:

uv run pytest tests/test_integration_event_bridge.py tests/test_issue_sink.py

For a live environment, run with ISSUE_SINK_TYPE=null first and confirm task_spawn_log contains the expected source id, condition, triggering event id, and synthetic task reference. Then switch to ISSUE_SINK_TYPE=rest only after a single known-safe rule match creates one issue-core task with the same fields.