activity-core/docs/ops-run-queue.md
tegwick b933cf52c8
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 1s
Build and Publish Container Image / build-and-push (push) Successful in 22s
Fail closed on invalid Glas profiles
Assistant: codex
Assistant-Model: gpt-5.6-sol
Assistant-Session: 01a028de-e2c8-7732-8521-46a7fc5db82f
2026-08-22 23:02:41 +02:00

7.8 KiB

Ops run claim queue

Status: implemented (ACTIVITY-WP-0026 code path; railiance rollout = T07)
Architecture: ACT-ADR-005
Deploy checklist: deploy-ops-run-queue-railiance.md

Durable, claimable work instances for scheduled / automation runs. Not a workplan task file. Not an issue-core or Forgejo ticket.

Model

Field Type Notes
id UUID Primary key (uuid4; UUIDv7 optional later)
activity_definition_id UUID FK → activity_definitions
idempotency_key text Unique; default {activity_id}:{source_id}:{triggering_event_id} — for cron fires, triggering_event_id must be per-fire (run_id or scheduled:{iso}), never bare scheduled
target_repo text From TaskSpec
title text
description text
labels JSONB array e.g. ["automated","research-brief"]
priority text low | medium | high
state text open | claimed | succeeded | failed | expired
claim_owner text nullable Worker identity
lease_until timestamptz nullable Claim lease deadline
attempt int Starts at 0; incremented on each claim
source_type text rule | instruction
source_id text Rule id
triggering_event_id text Event or workflow key
approach_hint text nullable Legacy definition-matching hint (ACT-ADR-006)
harness_profile_ref text nullable Authoritative execution selector, pinned <id>@<version>
execution_refs jsonb Attribution refs carried through, not authored here
result JSONB Completion metadata
created_at / updated_at timestamptz

API (actcore-api)

Method Path Role
GET /ops-runs List/filter
GET /ops-runs/{id} One row
POST /ops-runs/claim Claim open runs (lease)
POST /ops-runs/{id}/heartbeat Extend lease
POST /ops-runs/{id}/complete succeeded
POST /ops-runs/{id}/fail failed (+ optional reopen)
POST /ops-runs/expire-leases Reopen or expire stale claims

Claim body

{
  "worker_id": "rein-aharness@railiance01",
  "labels": ["automated"],
  "labels_mode": "any",
  "limit": 1,
  "lease_seconds": 900
}
  • labels_mode: any (default) — run must contain at least one listed label; all — run must contain every listed label; omit labels to claim any open run.
  • Claim uses FOR UPDATE SKIP LOCKED for concurrency safety.
  • Stale claims (state=claimed and lease_until < now()) are reopened before select.

Complete / fail body

{ "result": { "path": "briefs/…", "ok": true }, "worker_id": "…" }
{ "error": "llm timeout", "worker_id": "…", "reopen": false }

If reopen: true and attempt < max_attempts (env OPS_RUN_MAX_ATTEMPTS, default 3), state returns to open; else failed.

Emit path

On emit_tasks (when OPS_RUN_QUEUE_ENABLED is truthy, default true):

  1. Insert ops_run with state=open (idempotent on unique key).
  2. Dual-write existing IssueSink (state-hub progress by default).
  3. Write task_spawn_log audit as today.

Cron trigger keys: RunActivityWorkflow maps the schedule sentinel trigger_key="scheduled" to a per-fire triggering_event_id via emit_triggering_event_id() (scheduled:{scheduled_for} when known, else run_id). Using bare scheduled for every fire made only the first weekday create an ops_run; later days recorded activity_runs but left the claim queue empty.

Never requires Forgejo or issue-core for the claim path.

Run artefacts (ACTIVITY-WP-0027)

GET /ops/automations/{id}/runs joins each activity_run to related ops_runs and exposes deliverable links for the ops UI:

Field Source
ops_runs[] matched by triggering_event_id == run_id (or contains), else time window
artifacts[] from ops_runs.result (path, head_after, target_repo)
Forgejo URL FORGEJO_WEB_BASE (default https://forgejo.coulomb.social) + org + repo + path + ref

Run detail: GET /ops/automations/{id}/runs/{run_id} and /ops/ui/automations/{id}/runs/{run_id}.

Executor result should include at least: ok, path, head_after, target_repo, committed. No prompts or raw model output.

Auth

  • Worker: ACTIVITY_CORE_WORKER_TOKEN via X-Worker-Token or Authorization: Bearer (same value accepted on claim/complete/fail/heartbeat).
  • Operator: existing ops SSO / ACTIVITY_CORE_OPERATOR_TOKEN for list/status.
  • Local dev: ACTIVITY_CORE_OPS_ALLOW_UNAUTH_MUTATIONS=1 when no tokens set.

Env

Variable Default Meaning
OPS_RUN_QUEUE_ENABLED true Create ops_run on emit
OPS_RUN_LEASE_SECONDS 900 Default claim lease
OPS_RUN_MAX_ATTEMPTS 3 Fail permanently after N claims
ACTIVITY_CORE_WORKER_TOKEN unset Harness claim credential

Consumer (rein-aharness)

See REIN-A-0002. Claim loop → approach table → execute → complete + domain completion event (fi_daily_brief, etc.).

Not this queue

Concern Home
Multi-day engineering tasks Workplan files + State Hub
External tracker tickets issue-core → Forgejo (optional projection)
Schedule truth Temporal + activity definitions

Execution selection (ACT-ADR-006)

harness_profile_ref names an approved, version-pinned glas-harness profile (e.g. harness.agent-dev-local@1.0.0). The claiming executor passes the queued request into Glas, which resolves or refuses it before sandbox creation.

approach_hint and harness_profile_ref coexist with distinct semantics:

  • harness_profile_ref is the authoritative execution-constellation selector.
  • approach_hint is a legacy definition-matching hint only. It must never override, synthesize, or fall back from an absent or invalid profile ref. A malformed ref is an error at emission, not an invitation to route on the hint.

activity-core does not mirror the glas-harness profile catalogue — it is authoritative there, and Glas exposes no network validation service. So we validate structure only (present, no whitespace, <id>@<version> pinned). The pin matters: GlasProfiles.resolve treats an unpinned ref as matching every version and refuses it as ambiguous, so requiring the pin locally converts a late failure into an emission-time error without knowing any profile id.

A well-formed but unknown profile is still caught by the execution-side Glas resolver rather than at emission. That residual gap is accepted and recorded in ACT-ADR-006; closing it needs a scoped glas-harness API, not a local catalogue.

Task-emitting rules declare the selector and optional attribution refs on the action; instructions use the same fields at instruction level:

action:
  task_template: Run controlled maintenance
  harness_profile_ref: harness.agent-dev@1.0.0
  approach_hint: legacy-definition-match
  execution_refs:
    correlation_id: context.request.correlation_id
    goal_refs: [context.request.goal_ref]

File sync validates every declared profile structurally and rejects malformed or unversioned refs. emit_tasks repeats that validation across the complete batch before opening the database or IssueSink. A profile policy failure is a non-retryable activity error; no earlier item in that batch is emitted. Only the allowlisted attribution keys are carried to the queue.

ACTIVITY_CORE_REQUIRE_HARNESS_PROFILE=true makes a missing profile ref an error both during file sync and emission. Deterministic report-only instructions are exempt because they create no execution request. The flag stays off during coexistence while definitions adopt refs one at a time; turn it on once no caller depends on approach_hint for routing.