activity-core/docs/ops-run-queue.md
tegwick 21dc228cb1
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 2s
Build and Publish Container Image / build-and-push (push) Successful in 21s
fix(ops_run): unique triggering_event_id per cron fire
Cron schedules always passed trigger_key="scheduled", so ops_run
idempotency collapsed every weekday into one key. After the first fire,
create_ops_run was a silent no-op, the claim loop starved, and dual-clock
host timers produced empty FI briefs.

Map scheduled fires to run_id (or scheduled:{iso}) via
emit_triggering_event_id; log duplicate skips; document the contract.
2026-08-05 15:30:00 +02:00

119 lines
4.5 KiB
Markdown

# Ops run claim queue
**Status:** implemented (ACTIVITY-WP-0026 code path; railiance rollout = T07)
**Architecture:** [ACT-ADR-005](adr/adr-005-ops-runs-vs-dev-work-records.md)
**Deploy checklist:** [deploy-ops-run-queue-railiance.md](deploy-ops-run-queue-railiance.md)
Durable, claimable work instances for **scheduled / automation** runs. Not a
workplan task file. Not an issue-core or Forgejo ticket.
## Model
| Field | Type | Notes |
| ----- | ---- | ----- |
| `id` | UUID | Primary key (uuid4; UUIDv7 optional later) |
| `activity_definition_id` | UUID | FK → activity_definitions |
| `idempotency_key` | text | **Unique**; default `{activity_id}:{source_id}:{triggering_event_id}` — for cron fires, `triggering_event_id` must be **per-fire** (`run_id` or `scheduled:{iso}`), never bare `scheduled` |
| `target_repo` | text | From TaskSpec |
| `title` | text | |
| `description` | text | |
| `labels` | JSONB array | e.g. `["automated","research-brief"]` |
| `priority` | text | low \| medium \| high |
| `state` | text | `open` \| `claimed` \| `succeeded` \| `failed` \| `expired` |
| `claim_owner` | text nullable | Worker identity |
| `lease_until` | timestamptz nullable | Claim lease deadline |
| `attempt` | int | Starts at 0; incremented on each claim |
| `source_type` | text | rule \| instruction |
| `source_id` | text | Rule id |
| `triggering_event_id` | text | Event or workflow key |
| `approach_hint` | text nullable | Optional from rule |
| `result` | JSONB | Completion metadata |
| `created_at` / `updated_at` | timestamptz | |
## API (actcore-api)
| Method | Path | Role |
| ------ | ---- | ---- |
| `GET` | `/ops-runs` | List/filter |
| `GET` | `/ops-runs/{id}` | One row |
| `POST` | `/ops-runs/claim` | Claim open runs (lease) |
| `POST` | `/ops-runs/{id}/heartbeat` | Extend lease |
| `POST` | `/ops-runs/{id}/complete` | succeeded |
| `POST` | `/ops-runs/{id}/fail` | failed (+ optional reopen) |
| `POST` | `/ops-runs/expire-leases` | Reopen or expire stale claims |
### Claim body
```json
{
"worker_id": "rein-aharness@railiance01",
"labels": ["automated"],
"labels_mode": "any",
"limit": 1,
"lease_seconds": 900
}
```
- `labels_mode`: `any` (default) — run must contain at least one listed label;
`all` — run must contain every listed label; omit `labels` to claim any open run.
- Claim uses `FOR UPDATE SKIP LOCKED` for concurrency safety.
- Stale claims (`state=claimed` and `lease_until < now()`) are reopened before select.
### Complete / fail body
```json
{ "result": { "path": "briefs/…", "ok": true }, "worker_id": "…" }
```
```json
{ "error": "llm timeout", "worker_id": "…", "reopen": false }
```
If `reopen: true` and `attempt < max_attempts` (env `OPS_RUN_MAX_ATTEMPTS`, default 3),
state returns to `open`; else `failed`.
## Emit path
On `emit_tasks` (when `OPS_RUN_QUEUE_ENABLED` is truthy, **default true**):
1. Insert `ops_run` with `state=open` (idempotent on unique key).
2. Dual-write existing IssueSink (`state-hub` progress by default).
3. Write `task_spawn_log` audit as today.
**Cron trigger keys:** `RunActivityWorkflow` maps the schedule sentinel
`trigger_key="scheduled"` to a per-fire `triggering_event_id` via
`emit_triggering_event_id()` (`scheduled:{scheduled_for}` when known, else
`run_id`). Using bare `scheduled` for every fire made only the first weekday
create an `ops_run`; later days recorded `activity_runs` but left the claim
queue empty.
Never requires Forgejo or issue-core for the claim path.
## Auth
- **Worker:** `ACTIVITY_CORE_WORKER_TOKEN` via `X-Worker-Token` or
`Authorization: Bearer` (same value accepted on claim/complete/fail/heartbeat).
- **Operator:** existing ops SSO / `ACTIVITY_CORE_OPERATOR_TOKEN` for list/status.
- **Local dev:** `ACTIVITY_CORE_OPS_ALLOW_UNAUTH_MUTATIONS=1` when no tokens set.
## Env
| Variable | Default | Meaning |
| -------- | ------- | ------- |
| `OPS_RUN_QUEUE_ENABLED` | `true` | Create ops_run on emit |
| `OPS_RUN_LEASE_SECONDS` | `900` | Default claim lease |
| `OPS_RUN_MAX_ATTEMPTS` | `3` | Fail permanently after N claims |
| `ACTIVITY_CORE_WORKER_TOKEN` | unset | Harness claim credential |
## Consumer (rein-aharness)
See REIN-A-0002. Claim loop → approach table → execute → complete + domain
completion event (`fi_daily_brief`, etc.).
## Not this queue
| Concern | Home |
| ------- | ---- |
| Multi-day engineering tasks | Workplan files + State Hub |
| External tracker tickets | issue-core → **Forgejo** (optional projection) |
| Schedule truth | Temporal + activity definitions |