activity-core/docs/ops-run-queue.md

236 lines
9.9 KiB
Markdown
Raw Normal View History

# Ops run claim queue
**Status:** implemented (ACTIVITY-WP-0026 code path; railiance rollout = T07)
**Architecture:** [ACT-ADR-005](adr/adr-005-ops-runs-vs-dev-work-records.md)
**Deploy checklist:** [deploy-ops-run-queue-railiance.md](deploy-ops-run-queue-railiance.md)
Durable, claimable work instances for **scheduled / automation** runs. Not a
workplan task file. Not an issue-core or Forgejo ticket.
## Model
| Field | Type | Notes |
| ----- | ---- | ----- |
| `id` | UUID | Primary key (uuid4; UUIDv7 optional later) |
| `activity_definition_id` | UUID | FK → activity_definitions |
| `idempotency_key` | text | **Unique**; default `{activity_id}:{source_id}:{triggering_event_id}` — for cron fires, `triggering_event_id` must be **per-fire** (`run_id` or `scheduled:{iso}`), never bare `scheduled` |
| `target_repo` | text | From TaskSpec |
| `title` | text | |
| `description` | text | |
| `labels` | JSONB array | e.g. `["automated","research-brief"]` |
| `priority` | text | low \| medium \| high |
| `state` | text | `open` \| `claimed` \| `succeeded` \| `failed` \| `expired` |
| `claim_owner` | text nullable | Worker identity |
| `lease_until` | timestamptz nullable | Claim lease deadline |
| `attempt` | int | Starts at 0; incremented on each claim |
| `source_type` | text | rule \| instruction |
| `source_id` | text | Rule id |
| `triggering_event_id` | text | Event or workflow key |
| `approach_hint` | text nullable | **Legacy** definition-matching hint (ACT-ADR-006) |
| `harness_profile_ref` | text nullable | Authoritative execution selector, pinned `<id>@<version>` |
| `execution_refs` | jsonb | Attribution refs carried through, not authored here |
| `result` | JSONB | Completion metadata |
| `created_at` / `updated_at` | timestamptz | |
## API (actcore-api)
| Method | Path | Role |
| ------ | ---- | ---- |
| `GET` | `/ops-runs` | List/filter |
| `GET` | `/ops-runs/{id}` | One row |
| `POST` | `/ops-runs/claim` | Claim open runs (lease) |
| `POST` | `/ops-runs/{id}/heartbeat` | Extend lease |
| `POST` | `/ops-runs/{id}/complete` | succeeded |
| `POST` | `/ops-runs/{id}/fail` | failed (+ optional reopen) |
| `POST` | `/ops-runs/expire-leases` | Reopen or expire stale claims |
### Claim body
```json
{
"worker_id": "rein-aharness@railiance01",
"labels": ["automated"],
"labels_mode": "any",
"limit": 1,
"lease_seconds": 900
}
```
- `labels_mode`: `any` (default) — run must contain at least one listed label;
`all` — run must contain every listed label; omit `labels` to claim any open run.
- Claim uses `FOR UPDATE SKIP LOCKED` for concurrency safety.
- Stale claims (`state=claimed` and `lease_until <= now()`) are reopened before select.
- Heartbeat, complete, and fail lock the row and require an active lease
(`lease_until > now()`). An expired worker cannot revive or close its claim.
### Complete / fail body
```json
{ "result": { "path": "briefs/…", "ok": true }, "worker_id": "…" }
```
Glas consumers submit its `GatewayResult` directly. Activity Core stores the
safe evidence envelope and drops direct-caller output:
```json
{
"worker_id": "glas-worker@railiance01",
"result": {
"ok": true,
"evidence": {
"request_id": "…",
"contract_version": "1.0",
"profile_ref": "harness.agent-dev@1.0.0",
"rein_id": "rein-aharness",
"rein_version": "0.1.0",
"resolved_model": "claude-sonnet-4-6",
"sandbox_profile": "profile.bwrap-local",
"outcome": "succeeded",
"duration_s": 12.4,
"tokens_spent": 321,
"tool_events_count": 0,
"tool_events_completeness": "unavailable"
},
"tool_output": "dropped before persistence"
}
}
```
```json
{ "error": "llm timeout", "worker_id": "…", "reopen": false }
```
If `reopen: true` and `attempt < max_attempts` (env `OPS_RUN_MAX_ATTEMPTS`, default 3),
state returns to `open`; else `failed`.
Both completion and failure normalize through the same allowlist. The persisted
key is `result.execution_evidence`; it contains the Glas 1.0 constellation,
outcome, bounded measurements, commit/artifact references, and allowlisted
organizational refs. Zero stays zero and unsupported event visibility stays
`unavailable`; absent measurements remain absent. `tool_output`, `tool_error`,
prompts, messages, provider responses, credential fields, unknown nested blobs,
and undeclared refs are never stored. Read projections normalize historic rows
again before returning them.
## Emit path
On `emit_tasks` (when `OPS_RUN_QUEUE_ENABLED` is truthy, **default true**):
1. Insert `ops_run` with `state=open` (idempotent on unique key).
2. Dual-write existing IssueSink (`state-hub` progress by default).
3. Write `task_spawn_log` audit as today.
**Cron trigger keys:** `RunActivityWorkflow` maps the schedule sentinel
`trigger_key="scheduled"` to a per-fire `triggering_event_id` via
`emit_triggering_event_id()` (`scheduled:{scheduled_for}` when known, else
`run_id`). Using bare `scheduled` for every fire made only the first weekday
create an `ops_run`; later days recorded `activity_runs` but left the claim
queue empty.
Never requires Forgejo or issue-core for the claim path.
## Run artefacts (ACTIVITY-WP-0027)
`GET /ops/automations/{id}/runs` joins each `activity_run` to related
`ops_runs` and exposes deliverable links for the ops UI:
| Field | Source |
| ----- | ------ |
| `ops_runs[]` | matched by `triggering_event_id == run_id` (or contains), else time window |
| `artifacts[]` | from `ops_runs.result` (`path`, `head_after`, `target_repo`) |
| `ops_runs[].execution_evidence` | compact Glas profile/rein/model/sandbox/outcome evidence |
| Forgejo URL | `FORGEJO_WEB_BASE` (default `https://forgejo.coulomb.social`) + org + repo + path + ref |
Run detail: `GET /ops/automations/{id}/runs/{run_id}` and
`/ops/ui/automations/{id}/runs/{run_id}`.
Executor `result` should include at least: `ok`, `path`, `head_after`,
`target_repo`, `committed`. A Glas-backed executor should submit its full
`GatewayResult`; Activity Core extracts the safe envelope. No prompts or raw
model output are persisted or returned.
## Auth
- **Worker:** `ACTIVITY_CORE_WORKER_TOKEN` via `X-Worker-Token` or
`Authorization: Bearer`, bound to the exact non-secret
`ACTIVITY_CORE_WORKER_ID`. The body `worker_id` is a compatibility field and
must match that authenticated identity on claim/complete/fail/heartbeat.
- **Operator:** existing ops SSO / `ACTIVITY_CORE_OPERATOR_TOKEN` for
list/status and explicit lease expiry; operator credentials are not accepted
as worker mutation identities.
- **Local dev:** `ACTIVITY_CORE_OPS_ALLOW_UNAUTH_MUTATIONS=1` is required when
no tokens are set. There is no implicit open mode.
## Env
| Variable | Default | Meaning |
| -------- | ------- | ------- |
| `OPS_RUN_QUEUE_ENABLED` | `true` | Create ops_run on emit |
| `OPS_RUN_LEASE_SECONDS` | `900` | Default claim lease |
| `OPS_RUN_MAX_ATTEMPTS` | `3` | Fail permanently after N claims |
| `ACTIVITY_CORE_WORKER_TOKEN` | unset | Harness claim credential |
| `ACTIVITY_CORE_WORKER_ID` | unset | Exact queue identity bound to the worker credential; required when the token is set |
## Consumer (rein-aharness)
See REIN-A-0002. Claim loop → approach table → execute → complete + domain
completion event (`fi_daily_brief`, etc.).
## Not this queue
| Concern | Home |
| ------- | ---- |
| Multi-day engineering tasks | Workplan files + State Hub |
| External tracker tickets | issue-core → **Forgejo** (optional projection) |
| Schedule truth | Temporal + activity definitions |
## Execution selection (ACT-ADR-006)
`harness_profile_ref` names an approved, **version-pinned** glas-harness profile
(e.g. `harness.agent-dev-local@1.0.0`). The claiming executor passes the queued
request into Glas, which resolves or refuses it **before sandbox creation**.
`approach_hint` and `harness_profile_ref` coexist with **distinct semantics**:
- `harness_profile_ref` is the authoritative execution-constellation selector.
- `approach_hint` is a legacy definition-matching hint only. It must **never**
override, synthesize, or fall back from an absent or invalid profile ref. A
malformed ref is an error at emission, not an invitation to route on the hint.
activity-core does **not** mirror the glas-harness profile catalogue — it is
authoritative there, and Glas exposes no network validation service. So we
validate structure only (present, no whitespace, `<id>@<version>` pinned). The
pin matters: `GlasProfiles.resolve` treats an unpinned ref as matching every
version and refuses it as ambiguous, so requiring the pin locally converts a
late failure into an emission-time error without knowing any profile id.
A well-formed but *unknown* profile is still caught by the execution-side Glas
resolver rather than at emission. That residual gap is accepted and recorded in
ACT-ADR-006; closing it needs a scoped glas-harness API, not a local catalogue.
Task-emitting rules declare the selector and optional attribution refs on the
action; instructions use the same fields at instruction level:
```yaml
action:
task_template: Run controlled maintenance
harness_profile_ref: harness.agent-dev@1.0.0
approach_hint: legacy-definition-match
execution_refs:
correlation_id: context.request.correlation_id
goal_refs: [context.request.goal_ref]
```
File sync validates every declared profile structurally and rejects malformed
or unversioned refs. `emit_tasks` repeats that validation across the complete
batch before opening the database or IssueSink. A profile policy failure is a
non-retryable activity error; no earlier item in that batch is emitted. Only the
allowlisted attribution keys are carried to the queue.
`ACTIVITY_CORE_REQUIRE_HARNESS_PROFILE=true` makes a missing profile ref an
error both during file sync and emission. Deterministic report-only
instructions are exempt because they create no execution request. The flag
stays off during coexistence while definitions adopt refs one at a time; turn it
on once no caller depends on `approach_hint` for routing.