coordination-engine/docs/worker-runtime.md

195 lines
9.5 KiB
Markdown
Raw Normal View History

# Worker coordination runtime
The local Python 3.11+ service observes State Hub and wakes existing TAMQ workers.
It does not create terminals, execute repository commands, or change task status.
TAMQ must be installed independently; its Python package is not a dependency.
## Install and configure
```sh
uv sync --locked
uv tool install .
# For development: uv tool install --editable .
```
Create `~/.config/coordination-engine/config.toml` using this example. The `repos`
list is an explicit operator selection of workers eligible for automatic wakes.
The default empty list refuses to start. All selected repos and
`coordination-engine` must already appear in `gita freeze`.
```toml
[coordination]
repos = ["net-kingdom"]
api_base = "http://127.0.0.1:8000"
# Use the SAME path as the independently configured TAMQ service:
tamq_socket = "/tmp/tamq.sock"
poll_interval = 15
lease_seconds = 30
renew_interval = 10
max_attempts = 4
retry_backoff = [5, 15, 60, 300]
timeout = 5
busy_timeout = 5
policy_profile = "default"
allow = ["repo_inspect", "repo_edit", "local_checks", "state_updates", "tmux_wake"]
# Explicitly choose an endpoint if multiple sessions host the same repo:
# [coordination.endpoints]
# net-kingdom = "tmux-amq-12345"
```
`COORDINATION_CONFIG`, `COORDINATION_STATE_DIR`, `COORDINATION_SOCKET`,
`TAMQ_SOCKET`, and `STATEHUB_API_BASE` override file settings. State defaults to
`${XDG_STATE_HOME:-~/.local/state}/coordination-engine`; the control socket is
`${XDG_RUNTIME_DIR}/coordination-engine.sock`, or
`/tmp/coordination-<uid>/coordination-engine.sock` without XDG_RUNTIME_DIR.
The state directory must be owned by the service user and mode `0700`.
Both sockets must be same-user mode `0600`; Linux SO_PEERCRED is required.
TAMQ's own config file is not read: supply its socket explicitly if customized.
No secret belongs in this config. The current HTTP adapter uses the already
accessible local State Hub endpoint without inventing a credential mechanism.
If that endpoint requires credentials, leave the service stopped and resolve
its access through `warden route` and the owning subsystem. Authentication
failure keeps projections pending and prevents fresh observation/wakes.
## Run and inspect
```sh
coordination-engine serve # foreground; supervise using the local user identity
coordination-engine ping
coordination-engine status # JSON: adapter health, leases and local schema
coordination-engine history # local leases, including checkpoint content
coordination-engine stop # preserves leases, history and TAMQ terminals
coordination-engine once # one observation/dispatch pass; also takes writer lock
coordination-engine completion bash
```
Use `--config /path/to/config.toml` before the subcommand if needed. No service
is automatically installed or enabled by package installation. A stopped or
unavailable TAMQ endpoint is never implicitly opened. `once` can wake workers;
it is not a dry run. Ordinary diagnostics omit remote error text and payloads.
JSON status/history are local inspection surfaces and may contain checkpoints.
Only one service/writer may use a state directory. SIGINT/SIGTERM stops new
work, finishes the bounded adapter call, preserves state and removes the control
socket. With the default five-second adapter timeout, shutdown fits the
contract's ten-second grace period under normal local I/O. Restarting preserves
idempotency keys and receipt backlog. The synchronous local-alpha observer can
briefly delay control requests while polling; keep the selected worker set small
and lengthen leases for slower hubs. It is not a high-throughput scheduler.
## Actionability
The observer joins `/repos/`, `/workplans/`, paginated `/tasks/`, per-workplan
`dependencies/`, and per-repository unread `/messages/`. A complete successful
snapshot is required before dispatch. Hub outage backs polling off to 300 seconds;
TAMQ's independent local delivery continues. Reconnection is checked at the next
backoff deadline; restarting the service resets the delay.
Tasks must be `todo` or `progress` under a `ready` or `active` workplan. `wait`
tasks remain waiting until their authoritative state is changed. Dependencies
use the Hub's direction: `from_workplan_id` depends on `to_workplan_id` or
`to_task_id`. Only finished/archived workplans and done tasks satisfy an edge;
missing targets and unknown relationship types wait conservatively. This API
has no task-to-task dependency endpoint; task-targeted workplan edges are handled.
Only subjects beginning with `[action]` request an inbox wake; ordinary unread
messages are informational. Bodies are never copied into the wake prompt.
The worker must inspect the referenced source and recheck scope and authority.
This marker requests inspection and does not grant authority to its body.
`needs_human`, intervention notes, blocking reasons, disallowed structured
`action_classes`, and conservative sensitive-action text detection stop work
with receipts. The text detector is an additional stop mechanism, not an
authorization classifier; it can produce false positives. Unknown action
classes stop. The only supported v0.1 profile is `default`, optionally narrowed
through `allow`; sensitive-action retries are always zero. A changed source
revision may be reconsidered, but the same stopped revision is never retried.
Changing runtime configuration stops old pending leases instead of silently
reinterpreting their authority.
## Worker contract
Workers use the same approved Unix user and configuration. The socket verifies
peer UID; repo claims must be selected and currently gita registered. This is a
same-user trust boundary, not isolation between mutually untrusted processes.
```sh
coordination-engine ack LEASE_ID --repo net-kingdom
coordination-engine renew LEASE_ID --repo net-kingdom
coordination-engine checkpoint LEASE_ID --repo net-kingdom --file checkpoint.json
# Or, after the repository's normal file/State Hub updates:
coordination-engine complete LEASE_ID --repo net-kingdom --file checkpoint.json
```
Acknowledge before working; renew every 10 seconds by default. Updates for expired,
terminal, unknown or mismatched leases are rejected. The worker owns its renewal
loop. Expiry is not permission to keep working: stop or reestablish authority.
A checkpoint file is a local JSON object, at most 8 KiB:
```json
{"summary":"Parser fix checked", "files_changed":["src/parser.py"], "next_action":"Run the remaining local checks", "blocked_reason":null}
```
Keep credentials and sensitive values out of checkpoints. Checkpoint contents
stay in the private SQLite database; receipts contain only transition metadata.
`blocked_reason` may be `human`, `secret`, `destructive`, `policy`, `dependency`,
or `unavailable`; any value stops automatic continuation. A normal checkpoint
ends the old lease and creates a new checkpoint trigger/wake on the next poll.
Its `source_id` points to the parent trigger in local `history`, whose checkpoint
must be inspected before resuming. Recurring checkpoints form a local chain.
Transport failures/lease expiry retry with the **same** lease ID and identical
prompt, up to four attempts by default (configurable 09). This recovers admission
without creating a second TAMQ message. An already injected message is not
reinjected merely because the worker lease expired; exhaustion stops for review.
Explicit checkpoint continuation gets a **new** lease/message identity.
Only one unexpired offered/acknowledged/running lease per repo is admitted.
Changing task revision during an active lease does not wake a second worker.
Delivery states are recorded separately: `injected`/`acknowledged` do not complete
coordination work. TAMQ `failed` stops coordination; there is no automatic terminal
transport reset. The worker's completion or authoritative task `done` is completion;
other removed/changed sources retire stale leases as stopped.
## Retention, backup and receipts
SQLite uses WAL, foreign keys, transactional transitions and a schema version.
History is retained indefinitely; there is no automatic deletion. Existing tables
are backed up before a forward migration; newer schemas are refused.
```sh
coordination-engine db-version
coordination-engine backup
```
Backups are mode `0600`, consistent SQLite snapshots alongside the database. Copy
an explicit backup file to the operator's approved backup destination. To restore,
stop the service, preserve the entire old state directory (including WAL/SHM), and
place the chosen snapshot as `coordination.sqlite3` in a fresh mode-0700 state
directory. Configure that directory before restarting. Do not mix an old database
with current WAL/SHM files. TAMQ message history has its own independent backup and
purge procedures.
Every state transition and shutdown creates an audit receipt. Projection uses
`POST /progress/` and buffers failures locally. Delivery is **at least once**:
a crash after remote acceptance can repeat a projection; consumers should use
`detail.id` to deduplicate. An edge-relay queued response remains pending local
evidence. No message body, checkpoint text, credentials, or terminal output is
included. Local `status` reports pending receipt counts (up to the 100-row batch).
## Verification
```sh
make check
uv build # optional packaging check
COORDINATION_LIVE_SMOKE=1 uv run pytest -m live
```
The ordinary tests use temporary SQLite, fake HTTP/gita peers and real temporary
Unix sockets, including foreground service stop/restart. The live smoke is opt-in,
read-only discovery: registry, Hub repos, and TAMQ capabilities. It never injects
a wake. A production deployment or end-to-end agent run has not been performed
by this implementation task.