Refresh SCOPE.md intent gap analysis; add ACTIVITY-WP-0023
All checks were successful
CI Smoke / host-smoke (push) Successful in 1s
CI Smoke / container-smoke (push) Successful in 4s

Document production posture after WP-0020/0021, list INTENT↔practice gaps
G1–G10, and open a workplan for gap closure plus operational follow-ups.
This commit is contained in:
tegwick 2026-07-21 21:19:47 +02:00
parent bb7842dcdc
commit e73575dcd6
2 changed files with 436 additions and 258 deletions

440
SCOPE.md
View file

@ -1,7 +1,7 @@
--- ---
domain: capabilities domain: capabilities
repo: activity-core repo: activity-core
updated: "2026-06-16" updated: "2026-07-21"
--- ---
# SCOPE # SCOPE
@ -26,20 +26,27 @@ lifecycle.
An `ActivityDefinition` (a markdown file checked into a repo) declares a trigger An `ActivityDefinition` (a markdown file checked into a repo) declares a trigger
(cron, one-off scheduled datetime, or named event type), context sources to (cron, one-off scheduled datetime, or named event type), context sources to
resolve before evaluation, and a set of rules and instructions that determine resolve before evaluation, and a set of rules and instructions that determine
what tasks to create. When triggered, a durable Temporal workflow loads the what work or evidence to produce. When triggered, a durable Temporal workflow
definition, resolves context, evaluates the rule/instruction set, and emits task loads the definition, resolves context, evaluates the rule/instruction set, and
creation requests to issue-core or configured dry-run/audit sinks. Instructions emits outputs via configured sinks:
may also emit validated reports, and selected context resolvers may emit compact
non-secret evidence. Everything is auditable: the spawn log records the - **Tasks** — via `IssueSink` (`rest` → issue-core, `state-hub` progress, or
triggering event, matched rule/instruction metadata, model/prompt hash where `null` dry-run).
applicable, and resulting task references. - **Reports** — instruction reports to working memory and/or State Hub progress.
- **Evidence** — compact non-secret summaries from selected resolvers (ops
inventory, package prune, legacy-meter, etc.).
Everything is auditable: `activity_runs` records the activation; `task_spawn_log`
records emissions; report sinks leave progress/working-memory artifacts.
The two evaluation modes: The two evaluation modes:
- **Rule** — deterministic condition (sandboxed Python-like DSL) → fixed task - **Rule** — deterministic condition (sandboxed Python-like DSL) → fixed task
templates. Fast, testable, no LLM cost. templates. Fast, testable, no LLM cost.
- **Instruction** — optional pre-filter condition → LLM prompt with trusted - **Instruction** — optional pre-filter → LLM or **deterministic** report path
fields only → structured task list. For cases where the right tasks depend (`model: deterministic` / `none` / `unused`) → structured tasks and/or
on context that is easier to describe than to enumerate. reports. LLM reserved for genuine judgement; deterministic for pure context
summaries.
--- ---
@ -47,315 +54,232 @@ The two evaluation modes:
- **ActivityDefinition model**: trigger config (cron / scheduled / event), - **ActivityDefinition model**: trigger config (cron / scheduled / event),
context sources, rules (condition + action), instructions (trusted-field context sources, rules (condition + action), instructions (trusted-field
prompt + model + output schema). prompt + model + output schema + report sinks).
- **Event type registry**: publisher-declared markdown definitions with - **Event type registry**: publisher-declared markdown definitions with
attribute schemas, example payloads, and intent documentation. attribute schemas, example payloads, and intent documentation.
Curator-gating configurable per runtime environment. Curator-gating configurable per runtime environment.
- **Trigger types**: 5-field cron with timezone and misfire policy; one-off - **Trigger types**: 5-field cron with timezone and misfire policy; one-off
scheduled datetime; event-type subscription via NATS; manual one-shot API scheduled datetime; event-type subscription via NATS; manual one-shot API
trigger; one-shot schedule smoke tests for recurring definitions. trigger; one-shot schedule smoke tests for recurring definitions.
- **Context resolution adapters**: repo-scoping (repository capability queries), - **Schedule reconcile**: Temporal Schedule upsert/pause with per-row error
State Hub (domain/workplan state, SBOM status, daily triage digest, coding isolation (`sync_schedules` continues after `ScheduleAlreadyRunningError`).
retro read model), and ops inventory (bounded HTTP/HTTPS probes of a - **Context resolution adapters**: repo-scoping, State Hub (domain/workplan
non-secret service inventory). The adapter registry is extensible for other state, SBOM status, daily triage digest, coding retro, Binky rhythm status,
sources. consistency sweep, RecentlyOnScope hourly), shell queries (e.g. Forgejo
- **Rule evaluator**: sandboxed AST walker for Python-like boolean expressions package prune), and ops inventory probes. Registry is extensible.
over event attributes and resolved context. Rule actions support safe - **Rule evaluator**: sandboxed AST walker; `for_each` binding; no `exec()`.
`context.*` / `event.*` interpolation and explicit `for_each` per-item - **Instruction executor**: trusted-field rendering, llm-connect, structured
binding. No `exec()`. validation, item-granular recovery, deterministic report path, report sinks.
- **Instruction executor**: trusted-field prompt rendering, LLM call via - **Task emission adapter**: `IssueCoreRestSink`, `StateHubProgressSink`,
llm-connect, structured output validation, item-granular recovery with a `NullSink` selected by `ISSUE_SINK_TYPE`. Production railiance01 currently
quarantine lane and producer guardrails (count/length/depth caps, reference uses **`state-hub`** as the safe default for internal findings (see
allow-list) at the producer trust boundary, bounded validation-failure ACTIVITY-WP-0021/0022); restore `rest` only when issue-core→Forgejo is healthy
artifacts for report instructions, review-required audit metadata, and **and** policy allows external issues.
deterministic report sinks. A real downstream review queue is not implemented - **Report / evidence sinks**: working memory, State Hub progress
in this repo. (`daily_triage`, `sbom_staleness`, `forgejo_package_prune`,
- **Task emission adapter**: abstraction over issue-core; current transport is `ops_inventory_probe`, `activity_task_spawn`, etc.).
REST, with `ISSUE_SINK_TYPE=null` for dry-run/audit mode. It is designed to - **Spawn audit log**: `task_spawn_log` with rule/instruction id and sink ref.
migrate to a durable issue-core-owned NATS command boundary when issue-core - **Webhook receiver**: Gitea/GitHub → EventEnvelope → NATS.
provides that contract. - **Worker and workflow infrastructure**: Temporal `RunActivityWorkflow`
- **Report sinks**: instruction report outputs can be persisted to bounded (load → resolve → evaluate → report → **log_run** → emit). Emit failures no
local working memory and posted as State Hub progress events. These are longer skip `activity_runs` rows.
reporting outputs, not task lifecycle ownership. - **REST admin API** (FastAPI): CRUD, manual trigger, admin sync, metrics.
- **Ops evidence sinks**: `ops-inventory` context sources can post compact - **Automation status surface**: `make automation-status` and
non-secret `ops_inventory_probe` summaries to State Hub. Inter-Hub submission `make prod-automation-status` / `scripts/prod_automation_status.sh` for
is present only as a gated/deferred sink result until operator-owned railiance01 evidence without LLM authority.
`OPS_HUB_KEY` custody and widget mapping are ready. - **Operational runbook**: `docs/runbook.md`, emission boundary docs.
- **Spawn audit log**: every task emission recorded with rule/instruction id,
triggering event id, model and prompt hash (instructions), issue-core task ref.
- **Webhook receiver**: HTTP endpoint normalising inbound Gitea/GitHub webhook
payloads to EventEnvelope format and publishing to NATS.
- **Worker and workflow infrastructure**: Temporal-based durable execution —
`RunActivityWorkflow` orchestrates load → resolve → evaluate → emit.
- **REST admin API** (FastAPI): CRUD for ActivityDefinitions, manual trigger,
event type registry queries.
- **Prometheus metrics**: Temporal SDK metrics exposed for scraping.
- **Automation status surface**: deterministic, non-LLM status reporting via
`make automation-status` / `scripts/automation_status.py`, using repo-owned
evidence sources rather than coding assistant scheduler state.
- **Operational runbook**: `docs/runbook.md`.
--- ---
## Out of Scope ## Out of Scope
- **Task lifecycle** — creating, assigning, tracking, and closing tasks is - **Task lifecycle** — create/assign/track/close remains issue-core (or fleet
issue-core's responsibility. activity-core holds a spawn audit trail only. work-record connectors). activity-core holds a spawn audit trail only.
- **Project and initiative management** — phased, completion-gated, multi-step - **Project and initiative management** — project-core (future).
coordinated changes belong to project-core (future). - **General execution workers** — scans, patches, privileged API writes that
- **Execution of automatable tasks** — Temporal Activities that do real work are not explicitly declared as **bounded ActivityDefinition side-effects**
(run a scan, apply a patch, call an API) live in per-repo workers, not here. (e.g. retention prune with dry-run/apply gates) belong to per-repo workers
- **General ops execution** — Kubernetes, SSH, tunnel, authenticated service or platform tools.
checks, secret custody, OpenBao writes, and Inter-Hub widget/API-key - **General ops control plane** — Kubernetes, SSH, tunnels, secret custody,
provisioning belong to the owning operational repos and operator workflows. OpenBao policy administration.
activity-core may record non-secret probe evidence; it must not become the ops - **Event broker / Temporal server hosting** — consume, do not own lifecycle.
control plane. - **End-user task UI**.
- **Service inventory authority** — the Custodian inventory remains owned by - **Coding assistant schedulers** as production authority.
the custodian/state-hub surface. activity-core may read a projected
non-secret snapshot. **Boundary note (side-effect resolvers):** A small set of shell/context
- **Event broker hosting** — NATS JetStream is org infrastructure; activity-core resolvers intentionally invoke platform tools (e.g. Forgejo package prune).
consumes it but does not own its lifecycle. That is allowed only when the side-effect is the *declared purpose* of the
- **Temporal server hosting** — activity-core uses the Temporal SDK; the server ActivityDefinition, is credential-routed, evidence-posted, and not general
runs on Railiance infrastructure (or Docker Compose for dev). task execution. It must not expand into an unbounded ops executor.
- **End-user task UI** — tasks land in issue-core; presentation is separate.
- **Synchronous request-response patterns** — Temporal is async-first.
- **Coding assistant automation infrastructure** — assistant-provided reminders,
heartbeats, or scheduled jobs are not the execution or evidence authority for
activity-core automations. Assistants may run and summarize repo-native
commands only.
--- ---
## Relevant When ## Relevant When
- You need org-wide recurring maintenance automation (dependency scans, SBOM - Org-wide recurring maintenance (SBOM, package prune, legacy-meter, Binky
checks, staleness audits) without bespoke per-service cron jobs. rhythm, consistency sweeps) without bespoke per-service cron.
- You need reactive task generation: "when X happens across the org, create - Reactive task/report generation from org events and context.
structured tasks in the right repos." - One-off future scheduling.
- You need one-off future task scheduling without a separate reminder system. - Auditable activation history (runs, spawn log, progress).
- You want an auditable record of what triggered what and why. - Answering “how did automations go since Sunday?” from deterministic
- You need a scheduled, non-secret evidence note proving that declared service repo-native / prod evidence.
endpoints or access paths were observed, without executing privileged ops
commands.
- You are replacing scattered bespoke cron jobs and manual coordination with
a governed, observable automation layer.
- You need to answer "how did our automations go since Friday?" from
deterministic repo-native evidence before any optional LLM summary.
---
## Not Relevant When ## Not Relevant When
- You need to track whether a task is done, blocked, or reassigned → issue-core. - Tracking task completion/reassignment → issue-core / work records.
- You need to coordinate a multi-phase project with dependencies → project-core. - Multi-phase project orchestration → project-core.
- You need a simple system cron with no durability requirement. - Simple host cron with no durability or evidence needs.
- You need synchronous request-response patterns.
--- ---
## Current State ## Current State (2026-07-21)
- **Status**: active production-backed service with two visible open gates: ### Status
`ACTIVITY-WP-0006` still waits on three clean consecutive scheduled daily
triage runs and calibration feedback, and `ACTIVITY-WP-0008` is blocked until Production on **railiance01** (`activity-core` namespace). Recent finished
Helix Forge publishes the upstream `coding_retro` read model needed to enable workplans:
the Saturday schedule. `ACTIVITY-WP-0007` is finished: the bounded
ops-inventory probe/evidence slice has live Railiance evidence. | WP | Outcome |
- **Implementation**: core is functional. `RunActivityWorkflow`, | --- | --- |
`TaskExecutorWorkflow` (stub), PostgreSQL schema, Temporal Schedules and smoke | ACTIVITY-WP-0018 | Own-infra automation status CLI |
schedules, NATS Event Router, FastAPI admin API, Prometheus metrics, event | ACTIVITY-WP-0019 | Automation inventory targets |
type registry, markdown ActivityDefinition parser/sync, rule evaluator, | ACTIVITY-WP-0020 | Weekly Forgejo package prune **enabled**; first apply 38 deletes |
instruction executor, context resolvers, issue sink, report sinks, ops | ACTIVITY-WP-0021 | Prod reliability: state-hub sink, log_run-before-emit, schedule harden, DB probes, prod status path |
evidence sink, Kubernetes deployment, and operational runbook are all | ACTIVITY-WP-0006/0008 | Marked finished historically; enablement evidence has moved |
implemented.
- **Current definitions**: `weekly-sbom-staleness` is enabled and demonstrates Open product/policy workplan: **ACTIVITY-WP-0022** (IssueSink no-default-Forgejo).
the deterministic rule/fan-out path. `weekly-coding-retro` is present and
tested but intentionally disabled until live `coding_retro` evidence exists. ### Live posture (railiance01)
Railiance projects the daily State Hub WSJF triage definition and the disabled
ops-service-inventory probe definition from the runtime bundle. - Image: `activity-core:railiance01-prod` (locally imported).
- **Operational proof**: the State Hub daily WSJF triage path has produced - `ISSUE_SINK_TYPE=state-hub` (internal findings → State Hub progress, not
validated reports and working-memory notes, but the calibration gate is not Forgejo issues).
closed. A 2026-06-16 recheck found State Hub `daily_triage` progress and - Edge relay: `actcore-statehub-edge-relay` → in-cluster `state-hub`.
working-memory `daily-triage-*` notes only through 2026-06-06, so there is not - Weekly prune: worker has `FORGEJO_TOKEN` + hostPath `/opt/railiance-platform`
yet evidence for three clean consecutive scheduled runs after the June 7 + `live_images_file` (kubectl absent in pod).
runtime projection failure. The ops inventory probe path has live fallback - Proven one-shots (2026-07-21): Binky daily/mail COMPLETED with tasks_spawned=1;
evidence in State Hub; Inter-Hub per-entity submission remains deferred. SBOM deterministic report; daily triage progress; prune apply.
- **Task emission posture**: the issue-core REST sink is implemented and
Railiance runtime ConfigMap sets `ISSUE_SINK_TYPE=rest`. `ISSUE_CORE_API_KEY` ### Definitions (repo files + prod DB)
is synced into `actcore-runtime-secret` by ExternalSecret
`actcore-issue-core-runtime` (OpenBao path shared with issue-core). Bootstrap | Definition | Role |
the ESO token with `make openbao-eso-token-apply`. `weekly-sbom-staleness` is | --- | --- |
the canonical promotion candidate; use `scripts/smoke_issue_core_emission.py` | `weekly-sbom-staleness` | Deterministic SBOM report; task emit gated off |
before trusting production task creation. | `weekly-forgejo-package-prune` | Sunday 03:30 UTC apply prune + evidence |
- **Stability**: construction risk has shifted to operational hardening and | `weekly-legacy-meter-review` / `legacy-meter-8h-capture` | Meter evidence |
adoption risk. The last recorded full-suite pass in the workplans was | `weekly-coding-retro` | Present; enablement depends on `coding_retro` feed |
2026-06-04 (`128 passed, 1 skipped`), with later targeted coverage added for | Binky daily / mail / review-prep | Rhythm tasks via IssueSink |
ops inventory, ops evidence sinks, Railiance projection wiring, and weekly | Railiance ConfigMap bundle | Daily triage, ROS hourly, consistency sweep, etc. |
coding retro parsing/rule behavior.
- **Next**: close `ACTIVITY-WP-0006-T03` with real scheduled-run calibration ### Task emission posture
evidence; close `ACTIVITY-WP-0008-T03` once upstream `coding_retro` publication
exists and the dry-run/duplicate check passes; decide when to move selected - **Canonical REST sink** to issue-core remains implemented and documented.
task/report/evidence sinks from dry-run or fallback mode to their intended - **Production default is `state-hub`** after Forgejo PAT/backend failures and
live backends. fleet policy against silent Forgejo issues (WP-0021 path B, WP-0022).
- Restoring `rest` requires healthy issue-core→Forgejo **and** explicit policy
opt-in per definition.
--- ---
## Assessment Against Intent ## Assessment Against Intent
activity-core now matches the core intent: it answers **when** coordination ### Alignment
work should happen, **what** work should be created from current org context,
and **where** each work item should land. The daily WSJF triage is the clearest
judgement-oriented proof point; weekly SBOM staleness is the clearest
deterministic-rule proof point.
The governing boundary still matters. activity-core should keep owning trigger | INTENT promise | Status |
durability, context resolution, rule/instruction evaluation, report/task | --- | --- |
emission, and spawn/report audit. It should not become the task lifecycle | Answers **when / what / where** | Met — Temporal schedules, rules/instructions, sinks |
database, the project planner, or a general execution worker. The local | Markdown definitions | Met — `activity-definitions/`, event-types, external ConfigMaps |
`TaskExecutorWorkflow` remains a stub and should stay that way unless a future | Rules before instructions | Met — SBOM/prune/Binky rules; triage uses LLM when needed |
workplan explicitly rehomes execution responsibility. | Durable recurring automation | Met — railiance01 schedules fire; status surface exists |
| Audit trail | Met — `activity_runs`, spawn log, progress/working-memory |
| Does not own task lifecycle | Met — no completion tracking |
| Does not own project phases | Met |
One boundary nuance is now explicit: activity-core may post State Hub progress ### Gaps and tensions (INTENT ↔ practice)
events as a configured report or evidence sink. That is acceptable because it
records the result of an activity-core activation; it is not ownership of State
Hub state, task lifecycle, or workplan planning.
The main drift risk is convenience creep: adding direct task tracking, | Gap | Severity | Notes |
project-phase state, or bespoke operational scripts because the Temporal | --- | --- | --- |
substrate is already nearby. Future work should prefer declarative | **G1. Sink default vs INTENT “emit to issue-core”** | High (policy) | INTENT names issue-core as the task landing zone. Practice and work-record canon push **State Hub / repo work records** for internal findings. WP-0022 owns policy; code already supports `state-hub` / `null` / `rest`. |
ActivityDefinitions, bounded context resolvers, and outbound adapters over | **G2. Executor gap** | High (value) | Tasks are emitted (progress refs or issue-core), but **no per-repo executor** reliably consumes them (Binky briefs still need harness/agent pickup). INTENT assigns execution to per-repo Temporal workers — that substrate is thin. |
new one-off control paths. | **G3. Side-effect resolvers** | Medium | Forgejo prune `apply: true` runs platform deletes inside a context resolver. Allowed only as declared, gated side-effect; needs clearer SCOPE language (above) and operator refresh of live-images protection. |
| **G4. Review queue** | Medium | `review_required` remains metadata-only; no downstream review queue. |
| **G5. TaskExecutorWorkflow stub** | Lowmedium | Still registered; risk of attracting real execution. |
| **G6. ROS digest completeness** | Medium (ops) | Template file mounted; MarkiTect CLI missing in state-hub image → digests still fail per domain. |
| **G7. Credential delivery** | Medium (ops) | `FORGEJO_TOKEN` bootstrap-injected; ESO 403 on forgejo-admin path. issue-core `GITEA_BACKEND_TOKEN` still broken for path A rest sink. |
| **G8. Live-images hygiene** | Medium (ops) | Multi-cluster `live-images-all.txt` must be refreshed after deploys or prune can delete live tags (incident 2026-07-21, restored). |
| **G9. Evidence federation** | Low | Progress often lands on railiance01 edge/hub; workstation primary hub may not show the same feed without tunnel/outbox health. |
| **G10. API external access** | Low | ClusterIP-only; intentional until auth policy. |
## Known Gaps Against Intent ### Drift risks
- **Scheduled-run trust gap**: INTENT promises recurring coordination work that 1. **Convenience execution** — growing shell resolvers into a general ops runner.
runs without Bernd as the manual coordination layer. The daily triage path is 2. **Silent sink restore** — flipping `ISSUE_SINK_TYPE=rest` without policy +
implemented, but its current calibration task still lacks three clean healthy Forgejo backend reintroduces issue spam / 503s.
consecutive scheduled runs after the June 7 runtime failure. Until that closes, 3. **Protection file staleness** — weekly prune with `apply: true` depends on
daily triage remains a production-backed capability with an evidence gap, not fresh live-image lists.
a fully proven standing substrate. 4. **Stub workflow attraction**`TaskExecutorWorkflow` looks like an
- **Task creation gap**: INTENT says activations emit task creation requests to execution home.
issue-core. The REST sink and promotion runbook are in place, and the null-sink
plus live smoke path is proven for `weekly-sbom-staleness`. Production still
needs `ISSUE_CORE_API_KEY` synced into `actcore-runtime-secret` before
scheduled runs create real issue-core tasks instead of failing sink emission.
- **Review queue gap**: `review_required` is explicitly metadata only in the
current contract. No issue-core review queue integration exists here, so any
future queue routing needs a downstream issue-core contract before high-impact
instruction outputs rely on it.
- **Evidence backend posture**: the State Hub fallback evidence path is the
accepted current backend for `ops_inventory_probe`. Inter-Hub/ops-hub
submission is deliberately deferred behind `OPS_HUB_KEY`, widget mapping, and
operator approval, so per-entity ops evidence publication is future work.
- **Execution-boundary residue**: `TaskExecutorWorkflow` is still registered as
a stub that writes a done `task_instances` row. It should remain inert or be
removed/re-homed before it attracts real execution work, because execution is
explicitly outside activity-core's intent.
- **API exposure posture**: the FastAPI surface stays ClusterIP-only for now.
External ingress remains future work until an authenticated access policy is
designed.
--- ---
## How It Fits ## How It Fits
``` ```
[NATS JetStream] ← publishers: State Hub, Gitea webhooks, Temporal signals, cron [NATS JetStream / Temporal Schedules / manual trigger]
[activity-core] ← event type registry, rule evaluator, instruction executor [activity-core] Event type registry · rules · instructions · context resolvers
[activity-core] → [issue-core] → [repos/services]
[activity-core] → [report/evidence sinks] → [State Hub / working memory / future Inter-Hub] [IssueSink: rest | state-hub | null] → issue-core or State Hub progress
[report/evidence sinks] → State Hub / working memory
[bounded shell side-effects] → platform tools (e.g. package prune)
``` ```
- **Upstream**: NATS (event bus), Temporal (durable workflow engine), PostgreSQL - **Upstream**: NATS, Temporal, PostgreSQL, State Hub, repo-scoping, llm-connect.
(definitions and audit log), repo-scoping (context adapter), State Hub (context - **Downstream**: issue-core (optional), State Hub, platform tools, future
adapter and event publisher). per-repo executors / agent-harness.
- **Downstream**: issue-core (task management) and configured report/evidence sinks. - **Ops**: railiance01 k3s; workstation status via SSH helper.
Agents and humans pick up tasks from issue-core and do the actual work.
Railiance may use the null sink for dry-run/audit mode until live issue-core
emission is approved.
- **Coordinates with**: the state hub delegates maintenance automations to
activity-core by publishing lifecycle events or by being resolved as context.
activity-core may post progress events as report/evidence outputs, but it
does not own State Hub task/workplan state.
--- ---
## Terminology ## Terminology
- **ActivityDefinition** — a markdown file declaring trigger, context sources, - **ActivityDefinition** — markdown policy unit (trigger + context + rules/instructions).
rules, and instructions. The unit of automation policy. - **EventEnvelope** — canonical inbound event shape.
- **EventEnvelope** — the canonical internal event format; normalises all inbound - **Rule / Instruction** — deterministic vs judgement (or deterministic report) paths.
events (NATS, webhooks, cron signals) to a common structure. - **IssueSink** — task emission adapter (`rest` / `state-hub` / `null`).
- **Rule** — deterministic condition expression + task template action. Evaluated - **Report sink / Evidence sink** — non-task activation outputs.
by a sandboxed AST walker. - **Spawn audit trail** — local emission record; not authoritative task state.
- **Instruction** — LLM-evaluated task generation with trusted-field prompt - **Activity** (Temporal) vs **ActivityDefinition** (app policy) — different concepts.
interpolation and structured output schema enforcement.
- **Report sink** — configured persistence for instruction reports, currently
working-memory markdown notes and State Hub progress events.
- **Evidence sink** — configured persistence for compact non-secret resolver
evidence, currently State Hub progress for ops inventory probes; Inter-Hub is
a deferred gated target.
- **Event type** — a registered, schema-documented category of event (e.g.
`org.repo.registered`). Publisher-declared; curator-gated per environment.
- **Spawn audit trail** — activity-core's local record of what tasks were emitted,
to which issue-core backend, and under which rule/instruction. Not the
authoritative task record.
- Potentially confusing: **Activity** (Temporal concept — a single executable
step in a workflow) vs. **ActivityDefinition** (app concept — the policy record
that governs what gets spawned). These are different things.
--- ---
## Related / Overlapping ## Related / Overlapping
- `issue-core` (formerly issue-facade) — downstream task management; receives - `issue-core` — task connector / external tracker mapping (not fleet work-record origin).
all task emission from activity-core. - `repo-scoping`, State Hub / the-custodian, `llm-connect`, `agent-harness`.
- `repo-scoping` — context adapter for repository capability queries. - `railiance-platform` — Forgejo package prune tool and credential lanes.
- `the-custodian` / State Hub — context adapter for domain state; delegates - `ops-warden` — credential routing (not secret vending for non-SSH).
maintenance automation to activity-core via NATS events. - `project-core` (future), `ops-bridge` tunnels.
- `llm-connect` — instruction execution backend for judgement-oriented reports
such as daily State Hub WSJF triage.
- `core-hub` / `ops-hub` — production ops evidence intake (`hub.coulomb.social`);
historical Haskell Inter-Hub is `inter-hub-haskell` (retired 2026-07-08).
- `rules-core` (future extraction) — the rule evaluator and instruction executor
module, currently in `src/activity_core/rules/`.
- `project-core` (future) — project and initiative management; will use
activity-core to generate per-phase task sets.
- `ops-bridge` — SSH tunnel to remote Temporal server on Railiance.
--- ---
## Architecture Decisions ## Architecture Decisions
- `docs/adr/adr-001-event-bridge-architecture.md` — overall Event Bridge pattern, - `docs/adr/adr-001-event-bridge-architecture.md`
boundaries, state hub relationship, domain assignment. - `docs/adr/adr-002-definition-format.md`
- `docs/adr/adr-002-definition-format.md` — markdown-as-definition format, - `docs/adr/adr-003-rule-instruction-model.md`
governance model, event type schema, ActivityDefinition structure. - `docs/adr/adr-004-producer-trust-boundary.md`
- `docs/adr/adr-003-rule-instruction-model.md` — Rule DSL, Instruction safety
model, evaluation semantics, audit trail, testing strategy.
- `docs/adr/adr-004-producer-trust-boundary.md` — untrusted-producer premise,
trust-but-handle vs verify-and-mitigate postures, error-locality and
quarantine-with-provenance, producer guardrails for LLM/agent/human output.
--- ---
## Getting Oriented ## Getting Oriented
- Start with: `INTENT.md` (why), this file (what), `docs/adr/` (decisions). - Start: `INTENT.md`, this file, `docs/adr/`, `docs/runbook.md`.
- Key source files: `src/activity_core/models.py` (domain model), - Key code: `workflows.py`, `activities.py`, `issue_sink.py`, `report_sinks.py`,
`src/activity_core/workflows.py` (RunActivityWorkflow), `context_resolvers/`, `automation_status.py`, `sync_schedules.py`.
`src/activity_core/activities.py` (Temporal activities), - Definitions: `activity-definitions/`, `event-types/`, Railiance ConfigMaps.
`src/activity_core/event_router.py` (NATS → Temporal), - Status: `make automation-status SINCE=sunday` ·
`src/activity_core/schedule_manager.py` (Temporal Schedules), `make prod-automation-status SINCE=sunday`.
`src/activity_core/api.py` (FastAPI admin),
`src/activity_core/report_sinks.py` (instruction reports),
`src/activity_core/ops_evidence_sinks.py` (ops evidence),
and `src/activity_core/context_resolvers/` (external context adapters).
- Definition files: `event-types/`, `activity-definitions/`, and `tasks/`.
- Dev environment: `docker-compose.dev.yml` (Temporal + PostgreSQL + NATS).
- Entry points: `uv run python -m activity_core.worker` (Temporal worker),
`uv run uvicorn activity_core.api:app --port 8010` (admin API).
--- ---

View file

@ -0,0 +1,254 @@
---
id: ACTIVITY-WP-0023
type: workplan
title: "Intentscope gap closure and WP-0020/0021 operational follow-ups"
domain: infotech
repo: activity-core
status: proposed
owner: codex
topic_slug: activity-core
created: "2026-07-21"
updated: "2026-07-21"
---
# Intentscope gap closure and WP-0020/0021 operational follow-ups
## Origin
2026-07-21 gap analysis of `INTENT.md` against production practice and the
refreshed `SCOPE.md` (same day). Captures:
1. **Product/architecture gaps** vs the governing when/what/where principle.
2. **Non-blocking operational follow-ups** from ACTIVITY-WP-0020 (Forgejo prune)
and ACTIVITY-WP-0021 (prod reliability).
Does **not** re-implement ACTIVITY-WP-0022 (IssueSink default policy) — that
workplan owns G1 policy decisions; this workplan **consumes** its outcomes and
covers adjacent gaps.
## Gap summary (from SCOPE)
| ID | Gap | Intent tension |
| --- | --- | --- |
| G1 | Sink default vs “emit to issue-core” | Policy — **ACTIVITY-WP-0022** |
| G2 | Executor gap (tasks emit, work not executed) | INTENT assigns execution to per-repo workers |
| G3 | Side-effect resolvers (prune apply) | Must stay bounded/declared |
| G4 | `review_required` has no queue | Downstream contract missing |
| G5 | `TaskExecutorWorkflow` stub attraction | Execution boundary residue |
| G6 | ROS digests need MarkiTect | Ops completeness |
| G7 | Credential delivery (FORGEJO_TOKEN ESO; issue-core Gitea PAT) | Path A / weekly prune durability |
| G8 | `live-images-all.txt` hygiene | Prune safety after deploys |
| G9 | Evidence federation workstation vs edge hub | Operator visibility |
| G10 | ClusterIP-only API | Deferred auth |
## Goal
Close or explicitly disposition each gap with either shipped code/docs,
operator-owned policy (WP-0022), or a deliberate “wontfix / external owner”
note — so SCOPE stays truthful and INTENTs three questions remain load-bearing.
## Out of scope
- Implementing fleet work-record promotion (state-hub / the-custodian).
- Full agent-harness product roadmap (only the activity-core emission contract).
- OpenBao policy authorship for unrelated workloads.
- Re-opening finished WP-0006/0008 calibration narratives unless evidence is
still broken after WP-0021 triage fixes.
## Tasks
## Task: Align sink matrix with INTENT and WP-0022
```task
id: ACTIVITY-WP-0023-T01
status: wait
priority: high
```
**Depends on ACTIVITY-WP-0022-T01** (default sink policy decision).
1. After WP-0022 decides defaults, update `INTENT.md` wording if the fleet no
longer treats issue-core as the **default** landing zone (keep issue-core as
optional external connector).
2. Ensure `docs/issue-core-emission-boundary.md` + runbook sink matrix match:
`null` / `state-hub` / `rest` + per-definition opt-in.
3. Confirm railiance01 env and ConfigMap defaults match the decision (today:
`state-hub`).
**Done when:** INTENT + emission docs + prod defaults tell one story; WP-0022
tasks for docs/defaults are cross-linked as done.
## Task: Executor contract for emitted tasks (Binky first)
```task
id: ACTIVITY-WP-0023-T02
status: todo
priority: high
```
INTENT: execution lives in per-repo workers / harness — not activity-core.
1. Document the **consumer contract** for `activity_task_spawn` (state-hub) and
issue-core REST refs: required fields, idempotency, target_repo.
2. Prove one end-to-end path for **Binky daily brief**: emit → consumer
(agent-harness or documented human lane) → `binky_daily_brief` progress /
brief artifact.
3. File or link harness/railiance work if the consumer is out of this repo;
do not implement a real `TaskExecutorWorkflow` here.
**Done when:** A scheduled or one-shot Binky daily produces both a spawn record
**and** a Binky-side completion evidence note within one business day of fire.
## Task: Side-effect resolver guardrails
```task
id: ACTIVITY-WP-0023-T03
status: todo
priority: medium
```
1. Codify in runbook which shell queries may set `apply: true` (today: Forgejo
package prune only).
2. Hard-fail prune apply when `live_images_file` is missing or empty when
`apply: true` (prevents 2026-07-21 worker incident recurrence).
3. Optional: unit test that apply without protection file is rejected.
**Done when:** apply path cannot run without an explicit non-empty protection
source; docs list allowed side-effect definitions.
## Task: Live-images protection refresh automation
```task
id: ACTIVITY-WP-0023-T04
status: todo
priority: medium
```
Non-blocking follow-up from ACTIVITY-WP-0020.
1. Script (railiance-platform or activity-core `scripts/`) that exports live
forgejo images from coulombcore + railiance01 and merges to
`railiance-platform/docs/evidence/live-images-all.txt` on the worker host.
2. Document cadence (post-deploy checklist or weekly pre-prune cron on host).
3. Prefer non-secret output only (image refs, no tokens).
**Done when:** operator can refresh protection with one documented command;
runbook links it from the prune section.
## Task: FORGEJO_TOKEN via ESO (retire bootstrap inject)
```task
id: ACTIVITY-WP-0023-T05
status: todo
priority: medium
```
1. OpenBao / railiance-platform: grant activity-core ESO token read on
`platform/workloads/forgejo/forgejo-admin` **or** a dedicated
`activity-core` prune token with package write only.
2. Land ExternalSecret merge into `actcore-runtime-secret` `FORGEJO_TOKEN`.
3. Remove reliance on one-shot `kubectl patch secret` bootstrap.
**Done when:** ExternalSecret Ready=True and worker restarts still see a
working token without manual patch.
## Task: issue-core path A (optional rest sink)
```task
id: ACTIVITY-WP-0023-T06
status: todo
priority: low
```
Non-blocking follow-up from ACTIVITY-WP-0021.
1. Rotate/fix issue-core `GITEA_BACKEND_TOKEN` so `POST /issues/` returns 201
(not 503 forgejo-inbox).
2. Smoke from actcore-worker with `ISSUE_SINK_TYPE=rest` only for definitions
that opt in (per WP-0022).
3. Do **not** flip global default to rest without WP-0022.
**Done when:** documented smoke 201 for one opt-in definition, or explicit
wontfix with issue-core owning the backend fix only.
## Task: ROS MarkiTect completeness
```task
id: ACTIVITY-WP-0023-T07
status: todo
priority: medium
```
1. Coordinate with state-hub image: ship MarkiTect / set
`STATE_HUB_MARKITECT_CLI_PATH` so `/recently-on-scope/hourly` does not fail
per-domain with MarkiTect-required errors.
2. Template `COPY templates/` already landed; verify after image roll.
3. Re-trigger ROS hourly; confirm `failed` list empty or only non-template
reasons.
**Done when:** one clean ROS hourly context snapshot without missing-template
or MarkiTect-missing errors for configured domains.
## Task: Retire or quarantine TaskExecutorWorkflow
```task
id: ACTIVITY-WP-0023-T08
status: todo
priority: low
```
1. Decide: remove registration from worker, or rename/document as
legacy-no-op with metrics alert if used.
2. Ensure no ActivityDefinition or API path depends on it.
3. Update SCOPE if removed.
**Done when:** stub cannot attract production execution by accident.
## Task: Review-required disposition
```task
id: ACTIVITY-WP-0023-T09
status: todo
priority: low
```
1. Document that `review_required` is metadata-only until issue-core or
work-record review lane exists.
2. Either file a cross-repo task for a real review queue **or** mark as
wontfix / future with link in SCOPE known gaps.
**Done when:** SCOPE G4 has a disposition date and owner, not an open lie.
## Task: Evidence visibility note (edge vs workstation hub)
```task
id: ACTIVITY-WP-0023-T10
status: todo
priority: low
```
1. Document in runbook: prod progress is on railiance01 state-hub/edge; use
edge query or tunnel, not only workstation `127.0.0.1:8000` history.
2. Optional: `prod-automation-status` prints last N edge progress event types
for `daily_triage` / `forgejo_package_prune` / `activity_task_spawn`.
**Done when:** an operator following runbook can find post-fire evidence without
guessing hub topology.
## Success criteria
- [ ] G1 closed via WP-0022 + T01 doc alignment
- [ ] G2 has a proven Binky executor path or explicit external workplan
- [ ] G3 apply-without-protection is impossible
- [ ] G6G8 operational follow-ups shipped or scheduled with owners
- [ ] G5/G9/G10 dispositioned
- [ ] `SCOPE.md` re-checked after close; no stale “rest default / WP-0006 gate”
language
## References
- `INTENT.md`, `SCOPE.md` (2026-07-21 refresh)
- `ACTIVITY-WP-0020` closeout (prune enable, live-images, token bootstrap)
- `ACTIVITY-WP-0021` closeout (state-hub sink, schedule harden, ROS template)
- `ACTIVITY-WP-0022` IssueSink no-default-Forgejo
- `docs/issue-core-emission-boundary.md`