Repair production automation truth and schedule cleanup
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 1s
Build and Publish Container Image / build-and-push (push) Successful in 18s

This commit is contained in:
tegwick 2026-08-20 11:20:23 +02:00
parent 8bcb416285
commit 944fd158de
19 changed files with 441 additions and 64 deletions

View file

@ -34,7 +34,9 @@
| workplan | ACTIVITY-WP-0028 | finished | — | workplans/ACTIVITY-WP-0028-activity-review-cli.md |
| workplan | ACTIVITY-WP-0029 | active | — | workplans/ACTIVITY-WP-0029-hub-port-alignment.md |
| workplan | ACTIVITY-WP-0030 | active | — | workplans/ACTIVITY-WP-0030-daily-sbom-catchup.md |
| workplan | ACTIVITY-WP-0031 | active | — | workplans/ACTIVITY-WP-0031-production-execution-reliability-cleanup.md |
| workplan | ADHOC-2026-06-01 | finished | — | workplans/ADHOC-2026-06-01.md |
| workplan | ADHOC-2026-08-20 | finished | — | workplans/ADHOC-2026-08-20.md |
| workplan | custodian-WP-0001 | done | — | workplans/custodian-WP-0001-temporal-backbone.md |
| workplan | custodian-WP-0002 | done | — | workplans/custodian-WP-0002-triggers-ops.md |
| workplan | custodian-WP-0003a | done | — | workplans/custodian-WP-0003a-event-bridge-model-rules-registry.md |
@ -184,11 +186,17 @@
| task | ACTIVITY-WP-0029-T04 | done | — | workplans/ACTIVITY-WP-0029-hub-port-alignment.md |
| task | ACTIVITY-WP-0030-T01 | wait | — | workplans/ACTIVITY-WP-0030-daily-sbom-catchup.md |
| task | ACTIVITY-WP-0030-T02 | wait | — | workplans/ACTIVITY-WP-0030-daily-sbom-catchup.md |
| task | ACTIVITY-WP-0030-T03 | wait | — | workplans/ACTIVITY-WP-0030-daily-sbom-catchup.md |
| task | ACTIVITY-WP-0030-T03 | progress | — | workplans/ACTIVITY-WP-0030-daily-sbom-catchup.md |
| task | ACTIVITY-WP-0030-T04 | wait | — | workplans/ACTIVITY-WP-0030-daily-sbom-catchup.md |
| task | ACTIVITY-WP-0031-T01 | wait | — | workplans/ACTIVITY-WP-0031-production-execution-reliability-cleanup.md |
| task | ACTIVITY-WP-0031-T02 | done | — | workplans/ACTIVITY-WP-0031-production-execution-reliability-cleanup.md |
| task | ACTIVITY-WP-0031-T03 | done | — | workplans/ACTIVITY-WP-0031-production-execution-reliability-cleanup.md |
| task | ACTIVITY-WP-0031-T04 | done | — | workplans/ACTIVITY-WP-0031-production-execution-reliability-cleanup.md |
| task | ACTIVITY-WP-0031-T05 | wait | — | workplans/ACTIVITY-WP-0031-production-execution-reliability-cleanup.md |
| task | ADHOC-2026-06-01-T01 | done | — | workplans/ADHOC-2026-06-01.md |
| task | ADHOC-2026-06-01-T02 | done | — | workplans/ADHOC-2026-06-01.md |
| task | ADHOC-2026-06-01-T03 | done | — | workplans/ADHOC-2026-06-01.md |
| task | ADHOC-2026-08-20-T01 | done | — | workplans/ADHOC-2026-08-20.md |
| task | T01 | done | — | workplans/custodian-WP-0001-temporal-backbone.md |
| task | T02 | done | — | workplans/custodian-WP-0001-temporal-backbone.md |
| task | T03 | done | — | workplans/custodian-WP-0001-temporal-backbone.md |

View file

@ -1,10 +1,10 @@
---
id: weekly-sbom-staleness
name: Weekly SBOM Staleness Check
enabled: true
enabled: false
owner: custodian-agent
governance: custodian
status: active
status: paused
trigger:
type: cron
cron_expression: "0 9 * * 1"
@ -23,8 +23,15 @@ context_sources:
# Weekly SBOM Staleness Check
Runs every Monday at 09:00 Berlin time. Checks all tracked repositories for
SBOM staleness and flags any repository whose SBOM is older than 30 days.
> **Emergency-paused 2026-08-20 (ACTIVITY-WP-0031).** The 2026-08-17 fire
> emitted 75 independent rescan tasks while the bounded replacement in
> ACTIVITY-WP-0030 is blocked on CUST-WP-0062. Keep this definition disabled in
> source and production. Its deterministic weekly summary is not sufficient
> reason to retain the unbounded `for_each` task fan-out.
When enabled, this runs every Monday at 09:00 Berlin time. It checks all
tracked repositories for SBOM staleness and flags any repository whose SBOM is
older than 30 days.
ACTIVITY-WP-0021: the fleet no longer treats Forgejo issues as the primary
landing zone for automated tasks. The deterministic **state-hub-progress**

View file

@ -24,15 +24,24 @@ produces at most one workflow execution.
A *misfire* occurs when a scheduled trigger fires after its nominal time because the
worker was down, overloaded, or the schedule was paused.
| Policy | `dedupe_key_strategy` value | Behaviour |
| Policy | `trigger.misfire_policy` value | Behaviour |
|---|---|---|
| Skip | `"skip"` | Missed runs are discarded. Temporal `ScheduleOverlapPolicy.SKIP`. |
| Catch up | `"catchup"` | Missed runs are replayed up to 10 times (configurable). Uses `schedule.handle.backfill()`. |
| Compress | `"compress"` | One run is executed with a widened context window covering all missed intervals. |
The `dedupe_key_strategy` on `ActivityDefinition` must match the `misfire_policy` on
`CronTriggerConfig`. They are separate fields to allow the schedule configuration
(how Temporal fires) to be decoupled from the workflow logic (how missed runs are handled).
`trigger.misfire_policy` is the runtime control. The persisted
`dedupe_key_strategy` field is legacy API/DB compatibility metadata; schedule
execution does not read it. In particular, `dedupe_key_strategy: skip` does
**not** suppress a later scheduled fire whose resolved context or emitted task
content matches an earlier fire. New definitions should configure only the
trigger policy and should not rely on `dedupe_key_strategy` for content
deduplication.
Each nominal cron time produces a different workflow ID. If an hourly rule sees
the same unread message on consecutive hours, both hours may emit work. That
requires an explicit domain idempotency key or consumer-side state transition;
it is not a Temporal misfire concern.
---

View file

@ -47,12 +47,13 @@ data:
id: "e274defb-28f2-571e-abcb-c17b57eab473"
name: "RMASTER-WP-0020 OpenBao Retention Closeout"
type: activity-definition
version: "1.0"
enabled: true
version: "1.1"
enabled: false
owner: railiance-master
governance: custodian
status: active
status: paused
created: "2026-08-04"
updated: "2026-08-20"
trigger:
type: scheduled
at: "2026-08-17T08:00:00+02:00"
@ -61,7 +62,11 @@ data:
# RMASTER-WP-0020 OpenBao retention closeout
Emits a claimable ops run after the 14-day retention window. It never
The one-shot fired on 2026-08-17 and is now disabled so schedule sync does
not continually try to recreate a completed Temporal schedule. Its open
ops run remains operator-visible, but is intentionally not labelled
`automated`: railiance01 has neither a railiance-master checkout nor the
Claude CLI required by rein-aharness's agent-session approach. It never
deletes retained CoulombCore resources automatically.
```rule
@ -250,13 +255,13 @@ data:
id: "b8e4f1a2-3c6d-4e9f-a1b2-7d8e9f0a1b2c"
name: "Daily TODO.md Stale Review"
type: activity-definition
version: "1.1"
version: "1.2"
enabled: false
owner: custodian
governance: custodian
status: paused
created: "2026-07-08"
updated: "2026-07-20"
updated: "2026-08-20"
trigger:
type: cron
cron_expression: "30 8 * * *"
@ -277,27 +282,26 @@ data:
# Daily TODO.md Stale Review
> **Paused 2026-07-20.** The matched rule below always produces a
> `TaskSpec` (`rules/actions.py::_task_spec_for_rule`), which
> `RunActivityWorkflow` unconditionally routes through `emit_tasks` →
> the deployment-wide `IssueSink` (`ISSUE_SINK_TYPE`, default `rest` →
> issue-core → Forgejo). There is no per-activity sink override in the
> current schema, and this activity is the fleet's *only* consumer of
> `task_template` — it was the sole source of 5 stale-review Forgejo
> issues (`the-custodian` #1#5, opened 2026-07-08, all reviewed and
> closed 2026-07-20). Per current policy, Forgejo issue tracking is not
> the fleet coordination mechanism (binky-control housekeeping session,
> 2026-07-20) — so this activity is disabled (`enabled: false`) rather
> than left to recreate the same sidetrack daily.
> **Paused 2026-07-20; routing note corrected 2026-08-20.** This
> definition was paused after its matched rules created five unexpected
> Forgejo issues (`the-custodian` #1#5). That was the behavior at the
> time, when the deployment-wide `IssueSink` defaulted to `rest` →
> issue-core → Forgejo.
>
> **Re-enable path:** redirect the rule's output onto the work-record
> backbone (`work-record-types_v0.1.md`, `kind: intake` — "a stale TODO.md
> is exactly a finding/directive") instead of `IssueSink`. This needs
> either a new report-sink type in activity-core (file-based intake-item
> authoring in the target repo, mirroring `state-hub-progress`) or routing
> through the stage-3 promotion tooling seeded in `CUST-WP-0060`'s closure
> review (item 6, added 2026-07-20). Do not re-enable with `enabled: true`
> alone; that only restores the Forgejo-issue behavior.
> ACTIVITY-WP-0022 subsequently changed the fleet and code default to
> `ISSUE_SINK_TYPE=state-hub`. A matched rule now emits an
> `activity_task_spawn` progress event (and, when the ops-run queue is
> enabled, a claimable `ops_run`); it does **not** create a Forgejo issue
> unless an operator explicitly opts the deployment into
> `ISSUE_SINK_TYPE=rest`.
>
> **Current state:** `enabled: false` is retained so this documentation
> correction does not silently restore a production cadence. Before
> re-enabling, the owner should confirm that one task per stale repo per
> daily run is the intended fan-out and that the ops-run consumer will
> claim it. If that is the intended behavior, changing `enabled: true` is
> sufficient; no new sink type is required. State Hub progress remains
> visibility evidence, not claim authority.
>
> This file change is the ADR-001 source of truth; the live activity-core
> DB row picks it up on the next `make sync-activity-definitions` run
@ -331,12 +335,13 @@ data:
id: "f3a8c2e1-9b4d-4a6f-8e2d-1c5b7a9e3f04"
name: "Phase 5 Stabilization Daily Check"
type: activity-definition
version: "1.0"
enabled: true
version: "1.1"
enabled: false
owner: custodian
governance: custodian
status: active
status: paused
created: "2026-07-06"
updated: "2026-08-20"
trigger:
type: cron
cron_expression: "0 9 * * *"
@ -367,6 +372,9 @@ data:
---
# ActivityDefinition: Phase 5 Stabilization Daily Check
Disabled after review on 2026-08-20: its fixed stabilization window ended
on 2026-07-09 and the referenced State Hub task/workstream no longer exists.
weekly-legacy-meter-review.md: |
---
id: weekly-legacy-meter-review
@ -442,12 +450,13 @@ data:
id: "e7d2b5a8-4c1f-4e9a-b6d3-8f2a1c4e6b09"
name: "Phase 5 Stabilization Closeout Check"
type: activity-definition
version: "1.0"
enabled: true
version: "1.1"
enabled: false
owner: custodian
governance: custodian
status: active
status: paused
created: "2026-07-06"
updated: "2026-08-20"
trigger:
type: scheduled
at: "2026-07-09T17:35:00+00:00"
@ -477,17 +486,21 @@ data:
---
# ActivityDefinition: Phase 5 Stabilization Closeout Check
The one-shot fire date passed on 2026-07-09. Keep this disabled so schedule
sync does not retry a completed Temporal schedule.
core-hub-stabilization-daily.md: |
---
id: "b4c8e2f1-6a3d-4e5b-9f0c-2d7e8a1b3c4d"
name: "Core Hub Stabilization Daily Check"
type: activity-definition
version: "1.0"
enabled: true
version: "1.1"
enabled: false
owner: core-hub
governance: core-hub
status: active
status: paused
created: "2026-07-07"
updated: "2026-08-20"
trigger:
type: cron
cron_expression: "0 9 * * *"
@ -514,17 +527,21 @@ data:
---
# ActivityDefinition: Core Hub Stabilization Daily Check
Disabled after review on 2026-08-20: the fixed evidence window ended on
2026-07-10 and CORE-WP-0007 is finished and archived.
core-hub-stabilization-closeout.md: |
---
id: "c5d9f3a2-7b4e-5f6c-0a1d-3e8f9b2c4d5e"
name: "Core Hub Stabilization Closeout Check"
type: activity-definition
version: "1.0"
enabled: true
version: "1.1"
enabled: false
owner: core-hub
governance: core-hub
status: active
status: paused
created: "2026-07-07"
updated: "2026-08-20"
trigger:
type: scheduled
at: "2026-07-10T17:35:00+00:00"
@ -550,6 +567,9 @@ data:
---
# ActivityDefinition: Core Hub Stabilization Closeout Check
The one-shot fire date passed on 2026-07-10. Keep this disabled so schedule
sync does not retry a completed Temporal schedule.
ops-service-inventory-probes.md: |
---
id: "40d15a87-7ff6-4d8e-992c-37df15f95110"

View file

@ -46,6 +46,60 @@ GROUP BY d.name, d.enabled
ORDER BY runs DESC, d.name;
\"
echo
echo '--- execution outcomes by activity ---'
kubectl -n ${NS} exec actcore-app-db-0 -- psql -U actcore -d actcore -c \"
SELECT d.name, o.state, count(*) AS runs,
min(o.created_at) AS first_created,
max(o.updated_at) AS last_updated
FROM ops_runs o
JOIN activity_definitions d ON d.id = o.activity_definition_id
WHERE o.created_at >= timestamptz '${SINCE_UTC}'
GROUP BY d.name, o.state
ORDER BY d.name, o.state;
\"
echo
echo '--- failed or unclaimed execution detail (bounded) ---'
kubectl -n ${NS} exec actcore-app-db-0 -- psql -U actcore -d actcore -c \"
SELECT o.created_at, d.name, o.state, o.attempt, o.claim_owner,
left(o.title, 72) AS title,
left(coalesce(o.result->>'error', o.result->>'reason', ''), 240) AS failure
FROM ops_runs o
JOIN activity_definitions d ON d.id = o.activity_definition_id
WHERE o.created_at >= timestamptz '${SINCE_UTC}'
AND o.state IN ('failed', 'open', 'claimed')
ORDER BY o.created_at DESC
LIMIT 30;
\"
echo
echo '--- daily triage validation evidence (bounded) ---'
kubectl -n ${NS} exec deploy/actcore-api -- /app/.venv/bin/python3 -c '
import json, urllib.parse, urllib.request
from datetime import datetime
since = datetime.fromisoformat(\"${SINCE_UTC}\")
url = \"http://actcore-statehub-edge-relay:8000/progress/?\" + urllib.parse.urlencode(
{\"event_type\": \"daily_triage\", \"limit\": 100}
)
items = json.load(urllib.request.urlopen(url, timeout=10))
rows = []
for item in items:
created_raw = item.get(\"created_at\")
if not created_raw:
continue
created = datetime.fromisoformat(created_raw.replace(\"Z\", \"+00:00\"))
if created < since:
continue
detail = item.get(\"detail\") or {}
rows.append({
\"created_at\": created_raw,
\"run_id\": detail.get(\"activity_core_run_id\"),
\"output_validated\": detail.get(\"output_validated\"),
\"summary\": str(item.get(\"summary\") or \"\")[:200],
\"validation_error\": str(detail.get(\"validation_error\") or \"\")[:240],
})
print(json.dumps(rows[:14], indent=2))
'
echo
echo '--- non-high-frequency fires ---'
kubectl -n ${NS} exec actcore-app-db-0 -- psql -U actcore -d actcore -c \"
SELECT d.name, r.fired_at, r.tasks_spawned, r.run_id

View file

@ -352,7 +352,20 @@ async def evaluate_instructions(payload: dict) -> dict:
reports: list[dict] = []
for raw_instruction in instructions:
try:
instruction = InstructionDef.model_validate(raw_instruction)
instruction_data = dict(raw_instruction)
output_schema = instruction_data.get("output_schema")
if isinstance(output_schema, str) and output_schema.startswith(
("custodian://", "activity-core://")
):
# Resolve deployment-neutral URIs at the workflow/activity
# boundary. The pure rules package only accepts filesystem
# paths and must not import runtime integration modules.
from activity_core.runtime_paths import resolve_runtime_path
instruction_data["output_schema"] = str(
resolve_runtime_path(output_schema)
)
instruction = InstructionDef.model_validate(instruction_data)
except Exception as exc:
activity.logger.warning("instruction definition invalid — %s", exc)
continue

View file

@ -182,6 +182,13 @@ class ActivityDefinition(BaseModel):
instructions: list[InstructionDef] = Field(default_factory=list)
# Legacy — ignored when rules is non-empty
task_templates: list[TaskTemplate] = Field(default_factory=list)
dedupe_key_strategy: Literal["skip", "catchup", "compress"] = Field(default="skip")
dedupe_key_strategy: Literal["skip", "catchup", "compress"] = Field(
default="skip",
description=(
"Legacy persistence compatibility metadata; runtime scheduling uses "
"trigger_config.misfire_policy and this field does not deduplicate content"
),
deprecated=True,
)
version: int = Field(default=1, ge=1)
status: str = Field(default="active")

View file

@ -877,9 +877,7 @@ def _load_output_schema(schema_path: str) -> dict[str, Any] | None:
if not schema_path:
return None
from activity_core.runtime_paths import resolve_runtime_path
path = resolve_runtime_path(schema_path)
path = Path(schema_path).expanduser()
if not path.exists():
return None

View file

@ -72,6 +72,9 @@ async def sync(session_factory: async_sessionmaker[AsyncSession]) -> int:
task_templates=[],
rules_json=d.rules,
instructions_json=d.instructions,
# Legacy DB compatibility only. Schedule behavior comes
# from trigger_config.misfire_policy; this value does not
# suppress repeated content across nominal fire times.
dedupe_key_strategy="skip",
version=1,
)

View file

@ -24,7 +24,12 @@ from temporalio.client import Client
from activity_core.models import ActivityDefinition, CronTriggerConfig, ScheduledTriggerConfig
from activity_core.orm import ActivityDefinition as ActivityDefinitionRow
from activity_core.schedule_manager import delete_schedule, list_schedules, upsert_schedule
from activity_core.schedule_manager import (
cancel_scheduled,
delete_schedule,
list_schedules,
upsert_schedule,
)
logger = logging.getLogger(__name__)
@ -107,7 +112,14 @@ async def sync_schedule_rows(
valid_schedule_activity_ids.add(_valid_schedule_activity_id(defn))
try:
await upsert_schedule(client, defn)
if isinstance(defn.trigger_config, ScheduledTriggerConfig) and not defn.enabled:
# A disabled one-shot has no future recurrence to preserve. A
# completed Temporal schedule can reject in-place updates with
# ScheduleAlreadyRunningError, leaving every later sync noisy.
# Delete it idempotently instead; re-enabling recreates it.
await cancel_scheduled(client, defn.id)
else:
await upsert_schedule(client, defn)
except Exception as exc: # noqa: BLE001 — continue reconcile for other rows
result.errors += 1
detail = f"{defn.id} ({defn.name}): {type(exc).__name__}: {exc}"
@ -120,7 +132,10 @@ async def sync_schedule_rows(
logger.info("upserted schedule for activity %s (%s)", defn.id, defn.name)
else:
result.paused += 1
logger.info("upserted paused schedule for disabled activity %s", defn.id)
if isinstance(defn.trigger_config, ScheduledTriggerConfig):
logger.info("removed schedule for disabled one-shot activity %s", defn.id)
else:
logger.info("upserted paused schedule for disabled activity %s", defn.id)
# Tombstone cleanup: remove Temporal Schedules with no matching DB row.
try:

View file

@ -89,7 +89,7 @@ async def test_evaluate_instructions_returns_report_payload(monkeypatch) -> None
"trusted_fields": [],
"model": "test-model",
"prompt": "Run report.",
"output_schema": "schemas/daily-triage-report.json",
"output_schema": "activity-core://schemas/daily-triage-report.json",
"review_required": False,
}
],

View file

@ -201,6 +201,15 @@ async def test_status_endpoint_wraps_report(
}
monkeypatch.setattr(ops_api, "ops_status", fake_status)
session = AsyncMock()
session.__aenter__ = AsyncMock(return_value=session)
session.__aexit__ = AsyncMock(return_value=None)
count_result = MagicMock()
count_result.all.return_value = []
stuck_result = MagicMock()
stuck_result.scalar_one.return_value = 0
session.execute = AsyncMock(side_effect=[count_result, stuck_result])
ops_app.state.session_factory.return_value = session
transport = ASGITransport(app=ops_app)
async with AsyncClient(transport=transport, base_url="http://test") as client:
res = await client.get("/ops/automations/status", params={"since": "sunday"})

View file

@ -79,6 +79,25 @@ def test_external_configmap_projects_disabled_ops_probe_definition(tmp_path) ->
]
def test_daily_todo_pause_note_matches_safe_sink_default(tmp_path) -> None:
runtime_config = _by_kind_name("ConfigMap", "actcore-runtime-config")
definition_config = _by_kind_name(
"ConfigMap", "actcore-external-activity-definitions"
)
raw_definition = definition_config["data"]["daily-todo-md-stale-review.md"]
definition_path = tmp_path / "daily-todo-md-stale-review.md"
definition_path.write_text(raw_definition, encoding="utf-8")
definition = parse_file(definition_path)
assert runtime_config["data"]["ISSUE_SINK_TYPE"] == "state-hub"
assert definition.enabled is False
assert "ACTIVITY-WP-0022" in raw_definition
assert "`ISSUE_SINK_TYPE=state-hub`" in raw_definition
assert "default `rest`" not in raw_definition
assert "does **not** create a Forgejo issue" in raw_definition
def test_external_configmap_projects_weekly_legacy_meter_review(tmp_path) -> None:
config = _by_kind_name("ConfigMap", "actcore-external-activity-definitions")
raw_definition = config["data"]["weekly-legacy-meter-review.md"]
@ -146,7 +165,7 @@ def test_statehub_edge_relay_deployment_uses_state_hub_image_and_outbox_pvc() ->
container = deployment["spec"]["template"]["spec"]["containers"][0]
env = {item["name"]: item["value"] for item in container["env"]}
assert container["image"] == "forgejo.coulomb.social/coulomb/state-hub:main-d8808bf"
assert container["image"] == "forgejo.coulomb.social/coulomb/state-hub:main-1cf949b"
assert env["STATEHUB_UPSTREAM_URL"] == "http://state-hub.state-hub.svc.cluster.local:8000"
assert env["STATEHUB_OUTBOX_PATH"] == "/var/statehub/edge-outbox.sqlite3"
assert env["STATEHUB_READ_CACHE_PATH"] == "/var/statehub/edge-read-cache.sqlite3"

View file

@ -76,6 +76,8 @@ def test_post_missed_fire_alert_dual_writes_workplan_scope(monkeypatch) -> None:
posts: list[dict] = []
class _Resp:
status_code = 201
def raise_for_status(self) -> None: ...
def json(self) -> dict[str, str]:
return {"id": "progress-1"}

View file

@ -38,6 +38,7 @@ async def test_sync_schedule_rows_reports_drift_counts_and_preserves_one_shots(
orphan_id = uuid.uuid4()
upserted: list[tuple[uuid.UUID, bool, str]] = []
deleted: list[str] = []
cancelled_one_shots: list[uuid.UUID] = []
async def fake_upsert_schedule(client: object, defn: object) -> None:
upserted.append((
@ -65,9 +66,13 @@ async def test_sync_schedule_rows_reports_drift_counts_and_preserves_one_shots(
async def fake_delete_schedule(client: object, activity_id: str) -> None:
deleted.append(activity_id)
async def fake_cancel_scheduled(client: object, activity_id: uuid.UUID) -> None:
cancelled_one_shots.append(activity_id)
monkeypatch.setattr(sync_schedules, "upsert_schedule", fake_upsert_schedule)
monkeypatch.setattr(sync_schedules, "list_schedules", fake_list_schedules)
monkeypatch.setattr(sync_schedules, "delete_schedule", fake_delete_schedule)
monkeypatch.setattr(sync_schedules, "cancel_scheduled", fake_cancel_scheduled)
result = await sync_schedules.sync_schedule_rows(
object(),
@ -125,6 +130,49 @@ async def test_sync_schedule_rows_reports_drift_counts_and_preserves_one_shots(
(one_shot_id, True, "scheduled"),
]
assert deleted == [str(orphan_id)]
assert cancelled_one_shots == []
@pytest.mark.asyncio
async def test_sync_schedule_rows_removes_disabled_one_shot(monkeypatch) -> None:
one_shot_id = uuid.uuid4()
cancelled: list[uuid.UUID] = []
async def fake_upsert_schedule(client: object, defn: object) -> None:
raise AssertionError("disabled one-shot must not be upserted")
async def fake_cancel_scheduled(client: object, activity_id: uuid.UUID) -> None:
cancelled.append(activity_id)
async def fake_list_schedules(client: object) -> list[dict[str, str]]:
return []
monkeypatch.setattr(sync_schedules, "upsert_schedule", fake_upsert_schedule)
monkeypatch.setattr(sync_schedules, "cancel_scheduled", fake_cancel_scheduled)
monkeypatch.setattr(sync_schedules, "list_schedules", fake_list_schedules)
result = await sync_schedules.sync_schedule_rows(
object(),
[
_row(
activity_id=one_shot_id,
enabled=False,
trigger_config={
"trigger_type": "scheduled",
"at": datetime(2026, 8, 17, 6, 0, tzinfo=timezone.utc),
"timezone": "UTC",
},
)
],
)
assert cancelled == [one_shot_id]
assert result.to_dict() == {
"upserted": 0,
"paused": 1,
"deleted_orphans": 0,
"errors": 0,
}
@pytest.mark.asyncio

View file

@ -52,6 +52,7 @@ async def test_run_sync_runs_requested_sections(monkeypatch) -> None:
"upserted": 3,
"paused": 1,
"deleted_orphans": 2,
"errors": 0,
}
assert result["errors"] == []

View file

@ -8,7 +8,7 @@ status: active
owner: grok
topic_slug: infotech
created: "2026-08-18"
updated: "2026-08-18"
updated: "2026-08-20"
parent_workplan: CUST-WP-0062
related:
- CUST-WP-0062
@ -92,16 +92,17 @@ oldest never-scanned repo.
```task
id: ACTIVITY-WP-0030-T03
status: wait
status: progress
priority: high
state_hub_task_id: "18bcfc21-616a-48a2-aac3-cdcef7579913"
```
After the daily job is proven on railiance01: set
`weekly-sbom-staleness` `enabled: false` (or remove `flag-stale-sbom`), keep
or drop the weekly summary report only if it no longer emits one task per
repo. Update runbook / playbook. Prove one Monday window with zero SBOM task
spawns and a successful daily catch-up the same week.
Emergency containment began 2026-08-20 under ACTIVITY-WP-0031: set
`weekly-sbom-staleness` `enabled: false` in source and production before the
bounded replacement is available. Keep or drop the weekly summary report only
if it no longer emits one task per repo. After the daily job is proven on
railiance01, update the runbook / playbook and prove one Monday window with zero
SBOM task spawns and a successful daily catch-up the same week.
### Prod enable and evidence

View file

@ -0,0 +1,120 @@
---
id: ACTIVITY-WP-0031
type: workplan
title: "Production automation execution reliability cleanup"
domain: infotech
repo: activity-core
status: active
owner: codex
topic_slug: activity-core
created: "2026-08-20"
updated: "2026-08-20"
related:
- ACTIVITY-WP-0021
- ACTIVITY-WP-0022
- ACTIVITY-WP-0026
- ACTIVITY-WP-0030
---
# Production automation execution reliability cleanup
## Goal
Make production automation health reflect completed execution rather than only
successful schedule fires, stop known task floods, and repair the repository
contracts found during the 2026-08-20 activity-core review.
## Restore provider-backed execution
```task
id: ACTIVITY-WP-0031-T01
status: wait
priority: high
```
Production FI, Binky, and daily-triage runs reach `llm-connect` but fail after
three attempts. An in-cluster fixture returned a sanitized upstream OpenRouter
HTTP 401. Credential ownership is the `railiance-platform` OpenBao lane
(`openrouter-llm-connect`); activity-core must not bypass that custody boundary.
Track rotation/reconciliation, restart, and a successful fixture smoke as the
exit evidence.
## Emergency-pause weekly SBOM fan-out
```task
id: ACTIVITY-WP-0031-T02
status: done
priority: high
```
Disable `weekly-sbom-staleness` in source and on railiance01 before the next
Monday window. The bounded daily replacement remains owned by ACTIVITY-WP-0030;
this task is the immediate containment measure for the observed 75-task fire.
Completed 2026-08-20: source, live DB, and Temporal schedule are disabled. The
new image and ConfigMap projection were deployed; schedule reconciliation
reports the definition as paused.
## Report execution outcomes in production status
```task
id: ACTIVITY-WP-0031-T03
status: done
priority: high
```
Extend `scripts/prod_automation_status.sh` with ops-run outcome counts, bounded
failed/open detail, and daily-triage progress validation evidence. Preserve
sanitized error output and keep the command deterministic and read-only.
Completed 2026-08-20. The live report now shows ops-run state counts, bounded
failure/open detail, and bounded State Hub `daily_triage` validation evidence.
It exposed today's FI/Binky 502 failures and triage `output_validated=false`
beside otherwise healthy schedule-fire counts.
## Reconcile implementation and test contracts
```task
id: ACTIVITY-WP-0031-T04
status: done
priority: medium
```
Repair the four reproducible suite failures, keep runtime-path resolution out of
the rules package boundary, and document that `dedupe_key_strategy` is retained
metadata rather than repeated-content suppression.
Completed 2026-08-20. Runtime URI resolution now happens before the pure rules
boundary; stale assertions/test doubles are reconciled; disabled one-shots are
deleted idempotently during schedule sync; and the dedupe contract is truthful.
Verification: `347 passed, 1 skipped`; a subsequent focused warning regression
passed `10 passed` without warnings.
## Repair consumer routing and State Hub metadata
```task
id: ACTIVITY-WP-0031-T05
status: wait
priority: medium
```
Make the OpenBao retention reactivation task explicitly claimable without
authorizing destructive cleanup. Route Binky's missing workload credential
front door and the duplicate legacy workstream identifiers to their canonical
owners. Clear stale human-needed flags only for already-terminal activity-core
tasks.
Activity-core cleanup is complete: three terminal tasks no longer retain stale
`needs_human` flags. External handoffs were sent to binky-control and
railiance-platform for mail credential materialization, repo-manager for the
duplicate registrar IDs, and railiance-master for the unclaimed retention
closeout. The expired retention one-shot is disabled; its existing open ops run
remains operator-visible and cannot authorize destructive cleanup.
## Acceptance
- [x] Weekly SBOM fan-out is disabled in source and production
- [x] Production status distinguishes schedule fires from execution failures
- [x] Full test suite is green
- [ ] Provider and consumer prerequisites have named owners and evidence
- [x] Workplan and State Hub task state are consistent

View file

@ -0,0 +1,43 @@
---
id: ADHOC-2026-08-20
type: workplan
title: "Ad hoc — activity-core opportunistic fixes 2026-08-20"
domain: infotech
repo: activity-core
status: finished
owner: codex
topic_slug: activity-core
created: "2026-08-20"
updated: "2026-08-20"
---
# ADHOC-2026-08-20 — activity-core opportunistic fixes
## Reconcile the daily TODO review sink warning
```task
id: ADHOC-2026-08-20-T01
status: done
priority: low
```
The projected `daily-todo-md-stale-review` definition still said that
`IssueSink` defaulted to `rest` and that re-enabling the definition would
unconditionally recreate Forgejo issues. That warning described the state before
ACTIVITY-WP-0022 and contradicted the current code, deployment configuration,
and `weekly-sbom-staleness` definition.
Correct the warning to state that `ISSUE_SINK_TYPE=state-hub` is the safe fleet
default, `rest` is explicit opt-in, and State Hub progress is visibility rather
than claim authority. Keep the schedule disabled so a documentation correction
does not silently change production cadence, and document the operational checks
needed before it is deliberately re-enabled. Add a projection test so the stale
`default rest` claim cannot recur unnoticed.
Verification:
- Focused regression and IssueSink tests: `7 passed`.
- Full suite: `342 passed, 1 skipped, 4 failed`. The four failures are unrelated
existing inconsistencies: rules boundary allowlisting, the State Hub image pin
assertion, a schedule-health response double missing `status_code`, and the
sync-service expectation omitting the returned `errors` count.