# Work Orchestration Best Practices — External Research > Research artefact — 2026-07-19. Phase 2 of the work-orchestration > reconciliation effort. Surveys current (2025/2026) practice for organizing > human, agent, and classic-automation workflows, filtered for relevance to > a file-first, repo-anchored, solo-founder-plus-agent-fleet factory. > Companion: `2026-07-19-work-orchestration-infrastructure-survey.md`. ## 1. GitOps / everything-as-code The industry consensus for declarative operations maps directly onto ADR-001 and generalizes it: - **Four GitOps principles**: declarative desired state; versioned & immutable storage; state pulled automatically by agents; continuous reconciliation between desired and actual state ([GitLab](https://about.gitlab.com/topics/gitops/), [Spacelift](https://spacelift.io/blog/gitops-best-practices)). Our `fix-consistency` is exactly a reconciliation loop; the gap is that it runs on demand rather than continuously, and only covers workplans/tasks — not queues, suggestions, or decisions. - **Validate before apply**: schema validation, policy compliance, dry-run in CI before a change reaches the runtime system ([Pulumi](https://www.pulumi.com/blog/gitops-best-practices-i-wish-i-had-known-before/)). Applied to work items: queue/workplan YAML should be schema-validated in Forgejo CI, not only at hub-ingest time. - **Separate code and state repos** where lifecycles differ ([GitHub blog](https://github.blog/enterprise-software/devops/applying-gitops-principles-to-your-operations/)) — consistent with ADR-005 project repos and control-plane repos like binky-control. - **Ops-as-issues / IssueOps**: operations triggered and audited through issue lifecycles. Validation of our issue-core direction; the key lesson is that the issue is the *interface*, not the source of truth. ## 2. Human-in-the-loop tiering for agentic work 2026 agentic-governance literature converges on patterns our AutonomyPolicy already anticipates — with refinements worth adopting: - **Risk-tiered approval, not blanket oversight**: pre-approval (human-in-the-loop) for irreversible actions (payments, contracts, deletion, production changes); async monitoring (human-on-the-loop) for high-volume recoverable work ([Strata](https://www.strata.io/blog/agentic-identity/practicing-the-human-in-the-loop/), [StackAI](https://www.stackai.com/insights/human-in-the-loop-ai-agents-how-to-design-approval-workflows-for-safe-and-scalable-automation)). Maps to Green/Blue (on-the-loop) vs. Yellow/Orange/Red (in-the-loop). - **Approval gates before side effects — propose-then-commit**: block the side-effecting call until approval is recorded ([Velt](https://velt.dev/blog/why-ai-agents-need-approval-layer)). Our approval-package pattern satisfies this at the process level; the harness tool-profile mechanism is where it becomes technically enforced. - **Identity-aware orchestration**: pause execution, route to authorized humans, time-boxed decision windows, log every intervention ([Strata](https://www.strata.io/blog/agentic-identity/practicing-the-human-in-the-loop/)). `needed_by` / `fallback_if_no_response` in DecisionQueue are exactly time-boxed windows; the hub should own the clock, not the file reader. - **Every agent has an owner, risk profile, and documented decision boundary**; outputs traceable to model versions and configs ([CMR "Governing the Agentic Enterprise"](https://cmr.berkeley.edu/2026/03/governing-the-agentic-enterprise-a-new-operating-model-for-autonomous-ai-at-scale/), [McKinsey](https://www.mckinsey.com/capabilities/risk-and-resilience/our-insights/trust-in-the-age-of-agents)). Matches `agt-…` hub identities + manifests; argues for making `lane` and `owner` mandatory fields on every work item, not just on agents. - **Durable pause/resume**: graph-based orchestrators persist state so a workflow can wait days for human input and resume losslessly ([Virtido](https://virtido.com/blog/agentic-workflows-patterns-best-practices-enterprise)). Our equivalent is the file+hub pair: a Yellow item *is* a paused workflow; no in-memory orchestrator state should be required to resume it. - **Reviewer quality matters**: "a poorly trained reviewer approving flawed agent outputs is worse than no checkpoint" ([nhimg.org](https://nhimg.org/faq/how-should-security-teams-implement-human-in-the-loop-controls-for-ai-agents/)) — supports the approval-*package* discipline (evidence, recommendation, fallback) over bare questions, and the attention-budget cap. ## 3. Flow management: WIP limits, metrics, prioritization Kanban/lean portfolio practice supplies the quantitative layer the fleet lacks: - **Four flow metrics**: WIP, cycle time, work-item age, throughput ([ProKanban](https://www.prokanban.org/blog/https-prokanban-org-blog-the-kanban-pocket-guide-chapter-6-the-basic-metrics-of-flow)). All four are computable from hub data (timestamped UUIDs + status transitions) — none is surfaced today. Work-item age is the direct, data-driven form of the "stalled" tag (aspect 6). - **WIP limits at every state**, including portfolio level; WSJF (cost-of-delay / job size) chooses the next item, WIP limits decide whether the system may accept it ([Agility at Scale — WIP](https://agility-at-scale.com/safe/lpm/work-in-progress/), [WSJF](https://agility-at-scale.com/safe/lpm/wsjf-weighted-shortest-job-first/)). The founder attention budget (max 7 decisions/day) is a WIP limit on the human workstation; the same construct should apply to agent lanes (e.g. max active workplans per repo). A WSJF calibration experiment already exists (`docs/daily-statehub-wsjf-calibration-2026-06-04.md`). - **Epics as economic bets** with testable benefit hypotheses, decomposed only for the near horizon; close the measurement loop on predicted outcomes ([Agility at Scale — Epics](https://agility-at-scale.com/safe/lpm/epics/)). Maps to domain/repo goals → workplans; argues for a lightweight `hypothesis`/`expected_outcome` field on goals, checked at closure review. - **Lean budget guardrails**: fund value streams within policy boundaries instead of approving each piece of work; segment budgets per area of risk ([Atomic Object](https://spin.atomicobject.com/budget-guardrails-software-project/), [Agility at Scale — Guardrails](https://agility-at-scale.com/safe/lpm/portfolio-guardrails/)). Direct template for token/cost budgets: per-lane and per-workplan budget envelopes with escalation on breach, instead of per-request permission. ## 4. Dependency and risk practice - **Dependency mapping to see cascade risk**; identify risks early, attach mitigation to the item ([Asana](https://asana.com/resources/project-risks), [Atlassian](https://www.atlassian.com/work-management/project-management/project-risk-management)). The hub has dependency edges (task- and workplan-level); the gap is surfacing them (blocked-cascade views) and the absence of edges for queue-item species. - **Risk classification determines autonomy level, data access, and approval requirements per agent** ([TM Forum](https://inform.tmforum.org/features-and-opinion/ai-autonomy-governance-a-governance-framework-for-agentic-ai-enabling-safe-accountable-and-scalable-autonomous-intelligence), [Trend Micro](https://www.trendmicro.com/vinfo/us/security/news/cybercrime-and-digital-threats/from-anarchy-to-authority-closing-the-governance-gap-in-agentic-ai)). Confirms lanes-as-risk-classes; suggests lane should also gate credential catalog access (ops-warden already keyed per-lane). ## 5. Patterns worth *not* importing - **Heavy portfolio ceremony** (SAFe boards, quarterly PI planning): the attention budget is 90 min/day; anything requiring synchronous ceremonies fails the constraint. Take the *metrics and limits*, skip the meetings. - **Ticket-system-as-source-of-truth** (Jira-style): contradicts ADR-001; issues/tickets remain interface projections. - **Generic workflow engines as the primary model** (BPMN, Airflow-style DAGs for knowledge work): our work is document-anchored and interruption-driven; the task-flow-engine's assertion model (state = provable properties, not engine position) is the better fit and is already specced. - **Per-item human approval as default**: the literature is unanimous that blanket approval queues collapse under volume; risk-tiering is the scalable form — the AutonomyPolicy default ("ask Bernd is never the default") is the correct opinion. ## 6. Condensed implications for the draft 1. Treat *work intake* and *work execution* as separate concerns with different friction budgets — intake must cost one YAML block or less (the queues proved this), execution structure can cost a workplan file. 2. Every work-item species must live in the same reconciliation loop (file → CI validation → hub index → derived views) or it becomes a sidetrack by construction. 3. Human attention is a WIP-limited workstation; model it exactly like an agent lane with a budget, a queue, and flow metrics. 4. States should be few and abstract (canon lifecycle); everything situational (stalled, unclear, automatable, needs-human) should be a **derived or policy-assigned tag**, not a state — states migrate badly, tags migrate trivially. 5. Timestamped UUIDs internally + stable canonical names externally is the correct layering and is industry-consistent (UUIDv7 for time-sortable internal keys). 6. Budgets become guardrail envelopes (per lane / workplan / run) enforced at the runtime layer (harness), metered to the hub, escalated on breach. ## Sources - https://about.gitlab.com/topics/gitops/ - https://www.pulumi.com/blog/gitops-best-practices-i-wish-i-had-known-before/ - https://spacelift.io/blog/gitops-best-practices - https://github.blog/enterprise-software/devops/applying-gitops-principles-to-your-operations/ - https://virtido.com/blog/agentic-workflows-patterns-best-practices-enterprise - https://www.strata.io/blog/agentic-identity/practicing-the-human-in-the-loop/ - https://velt.dev/blog/why-ai-agents-need-approval-layer - https://www.stackai.com/insights/human-in-the-loop-ai-agents-how-to-design-approval-workflows-for-safe-and-scalable-automation - https://nhimg.org/faq/how-should-security-teams-implement-human-in-the-loop-controls-for-ai-agents/ - https://cmr.berkeley.edu/2026/03/governing-the-agentic-enterprise-a-new-operating-model-for-autonomous-ai-at-scale/ - https://www.mckinsey.com/capabilities/risk-and-resilience/our-insights/trust-in-the-age-of-agents - https://inform.tmforum.org/features-and-opinion/ai-autonomy-governance-a-governance-framework-for-agentic-ai-enabling-safe-accountable-and-scalable-autonomous-intelligence - https://www.trendmicro.com/vinfo/us/security/news/cybercrime-and-digital-threats/from-anarchy-to-authority-closing-the-governance-gap-in-agentic-ai - https://agility-at-scale.com/safe/lpm/work-in-progress/ - https://agility-at-scale.com/safe/lpm/wsjf-weighted-shortest-job-first/ - https://agility-at-scale.com/safe/lpm/epics/ - https://agility-at-scale.com/safe/lpm/portfolio-guardrails/ - https://spin.atomicobject.com/budget-guardrails-software-project/ - https://www.prokanban.org/blog/https-prokanban-org-blog-the-kanban-pocket-guide-chapter-6-the-basic-metrics-of-flow - https://asana.com/resources/project-risks - https://www.atlassian.com/work-management/project-management/project-risk-management