the-custodian/docs/agent-autonomy-decision.md
codex db91818e84
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 3s
Python Tests / pytest (push) Successful in 25s
Advance supervised agent records and close verified Secret annotation guard
2026-09-28 18:15:27 +02:00

9.2 KiB

Agent-specific supervised and autopilot modes

Founder decision, 2026-09-28, GOVERN @ estate. Recorded under CUST-WP-0073-T01. This supersedes the proposed permanent choice between observation-only agents and unrestricted deployment access.

Accepted direction

Autonomy is a characteristic of an identified agent, scoped to the work it is trusted to perform. Start agents in supervised-mode. Record how often the supervisor accepts privileged-action proposals unchanged and how well those exact proposals work when executed. Only a demonstrated record of successful operation without adaptation or refinement supports promotion to autopilot-mode. Autopilot has both a cost budget and a risk limit in EUR, which the supervisor can tune per agent and scope.

This decision accepts the model. It does not promote an agent, choose numerical limits, authorize a live credential cutover or claim runtime enforcement exists.

Minimal operating contract

An agent instance/assignment carries its stable identity, mode, supervisor, allowed action/target scope, execution-profile revision and reference to its promotion/limit decision. Mode persists across sessions. Trust in one scope does not automatically grant trust in another. New scopes start supervised; material model, tool-profile or execution-environment changes require a supervisor review of whether prior evidence still applies.

In supervised-mode, ordinary already-authorized preparation, reads, local edits and tests continue. For a privileged action the agent prepares the exact action, target, expected result, verification/rollback, cost estimate and risk estimate. The supervisor approves that proposal or performs it through the privileged execution path. The receipt records whether approval or human execution was needed. Approval is bound to the reviewed proposal revision; changing that revision requires review again. The agent cannot approve itself.

In autopilot-mode, a privileged action may run without per-action approval only inside the agent's granted scope and both monetary limits. Missing estimates, expired grants, insufficient remaining budget, unbounded risk, or actions outside scope return that action to supervision. Existing human-only lanes remain human-only unless explicitly changed by the governing authority. Mode and work record lane are separate dimensions; effective authority is their intersection.

Evidence for promotion

Use existing approval and execution receipts, keyed by agent, scope, profile revision and a stable proposal ID/digest. Record:

  • the original proposal, supervisor disposition (unchanged / revised / rejected), revisions and reason; pending/withdrawn proposals stay visible;
  • approval and execution identities, timestamps and the exact executed revision;
  • outcome verification, interventions, rollback/recovery, observed cost and realized loss/incident evidence. Store references, not credentials.

For a declared review window, report counts as well as rates:

  • Unchanged acceptance rate: proposals accepted without changes divided by all adjudicated original proposals. Revised and rejected proposals remain in the denominator; retries do not become fresh successes.
  • Unchanged execution success rate: original proposals executed as accepted, passing the agreed verification with no corrective refinement or rescue, divided by all executed original proposals with completed verification. Execution of a supervisor-revised proposal does not earn an unchanged success.
  • Also report pending/unverified outcomes and supervisor interventions/time. Approval without execution is not a successful outcome. No observations means unknown, never 100%.

The supervisor promotes explicitly for a named scope using an agreed minimum sample, review window, acceptance/success thresholds and incident tolerance. These values are not set by this decision; no default percentage grants access. Successful harmless work alone does not establish competence for higher-risk privileged actions. Promotion, revocation and limit changes retain history. The supervisor can reduce autonomy immediately; failed verification or missing enforcement stops further autonomous privileged actions pending review.

Two distinct EUR controls

Cost budget bounds attributable spend and commitments over an explicit period, including execution costs and resources/services the action commits to. Reserve the estimated maximum cost before execution, reconcile actuals afterwards, and count concurrent reservations against the same remaining budget. Unknown cost is not free. Record the budget source, currency and reset period. Do not confuse a token limit or provider subscription with a complete euro-denominated budget.

Risk limit bounds potential loss, separately from normal spending. Proposed operational interpretation for each grant: a conservatively assessed credible loss bound per action plus aggregate outstanding exposure from concurrent or dependent actions. Include recovery expense, service interruption and data loss where applicable, with assumptions and uncertainty. An expected-loss average alone must not hide a much larger credible downside. An unpriced or unbounded consequence requires supervision. EUR limits do not price away human-only or other non-monetary prohibitions. The supervisor accepts the valuation method and the action/exposure limits when granting autopilot; the agent cannot raise its own limit or declare its own estimate authoritative.

Before enabling autopilot, verify that the execution path enforces reservations, limits and revocation across concurrent actions. Recording fields in a manifest does not enforce a budget. Until that proof exists, the mode remains supervised.

Credential boundary and existing owners

Keep admin credentials out of the agent's direct reach in both modes. A scoped privileged execution path performs approved actions in supervised-mode and policy-authorized actions in autopilot-mode, returning sanitized outcome evidence. Unrestricted sudo, admin kubeconfig access or unreviewed GitOps deployment would bypass that path. Separate identities/isolation remain necessary, but the observation-only profile is the supervised starting profile, not a permanent limit on what an agent may accomplish.

Reuse existing boundaries rather than build another supervisor service:

Concern Existing surface / limit
Agent identity, mode and performance Consumer-owned agent instance/assignment records; agentic-resources performance loop. Its broader workforce inventory/assignment contracts are still proposed.
Exact human approval and execution outcome approval-engine approval object, informed-decision supervisor presentation, execution receipts. State Hub decisions record governance; they are not runtime approval tokens.
Runtime enforcement and credential custody glas-harness, selected rein and sandbox; existing authorization and credential-owner paths. A prompt or mode label grants nothing.
Monetary authority fin-hub budget source, with execution accounting supplied by the relevant runtime/service.
Risk judgement Supervisor-approved estimates; risk-nexus can hold evidence but is not a live EUR risk gate.
Work and review evidence Existing CUST-WP-0073 tasks and State Hub progress; no new task/workplan or service.

Bounded application to CUST-WP-0073

T01's policy choice is resolved by this decision. T02 implements and proves the supervised starting boundary and records the existing execution path selected; T04 adds per-agent mode and proposal/outcome evidence to guidance and the chosen existing receipts. First implementation may use reviewed file records and existing receipts. No new dashboard, automatic promotion engine, monetary risk estimator or general workforce system is required to finish credential separation.

Autopilot is a promotion option requiring its own concrete grant and demonstrated enforcement, not an activation promised by this workplan. No such grant exists from this decision. T03's annotation policy and T05's deferred rotation remain unchanged. CUST-WP-0071 retains its measurement and weekly-review scope.

Supervised record in use

The initial consumer-owned record is .kaizen/agents/custodian-codex/supervision.json. Generate its descriptive report:

python3 scripts/summarize_agent_supervision.py .kaizen/agents/custodian-codex/supervision.json

For each future scored proposal, retain the original digest before submission, the supervisor's approval reference and approved digest, then the executed digest and verification receipt. Use one stable proposal ID across revisions. Record refinement_or_rescue explicitly. Accepted revised proposals, rejections and rescued executions cannot earn unchanged success. Pending/unverified outcomes remain visible. These records contain references and outcomes, never credential values or executable approval tokens.

Earlier conversation authorization is retained as unscored where an exact submitted revision was not recorded. It is not backfilled into promotion evidence. The initial rates are unknown. The utility is descriptive and grants no authority; autopilot remains disabled and the interactive runtime has not been migrated. The existing sand-boxer isolation mechanism has a separate non-model proof in docs/evidence/2026-09-28-supervised-sandbox-proof.json.