the-custodian/docs/agent-autonomy-decision.md

158 lines
9.2 KiB
Markdown
Raw Normal View History

# Agent-specific supervised and autopilot modes
Founder decision, 2026-09-28, `GOVERN @ estate`.
Recorded under CUST-WP-0073-T01. This supersedes the proposed permanent choice
between observation-only agents and unrestricted deployment access.
## Accepted direction
Autonomy is a characteristic of an identified agent, scoped to the work it is
trusted to perform. Start agents in **supervised-mode**. Record how often the
supervisor accepts privileged-action proposals unchanged and how well those
exact proposals work when executed. Only a demonstrated record of successful
operation without adaptation or refinement supports promotion to
**autopilot-mode**. Autopilot has both a cost budget and a risk limit in EUR,
which the supervisor can tune per agent and scope.
This decision accepts the model. It does not promote an agent, choose numerical
limits, authorize a live credential cutover or claim runtime enforcement exists.
## Minimal operating contract
An agent instance/assignment carries its stable identity, mode, supervisor,
allowed action/target scope, execution-profile revision and reference to its
promotion/limit decision. Mode persists across sessions. Trust in one scope does
not automatically grant trust in another. New scopes start supervised; material
model, tool-profile or execution-environment changes require a supervisor review
of whether prior evidence still applies.
In supervised-mode, ordinary already-authorized preparation, reads, local edits
and tests continue. For a privileged action the agent prepares the exact action,
target, expected result, verification/rollback, cost estimate and risk estimate.
The supervisor approves that proposal or performs it through the privileged
execution path. The receipt records whether approval or human execution was
needed. Approval is bound to the reviewed proposal revision; changing that
revision requires review again. The agent cannot approve itself.
In autopilot-mode, a privileged action may run without per-action approval only
inside the agent's granted scope and both monetary limits. Missing estimates,
expired grants, insufficient remaining budget, unbounded risk, or actions outside
scope return that action to supervision. Existing human-only lanes remain
human-only unless explicitly changed by the governing authority. Mode and work
record `lane` are separate dimensions; effective authority is their intersection.
## Evidence for promotion
Use existing approval and execution receipts, keyed by agent, scope, profile
revision and a stable proposal ID/digest. Record:
- the original proposal, supervisor disposition (unchanged / revised /
rejected), revisions and reason; pending/withdrawn proposals stay visible;
- approval and execution identities, timestamps and the exact executed revision;
- outcome verification, interventions, rollback/recovery, observed cost and
realized loss/incident evidence. Store references, not credentials.
For a declared review window, report counts as well as rates:
- **Unchanged acceptance rate:** proposals accepted without changes divided by
all adjudicated original proposals. Revised and rejected proposals remain in
the denominator; retries do not become fresh successes.
- **Unchanged execution success rate:** original proposals executed as accepted,
passing the agreed verification with no corrective refinement or rescue,
divided by all executed original proposals with completed verification.
Execution of a supervisor-revised proposal does not earn an unchanged success.
- Also report pending/unverified outcomes and supervisor interventions/time.
Approval without execution is not a successful outcome. No observations means
unknown, never 100%.
The supervisor promotes explicitly for a named scope using an agreed minimum
sample, review window, acceptance/success thresholds and incident tolerance.
These values are not set by this decision; no default percentage grants access.
Successful harmless work alone does not establish competence for higher-risk
privileged actions. Promotion, revocation and limit changes retain history.
The supervisor can reduce autonomy immediately; failed verification or missing
enforcement stops further autonomous privileged actions pending review.
## Two distinct EUR controls
**Cost budget** bounds attributable spend and commitments over an explicit period,
including execution costs and resources/services the action commits to. Reserve
the estimated maximum cost before execution, reconcile actuals afterwards, and
count concurrent reservations against the same remaining budget. Unknown cost is
not free. Record the budget source, currency and reset period. Do not confuse a
token limit or provider subscription with a complete euro-denominated budget.
**Risk limit** bounds potential loss, separately from normal spending. Proposed
operational interpretation for each grant: a conservatively assessed credible
loss bound per action plus aggregate outstanding exposure from concurrent or
dependent actions. Include recovery expense, service interruption and data loss
where applicable, with assumptions and uncertainty. An expected-loss average
alone must not hide a much larger credible downside. An unpriced or unbounded
consequence requires supervision. EUR limits do not price away human-only or
other non-monetary prohibitions. The supervisor accepts the valuation method
and the action/exposure limits when granting autopilot; the agent cannot raise
its own limit or declare its own estimate authoritative.
Before enabling autopilot, verify that the execution path enforces reservations,
limits and revocation across concurrent actions. Recording fields in a manifest
does not enforce a budget. Until that proof exists, the mode remains supervised.
## Credential boundary and existing owners
Keep admin credentials out of the agent's direct reach in both modes. A scoped
privileged execution path performs approved actions in supervised-mode and
policy-authorized actions in autopilot-mode, returning sanitized outcome evidence.
Unrestricted sudo, admin kubeconfig access or unreviewed GitOps deployment would
bypass that path. Separate identities/isolation remain necessary, but the
observation-only profile is the supervised starting profile, not a permanent
limit on what an agent may accomplish.
Reuse existing boundaries rather than build another supervisor service:
| Concern | Existing surface / limit |
|---|---|
| Agent identity, mode and performance | Consumer-owned agent instance/assignment records; `agentic-resources` performance loop. Its broader workforce inventory/assignment contracts are still proposed. |
| Exact human approval and execution outcome | `approval-engine` approval object, `informed-decision` supervisor presentation, execution receipts. State Hub decisions record governance; they are not runtime approval tokens. |
| Runtime enforcement and credential custody | `glas-harness`, selected rein and sandbox; existing authorization and credential-owner paths. A prompt or mode label grants nothing. |
| Monetary authority | `fin-hub` budget source, with execution accounting supplied by the relevant runtime/service. |
| Risk judgement | Supervisor-approved estimates; `risk-nexus` can hold evidence but is not a live EUR risk gate. |
| Work and review evidence | Existing CUST-WP-0073 tasks and State Hub progress; no new task/workplan or service. |
## Bounded application to CUST-WP-0073
T01's policy choice is resolved by this decision. T02 implements and proves the
supervised starting boundary and records the existing execution path selected;
T04 adds per-agent mode and proposal/outcome evidence to guidance and the chosen
existing receipts. First implementation may use reviewed file records and
existing receipts. No new dashboard, automatic promotion engine, monetary risk
estimator or general workforce system is required to finish credential separation.
Autopilot is a promotion option requiring its own concrete grant and demonstrated
enforcement, not an activation promised by this workplan. No such grant exists
from this decision. T03's annotation policy and T05's deferred rotation remain
unchanged. CUST-WP-0071 retains its measurement and weekly-review scope.
## Supervised record in use
The initial consumer-owned record is
`.kaizen/agents/custodian-codex/supervision.json`. Generate its descriptive report:
```bash
python3 scripts/summarize_agent_supervision.py .kaizen/agents/custodian-codex/supervision.json
```
For each future scored proposal, retain the original digest before submission,
the supervisor's approval reference and approved digest, then the executed digest
and verification receipt. Use one stable proposal ID across revisions. Record
`refinement_or_rescue` explicitly. Accepted revised proposals, rejections and
rescued executions cannot earn unchanged success. Pending/unverified outcomes
remain visible. These records contain references and outcomes, never credential
values or executable approval tokens.
Earlier conversation authorization is retained as `unscored` where an exact
submitted revision was not recorded. It is not backfilled into promotion evidence.
The initial rates are unknown. The utility is descriptive and grants no authority;
autopilot remains disabled and the interactive runtime has not been migrated.
The existing sand-boxer isolation mechanism has a separate non-model proof in
`docs/evidence/2026-09-28-supervised-sandbox-proof.json`.