Advance supervised agent records and close verified Secret annotation guard
This commit is contained in:
parent
b2f6713721
commit
db91818e84
44 changed files with 6868 additions and 54 deletions
157
docs/agent-autonomy-decision.md
Normal file
157
docs/agent-autonomy-decision.md
Normal file
|
|
@ -0,0 +1,157 @@
|
|||
# Agent-specific supervised and autopilot modes
|
||||
|
||||
Founder decision, 2026-09-28, `GOVERN @ estate`.
|
||||
Recorded under CUST-WP-0073-T01. This supersedes the proposed permanent choice
|
||||
between observation-only agents and unrestricted deployment access.
|
||||
|
||||
## Accepted direction
|
||||
|
||||
Autonomy is a characteristic of an identified agent, scoped to the work it is
|
||||
trusted to perform. Start agents in **supervised-mode**. Record how often the
|
||||
supervisor accepts privileged-action proposals unchanged and how well those
|
||||
exact proposals work when executed. Only a demonstrated record of successful
|
||||
operation without adaptation or refinement supports promotion to
|
||||
**autopilot-mode**. Autopilot has both a cost budget and a risk limit in EUR,
|
||||
which the supervisor can tune per agent and scope.
|
||||
|
||||
This decision accepts the model. It does not promote an agent, choose numerical
|
||||
limits, authorize a live credential cutover or claim runtime enforcement exists.
|
||||
|
||||
## Minimal operating contract
|
||||
|
||||
An agent instance/assignment carries its stable identity, mode, supervisor,
|
||||
allowed action/target scope, execution-profile revision and reference to its
|
||||
promotion/limit decision. Mode persists across sessions. Trust in one scope does
|
||||
not automatically grant trust in another. New scopes start supervised; material
|
||||
model, tool-profile or execution-environment changes require a supervisor review
|
||||
of whether prior evidence still applies.
|
||||
|
||||
In supervised-mode, ordinary already-authorized preparation, reads, local edits
|
||||
and tests continue. For a privileged action the agent prepares the exact action,
|
||||
target, expected result, verification/rollback, cost estimate and risk estimate.
|
||||
The supervisor approves that proposal or performs it through the privileged
|
||||
execution path. The receipt records whether approval or human execution was
|
||||
needed. Approval is bound to the reviewed proposal revision; changing that
|
||||
revision requires review again. The agent cannot approve itself.
|
||||
|
||||
In autopilot-mode, a privileged action may run without per-action approval only
|
||||
inside the agent's granted scope and both monetary limits. Missing estimates,
|
||||
expired grants, insufficient remaining budget, unbounded risk, or actions outside
|
||||
scope return that action to supervision. Existing human-only lanes remain
|
||||
human-only unless explicitly changed by the governing authority. Mode and work
|
||||
record `lane` are separate dimensions; effective authority is their intersection.
|
||||
|
||||
## Evidence for promotion
|
||||
|
||||
Use existing approval and execution receipts, keyed by agent, scope, profile
|
||||
revision and a stable proposal ID/digest. Record:
|
||||
|
||||
- the original proposal, supervisor disposition (unchanged / revised /
|
||||
rejected), revisions and reason; pending/withdrawn proposals stay visible;
|
||||
- approval and execution identities, timestamps and the exact executed revision;
|
||||
- outcome verification, interventions, rollback/recovery, observed cost and
|
||||
realized loss/incident evidence. Store references, not credentials.
|
||||
|
||||
For a declared review window, report counts as well as rates:
|
||||
|
||||
- **Unchanged acceptance rate:** proposals accepted without changes divided by
|
||||
all adjudicated original proposals. Revised and rejected proposals remain in
|
||||
the denominator; retries do not become fresh successes.
|
||||
- **Unchanged execution success rate:** original proposals executed as accepted,
|
||||
passing the agreed verification with no corrective refinement or rescue,
|
||||
divided by all executed original proposals with completed verification.
|
||||
Execution of a supervisor-revised proposal does not earn an unchanged success.
|
||||
- Also report pending/unverified outcomes and supervisor interventions/time.
|
||||
Approval without execution is not a successful outcome. No observations means
|
||||
unknown, never 100%.
|
||||
|
||||
The supervisor promotes explicitly for a named scope using an agreed minimum
|
||||
sample, review window, acceptance/success thresholds and incident tolerance.
|
||||
These values are not set by this decision; no default percentage grants access.
|
||||
Successful harmless work alone does not establish competence for higher-risk
|
||||
privileged actions. Promotion, revocation and limit changes retain history.
|
||||
The supervisor can reduce autonomy immediately; failed verification or missing
|
||||
enforcement stops further autonomous privileged actions pending review.
|
||||
|
||||
## Two distinct EUR controls
|
||||
|
||||
**Cost budget** bounds attributable spend and commitments over an explicit period,
|
||||
including execution costs and resources/services the action commits to. Reserve
|
||||
the estimated maximum cost before execution, reconcile actuals afterwards, and
|
||||
count concurrent reservations against the same remaining budget. Unknown cost is
|
||||
not free. Record the budget source, currency and reset period. Do not confuse a
|
||||
token limit or provider subscription with a complete euro-denominated budget.
|
||||
|
||||
**Risk limit** bounds potential loss, separately from normal spending. Proposed
|
||||
operational interpretation for each grant: a conservatively assessed credible
|
||||
loss bound per action plus aggregate outstanding exposure from concurrent or
|
||||
dependent actions. Include recovery expense, service interruption and data loss
|
||||
where applicable, with assumptions and uncertainty. An expected-loss average
|
||||
alone must not hide a much larger credible downside. An unpriced or unbounded
|
||||
consequence requires supervision. EUR limits do not price away human-only or
|
||||
other non-monetary prohibitions. The supervisor accepts the valuation method
|
||||
and the action/exposure limits when granting autopilot; the agent cannot raise
|
||||
its own limit or declare its own estimate authoritative.
|
||||
|
||||
Before enabling autopilot, verify that the execution path enforces reservations,
|
||||
limits and revocation across concurrent actions. Recording fields in a manifest
|
||||
does not enforce a budget. Until that proof exists, the mode remains supervised.
|
||||
|
||||
## Credential boundary and existing owners
|
||||
|
||||
Keep admin credentials out of the agent's direct reach in both modes. A scoped
|
||||
privileged execution path performs approved actions in supervised-mode and
|
||||
policy-authorized actions in autopilot-mode, returning sanitized outcome evidence.
|
||||
Unrestricted sudo, admin kubeconfig access or unreviewed GitOps deployment would
|
||||
bypass that path. Separate identities/isolation remain necessary, but the
|
||||
observation-only profile is the supervised starting profile, not a permanent
|
||||
limit on what an agent may accomplish.
|
||||
|
||||
Reuse existing boundaries rather than build another supervisor service:
|
||||
|
||||
| Concern | Existing surface / limit |
|
||||
|---|---|
|
||||
| Agent identity, mode and performance | Consumer-owned agent instance/assignment records; `agentic-resources` performance loop. Its broader workforce inventory/assignment contracts are still proposed. |
|
||||
| Exact human approval and execution outcome | `approval-engine` approval object, `informed-decision` supervisor presentation, execution receipts. State Hub decisions record governance; they are not runtime approval tokens. |
|
||||
| Runtime enforcement and credential custody | `glas-harness`, selected rein and sandbox; existing authorization and credential-owner paths. A prompt or mode label grants nothing. |
|
||||
| Monetary authority | `fin-hub` budget source, with execution accounting supplied by the relevant runtime/service. |
|
||||
| Risk judgement | Supervisor-approved estimates; `risk-nexus` can hold evidence but is not a live EUR risk gate. |
|
||||
| Work and review evidence | Existing CUST-WP-0073 tasks and State Hub progress; no new task/workplan or service. |
|
||||
|
||||
## Bounded application to CUST-WP-0073
|
||||
|
||||
T01's policy choice is resolved by this decision. T02 implements and proves the
|
||||
supervised starting boundary and records the existing execution path selected;
|
||||
T04 adds per-agent mode and proposal/outcome evidence to guidance and the chosen
|
||||
existing receipts. First implementation may use reviewed file records and
|
||||
existing receipts. No new dashboard, automatic promotion engine, monetary risk
|
||||
estimator or general workforce system is required to finish credential separation.
|
||||
|
||||
Autopilot is a promotion option requiring its own concrete grant and demonstrated
|
||||
enforcement, not an activation promised by this workplan. No such grant exists
|
||||
from this decision. T03's annotation policy and T05's deferred rotation remain
|
||||
unchanged. CUST-WP-0071 retains its measurement and weekly-review scope.
|
||||
|
||||
## Supervised record in use
|
||||
|
||||
The initial consumer-owned record is
|
||||
`.kaizen/agents/custodian-codex/supervision.json`. Generate its descriptive report:
|
||||
|
||||
```bash
|
||||
python3 scripts/summarize_agent_supervision.py .kaizen/agents/custodian-codex/supervision.json
|
||||
```
|
||||
|
||||
For each future scored proposal, retain the original digest before submission,
|
||||
the supervisor's approval reference and approved digest, then the executed digest
|
||||
and verification receipt. Use one stable proposal ID across revisions. Record
|
||||
`refinement_or_rescue` explicitly. Accepted revised proposals, rejections and
|
||||
rescued executions cannot earn unchanged success. Pending/unverified outcomes
|
||||
remain visible. These records contain references and outcomes, never credential
|
||||
values or executable approval tokens.
|
||||
|
||||
Earlier conversation authorization is retained as `unscored` where an exact
|
||||
submitted revision was not recorded. It is not backfilled into promotion evidence.
|
||||
The initial rates are unknown. The utility is descriptive and grants no authority;
|
||||
autopilot remains disabled and the interactive runtime has not been migrated.
|
||||
The existing sand-boxer isolation mechanism has a separate non-model proof in
|
||||
`docs/evidence/2026-09-28-supervised-sandbox-proof.json`.
|
||||
Loading…
Add table
Add a link
Reference in a new issue