Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
19 KiB
| id | type | title | domain | repo | status | owner | topic_slug | created | updated | depends_on | unblocks | state_hub_workstream_id | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| TEN-WP-0006 | workplan | Guardrail and quota policy for tenants | infotech | tenant-engine | finished | claude | tenant-guardrails | 2026-08-16 | 2026-08-16 |
|
bddb6699-cb53-472a-9757-a3ed0eb0ce27 |
TEN-WP-0006 - Guardrail and quota policy
Implement the guardrail/quota concern that SCOPE.md, INTENT.md, and
.claude/rules/architecture.md have reserved since the repo was created but
never built. tenant-engine owns policy — what a tenant is permitted to
spend and how many entities/actions it may hold — not metering, not billing,
and not the enforcement decision itself.
The concrete hole this closes: ADR-0013 specifies that trial-grouped tenants
default to a zero spend budget. No such default exists today, so every
trial tenant currently carries no ceiling at all.
Boundary constraints
These are not negotiable without a new ADR, and every task below is bound by them:
tenant-engineis a data source, never a PDP. It answers "what is this tenant's limit" and "what is its recorded consumption".flex-authdecides whether an action is allowed. Never invert this.- No metering. Consumption counters are written to this service by
whoever meters;
tenant-enginedoes not observe spend or count actions itself. - No plan terms. A guardrail may be derived from an
adaptive-pricingplan id, but plan definitions are never copied here. - Fail closed. An unknown tenant, an unreadable store, or a missing guardrail resolves to the most restrictive answer, never to "unlimited".
T01 - Specify the guardrail model and its boundary
id: TEN-WP-0006-T01
status: done
priority: high
state_hub_task_id: "b0693ec3-b3a7-48fb-9f4f-baeedb488e9d"
Write the guardrail domain contract before any code. Decide and document:
- The limit kinds in scope. Start with spend budget (a currency amount over a period) and entity/action counts; leave rate limiting out unless it falls out for free.
- Precedence. A tenant may inherit a limit from its grouping (ADR-0013), from its assigned plan, and from an explicit per-tenant override. Define the resolution order and make it total — every tenant resolves to exactly one effective guardrail set, with no ambiguity and no "unset means unlimited".
- The
trial= zero-spend default, and what the default is for each other grouping. An unmapped grouping must resolve restrictively, not openly. - Reserved identifiers.
tenant:platformandtenant:coulombare ungrouped (seedomain.py); state explicitly what they resolve to. - Whether consumption lives in this repo at all. If it does, it is an opaque counter written by an external meter, with the writer named in the contract.
Document the consumer-facing shape in docs/tenant-guardrail-policy.md,
matching the style of docs/tenant-lifecycle-api.md.
Done when the precedence rules are unambiguous, the fail-closed default is explicit for every grouping, and the doc states plainly what this service does not do (meter, bill, decide).
Done 2026-08-16: docs/tenant-guardrail-policy.md. Decisions worth carrying:
- Limit kinds:
spend(integer minor units + ISO-4217, never floats),entity_count,action_count. Rate limiting stays out — it is a gateway concern with different latency and storage needs, not a free by-product. - Registry, not open keys. A limit key must be registered to resolve; an
unregistered key errors (
unknown_limit_key) rather than resolving to zero or unlimited. Zero would break a caller who merely misspelled a key; unlimited would fail open. Startup validation requires a default for every grouping, so an unmapped grouping cannot reach production. - Precedence, per key: override → plan → grouping → fail-closed floor. Per
key, not per set, so a plan supplying
spend.monthlydoes not wipe out a grouping-derivedentity.*. The floor is only reachable if the registry is internally inconsistent, and reports provenancefail_closedso that state is visible rather than mistaken for policy. unlimitedis an explicit sentinel, never a default and never the result of absence. The rule is "no unset means unlimited" — that bans inferring an open ceiling, not deliberately declaring one.- Reserved identifiers get a
reservedprofile at layer 3 rather than falling through: spend 0 (infrastructure identities are not billable spenders), counts explicitly unlimited. Falling through would clamp the platform's own identity to zero and take the platform down with it. - Lifecycle clamp applies after resolution and may only reduce;
retiredclamps every limit to the floor while leaving reads working. - Consumption does not live in this repo. tenant-engine is a low-write
policy authority with audit and CAS semantics; consumption is high-frequency
telemetry with opposite needs, and holding it here would drag metering into a
service
SCOPE.mddisowns it from. Reads return limits plus provenance only.
Two things flagged rather than settled: the non-trial grouping ceilings are
conservative opening values needing product sign-off (only trial = 0 is
canon), and no repo owns metering — the same gap SCOPE.md records for
payment processing. Until canon names one, flex-auth can enforce
ceiling-presence semantics but not consumption-relative ones. A zero budget is
fully enforceable with no meter at all, which is why the ADR-0013 trial
default lands immediately.
T02 - Implement the guardrail domain model
id: TEN-WP-0006-T02
status: done
priority: high
state_hub_task_id: "1b28d836-11d9-4ef3-967b-05bfa74c304a"
Add a guardrail/ module (the namespace architecture.md reserves) holding
pure domain types with no framework dependency, mirroring how domain.py is
structured: frozen dataclasses, validation that raises typed errors, and
resolution logic as a pure function over (grouping, plan_id, overrides).
Guardrail changes are audited like role grants — append-only, carrying
actor, reason, and correlation id. A limit change is a privilege change and
must be as reconstructible as a RoleGrant.
Done when effective-guardrail resolution is a total, side-effect-free function with unit tests covering every grouping, the reserved identifiers, missing plans, and conflicting overrides.
Done 2026-08-16: src/tenant_engine/guardrail/ — model.py (frozen
dataclasses, typed errors), registry.py (the key registry and grouping
defaults), resolution.py (the pure resolution function). 42 new tests, 166
total, all passing.
Design notes worth carrying:
- Registry totality is enforced at construction.
LimitDefinitionraisesGuardrailRegistryInvalidErrorif any grouping inGROUPINGSlacks a default, so an unmapped grouping cannot reach production. Thefail_closedbranch inresolve_limitis the second line of defence, not the first — its test has to bypass the constructor to reach it. - A floor may never be
unlimited, also enforced at construction. A floor that fails open is not a floor. - Currency is deliberately excluded from
compatible_with. Precedence resolves per key and the winning layer takes the value whole, so two currencies never combine — which is why there is no exchange rate anywhere in this service, and why there must not be one. - The lifecycle clamp only relabels when it actually reduces. A retired
trialtenant already at zero keeps provenancegrouping; a retiredlargetenant becomeslifecycle. Relabelling unconditionally would hide where the value came from. boolis rejected as an amount —Trueis anintin Python and would otherwise pass the non-negative check as a ceiling of 1.
Lint: the new files are clean. make lint still fails on 19 pre-existing
E501s in files this workplan does not touch (tests/test_store.py and
friends) — that baseline predates TEN-WP-0006 and is left alone rather than
folded into this diff.
T03 - Persist guardrails in both stores
id: TEN-WP-0006-T03
status: done
priority: high
state_hub_task_id: "f1c93573-6322-4505-b738-7d66d67e60a8"
Extend the TenantStore Protocol and both implementations (InMemoryTenantStore
and the SQLite store) with guardrail persistence and its audit trail. Follow the
precedent set by mutate_tenant() in TEN-WP-0005: the mutation, the audit
event, and any idempotency receipt commit in one transaction — a crash must
never leave a changed limit with no audit record.
Add a forward-only, idempotent migration that backfills existing tenants with their grouping-derived defaults. Existing tenants must not silently gain a looser limit than a fresh tenant of the same grouping would get.
Interaction with lifecycle: decide and test what a retired tenant's
guardrails resolve to. Follow the TEN-WP-0005 precedent — operations that only
reduce privilege stay available while retired; loosening does not.
Done when the store-conformance suite (already parametrised over both backends) covers guardrails, so the durable store cannot diverge from the reference semantics.
Done 2026-08-16: TenantStore gains guardrail_overrides, guardrail_changes,
and set_guardrail_override, implemented in both backends.
tests/test_guardrail_store_conformance.py parametrises 35 tests over
in-memory and SQLite. 201 tests total, all passing.
Deviations and decisions, both deliberate:
- No backfill migration, because grouping defaults are resolved rather than
materialised. The task anticipated backfilling existing tenants with their
grouping-derived defaults. Storing them would have been worse: it duplicates
a value that already has a single source, and it makes changing a default a
data migration instead of a config edit. Since a tenant with no override
resolves through the same function whether it predates guardrails or not,
"existing tenants must not silently gain a looser limit than a fresh tenant
of the same grouping" holds by construction, not by a migration step that
could be skipped. The forward-only, idempotent part is the two new tables,
created by the same
CREATE TABLE IF NOT EXISTSscript fresh and existing databases both take — no second schema definition to drift. - A guardrail change bumps the tenant version. It reuses the lifecycle
ETag, so a guardrail write invalidates a stale reader's
If-Matchexactly as a metadata edit does, rather than introducing a second, separately versioned resource that a caller could race against the first.
Retired tenants, following the TEN-WP-0005 reduce-privilege precedent:
- Guardrails stay readable; it is the values that clamp, not the endpoint.
- A retired tenant's effective limits clamp to the floor, provenance
lifecycle. - Tightening is allowed while retired, loosening is refused. The check compares the before/after value computed as if the tenant were active — because the clamp makes every candidate look inert today, and the real effect of a loosening override lands the moment the tenant is reactivated. That deferred loosening is the thing worth refusing. Clearing an override is evaluated the same way, so clearing a tightening override on a retired tenant is refused too, since fall-through would loosen.
Stored amounts are TEXT NOT NULL with an explicit "unlimited" token, never
NULL. A NULL that meant unlimited would make an open ceiling the result of
absence — the one thing the contract forbids.
T04 - Expose guardrail read and write APIs
id: TEN-WP-0006-T04
status: done
priority: high
state_hub_task_id: "4a256517-2773-4679-ad57-2909f22ac8a4"
Add the API surface, authorized through the existing flex-auth WriteAuthorizer
with distinct actions so policy can permit reading a limit without
permitting a change to it — the same split TEN-WP-0005 used for
tenant.update / tenant.retire / tenant.reactivate:
- a read returning a tenant's effective guardrails plus their provenance
(which layer each limit came from), for
flex-authto consume; - a write to set or clear a per-tenant override, requiring
Idempotency-Key,If-Match, actor, reason, and correlation id, consistent with the lifecycle mutation contract.
Authorization runs before the store is touched, so an unauthorized caller
cannot probe which tenants exist. Errors are redacted: never reflect policy
internals or store paths in detail.
New flex-auth actions mean a policy-package change outside this repo. Name the required actions explicitly in the task record so the flex-auth side is a tracked handoff and not a surprise — this is exactly what stalled TEN-WP-0005-T05.
Done when the routes are authorized, versioned, idempotent, correlated, and provider-neutral, and the OpenAPI document makes the semantics unambiguous.
Done 2026-08-16. Three routes: GET /tenants/{id}/guardrails (effective
limits plus provenance), PUT and DELETE /tenants/{id}/guardrails/{limit_key}. 26 API tests; 227 total, all passing.
Contract documented in docs/tenant-guardrail-policy.md.
Required flex-auth actions — the tracked handoff this task exists to make explicit:
| Action | Resource type | Purpose |
|---|---|---|
tenant.guardrail.read |
guardrail |
read a tenant's effective ceilings |
tenant.guardrail.set |
guardrail |
set or clear a per-tenant override |
Both are new and need a policy-package change in flex-auth before this
surface functions in production. Until then every check correctly resolves to
deny. See the regression note below — these land on a package that is
currently behind, not merely one that needs extending.
Decisions:
- The read is authorized too, unlike the existing
/rolesand/roles/livereads. A ceiling is policy about a tenant's commercial exposure, not a capability claim, and the task's whole point was that policy must be able to permit reading a limit without permitting a change.actoris a required query parameter, so there is no anonymous caller. Tested both ways: read-only permission cannot write, write-only permission cannot read. amountis a string on the wire. It lets theunlimitedsentinel and an integer share one field without a JSON union, and keeps a spend amount in minor units from ever round-tripping through a float.change_idis derived from(tenant_id, limit_key, Idempotency-Key)rather than random, so a genuine retry replays the original audit record instead of minting a second one for a mutation that happened once.- Authorization precedes every store touch, so a denied read of a real tenant and of a nonexistent one are byte-identical — no probing existence or registry contents through status codes.
- A retired tenant's refused loosening gets its own code,
guardrail_loosening_denied, rather than reusinginvalid_lifecycle_transition: the tenant's lifecycle is not in transition, and a caller should be able to tell "your tenant is retired, tighten instead" from "your retire/reactivate call was invalid".
T05 - Conformance and consumer handoff
id: TEN-WP-0006-T05
status: done
priority: medium
state_hub_task_id: "92036fa3-9031-4b42-a67e-93196e236e08"
Cover: every grouping's default, trial resolving to zero spend, precedence
between grouping/plan/override, unmapped grouping failing closed, reserved
identifiers, guardrail reads on a retired tenant, stale-version conflict,
idempotent replay, cross-tenant authorization denial, store outage, and error
redaction. Prove the existing create/role/plan/lifecycle clients are unaffected.
Then write the consumer handoff for flex-auth — the guardrail read is only
useful once a PDP consults it — naming the API version, the authorization
actions, and the contract doc, in the table format TEN-WP-0005-T05 used.
Do not deploy to production as part of this workplan. Cluster rollout is a separate, human-authorized step; see the note below.
Done when the full suite passes and the handoff is sent.
Done 2026-08-16. 231 tests pass (was 166 before this workplan; 65 new). Every item on the coverage list is exercised:
| Case | Where |
|---|---|
| every grouping's default | test_guardrail_domain (parametrised over GROUPINGS) |
trial resolves to zero spend |
test_guardrail_domain, test_api_guardrails |
| precedence grouping/plan/override | test_guardrail_domain |
| precedence is per key, not per set | test_guardrail_domain |
| unmapped grouping fails closed | test_guardrail_domain (bypasses the constructor to reach it) |
| reserved identifiers | test_guardrail_domain (parametrised over RESERVED_IDENTIFIERS) |
| guardrail read on a retired tenant | test_guardrail_store_conformance, test_api_guardrails |
| stale-version conflict | both conformance suites |
| idempotent replay | both conformance suites |
| cross-tenant authorization denial | test_api_guardrails |
| store outage, read and write path | test_api_guardrails |
| error redaction | test_api_guardrails |
| existing clients unaffected | test_api_guardrails, plus all 166 prior tests unchanged |
Consumer handoff to flex-auth:
| Fact | Value |
|---|---|
| Contract | docs/tenant-guardrail-policy.md |
| API version | 0.1.0 (unchanged — additive routes only) |
| Read route | GET /tenants/{id}/guardrails?actor=<actor> |
| Write routes | PUT / DELETE /tenants/{id}/guardrails/{limit_key} |
| Required actions | tenant.guardrail.read, tenant.guardrail.set (resource type guardrail) |
| Policy package | tenant-engine.write-api.mutate — extends the existing seven actions to nine |
| Source revision | d1d9c9a |
| Immutable image | none — not built. Production rollout is out of scope; see below |
Two things the consumer must know, both stated plainly in the message sent:
flex-authcannot use this yet. Both actions are new, so until the policy package carries them every check denies — correctly, but the surface is inert in production until that lands.- The read alone is not enforcement. No repo owns metering, so a PDP can
enforce ceiling-presence semantics (is there a limit? is it zero?) but not
consumption-relative ones (has the tenant used it up?). The ADR-0013
trial= zero default is fully enforceable today precisely because it needs no meter.
Out of scope / explicitly deferred
- Production rollout. Image build and cluster apply need credentials and a human authorization decision. Track separately.
- The live flex-auth regression (
flex-auth-tenant-enginerolled back to the four-action image, sotenant.update/tenant.retire/tenant.reactivatecurrently deny withunknown_action). Unrelated to this workplan, but it means any new actions from T04 land on a policy package that is already behind. Resolve the regression before shipping T04's actions. - Metering and billing. Not this repo, at any point.