28 KiB
Architecture Blueprint — Governed Qonto Assistant
Status: blueprint v0.2 — 2026-07-21 (revised after repo review)
Repo: qonto-assistant (canonical home).
Context: BINKY-WP-0005 finished (live OpenBao lane + first read-only pull).
Question: how do multiple coding agents / harnesses interact with Qonto conveniently and under one policy, without wiring vendor MCP into each client and accepting policy drift?
Related:INTENT.md,research/2026-07-21-mcp-gateway-and-governed-domain-assistant.md, binky-controlintegrations/qonto-mcp.md, AutonomyPolicy, agent-harness ADR-001, ops-warden catalogbinky-qonto-api, CCR-2026-0008.
1. Problem
We already can:
- Fetch company Qonto credentials from OpenBao (
tenants/binky/qonto-api). - Call Qonto’s thirdparty REST API (proven: first pull, CostRunRate v2).
- Map those credentials into the self-hosted
qonto/qonto-mcp-serverenv shape (QONTO_API_KEY/QONTO_ORGANIZATION_ID).
What we must not do as the long-term pattern:
- Install vendor MCP (or raw API keys) into Claude Desktop, Claude Code, Codex, Cursor, Grok, agent-harness, executor-worker, etc. separately.
- Rely on each harness’s local tool allow-list as the only spend control.
- Hope every client remembers “read-only forever.”
Why that fails:
| Failure mode | Effect |
|---|---|
| Policy drift | One harness allows a write tool; another forbids it. Same company key, different blast radius. |
| Credential sprawl | Keys appear in N configs, Docker args, session env dumps. |
| Audit fragmentation | No single log of “which agent did what to the bank.” |
| Qonto cannot help | Business API keys are not scope-limited server-side (confirmed in design). There is no “read-only API key” or MCP-side spend ACL from Qonto. |
| Volume-cost tools | Cards, invoicing, payment requests, transfers, and some “create” operations may incur fees or move money — must be blocked before the vendor sees them. |
Founder policy for now (normative intent):
Agents may inspect balances, transactions, and statements for control-plane and finance awareness. Agents must not initiate spending, transfers, card operations, or any action with volume/transaction cost. Payments stay Red lane (human in Qonto app).
That is stronger than “read-only MCP tools” in the abstract: it is a business policy that must be enforced at a choke point, not a client preference.
2. What industry patterns say
Research (2025–2026 MCP ecosystem) converges on the same shape your proxy intuition already points at.
2.1 MCP gateway / proxy as the control plane
Production patterns treat MCP clients (agent harnesses) as untrusted for tool policy, and put a gateway between them and tools:
- Envoy AI Gateway MCPRoute — multiplex backends, toolSelector
allow/deny, OAuth on the gateway, upstream API-key injection so clients
never hold vendor secrets, CEL rules on
tools/callparams. - Kong AI Gateway MCP Tool ACLs — default-deny, per-consumer tool subsets at the gateway (solves “all-or-nothing” MCP exposure).
- Speakeasy / Portkey / Permit MCP Gateway framing — authn, RBAC, audit of every tool call, policy evaluated on each call, not only at connect time.
- Harness’s own MCP guidance — gateways for centralized credentials, governance, audit, single endpoint vs N direct MCP connections.
Common lesson: clients connect to one governed endpoint; the gateway owns credentials, tool inventory, and policy.
2.2 MCP security best practices (protocol-level)
From MCP security guidance and industry write-ups:
| Practice | Application to Qonto |
|---|---|
| No token passthrough | Clients must not present the Qonto bank key. The assistant service is the only OAuth/API client to Qonto. |
| Tool-level scopes, not server-level | “Access to Qonto MCP” is too coarse; allow list_transactions, deny create_transfer. |
| Default deny | Unknown tools fail closed. |
| JIT escalation for high privilege | Any future spend path is Yellow/Orange/Red with explicit re-authorization — never session-global “write mode.” |
| Confused-deputy / static client ID | If the service ever OAuth-proxies, consent and audience binding matter; for API-key custody we avoid OAuth to Qonto entirely. |
| Audit end-to-end | Actor (harness/agent id) + tool + args hash + decision + latency; no secrets in logs. |
2.3 Domain capability service (not “another harness”)
A second pattern, older than MCP: BFF / capability API for a domain (banking, email, CRM). Agents call a narrow, policy-shaped interface (“get runway snapshot”) rather than the full vendor surface. That is how email-connect already behaves for IMAP: consumer never holds IMAP secrets in the product path; it goes through custody + a purpose-built reader.
llm-connect is the wrong home for this: it is multi-provider LLM abstraction (adapters, routing, cost models). Putting bank tools there blurs the layer map (ARCHITECTURE-LAYERS.md) and couples finance policy to model routing.
agent-harness is the wrong home for Qonto policy itself, but the right home for session binding: harness ADR-001 already says the harness is the credential holder and policy enforcement point for agent sessions. That must not mean “each harness reimplements bank policy.” It means the harness only grants agents a named tool profile that talks to already-governed domain services. Domain policy lives below the harness.
2.4 Synthesis
| Pattern | Use for Qonto? |
|---|---|
| Wire vendor MCP into every harness | No — policy drift, key sprawl |
| Generic MCP gateway only (Envoy/Kong in front of vendor MCP) | Partial — good transport + tool filter; weak on semantic spend rules and finance-shaped tools |
| Governed domain assistant (our MCP + REST, policy inside) | Yes — recommended |
| Flex-auth decisions on every call | Yes — who may call; lane/risk attributes |
| OpenBao + ops-warden for bank key | Yes — already live; only the assistant fetches |
Your “qonto-proxy governed API & MCP service” is the right class of solution. Name it as a capability assistant, not only a dumb reverse proxy: it should own a safe tool surface, not only filter someone else’s.
3. Goals and non-goals
Goals
- One policy for “no spend / no volume-cost actions” across all agents.
- One credential path (existing OpenBao lane); harnesses never see bank keys.
- Convenient agent UX: natural “ask the finance assistant” via MCP and/or small REST helpers used by scripts and rhythm sessions.
- Auditability: every call attributable to agent/harness/session.
- Fits NetKingdom: flex-auth, OpenBao, ops-warden, agent-harness instances, activity-core schedules — no parallel IAM.
- Dogfood-ready: internal use is product evidence (DogfoodPolicy).
Non-goals (v1)
- Replacing DATEV / StB / DUO accounting.
- Automating payments, SEPA, card issuance, expense submission.
- Hosted Qonto OAuth connector (
mcp.qonto.com) as corporate custody path. - Making llm-connect a tool gateway.
- Per-harness custom Qonto integrations.
4. Recommended architecture
4.1 One-liner
Build a small “Qonto Governed Assistant” service that is the only component allowed to hold the bank API key and speak Qonto’s protocol. Expose a policy-shaped MCP server and a thin REST API. All harnesses are dumb clients of that service. Policy is code + tests in one place.
4.2 Context diagram
┌──────────────────────────────────────────────────────────────────────────┐
│ Agent clients │
│ Claude Code · Codex · Cursor · Grok · agent-harness sessions · scripts │
└───────────────┬───────────────────────────────┬──────────────────────────┘
│ MCP (streamable-HTTP) │ REST (JSON)
▼ ▼
┌──────────────────────────────────────────────────────────────────────────┐
│ Qonto Governed Assistant («qonto-assistant») │
│ │
│ Authn/Authz edge │
│ OIDC / workload id / mTLS + optional flex-auth │
│ │ │
│ ▼ │
│ Protocol adapters │
│ MCP adapter REST adapter │
│ │ │ │
│ └──────────┬────────┘ │
│ ▼ │
│ Capability service core │
│ canonical capability ids + response contracts │
│ │ │
│ ▼ │
│ Policy engine │
│ default-deny, tenant scope, arg constraints, data class rules │
│ │ │
│ ┌─────┴──────────┐ │
│ ▼ ▼ │
│ Qonto adapter Audit emitter │
│ REST client structured events │
│ │ │ │
└──────┼────────────────┼──────────────────────────────────────────────────┘
│ │
▼ ▼
OpenBao tenants/<tenant>/qonto-api stdout / log sink / metrics
│
▼
Qonto thirdparty API
Optional later: put Envoy AI Gateway (or similar) in front of
qonto-assistant for fleet-wide MCP aggregation with other domain assistants
(email, issue-core, …). Do not put the gateway directly in front of
vendor qonto-mcp-server as the sole control — semantic policy is thinner
there.
4.3 Why a domain assistant beats “proxy the vendor MCP only”
| Approach | Pros | Cons |
|---|---|---|
| A. Gateway + vendor MCP | Fast; toolSelector; industry-standard | Vendor tool names/args still drive the model; hard to encode “no volume cost”; vendor MCP is unmaintained / dual-use with write tools |
| B. Domain assistant (recommended) | Safe tool catalog we design; semantic policy; dual MCP+REST; dogfood product surface | Build cost |
| C. Per-harness allow-lists | Zero new service | Guaranteed drift; keys leak into clients |
Hybrid that works: implement B; if we still want vendor MCP for parity experiments, run it only as an internal backend behind the same policy gate (never published to harnesses). Prefer direct REST client inside the assistant (we already proved REST) — fewer moving parts than shelling out to unmaintained MCP.
4.4 Component responsibilities
| Component | Owns | Does not own |
|---|---|---|
| qonto-assistant | Tool catalog, spend/volume policy, Qonto REST client, dual surfaces (MCP+REST), call audit | LLM routing, scheduling, identity issuance |
| OpenBao + ops-warden | Custody path, rotation guide, high-risk agent boundary | Business rules about transfers |
| flex-auth | “May actor X call capability finance.qonto.read?” |
Bank credential material |
| agent-harness | Session lifecycle; grants profile finance-qonto-read; injects assistant URL/token only |
Bank policy implementation |
| llm-connect | Which model answers | Tool execution |
| activity-core | When Finance Steward runs | How bank is called |
| binky-control | CostRunRate, queues, AutonomyPolicy, dogfood evidence | Runtime of the assistant |
| Qonto | Bank of record | Our agent policy |
4.5 Protocol adapters over one service core
The critical implementation rule is:
REST routes and MCP tools are only transport adapters. They must translate into the same internal
CapabilityRequestmodel before policy runs.
That avoids the most likely drift bug in this repo: REST enforcing one set of rules while MCP grows a slightly different set later.
Canonical request shape:
CapabilityRequest
capability_id
tenant_id
actor_claims
resource_scope
request_args
protocol # rest | mcp
Recommended v1 capability map:
| Internal capability id | REST surface | MCP surface | Notes |
|---|---|---|---|
org_summary |
GET /v1/accounts |
qonto_org_summary |
Core balances/org read |
list_transactions |
GET /v1/transactions |
qonto_list_transactions |
Capped, filtered history |
cost_run_rate_hints |
used by GET /v1/snapshot |
qonto_cost_run_rate_hints (Phase 2) |
Normalized finance hints, not raw export |
snapshot_bundle |
GET /v1/snapshot |
optional later | Composite read; not a separate privilege if it only orchestrates allowed reads |
find_counterparties |
defer | defer | Candidate future read capability, not Phase 1 MVP |
Rules:
snapshot_bundleis an orchestrator over already-allowed reads, not a way to bypass policy with a "bigger" internal export.- REST and MCP names may differ, but both must map to the same
capability_id, deny reasons, and response redaction rules. - Vendor payloads should be normalized into repo-owned response schemas; do not stream raw Qonto JSON into agent contexts by default.
4.6 Policy model (v1)
Encode as code + declarative YAML, tested in CI — not wiki prose alone. Policy should evaluate this tuple:
(capability_id, actor_claims, tenant_id, resource_scope, request_args, response_class)
Sketch:
# policy/qonto-v1.yaml (sketch)
version: 1
default: deny
capabilities:
org_summary:
lanes: [green, blue]
list_transactions:
lanes: [green, blue]
constraints:
max_per_page: 100
max_pages_per_call: 5
max_window_days: 93
cost_run_rate_hints:
lanes: [green, blue]
snapshot_bundle:
compose_only: [org_summary, cost_run_rate_hints]
deny_classes:
spend:
match_prefixes: [create_, issue_, transfer_, payout_]
volume_cost:
match_tags: [card_operation, invoice_send, payment_link, subscription_change]
credential_exfil:
response_fields: [api_key, authorization_header, full_iban]
lanes:
green_blue:
allow_set: [org_summary, list_transactions, cost_run_rate_hints, snapshot_bundle]
yellow_plus:
same_as: green_blue
red:
human_only_in_qonto_app: true
Semantic rules beyond tool name:
- Reject unknown capability ids before touching Qonto.
- Enforce tenant binding before resource lookup; callers do not choose an arbitrary OpenBao path or tenant id.
- Inspect arguments for suspicious write semantics even if a route/tool is mis-wired later.
- Treat response shaping as policy, not presentation polish. Redaction rules should be versioned and testable like allow/deny rules.
- Emit stable deny reasons such as
unknown_capability,tenant_scope,arg_constraint,volume_cost,credential_exfil, andauthz_denied.
Map to AutonomyPolicy:
| Action class | Lane |
|---|---|
| Balance / txn read, CostRunRate update drafts | Green / Blue |
| Recommend plan downgrade / keep Qonto | Green (recommend only) |
| Change Qonto plan, create API key, transfer | Red |
4.7 Authentication and identity
Separate three identities clearly:
- Caller identity: human or workload calling REST/MCP.
- Assistant service identity: the only identity allowed to read
tenants/<tenant>/qonto-api. - Qonto upstream identity: the company credential presented to Qonto.
Caller
-> authenticates to qonto-assistant (OIDC workload identity, mTLS, or similar)
-> assistant checks claims + optional flex-auth grant (finance.qonto.read)
-> assistant resolves tenant from claims/policy
-> assistant fetches secret using its own runtime role
-> Qonto sees only the tenant's API identity
Rules:
- Never put bank credentials in harness env for general sessions.
- Do not reuse an
agent-harness-*OpenBao role for assistant secret fetch. Use a dedicated assistant runtime identity such asqonto-assistant-runtime. - Even in single-tenant dogfood, keep
tenant_idexplicit in claims and audit events so productization does not require reworking the core contract.
4.8 Secret and upstream session lifecycle
The blueprint needs an explicit secret lifecycle because this service is the sole credential choke point.
- OpenBao read happens inside the service only.
- Secrets may be cached in memory only with a short TTL to avoid fetching on every request.
- Cache invalidation should occur on startup, TTL expiry, and upstream 401/403.
- No secret material is written to disk, progress logs, traces, or chat output.
- Local developer mode may use env inject for mocks/tests, but production path remains OpenBao-first.
This keeps rotation and incident response tractable without coupling every request to a secret store round-trip.
4.9 Data handling and response shaping
The current blueprint mentions redaction, but this needs to be a first-class contract because the main exfil path is not only "the key" but over-sharing financial data into agent chat contexts.
| Data class | Examples | Default handling |
|---|---|---|
secret |
API key, Authorization header, OpenBao token | Never return, never log |
sensitive_financial |
full IBAN, account owner identifiers, raw vendor refs | Redact or truncate by default |
operational_summary |
balances, normalized counterparties, capped txn summaries | Return when capability allows |
Guidelines:
- Prefer normalized summary DTOs over raw vendor objects.
- Default to IBAN last4 or account alias, not full identifiers.
- Bulk export, statement download, and raw vendor dumps are out of scope for v1.
- If a future higher-assurance mode needs fuller data, it should be a separate capability with separate policy and audit semantics, not a flag hidden inside an existing read.
4.10 Observability and audit
Per-request audit should not depend on State Hub progress writes. State Hub is appropriate for work/progress records, not as the hot-path sink for every bank read.
Recommended split:
- Hot path: structured stdout/file/OTLP audit events for every request.
- Cold path: State Hub progress notes or operator summaries for notable sessions, incidents, or rollout evidence.
Suggested audit envelope:
request_id: req-...
timestamp: ...
actor: agt-... / workload id
tenant_id: binky
capability: list_transactions
protocol: rest | mcp
decision: allow | deny
deny_reason: volume_cost | unknown_capability | tenant_scope | authz_denied
policy_version: 1
latency_ms: ...
qonto_http_status: ...
result_count: ...
# never: Authorization header, API_KEY, OpenBao token, full IBAN by default
4.11 Operational guardrails (v1, not "later if we remember")
Basic safety and reliability controls belong in Phase 1 because one noisy agent session can otherwise degrade the single bank integration for everyone.
| Guardrail | Minimum requirement |
|---|---|
| Upstream timeout | bounded per request; fail fast and explicitly |
| Pagination | hard caps in policy, not caller preference |
| Concurrency | bounded per actor/tenant to avoid Qonto burst abuse |
| Rate limiting | basic per-actor or per-token limit from first deploy |
| Retry policy | conservative; no blind retries on ambiguous write-like failures |
| Persistence | no raw Qonto payload store in v1; ephemeral in-memory only |
| Failure mode | fail closed on policy/authn/authz uncertainty |
These controls move from "nice operational follow-up" to "architecture requirement" because the repo's entire point is governed centralization.
4.12 Placement in the repo map (options)
| Option | Repo | When |
|---|---|---|
| A (preferred dogfood) | New small service under Coulomb, e.g. qonto-assistant or generalize later to finance-connect |
Clean boundary, productizable |
| B | Module inside a future domain-connectors monorepo |
If many thin connectors share skeleton |
| C | Package under binky-control only |
Too tenant-specific; resists reuse |
Start as tenant-backed dogfood with a clear INTENT that it may graduate to a multi-tenant “governed bank connector” offer. Keep binky-control as consumer + policy source of business truth (CostRunRate, AutonomyPolicy), not as the runtime host long-term.
5. Alternatives considered
5.1 “Anchor policy at Qonto”
Unavailable for our custody model. Self-hosted MCP + API keys are not fine-scoped. Hosted OAuth MCP is the corporate anti-pattern we already rejected (credentials outside OpenBao, write tools bundled). Even with OAuth, spend prevention would still need our gateway for defense in depth.
5.2 Generic MCP gateway only
Good complement, insufficient alone. Use later for fleet aggregation. v1 needs semantic finance policy and a stable REST face for non-MCP clients — easier as our service.
5.3 agent-harness-only enforcement
Harness tool profiles remain necessary (don’t hand agents unrestricted network). But bank policy must not live only there: interactive Claude on a laptop would bypass harness profiles unless it also only sees the assistant endpoint. Network/policy gravity toward the assistant.
5.4 llm-connect as tool hub
Reject for layering reasons (LLM vs domain tools). Route models through llm-connect; route bank actions through qonto-assistant.
6. Phased delivery
Phase 0 — Already done (BINKY-WP-0005)
- Custody lane live
- REST pull proven
- CostRunRate TBC resolved for desk + plan
- Catalog + CCR active
Phase 1 — Policy kernel + REST (minimum useful product)
- Service skeleton with a protocol-neutral capability core and shared
decide(request, claims) -> Allow|Deny. - REST adapters for
GET /v1/accounts,GET /v1/transactions,GET /v1/snapshot. - Hard deny list for any write/spend path (even if not implemented).
- OpenBao fetch only inside service using a dedicated assistant runtime role.
- Baseline guardrails ship in Phase 1: bounded pagination, timeouts, basic rate limiting, bounded concurrency, and redaction tests.
- Smoke: CI uses mock Qonto; manual live read stays explicitly opt-in.
- Script replaces ad-hoc first-pull for CostRunRate refresh.
Phase 2 — MCP surface for all harnesses
- Streamable-HTTP MCP adapter on the same capability core and policy engine.
- Document one client config snippet for Claude/Codex/Cursor/Grok: URL + OIDC/workload auth — no bank secrets.
- agent-harness tool profile
finance-qonto-read→ assistant only. - MCP and REST emit the same structured audit schema; State Hub receives operator-level progress notes, not per-call hot-path events.
Phase 3 — Flex-auth + fleet
- flex-auth resource
finance.qonto.read(+ later.export). - Graduated quotas, differentiated response redaction profiles, and finer
capability scopes such as
.exportif ever approved. - Optional Envoy/gateway in front for multi-assistant mesh.
Phase 4 — Productization (dogfood → offer)
- Multi-tenant path layout
tenants/<slug>/qonto-api. - Policy packs per autonomy lane.
- OfferCatalog entry: “Governed bank assistant for agent fleets.”
7. Harness integration pattern (the consistency trick)
Wrong:
each harness → OpenBao qonto key → vendor MCP (local allow-list hope)
Right:
each harness → auth to qonto-assistant → policy → Qonto
↑
single tool catalog + single deny rules
Checklist for any new harness:
- Does not get
binky-qonto-apibank secret. - Gets only assistant base URL + client credentials for
finance.qonto.read. - Tool profile names assistant tools, not vendor tool names.
- Offline/dev: point at assistant mock, not a copy of production key.
8. Risk register (architecture-level)
| Risk | Mitigation |
|---|---|
| Assistant is a high-value target | High-risk path already in agent-high-risk-boundary; no raw stream to coding agents; short-lived OpenBao tokens; network policy |
| Model jailbreak asks for transfer | Tool not registered + deny classes; never implement spend in v1 |
| Vendor adds new MCP write tools | We do not auto-mirror vendor catalog; allow-list is ours |
| Bypass via direct thirdparty from laptop | Platform policy + founder discipline; optional egress controls later; secrets not on laptops |
| Over-filtering legitimate finance work | Explicit allow tools + CostRunRate helpers; expand by policy PR with tests |
| Cross-tenant data mix-up later | Tenant id explicit in claims, policy, audit, and OpenBao path resolution from v1 |
| Audit/log sink leaks sensitive fields | Redaction-by-class policy + tests; State Hub not used as raw per-request event store |
9. Decision requests (for later DEC / workplan)
- Adopt domain-assistant architecture (B) for Qonto (this blueprint).
- Repo home: new
qonto-assistantvs wait for genericfinance-connect. - v1 policy freeze: no spend / no volume-cost tools — hard deny.
- Service identity: dedicated assistant runtime role; no harness role reuse.
- MCP transport: remote streamable-HTTP only for production clients.
- Whether to ever run vendor MCP as internal backend (default: no).
Suggested workplan slug: BINKY-WP-0006 or a Coulomb-side
QONTO-WP-0001 once the repo exists.
10. References
Internal
integrations/qonto-mcp.md— variant decision, custody, first-pull resultsfinance/CostRunRate.md— consumer of read pathAutonomyPolicy.md— Green…Red lanes; spending defaultsagent-harnessINTENT + ADR-001 — single runtime; instances declarativellm-connectARCHITECTURE-LAYERS — LLM only, not domain toolsflex-authINTENT — authorization registry- ops-warden catalog
binky-qonto-api, CCR-2026-0008
External patterns (research snapshot 2026-07)
- Envoy AI Gateway MCPRoute: tool filtering, upstream key injection, OAuth, CEL on MCP params
- Kong AI Gateway MCP Tool ACLs: default-deny per-consumer tool grants
- MCP gateway product literature (Portkey, Speakeasy, Permit): central policy + audit for multi-client agents
- MCP security guidance: no token passthrough; tool-level scopes; proxy consent / confused-deputy awareness
- CSA / industry agentic MCP guides: inventory of servers; JIT escalation for high privilege
11. Summary recommendation
- Do not wire Qonto MCP into each harness.
- Do build a Qonto Governed Assistant with:
- sole possession of bank credentials,
- a protocol-neutral capability core behind REST and MCP adapters,
- default-deny tool catalog,
- no spend / no volume-cost policy as code,
- data classification/redaction rules,
- MCP + REST for all clients,
- flex-auth + OpenBao + audit.
- Treat generic MCP gateways as a later mesh layer, not a substitute for domain policy.
- Align agent-harness so it only ever grants access to that assistant — preserving ADR-001 without re-implementing bank rules per runtime.
This matches both industry MCP-gateway practice and your existing platform boundaries (custody, flex-auth, harness instances, dogfood).