Author INTENT and SCOPE; persist ArchitectureBlueprint under specs/ and MCP gateway research under research/. Register with State Hub (QONTO-WP- prefix); seed QONTO-WP-0001 (bootstrap, finished) and QONTO-WP-0002 (Phase 1 REST policy kernel, ready).
23 KiB
Architecture Blueprint — Governed Qonto Assistant
Status: blueprint v0.1 — 2026-07-21
Repo: qonto-assistant (canonical home).
Context: BINKY-WP-0005 finished (live OpenBao lane + first read-only pull).
Question: how do multiple coding agents / harnesses interact with Qonto conveniently and under one policy, without wiring vendor MCP into each client and accepting policy drift?
Related:INTENT.md,research/2026-07-21-mcp-gateway-and-governed-domain-assistant.md, binky-controlintegrations/qonto-mcp.md, AutonomyPolicy, agent-harness ADR-001, ops-warden catalogbinky-qonto-api, CCR-2026-0008.
1. Problem
We already can:
- Fetch company Qonto credentials from OpenBao (
tenants/binky/qonto-api). - Call Qonto’s thirdparty REST API (proven: first pull, CostRunRate v2).
- Map those credentials into the self-hosted
qonto/qonto-mcp-serverenv shape (QONTO_API_KEY/QONTO_ORGANIZATION_ID).
What we must not do as the long-term pattern:
- Install vendor MCP (or raw API keys) into Claude Desktop, Claude Code, Codex, Cursor, Grok, agent-harness, executor-worker, etc. separately.
- Rely on each harness’s local tool allow-list as the only spend control.
- Hope every client remembers “read-only forever.”
Why that fails:
| Failure mode | Effect |
|---|---|
| Policy drift | One harness allows a write tool; another forbids it. Same company key, different blast radius. |
| Credential sprawl | Keys appear in N configs, Docker args, session env dumps. |
| Audit fragmentation | No single log of “which agent did what to the bank.” |
| Qonto cannot help | Business API keys are not scope-limited server-side (confirmed in design). There is no “read-only API key” or MCP-side spend ACL from Qonto. |
| Volume-cost tools | Cards, invoicing, payment requests, transfers, and some “create” operations may incur fees or move money — must be blocked before the vendor sees them. |
Founder policy for now (normative intent):
Agents may inspect balances, transactions, and statements for control-plane and finance awareness. Agents must not initiate spending, transfers, card operations, or any action with volume/transaction cost. Payments stay Red lane (human in Qonto app).
That is stronger than “read-only MCP tools” in the abstract: it is a business policy that must be enforced at a choke point, not a client preference.
2. What industry patterns say
Research (2025–2026 MCP ecosystem) converges on the same shape your proxy intuition already points at.
2.1 MCP gateway / proxy as the control plane
Production patterns treat MCP clients (agent harnesses) as untrusted for tool policy, and put a gateway between them and tools:
- Envoy AI Gateway MCPRoute — multiplex backends, toolSelector
allow/deny, OAuth on the gateway, upstream API-key injection so clients
never hold vendor secrets, CEL rules on
tools/callparams. - Kong AI Gateway MCP Tool ACLs — default-deny, per-consumer tool subsets at the gateway (solves “all-or-nothing” MCP exposure).
- Speakeasy / Portkey / Permit MCP Gateway framing — authn, RBAC, audit of every tool call, policy evaluated on each call, not only at connect time.
- Harness’s own MCP guidance — gateways for centralized credentials, governance, audit, single endpoint vs N direct MCP connections.
Common lesson: clients connect to one governed endpoint; the gateway owns credentials, tool inventory, and policy.
2.2 MCP security best practices (protocol-level)
From MCP security guidance and industry write-ups:
| Practice | Application to Qonto |
|---|---|
| No token passthrough | Clients must not present the Qonto bank key. The assistant service is the only OAuth/API client to Qonto. |
| Tool-level scopes, not server-level | “Access to Qonto MCP” is too coarse; allow list_transactions, deny create_transfer. |
| Default deny | Unknown tools fail closed. |
| JIT escalation for high privilege | Any future spend path is Yellow/Orange/Red with explicit re-authorization — never session-global “write mode.” |
| Confused-deputy / static client ID | If the service ever OAuth-proxies, consent and audience binding matter; for API-key custody we avoid OAuth to Qonto entirely. |
| Audit end-to-end | Actor (harness/agent id) + tool + args hash + decision + latency; no secrets in logs. |
2.3 Domain capability service (not “another harness”)
A second pattern, older than MCP: BFF / capability API for a domain (banking, email, CRM). Agents call a narrow, policy-shaped interface (“get runway snapshot”) rather than the full vendor surface. That is how email-connect already behaves for IMAP: consumer never holds IMAP secrets in the product path; it goes through custody + a purpose-built reader.
llm-connect is the wrong home for this: it is multi-provider LLM abstraction (adapters, routing, cost models). Putting bank tools there blurs the layer map (ARCHITECTURE-LAYERS.md) and couples finance policy to model routing.
agent-harness is the wrong home for Qonto policy itself, but the right home for session binding: harness ADR-001 already says the harness is the credential holder and policy enforcement point for agent sessions. That must not mean “each harness reimplements bank policy.” It means the harness only grants agents a named tool profile that talks to already-governed domain services. Domain policy lives below the harness.
2.4 Synthesis
| Pattern | Use for Qonto? |
|---|---|
| Wire vendor MCP into every harness | No — policy drift, key sprawl |
| Generic MCP gateway only (Envoy/Kong in front of vendor MCP) | Partial — good transport + tool filter; weak on semantic spend rules and finance-shaped tools |
| Governed domain assistant (our MCP + REST, policy inside) | Yes — recommended |
| Flex-auth decisions on every call | Yes — who may call; lane/risk attributes |
| OpenBao + ops-warden for bank key | Yes — already live; only the assistant fetches |
Your “qonto-proxy governed API & MCP service” is the right class of solution. Name it as a capability assistant, not only a dumb reverse proxy: it should own a safe tool surface, not only filter someone else’s.
3. Goals and non-goals
Goals
- One policy for “no spend / no volume-cost actions” across all agents.
- One credential path (existing OpenBao lane); harnesses never see bank keys.
- Convenient agent UX: natural “ask the finance assistant” via MCP and/or small REST helpers used by scripts and rhythm sessions.
- Auditability: every call attributable to agent/harness/session.
- Fits NetKingdom: flex-auth, OpenBao, ops-warden, agent-harness instances, activity-core schedules — no parallel IAM.
- Dogfood-ready: internal use is product evidence (DogfoodPolicy).
Non-goals (v1)
- Replacing DATEV / StB / DUO accounting.
- Automating payments, SEPA, card issuance, expense submission.
- Hosted Qonto OAuth connector (
mcp.qonto.com) as corporate custody path. - Making llm-connect a tool gateway.
- Per-harness custom Qonto integrations.
4. Recommended architecture
4.1 One-liner
Build a small “Qonto Governed Assistant” service that is the only component allowed to hold the bank API key and speak Qonto’s protocol. Expose a policy-shaped MCP server and a thin REST API. All harnesses are dumb clients of that service. Policy is code + tests in one place.
4.2 Context diagram
┌──────────────────────────────────────────────────────────────────────────┐
│ Agent clients (many) │
│ Claude Code · Codex · Cursor · Grok · agent-harness sessions · scripts │
└───────────────┬───────────────────────────────┬──────────────────────────┘
│ MCP (streamable-HTTP) │ REST (JSON)
│ tools/list · tools/call │ /v1/snapshot, /v1/txns, …
▼ ▼
┌──────────────────────────────────────────────────────────────────────────┐
│ Qonto Governed Assistant («qonto-assistant») │
│ ┌─────────────┐ ┌──────────────────┐ ┌─────────────────────────────┐ │
│ │ Authn edge │→ │ Policy gate │→ │ Capability tools │ │
│ │ (OIDC/mTLS │ │ (default-deny │ │ (finance-shaped, not 1:1 │ │
│ │ flex-auth) │ │ tool+args+lane) │ │ vendor dump) │ │
│ └─────────────┘ └────────┬─────────┘ └──────────────┬──────────────┘ │
│ │ audit │ │
│ ▼ ▼ │
│ State Hub / logs Qonto client (REST) │
│ (metadata only) Authorization: user:key │
└───────────────────────────────────────────────┬──────────────────────────┘
│ only this process
▼
OpenBao tenants/binky/qonto-api
(via warden / AppRole — short TTL)
│
▼
Qonto thirdparty API
Optional later: put Envoy AI Gateway (or similar) in front of
qonto-assistant for fleet-wide MCP aggregation with other domain assistants
(email, issue-core, …). Do not put the gateway directly in front of
vendor qonto-mcp-server as the sole control — semantic policy is thinner
there.
4.3 Why a domain assistant beats “proxy the vendor MCP only”
| Approach | Pros | Cons |
|---|---|---|
| A. Gateway + vendor MCP | Fast; toolSelector; industry-standard | Vendor tool names/args still drive the model; hard to encode “no volume cost”; vendor MCP is unmaintained / dual-use with write tools |
| B. Domain assistant (recommended) | Safe tool catalog we design; semantic policy; dual MCP+REST; dogfood product surface | Build cost |
| C. Per-harness allow-lists | Zero new service | Guaranteed drift; keys leak into clients |
Hybrid that works: implement B; if we still want vendor MCP for parity experiments, run it only as an internal backend behind the same policy gate (never published to harnesses). Prefer direct REST client inside the assistant (we already proved REST) — fewer moving parts than shelling out to unmaintained MCP.
4.4 Component responsibilities
| Component | Owns | Does not own |
|---|---|---|
| qonto-assistant | Tool catalog, spend/volume policy, Qonto REST client, dual surfaces (MCP+REST), call audit | LLM routing, scheduling, identity issuance |
| OpenBao + ops-warden | Custody path, rotation guide, high-risk agent boundary | Business rules about transfers |
| flex-auth | “May actor X call capability finance.qonto.read?” |
Bank credential material |
| agent-harness | Session lifecycle; grants profile finance-qonto-read; injects assistant URL/token only |
Bank policy implementation |
| llm-connect | Which model answers | Tool execution |
| activity-core | When Finance Steward runs | How bank is called |
| binky-control | CostRunRate, queues, AutonomyPolicy, dogfood evidence | Runtime of the assistant |
| Qonto | Bank of record | Our agent policy |
4.5 Surfaces (same policy, two protocols)
MCP (agent-native)
- Transport: streamable-HTTP (remote), optional stdio only for local dev with the same binary and policy module.
tools/listreturns only approved tools (default-deny).tools/callre-checks policy (tool id + arguments + actor claims) — connect-time allow-list alone is insufficient.
REST (script / rhythm / tests)
- Stable JSON for CostRunRate refresh, CI smoke, non-MCP agents.
- Same policy middleware as MCP (shared library, one decision function).
Example tool catalog (v1 — illustrative):
| Tool / endpoint | Purpose | Policy |
|---|---|---|
qonto_org_summary |
Org + account balances | Allow (Green/Blue) |
qonto_list_transactions |
Filtered history | Allow; hard caps on page size; no export bulk to chat by default |
qonto_cost_run_rate_hints |
Map debits to known CostRunRate rows | Allow |
qonto_find_counterparties |
Search labels | Allow |
| (any transfer / payment / card / invoice create) | — | Deny always (v1) |
| (vendor MCP write tools) | — | Not registered |
4.6 Policy model (v1)
Encode as code + declarative YAML, tested in CI — not wiki prose alone.
# policy/qonto-v1.yaml (sketch)
version: 1
default: deny
allow:
- id: org_summary
- id: list_transactions
constraints:
max_per_page: 100
max_pages_per_call: 5
- id: cost_run_rate_hints
deny_classes:
- spend # transfers, payouts, direct debits initiation
- volume_cost # card ops, invoicing sends, paid features that bill per use
- credential_exfil # tools that return raw API keys or full IBANs if avoidable
lanes:
green_blue: allow_set: [org_summary, list_transactions, cost_run_rate_hints]
yellow_plus: same_as_green_blue # no spend even if human is "nearby"
red: human_only_in_qonto_app
Semantic rules (beyond tool name):
- Inspect arguments: reject if tool is “create_*”, amounts > 0 with side effects, or known write operation types.
- Prefer response shaping: return IBAN last4, not full IBAN, in agent channels unless a higher assurance mode is granted later.
- Volume-cost: deny anything that creates a fee-bearing Qonto operation (subscription change is also Red — out of band).
Map to AutonomyPolicy:
| Action class | Lane |
|---|---|
| Balance / txn read, CostRunRate update drafts | Green / Blue |
| Recommend plan downgrade / keep Qonto | Green (recommend only) |
| Change Qonto plan, create API key, transfer | Red |
4.7 Authentication and identity
Harness / human agent
→ authenticates to qonto-assistant (OIDC netkingdom or mTLS workload id)
→ flex-auth: finance.qonto.read (or finer)
→ assistant fetches bank secret with its own AppRole / short-lived token
→ Qonto sees only the company API identity (not the coding agent)
- Never put
API_KEYin harness env for general sessions. - Optional AppRole
agent-harness-binky-qontomirroring mail lane — but scoped so onlyqonto-assistantcan read the path (not every agent process). Agents authenticate to the assistant, not to OpenBao for bank secrets.
4.8 Placement in the repo map (options)
| Option | Repo | When |
|---|---|---|
| A (preferred dogfood) | New small service under Coulomb, e.g. qonto-assistant or generalize later to finance-connect |
Clean boundary, productizable |
| B | Module inside a future domain-connectors monorepo |
If many thin connectors share skeleton |
| C | Package under binky-control only |
Too tenant-specific; resists reuse |
Start as tenant-backed dogfood with a clear INTENT that it may graduate to a multi-tenant “governed bank connector” offer. Keep binky-control as consumer + policy source of business truth (CostRunRate, AutonomyPolicy), not as the runtime host long-term.
4.9 Observability and evidence
Every call records (State Hub progress or dedicated audit log):
actor: agt-… / harness session id
capability: finance.qonto.read
tool: list_transactions
decision: allow | deny
deny_reason: volume_cost | unknown_tool | flex_auth | rate_limit
latency_ms: …
qonto_http_status: …
# never: Authorization header, API_KEY, full account numbers if avoidable
Finance Steward rhythm writes metadata into finance/ only (same rule
as first pull).
5. Alternatives considered
5.1 “Anchor policy at Qonto”
Unavailable for our custody model. Self-hosted MCP + API keys are not fine-scoped. Hosted OAuth MCP is the corporate anti-pattern we already rejected (credentials outside OpenBao, write tools bundled). Even with OAuth, spend prevention would still need our gateway for defense in depth.
5.2 Generic MCP gateway only
Good complement, insufficient alone. Use later for fleet aggregation. v1 needs semantic finance policy and a stable REST face for non-MCP clients — easier as our service.
5.3 agent-harness-only enforcement
Harness tool profiles remain necessary (don’t hand agents unrestricted network). But bank policy must not live only there: interactive Claude on a laptop would bypass harness profiles unless it also only sees the assistant endpoint. Network/policy gravity toward the assistant.
5.4 llm-connect as tool hub
Reject for layering reasons (LLM vs domain tools). Route models through llm-connect; route bank actions through qonto-assistant.
6. Phased delivery
Phase 0 — Already done (BINKY-WP-0005)
- Custody lane live
- REST pull proven
- CostRunRate TBC resolved for desk + plan
- Catalog + CCR active
Phase 1 — Policy kernel + REST (minimum useful product)
- Service skeleton with shared
decide(tool, args, claims) -> Allow|Deny. - REST:
GET /v1/accounts,GET /v1/transactions,GET /v1/snapshot. - Hard deny list for any write/spend path (even if not implemented).
- OpenBao fetch only inside service (AppRole or OIDC role).
- Smoke: CI uses mock Qonto; manual: live read against dogfood account.
- Script replaces ad-hoc first-pull for CostRunRate refresh.
Phase 2 — MCP surface for all harnesses
- Streamable-HTTP MCP on the same decision function.
- Document one client config snippet for Claude/Codex/Cursor/Grok: URL + OIDC/workload auth — no bank secrets.
- agent-harness tool profile
finance-qonto-read→ assistant only. - Audit events to State Hub.
Phase 3 — Flex-auth + fleet
- flex-auth resource
finance.qonto.read(+ later.export). - Rate limits, concurrency caps, optional response redaction modes.
- Optional Envoy/gateway in front for multi-assistant mesh.
Phase 4 — Productization (dogfood → offer)
- Multi-tenant path layout
tenants/<slug>/qonto-api. - Policy packs per autonomy lane.
- OfferCatalog entry: “Governed bank assistant for agent fleets.”
7. Harness integration pattern (the consistency trick)
Wrong:
each harness → OpenBao qonto key → vendor MCP (local allow-list hope)
Right:
each harness → auth to qonto-assistant → policy → Qonto
↑
single tool catalog + single deny rules
Checklist for any new harness:
- Does not get
binky-qonto-apibank secret. - Gets only assistant base URL + client credentials for
finance.qonto.read. - Tool profile names assistant tools, not vendor tool names.
- Offline/dev: point at assistant mock, not a copy of production key.
8. Risk register (architecture-level)
| Risk | Mitigation |
|---|---|
| Assistant is a high-value target | High-risk path already in agent-high-risk-boundary; no raw stream to coding agents; short-lived OpenBao tokens; network policy |
| Model jailbreak asks for transfer | Tool not registered + deny classes; never implement spend in v1 |
| Vendor adds new MCP write tools | We do not auto-mirror vendor catalog; allow-list is ours |
| Bypass via direct thirdparty from laptop | Platform policy + founder discipline; optional egress controls later; secrets not on laptops |
| Over-filtering legitimate finance work | Explicit allow tools + CostRunRate helpers; expand by policy PR with tests |
9. Decision requests (for later DEC / workplan)
- Adopt domain-assistant architecture (B) for Qonto (this blueprint).
- Repo home: new
qonto-assistantvs wait for genericfinance-connect. - v1 policy freeze: no spend / no volume-cost tools — hard deny.
- MCP transport: remote streamable-HTTP only for production clients.
- Whether to ever run vendor MCP as internal backend (default: no).
Suggested workplan slug: BINKY-WP-0006 or a Coulomb-side
QONTO-WP-0001 once the repo exists.
10. References
Internal
integrations/qonto-mcp.md— variant decision, custody, first-pull resultsfinance/CostRunRate.md— consumer of read pathAutonomyPolicy.md— Green…Red lanes; spending defaultsagent-harnessINTENT + ADR-001 — single runtime; instances declarativellm-connectARCHITECTURE-LAYERS — LLM only, not domain toolsflex-authINTENT — authorization registry- ops-warden catalog
binky-qonto-api, CCR-2026-0008
External patterns (research snapshot 2026-07)
- Envoy AI Gateway MCPRoute: tool filtering, upstream key injection, OAuth, CEL on MCP params
- Kong AI Gateway MCP Tool ACLs: default-deny per-consumer tool grants
- MCP gateway product literature (Portkey, Speakeasy, Permit): central policy + audit for multi-client agents
- MCP security guidance: no token passthrough; tool-level scopes; proxy consent / confused-deputy awareness
- CSA / industry agentic MCP guides: inventory of servers; JIT escalation for high privilege
11. Summary recommendation
- Do not wire Qonto MCP into each harness.
- Do build a Qonto Governed Assistant with:
- sole possession of bank credentials,
- default-deny tool catalog,
- no spend / no volume-cost policy as code,
- MCP + REST for all clients,
- flex-auth + OpenBao + audit.
- Treat generic MCP gateways as a later mesh layer, not a substitute for domain policy.
- Align agent-harness so it only ever grants access to that assistant — preserving ADR-001 without re-implementing bank rules per runtime.
This matches both industry MCP-gateway practice and your existing platform boundaries (custody, flex-auth, harness instances, dogfood).