qonto-assistant/specs/ArchitectureBlueprint.md
tegwick e1ab23f83e Bootstrap qonto-assistant: intent, blueprint, research, workplans
Author INTENT and SCOPE; persist ArchitectureBlueprint under specs/ and MCP
gateway research under research/. Register with State Hub (QONTO-WP- prefix);
seed QONTO-WP-0001 (bootstrap, finished) and QONTO-WP-0002 (Phase 1 REST
policy kernel, ready).
2026-07-21 23:32:24 +02:00

499 lines
23 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Architecture Blueprint — Governed Qonto Assistant
> Status: blueprint v0.1 — 2026-07-21
> Repo: **qonto-assistant** (canonical home).
> Context: BINKY-WP-0005 finished (live OpenBao lane + first read-only pull).
> Question: how do multiple coding agents / harnesses interact with Qonto
> **conveniently and under one policy**, without wiring vendor MCP into each
> client and accepting policy drift?
> Related: `INTENT.md`, `research/2026-07-21-mcp-gateway-and-governed-domain-assistant.md`,
> binky-control `integrations/qonto-mcp.md`, AutonomyPolicy, agent-harness
> ADR-001, ops-warden catalog `binky-qonto-api`, CCR-2026-0008.
---
## 1. Problem
We already can:
1. Fetch company Qonto credentials from OpenBao (`tenants/binky/qonto-api`).
2. Call Qontos thirdparty REST API (proven: first pull, CostRunRate v2).
3. Map those credentials into the self-hosted `qonto/qonto-mcp-server` env
shape (`QONTO_API_KEY` / `QONTO_ORGANIZATION_ID`).
What we must **not** do as the long-term pattern:
- Install vendor MCP (or raw API keys) into Claude Desktop, Claude Code,
Codex, Cursor, Grok, agent-harness, executor-worker, etc. **separately**.
- Rely on each harnesss local tool allow-list as the only spend control.
- Hope every client remembers “read-only forever.”
Why that fails:
| Failure mode | Effect |
| --- | --- |
| **Policy drift** | One harness allows a write tool; another forbids it. Same company key, different blast radius. |
| **Credential sprawl** | Keys appear in N configs, Docker args, session env dumps. |
| **Audit fragmentation** | No single log of “which agent did what to the bank.” |
| **Qonto cannot help** | Business API keys are **not** scope-limited server-side (confirmed in design). There is no “read-only API key” or MCP-side spend ACL from Qonto. |
| **Volume-cost tools** | Cards, invoicing, payment requests, transfers, and some “create” operations may incur fees or move money — must be blocked **before** the vendor sees them. |
**Founder policy for now (normative intent):**
> Agents may inspect balances, transactions, and statements for control-plane
> and finance awareness. Agents must **not** initiate spending, transfers,
> card operations, or any action with volume/transaction cost. Payments stay
> Red lane (human in Qonto app).
That is stronger than “read-only MCP tools” in the abstract: it is a **business
policy** that must be enforced at a choke point, not a client preference.
---
## 2. What industry patterns say
Research (20252026 MCP ecosystem) converges on the same shape your proxy
intuition already points at.
### 2.1 MCP gateway / proxy as the control plane
Production patterns treat MCP **clients** (agent harnesses) as untrusted for
tool policy, and put a **gateway** between them and tools:
- **Envoy AI Gateway MCPRoute** — multiplex backends, **toolSelector**
allow/deny, OAuth on the gateway, upstream API-key injection so clients
never hold vendor secrets, CEL rules on `tools/call` params.
- **Kong AI Gateway MCP Tool ACLs** — default-deny, per-consumer tool
subsets at the gateway (solves “all-or-nothing” MCP exposure).
- **Speakeasy / Portkey / Permit MCP Gateway** framing — authn, RBAC, audit
of every tool call, policy evaluated **on each call**, not only at
connect time.
- **Harnesss own MCP guidance** — gateways for centralized credentials,
governance, audit, single endpoint vs N direct MCP connections.
Common lesson: **clients connect to one governed endpoint; the gateway owns
credentials, tool inventory, and policy.**
### 2.2 MCP security best practices (protocol-level)
From MCP security guidance and industry write-ups:
| Practice | Application to Qonto |
| --- | --- |
| **No token passthrough** | Clients must not present the Qonto bank key. The assistant service is the only OAuth/API client to Qonto. |
| **Tool-level scopes, not server-level** | “Access to Qonto MCP” is too coarse; allow `list_transactions`, deny `create_transfer`. |
| **Default deny** | Unknown tools fail closed. |
| **JIT escalation for high privilege** | Any future spend path is Yellow/Orange/Red with explicit re-authorization — never session-global “write mode.” |
| **Confused-deputy / static client ID** | If the service ever OAuth-proxies, consent and audience binding matter; for API-key custody we avoid OAuth to Qonto entirely. |
| **Audit end-to-end** | Actor (harness/agent id) + tool + args hash + decision + latency; no secrets in logs. |
### 2.3 Domain capability service (not “another harness”)
A second pattern, older than MCP: **BFF / capability API** for a domain
(banking, email, CRM). Agents call a **narrow, policy-shaped interface**
(“get runway snapshot”) rather than the full vendor surface. That is how
email-connect already behaves for IMAP: consumer never holds IMAP secrets
in the product path; it goes through custody + a purpose-built reader.
**llm-connect is the wrong home** for this: it is multi-provider **LLM**
abstraction (adapters, routing, cost models). Putting bank tools there
blurs the layer map (ARCHITECTURE-LAYERS.md) and couples finance policy to
model routing.
**agent-harness is the wrong home for Qonto policy itself**, but the right
home for *session* binding: harness ADR-001 already says the harness is the
credential holder and policy enforcement point **for agent sessions**. That
must not mean “each harness reimplements bank policy.” It means the harness
only grants agents a **named tool profile** that talks to **already-governed**
domain services. Domain policy lives **below** the harness.
### 2.4 Synthesis
| Pattern | Use for Qonto? |
| --- | --- |
| Wire vendor MCP into every harness | **No** — policy drift, key sprawl |
| Generic MCP gateway only (Envoy/Kong in front of vendor MCP) | **Partial** — good transport + tool filter; weak on **semantic** spend rules and finance-shaped tools |
| **Governed domain assistant** (our MCP + REST, policy inside) | **Yes — recommended** |
| Flex-auth decisions on every call | **Yes** — who may call; lane/risk attributes |
| OpenBao + ops-warden for bank key | **Yes** — already live; only the assistant fetches |
Your “qonto-proxy governed API & MCP service” is the right **class** of
solution. Name it as a **capability assistant**, not only a dumb reverse
proxy: it should *own* a safe tool surface, not only filter someone elses.
---
## 3. Goals and non-goals
### Goals
1. **One policy** for “no spend / no volume-cost actions” across all agents.
2. **One credential** path (existing OpenBao lane); harnesses never see bank keys.
3. **Convenient agent UX**: natural “ask the finance assistant” via MCP and/or
small REST helpers used by scripts and rhythm sessions.
4. **Auditability**: every call attributable to agent/harness/session.
5. **Fits NetKingdom**: flex-auth, OpenBao, ops-warden, agent-harness
instances, activity-core schedules — no parallel IAM.
6. **Dogfood-ready**: internal use is product evidence (DogfoodPolicy).
### Non-goals (v1)
- Replacing DATEV / StB / DUO accounting.
- Automating payments, SEPA, card issuance, expense submission.
- Hosted Qonto OAuth connector (`mcp.qonto.com`) as corporate custody path.
- Making llm-connect a tool gateway.
- Per-harness custom Qonto integrations.
---
## 4. Recommended architecture
### 4.1 One-liner
> **Build a small “Qonto Governed Assistant” service that is the only
> component allowed to hold the bank API key and speak Qontos protocol.
> Expose a **policy-shaped** MCP server and a thin REST API. All harnesses
> are dumb clients of that service. Policy is code + tests in one place.**
### 4.2 Context diagram
```text
┌──────────────────────────────────────────────────────────────────────────┐
│ Agent clients (many) │
│ Claude Code · Codex · Cursor · Grok · agent-harness sessions · scripts │
└───────────────┬───────────────────────────────┬──────────────────────────┘
│ MCP (streamable-HTTP) │ REST (JSON)
│ tools/list · tools/call │ /v1/snapshot, /v1/txns, …
▼ ▼
┌──────────────────────────────────────────────────────────────────────────┐
│ Qonto Governed Assistant («qonto-assistant») │
│ ┌─────────────┐ ┌──────────────────┐ ┌─────────────────────────────┐ │
│ │ Authn edge │→ │ Policy gate │→ │ Capability tools │ │
│ │ (OIDC/mTLS │ │ (default-deny │ │ (finance-shaped, not 1:1 │ │
│ │ flex-auth) │ │ tool+args+lane) │ │ vendor dump) │ │
│ └─────────────┘ └────────┬─────────┘ └──────────────┬──────────────┘ │
│ │ audit │ │
│ ▼ ▼ │
│ State Hub / logs Qonto client (REST) │
│ (metadata only) Authorization: user:key │
└───────────────────────────────────────────────┬──────────────────────────┘
│ only this process
OpenBao tenants/binky/qonto-api
(via warden / AppRole — short TTL)
Qonto thirdparty API
```
Optional later: put **Envoy AI Gateway** (or similar) **in front of**
`qonto-assistant` for fleet-wide MCP aggregation with other domain assistants
(email, issue-core, …). Do **not** put the gateway *directly* in front of
vendor `qonto-mcp-server` as the sole control — semantic policy is thinner
there.
### 4.3 Why a domain assistant beats “proxy the vendor MCP only”
| Approach | Pros | Cons |
| --- | --- | --- |
| **A. Gateway + vendor MCP** | Fast; toolSelector; industry-standard | Vendor tool names/args still drive the model; hard to encode “no volume cost”; vendor MCP is unmaintained / dual-use with write tools |
| **B. Domain assistant (recommended)** | Safe tool catalog we design; semantic policy; dual MCP+REST; dogfood product surface | Build cost |
| **C. Per-harness allow-lists** | Zero new service | Guaranteed drift; keys leak into clients |
**Hybrid that works:** implement **B**; if we still want vendor MCP for
parity experiments, run it **only as an internal backend** behind the same
policy gate (never published to harnesses). Prefer **direct REST client**
inside the assistant (we already proved REST) — fewer moving parts than
shelling out to unmaintained MCP.
### 4.4 Component responsibilities
| Component | Owns | Does not own |
| --- | --- | --- |
| **qonto-assistant** | Tool catalog, spend/volume policy, Qonto REST client, dual surfaces (MCP+REST), call audit | LLM routing, scheduling, identity issuance |
| **OpenBao + ops-warden** | Custody path, rotation guide, high-risk agent boundary | Business rules about transfers |
| **flex-auth** | “May actor X call capability `finance.qonto.read`?” | Bank credential material |
| **agent-harness** | Session lifecycle; grants profile `finance-qonto-read`; injects *assistant* URL/token only | Bank policy implementation |
| **llm-connect** | Which model answers | Tool execution |
| **activity-core** | When Finance Steward runs | How bank is called |
| **binky-control** | CostRunRate, queues, AutonomyPolicy, dogfood evidence | Runtime of the assistant |
| **Qonto** | Bank of record | Our agent policy |
### 4.5 Surfaces (same policy, two protocols)
**MCP (agent-native)**
- Transport: streamable-HTTP (remote), optional stdio only for local dev
with the **same** binary and policy module.
- `tools/list` returns **only** approved tools (default-deny).
- `tools/call` re-checks policy (tool id + arguments + actor claims) —
connect-time allow-list alone is insufficient.
**REST (script / rhythm / tests)**
- Stable JSON for CostRunRate refresh, CI smoke, non-MCP agents.
- Same policy middleware as MCP (shared library, one decision function).
Example tool catalog (v1 — illustrative):
| Tool / endpoint | Purpose | Policy |
| --- | --- | --- |
| `qonto_org_summary` | Org + account balances | Allow (Green/Blue) |
| `qonto_list_transactions` | Filtered history | Allow; hard caps on page size; no export bulk to chat by default |
| `qonto_cost_run_rate_hints` | Map debits to known CostRunRate rows | Allow |
| `qonto_find_counterparties` | Search labels | Allow |
| *(any transfer / payment / card / invoice create)* | — | **Deny always** (v1) |
| *(vendor MCP write tools)* | — | Not registered |
### 4.6 Policy model (v1)
Encode as **code + declarative YAML**, tested in CI — not wiki prose alone.
```yaml
# policy/qonto-v1.yaml (sketch)
version: 1
default: deny
allow:
- id: org_summary
- id: list_transactions
constraints:
max_per_page: 100
max_pages_per_call: 5
- id: cost_run_rate_hints
deny_classes:
- spend # transfers, payouts, direct debits initiation
- volume_cost # card ops, invoicing sends, paid features that bill per use
- credential_exfil # tools that return raw API keys or full IBANs if avoidable
lanes:
green_blue: allow_set: [org_summary, list_transactions, cost_run_rate_hints]
yellow_plus: same_as_green_blue # no spend even if human is "nearby"
red: human_only_in_qonto_app
```
**Semantic rules (beyond tool name):**
- Inspect arguments: reject if tool is “create_*”, amounts > 0 with side
effects, or known write operation types.
- Prefer **response shaping**: return IBAN last4, not full IBAN, in agent
channels unless a higher assurance mode is granted later.
- **Volume-cost**: deny anything that creates a fee-bearing Qonto operation
(subscription change is also Red — out of band).
Map to AutonomyPolicy:
| Action class | Lane |
| --- | --- |
| Balance / txn read, CostRunRate update drafts | Green / Blue |
| Recommend plan downgrade / keep Qonto | Green (recommend only) |
| Change Qonto plan, create API key, transfer | Red |
### 4.7 Authentication and identity
```text
Harness / human agent
→ authenticates to qonto-assistant (OIDC netkingdom or mTLS workload id)
→ flex-auth: finance.qonto.read (or finer)
→ assistant fetches bank secret with its own AppRole / short-lived token
→ Qonto sees only the company API identity (not the coding agent)
```
- **Never** put `API_KEY` in harness env for general sessions.
- Optional **AppRole** `agent-harness-binky-qonto` mirroring mail lane — but
scoped so only `qonto-assistant` can read the path (not every agent
process). Agents authenticate *to the assistant*, not to OpenBao for
bank secrets.
### 4.8 Placement in the repo map (options)
| Option | Repo | When |
| --- | --- | --- |
| **A (preferred dogfood)** | New small service under Coulomb, e.g. `qonto-assistant` or generalize later to `finance-connect` | Clean boundary, productizable |
| **B** | Module inside a future `domain-connectors` monorepo | If many thin connectors share skeleton |
| **C** | Package under `binky-control` only | Too tenant-specific; resists reuse |
Start as **tenant-backed dogfood** with a clear INTENT that it may graduate
to a multi-tenant “governed bank connector” offer. Keep binky-control as
*consumer + policy source of business truth* (CostRunRate, AutonomyPolicy),
not as the runtime host long-term.
### 4.9 Observability and evidence
Every call records (State Hub progress or dedicated audit log):
```yaml
actor: agt-… / harness session id
capability: finance.qonto.read
tool: list_transactions
decision: allow | deny
deny_reason: volume_cost | unknown_tool | flex_auth | rate_limit
latency_ms:
qonto_http_status:
# never: Authorization header, API_KEY, full account numbers if avoidable
```
Finance Steward rhythm writes **metadata** into `finance/` only (same rule
as first pull).
---
## 5. Alternatives considered
### 5.1 “Anchor policy at Qonto”
**Unavailable for our custody model.** Self-hosted MCP + API keys are not
fine-scoped. Hosted OAuth MCP is the corporate **anti-pattern** we already
rejected (credentials outside OpenBao, write tools bundled). Even with
OAuth, spend prevention would still need our gateway for defense in depth.
### 5.2 Generic MCP gateway only
**Good complement, insufficient alone.** Use later for fleet aggregation.
v1 needs **semantic** finance policy and a stable REST face for non-MCP
clients — easier as our service.
### 5.3 agent-harness-only enforcement
Harness tool profiles remain necessary (dont hand agents unrestricted
network). But **bank policy must not live only there**: interactive Claude
on a laptop would bypass harness profiles unless it also only sees the
assistant endpoint. **Network/policy gravity toward the assistant.**
### 5.4 llm-connect as tool hub
Reject for layering reasons (LLM vs domain tools). Route models through
llm-connect; route **bank actions** through qonto-assistant.
---
## 6. Phased delivery
### Phase 0 — Already done (BINKY-WP-0005)
- [x] Custody lane live
- [x] REST pull proven
- [x] CostRunRate TBC resolved for desk + plan
- [x] Catalog + CCR active
### Phase 1 — Policy kernel + REST (minimum useful product)
1. Service skeleton with shared `decide(tool, args, claims) -> Allow|Deny`.
2. REST: `GET /v1/accounts`, `GET /v1/transactions`, `GET /v1/snapshot`.
3. Hard deny list for any write/spend path (even if not implemented).
4. OpenBao fetch only inside service (AppRole or OIDC role).
5. Smoke: CI uses mock Qonto; manual: live read against dogfood account.
6. Script replaces ad-hoc first-pull for CostRunRate refresh.
### Phase 2 — MCP surface for all harnesses
1. Streamable-HTTP MCP on the same decision function.
2. Document **one** client config snippet for Claude/Codex/Cursor/Grok:
URL + OIDC/workload auth — **no bank secrets**.
3. agent-harness tool profile `finance-qonto-read` → assistant only.
4. Audit events to State Hub.
### Phase 3 — Flex-auth + fleet
1. flex-auth resource `finance.qonto.read` (+ later `.export`).
2. Rate limits, concurrency caps, optional response redaction modes.
3. Optional Envoy/gateway in front for multi-assistant mesh.
### Phase 4 — Productization (dogfood → offer)
1. Multi-tenant path layout `tenants/<slug>/qonto-api`.
2. Policy packs per autonomy lane.
3. OfferCatalog entry: “Governed bank assistant for agent fleets.”
---
## 7. Harness integration pattern (the consistency trick)
**Wrong:**
```text
each harness → OpenBao qonto key → vendor MCP (local allow-list hope)
```
**Right:**
```text
each harness → auth to qonto-assistant → policy → Qonto
single tool catalog + single deny rules
```
Checklist for any new harness:
1. Does **not** get `binky-qonto-api` bank secret.
2. Gets only assistant base URL + client credentials for `finance.qonto.read`.
3. Tool profile names assistant tools, not vendor tool names.
4. Offline/dev: point at assistant mock, not a copy of production key.
---
## 8. Risk register (architecture-level)
| Risk | Mitigation |
| --- | --- |
| Assistant is a high-value target | High-risk path already in agent-high-risk-boundary; no raw stream to coding agents; short-lived OpenBao tokens; network policy |
| Model jailbreak asks for transfer | Tool not registered + deny classes; never implement spend in v1 |
| Vendor adds new MCP write tools | We do not auto-mirror vendor catalog; allow-list is ours |
| Bypass via direct thirdparty from laptop | Platform policy + founder discipline; optional egress controls later; secrets not on laptops |
| Over-filtering legitimate finance work | Explicit allow tools + CostRunRate helpers; expand by policy PR with tests |
---
## 9. Decision requests (for later DEC / workplan)
1. **Adopt domain-assistant architecture (B)** for Qonto (this blueprint).
2. **Repo home:** new `qonto-assistant` vs wait for generic `finance-connect`.
3. **v1 policy freeze:** no spend / no volume-cost tools — hard deny.
4. **MCP transport:** remote streamable-HTTP only for production clients.
5. **Whether** to ever run vendor MCP as internal backend (default: **no**).
Suggested workplan slug: `BINKY-WP-0006` or a Coulomb-side
`QONTO-WP-0001` once the repo exists.
---
## 10. References
### Internal
- `integrations/qonto-mcp.md` — variant decision, custody, first-pull results
- `finance/CostRunRate.md` — consumer of read path
- `AutonomyPolicy.md` — Green…Red lanes; spending defaults
- `agent-harness` INTENT + ADR-001 — single runtime; instances declarative
- `llm-connect` ARCHITECTURE-LAYERS — LLM only, not domain tools
- `flex-auth` INTENT — authorization registry
- ops-warden catalog `binky-qonto-api`, CCR-2026-0008
### External patterns (research snapshot 2026-07)
- Envoy AI Gateway MCPRoute: tool filtering, upstream key injection, OAuth,
CEL on MCP params
- Kong AI Gateway MCP Tool ACLs: default-deny per-consumer tool grants
- MCP gateway product literature (Portkey, Speakeasy, Permit): central
policy + audit for multi-client agents
- MCP security guidance: no token passthrough; tool-level scopes; proxy
consent / confused-deputy awareness
- CSA / industry agentic MCP guides: inventory of servers; JIT escalation
for high privilege
---
## 11. Summary recommendation
1. **Do not** wire Qonto MCP into each harness.
2. **Do** build a **Qonto Governed Assistant** with:
- sole possession of bank credentials,
- default-deny tool catalog,
- **no spend / no volume-cost** policy as code,
- MCP + REST for all clients,
- flex-auth + OpenBao + audit.
3. Treat generic MCP gateways as a **later mesh layer**, not a substitute
for domain policy.
4. Align agent-harness so it only ever grants access to that assistant —
preserving ADR-001 without re-implementing bank rules per runtime.
This matches both industry MCP-gateway practice and your existing
platform boundaries (custody, flex-auth, harness instances, dogfood).