WEEK 5 — Testing MCP: When Your AI Agent's Tools Live on a Server
Meta description: How FinBot tests its MCP layer: repositories, session scoping, and the server factory and tool provider, plus the vendor access gaps the tests exposed.
Most people building with AI agents are familiar with function calling. You define a list of tools, pass them to the LLM, and the LLM calls them by name.
Model Context Protocol (MCP) goes a step further. Instead of defining tools inside the application, the agent connects to an MCP server, a separate component that offers tools, resources and prompts through a standard protocol. FinBot uses MCP servers for several of its capabilities, including payments (FinStripe), email (FinMail) and file storage (FinDrive). PRs #311 through #315 add tests for the foundation of that layer: the repositories, the FinStripe server's session access, and the factory and provider that create servers and hand their tools to agents. What Is Being Tested MCP testing splits into two layers that look similar but fail in different ways. The repository layer is data access.
A repository answers questions like:
Which transfers belong to this vendor?
What files are in this vendor's FinDrive folder?
Which emails are unread?
Repository tests check that the right data comes back for the right session, and that a session can't reach data outside its scope. The server layer is what the LLM talks to. It lists the tools the LLM can call, handles the MCP message format and passes requests to the repository for the actual data. Server tests check that the right tools are listed, that tool calls behave correctly, and that errors come back in a form the LLM can understand.
A repository bug gives you wrong data. A server bug can give you the right data in the wrong format, the wrong tools, or data the caller should never have seen. And as the FinStripe tests show, a repository layer that passes every test does not make the server above it safe.
Why the Tool Description Is Part of the Contract One part of MCP testing has no real equivalent in standard API testing: the tool description is part of the contract. When an MCP server lists a tool, it sends the LLM a name, a description and a parameter schema. The LLM reads the description to decide when and how to call the tool. A description that is vague, or that says more than it should, is a defect, even if the code behind it works. That's why PR #315 tests _apply_tool_overrides, which the PR calls FinBot's "supply chain attack surface." FinBot lets configuration override a tool's description. If someone can change what a description says, they can change when and how the LLM uses the tool, without touching any code. The tests confirm that:
A description override is applied to the matching tool.
An override naming a tool that doesn't exist doesn't crash the provider.
An override without a description is skipped.
No overrides, or no provider, means nothing is changed.
Once a description can be changed from outside the code, it's input, and it needs tests like any other input. FinStripe: Vendor Session Access (PR #311) FinStripe is FinBot's mock Stripe payment processor. Its MCP server has four tools: create_transfer, get_transfer, get_account_balance and list_transfers. PR #311 tests all four, with a focus on who can see and move what. The session is passed to the server when the server is created, not with each call: python server = create_finstripe_server(session) Every tool call on that server runs in that session's context. So the tests create a separate server for each session and check that the boundary holds: python class TestGetTransfer: async def test_mcp_get_003_namespace_isolation(self, db): """MCP-GET-003: get_transfer does not return transfers from other namespaces""" session_a = session_manager.create_session(email="a@example.com") session_b = session_manager.create_session(email="b@example.com") vendor = make_vendor(db, session_a, email="vendor@test.com") invoice = make_invoice(db, session_a, vendor.id) server_a = create_finstripe_server(session_a) server_b = create_finstripe_server(session_b) created = await call( server_a, "create_transfer", vendor_account="123456789012", amount=1000.0, invoice_reference="INV-001", vendor_id=vendor.id, invoice_id=invoice.id, ) result = await call(server_b, "get_transfer", transfer_id=created["transfer_id"]) assert "error" in result Session A creates a transfer, session B asks for it, and session B gets an error. Isolation between namespaces works. Isolation between vendors does not. PR #311 includes bug-exposing tests that document these defects in the FinStripe server: A vendor session can retrieve another vendor's transfer. A vendor session can pay a different vendor, which opens the door to cross-vendor fraud. A vendor session with no vendor_id skips the identity check entirely. The same invoice can be paid twice. An invoice that hasn't been approved is accepted without a status check. A zero-amount transfer is allowed. This is the same scoping rule from Week 4's tool tests, now checked at the MCP level, and the result is different. The tools layer and the MCP layer are separate ways into the same data, so each one has to enforce the rules on its own. Passing tests in one layer tell you nothing about the other. The Repositories: FinStripe, FinDrive and FinMail (PRs #312–#314) The repository tests are less dramatic, and that's a good result. FinStripe (#313) tests transactions: listing them by invoice and by vendor, updating status, looking them up by transfer ID, and paging with limit and offset. It checks that queries are scoped to the vendor and isolated between namespaces, and that unknown IDs are handled. FinDrive (#314) tests files: deleting, updating, counting, listing by path and limiting results. It checks that namespaces are isolated and that missing files are handled. FinMail (#312) tests email queries: vendor and admin filters, unread counts, marking messages as read in bulk, paging and isolation between vendors. It also tests address routing: mail sent to a user address goes to the admin inbox, and mail sent to an unknown address goes to the external inbox. All three PRs report the same result: every test passes, and no bugs were found in the repository layer. The routing tests are worth pointing out. Sending unknown addresses to the external inbox is a design decision, not a bug. Writing a test for it makes the decision visible. If someone later changes where unknown mail goes, a test fails and the change has to be deliberate. The Factory and Provider (PR #315) The MCP factory creates server instances. The MCPToolProvider connects to a server, discovers its tools and hands them to the agent. Neither contains business logic, but every agent depends on both at startup. A mistake here fails quietly. If the agent gets the wrong server or the wrong tools, nothing crashes. It just works with fewer abilities, or different ones, than expected. The tests cover: Server creation: asking for an unknown server type returns None instead of quietly substituting another server. The caller has to handle the missing server, and a wrong server can't slip in unnoticed. Connecting and disconnecting: the provider opens and closes its connection to the server. Tool definitions: get_tool_definitions returns tools in the OpenAI function format, and get_callables returns exactly one callable per registered tool. Tool overrides: the supply chain tests described above.
What This Week Shows The layers of FinBot's MCP stack don't all fail in the same place. The repositories passed every test. Isolation between namespaces held at the server layer. The real gaps were in vendor access checks in the FinStripe server, the layer the LLM calls directly. Only testing each layer on its own finds that. Next week: testing the individual MCP servers, including FinMail, FinDrive, SystemUtils and TaxCalc.
Comments