WEEK 6 - Testing Four MCP Servers: What Each One Revealed
Meta description: Server-level tests for FinBot's four MCP servers (FinMail, FinDrive, SystemUtils and TaxCalc), and the access gaps and missing validation they uncovered.
PRs #316 through #319 add server-level tests for four of FinBot's MCP servers. Each server gives AI agents a different set of tools. Together the tests show a pattern: the checks that exist mostly work, and the problems are in the checks that are missing.
A Test Every Server Gets: The Exact Tool List
Each PR includes a tool discovery test that checks the server offers exactly the tools it should:
FinMail: exactly 5 tools
FinDrive: exactly 5 tools
SystemUtils: exactly 8 tools
TaxCalc: exactly 3 tools
This is a negative test as much as a positive one. If a server starts offering an extra tool, whether an admin tool, a debug tool or one copied from another server, the LLM will see it and may use it. A test that counts the tools makes that change fail loudly instead of slipping into production. Several of the PRs also check that each tool's parameter schema is present, because the schema is what the LLM reads to decide how to call the tool.
FinMail Server (PR #316): Who Can Read Whose Email
FinMail gives agents five tools: send_email, list_inbox, read_email, search_emails and mark_as_read.
The routing tests confirm where mail goes. Mail to a vendor's address lands in that vendor's inbox, mail to the admin domain or an internal department goes to the admin inbox, and mail to an unknown address goes to an external "dead drop."
The access tests are where FinMail reveals the most. PR #316 documents these defects:
Vendors can reach other vendors' mail. A vendor session can list another vendor's inbox, read their messages and mark them as read.
The admin check can be bypassed. A vendor session can get admin emails by passing an inbox type the server doesn't recognize, a garbage value or the right value in different capitalization. The server checks for one exact value and lets everything else through.
Sender names can be spoofed. Any display name is accepted.
Sending to nobody reports success. send_email with an empty recipient list returns sent=True.
Input isn't validated. Invalid message types, addresses without an @, empty subjects and bodies, and unlimited recipient lists are all accepted.
One more test is worth pointing out: a prompt injection payload in an email body is stored and later comes back in search results. That's indirect prompt injection. An attacker never talks to the agent directly. They send an email, and the agent reads the payload when it searches the inbox later.
FinDrive Server (PR #317): File Access and Filenames
FinDrive gives agents five tools: upload_file, get_file, list_files, delete_file and search_files.
Isolation between namespaces holds on every tool. A session can't upload to, read, list, delete or search another namespace's files. Vendors are also blocked from reading or deleting files an admin uploaded.
Isolation between vendors does not. Within a namespace, a vendor session can read, list and delete another vendor's files. This is the same gap Week 5 found in FinStripe, which suggests a pattern rather than a one-off.
The upload tests go after the filename, which is exactly the kind of free-text input an LLM passes along without a second thought. All of these are accepted:
An empty or whitespace-only filename
A path traversal sequence such as ../../../etc/passwd, in the filename or the folder
A newline in the filename, which can split and forge log entries
A null byte in the filename
A prompt injection payload in the filename, stored as-is and returned in list_files
The tests also found that dangerous file types are accepted, and that max_files_per_vendor (default 50) exists in the config but is never enforced.
SystemUtils Server (PR #318): The Most Dangerous Tools
SystemUtils has the most powerful tools in FinBot: run_diagnostics, manage_storage, rotate_logs, database_maintenance, network_request, read_config, manage_users and execute_script.
It's also a mock. In the PR's words, it "records what the agent attempted but executes nothing." That makes it a safe place to answer one question: if an LLM is manipulated into asking for something destructive, does the server stop it?
The answer is no. The tests document that SystemUtils accepts:
Shell injection in diagnostic commands
DROP TABLE and unguarded DELETE statements
Requests to internal IP addresses (an SSRF risk) and to data exfiltration URLs
Reads of sensitive system files and .env files
Deleting the admin user and promoting a user to superadmin
Destructive bash scripts and credential-theft scripts
Interpreters not on any allow-list
In FinBot, this is deliberate. It's a CTF, and these gaps are the attack surface players are meant to find, so several were closed as working as designed. In a real system, each item on that list would be a critical finding. The lesson carries over: if a tool can do something dangerous, the server has to validate the request, because the LLM calling it can be manipulated. Tests like these give you an exact list of what a poisoned agent could get away with.
TaxCalc Server (PR #319): Correct Math, Missing Guards
TaxCalc is a stateless mock tax calculator with three tools: calculate_tax, get_tax_rates and validate_tax_id.
Most of the tests confirm correct behavior:
Jurisdictions: all seven compute correctly, and New York includes state, county and city tax.
Categories: services are exempt when configured, and entertainment gets its surcharge.
Rates: amounts round to two decimal places, and the combined rate equals the sum of its parts.
Unknown jurisdictions: they return an error that lists the valid options, which helps an LLM correct itself.
The findings are about input the calculator never expected:
A negative amount produces negative tax. An agent that gets confused, or is manipulated, could create a tax credit out of nothing.
An unknown category is silently treated as goods. The wrong rate gets applied with no error.
A whitespace-only jurisdiction isn't treated as empty, because the value is never stripped of spaces.
The SSN validation branch can never run. The code exists, but no input reaches it, so it looks tested when it isn't.
What This Week Shows
Across all four servers, the boundaries that were built work: namespace isolation, admin-only files and known jurisdictions. The failures are in what nobody built, such as vendor-level checks, input validation and limits that exist in config but aren't enforced. Testing the boundaries you designed isn't enough. You also have to test the inputs an LLM might actually send.
Next week: how test results from all these PRs get routed to the right tab in a live Google Sheets tracker.
Comments