top of page

WEEK 4 - Testing the Tools Between AI Decisions and Real Actions

carocsteads
Sep 16
5 min read


Meta description: 

How FinBot tests AI agent tools: why schemas and functions need separate tests, plus namespace isolation, payment safeguards and the bugs the tests found.


When an AI agent decides to take an action, such as looking up a vendor, approving an invoice or starting a payment, it does so by calling a tool. Tools connect what the LLM decides to what actually happens in the system. They are also one of the least-tested layers in AI applications.

PR #227 adds FinBot's first complete tool unit test suite. PR #279 adds tests for payment tools. Together they cover the layer between an agent's decisions and real database writes.


What a Tool Actually Is


In FinBot, a tool is an async Python function that the LLM can request by name. Say a vendor asks, "What is the status of my invoice?" The LLM does not query the database, because it can't. It has no database connection, no session and no knowledge of what data exists. It has a list of tools it may call and a description of each one.

The LLM reads those descriptions, picks the tool that matches what the user wants and generates a tool call as a JSON object:

json

{
  "name": "get_invoice_status",
  "arguments": { "invoice_id": 42 }
}

The agent framework receives that JSON, checks the arguments against the tool's parameter schema and runs the Python function. The function queries the database and returns the result. The LLM then uses that result to write the response the user sees.

Every step from the user's message to the final response goes through tools. The LLM never touches the database directly.

That means every tool has two interfaces, and each needs its own tests.


The schema interface


Before the function ever runs, the LLM has read a JSON definition of the tool: its name, what it does, which arguments it accepts and their types. The LLM uses that definition to decide when to call the tool and what to pass.

  • If the name is too close to another tool's name, the LLM may call the wrong one.

  • If the description is vague, the LLM may call the tool when it shouldn't, or skip it when it should.

  • If a required argument is missing from the schema, the LLM won't know to send it.

None of these mistakes causes an obvious error when you write it.

The schema is a Python dictionary that gets turned into JSON and sent to the model. It looks like configuration, not code. No type checker will tell you a description is misleading, and no linter will flag a missing required field. The only way to know the schema is right is to test it.


The execution interface


This is the Python function itself, and it's what most people think of when they hear "testing a tool."

  • Does the function return the right data?

  • Does it handle bad input without crashing?

  • Does it fail clearly when a record isn't found?

Here's a test from PR #227 that covers the last question:

python

async def test_inv_get_002_raises_on_missing_invoice(self, db):
    """INV-GET-002: get_invoice_details raises ValueError for missing invoice"""
    session = session_manager.create_session(email="test@example.com")

    with pytest.raises(ValueError, match="Invoice not found"):
        await get_invoice_details(99999, session)

These questions matter, but they only cover half the picture:

  • A function that works perfectly but has a wrong description will be called at the wrong time, or never.

  • A correct schema that points to a function with a type-handling bug will fail at runtime. The failure is hard to trace because the LLM's tool call looked correct.

The two interfaces can break separately. Changing the schema doesn't update the function, and refactoring the function doesn't update the schema. Over time they can drift apart without either one raising an error. PRs #227 and #279 cover the execution interface. Schema tests that check names, descriptions, required fields, parameter types and enum values are the next layer to add, because they catch that drift before the LLM starts behaving in ways nobody can explain.


Namespace Isolation


One rule these tests enforce: a tool must only reach data in the current session's namespace.

get_vendor_details takes a vendor ID as an argument. If it looked up any ID it was given, any session could read any vendor's data. That's broken object-level authorization (BOLA), #1 on the OWASP API Security Top 10. The fix is to search for the ID only within the current session's namespace, so a vendor that belongs to another session comes back as "not found."

Here's the test from PR #227:

python

class TestGetVendorDetails:
    async def test_vnd_get_003_namespace_isolation(self, db):
        """VND-GET-003: get_vendor_details cannot access vendor from another namespace"""
        session_a = session_manager.create_session(email="user_a@example.com")
        session_b = session_manager.create_session(email="user_b@example.com")
        vendor = make_vendor(db, session_a)

        with pytest.raises(ValueError, match="Vendor not found"):
            await get_vendor_details(vendor.id, session_b)

Session A creates a vendor, and session B asks for it by ID. From session B's point of view, the vendor doesn't exist.

This is a common gap in agentic AI systems. You can instruct the LLM to request only the current vendor's data, but the tool is where that rule is actually enforced. If the tool doesn't enforce it, the instruction alone won't protect anyone.


Payment Tool Testing



PR #279 adds tests for payment tools, where a mistake means money moves.

Namespace isolation. The same rule applies to payments. The test creates an invoice in one session, tries to pay it from a different session and asserts that the payment fails.

No double payments. Paying the same invoice twice should never send the money twice. The test pays an approved invoice, then calls process_payment again on that invoice with a new payment reference. The second call must raise an error. Once an invoice is marked paid, it can't be paid again, even when the request looks different.

These rules sound obvious once you state them, but they are easy to skip. With an LLM tool, the caller is a model that can make mistakes or be manipulated, so the tool itself has to enforce them.


What the Tests Found


PR #227 includes tests that document real defects in FinBot's tools:

  • Any status is accepted. update_vendor_status accepts any status string, including "hacked", without checking it against the allowed values. A manipulated agent could put a vendor into a state the rest of the system doesn't expect.

  • None is saved as text. Calling update_vendor_agent_notes with None writes the literal string "None" to the database instead of leaving the field empty.

  • Database sessions are left open. When a tool raises an error partway through, its database session isn't closed.

None of these bugs produces an error, so the system looks fine until bad data builds up or connections run out. Tests that feed tools the inputs an LLM might actually send, not just the ideal ones, are how you find them.


Next week: what happens when AI agents get their tools from a remote server, and how to test a Model Context Protocol (MCP) implementation.



 
 
 

Recent Posts

See All
Week 3: Testing the Challenge Detection Layer

Most test suites verify that a system does the right thing when inputs are normal. FinBot adds a harder requirement: verify that the system notices when something wrong is happening, even when the wro

 
 
 

Comments


© 2023 by Carolina Steadham. All rights reserved.

bottom of page