Skip to main content

Command Palette

Search for a command to run...

The Tool Call Problem

Updated
•9 min read•View as Markdown
M
I write about anything in the tech space

One of the things I spend a lot of time thinking about at work is what happens when an AI assistant decides to call a tool it should not, or calls the right tool with the wrong arguments. The assistant has access to real operations: creating products, editing data and code. A wrong tool call is not a bad response you can ignore. It is an action that happened.

This is the tool call problem and it is more common than people talk about.

How Bad Is It Actually

Benchmarks on tool call accuracy look decent on paper. The BFCL v4 leaderboard, updated July 2026, tracks 13 models and puts the current average score at 0.6, with the best model at 0.8. That means even top models get tool calls wrong 20% of the time under controlled benchmark conditions. In production it gets worse.

BFCL v4 shifted to a more realistic evaluation model covering multi-step agentic tasks, multi-turn conversations, and a hallucination measurement category that specifically tests whether models correctly refuse to call any tool when nothing in the schema matches the user's query. Models that score well on simple single-tool calls often fall apart in multi-turn scenarios where they need to track prior results across several rounds.

Error recovery rates are where the small vs large model gap shows up most clearly. On BFCL v4 multi-turn, larger models recover from a failed tool call more than 50% of the time. Smaller general-purpose models like Qwen3-8B average around 20%, typically falling into hallucinated retry loops until the conversation collapses.

Small models fine-tuned specifically for tool use tell a different story though. Research on the xLAM family shows a 1B parameter model fine-tuned for function calling outperforming much larger general-purpose models on BFCL. For well-scoped, predictable tool sets, a small specialised model can be more reliable and significantly cheaper than routing everything through a frontier model.

The specific failure modes you see in practice are:

Hallucinated tools. The model calls a function that does not exist. It invented a tool name that sounds plausible given the context.

Wrong arguments. The model calls the right tool but passes incorrect values, missing fields, or fields with the wrong type.

Phantom tool calls. The model claims it called a tool in its response text without actually emitting a tool call. You get something like "I've updated your order status" in the response but no tool call object anywhere. This is more common with Gemini and many open source models than it is with frontier models. It is also one of the harder failures to catch because your validation layer never sees a tool call to reject. The operation simply did not happen and the user thinks it did.

One possible approach is to cross-check the model's response text against the actual tool calls that ran. If the response claims an action was taken but no corresponding tool call exists in the turn, treat it as a failure.

def detect_phantom_tool_call(response_text, tool_calls_executed, tool_action_phrases):
    for phrase in tool_action_phrases:
        if phrase in response_text.lower():
            if not tool_calls_executed:
                return True, f"Response claims action was taken but no tool call was executed"
    return False, None

# Example usage
action_phrases = ["i've updated", "i've deleted", "i've created", "done, i've", "i have added"]
phantom, reason = detect_phantom_tool_call(response.text, response.tool_calls, action_phrases)
if phantom:
    # Re-prompt the model or surface the failure to the user
    pass

Unnecessary tool calls. The model calls a tool when it could have answered from context. This costs tokens and latency and sometimes causes unintended side effects.

Chained failures. One wrong tool call produces bad output that feeds into the next tool call. By the time you notice, several operations have run on bad data.

Why Safety Nets Matter More Than Accuracy

The instinct when you see accuracy numbers like these is to switch models or tune prompts until the numbers improve. That helps but it does not solve the problem. Even a 95% accurate model fails 1 in 20 calls. At scale that is a lot of failed tool calls.

The more reliable approach is to build your system so that wrong tool calls are caught before they cause damage, or fail gracefully when they do run. Accuracy improvements and safety nets work together, not as alternatives.

Building Safety Nets Without Burning Tokens

The obvious approach is to run a second LLM call to validate every tool call before executing it. That works but it doubles your token usage on every interaction.

Schema validation first. Before any tool call reaches execution, validate it against your tool schemas. Check that the function name exists, that all required arguments are present, and that argument types match. This can catch hallucinated tools and malformed calls instantly with no LLM call required.

One thing the happy path misses: most LLM SDKs return tool call arguments as a raw JSON string. If the model returns malformed JSON, your validation crashes before it even runs. Wrapping the parse in a try/except handles this.

def validate_tool_call(tool_call, tool_schemas):
    name = tool_call.function.name
    if name not in tool_schemas:
        return False, f"Tool '{name}' does not exist"

    try:
        args = json.loads(tool_call.function.arguments)
    except json.JSONDecodeError as e:
        return False, f"Malformed arguments JSON: {e}"

    schema = tool_schemas[name]
    required = schema.get("required", [])

    for field in required:
        if field not in args:
            return False, f"Missing required field: {field}"

    return True, None

Rule-based checks for high-risk tools. After schema validation, you can run lightweight checks for operations that are hard to reverse. Keyword matching on user messages is too brittle to rely on here. A user saying "clean up my old products" contains no delete keyword but clearly implies deletion. A user saying "don't delete anything" contains the word delete but is the opposite of a deletion intent.

One option that works better is checking the relationship between the tool being called and the tools that ran immediately before it. A delete_product call that follows a list_products call is plausible. A delete_product call as the very first tool call in a session with a vague user message is suspicious. That context is already available without any extra model call.

HIGH_RISK_TOOLS = {"delete_product", "update_order_status", "delete_page"}

def plausibility_check(tool_call, prior_tool_calls):
    name = tool_call.function.name
    if name not in HIGH_RISK_TOOLS:
        return True, None

    prior_names = [t.function.name for t in prior_tool_calls]
    if not prior_names:
        return False, f"High-risk tool '{name}' called with no prior context"

    return True, None

Semantic and embedding-based intent matching. A lighter alternative to a full LLM validation call is using a small embedding model or sentence transformer to compare the user's message against the description of the tool being called. If the semantic similarity is below a threshold, flag it for review or block it. This is cheaper than a full LLM call and more reliable than keyword matching. A small QA model like DistilBERT fine-tuned on intent classification can also work here for well-defined tool sets.

from sentence_transformers import SentenceTransformer, util

model = SentenceTransformer("all-MiniLM-L6-v2")

def intent_matches_tool(user_message, tool_description, threshold=0.4):
    embeddings = model.encode([user_message, tool_description])
    score = util.cos_sim(embeddings[0], embeddings[1]).item()
    return score >= threshold, score

This is not a definitive signal on its own but combined with schema validation and context checks it adds a meaningful layer with minimal cost.

Reserve LLM validation for high-risk calls only. Only escalate to a second LLM call for tools that touch sensitive operations, things that are hard to reverse. Using structured output here avoids the free-text parsing problems that come with asking a model for a yes/no answer.

from pydantic import BaseModel

class ToolCallDecision(BaseModel):
    approved: bool
    reason: str

async def validate_high_risk(tool_call, user_message, conversation_history):
    prompt = f"""
    Conversation so far: {conversation_history}
    Latest user message: {user_message}
    Tool called: {tool_call.function.name}
    Arguments: {tool_call.function.arguments}

    Does this tool call clearly match what the user wants?
    """
    response = await llm.complete(
        prompt,
        response_model=ToolCallDecision,
        max_tokens=100
    )
    return response.approved

Retry with error feedback before giving up. When a tool call fails validation, feeding the error back to the model and retrying once can be worth doing before returning a fallback. Models often self-correct when told specifically what went wrong.

async def execute_with_retry(tool_call, user_message, prior_tool_calls):
    valid, error = validate_tool_call(tool_call, TOOL_SCHEMAS)
    if not valid:
        correction_prompt = f"Your tool call failed: {error}. Please correct it."
        corrected = await llm.complete(correction_prompt)
        tool_call = corrected.tool_calls[0] if corrected.tool_calls else None
        if not tool_call:
            return {"status": "failed", "action": "ask_clarification"}
        valid, error = validate_tool_call(tool_call, TOOL_SCHEMAS)
        if not valid:
            return {"status": "failed", "action": "ask_clarification"}

    plausible, reason = plausibility_check(tool_call, prior_tool_calls)
    if not plausible:
        return {"status": "failed", "action": "request_confirmation", "reason": reason}

    if tool_call.function.name in HIGH_RISK_TOOLS:
        approved = await validate_high_risk(tool_call, user_message, prior_tool_calls)
        if not approved:
            return {"status": "failed", "action": "request_confirmation"}

    result = await execute_tool(tool_call)
    return {"status": "success", "result": result}

The action field tells your agent loop what to do next. ask_clarification means the agent asks the user to rephrase. request_confirmation means it describes what it was about to do and asks for explicit approval before trying again. The agent loop consuming this needs to handle both cases and produce a sensible user-facing message, which is where most implementations fall short.

Detection After the Fact

Schema validation catches structurally broken calls. It does not catch the harder failure: a tool call that ran successfully but did the wrong thing. No exception was raised, the tool returned a result, and the wrong operation happened.

The signals to watch for in your logs are unexpected tool sequences and mismatches between what the user said and which tool ran. Neither of these are errors your system will surface automatically. You have to look for them.

LangFuse makes this practical. Every tool call shows up in the trace with the full conversation context around it, so you can spot patterns across sessions rather than debugging individual interactions in isolation. Setting up a regular review of high-risk tool calls specifically is one of the more useful habits to build. The patterns you find there will tell you where your validation rules need tightening more reliably than any benchmark will.

The Broader Point

Tool calls are where LLM errors stop being abstract and start having consequences. Schema validation and a retry loop handle most structural failures cheaply. Selective LLM validation covers the ambiguous high-risk cases. Logging catches what gets through all of that.

None of these steps are expensive individually. The expensive mistake is skipping them and finding out through a user complaint that your agent deleted something it should not have.