LAB/015 · Running · Self-directed, not client work

Tool-Calling Reliability of Local Models

Three local models were given tools and 1,698 requests from the Berkeley Function Calling Leaderboard. About 85% of answers were right. Almost every wrong answer was a well-formed call, so a schema check fixed nothing.

The question
If we let a local model call our systems (an API, a database, a form), how often does it press the right button with the right values, and does validating its calls make it safe?
What it showed
About 85% of answers were correct for all three model sizes; the 14B model was no better than the 4B. Real user requests were much harder than benchmark-style ones (67–72% vs 94–96%). Fewer than 2 in 1,000 answers failed a schema check, so the guardrail changed nothing.
What it means for you
The dangerous errors are valid calls with wrong values: a percentage as 7 instead of 0.07, a time in the wrong format. Guard tool calls with business rules and confirmations, write precise tool descriptions, and test on your own requests.

Want this measured on your own data? AI feasibility diagnostic, €1,900 · about one week.

An AI agent is a language model that can call tools: look up an order, book a meeting, update a record. Whether it is safe to connect one to real systems comes down to a plain question: how often does it call the right tool, with the right values, and how often does it call a tool when it should not? This Lab measures that for three local models, and tests the most common safeguard: checking each call against the tool’s schema.

The test

  • Requests: 1,698 items from the Berkeley Function Calling Leaderboard v3 (Apache-2.0): benchmark-style requests with one or several tools offered, requests contributed by real users with their own tools, and requests where no offered tool fits and the right answer is not to call one.
  • Models: Qwen3-4B-Instruct-2507, Qwen3-8B and Qwen3-14B, all compressed to 4 bits and running locally on one RTX 4070 with native tool calling.
  • Scoring: right function, every argument within the accepted values, no invented arguments; for “no tool fits” items, correct means no call. A simplified version of the leaderboard’s own check, on a subset, so the numbers are not directly comparable with the official leaderboard.
  • Guardrail: each call is checked against the tool’s own schema (does the function exist, are required arguments present, are types and allowed values right). If not, the model is shown the error and asked once more.

Results

Correct behaviour (%), by kind of request, raw tool calling. Median time per request on one RTX 4070.
One tool offeredPick from several toolsReal-world, one toolNo tool fits (should not call)Real-world, several toolsReal-world, no tool fitsAllMedian time
Qwen3-4B Instruct-2507 Q4_K_M96.0%92.0%72.1%91.2%73.7%88.3%85.9%0.37 s
Qwen3-8B Q4_K_M94.0%92.0%66.7%89.6%70.7%83.7%83.0%0.57 s
Qwen3-14B Q4_K_M95.5%90.5%71.3%90.8%72.0%88.0%85.1%0.99 s
Subset and simplified checker, so not directly comparable with the official leaderboard · LAB/015 · measured 2026-10-09 · RTX 4070 12 GB

Size did not buy reliability. The 4B model was as accurate as the 14B model, and almost three times faster. The 8B model was slightly worse than both.

Real requests are harder. With one tool offered, benchmark-style requests were answered correctly 94–96% of the time; requests written by real users, for their own tools, only 67–72%:

Benchmark-style vs real-world requests, one tool offeredCorrect (%) · real-world requests and tools were contributed by users

Qwen3-4B Instruct-2507 Q4_K_M

  1. Benchmark-style96.0%
  2. Real-world72.1%

Qwen3-8B Q4_K_M

  1. Benchmark-style94.0%
  2. Real-world66.7%

Qwen3-14B Q4_K_M

  1. Benchmark-style95.5%
  2. Real-world71.3%

The guardrail that did nothing

The schema guardrail: answers that failed the schema check (and were retried), and overall accuracy with and without it.
Failed the schema checkAccuracy, rawAccuracy, with guardrail
Qwen3-4B Instruct-2507 Q4_K_M285.9%85.9%
Qwen3-8B Q4_K_M383.0%83.0%
Qwen3-14B Q4_K_M185.1%85.1%
LAB/015 · measured 2026-10-09 · RTX 4070 12 GB

Out of 1,698 answers per model, between one and three failed the schema check. The calls were almost always well-formed, so validating them caught almost nothing. The mistakes were in the meaning:

Why answers were wrong, all three models together (5,094 answers).
Answers
Right function, wrong argument value337
Called a tool when none fitted188
No call when a tool fitted138
Missing a required argument73
Several calls instead of one24
Wrong function18
Invented an argument2
LAB/015 · measured 2026-10-09 · RTX 4070 12 GB

The largest group is a call to the right function with a wrong value. A schema cannot catch these: the value has the right type, it is just not what the user meant.

A valid call, 100 times wrong (Qwen3-14B Q4_K_M)Wrong
Request
Predict the total expected profit of stocks XYZ in 5 years given I have invested $5000 and annual return rate is 7%.
Tool argument, as described to the model
annual_return (float): The annual return rate of the investment.
Qwen3-14B Q4_K_M called
investment_predictProfit({"investment_amount": 5000, "annual_return": 7, "years": 5})
Expected
annual_return = 0.07. Qwen3-4B Instruct-2507 Q4_K_M and Qwen3-8B Q4_K_M got it right.

BFCL v3 item simple_134 (Apache-2.0).

The tool’s description did not say whether the rate is a fraction or a percentage, so the model had to guess. The two smaller models guessed right; the largest did not. Many wrong values are like this: a date or time in a different format, a city written differently, a default the user did not ask for. Some are strict labels in the benchmark itself, which the official leaderboard also counts as wrong.

The second-largest group is calling a tool when none fits: 188 answers across the three models, about 12% of the requests where the right answer was not to call anything.

What this means in practice

  • Schema validation is necessary, not sufficient. Keep it, but do not count it as a safety measure. Real protection comes from business rules (is this amount plausible, does this customer exist), confirmation before anything irreversible, and limits on what each tool can do.
  • Tool descriptions are part of the program. State units, formats and ranges (“rate as a fraction, 0.07 for 7%”). Ambiguity in a description becomes a wrong value in a call.
  • A bigger model is not automatically a safer one. Measure on your own requests before paying for the extra memory and latency.
  • Test with real requests. Benchmark-style requests made every model look far better than real users’ requests did.
Method and environment
data
Berkeley Function Calling Leaderboard v3 (Apache-2.0): simple, multiple, live simple, irrelevance (all), live multiple and live irrelevance (300 each, seed 15); 1,698 items, first turn only
models
Qwen3-4B-Instruct-2507, Qwen3-8B, Qwen3-14B (Q4_K_M, llama.cpp native tool calling, thinking off, greedy)
system prompt
You are an assistant with access to tools. Call a tool only if it can answer the request; otherwise reply in text.
scoring
simplified BFCL AST check: right function, every ground-truth argument within its accepted values (strings standardised as BFCL does, nested dicts key by key), no undefined arguments; irrelevance: no call. The first scoring pass was too strict on nested dicts and punctuation and was corrected before publication (rescore.py)
guardrail
the same first answer validated against the tool's own schema (known function, required arguments, types, enums); on failure the error is returned and the model asked once more
environment
gpu: RTX 4070 12 GB
measured
2026-10-09

Limits

First turn only, a subset of the leaderboard, and a simplified checker. The first scoring pass compared nested arguments and punctuation more strictly than the leaderboard does; it was corrected and all answers rescored from the saved calls before publication (150 of 10,188 answers changed). The guardrail checks types, required arguments and allowed values only; with other servers or APIs, malformed calls may be more common and the guardrail more useful. Thinking mode was off for all models. The code is in labs/tool-calling in this site’s repository.

New Lab results by email

A short email when a new measurement is published, a few times a year. No newsletter fluff, reply to unsubscribe.