Skip to content

Function calling and tool use

Goal

Teach the model when and how to call external tools.

Tools may include:

  • Search
  • Calculator
  • Calendar
  • Database
  • Code interpreter
  • Browser
  • File retrieval
  • CRM
  • Email
  • Weather API
  • Medical knowledge base
  • Vector search
  • Internal company APIs

Toolformer showed that language models can be trained to decide which APIs to call, when to call them, what arguments to pass, and how to incorporate tool results.

Function-calling competencies

A good tool-using model must learn:

  1. Whether a tool is needed.
  2. Which tool to use.
  3. How to fill arguments.
  4. How to obey schema.
  5. How to handle missing information.
  6. How to handle tool errors.
  7. How to combine multiple tool results.
  8. How to avoid calling tools unnecessarily.
  9. How to explain results to the user.
  10. How not to fabricate tool results.

Training data patterns

No-tool examples

Important. Otherwise the model overuses tools.

{
  "user": "Write a haiku about winter.",
  "assistant": "No tool call; answer directly."
}

Without explicit no-tool examples, models learn to call tools for everything — "let me search for haiku conventions" — which kills latency and quality.

Single-tool examples

{
  "user": "What is 17.5% of 240?",
  "assistant_tool_call": {
    "name": "calculator",
    "arguments": { "expression": "0.175 * 240" }
  },
  "tool_result": "42",
  "assistant": "17.5% of 240 is 42."
}

Multi-tool examples

Example: search web → open page → extract facts → calculate → summarize.

These trajectories teach planning across tools and information flow between them.

Error examples

The model should learn robust recovery:

  • API timeout → retry or apologize and stop
  • Invalid argument → fix the argument and retry
  • Empty result → broaden the query
  • Permission denied → explain to user, don't keep trying
  • Ambiguous query → ask the user a clarifying question
  • Conflicting sources → present the disagreement, don't pick one silently

Schema-constrained output

OpenAI's Structured Outputs and similar features (Pydantic-validated outputs in many SDKs, JSON-schema-constrained sampling in vLLM/Outlines) enforce JSON Schema adherence for model outputs, including function calling with strict schema matching. See Structured output.

Tool-use failure modes

  • Hallucinated tool calls — calling a tool that wasn't defined.
  • Invalid JSON — unparseable arguments.
  • Wrong function — calling search when calculator was the answer.
  • Missing required arguments — calling with partial input.
  • Calling tools when not needed — over-tooling.
  • Not calling tools when needed — under-tooling.
  • Ignoring tool results — generating an answer that contradicts the tool output.
  • Fabricating tool results — generating fake tool outputs after the call.
  • Security issues through prompt injection — see Red teaming.
  • Leaking hidden tool instructions — exposing system-prompt tool definitions.
  • Performing irreversible actions without confirmation — sending emails, deleting records.

Agentic patterns

Beyond single-turn function calling, modern agents need:

  • Planning — break a user goal into a tool-call sequence.
  • Reflection — recognize when a path isn't working and try another.
  • Memory — carry state across many tool calls.
  • Confirmation — pause before destructive actions.
  • Cost awareness — don't loop expensive tool calls.

The data for this is hard to scale. Common sources:

  • ReAct-style traces (Yao et al., 2022) — interleaved reasoning + actions.
  • SWE-bench-style task trajectories — long agent rollouts on real software-engineering tasks.
  • WebArena / VisualWebArena — browser-agent tasks.
  • τ-bench / BFCL — function-calling benchmarks.

Practical tips

  • Treat tool definitions like a schema, not prose. A consistent JSON-schema representation in training data (and at inference) drastically reduces malformed calls.
  • Include diverse tool sets in training. A model trained on 5 tools generalizes poorly to 50. Train on hundreds of synthetic tools so the model learns to read tool definitions, not memorize them.
  • Train recovery, not just success. Models that have only seen successful trajectories collapse on the first error. Include errored trajectories with successful recovery.
  • Test prompt injection. Include retrieved-content with injection attempts in your eval set; see Red teaming.
  • Evaluate end-to-end task success, not just call validity. "JSON was valid" is a necessary but insufficient signal.

Further reading