API / disable-parallel-tool-calls

API

Disable parallel tool calls

Make Grok request at most one function call per Responses turn when your backend must run tools serially — for example a booking flow that needs step A before step B, or a worker pool that cannot safely execute two side effects at once. Official Function Calling documents Parallel Function Calling: by default the model may return multiple tool calls in a single response, and you disable that behavior by setting parallel_tool_calls to false on the request. Create a key and load credits at console.x.ai.

What you need

An XAI_API_KEY, at least two custom function tools (or one tool the model might call twice with different args), and a product rule that requires ordered execution. Neighboring jobs include Force a specific tool with tool_choice when you also need a named tool on this turn, Call a function with the Grok API for the define-and-return loop, and Mix client-side and server-side tools on Grok when web search sits beside local functions. More API jobs live on the API hub.

Set parallel_tool_calls false

  1. Export the key outside of source control:
export XAI_API_KEY="your_api_key"
  1. Call /v1/responses with your tools array and "parallel_tool_calls": false. Keep tool parameter schemas rooted at an object (or a oneOf/anyOf of objects):
curl https://api.x.ai/v1/responses \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $XAI_API_KEY" \
  -d '{
  "model": "grok-4.7",
  "input": [
    {"role": "user", "content": "Get the temperature and cloud ceiling in San Francisco."}
  ],
  "tools": [
    {
      "type": "function",
      "name": "get_temperature",
      "description": "Get current temperature for a location",
      "parameters": {
        "type": "object",
        "properties": {
          "location": {"type": "string", "description": "City name"}
        },
        "required": ["location"]
      }
    },
    {
      "type": "function",
      "name": "get_ceiling",
      "description": "Get current cloud ceiling for a location",
      "parameters": {
        "type": "object",
        "properties": {
          "location": {"type": "string", "description": "City name"}
        },
        "required": ["location"]
      }
    }
  ],
  "parallel_tool_calls": false
}'
  1. In the OpenAI-compatible SDK pointed at https://api.x.ai/v1, pass the same flag on responses.create, execute the single function_call you receive, return function_call_output, then continue the conversation (often via previous_response_id) so the model can request the next tool on a later turn:
import json
import os
from openai import OpenAI

client = OpenAI(api_key=os.getenv("XAI_API_KEY"), base_url="https://api.x.ai/v1")

tools = [
    {
        "type": "function",
        "name": "get_temperature",
        "description": "Get current temperature for a location",
        "parameters": {
            "type": "object",
            "properties": {
                "location": {"type": "string", "description": "City name"}
            },
            "required": ["location"],
        },
    },
    {
        "type": "function",
        "name": "get_ceiling",
        "description": "Get current cloud ceiling for a location",
        "parameters": {
            "type": "object",
            "properties": {
                "location": {"type": "string", "description": "City name"}
            },
            "required": ["location"],
        },
    },
]

response = client.responses.create(
    model="grok-4.7",
    input=[
        {
            "role": "user",
            "content": "Get the temperature and cloud ceiling in San Francisco.",
        }
    ],
    tools=tools,
    parallel_tool_calls=False,
)

for item in response.output:
    if item.type == "function_call":
        args = json.loads(item.arguments)
        print(item.name, args)
        # Run one tool, send function_call_output, then sample again for the next call.

When to leave parallel on

Keep the default (true / omitted) when independent lookups can run together — temperature plus ceiling, or two database reads with no ordering constraint. Official docs show iterating every entry in response.tool_calls (or every function_call output item) and appending each result before you ask the model to continue. Parallelism pairs cleanly with "tool_choice": "required" or "auto"; it does not replace forcing a named tool when only one function is allowed.

Pitfalls

Leaving parallel enabled while your handler only processes the first function_call drops sibling results and starves the next turn of context. Setting parallel_tool_calls: false does not skip tool execution — you still run the one call and return output. Built-in server-side tools (web search and similar) still execute on xAI’s side when selected; this flag mainly shapes how many client-side function calls arrive together. Parameter schemas with a scalar root still fail with 400 regardless of parallelism.