
Disable parallel tool calls
Make Grok request at most one function call per Responses turn when your backend must run tools serially — for example a booking flow that needs step A before step B, or a worker pool that cannot safely execute two side effects at once. Official Function Calling documents Parallel Function Calling: by default the model may return multiple tool calls in a single response, and you disable that behavior by setting parallel_tool_calls to false on the request. Create a key and load credits at console.x.ai.
What you need
An XAI_API_KEY, at least two custom function tools (or one tool the model might call twice with different args), and a product rule that requires ordered execution. Neighboring jobs include Force a specific tool with tool_choice when you also need a named tool on this turn, Call a function with the Grok API for the define-and-return loop, and Mix client-side and server-side tools on Grok when web search sits beside local functions. More API jobs live on the API hub.
Set parallel_tool_calls false
- Export the key outside of source control:
export XAI_API_KEY="your_api_key"
- Call
/v1/responseswith yourtoolsarray and"parallel_tool_calls": false. Keep tool parameter schemas rooted at an object (or aoneOf/anyOfof objects):
curl https://api.x.ai/v1/responses \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $XAI_API_KEY" \
-d '{
"model": "grok-4.7",
"input": [
{"role": "user", "content": "Get the temperature and cloud ceiling in San Francisco."}
],
"tools": [
{
"type": "function",
"name": "get_temperature",
"description": "Get current temperature for a location",
"parameters": {
"type": "object",
"properties": {
"location": {"type": "string", "description": "City name"}
},
"required": ["location"]
}
},
{
"type": "function",
"name": "get_ceiling",
"description": "Get current cloud ceiling for a location",
"parameters": {
"type": "object",
"properties": {
"location": {"type": "string", "description": "City name"}
},
"required": ["location"]
}
}
],
"parallel_tool_calls": false
}'
- In the OpenAI-compatible SDK pointed at
https://api.x.ai/v1, pass the same flag onresponses.create, execute the singlefunction_callyou receive, returnfunction_call_output, then continue the conversation (often viaprevious_response_id) so the model can request the next tool on a later turn:
import json
import os
from openai import OpenAI
client = OpenAI(api_key=os.getenv("XAI_API_KEY"), base_url="https://api.x.ai/v1")
tools = [
{
"type": "function",
"name": "get_temperature",
"description": "Get current temperature for a location",
"parameters": {
"type": "object",
"properties": {
"location": {"type": "string", "description": "City name"}
},
"required": ["location"],
},
},
{
"type": "function",
"name": "get_ceiling",
"description": "Get current cloud ceiling for a location",
"parameters": {
"type": "object",
"properties": {
"location": {"type": "string", "description": "City name"}
},
"required": ["location"],
},
},
]
response = client.responses.create(
model="grok-4.7",
input=[
{
"role": "user",
"content": "Get the temperature and cloud ceiling in San Francisco.",
}
],
tools=tools,
parallel_tool_calls=False,
)
for item in response.output:
if item.type == "function_call":
args = json.loads(item.arguments)
print(item.name, args)
# Run one tool, send function_call_output, then sample again for the next call.
When to leave parallel on
Keep the default (true / omitted) when independent lookups can run together — temperature plus ceiling, or two database reads with no ordering constraint. Official docs show iterating every entry in response.tool_calls (or every function_call output item) and appending each result before you ask the model to continue. Parallelism pairs cleanly with "tool_choice": "required" or "auto"; it does not replace forcing a named tool when only one function is allowed.
Pitfalls
Leaving parallel enabled while your handler only processes the first function_call drops sibling results and starves the next turn of context. Setting parallel_tool_calls: false does not skip tool execution — you still run the one call and return output. Built-in server-side tools (web search and similar) still execute on xAI’s side when selected; this flag mainly shapes how many client-side function calls arrive together. Parameter schemas with a scalar root still fail with 400 regardless of parallelism.