
Set max_turns on agentic Responses
Limit how many assistant / server-side tool-call turns an agentic request may run by passing max_turns, so quick lookups stay cheap and deep research only spends the budget you chose. Official Tool Usage Details documents the parameter under Limiting Tool Call Turns, and Advanced Usage explains how client-side tool pauses reset the counter on the next request. Authenticate with a Bearer inference key from console.x.ai. Pair max_turns with server-side tools such as web search or X search on the same request.
What you need
An XAI_API_KEY, the xAI SDK (or an equivalent Responses client that accepts max_turns), and at least one server-side tool in the request so the agentic loop has something to iterate. Neighboring jobs include Mix client-side and server-side tools on Grok when local functions share the loop, Add web search to an API request when you only need browse, and Call a function with Grok API when the turn budget must leave room for your own tools. More API jobs live on the API hub.
Cap the agentic loop
- Export the inference key and import the SDK helpers:
export XAI_API_KEY="your_api_key"
- Create a chat (or Responses) request with tools and an explicit
max_turns. Official guidance: a turn is one assistant iteration that may invoke several tools in parallel — the parameter does not mean “exactly N tool calls.”
import os
from xai_sdk import Client
from xai_sdk.chat import user
from xai_sdk.tools import web_search, x_search
client = Client(api_key=os.getenv("XAI_API_KEY"))
chat = client.chat.create(
model="grok-4.6",
tools=[
web_search(),
x_search(),
],
max_turns=3, # at most 3 assistant/tool-call turns
)
chat.append(user("What is the latest news from xAI?"))
response = chat.sample()
print(response.content)
Pick a budget that matches the job. Official examples frame roughly 1–2 turns for quick lookups, 3–5 for balanced research, and 10+ (or unset) for deep research that accepts more latency and cost. When the agent hits the limit, it stops making additional tool calls and answers from what it already gathered.
Remember the client-side checkpoint rule from Advanced Usage: when the model returns a client-side function call, that request ends and your follow-up starts with a fresh
max_turnscount. Re-pass the same ceiling on each follow-up if you still want the cap.
Pitfalls
Treating max_turns=3 as “three tool calls” undercounts parallel tool use inside a single turn. Omitting the field does not mean unlimited looping forever — the server still applies a global default cap. Expecting the same remaining budget after you handle a client-side tool and call again fails; that follow-up resets the counter unless you set max_turns again.