
Set max_output_tokens on Responses API
Cap how many tokens a Responses call may generate by passing max_output_tokens on POST https://api.x.ai/v1/responses, so short summaries, UI snippets, and cost-bounded jobs stop before the model writes a novel. Official Comparison with Chat Completions maps the legacy Chat Completions field max_tokens to Responses max_output_tokens and marks Responses as the recommended inference surface. Authenticate with a Bearer inference key from console.x.ai. The same body still takes model, input, and the other Responses fields you already use for chat and tools.
What you need
An XAI_API_KEY with Responses access, a terminal with curl (or the OpenAI / xAI SDK pointed at https://api.x.ai/v1), and a clear target length for the job so the cap matches the product, not a guess. Neighboring jobs include Migrate Chat Completions calls to Responses when you are renaming the whole request shape, Disable store on Responses when you also need to turn off server-side retention, and Chain Responses with previous_response_id when a capped turn continues a stored conversation. More API jobs live on the API hub.
Cap the output
- Export the inference key outside of source control:
export XAI_API_KEY="your_api_key"
- Call Responses with
max_output_tokensset to the ceiling you want for this turn. Use the Responses field name — notmax_tokens— or a Chat Completions-shaped body will miss the intended limit on/v1/responses:
curl https://api.x.ai/v1/responses \
-H "Authorization: Bearer $XAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "grok-4.6",
"input": "Summarize the Riemann hypothesis in two sentences.",
"max_output_tokens": 128
}'
- Prefer the OpenAI-compatible SDK when your app already speaks Responses. The parameter name stays
max_output_tokensonclient.responses.create:
import os
from openai import OpenAI
client = OpenAI(
api_key=os.getenv("XAI_API_KEY"),
base_url="https://api.x.ai/v1",
)
response = client.responses.create(
model="grok-4.6",
input="Summarize the Riemann hypothesis in two sentences.",
max_output_tokens=128,
)
print(response.output_text)
- Treat the value as a hard generation ceiling for that request. The model may finish earlier when it hits a natural stop; if the response ends because of length, raise the cap or tighten the prompt instead of assuming the full answer landed under a tiny budget.
Pitfalls
Sending max_tokens on /v1/responses after a Chat Completions migration leaves the Responses-native ceiling unset. Setting an extremely low cap on reasoning-heavy prompts can truncate mid-thought relative to what you expected from an uncapped run. Forgetting that each follow-up with previous_response_id is its own request means you must pass max_output_tokens again on every turn that needs the same ceiling.