API / set-max-output-tokens-on-responses

API

Set max_output_tokens on Responses API

Cap how many tokens a Responses call may generate by passing max_output_tokens on POST https://api.x.ai/v1/responses, so short summaries, UI snippets, and cost-bounded jobs stop before the model writes a novel. Official Comparison with Chat Completions maps the legacy Chat Completions field max_tokens to Responses max_output_tokens and marks Responses as the recommended inference surface. Authenticate with a Bearer inference key from console.x.ai. The same body still takes model, input, and the other Responses fields you already use for chat and tools.

What you need

An XAI_API_KEY with Responses access, a terminal with curl (or the OpenAI / xAI SDK pointed at https://api.x.ai/v1), and a clear target length for the job so the cap matches the product, not a guess. Neighboring jobs include Migrate Chat Completions calls to Responses when you are renaming the whole request shape, Disable store on Responses when you also need to turn off server-side retention, and Chain Responses with previous_response_id when a capped turn continues a stored conversation. More API jobs live on the API hub.

Cap the output

  1. Export the inference key outside of source control:
export XAI_API_KEY="your_api_key"
  1. Call Responses with max_output_tokens set to the ceiling you want for this turn. Use the Responses field name — not max_tokens — or a Chat Completions-shaped body will miss the intended limit on /v1/responses:
curl https://api.x.ai/v1/responses \
  -H "Authorization: Bearer $XAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "grok-4.6",
    "input": "Summarize the Riemann hypothesis in two sentences.",
    "max_output_tokens": 128
  }'
  1. Prefer the OpenAI-compatible SDK when your app already speaks Responses. The parameter name stays max_output_tokens on client.responses.create:
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.getenv("XAI_API_KEY"),
    base_url="https://api.x.ai/v1",
)

response = client.responses.create(
    model="grok-4.6",
    input="Summarize the Riemann hypothesis in two sentences.",
    max_output_tokens=128,
)
print(response.output_text)
  1. Treat the value as a hard generation ceiling for that request. The model may finish earlier when it hits a natural stop; if the response ends because of length, raise the cap or tighten the prompt instead of assuming the full answer landed under a tiny budget.

Pitfalls

Sending max_tokens on /v1/responses after a Chat Completions migration leaves the Responses-native ceiling unset. Setting an extremely low cap on reasoning-heavy prompts can truncate mid-thought relative to what you expected from an uncapped run. Forgetting that each follow-up with previous_response_id is its own request means you must pass max_output_tokens again on every turn that needs the same ceiling.