API / stream-structured-outputs-from-grok-api

API

Stream structured outputs from the Grok API

Stream a JSON object that still matches your schema so UIs can show partial fields while Grok fills an invoice, summary, or entity list — without waiting for the full parse round-trip. Official Structured Outputs documents passing a Pydantic model as response_format on chat.create, then calling stream() so chunks build up the JSON string in response.content, which you validate with Model.model_validate_json after the stream ends. For one-shot typed parse without streaming, use Return structured JSON from the Grok API. Create a key at console.x.ai.

What you need

An XAI_API_KEY, the xAI Python SDK (or an equivalent client that streams under a JSON schema), and a Pydantic (or Zod) schema for the object you want. Neighboring jobs include Combine structured outputs with tools on the Grok API when web search or functions must run first, and Return structured JSON from the Grok API when chat.parse is enough. More API jobs live on the API hub.

Stream with response_format

  1. Export the key:
export XAI_API_KEY="your_api_key"
  1. Create the chat with your schema as response_format, append messages, then stream. Chunks are partial JSON; parse only after the stream completes:
import os

from pydantic import BaseModel, Field
from xai_sdk import Client
from xai_sdk.chat import system, user


class Summary(BaseModel):
    title: str = Field(description="A brief title")
    key_points: list[str] = Field(description="Main points from the text")
    sentiment: str = Field(description="Overall sentiment: positive, negative, or neutral")


client = Client(api_key=os.getenv("XAI_API_KEY"))

chat = client.chat.create(
    model="grok-4.7",
    response_format=Summary,
)

chat.append(system("Analyze the following text and provide a structured summary."))
chat.append(user("The new product launch exceeded expectations with record sales..."))

for response, chunk in chat.stream():
    print(chunk.content, end="", flush=True)

summary = Summary.model_validate_json(response.content)
print(f"\nTitle: {summary.title}")
print(f"Sentiment: {summary.sentiment}")
  1. Prefer chat.parse(Model) when you do not need progressive UI updates — it returns (Response, Model) already parsed. Use response_format + sample() the same way as stream when you want the raw JSON string without streaming.

  2. Keep schemas inside the supported JSON Schema subset (Draft 2020-12 works best). additionalProperties defaults to false. Rejected shapes (empty enum, boolean property schemas, and similar) return 400 before any tokens stream.

When streaming helps

Progressive dashboards, long extraction jobs where operators want to see keys appear, and pipelines that tee the JSON string to logs before validation. Do not parse incomplete chunks as final objects — wait for the finished response.content. Structured outputs with tools on supported Grok 4 family models still apply after tool loops; streaming the typed final answer follows the same response_format pattern once tools are done.

Pitfalls

Calling model_validate_json on every partial chunk raises validation errors on truncated JSON. Assuming every SDK exposes streaming structured parse the same way mixes OpenAI beta.chat.completions.parse, Responses text.format, and xAI response_format. Omitting required fields in the schema while your prompt asks for them yields incomplete objects. Logging full prompts that contain PII into public tickets creates a compliance incident separate from the stream itself.