API / stream-image-generation-tool-progress

API

Stream image generation tool progress

Watch an image_generation tool call move from in_progress through generating to completed on a streamed Responses request so your UI can show status before the base64 still arrives. Official Image Generation Tool Streaming section documents the event sequence: each image generation call emits progress events (in_progress, then generating, then completed), followed by a response.output_item.done event whose item carries the base64 result, and partial image previews are not emitted. Create a key and load credits at console.x.ai before you stream tool-enabled turns.

What you need

An XAI_API_KEY with Responses access, a client that supports stream=True (or the xAI SDK chat stream with include=["verbose_streaming"]), and a front end or log sink that can render status text without expecting progressive JPEG frames from the tool. You also need enough timeout headroom for an agentic turn that may think before it starts Imagine. Neighboring jobs include Use the image generation tool in a conversation for the non-streaming baseline, Chain multi-turn edits with the image generation tool when follow-up edits reuse prior stills, and Combine web search and the image generation tool when the same stream may also surface web_search_call items. Browse more API jobs on the API hub when you are wiring related tool calls.

Stream status, then save the still

  1. Export the key outside of source control before you open a long-lived stream:
export XAI_API_KEY="your_api_key"
  1. With the OpenAI-compatible Responses client pointed at https://api.x.ai/v1, enable streaming and branch on image-generation event types, writing the file only when response.output_item.done carries an image_generation_call item:
import base64
import os
from openai import OpenAI

client = OpenAI(api_key=os.getenv("XAI_API_KEY"), base_url="https://api.x.ai/v1")

stream = client.responses.create(
    model="grok-4.7",
    input="Generate an image of an origami fox in a paper forest",
    tools=[{"type": "image_generation"}],
    stream=True,
)

for event in stream:
    if event.type.startswith("response.image_generation_call."):
        # in_progress -> generating -> completed
        print(f"Image generation status: {event.type.rsplit('.', 1)[-1]}")
    elif event.type == "response.output_item.done" and event.item.type == "image_generation_call":
        with open("origami_fox.jpg", "wb") as f:
            f.write(base64.b64decode(event.item.result))
    elif event.type == "response.output_text.delta":
        print(event.delta, end="", flush=True)
  1. In the xAI Python SDK, pass include=["verbose_streaming"], watch chunk.tool_calls for image_generation_tool via get_tool_call_type, and read decoded bytes from the accumulated response.image_outputs after the stream ends:
import os
from xai_sdk import Client
from xai_sdk.chat import user
from xai_sdk.tools import get_tool_call_type, image_generation

client = Client(api_key=os.getenv("XAI_API_KEY"))
chat = client.chat.create(
    model="grok-4.7",
    tools=[image_generation()],
    include=["verbose_streaming"],
)
chat.append(user("Generate an image of an origami fox in a paper forest"))

for response, chunk in chat.stream():
    for tool_call in chunk.tool_calls:
        if get_tool_call_type(tool_call) == "image_generation_tool":
            print(f"Generating image: {tool_call.function.arguments}")
    if chunk.content:
        print(chunk.content, end="", flush=True)

with open("image.jpeg", "wb") as f:
    f.write(response.image_outputs[0].image)

What streaming does not give you

The docs are explicit that partial image previews are not emitted during generation, so design the UI around status labels and the final decode rather than progressive rendering of half-drawn frames. Text deltas for the assistant message can still arrive alongside tool progress, and you should keep those streams separate in your renderer so a status toast does not swallow the prose the model streams after the still. Edit calls that carry ie_ id prefixes follow the same progress event family when the model refines a prior still inside a streamed multi-turn chat, which means your status component can stay shared across generate and edit turns.

Pitfalls

Closing the stream early after the first generating event leaves you without result bytes on the final output item done event, so keep the iterator open until the response finishes. Mixing this conversational tool stream with a direct /v1/images/generations call in the same handler can confuse timeout budgets, because the agentic loop may run other tools before Imagine starts. When you also enable web search in the same streamed request, expect interleaved tool events and only decode items whose type is image_generation_call. Log the progress event names your client actually receives in staging once, because SDK wrappers sometimes rename the trailing status segment even when the underlying Responses stream matches the docs.