API / combine-web-search-and-image-generation-tool

API

Combine web search and the image generation tool

Ask Grok to look up a fact with web_search, then turn what it found into a still with the server-side image_generation tool in the same Responses request so the poster matches the live answer. Official Image Generation Tool documents this combination under Combining with other tools: put both tool objects in tools, send a prompt that needs a lookup-then-draw workflow, and let the agentic loop interleave a web_search_call, a cited message, and an image_generation_call whose result carries base64 image bytes. Advanced Usage covers the same multi-tool pattern for other stacks, and you should create a key and load credits at console.x.ai before you bill both tools on one turn.

What you need

An XAI_API_KEY with Responses access, a chat model that supports server-side tools (examples use grok-4.7), and a user prompt that truly needs a live fact before the image — for example “find the latest FIFA World Cup winner, then generate a vintage travel-poster celebration for that team.” Neighboring jobs include Use the image generation tool in a conversation for a single-tool baseline, Add web search to an API request when you only need citations, and Configure image generation tool action when later turns must stay generate-only or edit-only. Browse more API jobs on the API hub when you are wiring related tool calls.

Put both tools on one request

  1. Export the key outside of source control so the value never lands in git history or chat logs:
export XAI_API_KEY="your_api_key"
  1. Call /v1/responses with both tools listed and a prompt that sequences research before drawing, then decode the image_generation_call result from base64 when the request finishes:
curl https://api.x.ai/v1/responses \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $XAI_API_KEY" \
  -d '{
  "model": "grok-4.7",
  "input": "Find out which team won the most recent FIFA World Cup, then generate an image of a celebratory poster for that team, in a vintage travel-poster style.",
  "tools": [
    { "type": "web_search" },
    { "type": "image_generation" }
  ]
}' | jq -r '.output[] | select(.type == "image_generation_call") | .result' \
  | base64 --decode > champions_poster.jpg
  1. In the xAI Python SDK, pass web_search() and image_generation() together, then read response.image_outputs and response.server_side_tool_usage so you can confirm both tools ran on the turn:
import os
from xai_sdk import Client
from xai_sdk.chat import user
from xai_sdk.tools import image_generation, web_search

client = Client(api_key=os.getenv("XAI_API_KEY"))
chat = client.chat.create(
    model="grok-4.7",
    tools=[web_search(), image_generation()],
)
chat.append(
    user(
        "Find out which team won the most recent FIFA World Cup, then generate an "
        "image of a celebratory poster for that team, in a vintage travel-poster style."
    )
)
response = chat.sample()
print(response.content)
with open("image.jpeg", "wb") as f:
    f.write(response.image_outputs[0].image)
print(response.server_side_tool_usage)

How to read the interleaved output

Responses API output items arrive in the order the agent ran them: a web_search_call with the query details, a message that answers the factual question (often with citations you can surface), and an image_generation_call whose prompt field shows the Imagine prompt the chat model wrote from those facts. Generation ids use an ig_ prefix, and if a later turn edits the same still, edit ids use ie_. The same pairing pattern works with x_search, code_interpreter / code_execution, and your own client-side functions when you need social signal or local side effects after the research step.

Pitfalls

A prompt that never asks for a lookup can skip web search and invent a subject for the poster, so write the user text as an explicit find-then-draw instruction when the image must match a current fact. Billing includes each server-side tool invocation, so watch server_side_tool_usage when you ship this pattern inside a loop. When you need fixed aspect_ratio, resolution, or quality knobs instead of conversational framing, call the direct Imagine image endpoints after you finish the search turn rather than relying on the conversational tool alone.