API / include-an-image-in-an-agentic-tool-request

API

Include an image in an agentic tool request

Bootstrap an agentic chat with a picture so Grok can look at the image and still call server-side tools such as web search or X search in the same turn. Official Advanced Usage shows appending a user message that includes both text and an image(...) URL before you stream a tool-enabled chat — the model can identify what is in the frame and then search for related facts. Create a key and load credits at console.x.ai.

What you need

An XAI_API_KEY, a publicly reachable image URL (or another image form your SDK accepts on image(...)), and at least one server-side tool on the chat. Neighboring jobs include Continue an agentic conversation with store_messages when you want a follow-up after the first research turn, Use encrypted content for agentic multi-turn under ZDR, and Mix client-side and server-side tools on Grok when a local function should join the loop. More API jobs live on the API hub.

Attach an image, then stream with tools

  1. Export the key and create a chat with the tools you want active:
import os

from xai_sdk import Client
from xai_sdk.chat import image, user
from xai_sdk.tools import web_search, x_search

client = Client(api_key=os.getenv("XAI_API_KEY"))

chat = client.chat.create(
    model="grok-4.7",
    tools=[web_search(), x_search()],
    include=["verbose_streaming"],
)
  1. Append a user message that carries both instructions and the image URL, then stream so you can watch server-side tool calls as they fire:
chat.append(
    user(
        "Search the internet and tell me what kind of dog is in the image below.",
        "And what is the typical lifespan of this dog breed?",
        image(
            "https://pbs.twimg.com/media/G3B7SweXsAAgv5N?format=jpg&name=900x900"
        ),
    )
)

is_thinking = True
for response, chunk in chat.stream():
    for tool_call in chunk.tool_calls:
        print(
            f"\nCalling tool: {tool_call.function.name} "
            f"with arguments: {tool_call.function.arguments}"
        )
    if response.usage.reasoning_tokens and is_thinking:
        print(
            f"\rThinking... ({response.usage.reasoning_tokens} tokens)",
            end="",
            flush=True,
        )
    if chunk.content and is_thinking:
        print("\n\nFinal Response:")
        is_thinking = False
    if chunk.content and not is_thinking:
        print(chunk.content, end="", flush=True)

print("\n\nCitations:")
print(response.citations)
print("\nUsage:")
print(response.usage)
print(response.server_side_tool_usage)
  1. For a multi-turn research follow-up, either set store_messages=True on create and resume with previous_response_id, or set use_encrypted_content=True and chat.append(response) before the next user line. Keep the image only on the turns that need visual context — later turns inherit agentic state from the continuation path you chose.

  2. Prefer a stable HTTPS URL the API can fetch. When the binary must stay private, upload through the Files API first and use the attachment patterns from Attach files to a Grok chat request instead of a public CDN link.

After the first answer

Read response.citations and server_side_tool_usage to see which search calls grounded the breed facts. If the model misread the image, send a corrective user turn with a tighter crop or a second image rather than restarting without tools. Pair with Send a safety_identifier on Grok API requests when many end users share one key on an image-upload product.

Pitfalls

Passing a URL that requires cookies or auth returns a fetch failure and the agent proceeds without visual context. Enabling tools without an image-capable model wastes the attachment. Expecting the image alone to trigger search without an instruction that asks for web research leaves the model describing the frame only. Mixing Console API credits with SuperGrok weekly pools on grok.com confuses two different meters the docs keep separate.