
Configure image generation tool action
Pin the server-side image_generation tool to generate-only, edit-only, or both so Grok cannot drift into the wrong mode inside a Responses conversation. Official Image Generation Tool documents the optional action field on the tool object with values auto (default: generate and edit), generate (text-to-image only), and edit (image editing only). The tool still uses grok-imagine-image-2.0 under the hood and still picks aspect ratio from the user text; action only gates which capabilities the model may invoke. Create a key and load credits at console.x.ai before you bill a tool-enabled turn.
What you need
An XAI_API_KEY with Responses + Imagine access, a chat model that supports tools (examples use grok-4.7), and a clear reason to restrict the tool — for example locking a support bot to fresh stills with generate, or forcing refinement of an attached photo with edit. Neighboring jobs include Use the image generation tool in a conversation, Chain multi-turn edits with the image generation tool, and Edit an image with the Imagine API when you want direct /v1/images/edits control instead of the conversational tool. More API jobs live on the API hub.
Set action on the tool object
- Export the key outside of source control:
export XAI_API_KEY="your_api_key"
- Restrict the tool to text-to-image only by setting
"action": "generate"inside theimage_generationtool entry (omitactionor pass"auto"when both generate and edit should stay available):
curl https://api.x.ai/v1/responses \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $XAI_API_KEY" \
-d '{
"model": "grok-4.7",
"input": "Generate an image of a hot air balloon over the desert",
"tools": [
{
"type": "image_generation",
"action": "generate"
}
]
}'
- Or pin edit-only when the user message already carries an
input_image(or a priorimage_generation_callin the thread) and you do not want new text-to-image calls:
curl https://api.x.ai/v1/responses \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $XAI_API_KEY" \
-d '{
"model": "grok-4.7",
"input": [
{
"role": "user",
"content": [
{
"type": "input_text",
"text": "Edit this image so it looks like a watercolor painting."
},
{
"type": "input_image",
"image_url": "https://docs.x.ai/assets/api-examples/images/style-realistic.png"
}
]
}
],
"tools": [
{
"type": "image_generation",
"action": "edit"
}
]
}'
- In the xAI Python SDK, pass the same gate into
image_generation(action="generate")orimage_generation(action="edit")when you create the chat:
import os
from xai_sdk import Client
from xai_sdk.chat import user
from xai_sdk.tools import image_generation
client = Client(api_key=os.getenv("XAI_API_KEY"))
chat = client.chat.create(
model="grok-4.7",
tools=[image_generation(action="generate")],
)
chat.append(user("Generate an image of a hot air balloon over the desert"))
response = chat.sample()
with open("balloon.jpeg", "wb") as f:
f.write(response.image_outputs[0].image)
How to read the result
Completed tool calls land as image_generation_call output items, where generation ids are prefixed ig_ and edit ids are prefixed ie_, which helps you confirm the model stayed inside the action you set. The prompt field on that item shows the Imagine prompt the chat model wrote, and the result field on the Responses API carries base64 image bytes with no data-URL prefix. Because the tool does not accept size or format parameters, ask for a ratio in the user text whenever framing matters for the placement you will ship.
Pitfalls
Leaving action unset keeps auto, so a follow-up that says “make it darker” may edit a prior still even if your product copy assumed generate-only, and setting action: "edit" without any image in the conversation (attached input or prior tool output) leaves the model with nothing legal to call. When you need pixel-level control over aspect_ratio, resolution, and quality outside the conversational tool, call the direct Imagine image endpoints instead of relying on action alone.