API / chain-multi-turn-edits-with-image-generation-tool

API

Chain multi-turn edits with the image generation tool

Refine a still across conversation turns by keeping prior image_generation_call outputs in the thread so Grok can edit them without re-uploading bytes on every follow-up. Official Image Generation Tool documents multi-turn editing in plain steps: generate on turn one, then continue with previous_response_id on the Responses API or by appending the previous response object to the chat in the xAI SDK, and ask for the next change in natural language. Images produced on earlier turns stay editable inside that conversation, and edit calls produce image_generation_call items with an ie_ id prefix so you can tell refinement apart from a fresh ig_ generation. Create a key and load credits at console.x.ai before you burn multiple Imagine calls in one chat loop.

What you need

An XAI_API_KEY, Responses access with the image_generation tool enabled (the default action: "auto" is enough for generate-then-edit), and a conversation store that either keeps previous_response_id or passes prior output items — including every image_generation_call — back verbatim so the model still sees the image bytes. Neighboring jobs include Use the image generation tool in a conversation for the single-turn baseline, Configure image generation tool action when later turns must stay edit-only, and Edit an image with the Imagine API for a single-shot /v1/images/edits call that sits outside chat state. More API jobs live on the API hub.

Generate, then edit on the next turn

  1. Export the key:
export XAI_API_KEY="your_api_key"
  1. Create the first still with the tool on POST https://api.x.ai/v1/responses, then save the response id for the follow-up:
import base64
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.getenv("XAI_API_KEY"),
    base_url="https://api.x.ai/v1",
)

response = client.responses.create(
    model="grok-4.7",
    input="Generate an image of a lighthouse on a rocky coast",
    tools=[{"type": "image_generation"}],
)

image_data = [
    output.result
    for output in response.output
    if output.type == "image_generation_call"
]
if image_data:
    with open("lighthouse.jpg", "wb") as f:
        f.write(base64.b64decode(image_data[0]))
  1. Continue the same conversation with previous_response_id and a refinement prompt so the model edits the prior still instead of starting from a blank canvas:
followup = client.responses.create(
    model="grok-4.7",
    previous_response_id=response.id,
    input="Make it night time with a full moon",
    tools=[{"type": "image_generation"}],
)

image_data_followup = [
    output.result
    for output in followup.output
    if output.type == "image_generation_call"
]
if image_data_followup:
    with open("lighthouse_night.jpg", "wb") as f:
        f.write(base64.b64decode(image_data_followup[0]))
  1. With the xAI SDK, append the first response object to the chat before the next user message so the decoded image remains available for editing:
import os
from xai_sdk import Client
from xai_sdk.chat import user
from xai_sdk.tools import image_generation

client = Client(api_key=os.getenv("XAI_API_KEY"))
chat = client.chat.create(model="grok-4.7", tools=[image_generation()])

chat.append(user("Generate an image of a lighthouse on a rocky coast"))
response = chat.sample()
with open("image.jpeg", "wb") as f:
    f.write(response.image_outputs[0].image)

chat.append(response)
chat.append(user("Make it night time with a full moon"))
followup = chat.sample()
with open("edited_image.jpeg", "wb") as f:
    f.write(followup.image_outputs[0].image)

When you manage state yourself

If your app does not use previous_response_id, pass the previous turn’s output items — including every image_generation_call — back in input unchanged so the server still has the earlier still available for editing. Dropping those items removes the editable image from context, which usually makes the model refuse the edit or quietly generate a brand-new still that ignores the thumbnail your UI is still showing. Pin action: "edit" on later turns when product rules forbid brand-new text-to-image calls during a refinement loop, and keep action: "auto" only when a turn is allowed to either generate or edit depending on the user wording.

Pitfalls

Starting a brand-new Responses request without previous_response_id (and without replaying prior output items) breaks the edit chain even when your UI still shows the old thumbnail from local storage. Asking for a ratio or style change in the follow-up text still works because the tool has no separate size fields, but stacking conflicting instructions across many turns can produce larger visual jumps than a single carefully written edit prompt would. When you need pixel-level aspect_ratio, resolution, or quality control, leave the conversational tool and call the Imagine image endpoints directly instead of expecting tool_choice or multi-turn chat state to expose those knobs.