API / continue-agentic-conversation-with-store-messages

API

Continue an agentic conversation with store_messages

Keep web search, X search, and other server-side tool history alive across follow-up prompts by storing the first agentic turn on xAI and resuming with previous_response_id. Official Advanced Usage documents two steps: set store_messages=True on the first chat.create, then pass previous_response_id=response.id when you open the next chat so reasoning, tool calls, and tool responses hydrate into the follow-up. The second turn does not need the same tools or model parameters as the first. Create a key and load credits at console.x.ai.

What you need

An XAI_API_KEY, at least one server-side tool (web_search, x_search, or similar), and a product flow that asks a follow-up after the first agentic answer. This path stores conversation state on xAI servers — teams on Zero Data Retention should use Use encrypted content for agentic multi-turn instead. Neighboring jobs include Mix client-side and server-side tools on Grok when local functions join the loop, and Include an image in an agentic tool request when the first turn also carries a picture. More API jobs live on the API hub.

Store the first turn, resume the second

  1. Export the key, create a chat with tools and store_messages=True, then stream the first user question:
import os

from xai_sdk import Client
from xai_sdk.chat import user
from xai_sdk.tools import web_search, x_search

client = Client(api_key=os.getenv("XAI_API_KEY"))

chat = client.chat.create(
    model="grok-4.7",
    tools=[web_search(), x_search()],
    store_messages=True,
)
chat.append(user("What is xAI?"))

print("##### First turn #####")
for response, chunk in chat.stream():
    print(chunk.content, end="", flush=True)

print("\nUsage for first turn:", response.server_side_tool_usage)
first_id = response.id
  1. Open a new chat that points at the stored response id, append the follow-up user message, and stream again — the server rehydrates prior tool state even if you change the tool list:
chat = client.chat.create(
    model="grok-4.7",
    tools=[web_search(), x_search()],
    previous_response_id=first_id,
)
chat.append(user("What is its latest mission?"))

print("\n##### Second turn #####")
for response, chunk in chat.stream():
    print(chunk.content, end="", flush=True)

print("\nUsage for second turn:", response.server_side_tool_usage)
  1. On the Responses API with the OpenAI SDK, the same idea is previous_response_id after a first call that stored state; keep using the id returned from the turn you want to continue. Log server_side_tool_usage per turn when you need cost attribution for search and code tools separately from token usage.

When to pick store_messages

Use remote storage when your backend can keep a response id between turns and ZDR is off. Prefer encrypted client-side content when retention policies forbid server-side conversation history. For long tool-heavy loops on one socket, also consider Use the Responses API over WebSocket, which can chain previous_response_id in memory without reopening HTTP each turn.

Pitfalls

Forgetting store_messages=True on the first turn leaves nothing for previous_response_id to resume. Re-sending the entire history manually while also passing previous_response_id double-counts context and can confuse tool state. Expecting this pattern under Zero Data Retention fails because ZDR blocks the stored Responses path — switch to encrypted content. Mixing Console API credits with SuperGrok weekly pools on grok.com confuses two different meters the docs keep separate.