
Compact a chat in place with the xAI SDK
Call chat.compact() on a live xAI SDK Chat so a long agent loop replaces its accumulated messages with one compaction item, then keep calling chat.sample() without rebuilding the conversation by hand. Official Context Compaction documents in-place compaction for long-running agent loops: the method runs compaction against the chat’s current messages and swaps them for the compaction blob; the server rehydrates that prefix on the next request. Create a key and load credits at console.x.ai.
What you need
Python with the xai-sdk package, an XAI_API_KEY, and a multi-turn loop that already uses client.chat.create plus chat.append / chat.sample(). Neighboring jobs include Use context compaction for the REST POST /v1/responses/compact path and compact_context helper, Avoid breaking prompt cache when you also want prefix reuse, and Chain responses with previous_response_id when server-stored turns are enough without a local chat object. More API jobs live on the API hub.
Compact inside the agent loop
- Create a chat with encrypted content when you use a reasoning model, so prior reasoning survives compaction:
import os
from xai_sdk import Client
from xai_sdk.chat import system, user
client = Client(api_key=os.environ["XAI_API_KEY"])
chat = client.chat.create(model="grok-4.7", use_encrypted_content=True)
chat.append(system("You are a helpful assistant. Keep answers brief."))
- Run turns as usual — append the user message, sample, append the response — and every N turns call
chat.compact():
compact_every = 5
for turn in range(1, 100):
chat.append(user(input("You: ")))
response = chat.sample()
print(f"Grok: {response.content}")
chat.append(response)
if turn % compact_every == 0:
before = len(chat.messages)
compact = chat.compact()
print(
f"[compacted {before} → {len(chat.messages)} messages | "
f"dropped {compact.dropped_message_count} | "
f"tokens used: {compact.usage.total_tokens}]"
)
After
chat.compact()returns, keep sampling. Do not parse or editencrypted_content. Treat the compaction item as the new head of the conversation and only append new user turns after it. OnAsyncClient, useawait chat.compact()the same way.Compact only while the conversation still fits the model context window — compaction shrinks history; it does not rescue a request that already hit
context_length_exceeded. Re-compacting later is fine when the chat grows again after a prior compact.
Pitfalls
Calling compact on an empty or tiny chat wastes a compaction request without lowering later input cost. Editing the returned blob or reordering items before the next sample breaks the chain. Forgetting use_encrypted_content=True on reasoning models drops reasoning content that the docs recommend preserving across compacted turns. Expecting chat.compact() to exist on the OpenAI-compatible client is wrong — that path uses client.responses.compact and spreads compacted.output into the next input instead.