
Cancel a streaming TTS utterance with text.clear
Interrupt the current streaming TTS turn when the user barges in, without tearing down the WebSocket and paying another handshake on every correction. Official Text to Speech Cancellation (Barge-in) documents the client event {"type":"text.clear"} on wss://api.x.ai/v1/tts, after which the server stops synthesis, discards buffered audio for that utterance, and answers with {"type":"audio.clear"} so your client can flush local playback and start a fresh text.delta → text.done sequence on the same connection.
What you need
You need a backend that already opens streaming TTS as in Stream text to speech with the Voice API, plus a UI or voice-activity path that can detect user interrupt and send text.clear while audio is still arriving. Neighboring jobs include Set TTS speed via the xAI API, Enable text normalization in TTS, and Convert text to speech with the Voice API. More voice jobs live on the Voice hub.
Barge-in flow
Connect to wss://api.x.ai/v1/tts?language=en&voice=eve&codec=mp3 with Authorization: Bearer $XAI_API_KEY on the upgrade, and keep that socket behind your own proxy so the key never lands in a browser. Start an utterance with one or more text.delta messages followed by text.done, then decode each audio.delta base64 chunk into your play queue while the turn is live. When the user interrupts, send {"type":"text.clear"} and wait for {"type":"audio.clear"} before you flush the local buffer and speak the replacement turn with a new text.delta / text.done pair that produces fresh audio.delta frames until audio.done.
Minimal Python sketch
import asyncio, base64, json, os
import websockets
async def barge_in():
uri = "wss://api.x.ai/v1/tts?language=en&voice=eve&codec=mp3"
async with websockets.connect(
uri,
additional_headers={"Authorization": f"Bearer {os.environ['XAI_API_KEY']}"},
) as ws:
await ws.send(json.dumps({
"type": "text.delta",
"delta": "The answer to your question is a long explanation..."
}))
await ws.send(json.dumps({"type": "text.done"}))
first = json.loads(await ws.recv())
await ws.send(json.dumps({"type": "text.clear"}))
async for raw in ws:
if json.loads(raw)["type"] == "audio.clear":
break
await ws.send(json.dumps({
"type": "text.delta",
"delta": "Actually, let me start over."
}))
await ws.send(json.dumps({"type": "text.done"}))
audio = bytearray()
async for raw in ws:
event = json.loads(raw)
if event["type"] == "audio.delta":
audio.extend(base64.b64decode(event["delta"]))
elif event["type"] == "audio.done":
break
return bytes(audio)
asyncio.run(barge_in())
Session behavior after clear
text.clear remains safe when no utterance is active because the server still returns audio.clear immediately, which keeps client state machines simple. Pronunciation replace maps set via session.update survive clear and continue to apply to later utterances under the same multi-turn rules as a normal streaming session. Concurrent streaming TTS sessions stay capped at 50 per team, and cancelling avoids burning a reconnect that can cost roughly 600 ms for distant clients on every barge-in.
Pitfalls
Ignoring audio.clear and continuing to play queued audio.delta bytes produces overlapping speech that sounds like a failed interrupt even though the server already discarded the turn. Closing the WebSocket instead of clearing works functionally but reintroduces handshake latency on every barge-in, which defeats the point of the clear protocol. Sending a new text.delta before audio.clear can race with discarded buffers, so wait for the clear event and only then speak the replacement turn.