VOICE / optimize-tts-streaming-latency

Voice

Optimize TTS streaming latency via the xAI API

Shorten time-to-first-audio on Grok Voice Text to Speech when interactive agents need sound sooner than the default quality-first buffering, using the documented optimize_streaming_latency levels instead of truncating the script. Official Text to Speech describes the field for streaming synthesis: 0 (default) keeps the best audio quality with no optimization, 1 reduces first-chunk size for lower time-to-first-audio with a minor quality tradeoff at chunk boundaries, and 2 reduces first-chunk size further for the lowest time-to-first-audio with a more noticeable boundary tradeoff. The REST streaming reference documents the same knob as a WebSocket query parameter (default 0), so you set it when you open wss://api.x.ai/v1/tts rather than after the first text.delta.

What you need

An xAI API key, a streaming client that upgrades to wss://api.x.ai/v1/tts with Bearer auth on your backend, and a clear choice between quality (0) and earlier first audio (1 or 2 per the guide). Neighboring jobs include Stream text to speech with the Voice API, Set TTS output format via the xAI API, and Set TTS speed via the xAI API. More voice jobs live on the Voice hub.

Choose a latency level

Level Guide meaning When to use it
0 No optimization; best audio quality (default) Offline render, final narration, quality QA
1 Smaller first chunk; minor boundary tradeoff Live agents that need a quicker start without aggressive artifacts
2 Smallest first chunk; more noticeable boundary tradeoff Hard latency budgets where first sound beats chunk polish

The parameter does not change speed, codec, sample rate, or bit rate. Raise speed when you need shorter finished audio; raise latency optimization when you need the first audible chunk sooner.

Set the level on the streaming socket

Pass optimize_streaming_latency on the WebSocket URL with required language and your chosen voice. Keep the key off the browser:

wss://api.x.ai/v1/tts?language=en&voice=eve&codec=mp3&sample_rate=24000&optimize_streaming_latency=1
import asyncio
import base64
import json
import os

import websockets

async def stream_low_latency():
    uri = (
        "wss://api.x.ai/v1/tts"
        "?language=en&voice=ara&codec=mp3"
        "&optimize_streaming_latency=1"
    )
    headers = {"Authorization": f"Bearer {os.environ['XAI_API_KEY']}"}
    async with websockets.connect(uri, additional_headers=headers) as ws:
        await ws.send(json.dumps({"type": "text.delta", "text": "Thanks for waiting. Here is your answer."}))
        await ws.send(json.dumps({"type": "text.done"}))
        async for raw in ws:
            msg = json.loads(raw)
            if msg.get("type") == "audio.delta":
                chunk = base64.b64decode(msg["delta"])
                # append chunk to your playback buffer
            if msg.get("type") == "audio.done":
                break

asyncio.run(stream_low_latency())

Measure time from the first text.delta to the first audio.delta under 0 and 1 (and 2 when your client follows the guide’s three-level table) on the same text before you lock a product default. Changing the level mid-utterance requires a new connection because query parameters are fixed at upgrade.

Unary requests and related flags

The Text to Speech request-body table also lists optimize_streaming_latency beside unary fields. Prefer the streaming path when the job is time-to-first-audio; a full unary response still waits for synthesis of the whole utterance before you hear anything. Do not confuse this knob with with_timestamps, which adds a post-synthesis alignment pass and can increase latency when you enable it for captions.

Pitfalls

Leaving the default 0 in a voice agent and blaming the model for slow starts misses the documented first-chunk control. Jumping to 2 without listening for boundary artifacts can ship choppy IVR audio that users hear as dropouts. Combining aggressive latency optimization with extreme speed and heavy speech tags stacks delivery effects that are hard to debug separately. Never put the API key in a browser WebSocket; proxy the upgrade through your server the same way as other streaming TTS jobs.