
Set TTS output format via the xAI API
Pin the audio codec, sample rate, and MP3 bit rate on Grok Voice Text to Speech when the default MP3 at 24 kHz / 128 kbps does not match your player, telephony trunk, or post-production pipeline. Official Text to Speech documents the optional output_format object on unary POST https://api.x.ai/v1/tts, and streaming wss://api.x.ai/v1/tts takes the same choices as codec, sample_rate, and bit_rate query parameters when you open the socket. Omitting the object (or the query params) keeps the documented default of MP3 at 24 kHz with a 128 kbps bit rate, which is the right starting point for most web and mobile playback.
What you need
An xAI API key from console.x.ai, required text and language (BCP-47 or auto), and a voice_id such as a built-in from List TTS voices via the API or a custom clone id. Neighboring knobs include Set TTS speed via the xAI API, Optimize TTS streaming latency, and Convert text to speech with the Voice API. More voice jobs live on the Voice hub.
Supported codecs, rates, and bit rates
| Field | Documented values | Typical use |
|---|---|---|
codec |
mp3, wav, pcm, mulaw, alaw |
Web playback, lossless edit, raw pipelines, G.711 telephony |
sample_rate |
8000, 16000, 22050, 24000 (default), 44100, 48000 |
Narrowband trunks through studio masters |
bit_rate |
32000, 64000, 96000, 128000 (default), 192000 |
MP3 only; ignored for other codecs |
Content types map as audio/mpeg for MP3, audio/wav for WAV, audio/pcm for PCM, audio/basic for μ-law, and audio/alaw for A-law per the same guide.
Set format on unary TTS
Export the inference API key outside of source control, then send output_format in the JSON body. High-fidelity narration often uses MP3 at 44.1 kHz with a 192 kbps bit rate:
export XAI_API_KEY="your_api_key"
curl -X POST https://api.x.ai/v1/tts \
-H "Authorization: Bearer ${XAI_API_KEY}" \
-H "Content-Type: application/json" \
-d '{
"text": "Crystal clear audio at maximum quality.",
"voice_id": "rex",
"language": "en",
"output_format": {
"codec": "mp3",
"sample_rate": 44100,
"bit_rate": 192000
}
}' \
--output studio.mp3
For G.711 μ-law telephony, drop to 8 kHz and omit bit rate because it applies only to MP3:
curl -X POST https://api.x.ai/v1/tts \
-H "Authorization: Bearer ${XAI_API_KEY}" \
-H "Content-Type: application/json" \
-d '{
"text": "Hello, thank you for calling. How can I help you today?",
"voice_id": "ara",
"language": "en",
"output_format": {
"codec": "mulaw",
"sample_rate": 8000
}
}' \
--output ivr.ulaw
Save the response body with the matching extension so your player or PBX does not mis-decode PCM or μ-law as MP3.
Set format on streaming TTS
Codec, sample rate, and bit rate are fixed for a WebSocket connection at upgrade time. Pass them as query parameters alongside required language, then send text.delta / text.done as in Stream text to speech with the Voice API:
wss://api.x.ai/v1/tts?language=en&voice=eve&codec=wav&sample_rate=48000
Open a new socket when you need a different format mid-session. Keep the API key on your backend proxy.
Pitfalls
Sending bit_rate with wav, pcm, mulaw, or alaw does not change those codecs and can confuse client validation that assumes every field is honored. Pairing mulaw or alaw with a sample rate other than 8000 fights typical PSTN expectations even when the API accepts the combination. Saving a non-MP3 body as .mp3 produces silent or garbled playback in browsers that trust the extension over the bytes. Format is independent of speed and optimize_streaming_latency, so fix the container first, then tune pace and first-chunk latency separately.