VOICE / set-tts-speed-via-api

Voice

Set TTS speed via the xAI API

Speed up or slow down Grok Voice TTS when the default pacing is too slow for IVR menus or too fast for accessibility playback, using the documented speed multiplier instead of editing the script or post-processing the WAV. Official Text to Speech places speed on unary POST https://api.x.ai/v1/tts and as a query parameter on streaming wss://api.x.ai/v1/tts, with an accepted range of 0.7 to 1.5 and a default of 1.0 for normal pace. Values below 1.0 stretch delivery for clearer listening, while values above 1.0 compress wall-clock duration without changing the characters you send or the voice you selected.

What you need

You need an xAI API key, a chosen voice_id (a built-in such as eve or a custom clone id from the console or Custom Voices API), and a target multiplier that stays inside the documented range so validation does not reject the call. Neighboring jobs include Convert text to speech with the Voice API, Enable text normalization in TTS, and Stream text to speech with the Voice API. More voice jobs live on the Voice hub.

Set speed on unary TTS

Export the inference API key outside of source control, then include speed in the JSON body alongside required text and language. The docs full-options example uses 1.2 for snappier narration while also setting a high-fidelity output_format; you can omit output_format when the default MP3 at 24 kHz / 128 kbps is enough.

export XAI_API_KEY="your_api_key"

curl -X POST https://api.x.ai/v1/tts \
  -H "Authorization: Bearer ${XAI_API_KEY}" \
  -H "Content-Type: application/json" \
  -d '{
    "text": "Hello! This is a high-fidelity text to speech example.",
    "voice_id": "ara",
    "language": "en",
    "speed": 1.2,
    "output_format": {
      "codec": "mp3",
      "sample_rate": 44100,
      "bit_rate": 192000
    }
  }' \
  --output faster.mp3

For slower accessibility playback, try 0.85 or 0.7 on the same text and compare file duration against a 1.0 baseline before you ship. Reject values outside 0.7–1.5 in your client so callers never hit a 400 from an accidental percentage such as 200.

Set speed on streaming TTS

Pass speed on the WebSocket URL when you open the session, because audio parameters including pace are fixed for that connection and changing them mid-conversation means opening a new socket with the new multiplier:

wss://api.x.ai/v1/tts?language=en&voice=eve&codec=mp3&speed=1.25

After the upgrade succeeds with Bearer auth, send text.delta then text.done and decode audio.delta chunks the same way as Stream text to speech with the Voice API. Keep the proxy on your backend so the key never reaches a browser.

Combine with latency and format knobs

speed stays independent of optimize_streaming_latency levels 0 / 1 / 2 and of codec, sample rate, and bit rate inside output_format, so you can raise speed when you need shorter wall-clock audio and raise latency optimization when you need earlier first-chunk audio on the stream. Telephony codecs such as mulaw or alaw at 8 kHz still need a multiplier that remains intelligible on narrowband lines after compression.

Pitfalls

Treating speed like a free-form percentage returns a client error instead of clamping to the documented bounds. Inline speech tags that alter local delivery still apply on top of the global multiplier, so a wrapping delivery tag plus an extreme speed value can stack into hard-to-follow audio that you should listen through once before shipping IVR copy. Never expose the API key from a browser; proxy both unary and streaming TTS through your server.