
Enable text normalization in TTS
Turn on written-to-spoken expansion when your TTS input still contains digits, currency symbols, abbreviations, or dates that should be read the way a person would say them aloud. Official Text to Speech documents the boolean text_normalization on unary POST https://api.x.ai/v1/tts and as a query parameter on the streaming WebSocket wss://api.x.ai/v1/tts. When the flag is true, the model rewrites written-form tokens into spoken-form before synthesis; the default remains false, so receipts, invoices, and support scripts keep sounding digit-by-digit until you opt in.
What you need
An xAI API key with Voice access, and text that benefits from Inverse Text Normalization style expansion (amounts, phone-adjacent numbers, abbreviated units). Neighboring jobs include Convert text to speech with the Voice API, Get TTS character timestamps, and Fix pronunciation with the TTS replace map. More voice jobs live on the Voice hub.
Enable normalization on unary TTS
- Export the inference API key outside of source control:
export XAI_API_KEY="your_api_key"
- POST with
text_normalizationset totruealongside requiredtextandlanguage:
curl -X POST https://api.x.ai/v1/tts \
-H "Authorization: Bearer ${XAI_API_KEY}" \
-H "Content-Type: application/json" \
-d '{
"text": "Your total is $1,250.00, due 6/12 at 3:30 PM.",
"voice_id": "eve",
"language": "en",
"text_normalization": true
}' \
--output normalized.mp3
- Play the file and confirm spoken expansions (currency and clock times) rather than digit spelling. Leave the flag off when you need literal character readout for codes or serials.
Enable normalization on streaming TTS
Open the WebSocket with text_normalization=true in the query string (same accepted values as unary: true / false, default false):
wss://api.x.ai/v1/tts?language=en&voice=eve&codec=mp3&text_normalization=true
Send text.delta then text.done as in Stream text to speech with the Voice API. Normalization applies for that connection configuration; reconnect if you need the opposite setting mid-session.
Pair with timestamps carefully
When with_timestamps is also true, graph_chars still mirrors the written characters you sent. Normalization expands symbols into multiple spoken words while keeping the original character count, so the $ in $5 can carry the full timing span for "five dollars" while the 5 interpolates inside that span. Walk graph_chars and graph_times in index order instead of slicing the input string by assumed spoken length. See Get TTS character timestamps for the JSON envelope shape.
Pitfalls
Leaving text_normalization unset keeps the default false, so demo scripts that look fine in the playground with normalization on will regress in production if the flag never ships in your client. Combining normalization with a dense replace map is valid, but replacements run on the text you send; phonetics in replace values stay unaffected by normalization per the same docs. Client-side calls that expose the API key remain unsupported — proxy unary and WebSocket traffic through your backend.