
Format STT numbers and currency with language
Turn spoken amounts into written forms such as $167,983.15 instead of "one hundred sixty-seven thousand…" by enabling Inverse Text Normalization on the Speech-to-Text API. Official Speech to Text documents format=true together with a language code (for example en); formatting numbers, currencies, and units requires both fields. The model still hears speech in supported languages when language is omitted — setting it unlocks written-form formatting. Create a key at console.x.ai.
What you need
An XAI_API_KEY, an audio file or URL to transcribe, and the BCP-47-style language code that matches the speech you want formatted. Neighboring jobs include Bias STT with a language hint and keyterms when product names must stick, Keep filler words in speech-to-text when you need disfluencies, and Transcribe audio with the Voice API for the base batch path. More Voice jobs live on the Voice hub.
Batch REST with format and language
- Export the key. Put option fields before
filein the multipart body (fields afterfilemay be ignored on streamable uploads):
export XAI_API_KEY="your_api_key"
curl -X POST https://api.x.ai/v1/stt \
-H "Authorization: Bearer $XAI_API_KEY" \
-F model=grok-voice-transcribe-2.0 \
-F format=true \
-F language=en \
-F file=@meeting.mp3
Read
textand thewordsarray from the JSON response. With formatting on, currency and number spans appear in written form inside both the full transcript and word timestamps.Or call the same fields from Python (
formatandlanguageindata, file last):
import os
import requests
response = requests.post(
"https://api.x.ai/v1/stt",
headers={"Authorization": f"Bearer {os.environ['XAI_API_KEY']}"},
files={"file": ("meeting.mp3", open("meeting.mp3", "rb"), "audio/mpeg")},
data=[
("model", "grok-voice-transcribe-2.0"),
("format", "true"),
("language", "en"),
],
)
response.raise_for_status()
print(response.json()["text"])
Streaming WebSocket
Add format is not a separate streaming flag — set language=en (or another supported code) on the WebSocket query string so text formatting applies to partial and final events:
wss://api.x.ai/v1/stt?model=grok-voice-transcribe-2.0&sample_rate=16000&encoding=pcm&interim_results=true&language=en
Proxy the socket through your backend so the API key never ships to the browser. Default model when model is omitted is grok-voice-transcribe-2.0; pin grok-voice-transcribe-1.0 only when you intentionally stay on the older slug.
Pitfalls
Sending format=true without language returns 400 — the docs require both. Expecting formatting for a language outside the documented table will not rewrite numbers even if the model transcribes the speech. Putting file before format / language in multipart can drop those options. Confusing Inverse Text Normalization with keyterm biasing mixes two different controls — keyterms bias recognition; format+language rewrite how numbers and currency appear in the text.