VOICE / format-stt-numbers-with-language

Voice

Format STT numbers and currency with language

Turn spoken amounts into written forms such as $167,983.15 instead of "one hundred sixty-seven thousand…" by enabling Inverse Text Normalization on the Speech-to-Text API. Official Speech to Text documents format=true together with a language code (for example en); formatting numbers, currencies, and units requires both fields. The model still hears speech in supported languages when language is omitted — setting it unlocks written-form formatting. Create a key at console.x.ai.

What you need

An XAI_API_KEY, an audio file or URL to transcribe, and the BCP-47-style language code that matches the speech you want formatted. Neighboring jobs include Bias STT with a language hint and keyterms when product names must stick, Keep filler words in speech-to-text when you need disfluencies, and Transcribe audio with the Voice API for the base batch path. More Voice jobs live on the Voice hub.

Batch REST with format and language

  1. Export the key. Put option fields before file in the multipart body (fields after file may be ignored on streamable uploads):
export XAI_API_KEY="your_api_key"

curl -X POST https://api.x.ai/v1/stt \
  -H "Authorization: Bearer $XAI_API_KEY" \
  -F model=grok-voice-transcribe-2.0 \
  -F format=true \
  -F language=en \
  -F file=@meeting.mp3
  1. Read text and the words array from the JSON response. With formatting on, currency and number spans appear in written form inside both the full transcript and word timestamps.

  2. Or call the same fields from Python (format and language in data, file last):

import os
import requests

response = requests.post(
    "https://api.x.ai/v1/stt",
    headers={"Authorization": f"Bearer {os.environ['XAI_API_KEY']}"},
    files={"file": ("meeting.mp3", open("meeting.mp3", "rb"), "audio/mpeg")},
    data=[
        ("model", "grok-voice-transcribe-2.0"),
        ("format", "true"),
        ("language", "en"),
    ],
)
response.raise_for_status()
print(response.json()["text"])

Streaming WebSocket

Add format is not a separate streaming flag — set language=en (or another supported code) on the WebSocket query string so text formatting applies to partial and final events:

wss://api.x.ai/v1/stt?model=grok-voice-transcribe-2.0&sample_rate=16000&encoding=pcm&interim_results=true&language=en

Proxy the socket through your backend so the API key never ships to the browser. Default model when model is omitted is grok-voice-transcribe-2.0; pin grok-voice-transcribe-1.0 only when you intentionally stay on the older slug.

Pitfalls

Sending format=true without language returns 400 — the docs require both. Expecting formatting for a language outside the documented table will not rewrite numbers even if the model transcribes the speech. Putting file before format / language in multipart can drop those options. Confusing Inverse Text Normalization with keyterm biasing mixes two different controls — keyterms bias recognition; format+language rewrite how numbers and currency appear in the text.