
Enable interim results on streaming speech-to-text
Show live captions that update while the speaker is still talking by turning on interim_results on the streaming STT WebSocket. Official Speech to Text lists interim_results under Streaming → Query Parameters: when true, the server emits partial transcripts with is_final=false about every 500 ms; when false (default), you only receive chunk-final and utterance-final events. Create a key and load credits at console.x.ai.
What you need
An xAI API key, a backend proxy to wss://api.x.ai/v1/stt (never put the key in the browser), and a UI that can redraw caption text as interim strings arrive. Neighboring jobs include Stream speech to text over WebSocket for the full connect-and-stream loop, Use finalize on streaming speech-to-text for push-to-talk locks, and Enable STT Smart Turn when you also want ML end-of-turn confidence. More voice jobs live on the Voice hub.
Turn on interim_results
- Open the WebSocket with
interim_results=truein the query string. Prefer 16 kHz PCM for the model’s native rate:
wss://api.x.ai/v1/stt?model=grok-voice-transcribe-2.0&sample_rate=16000&encoding=pcm&interim_results=true&language=en
- Wait for
transcript.created, then stream raw audio as binary frames (~100 ms chunks). Readtranscript.partialevents and branch onis_final/speech_final:
is_final |
speech_final |
Meaning |
|---|---|---|
false |
false |
Interim — text may still change (only with interim_results=true) |
true |
false |
Chunk final — roughly 3s of speech locked |
true |
true |
Utterance final — speaker paused / turn complete |
import asyncio
import json
import os
import websockets
WS_URL = (
"wss://api.x.ai/v1/stt"
"?model=grok-voice-transcribe-2.0"
"&sample_rate=16000"
"&encoding=pcm"
"&interim_results=true"
"&language=en"
)
async def live_captions(audio_path: str):
headers = {"Authorization": f"Bearer {os.environ['XAI_API_KEY']}"}
async with websockets.connect(WS_URL, additional_headers=headers) as ws:
assert json.loads(await ws.recv())["type"] == "transcript.created"
with open(audio_path, "rb") as f:
f.read(44) # skip WAV header for this PCM16 demo
chunk = 16000 * 2 // 10 # 100 ms
while data := f.read(chunk):
await ws.send(data)
await asyncio.sleep(0.1)
await ws.send(json.dumps({"type": "audio.done"}))
async for message in ws:
event = json.loads(message)
if event["type"] == "transcript.partial":
if not event["is_final"]:
print(f"[live] {event['text']}", end="\r", flush=True)
elif event.get("speech_final"):
print(f"\n[utterance] {event['text']}")
else:
print(f"\n[chunk] {event['text']}")
elif event["type"] == "transcript.done":
print(f"\n[done] {event['text']} ({event['duration']}s)")
break
asyncio.run(live_captions("audio.wav"))
- In the UI, treat interim lines as disposable overlays: replace the live string on each
is_final=falseevent, then commit text to the transcript only whenis_finalbecomestrue. That keeps captions responsive without writing every draft into storage.
Pair with other streaming knobs
interim_results stacks with language, keyterm, diarize, smart_turn, and Opus encoding. Smart Turn still demotes low-confidence silence to chunk-final instead of utterance-final; interim events continue during active speech. For bandwidth-constrained clients, set encoding=opus and keep interim_results=true so mobile captions stay live at roughly 4 KB/s instead of raw PCM rates.
Pitfalls
Omitting interim_results (or setting it false) means you will never see is_final=false rows — captions jump only when chunks lock. Do not wait for transcript.done to paint the first words; that event arrives after audio.done and closes the session. Dumping the entire file as one binary frame breaks pacing; send real-time-sized chunks. Keep the API key on the server and proxy the socket for browsers.