VOICE / enable-stt-interim-results

Voice

Enable interim results on streaming speech-to-text

Show live captions that update while the speaker is still talking by turning on interim_results on the streaming STT WebSocket. Official Speech to Text lists interim_results under Streaming → Query Parameters: when true, the server emits partial transcripts with is_final=false about every 500 ms; when false (default), you only receive chunk-final and utterance-final events. Create a key and load credits at console.x.ai.

What you need

An xAI API key, a backend proxy to wss://api.x.ai/v1/stt (never put the key in the browser), and a UI that can redraw caption text as interim strings arrive. Neighboring jobs include Stream speech to text over WebSocket for the full connect-and-stream loop, Use finalize on streaming speech-to-text for push-to-talk locks, and Enable STT Smart Turn when you also want ML end-of-turn confidence. More voice jobs live on the Voice hub.

Turn on interim_results

  1. Open the WebSocket with interim_results=true in the query string. Prefer 16 kHz PCM for the model’s native rate:
wss://api.x.ai/v1/stt?model=grok-voice-transcribe-2.0&sample_rate=16000&encoding=pcm&interim_results=true&language=en
  1. Wait for transcript.created, then stream raw audio as binary frames (~100 ms chunks). Read transcript.partial events and branch on is_final / speech_final:
is_final speech_final Meaning
false false Interim — text may still change (only with interim_results=true)
true false Chunk final — roughly 3s of speech locked
true true Utterance final — speaker paused / turn complete
import asyncio
import json
import os

import websockets

WS_URL = (
    "wss://api.x.ai/v1/stt"
    "?model=grok-voice-transcribe-2.0"
    "&sample_rate=16000"
    "&encoding=pcm"
    "&interim_results=true"
    "&language=en"
)

async def live_captions(audio_path: str):
    headers = {"Authorization": f"Bearer {os.environ['XAI_API_KEY']}"}
    async with websockets.connect(WS_URL, additional_headers=headers) as ws:
        assert json.loads(await ws.recv())["type"] == "transcript.created"
        with open(audio_path, "rb") as f:
            f.read(44)  # skip WAV header for this PCM16 demo
            chunk = 16000 * 2 // 10  # 100 ms
            while data := f.read(chunk):
                await ws.send(data)
                await asyncio.sleep(0.1)
        await ws.send(json.dumps({"type": "audio.done"}))
        async for message in ws:
            event = json.loads(message)
            if event["type"] == "transcript.partial":
                if not event["is_final"]:
                    print(f"[live] {event['text']}", end="\r", flush=True)
                elif event.get("speech_final"):
                    print(f"\n[utterance] {event['text']}")
                else:
                    print(f"\n[chunk] {event['text']}")
            elif event["type"] == "transcript.done":
                print(f"\n[done] {event['text']} ({event['duration']}s)")
                break

asyncio.run(live_captions("audio.wav"))
  1. In the UI, treat interim lines as disposable overlays: replace the live string on each is_final=false event, then commit text to the transcript only when is_final becomes true. That keeps captions responsive without writing every draft into storage.

Pair with other streaming knobs

interim_results stacks with language, keyterm, diarize, smart_turn, and Opus encoding. Smart Turn still demotes low-confidence silence to chunk-final instead of utterance-final; interim events continue during active speech. For bandwidth-constrained clients, set encoding=opus and keep interim_results=true so mobile captions stay live at roughly 4 KB/s instead of raw PCM rates.

Pitfalls

Omitting interim_results (or setting it false) means you will never see is_final=false rows — captions jump only when chunks lock. Do not wait for transcript.done to paint the first words; that event arrives after audio.done and closes the session. Dumping the entire file as one binary frame breaks pacing; send real-time-sized chunks. Keep the API key on the server and proxy the socket for browsers.