VOICE / stream-stt-with-opus-encoding

Voice

Stream speech-to-text with Opus encoding

Cut streaming bandwidth on mobile and other constrained clients by opening the speech-to-text WebSocket with encoding=opus so each binary frame carries one compressed Opus packet instead of raw PCM16. Official Speech to Text documents Opus under Streaming → Opus Streaming: roughly 4 KB/s versus about 48 KB/s for PCM16 at 24 kHz, with no client-side resampling step because Opus packets are sample-rate-agnostic. The default encoding remains pcm, so you must set the query parameter explicitly when your capture stack already emits Opus (WebRTC, many platform audio APIs).

What you need

An xAI API key with Voice access, a backend that proxies wss://api.x.ai/v1/stt (never put the key in the browser), and a producer that emits one raw Opus packet per WebSocket binary frame. Neighboring jobs include Stream speech to text over WebSocket for the PCM baseline, Enable STT Smart Turn, and Bias STT with a language hint and keyterms. More voice jobs live on the Voice hub.

Open the WebSocket with Opus

  1. Export the inference API key outside of source control:
export XAI_API_KEY="your_api_key"
  1. Connect with encoding=opus. Omit sample_rate — the docs state it is ignored for this encoding. Enable interim_results=true when the UI should refresh partial text:
wss://api.x.ai/v1/stt?model=grok-voice-transcribe-2.0&encoding=opus&interim_results=true
  1. Wait for the server transcript.created event before sending audio, then stream binary frames paced near real time (for example about 100 ms of speech per packet from your encoder).
import asyncio
import json
import os

import websockets

WS_URL = (
    "wss://api.x.ai/v1/stt"
    "?model=grok-voice-transcribe-2.0"
    "&encoding=opus"
    "&interim_results=true"
)

async def stream_opus_packets(packets: list[bytes]):
    headers = {"Authorization": f"Bearer {os.environ['XAI_API_KEY']}"}
    async with websockets.connect(WS_URL, additional_headers=headers) as ws:
        ready = json.loads(await ws.recv())
        assert ready["type"] == "transcript.created"
        for packet in packets:
            # Exactly one Opus packet per binary frame — no containers, no concat.
            await ws.send(packet)
            await asyncio.sleep(0.1)
        await ws.send(json.dumps({"type": "audio.done"}))
        async for message in ws:
            event = json.loads(message)
            if event["type"] == "transcript.partial":
                tag = "FINAL" if event.get("is_final") else "partial"
                print(f"[{tag}] {event['text']}")
            elif event["type"] == "transcript.done":
                print(event["text"], event.get("duration"))
                break

# packets = list of raw Opus packet bytes from your encoder (not Ogg/WebM)
# asyncio.run(stream_opus_packets(packets))

Framing rules that matter

Send raw Opus packets, not Ogg-Opus or WebM containers. Container files belong on the batch REST endpoint (POST /v1/stt with file or url), which auto-detects formats. Never concatenate several packets into one frame or split one packet across frames — Opus packets do not mark their own boundaries; the WebSocket frame is the boundary. Multichannel streaming is unsupported with encoding=opus (mono only); for agent/customer stereo, stay on PCM with multichannel=true as in the streaming STT docs.

Pitfalls

If a frame cannot be decoded, the server emits an error event and closes the session. Misframed packets often decode as noise rather than raising a clean error, so validate that your encoder emits one whole packet per frame before blaming the model. Leaving encoding unset keeps the PCM default and will reject or garble Opus bytes. Pair Opus with Use finalize on streaming STT when you need push-to-talk utterance locks without ending the socket via audio.done.