
Use finalize on streaming speech-to-text
Force the current streaming utterance to lock as speech_final immediately when the user releases a push-to-talk control, instead of waiting for silence endpointing or Smart Turn confidence. Official Speech to Text lists the client message {"type":"finalize"} under Streaming → Client Messages: it finalizes the active utterance for PTT while the WebSocket session stays open for more audio. That is different from {"type":"audio.done"}, which flushes the final transcript.done event and closes the socket.
What you need
An xAI API key, a backend proxy to wss://api.x.ai/v1/stt, and a client UX that knows when the speaker finished a turn (button release, foot pedal, or hotkey). Neighboring jobs include Stream speech to text over WebSocket, Set STT endpointing, Enable STT Smart Turn, and Stream speech-to-text with Opus encoding. More voice jobs live on the Voice hub.
Finalize on push-to-talk release
- Open the streaming STT WebSocket with the encodings and options your app needs (PCM example below). Wait for
transcript.createdbefore sending mic frames.
wss://api.x.ai/v1/stt?model=grok-voice-transcribe-2.0&sample_rate=16000&encoding=pcm&interim_results=true
While the PTT button is held, send raw audio as binary frames (PCM16 little-endian for
encoding=pcm, or one Opus packet per frame whenencoding=opus).On button release, send a JSON finalize message so the server promotes the current utterance to
speech_finalwithout ending the session:
{"type": "finalize"}
The docs' push-to-talk example also shows {"type":"Finalize"} — send the JSON object as a text WebSocket frame, then keep the connection open for the next hold.
import asyncio
import json
import os
import websockets
WS_URL = (
"wss://api.x.ai/v1/stt"
"?model=grok-voice-transcribe-2.0"
"&sample_rate=16000"
"&encoding=pcm"
"&interim_results=true"
)
async def ptt_session(pcm_chunks: list[bytes]):
headers = {"Authorization": f"Bearer {os.environ['XAI_API_KEY']}"}
async with websockets.connect(WS_URL, additional_headers=headers) as ws:
assert json.loads(await ws.recv())["type"] == "transcript.created"
for chunk in pcm_chunks:
await ws.send(chunk)
# Button released — lock this utterance, keep socket alive.
await ws.send(json.dumps({"type": "finalize"}))
# Continue reading transcript.partial events; send more audio for the next PTT.
# When the whole session ends, send {"type": "audio.done"} instead.
- Read
transcript.partialevents. After finalize you should seeis_final=trueandspeech_final=truefor the locked utterance while later presses can start fresh speech on the same connection.
Multichannel finalize
When multichannel=true and channels ≥ 2 (PCM only — Opus is mono), you can finalize one channel without touching the others:
{"type": "Finalize", "channel": 0}
Omit channel to finalize every channel at once. Typical call-center wiring puts the agent on channel 0 and the caller on channel 1 so each side can lock independently.
Choose finalize vs endpointing vs audio.done
Use finalize for explicit PTT or when the UI owns turn boundaries. Use endpointing / Smart Turn when hands-free VAD should decide silence. Use audio.done only when the client will send no more audio and wants transcript.done plus connection close. Mixing them is fine across a session — finalize mid-call, then audio.done at hang-up.
Pitfalls
Sending audio.done when you meant finalize will end the WebSocket and drop the next press unless you reconnect. Finalize does not replace waiting for transcript.created at connect time. Keep API keys on the server; browser clients should talk to your proxy, not directly to api.x.ai.