VOICE / use-finalize-on-streaming-stt

Voice

Use finalize on streaming speech-to-text

Force the current streaming utterance to lock as speech_final immediately when the user releases a push-to-talk control, instead of waiting for silence endpointing or Smart Turn confidence. Official Speech to Text lists the client message {"type":"finalize"} under Streaming → Client Messages: it finalizes the active utterance for PTT while the WebSocket session stays open for more audio. That is different from {"type":"audio.done"}, which flushes the final transcript.done event and closes the socket.

What you need

An xAI API key, a backend proxy to wss://api.x.ai/v1/stt, and a client UX that knows when the speaker finished a turn (button release, foot pedal, or hotkey). Neighboring jobs include Stream speech to text over WebSocket, Set STT endpointing, Enable STT Smart Turn, and Stream speech-to-text with Opus encoding. More voice jobs live on the Voice hub.

Finalize on push-to-talk release

  1. Open the streaming STT WebSocket with the encodings and options your app needs (PCM example below). Wait for transcript.created before sending mic frames.
wss://api.x.ai/v1/stt?model=grok-voice-transcribe-2.0&sample_rate=16000&encoding=pcm&interim_results=true
  1. While the PTT button is held, send raw audio as binary frames (PCM16 little-endian for encoding=pcm, or one Opus packet per frame when encoding=opus).

  2. On button release, send a JSON finalize message so the server promotes the current utterance to speech_final without ending the session:

{"type": "finalize"}

The docs' push-to-talk example also shows {"type":"Finalize"} — send the JSON object as a text WebSocket frame, then keep the connection open for the next hold.

import asyncio
import json
import os

import websockets

WS_URL = (
    "wss://api.x.ai/v1/stt"
    "?model=grok-voice-transcribe-2.0"
    "&sample_rate=16000"
    "&encoding=pcm"
    "&interim_results=true"
)

async def ptt_session(pcm_chunks: list[bytes]):
    headers = {"Authorization": f"Bearer {os.environ['XAI_API_KEY']}"}
    async with websockets.connect(WS_URL, additional_headers=headers) as ws:
        assert json.loads(await ws.recv())["type"] == "transcript.created"
        for chunk in pcm_chunks:
            await ws.send(chunk)
        # Button released — lock this utterance, keep socket alive.
        await ws.send(json.dumps({"type": "finalize"}))
        # Continue reading transcript.partial events; send more audio for the next PTT.
        # When the whole session ends, send {"type": "audio.done"} instead.
  1. Read transcript.partial events. After finalize you should see is_final=true and speech_final=true for the locked utterance while later presses can start fresh speech on the same connection.

Multichannel finalize

When multichannel=true and channels ≥ 2 (PCM only — Opus is mono), you can finalize one channel without touching the others:

{"type": "Finalize", "channel": 0}

Omit channel to finalize every channel at once. Typical call-center wiring puts the agent on channel 0 and the caller on channel 1 so each side can lock independently.

Choose finalize vs endpointing vs audio.done

Use finalize for explicit PTT or when the UI owns turn boundaries. Use endpointing / Smart Turn when hands-free VAD should decide silence. Use audio.done only when the client will send no more audio and wants transcript.done plus connection close. Mixing them is fine across a session — finalize mid-call, then audio.done at hang-up.

Pitfalls

Sending audio.done when you meant finalize will end the WebSocket and drop the next press unless you reconnect. Finalize does not replace waiting for transcript.created at connect time. Keep API keys on the server; browser clients should talk to your proxy, not directly to api.x.ai.