
Use a custom voice in speech-to-speech
Make a realtime voice agent speak with a cloned voice_id instead of a built-in roster name such as eve. Official Speech to Speech documents the voice session parameter as accepting any built-in voice or a custom voice ID from the Custom Voices API, and the Voice Overview quick start states that the same custom voice_id works on Speech to Speech session.update. Create or copy the id at console.x.ai (or via Enterprise POST /v1/custom-voices).
What you need
An xAI API key (or ephemeral client token for browsers), a custom voice_id (8 lowercase alphanumeric characters), and a WebSocket client for wss://api.x.ai/v1/realtime. Neighboring jobs include Create a custom voice for cloning the reference clip, Start a speech-to-speech session for the first-turn event loop, and Get a custom voice via the xAI API when you need metadata before you connect. More voice jobs live on the Voice hub.
Pass voice_id on session.update
Resolve the custom id (Console three-dot menu, or
GET /v1/custom-voicesfor Enterprise list). Do not look for customs onGET /v1/tts/voices— that endpoint is built-ins only.Open the realtime socket with a model query param, then send
session.updatewithvoiceset to the custom id (same field you would set to"eve"):
import asyncio
import json
import os
import websockets
VOICE_ID = "nlbqfwie" # replace with your custom voice_id
MODEL = "grok-voice-latest"
async def custom_voice_agent():
url = f"wss://api.x.ai/v1/realtime?model={MODEL}"
headers = {"Authorization": f"Bearer {os.environ['XAI_API_KEY']}"}
async with websockets.connect(url, additional_headers=headers) as ws:
await ws.send(
json.dumps(
{
"type": "session.update",
"session": {
"voice": VOICE_ID,
"instructions": "You are a helpful assistant.",
"turn_detection": {"type": "server_vad"},
"audio": {
"input": {"format": {"type": "audio/pcm", "rate": 24000}},
"output": {"format": {"type": "audio/pcm", "rate": 24000}},
},
},
}
)
)
await ws.send(
json.dumps(
{
"type": "conversation.item.create",
"item": {
"type": "message",
"role": "user",
"content": [
{"type": "input_text", "text": "Introduce yourself briefly."}
],
},
}
)
)
await ws.send(json.dumps({"type": "response.create"}))
async for raw in ws:
event = json.loads(raw)
print(event["type"])
if event["type"] in ("response.done", "error"):
break
asyncio.run(custom_voice_agent())
- Play assistant audio from
response.output_audio.delta(JSON/base64) or binary frames whenaudio.output.transportis"binary", using the same sample rate you configured. You can changevoicemid-session with anothersession.update; the applied config echoes onsession.updated.
Same id across TTS and realtime
One custom voice_id works on unary/streaming TTS and on Speech to Speech. Clone once, then reuse the id in phone agents, web agents, and batch narration. Custom Voices availability is United States only (Illinois excluded); Console create still works when the Enterprise multipart create returns 403. Cap is 30 voices per team.
Pitfalls
Passing a built-in name that was deleted or mistyped, or a custom id from another team, fails voice selection — verify with Console or GET /v1/custom-voices. Browsers cannot send Authorization on WebSocket; mint an ephemeral token server-side. Keep API keys off the client. Audio and text-input billing for Speech to Speech is separate from SuperGrok’s weekly pool on grok.com.