
Add preset voices to Imagine reference-to-video
Give subjects in a grok-imagine-video-1.5 clip a speaking voice by passing up to three preset voice_id values in reference_audios, then tag those voices in the prompt as <AUDIO_0>, <AUDIO_1>, and <AUDIO_2> so the model knows who speaks which line. Official Reference-to-Video documents preset voices as generally available on that model, drawn from the same built-in roster as Text to Speech. Identifiers are case-insensitive; an unknown voice_id returns 400 with the available list. Create a key and load credits at console.x.ai.
What you need
An XAI_API_KEY with Console balance, model grok-imagine-video-1.5, optional reference_images when a person or product should appear with the voice, and a prompt that names each voice with the <AUDIO_n> tags. Neighboring jobs include Generate a video from reference images for image-only reference-to-video, Pin keyframes on Imagine video 1.5 when mid-clip stills matter, and Pick built-in voices for the Grok Voice API for the TTS roster those ids come from. More Imagine jobs live on the API hub.
Generate with preset voices
- Export the key and keep it out of chat logs and public repos:
export XAI_API_KEY="your_api_key"
- Start an async generation that mixes reference images and preset voices (audio-only reference-to-video is also valid when you omit images):
REQUEST_ID=$(curl -s -X POST https://api.x.ai/v1/videos/generations \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $XAI_API_KEY" \
-d '{
"model": "grok-imagine-video-1.5",
"prompt": "The person from <IMAGE_0> presents the product from <IMAGE_1> on the set from <IMAGE_2>, speaking with the voice from <AUDIO_0>. A second speaker with the voice from <AUDIO_1> replies.",
"reference_images": [
{"url": "<IMAGE_URL_1>"},
{"url": "<IMAGE_URL_2>"},
{"url": "<IMAGE_URL_3>"}
],
"reference_audios": [
{"voice_id": "eve"},
{"voice_id": "leo"}
],
"duration": 8,
"aspect_ratio": "9:16",
"resolution": "720p"
}' | jq -r '.request_id')
echo "Request ID: $REQUEST_ID"
In the xAI Python SDK, pass the same
reference_audioslist of{"voice_id": "..."}objects alongsidereference_image_urlsonclient.video.generate, then read the completed URL from the returned response when the SDK polls for you.Poll
GET https://api.x.ai/v1/videos/$REQUEST_IDuntilstatusisdone, then download the temporary URL promptly. Persist withstorage_optionswhen you need a stable Files copy.
Pitfalls
Custom caller-supplied audio clips as voice references are partner-only; stick to preset voice_id values unless your team has that grant. Classic grok-imagine-video rejects this reference-audio path. Forgetting the <AUDIO_n> tags in the prompt leaves the model without an explicit speaker map. Pasting keys into tickets or committing them to git forces a rotate.