API / add-preset-voices-to-imagine-video

API

Add preset voices to Imagine reference-to-video

Give subjects in a grok-imagine-video-1.5 clip a speaking voice by passing up to three preset voice_id values in reference_audios, then tag those voices in the prompt as <AUDIO_0>, <AUDIO_1>, and <AUDIO_2> so the model knows who speaks which line. Official Reference-to-Video documents preset voices as generally available on that model, drawn from the same built-in roster as Text to Speech. Identifiers are case-insensitive; an unknown voice_id returns 400 with the available list. Create a key and load credits at console.x.ai.

What you need

An XAI_API_KEY with Console balance, model grok-imagine-video-1.5, optional reference_images when a person or product should appear with the voice, and a prompt that names each voice with the <AUDIO_n> tags. Neighboring jobs include Generate a video from reference images for image-only reference-to-video, Pin keyframes on Imagine video 1.5 when mid-clip stills matter, and Pick built-in voices for the Grok Voice API for the TTS roster those ids come from. More Imagine jobs live on the API hub.

Generate with preset voices

  1. Export the key and keep it out of chat logs and public repos:
export XAI_API_KEY="your_api_key"
  1. Start an async generation that mixes reference images and preset voices (audio-only reference-to-video is also valid when you omit images):
REQUEST_ID=$(curl -s -X POST https://api.x.ai/v1/videos/generations \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $XAI_API_KEY" \
  -d '{
    "model": "grok-imagine-video-1.5",
    "prompt": "The person from <IMAGE_0> presents the product from <IMAGE_1> on the set from <IMAGE_2>, speaking with the voice from <AUDIO_0>. A second speaker with the voice from <AUDIO_1> replies.",
    "reference_images": [
      {"url": "<IMAGE_URL_1>"},
      {"url": "<IMAGE_URL_2>"},
      {"url": "<IMAGE_URL_3>"}
    ],
    "reference_audios": [
      {"voice_id": "eve"},
      {"voice_id": "leo"}
    ],
    "duration": 8,
    "aspect_ratio": "9:16",
    "resolution": "720p"
  }' | jq -r '.request_id')

echo "Request ID: $REQUEST_ID"
  1. In the xAI Python SDK, pass the same reference_audios list of {"voice_id": "..."} objects alongside reference_image_urls on client.video.generate, then read the completed URL from the returned response when the SDK polls for you.

  2. Poll GET https://api.x.ai/v1/videos/$REQUEST_ID until status is done, then download the temporary URL promptly. Persist with storage_options when you need a stable Files copy.

Pitfalls

Custom caller-supplied audio clips as voice references are partner-only; stick to preset voice_id values unless your team has that grant. Classic grok-imagine-video rejects this reference-audio path. Forgetting the <AUDIO_n> tags in the prompt leaves the model without an explicit speaker map. Pasting keys into tickets or committing them to git forces a rotate.