VOICE / download-custom-voice-reference-audio

Voice

Download custom voice reference audio

Pull the original reference clip that produced a custom voice when you need to audit quality, archive the sample, or compare against a new recording before you delete and recreate the clone. Official Custom Voices documents GET https://api.x.ai/v1/custom-voices/{voice_id}/audio, which streams the stored reference file with an appropriate Content-Type header such as audio/wav or audio/mpeg. This path is separate from the metadata GET — it returns bytes, not the JSON voice object.

What you need

An xAI API key for the team that owns the voice, and a valid voice_id from Create a custom voice, List custom voices via the xAI API, or Get a custom voice via the xAI API. Custom Voices availability is United States only, except Illinois, per the same docs. Neighboring jobs include Update custom voice metadata via the xAI API, Delete a custom voice via the xAI API, and Convert text to speech with the Voice API. More voice jobs live on the Voice hub.

Stream the reference file

  1. Export the inference API key and the target voice id:
export XAI_API_KEY="your_api_key"
export VOICE_ID="nlbqfwie"
  1. GET the /audio subpath and write the body to disk. Follow redirects if your client requires it, and preserve the response Content-Type when choosing an extension:
curl -sL "https://api.x.ai/v1/custom-voices/${VOICE_ID}/audio" \
  -H "Authorization: Bearer ${XAI_API_KEY}" \
  -D /tmp/voice-audio.headers \
  -o /tmp/voice-reference.bin

# Inspect Content-Type, then rename (example for WAV):
grep -i '^content-type:' /tmp/voice-audio.headers
mv /tmp/voice-reference.bin /tmp/voice-reference.wav
import os
import requests

voice_id = os.environ["VOICE_ID"]
response = requests.get(
    f"https://api.x.ai/v1/custom-voices/{voice_id}/audio",
    headers={"Authorization": f"Bearer {os.environ['XAI_API_KEY']}"},
)
if response.status_code == 404:
    raise SystemExit("voice or audio not found for this team")
response.raise_for_status()
content_type = response.headers.get("Content-Type", "application/octet-stream")
ext = "wav" if "wav" in content_type else "bin"
path = f"reference-{voice_id}.{ext}"
with open(path, "wb") as f:
    f.write(response.content)
print(path, content_type, len(response.content), "bytes")
  1. Listen to the file and compare against your recording checklist from the Custom Voices guide: quiet room, single speaker, expressive delivery, preferably 90–120 seconds, WAV at about 24 kHz mono when you re-upload.

Re-record workflow

There is no API to replace the underlying audio on an existing voice_id. Metadata PATCH leaves the sample untouched. When the downloaded reference proves the clone was trained on a bad take, Delete a custom voice via the xAI API, then create a new voice from a cleaner clip and update any TTS or Speech-to-Speech clients to the new id.

Pitfalls

Do not confuse this endpoint with TTS synthesis — /audio returns the upload you (or your team) provided, not a freshly generated greeting. A 404 means the voice is missing for this team or already deleted. Lossy reference formats (MP3 and friends) can bake compression artifacts into the clone; prefer WAV on the next create. Keep downloads on a backend so the API key never ships to the client.