
Transcribe audio from a URL via speech-to-text
Hand the Speech-to-Text API a publicly reachable audio URL when the file already lives in object storage and you do not want to upload bytes from your app server. Official Speech to Text documents multipart POST https://api.x.ai/v1/stt with either file or url required, and when you pass url the service downloads the audio server-side before returning the same JSON transcript shape (text, duration, words, and optional channels) you already get from a local upload.
What you need
You need an xAI API key plus an HTTPS URL your team controls (or that allows the xAI fetch) pointing at a supported container such as MP3, WAV, OGG, Opus, FLAC, AAC, MP4, M4A, or MKV within the five-hundred-megabyte limit. Neighboring jobs include Transcribe audio with the Voice API, Format speech-to-text with language and Inverse Text Normalization, and Stream speech to text over WebSocket. More voice jobs live on the Voice hub.
Transcribe with the url field
Export the inference API key outside of source control, then POST multipart form fields that include url instead of file, keeping option fields first and omitting the file part entirely for this path:
export XAI_API_KEY="your_api_key"
curl -X POST https://api.x.ai/v1/stt \
-H "Authorization: Bearer ${XAI_API_KEY}" \
-F model=grok-voice-transcribe-2.0 \
-F format=true \
-F language=en \
-F "keyterm=Understand The Universe" \
-F "url=https://example.com/path/to/recording.mp3"
Read text, duration, and the words array from the JSON response the same way you would after a file upload. The default model is grok-voice-transcribe-2.0 when model is omitted; pin grok-voice-transcribe-1.0 only when you intentionally need the original slug for compatibility.
Optional flags that still apply
Pairing language with format=true enables Inverse Text Normalization so spoken numbers and currency land in written form, and you can repeat the keyterm field for product names up to one hundred terms of fifty characters each. Set diarize=true when you need per-word speaker integers, multichannel=true when discrete channels should return separately, or filler_words=true when uh/um tokens must remain in the transcript. Raw headerless audio still requires audio_format and sample_rate whether the bytes arrive through file or through url.
Errors specific to URL fetch
| Status | Meaning |
|---|---|
400 |
Missing both file and url, bad options, or format=true without language |
413 |
Payload larger than 500 MB after download |
502 |
Server-side URL download failed (bad host, timeout, or non-audio response) |
401 / 429 / 503 |
Auth, rate limit, or temporary backend issues — same as file uploads |
Retry 502 only after you confirm the URL returns audio with a normal HTTP client from your own network, and prefer stable object-storage signed URLs with enough TTL for the STT download to finish.
Pitfalls
Sending both an empty file part and a url can confuse multipart parsers, so use one source per request. Private URLs that require cookies or IP allowlists fail with 502 because the STT service fetches without your browser session. Streaming WebSocket STT still expects binary frames from your client; URL ingest remains a REST-only shortcut for batch jobs already sitting on a CDN.