
Transcribe raw PCM audio via speech-to-text
Send headerless PCM (or µ-law / A-law) bytes to the batch STT endpoint when your pipeline already has raw samples and no WAV/MP3 container. Official Speech to Text lists raw formats under Supported Audio Formats: for pcm, mulaw, and alaw you must pass audio_format plus sample_rate (and channels when multichannel raw audio needs it). Container formats such as WAV and MP3 are auto-detected — do not set audio_format for those. Create a key and load credits at console.x.ai.
What you need
An XAI_API_KEY, a raw audio file or buffer (signed 16-bit little-endian PCM is the common case), and the sample rate those samples were recorded at (8000, 16000, 22050, 24000, 44100, or 48000). Neighboring jobs include Transcribe audio from a URL via STT when the asset is already hosted, Stream speech to text over WebSocket for live PCM frames, and Use STT language and format for Inverse Text Normalization. More voice jobs live on the Voice hub.
Batch POST with audio_format and sample_rate
- Export the key, then POST multipart form fields before
file. Putaudio_format=pcmand the matchingsample_rateon the request; omit those fields when the file is a container:
export XAI_API_KEY="your_api_key"
curl -X POST https://api.x.ai/v1/stt \
-H "Authorization: Bearer $XAI_API_KEY" \
-F model=grok-voice-transcribe-2.0 \
-F audio_format=pcm \
-F sample_rate=16000 \
-F language=en \
-F format=true \
-F file=@utterance.pcm
- In Python, send the same fields with
requests. Keep option fields ahead of the file part so streamable uploads do not drop trailing form values:
import os
import requests
with open("utterance.pcm", "rb") as audio:
response = requests.post(
"https://api.x.ai/v1/stt",
headers={"Authorization": f"Bearer {os.environ['XAI_API_KEY']}"},
files={"file": ("utterance.pcm", audio, "application/octet-stream")},
data=[
("model", "grok-voice-transcribe-2.0"),
("audio_format", "pcm"),
("sample_rate", "16000"),
("language", "en"),
("format", "true"),
],
)
response.raise_for_status()
result = response.json()
print(result["text"])
print(f"Duration: {result['duration']}s")
for word in result.get("words", []):
print(f" {word['start']:.2f}s - {word['end']:.2f}s: {word['text']}")
- For telephony µ-law or A-law dumps, set
audio_format=mulaworaudio_format=alawwith the matchingsample_rate(often8000). PCM is signed 16-bit little-endian at 2 bytes per sample; µ-law and A-law are 1 byte per sample.
Multichannel raw audio
When the buffer interleaves more than one channel and you need per-channel transcripts, set multichannel=true and pass channels (2–8). Container formats auto-detect channel count; raw PCM needs the explicit channels field. Streaming STT can also ingest interleaved PCM over the WebSocket with encoding=pcm&multichannel=true&channels=2 — see the stream how-to for that path.
Pitfalls
Missing sample_rate on raw audio returns 400. Setting audio_format=pcm on a WAV/MP3 file fights auto-detection — leave audio_format unset for containers. format=true without language is rejected. Max upload size is 500 MB. Streaming WebSocket PCM uses query params (encoding / sample_rate) instead of multipart fields; do not confuse the two APIs when debugging.