VOICE / transcribe-raw-pcm-via-stt

Voice

Transcribe raw PCM audio via speech-to-text

Send headerless PCM (or µ-law / A-law) bytes to the batch STT endpoint when your pipeline already has raw samples and no WAV/MP3 container. Official Speech to Text lists raw formats under Supported Audio Formats: for pcm, mulaw, and alaw you must pass audio_format plus sample_rate (and channels when multichannel raw audio needs it). Container formats such as WAV and MP3 are auto-detected — do not set audio_format for those. Create a key and load credits at console.x.ai.

What you need

An XAI_API_KEY, a raw audio file or buffer (signed 16-bit little-endian PCM is the common case), and the sample rate those samples were recorded at (8000, 16000, 22050, 24000, 44100, or 48000). Neighboring jobs include Transcribe audio from a URL via STT when the asset is already hosted, Stream speech to text over WebSocket for live PCM frames, and Use STT language and format for Inverse Text Normalization. More voice jobs live on the Voice hub.

Batch POST with audio_format and sample_rate

  1. Export the key, then POST multipart form fields before file. Put audio_format=pcm and the matching sample_rate on the request; omit those fields when the file is a container:
export XAI_API_KEY="your_api_key"

curl -X POST https://api.x.ai/v1/stt \
  -H "Authorization: Bearer $XAI_API_KEY" \
  -F model=grok-voice-transcribe-2.0 \
  -F audio_format=pcm \
  -F sample_rate=16000 \
  -F language=en \
  -F format=true \
  -F file=@utterance.pcm
  1. In Python, send the same fields with requests. Keep option fields ahead of the file part so streamable uploads do not drop trailing form values:
import os
import requests

with open("utterance.pcm", "rb") as audio:
    response = requests.post(
        "https://api.x.ai/v1/stt",
        headers={"Authorization": f"Bearer {os.environ['XAI_API_KEY']}"},
        files={"file": ("utterance.pcm", audio, "application/octet-stream")},
        data=[
            ("model", "grok-voice-transcribe-2.0"),
            ("audio_format", "pcm"),
            ("sample_rate", "16000"),
            ("language", "en"),
            ("format", "true"),
        ],
    )
response.raise_for_status()
result = response.json()
print(result["text"])
print(f"Duration: {result['duration']}s")
for word in result.get("words", []):
    print(f"  {word['start']:.2f}s - {word['end']:.2f}s: {word['text']}")
  1. For telephony µ-law or A-law dumps, set audio_format=mulaw or audio_format=alaw with the matching sample_rate (often 8000). PCM is signed 16-bit little-endian at 2 bytes per sample; µ-law and A-law are 1 byte per sample.

Multichannel raw audio

When the buffer interleaves more than one channel and you need per-channel transcripts, set multichannel=true and pass channels (2–8). Container formats auto-detect channel count; raw PCM needs the explicit channels field. Streaming STT can also ingest interleaved PCM over the WebSocket with encoding=pcm&multichannel=true&channels=2 — see the stream how-to for that path.

Pitfalls

Missing sample_rate on raw audio returns 400. Setting audio_format=pcm on a WAV/MP3 file fights auto-detection — leave audio_format unset for containers. format=true without language is rejected. Max upload size is 500 MB. Streaming WebSocket PCM uses query params (encoding / sample_rate) instead of multipart fields; do not confuse the two APIs when debugging.