VOICE / handle-dtmf-on-sip-calls

Voice

Handle DTMF on SIP calls

Capture phone keypad presses on a SIP Speech to Speech session when callers enter account digits, choose a menu option, or submit a PIN with #. Official SIP Phone Calls documents DTMF under “DTMF phone keypresses”: tones are buffered server-side, flushed to the model as text input on documented triggers, and mirrored to your client as input_audio_buffer.dtmf_event_received audit events. Voice REST lists the same event on the realtime WebSocket and notes DTMF is SIP-only, so direct WebSocket sessions do not emit it.

What you need

You need an inbound SIP call already joined with wss://api.x.ai/v1/realtime?call_id=… after a verified realtime.call.incoming webhook, plus a client that reads WebSocket JSON events while the call is live. Neighboring jobs include Handle SIP phone calls with Speech to Speech for register → webhook → join, Hang up a SIP call when a digit path should end the call, and Transfer a SIP call with refer when a keypress should escalate to a human; more voice jobs live on the Voice hub.

Watch the audit events

Each keypress arrives on the client WebSocket as an audit trail while the server also buffers digits for the model:

{
  "type": "input_audio_buffer.dtmf_event_received",
  "event": "5",
  "received_at": 1730000000
}

Log or display event (the digit or symbol) and received_at for support tooling, and do not invent a separate DTMF REST endpoint because the SIP guide and Voice REST treat the WebSocket event as the client-facing surface. Direct realtime connections without call_id never see these events, which is the usual failure mode when you test menus on a browser-only socket.

Know when digits flush to the model

Buffered digits are submitted to the model when any of the following occurs, per the SIP guide: the user presses # (submit key), 2.5 seconds of idle time pass after the last keypress, or the user begins speaking and speech preempts the digit buffer. Design menus so callers know to press # after a PIN, or accept the 2.5s idle flush for single-digit trees. If the caller starts talking mid-entry, expect the buffer to flush early and prompt them to re-enter if your agent needs a complete code, and pair agent instructions with these flush rules so the model treats flushed text as keypad input rather than spoken words.

Pitfalls

Expecting DTMF on a browser-only Speech to Speech socket fails because docs limit DTMF to SIP sessions. Parsing spoken “five” as a substitute for the dtmf_event_received path mixes modalities and loses the audit trail. Assuming every digit becomes a model turn immediately ignores the buffer until #, idle, or speech. Closing the WebSocket before logging events loses keypad history for that call even though hangup or refer may still run from your backend.