
Use the US regional API endpoint
Point your client at https://us.api.x.ai/v1 when you need API request handling and model inference to happen in the United States. Official Regional Endpoints explains that the global host https://api.x.ai may route between regions for capacity, so processing location is not guaranteed there. The US endpoint is available to every team, and your existing API keys work on both hosts. Today it serves grok-4.7 and grok-4.6 only — image generation, video generation, and voice APIs stay on the global endpoint. Token usage costs 10% more than global rates. Create or reuse a key at console.x.ai.
What you need
An XAI_API_KEY for your team, a client that can set a custom base URL or api_host, and a workload that fits the US model list. Confirm current availability with GET https://us.api.x.ai/v1/models or the models page in the Console before you cut traffic over. Keep Imagine and Voice calls on the global endpoint. Pair this job with Call Grok 4.7 via API when you pick the model id, and with Enable Zero Data Retention on the xAI Console when retention is a separate compliance control. Browse neighboring API jobs on the API hub.
Point the client at us.api.x.ai
- Export the key locally and keep it out of chat logs and public repos:
export XAI_API_KEY="your_api_key"
- Call the Responses API on the US host with curl:
curl https://us.api.x.ai/v1/responses \
-H "Authorization: Bearer $XAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "grok-4.7",
"input": "Explain latency versus throughput in two sentences."
}'
- Or set the Python SDK
api_hostto the bare host (nohttps://):
import os
from xai_sdk import Client
from xai_sdk.chat import user
client = Client(
api_key=os.getenv("XAI_API_KEY"),
api_host="us.api.x.ai",
)
chat = client.chat.create(model="grok-4.7")
chat.append(user("Explain latency versus throughput in two sentences."))
print(chat.sample().content)
- OpenAI-compatible clients use
base_url="https://us.api.x.ai/v1". Requesting a model that is not on the US list, includinggrok-latest, returns404 Not Found— confirm the endpoint before you troubleshoot model ACLs.
What the US guarantee covers
When you call https://us.api.x.ai/v1, xAI guarantees that request handling, model inference, safety moderation, and retained request data stay in the United States. Files, Collections, and server-side tools such as web search, X search, and code execution still work on the US endpoint, but they sit outside that guarantee and may process data elsewhere. The network path from your systems to xAI is also outside the guarantee. Prompt cache hits are not guaranteed across endpoints, so keep each conversation on one host. Zero Data Retention is a team-level setting and applies on both endpoints once enabled.
Pitfalls
Sending Imagine or Voice traffic to us.api.x.ai fails because those APIs are not served there. Treating the regional endpoint alone as a full data-residency contract overstates the guarantee — contact sales@x.ai when contracts require more. Mixing the same conversation across global and US hosts loses prompt-cache locality. Ignoring the 10% premium on input, output, and cached tokens surprises the bill after cutover.