
Use the US regional xAI API endpoint
Point your client at https://us.api.x.ai/v1 when you need API request handling, inference, moderation, and retained request data to stay in the United States instead of the global https://api.x.ai host that may route between regions. Official Regional Endpoints documents that every team can use the US host with existing API keys, that token usage costs 10% more than global, and that the US host currently serves grok-4.7 and grok-4.6 only. Image, video, and voice APIs stay on the global endpoint. Create a key and load credits at console.x.ai.
What you need
An XAI_API_KEY that already works on the global host, a client that can override the base URL or SDK api_host, and a workload that fits the US model list. Neighboring jobs include List models via the xAI API against the US host’s GET /v1/models, Enable Zero Data Retention in the xAI Console when retention policy is separate from residency, and Chain responses with previous id for multi-turn work kept on one endpoint. More API jobs live on the API hub.
Point the client at the US host
- Export the same inference key you use globally:
export XAI_API_KEY="your_api_key"
- Call Responses on the US base URL:
curl https://us.api.x.ai/v1/responses \
-H "Authorization: Bearer $XAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "grok-4.7",
"input": "Explain latency versus throughput in two sentences."
}'
- OpenAI SDK and Vercel AI SDK clients set
baseURL/base_urltohttps://us.api.x.ai/v1. The xAI Python SDK takes the bare host:
import os
from xai_sdk import Client
from xai_sdk.chat import user
client = Client(
api_key=os.getenv("XAI_API_KEY"),
api_host="us.api.x.ai",
)
chat = client.chat.create(model="grok-4.7")
chat.append(user("Explain latency versus throughput in two sentences."))
print(chat.sample().content)
- Confirm availability with
GET https://us.api.x.ai/v1/modelsbefore hard-coding a slug. A model that exists globally but not on the US host returns404with a not-found body that can look like an ACL miss — check the host first.
What the guarantee covers
US handling covers API servers, inference, safety moderation, and retained request metadata/inputs/outputs described in the Security FAQ. Files, Collections, and server-side tools such as web search, X search, and code execution may still process data outside the United States even when you call the US host. The network path from your systems to SpaceXAI is outside the guarantee. Prompt cache hits are not guaranteed across endpoints, so keep each conversation on one host and set a prompt_cache_key when you need routing affinity.
Pitfalls
Requesting Imagine, video, or voice routes on us.api.x.ai fails because those APIs are not served there. Treating the regional endpoint alone as a full contractual data-residency package overshoots what the docs guarantee — contact sales@x.ai when contracts require more. Mixing turns of one conversation across global and US hosts breaks prompt-cache affinity and can surprise latency.