Key Information

Rate Limits

View as Markdown

Every xAI API team has per-model rate limits on two dimensions: requests per second (RPS) and tokens per minute (TPM). Your per-second limit is derived from your per-minute request budget (RPM / 60): you cannot spend a full minute's requests in a single second, which protects the API from sudden bursts. These limits scale with your team's tier, which is determined by cumulative spend on the API.

You can view your team's current tier and per-model limits on the Models page in the xAI Console.


Rate limit tiers

Your tier is based on cumulative spend on the xAI API since January 1, 2026. Tiers unlock automatically as your spend increases.

TierSpend threshold
Tier 0$0 (default)
Tier 1$50
Tier 2$250
Tier 3$1,000
Tier 4$5,000
EnterpriseAvailable on request

Qualification is based on total revenue received through prepaid credit purchases or successfully fulfilled invoices. Once you qualify for a tier, you stay there permanently; tiers never downgrade.


Per-model limits

Each tier sets hard RPS and TPM caps per model. Limits scale exponentially with tier. Exceeding any limit returns a 429 Too Many Requests error.

The table below lists RPS and TPM limits at each tier for every model. Voice & Audio endpoints follow, limited by requests per second and concurrent sessions (CST) rather than tokens; Speech to Speech is limited by concurrent sessions alone. You can also view your team's personalized limits on the Models page in the xAI Console.

ModelTierRequests per secondTokens per minute
Language Models
grok-4.7T0T1T2T3T415017220831250050M53M60M74M100M
grok-4.6T0T1T2T3T415017220831250050M53M60M74M100M
grok-4.5T0T1T2T3T415017220831250050M53M60M74M100M
grok-4.3T0T1T2T3T437507512520810M15M25M45M85M
grok-4.20-0309-reasoningT0T1T2T3T437507512520810M15M25M45M85M
grok-4.20-0309-non-reasoningT0T1T2T3T437507512520810M15M25M45M85M
grok-build-0.1T0T1T2T3T437507512520810M15M25M45M85M
grok-4.20-multi-agent-0309T0T1T2T3T49121831562.5M3.7M6.2M11M21M
Image Generation
grok-imagine-image-qualityT0T1T2T3T46122550100
grok-imagine-image-2.0T0T1T2T3T46122550100
grok-imagine-imageT0T1T2T3T46122550100
Video Generation
grok-imagine-video-1.5T0T1T2T3T410203979158
grok-imagine-videoT0T1T2T3T410203979158

Voice & Audio

ModelRPSConcurrent sessions
grok-voice-think-fast-2.0T0: 10, T1: 20, T2: 50, T3: 100, T4: 200
Text to SpeechT0: 50, T1: 50, T2: 100, T3: 250, T4: 500T0: 100, T1: 200, T2: 200, T3: 300, T4: 500
Speech to TextT0: 10, T1: 10, T2: 20, T3: 30, T4: 40T0: 100, T1: 200, T2: 200, T3: 300, T4: 500

What counts toward TPM

All tokens consumed by a request count toward the TPM limit for that model:

  • Prompt tokens (text, image, and audio)

  • Completion tokens

  • Reasoning tokens (on reasoning models)

  • Cached prompt tokens (still count toward TPM, though they are billed at a reduced rate)

For details on how tokens are counted and priced, see Models and Pricing. For per-request cost tracking, see Cost Tracking.


Handling rate limit errors

When you exceed your rate limit, the API returns HTTP 429. Implement exponential backoff to handle this gracefully:

import os
import time
from openai import OpenAI, RateLimitError

client = OpenAI(base_url="https://api.x.ai/v1", api_key=os.getenv("XAI_API_KEY"))

def request_with_backoff(messages, max_retries=5):
    for attempt in range(max_retries):
        try:
            return client.chat.completions.create(
                model="grok-4.7",
                messages=messages,
            )
        except RateLimitError:
            wait = 2 ** attempt
            time.sleep(wait)
    raise RateLimitError("Max retries exceeded")

Increasing your limits

  • Spend more. Tiers upgrade automatically based on cumulative spend. No action required on your part.

  • Request an increase. Submit a request through the xAI Console if you need higher limits without additional spend, or limits beyond Tier 4.

  • Contact sales. For enterprise-grade capacity, please email sales@x.ai.


Last updated:September 17, 2026