A 429 from Claude means a limit refused your request, and the fix depends on which limit it was. Check two things before you touch your code: whether the response carries a retry-after header, and what the error message says.
retry-afteris present: you hit a rate limit (requests, input tokens or output tokens per minute). Wait the number of seconds it gives, retry with a little random jitter, and spread your traffic out so it doesn't happen again.- No
retry-after, anderror.details.error_codeisenforced_spend_limit_reached: your organization reached its usage tier's monthly spend cap. Retrying fails until 00:00 UTC on the first day of next month or until Anthropic raises your tier, so stop the retry loop and use Request tier increase in the Console. - The message mentions Claude Code, long context, a gateway or OpenRouter: the limit isn't your API tier at all. The table below matches each message to its fix.
Anthropic's errors reference lists three official reasons for a 429 rate_limit_error: a rate limit, the tier's monthly spend cap, or a spend limit on the Claude Code workspace. The other messages below come from Claude Code subscriptions and from services that sit between you and Claude.
Which Claude 429 do you have? Match the message to the fix
| What you see | Which limit said no | Does retrying help? | Fix |
|---|---|---|---|
rate_limit_error naming the rate limit you exceeded, plus a retry-after header | Requests, input tokens or output tokens per minute for that model | Yes, after retry-after | Wait, retry with jitter, pace requests |
rate_limit_error right after a sharp jump in traffic, while you're under your published limits | Acceleration limit | Yes, once traffic grows more gradually | Ramp traffic up in steps |
"error_code": "enforced_spend_limit_reached", message starts "You have reached your API usage limits", no retry-after | Monthly spend cap of your usage tier | No | Request a tier increase or wait for 00:00 UTC on the 1st |
| HTTP 400 (not 429): "You have reached your specified API usage limits" | A spend limit you set yourself | No | Raise or remove the limit in Billing |
Claude Code on an API key gets a 429 with retry-after | Spend limit on the Claude Code workspace | After retry-after | Wait, or have an admin raise the workspace limit |
API Error: Request rejected (429) · Usage credits are required for long context requests. | Claude subscription using a [1m] model that needs usage credits | No | Turn on usage credits or switch model |
API Error: Request rejected (429) · Rate limited. Please try again later. or {"error":{"message":"Rate limited. Please try again later.","type":"rate_limit_error"}} | Claude Pro/Max login endpoints used by Claude Code | Sometimes, after a pause | Stop polling, wait, check /usage |
429 AI token rate limit exceeded for provider(s): anthropic | Your company's Kong AI Gateway | Only when the gateway window resets | Ask the gateway owner to raise your limit |
Rate limit exceeded: free-models-per-day | OpenRouter's cap on :free models | No, until the daily cap resets | Buy at least 10 credits or use a paid model |
The first five rows come from the Claude API itself; the fourth is a 400 that's easy to mistake for a spend-cap 429. The last four come from Claude Code subscriptions and third-party layers. Each row has its own section below.

Rate limit 429 with retry-after: RPM, ITPM or OTPM exceeded
This is the classic "Claude 429 rate limited" case. The Messages API measures three limits per model class: requests per minute (RPM), input tokens per minute (ITPM) and output tokens per minute (OTPM). Where many providers use one combined tokens-per-minute (TPM) limit, Claude counts input and output tokens separately. Going over any one of them returns a 429 that names the limit you exceeded and carries a retry-after header with the number of seconds to wait, according to the rate limits documentation. A native API error body looks like this:
{
"type": "error",
"error": {
"type": "rate_limit_error",
"message": "<which rate limit you exceeded>"
},
"request_id": "req_..."
}The fix has three parts:
- Honor
retry-after. Anthropic's header description says it plainly: "Earlier retries will fail." Waiting less than the header says only burns another request. - Add jitter. If several workers got the same 429, they also got the same
retry-after. A random extra delay of up to a second keeps them from all retrying in the same instant. - Smooth the traffic that caused it. Limits use a token bucket that refills continuously, and Anthropic warns that limits can be enforced over shorter intervals: 60 RPM may be enforced as 1 request per second. Your per-second budget is roughly the per-minute limit divided by 60, so 1,000 RPM works out to about 16.7 requests per second. A burst of 200 requests in two seconds can trip a 1,000 RPM limit even though your minute total is far below it.
Three facts explain most "but I'm under the limit" surprises. Limits apply to the whole organization, so every API key in it draws from the same pool. A workspace limit can be lower than the organization limit. And each model class has its own bucket, so Opus 5.5 traffic doesn't use up your Sonnet 5.5 allowance, while Opus 4.8, 4.7, 4.6 and 4.5 all share one Opus 4.x bucket.
Spend cap 429: enforced_spend_limit_reached means stop retrying
Each usage tier has a monthly spend cap: $500 on Start, $1,000 on Build and $200,000 on Scale (Custom has none). When your organization reaches it, API usage pauses until 00:00 UTC on the first day of the next month, and every request returns HTTP 429. The rate limits documentation shows this body:
{
"type": "error",
"error": {
"type": "rate_limit_error",
"message": "You have reached your API usage limits: your organization has crossed its monthly API usage threshold, set based on your organization's API tier. You will regain access on 2026-09-01 at 00:00 UTC.",
"details": { "error_code": "enforced_spend_limit_reached" }
},
"request_id": "req_018EeWyXxfu5pfWkrYcMdjWG"
}The error type is the same rate_limit_error as an ordinary rate limit, which is why so much retry code gets stuck here. Two signals tell them apart:
- The spend-cap 429 has no
retry-afterheader. error.details.error_codeisenforced_spend_limit_reachedon the Messages API.
Retrying won't help. That includes the SDK's automatic retries, which fail until access resumes. Your options are to wait for the 1st of the month or to use Request tier increase on the Rate limits page in the Claude Console; moving to a higher tier restores access. Anthropic's Help Center article on rate limits for the Claude API adds that you can request higher limits once you're using at least 50% of your current ones.
Don't confuse this with the spend limit you set yourself under Settings > Billing. That one returns HTTP 400 with invalid_request_error and a message starting "You have reached your specified API usage limits" (or "specified workspace API usage limits"). Raise or remove the limit to get access back.
Acceleration limit 429 after a sudden traffic spike
A 429 can also come from what Anthropic calls acceleration limits. If your organization's usage jumps sharply, for example when a batch job or a launch multiplies traffic in minutes, the API may return 429s even though you're under your published RPM and token limits. Anthropic's advice is to ramp traffic up gradually and keep usage patterns consistent.
Anthropic doesn't publish the thresholds, and there's no public third-party measurement to plan around. In practice, treat it as a ramp problem: start a new workload at a fraction of its target rate, raise the rate in steps, and watch the rate limit headers as you go. If 429s appear only during the ramp and disappear at a steady rate, acceleration was the likely cause.
Claude Code 429 errors: workspace limits, long context and "Rate limited"
A "Claude Code 429" can mean three different things, depending on whether Claude Code runs on an API key or on a Claude Pro or Max login.
Claude Code workspace spend limit returns 429 with retry-after
When Claude Code runs on a Console API key, its usage goes to the Claude Code workspace. Spend limits on that workspace are checked separately from your other workspaces. Requests over the limit can get a 429 with a retry-after header, rather than the 400 that other workspace spend limits return. Wait for the header, or ask a Console admin to raise the Claude Code workspace's limit. Organization rate limits and the tier spend cap still apply on top, so the sections above cover the remaining cases.
"Usage credits are required for long context requests" (429)
The full message reported by Claude Code users is:
API Error: Request rejected (429) · Usage credits are required for long context requests.It's a 429 rate_limit_error, but it has nothing to do with request volume. It appears when a Claude subscription user runs a 1M-context model variant that needs usage credits. Anthropic's Help Center article on context windows on paid Claude plans lists two such cases in Claude Code:
- Opus 4.6 with 1M context (
claude-opus-4-6[1m]): usage credits must be enabled on Pro. - Sonnet 4.6 with 1M context (
claude-sonnet-4-6[1m]): usage credits are required on every plan except usage-based Enterprise.
The same article lists Opus 5.5, Sonnet 5.5, Haiku 5.5 and Fable 5.1 at 1M tokens in Claude Code without that note. The fixes, in order of effort:
- Switch to a standard-context variant with
/model, or drop the[1m]suffix from the model ID. - Pick a model whose 1M window doesn't need credits, such as Sonnet 5.5 or Opus 5.5.
- Turn on usage credits at claude.ai/settings/usage if you want to keep the 4.6 1M variant.
Users also report the error without choosing a [1m] model on purpose. In anthropics/claude-code #61828 (May 23, 2026, Claude Code 2.1.145), Pro and Max users saw a newer wording, "Usage credits required for 1M context · turn on usage credits at claude.ai/settings/usage, or use --model to switch to standard context", on Sonnet 4.6 while their usage was low. One commenter traced it to a Sonnet 1M model in their settings JSON file. In #92456 (September 6, 2026), a user found that ANTHROPIC_DEFAULT_SONNET_MODEL=claude-sonnet-4-6[1M] made a subagent hit this 429. After that, Claude Code limited the whole session to 200k context. If you see the error unexpectedly, check settings.json and your ANTHROPIC_DEFAULT_*_MODEL environment variables for a [1m] ID.
"Rate limited. Please try again later." on Pro and Max
Claude Code users signed in with a Pro or Max plan report a third 429, in two forms:
API Error: Request rejected (429) · Rate limited. Please try again later.{"error":{"message":"Rate limited. Please try again later.","type":"rate_limit_error"}}The raw JSON lacks the top-level "type": "error" and request_id that Messages API errors carry. User reports trace it to the endpoints behind a Claude subscription login, not to API tiers:
- The
/usagecommand. anthropics/claude-code #32503 (March 9, 2026, Max plan) shows/usagefailing with exactly that JSON. The linked report #31637 (March 6, 2026) found that the/api/oauth/usageendpoint returned it after a short burst of polls, sent noRetry-Afterheader, and kept refusing for 30 minutes and more. Usage monitors and status-bar tools that poll this endpoint every 30 to 60 seconds can trigger it. Stop or slow the polling and wait. - Prompts themselves. In #69023 (June 17, 2026), a Pro user got the
Request rejected (429)form while/usageshowed 2% of the session and 44% of the week used. The issue was closed for inactivity without an official explanation.
Subscription usage limits are separate from API usage tiers, so raising an API tier won't change them. Check /usage first; Claude Code /usage: Read Limits, Resets, and Usage Monitors explains how to read it and what to do when its numbers disagree with what Claude Code enforces.
"AI token rate limit exceeded for provider(s): anthropic" is a Kong 429
If your requests go through a company gateway and you get:
429 AI token rate limit exceeded for provider(s): anthropicthe 429 came from the gateway, not from Anthropic. "AI token rate limit exceeded for provider(s): " is the default error_message of Kong AI Gateway's AI Rate Limiting Advanced plugin, and its default error_code is 429. The plugin counts tokens per consumer by default, and other scopes can be configured: consumer group, credential, header, IP, path or service. Limits are set per provider through llm_providers, which is why the message ends with anthropic.
Nothing in your Anthropic Console will change this error. Ask whoever runs the gateway which limit your consumer has for the anthropic provider and how long its window is. Then either get the limit raised or pace your client to fit inside it. Your retry code still helps, but expect no anthropic-ratelimit-* headers on this response.
OpenRouter 429 "free-models-per-day": add 10 credits or go paid
The message starting Rate limit exceeded: free-models-per-day. Add 10 credits to unlock… comes from OpenRouter, not from Anthropic. OpenRouter's limits documentation sets these caps for model IDs ending in :free:
| OpenRouter account | :free requests per minute | :free requests per day |
|---|---|---|
| Fewer than 10 credits purchased (all time) | 20 | 50 |
| At least 10 credits purchased | 20 | 1,000 |
The fix is to buy at least 10 credits, which raises the daily cap to 1,000, or to switch to the paid variant of the model, which has no platform-level request cap. Claude models on OpenRouter are paid models, so if your tool shows this error while you expected Claude, check the model ID: a :free suffix means the request never went to Claude.
OpenRouter can also return a 429 when the upstream provider is rate limiting or at capacity. A negative credit balance produces 402 errors, even on free models.
Retry Claude 429s in Python and TypeScript: retry-after plus full jitter
The official Anthropic SDKs already retry connection errors, 429s and 5xx errors twice by default with exponential backoff, and they honor retry-after. Set max_retries (maxRetries in TypeScript) to change or disable that. For a batch job or a busy service you usually want more control: more attempts, a cap on total wait, and an immediate stop on spend-cap 429s that will never succeed.
The two examples below do that. They turn off the SDK's built-in retries so the attempts don't multiply. On a 429 they use retry-after plus up to one second of jitter. Without the header, they fall back to full-jitter exponential backoff. They raise immediately on enforced_spend_limit_reached, and they retry 5xx errors, including 529, with backoff.

Python retry with the anthropic SDK
import random
import time
import anthropic
client = anthropic.Anthropic(max_retries=0) # this loop owns retries
MAX_ATTEMPTS = 6
BASE_DELAY = 1.0 # seconds
MAX_DELAY = 60.0 # cap for computed backoff
class SpendCapReached(Exception):
"""Monthly tier spend cap: retrying fails until access resumes."""
def is_spend_cap(err: anthropic.RateLimitError) -> bool:
body = err.body if isinstance(err.body, dict) else {}
details = (body.get("error") or {}).get("details") or {}
return details.get("error_code") == "enforced_spend_limit_reached"
def retry_after_seconds(err: anthropic.APIStatusError) -> float | None:
value = err.response.headers.get("retry-after")
try:
return float(value) if value is not None else None
except ValueError:
return None
def full_jitter(attempt: int) -> float:
return random.uniform(0, min(MAX_DELAY, BASE_DELAY * 2 ** attempt))
def create_with_retry(**params) -> anthropic.types.Message:
for attempt in range(MAX_ATTEMPTS):
last = attempt == MAX_ATTEMPTS - 1
try:
return client.messages.create(**params)
except anthropic.RateLimitError as err:
if is_spend_cap(err):
raise SpendCapReached(err.message) from err
if last:
raise
wait = retry_after_seconds(err)
wait = full_jitter(attempt) if wait is None else wait + random.uniform(0, 1)
except anthropic.APIStatusError as err: # 529, 500 and other 5xx
if err.status_code < 500 or last:
raise
wait = full_jitter(attempt)
except anthropic.APIConnectionError:
if last:
raise
wait = full_jitter(attempt)
time.sleep(wait)
raise RuntimeError("unreachable")
message = create_with_retry(
model="claude-sonnet-5-5",
max_tokens=1024,
messages=[{"role": "user", "content": "Hello, Claude"}],
)
print(next(block.text for block in message.content if block.type == "text"))RateLimitError is a subclass of APIStatusError, so it has to be caught first. err.body is the decoded JSON body, which is where details.error_code lives. In the Python SDK, 529 raises OverloadedError and other 5xx codes raise InternalServerError, and both are caught by the status_code >= 500 branch.
TypeScript retry with @anthropic-ai/sdk
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({ maxRetries: 0 }); // this loop owns retries
const MAX_ATTEMPTS = 6;
const BASE_DELAY_MS = 1_000;
const MAX_DELAY_MS = 60_000;
export class SpendCapReached extends Error {}
const sleep = (ms: number) => new Promise((resolve) => setTimeout(resolve, ms));
const fullJitter = (attempt: number) =>
Math.random() * Math.min(MAX_DELAY_MS, BASE_DELAY_MS * 2 ** attempt);
type ErrorBody = { error?: { details?: { error_code?: string } } };
function isSpendCap(err: InstanceType<typeof Anthropic.RateLimitError>): boolean {
const body = err.error as ErrorBody | undefined;
return body?.error?.details?.error_code === "enforced_spend_limit_reached";
}
function retryAfterMs(err: InstanceType<typeof Anthropic.APIError>): number | null {
const value = err.headers?.get("retry-after");
const seconds = value == null ? Number.NaN : Number(value);
return Number.isFinite(seconds) ? seconds * 1000 : null;
}
export async function createWithRetry(
params: Anthropic.MessageCreateParamsNonStreaming,
): Promise<Anthropic.Message> {
for (let attempt = 0; ; attempt++) {
const last = attempt >= MAX_ATTEMPTS - 1;
try {
return await client.messages.create(params);
} catch (err) {
if (err instanceof Anthropic.RateLimitError) {
if (isSpendCap(err)) throw new SpendCapReached(err.message);
if (last) throw err;
const wait = retryAfterMs(err);
await sleep(wait === null ? fullJitter(attempt) : wait + Math.random() * 1000);
} else if (err instanceof Anthropic.APIConnectionError && !last) {
await sleep(fullJitter(attempt));
} else if (err instanceof Anthropic.APIError && (err.status ?? 0) >= 500 && !last) {
await sleep(fullJitter(attempt));
} else {
throw err;
}
}
}
}
const message = await createWithRetry({
model: "claude-sonnet-5-5",
max_tokens: 1024,
messages: [{ role: "user", content: "Hello, Claude" }],
});
console.log(message.content);In current versions of the TypeScript SDK, err.headers is a standard Headers object, so read it with .get("retry-after"). err.error holds the parsed JSON body, and every status of 500 and above, including 529, maps to InternalServerError.
If you'd rather keep the SDK's built-in retries, leave max_retries at its default or raise it, and add only the spend-cap check around the call. The SDK will still try a spend-cap 429 two more times before raising. That only delays the failure by a few seconds, and your code then stops on the enforced_spend_limit_reached code instead of looping.
Why full jitter instead of a fixed delay
With plain exponential backoff, every client that failed at the same moment waits 1, 2, 4, 8 seconds in lockstep and hits the limit together again. Full jitter picks a random wait between zero and the current backoff ceiling (random(0, min(cap, base × 2^attempt))). The retries spread across the whole window, so the bucket refills between them. The ceiling still grows exponentially, which keeps a persistent problem from turning into a tight retry loop. When retry-after is present, use it as the floor and add only a small random spread on top. Retrying earlier than the header allows fails by definition.
Streaming needs one more check. An error can arrive after the API has already returned HTTP 200, as a server-sent error event in the stream, so status-code retry logic around the initial call won't see it. Handle error events in your stream consumer as well.
Claude rate limit headers: see the 429 coming
Every Messages API response carries headers that show the limit being enforced, what's left and when it refills, per the response headers table:
| Header | What it tells you |
|---|---|
retry-after | Seconds to wait before retrying; earlier retries fail. Not sent with the spend-cap 429 |
anthropic-ratelimit-requests-limit / -remaining / -reset | Request limit, requests left, and when the bucket is full again |
anthropic-ratelimit-tokens-limit / -remaining / -reset | Values for the most restrictive token limit in effect right now |
anthropic-ratelimit-input-tokens-limit / -remaining / -reset | Input token (ITPM) limit, remaining and refill time |
anthropic-ratelimit-output-tokens-limit / -remaining / -reset | Output token (OTPM) limit, remaining and refill time |
anthropic-priority-input-tokens-* / anthropic-priority-output-tokens-* | Same fields for Priority Tier capacity (Priority Tier only) |
Reset times are RFC 3339 timestamps, and token "remaining" values are rounded to the nearest thousand. If a workspace limit is the tightest one, the tokens-* headers show the workspace values. The anthropic-workspace-id header tells you which workspace a request counted against.
The SDKs return parsed objects by default, so ask for the raw response to read headers:
raw = client.messages.with_raw_response.create(
model="claude-sonnet-5-5",
max_tokens=1024,
messages=[{"role": "user", "content": "ping"}],
)
h = raw.headers
for kind in ("requests", "input-tokens", "output-tokens"):
print(
f"{kind}: {h.get(f'anthropic-ratelimit-{kind}-remaining')} "
f"of {h.get(f'anthropic-ratelimit-{kind}-limit')}, "
f"full at {h.get(f'anthropic-ratelimit-{kind}-reset')}"
)
message = raw.parse()In TypeScript, .withResponse() on the create call returns { data, response }, and response.headers.get(...) reads the same fields. A simple guard is to slow down or queue work when remaining drops below about 10% of limit. The SDK-level retries then become a backstop rather than your main control.
For teams with many workers, a shared pacer keeps bursts under the per-second budget. This asyncio sketch spaces requests to the RPM limit divided by 60 per second:
import asyncio
import time
class Pacer:
"""Spaces calls evenly so bursts stay under rpm / 60 per second."""
def __init__(self, rpm: int, safety: float = 0.8):
self.interval = 60.0 / (rpm * safety)
self.next_slot = time.monotonic()
self.lock = asyncio.Lock()
async def wait(self) -> None:
async with self.lock:
now = time.monotonic()
delay = max(0.0, self.next_slot - now)
self.next_slot = max(now, self.next_slot) + self.interval
await asyncio.sleep(delay)
pacer = Pacer(rpm=1_000) # Start-tier RPM for standard models
# before each request: await pacer.wait()At 1,000 RPM and a safety factor of 0.8, the interval is 60 / 800 = 0.075 seconds, or about 13 requests per second. It paces requests only. If your prompts are large, ITPM will be the tighter limit, so size the pacer from tokens per request as well.
Claude 429 vs 529 vs 500: which errors to retry and how
| Status and type | What it means | Retry? | Where to go next |
|---|---|---|---|
429 rate_limit_error, with retry-after | Your organization or workspace hit a rate limit | Yes, after retry-after, with jitter | This page |
429 rate_limit_error, enforced_spend_limit_reached | Monthly tier spend cap reached | No | Request tier increase |
529 overloaded_error | The API is temporarily overloaded across all users | Yes, with exponential backoff | Claude 529 Overloaded Error: What It Means and How to Fix It (2026) |
500 api_error | Unexpected error inside Anthropic's systems | Yes, with backoff; contact support with the request ID if it persists | Claude API Error 500: Fix api_error Without Blind Retries |
400 invalid_request_error, "specified API usage limits" | Spend limit you set yourself | No | Raise the limit in Billing |
413 request_too_large | Request larger than 32 MB on the Messages API | No | Shrink the request |
The practical difference: a 429 is about your usage, so pacing, caching and tier changes fix it. A 529 is about Anthropic's capacity, so nothing in your configuration prevents it and only waiting helps. The errors documentation adds one overlap: a sharp increase in your own traffic can produce 429s from acceleration limits rather than 529s. Always log the request_id from the error body or the request-id header; support needs it for any 500 or unexplained 429.
How to prevent Claude 429 errors: pacing, caching, Batch and tier increases
- Pace instead of bursting. Size concurrency so requests stay near the per-minute limit divided by 60 per second, and ramp new workloads up gradually to stay clear of acceleration limits.
- Cache repeated input. For most current models, only uncached input counts toward ITPM:
input_tokenspluscache_creation_input_tokenscount, andcache_read_input_tokensdon't. Claude Haiku 3.5 is the exception. Anthropic's own example: with a 2,000,000 ITPM limit and an 80% cache hit rate, you can process 10,000,000 total input tokens per minute, because 2,000,000 ÷ (1 − 0.8) = 10,000,000. How to Use Prompt Caching in Claude API: Complete 2026 Guide with Code Examples shows how to set the cache breakpoints. - Don't shrink
max_tokensto dodge OTPM. OTPM counts only the output tokens actually generated.max_tokensdoesn't factor into the calculation, so lowering it changes nothing for rate limits. - Move offline work to the Message Batches API. Batches have their own limits, shared across models: 1,000 RPM and 200,000 queued batch requests on Start, 2,000 and 300,000 on Build, 4,000 and 500,000 on Scale, with up to 100,000 requests per batch. Batch processing is also billed at 50% of standard prices; see Claude API Pricing Per Million Tokens (October 2026): $0.10–$50.
- Spread load across model classes. Limits are per model class, so routing simple tasks to Haiku 5.5 frees Sonnet 5.5 or Opus 5.5 capacity. Opus 4.x and Sonnet 4.x versions share buckets within their family, so that only works across classes.
- Fence off noisy projects with workspace limits. A workspace can be capped below the organization limit so one job can't starve the rest. The organization limit always applies on top.
- Request a tier increase before you need it. Once you're using at least 50% of your limits, the Console's Request tier increase button is the documented path to both higher rate limits and a higher spend cap.
- Watch the Console's rate limit charts. The Usage page plots hourly peak uncached input tokens per minute and output tokens per minute against your limit, plus your cache rate.
Claude API tiers in brief: Start, Build and Scale limits
Anthropic places organizations on usage tiers automatically, based on usage history and account standing. These are the standard limits for the Messages API as of October 8, 2026, from the rate limits documentation:
| Tier | Monthly spend cap | Standard models: RPM / ITPM / OTPM | Fable 5.x: RPM / ITPM / OTPM |
|---|---|---|---|
| Start | $500 | 1,000 / 2,000,000 / 400,000 | 1,000 / 500,000 / 100,000 |
| Build | $1,000 | 5,000 / 5,000,000 / 1,000,000 | 2,000 / 1,500,000 / 300,000 |
| Scale | $200,000 | 10,000 / 10,000,000 / 2,000,000 | 4,000 / 4,000,000 / 800,000 |
| Custom | None | Arranged with Anthropic sales | Arranged with Anthropic sales |
"Standard models" covers Opus 5.5, Opus 5, Opus 4.x, Sonnet 5.5, Sonnet 5, Sonnet 4.x, Haiku 5.5 and Haiku 4.5, each class with its own bucket. New organizations and those with little usage history may start in an Evaluation tier with lower limits that Anthropic doesn't publish. If your Console shows numbers below the Start row, that's why early 429s appear.
The old Tier 1 to Tier 4 system, where a $5 to $400 credit purchase unlocked limits as low as 50 RPM, no longer exists. Anthropic's release note of June 26, 2026 consolidated tiers into Start, Build and Scale, said most organizations moved to a higher tier and none received lower limits. Old Tier 1–4 numbers in libraries, forum posts and screenshots are obsolete. Per-model detail, reported old-to-new tier mappings and Evaluation-tier tells are in Claude API Quota Tiers and Limits Explained: Start, Build, Scale.
FAQ about Claude API 429 errors
What does the 429 error code mean in the Claude API?
HTTP 429 is "Too Many Requests." In the Claude API it comes with the error type rate_limit_error and means your organization hit a rate limit, reached its usage tier's monthly spend cap, or reached a spend limit on the Claude Code workspace. The retry-after header and error.details.error_code tell you which.
How do I fix API error 429 from Claude?
Read the response first. With a retry-after header, wait that long, retry with jitter and pace your traffic. With enforced_spend_limit_reached and no header, stop retrying and request a tier increase or wait for 00:00 UTC on the 1st. With a Claude Code, Kong or OpenRouter message, fix it at that layer as described in the matching section above.
Why is Claude saying rate limit exceeded when my usage looks low?
The most common causes are bursts, because limits can be enforced per second rather than per minute. Other API keys or workspaces in the same organization share your pool, and a workspace may have a lower limit of its own. On a new account, the Evaluation tier may sit below the published Start numbers. In Claude Code on a Pro or Max plan, a 429 can come from a [1m] model that needs usage credits, or from subscription endpoints, rather than from your usage.
How long should I wait after a Claude 429 error?
Exactly as long as retry-after says, plus a small random delay. Without the header and without a spend-cap code, use full-jitter exponential backoff starting around one second and capped around a minute. A spend-cap 429 lasts until 00:00 UTC on the first day of next month unless your tier is raised.
Do more API keys raise my Claude rate limit?
No. Limits are set at the organization level, so every key in the organization draws from the same buckets. Separate workspaces let you cap a project lower, not higher, than the organization limit.
Does the Anthropic SDK retry 429 errors automatically?
Yes. The official SDKs retry connection errors, rate limits and 5xx errors twice by default with exponential backoff and honor retry-after. Change this with max_retries (maxRetries in TypeScript). They also retry the spend-cap 429, which can't succeed, so add a check for enforced_spend_limit_reached.
Does max_tokens count against Claude's output token limit?
No. OTPM is evaluated in real time on the output tokens actually generated, and max_tokens doesn't factor into it. A high max_tokens has no rate limit downside.
Is "Claude Code API Error: rate limit reached" the same as an API 429?
Only when Claude Code runs on an API key, in which case your organization's API limits and the Claude Code workspace's spend limit apply. On a Pro or Max login, Claude Code uses your plan's usage limits, which are separate from API tiers. Check /usage before changing anything in the Console.



