Skip to content

Claude API Quota Tiers and Limits Explained: Start, Build, Scale

Anthropic assigns usage tiers automatically: Start caps monthly spend at $500, Build at $1,000, Scale at $200,000. Cache reads don't count toward ITPM.

•••21 min read•API Guides
Isometric stacks comparing Claude API standard-model RPM by usage tier: Start 1,000, Build 5,000, Scale 10,000

As of October 8, 2026, the Claude API has three standard usage tiers (Start, Build and Scale) plus a Custom tier arranged with Anthropic's account team. You don't buy a tier with a deposit. Anthropic places your organization on one automatically based on usage history and account standing, and moves it up over time. New organizations may begin in a lower Evaluation tier first. Each tier has a monthly spend cap ($500 for Start, $1,000 for Build, $200,000 for Scale, none for Custom) and per-model rate limits measured in RPM, ITPM and OTPM. Your current tier is on the Rate limits page of the Claude Console. Source: Anthropic's rate limits documentation.

Tier 1, Tier 2, Tier 3 and Tier 4 are the old names. Anthropic replaced them on June 26, 2026 and said most organizations moved to a higher tier and none got lower limits, but it hasn't published a tier-by-tier mapping. Third-party reports suggest one, covered in the section on the old names.

What Are Claude API Quotas and Limits? Spend Caps and Rate Limits

Two kinds of limits stop Claude API requests, and they fail differently:

  • Spend limits cap what your organization can spend on the API in a calendar month. Every tier except Custom has one, and you can also set your own lower limit.
  • Rate limits cap how fast you can send work: requests per minute (RPM), input tokens per minute (ITPM) and output tokens per minute (OTPM), counted separately for each model class.

Both are enforced at the organization level, not per API key. Every key, teammate and app in one organization draws from the same pool. Inside an organization, you can give individual workspaces lower limits so one project can't starve the others (more on that below).

The published numbers are maximums, not guarantees. Anthropic notes that limits can be enforced over shorter intervals than a minute. A 60 RPM limit, for example, may behave like 1 request per second, so a tight burst can trigger a 429 even when your per-minute total looks fine.

Claude API Usage Tiers: Start, Build, Scale and Custom

Usage tierHow you get thereMonthly spend capStandard-model RPM / ITPM / OTPM
EvaluationPossible starting point for new or low-history organizationsNot publishedBelow the Start limits, not published
StartAutomatic placement$5001,000 / 2,000,000 / 400,000
BuildAutomatic, as usage history grows$1,0005,000 / 5,000,000 / 1,000,000
ScaleAutomatic, as usage history grows$200,00010,000 / 10,000,000 / 2,000,000
CustomArranged with the account teamNo capNegotiated

"Standard model" here means every current model class except Fable 5.x, which has lower limits (see the per-model table). Figures are from Anthropic's rate limits documentation as of October 8, 2026.

Monthly spend caps: $500, $1,000 and $200,000

The spend cap is the most your organization can spend on the API in one calendar month. When you reach it, usage pauses until 00:00 UTC on the first day of the next month, unless Anthropic grants a higher limit sooner. Moving to a higher tier restores access.

Separately, you can set your own spend limit under Settings > Billing > Spend limits in the Console. It can be lower than your tier's cap but never higher. It's the right tool if you want a budget alarm that stops traffic at, say, $200 on a Start-tier account.

Evaluation tier: lower starting limits for new organizations

Anthropic says new organizations and organizations with limited usage history may start in an Evaluation tier whose limits sit below the standard Start numbers. The limits rise automatically as the organization builds history. Anthropic doesn't publish the Evaluation numbers, and no third-party source found for this guide lists them either. The practical tell, as OfoxAI's 429 guide puts it: if the Rate limits page in the Console shows numbers below the Start row (1,000 RPM and 2,000,000 ITPM on standard models), your organization is still in Evaluation. Check that page before debugging your code when a brand-new account hits 429s early.

Where to see your current tier: Console and Rate Limits API

  • The Rate limits page in the Claude Console shows your tier and current limits per model.
  • The Billing page shows your monthly spend cap and any spend limit you set.
  • The Rate Limits API returns organization and workspace limits programmatically, which helps if you want to size worker pools from real values instead of hard-coded ones.

Anthropic doesn't publish the exact criteria or thresholds for moving from Start to Build or from Build to Scale. The documented rule is automatic movement based on usage history and account standing, plus the option to request an increase. Anthropic's Help Center adds that you can request higher limits in the Console once you're using at least 50% of your current limits. A ClaudeDevs post on X (July 2, 2026, as quoted by AI Catchup) put it more bluntly: tiers are "no longer based on API spend." Topping up credits won't move you up the way the old $40/$200/$400 thresholds did.

What Happened to Claude API Tier 1, Tier 2, Tier 3 and Tier 4?

Earlier versions of Anthropic's documentation described four numbered tiers that you unlocked by buying credits:

Old tier (no longer documented)Cumulative credit purchaseMonthly spend limit
Tier 1$5$100
Tier 2$40$500
Tier 3$200$1,000
Tier 4$400$5,000

As of October 8, 2026, that table and its deposit thresholds are gone from the rate limits documentation. Tiers are now named Start, Build, Scale and Custom, and placement is based on usage history and account standing rather than on how much you prepaid.

What this means in practice:

  • "Claude Tier 4" isn't a current tier. If a library, forum post or old guide says "you need Tier 4," read it as "you need higher limits." Then check what your organization actually has on the Rate limits page.
  • There is no official old-to-new mapping, only the reports below. Anthropic hasn't said whether a former Tier 2 organization became Start or Build, or whether Tier 4 became Scale. Your Console shows the result for your account.
  • Old Tier 1–4 rate limit numbers are obsolete. The current Start tier alone allows 1,000 RPM and 2,000,000 ITPM on standard models. Plan against the current per-model tables, not old screenshots.
  • The old "1M context only on Tier 4" rule no longer applies. See the long-context section.

What reports say about how old tiers became Start, Build and Scale

The only first-party statement is Anthropic's release note of June 26, 2026: rate limits were raised, Sonnet and Haiku limits now match Opus at every tier, usage tiers were consolidated into Start, Build and Scale, "most organizations move to a higher tier," no organization received lower limits, and no action is required.

Two third-party reports fill in more detail. Treat them as reports, not Anthropic policy:

  • A May 2026 limit increase on the old tiers. MindStudio reported on May 8, 2026 that Opus input limits jumped to 500,000 ITPM on Tier 1 (from 30,000), 2,000,000 on Tier 2, 5,000,000 on Tier 3 and 10,000,000 on Tier 4, with Tier 1 output going from 8,000 to 80,000 OTPM.
  • "No longer based on spend" and "5x at the highest tier." AI Catchup quoted a ClaudeDevs post on X (July 2, 2026) saying tiers are no longer based on API spend and that the latest Sonnet and Haiku models got 5x higher limits at the highest tier. AI Catchup later noted that Anthropic's docs describe the change as parity with Opus and don't state a 5x figure.

Put the May numbers next to today's table and a pattern appears. This is an inference, not a published mapping:

Old tierOpus ITPM reported in May 2026Current tier with the same ITPMLikely outcome (unconfirmed)
Tier 1500,000None (below Start's 2,000,000)Moved up to Start, or held in Evaluation if usage history was thin
Tier 22,000,000StartStart or higher
Tier 35,000,000BuildBuild or higher
Tier 410,000,000ScaleScale

Read it this way: if your organization was on Tier 4 before June, Scale-level limits are the likely result. If you were on Tier 1, the release note's "no organization receives lower limits" means you should see at least the old Tier 1 numbers, and most likely Start. The Console is still the only authoritative answer for your account.

Rate Limits Explained: RPM, ITPM, and OTPM per Model

Each model class has its own RPM, ITPM and OTPM limit. Exceeding any one of the three returns HTTP 429 with a retry-after header that tells you how long to wait. Because limits are per model, you can run Sonnet 5.5 and Haiku 5.5 at their full limits at the same time.

Start, Build and Scale rate limits by model

Model classStart (RPM / ITPM / OTPM)Build (RPM / ITPM / OTPM)Scale (RPM / ITPM / OTPM)
Opus 5.5, Opus 5, Opus 4.x1,000 / 2,000,000 / 400,0005,000 / 5,000,000 / 1,000,00010,000 / 10,000,000 / 2,000,000
Sonnet 5.5, Sonnet 5, Sonnet 4.x1,000 / 2,000,000 / 400,0005,000 / 5,000,000 / 1,000,00010,000 / 10,000,000 / 2,000,000
Haiku 5.5, Haiku 4.51,000 / 2,000,000 / 400,0005,000 / 5,000,000 / 1,000,00010,000 / 10,000,000 / 2,000,000
Fable 5.x1,000 / 500,000 / 100,0002,000 / 1,500,000 / 300,0004,000 / 4,000,000 / 800,000
Haiku 3.5 (retired except on Google Cloud)1,000 / 100,000 / 20,0002,000 / 200,000 / 40,0004,000 / 400,000 / 80,000

Each row lists model classes that have the same numbers. They don't share one limit, though: each model class gets its own bucket at those values, as the next section explains. Custom-tier limits aren't published; you request them from sales through the Rate limits page. Source: Anthropic rate limits documentation, Messages API tables, as of October 8, 2026.

Which Claude models share a rate limit bucket

A few model families are grouped into one combined limit, so traffic to any member draws from the same bucket:

  • Fable 5.x: Fable 5.1 and Fable 5 share one limit. Mythos 5.1 and Mythos 5 share a separate combined limit on the same terms.
  • Opus 4.x: Opus 4.8, 4.7, 4.6 and 4.5 share one limit. Opus 5.5 and Opus 5 each have their own.
  • Sonnet 4.x: Sonnet 4.6 and 4.5 share one limit. Sonnet 5.5 and Sonnet 5 each have their own.

Limits are also shared across inference_geo values, so "us" and "global" requests draw from the same pool. If you can't decide which model to run in the first place, Claude Sonnet vs Opus vs Haiku vs Fable: Which Model to Use compares them by task.

Cache-aware ITPM: cache reads don't count toward input limits

For every current model except Haiku 3.5, only two usage fields count toward ITPM:

  • input_tokens: tokens after your last cache breakpoint
  • cache_creation_input_tokens: tokens being written to the cache

cache_read_input_tokens don't count. Anthropic's own example: with a 2,000,000 ITPM limit and an 80% cache hit rate, you can push about 10,000,000 total input tokens per minute (2M uncached + 8M read from cache). The general form is effective input per minute ≈ ITPM ÷ (1 − cache hit rate). At a 50% hit rate the same limit stretches to about 4,000,000.

Haiku 3.5 is the exception: its cache reads do count toward ITPM.

Bar chart: with a 2,000,000 ITPM limit, total input per minute is 2M with no cache hits, 4M at a 50% hit rate and 10M at 80%, because cache reads don't count

OTPM counts generated tokens, not max_tokens

OTPM is evaluated as output is produced and counts only tokens actually generated. The max_tokens value you send doesn't reserve OTPM capacity, so setting it high has no rate-limit cost. There's no need to shrink it to save OTPM. Keep it as a guard against runaway responses and cost, not as a rate-limit tactic.

Token bucket replenishment and acceleration limits

Anthropic uses a token bucket: capacity refills continuously up to your maximum instead of resetting at the top of each minute. After a 429 you don't wait for a fixed window; retry-after tells you when enough capacity is back.

There's one more source of 429s: acceleration limits. A sharp jump in usage can trigger them even when you're under your per-minute limits. Anthropic's advice is to ramp traffic up gradually and keep usage patterns consistent, which matters for launches, backfills and load tests.

Choosing the Right Tier for Your Use Case: Spend Cap vs Throughput

You can't pick a tier the way you could buy into Tier 4. What you can do is figure out which limit you'll hit first and plan around it. A quick calculation at Anthropic's list prices (no caching, standard global pricing) shows how fast saturated traffic reaches each monthly cap:

Tier and modelCost per minute at full ITPM + OTPMMinutes of saturated traffic until the monthly cap
Start, Sonnet 5.5 ($2 / $10 per 1M tokens)2M × $2 + 0.4M × $10 = $8.00$500 ÷ $8 ≈ 62 minutes
Start, Haiku 5.5 ($0.10 / $0.50, prompts ≤100K)2M × $0.10 + 0.4M × $0.50 = $0.40$500 ÷ $0.40 = 1,250 minutes (≈21 hours)
Build, Sonnet 5.55M × $2 + 1M × $10 = $20.00$1,000 ÷ $20 = 50 minutes
Scale, Sonnet 5.510M × $2 + 2M × $10 = $40.00$200,000 ÷ $40 = 5,000 minutes (≈83 hours)

How to read it:

  • On Start and Build, the monthly spend cap usually stops you before the rate limits do. That holds for any sustained workload. The per-minute limits are generous; the dollar ceiling is not. If you expect more than $500 or $1,000 of usage in a month, request a higher limit before launch rather than after the cap pauses production.
  • Rate limits matter most for bursts. A batch job that fires thousands of requests at once, or one huge prompt per call, hits RPM or ITPM long before the monthly bill becomes a problem.
  • Scale is for real production volume. The $200,000 cap leaves room for days of saturated traffic. If you need more than that, or limits above Scale, you're in Custom territory.

Prices are for illustration and come from Anthropic's pricing page as of October 8, 2026. For the full per-model price list, cache and Batch rates, see Claude API Pricing Per Million Tokens.

How to Get Higher Claude API Limits: Request Tier Increase

  1. Open the Rate limits page in the Claude Console and use Request tier increase. The same flow covers higher rate limits and a higher monthly spend cap.
  2. Show sustained usage. Anthropic's Help Center article on rate limits says you can request higher limits once you're using at least 50% of your current limits. The Usage page charts (hourly peak uncached input tokens per minute, output tokens per minute and cache rate) are the evidence to look at before asking.
  3. For urgent needs, contact Anthropic support. Support can raise limits too.
  4. For limits above Scale, contact sales through the Rate limits page. That's the Custom tier.

Claude Platform on AWS is different. Organizations there start on the Start tier and move up as they build a history of paid AWS Marketplace invoices. The Request tier increase button isn't available; contact your Anthropic account representative or support instead.

Workspace limits: give one project less, not more

You can set custom spend and rate limits per workspace to protect the rest of the organization. For example, you can cap a staging workspace so a runaway test can't eat the production budget. Three rules apply. You can't set limits on the default workspace. A workspace without its own limits inherits the organization's. And organization-wide limits always apply, even if workspace limits add up to more. Workspace limits can only divide what your tier gives you; they never raise it.

Handling Claude API 429 Errors: Rate Limit vs Spend Cap

Not every 429 from the Claude API means "slow down." The spend-cap error uses the same HTTP status and error type as a rate limit, but retrying it is useless until access resumes.

Decision diagram: a Claude API 429 with error_code enforced_spend_limit_reached and no retry-after means stop retrying, while any other 429 with retry-after means wait, add jitter and retry; a 400 usage-limits message means your own spend limit
What you hitHTTP status and errorretry-after header?What to do
RPM, ITPM, OTPM or acceleration limit429 rate_limit_errorYesWait for retry-after, then retry with backoff and jitter
Tier's monthly spend cap429 rate_limit_error with error.details.error_code = enforced_spend_limit_reachedNoStop retrying. Access returns at 00:00 UTC on the 1st, or after a tier increase
Spend limit you set yourself400 invalid_request_error, message starts "You have reached your specified API usage limits"NoRaise or remove the limit under Settings > Billing
Claude Code workspace limit429YesWait; this limit is checked separately from the organization cap

Fast mode on Opus 5.5, Opus 5 and Opus 4.8 has its own dedicated limits. Exceeding them also returns a 429 with retry-after. For the full diagnosis path, including how 429 differs from 529, see How to Fix Claude API 429 Rate Limit Error.

Retry with exponential backoff and jitter in Python

The official SDKs retry some errors automatically, but automatic retries can't fix a spend-cap 429. The pattern below turns off the SDK's own retries so one loop controls them. It honors retry-after when present, otherwise uses exponential backoff with full jitter, and stops immediately on enforced_spend_limit_reached.

python
import random
import time

import anthropic

client = anthropic.Anthropic(max_retries=0)  # this loop owns the retries


class SpendCapReached(Exception):
    """Monthly spend cap hit: retrying will fail until access resumes."""


def create_with_backoff(max_attempts=6, base=1.0, cap=60.0, **params):
    for attempt in range(max_attempts):
        try:
            return client.messages.create(**params)
        except anthropic.RateLimitError as err:
            body = err.body if isinstance(err.body, dict) else {}
            error = body.get("error", body)
            code = (error.get("details") or {}).get("error_code")
            if code == "enforced_spend_limit_reached":
                raise SpendCapReached(error.get("message", "spend cap reached")) from err
            if attempt == max_attempts - 1:
                raise

            retry_after = err.response.headers.get("retry-after")
            if retry_after is not None:
                delay = float(retry_after) + random.uniform(0, 1)  # small jitter
            else:
                delay = random.uniform(0, min(cap, base * 2 ** attempt))  # full jitter
            time.sleep(delay)


# Usage: pass the same arguments you would give client.messages.create()
# reply = create_with_backoff(model=MODEL_ID, max_tokens=1024,
#                             messages=[{"role": "user", "content": "Hello"}])

MODEL_ID is whichever model ID you already use. Jitter matters when many workers share one organization: without it, they all wake up at the same moment and trigger the next 429 together. Catch SpendCapReached at the job level. Alert someone, pause the queue and don't requeue the request.

Rate limit response headers to monitor

Every response carries headers that show your remaining capacity, so you can slow down before you hit a 429:

HeaderWhat it tells you
retry-afterSeconds to wait before retrying (absent on the spend-cap 429)
anthropic-ratelimit-requests-limit / -remaining / -resetRPM limit, requests left, full-refill time (RFC 3339)
anthropic-ratelimit-input-tokens-limit / -remaining / -resetITPM limit, input tokens left (rounded to the nearest thousand), refill time
anthropic-ratelimit-output-tokens-limit / -remaining / -resetOTPM limit, output tokens left, refill time
anthropic-ratelimit-tokens-limit / -remaining / -resetThe most restrictive token limit currently in effect, including a workspace limit
anthropic-priority-*Priority Tier capacity (Priority Tier customers only)

In the Python SDK, read them through the raw-response wrapper:

python
raw = client.messages.with_raw_response.create(**params)
left = raw.headers.get("anthropic-ratelimit-input-tokens-remaining")
message = raw.parse()  # the normal Message object

Optimizing Your Rate Limits with Prompt Caching and Batch

Once you know which limit binds, two features buy the most headroom without a tier change.

Prompt caching: stretch ITPM and cut input cost

Because cache reads don't count toward ITPM on current models, caching the repeated part of every request multiplies usable input throughput. Good candidates are system instructions, long reference documents, tool definitions and earlier conversation turns. With an 80% hit rate, a Build-tier 5,000,000 ITPM limit covers about 25,000,000 total input tokens per minute (5M ÷ 0.2). Cache reads are also billed at a fraction of the base input price, so the same change lowers your spend against the monthly cap.

Check your actual hit rate in the Rate Limit - Input Tokens chart on the Console Usage page, which plots it next to your ITPM limit. For cache_control placement, cache lifetimes and code, see How to Use Prompt Caching in Claude API.

Message Batches API limits per tier

The Message Batches API has its own limits, shared across all models and separate from the per-model Messages limits. Batch processing is billed at 50% of standard input and output prices, which also makes the monthly spend cap go twice as far for work that can wait.

TierBatch API RPM (all endpoints)Max batch requests in processing queueMax requests per batch
Start1,000200,000100,000
Build2,000300,000100,000
Scale4,000500,000100,000
CustomContact salesContact salesContact sales

A "batch request" is one item inside a Message Batch. A single batch of 100,000 items uses half of the Start tier's 200,000-item queue until those items are processed. Batch suits classification, moderation, bulk summarization, evaluations and nightly reports: anything without a user waiting on the answer.

Other Claude API Limits: Managed Agents, Fast Mode, Files and 1M Context

Managed Agents, Fast mode and Files API limits

  • Claude Managed Agents: create endpoints (agents, sessions, environments) allow 300 requests per minute per organization, and read endpoints (retrieve, list, stream) allow 1,200 per minute. Both are separate from Messages API limits.
  • Fast mode (research preview, speed: "fast" on Opus 5.5, Opus 5 and Opus 4.8) has dedicated limits separate from standard Opus limits. Its status comes back in anthropic-fast-* headers.
  • Files API requests have their own per-organization limit, shared across upload, list, retrieve, download and delete.

1M token context window: no longer a Tier 4 feature

Under the old numbered tiers, the 1M-token context window was a beta limited to Tier 4 organizations and enabled with a beta header. That restriction is gone. Anthropic's pricing documentation now states that Claude 4.6 and later models (except Haiku 5.5) include the full 1M-token context window at standard pricing, with no tier requirement. A 900K-token request is billed at the same per-token rate as a 9K-token one.

Two things still apply. Huge prompts consume ITPM quickly, so a few 900K-token requests can use up a Start-tier minute unless most of the prompt is a cache read. And Haiku 5.5 prices by prompt length: prompts over 100,000 tokens cost $0.50 instead of $0.10 per million input tokens. Source: Anthropic pricing documentation.

FAQ: Common Questions About Claude API Limits

What are the different tiers for the Claude API?

Start, Build and Scale are the standard usage tiers, with monthly spend caps of $500, $1,000 and $200,000. Custom is for organizations whose limits are arranged with Anthropic's account team and has no spend cap. New organizations may begin in an Evaluation tier with lower, unpublished limits.

What is Claude API Tier 4 now?

Tier 4 was the top self-serve tier in Anthropic's older scheme (a $400 cumulative credit purchase, $5,000 monthly spend limit). The current documentation no longer uses it, and Anthropic hasn't published a mapping. Reported May 2026 limits put old Tier 4 at 10,000,000 Opus ITPM, the same as today's Scale tier, so former Tier 4 organizations most likely landed on Scale. That's an inference from third-party reports; check the Rate limits page in the Console for your organization's actual tier.

Is there a free tier for the Claude API?

No. There's no free usage tier. Anthropic's pricing page says new users get a small amount of free credits to test the API, without stating the amount. Limits for the consumer Claude apps and Claude Code subscriptions are a separate system; for Claude Code, see Claude Code /usage: Read Limits, Resets, and Usage Monitors.

How quickly do Claude API rate limits reset?

Continuously. The token bucket refills capacity at a steady rate up to your maximum rather than resetting on the minute. The anthropic-ratelimit-*-reset headers give the time of a full refill, and retry-after on a 429 says how long to wait. The monthly spend cap is different: it resets at 00:00 UTC on the first day of the next month.

Can I have different limits for different API keys?

Not per key. Limits apply to the whole organization. The closest tool is workspaces: put keys in separate workspaces and give a workspace lower spend or rate limits. You can't raise a workspace above the organization's limits.

Do cached tokens count toward Claude API rate limits?

Cache reads (cache_read_input_tokens) don't count toward ITPM on current models; Haiku 3.5 is the exception. Uncached input and cache writes do count. Output tokens always count toward OTPM.

What happens when I hit my monthly spend limit?

If it's your tier's cap, requests return HTTP 429 with error_code enforced_spend_limit_reached and no retry-after. Usage stays paused until 00:00 UTC on the 1st, unless you get a higher limit through Request tier increase. If it's a limit you set yourself, requests return HTTP 400 and you can raise or remove the limit under Settings > Billing.

How do I contact sales for custom Claude API limits?

Go to the Rate limits page in the Claude Console and use the contact-sales option for limits above the Scale tier. Anthropic's pricing page also lists sales@anthropic.com for enterprise pricing and custom rate limits.

Are Claude API rate limits shared across models?

No. Each model class has its own limits, so different models can run at full limits simultaneously. The exceptions are the combined buckets: Fable 5.1 + Fable 5, Mythos 5.1 + Mythos 5, Opus 4.8/4.7/4.6/4.5, and Sonnet 4.6 + 4.5.

What's the difference between Priority Tier and usage tiers?

They're different things. Usage tiers (Start, Build, Scale, Custom) set your organization's limits and spend cap. Service tiers (Priority, Standard and Batch) describe how a request is served. Priority Tier customers also get their own anthropic-priority-* rate limit headers.

Summary and Quick Reference: Claude API Tiers at a Glance

LimitStartBuildScaleCustom
Monthly spend cap$500$1,000$200,000None
Opus / Sonnet / Haiku 5.x and 4.x RPM1,0005,00010,000Negotiated
ITPM (uncached input)2,000,0005,000,00010,000,000Negotiated
OTPM400,0001,000,0002,000,000Negotiated
Fable 5.x RPM / ITPM / OTPM1,000 / 500K / 100K2,000 / 1.5M / 300K4,000 / 4M / 800KNegotiated
Batch queue (requests)200,000300,000500,000Contact sales

Key rules as of October 8, 2026:

  1. Tiers are assigned automatically from usage history and account standing. There are no deposit thresholds, and Tier 1–4 are legacy names with no published mapping.
  2. Cached reads don't count toward ITPM (except Haiku 3.5), and max_tokens doesn't count toward OTPM.
  3. A 429 with enforced_spend_limit_reached and no retry-after is the monthly cap. Don't retry it.
  4. Every other 429: honor retry-after, add jitter and ramp traffic gradually.
  5. Need more? Use Request tier increase once you're using at least 50% of your current limits, or contact sales for Custom.

Limits change, so treat the official rate limits page and your Console as the final word for your organization.