AI & Agents

Claude 3.5 Sonnet Rate Limits: API Tiers, Message Caps, and Throughput

The Claude 3.5 Sonnet rate limit defines Anthropic throughput caps on requests per minute (RPM) and tokens per minute (TPM) across API tiers, alongside Claude.ai session message limits. Tier 1 accounts encounter strict ceilings of 50 RPM and 40,000 TPM, where a single prompt analyzing a large codebase can trigger immediate throttling. Managing token consumption requires prompt caching, request pacing, and decoupled file storage through external MCP workspaces.

Tom Langridge 14 min read Updated
Anthropic rate limits constrain API requests and tokens per minute across developer tiers.

How Anthropic Token Buckets Enforce Rate Limits

The Claude 3.5 Sonnet rate limit defines Anthropic throughput caps on requests per minute (RPM) and tokens per minute (TPM) across API tiers, alongside Claude.ai session message limits. Developers building autonomous coding agents and document processing pipelines encounter these caps when their systems begin making concurrent calls. A single coding prompt that inspects a 100,000-token repository consumes 100,000 input tokens in one shot, which instantly exhausts the Tier 1 per-minute allowance. The resulting HTTP 429 response halts automated workflows and stalls downstream tasks until capacity recovers.

Anthropic enforces API limits using a token bucket algorithm rather than a fixed-window counter. In a fixed-window architecture, limits reset abruptly at the start of a calendar minute, allowing sudden bursts of traffic that can strain model infrastructure. Under Anthropic's token bucket model, capacity replenishes continuously at a steady rate up to your maximum limit. If your organization possesses an allowance of 40,000 tokens per minute, the bucket refills steadily every second. When an agent fires a large request that drains the bucket, subsequent calls fail until the bucket accumulates enough tokens to cover the next payload.

Anthropic measures and constrains usage across three distinct dimensions on the Messages API:

  • Requests Per Minute (RPM): The total number of individual API requests initiated by your organization within a 60-second window.
  • Input Tokens Per Minute (ITPM): The cumulative volume of prompt tokens submitted to the model, evaluated at the start of each request.
  • Output Tokens Per Minute (OTPM): The volume of completion tokens generated by the model, metered in real time as tokens stream back.

A common misconception among developers is that setting a high value for the max_tokens parameter consumes ITPM or OTPM quota in advance. Anthropic evaluates OTPM limits based only on the tokens the model actually generates during the completion. Setting max_tokens: 4096 on a response that only generates 120 tokens draws only 120 tokens against your OTPM ceiling.

All API rate limits apply at the organization level, not per API key. If your team provisions separate keys for individual developers, testing environments, and production background workers, all of those keys draw from the same shared pool. Multiple autonomous coding loops running concurrently against the same Anthropic organization will exhaust your per-minute token allocation rapidly compared to a single developer working in isolation. Review Anthropic's rate limits documentation for organization-level settings.

How Claude 3.5 Sonnet Rate Limits Compare Across API Tiers

Anthropic groups API accounts into progressive usage tiers. Your organization moves through these tiers automatically as your cumulative prepaid credit purchases clear and your account builds standing. Each tier level expands your requests per minute, input tokens per minute, and output tokens per minute ceilings.

The following comparison details the throughput caps for Claude 3.5 Sonnet across standard developer tiers:

API Tier Level Credit Deposit Threshold Requests Per Minute (RPM) Input Tokens Per Minute (ITPM) Output Tokens Per Minute (OTPM)
Tier 1 $5 prepaid deposit 50 RPM 40,000 TPM 8,000 OTPM
Tier 2 $40 cumulative deposit 1,000 RPM 80,000 TPM 16,000 OTPM
Tier 3 $200 cumulative deposit 2,000 RPM 160,000 TPM 32,000 OTPM
Tier 4 $400 cumulative deposit 4,000 RPM 400,000 TPM 80,000 OTPM

In newer account dashboards, Anthropic also classifies enterprise production volume under named service tiers: Start, Build, and Scale, each offering progressively higher monthly spending ceilings followed by negotiated Custom enterprise agreements. Regardless of dashboard nomenclature, the underlying mechanics govern concurrency identically.

Throughput ceilings differ substantially between model families. Anthropic assigns separate rate limit pools to Claude Haiku, Claude Sonnet, and Claude Opus:

  • Claude Haiku: Because Haiku requires less computational cluster overhead per forward pass, Anthropic grants it much higher baseline limits. Tier 1 accounts receive higher initial ITPM allocations for Haiku, scaling to millions of tokens per minute at higher tiers.
  • Claude Opus: Because Opus demands massive hardware resources for deep reasoning, its throughput caps are the most restrictive across all tiers.
  • Claude 3.5 Sonnet: Sonnet represents the primary workhorse model for coding and multi-step reasoning. Its 40,000 TPM baseline on Tier 1 creates an immediate operational ceiling for multi-file code editing, where reading full modules quickly exceeds the threshold.

In addition to standard minute-level limits, Anthropic monitors traffic acceleration. If your organization averages two requests per minute and suddenly dispatches 45 requests within three seconds, you can encounter a 429 response even while staying below your 50 RPM cap. Distributing requests evenly over each 60-second window prevents acceleration throttling.

Fastio features

Keep Claude Agents Running Within API Limits

Connect your Claude agents to Fast.io workspaces through MCP. Index large file repositories for targeted retrieval instead of stuffing raw documents into context. Monthly plans start with a 30-day free trial.

Why Prompt Caching Multiplies Effective Token Throughput

The most effective mechanism for maximizing throughput under Anthropic rate limits is prompt caching. Many API providers evaluate rate limits against total tokens, treating cached and uncached input identically. Anthropic operates differently: for most Claude models, only uncached input tokens count toward your ITPM rate limits.

When an API call includes prompt caching headers, Anthropic categorizes the input payload into three distinct token groups:

  1. input_tokens: Uncached tokens submitted after the final cache breakpoint. These count directly toward your ITPM limit.
  2. cache_creation_input_tokens: Tokens written to the cache for the first time. These count toward your ITPM limit during that specific request.
  3. cache_read_input_tokens: Previously cached tokens read directly from memory. These do not count toward your ITPM rate limit.

This distinction alters the throughput equation for agentic development. Suppose an organization operates on Tier 2 with an 80,000 ITPM ceiling. If the developer structures an agent to inspect a 60,000-token codebase repeatedly, sending that codebase uncached would allow only one request per minute before triggering a 429 error.

By placing a cache control breakpoint on the codebase and static system prompt, the initial request writes the 60,000 tokens to cache. Subsequent requests that modify only the trailing 500-token user instruction register input_tokens: 500 and cache_read_input_tokens: 60,000. Because the 60,000 read tokens bypass the ITPM calculation, the agent can issue dozens of requests per minute without approaching the 80,000 token limit. The effective throughput multiplies while the underlying account tier remains unchanged.

Implementing prompt caching requires attention to technical constraints:

  • Minimum token threshold: Claude 3.5 Sonnet requires a minimum of 1,024 tokens to activate a cache block. Submitting smaller snippets will not trigger cache creation.
  • Cache lifetime: Anthropic maintains cached blocks in memory for 5 minutes following the most recent read. Every request that reads the cache refreshes that 5-minute timer. If an interactive session pauses for longer than 5 minutes while waiting for user approval or external network calls, the cache expires. The subsequent call must write the cache anew, consuming the full token count against ITPM.
  • Breakpoint placement: Anthropic permits up to four cache breakpoints per request using the {"type": "ephemeral"} control block. Place these markers on static components that rarely change: system prompts, tool declarations, and core project context. Dynamic user messages must always appear after the final cache breakpoint.

What Governs Claude.ai Message Limits and File Ceilings

Rate limits on the web and desktop chat interfaces at Claude.ai operate under different rules than the developer API. Instead of measuring continuous tokens per minute through programmatic buckets, Claude.ai enforces message quotas governed by a rolling 5-hour window.

On Claude Free, users receive a dynamic message budget that fluctuates based on overall system demand. On Claude Pro, subscribers receive an expanded allocation, allowing regular interactive exchanges every 5-hour window. Claude Team and Enterprise plans grant higher organizational allocations with centralized user management.

The rolling 5-hour window begins at the precise moment you send your first message in a session. If you submit a prompt at 9:00 AM, your budget recalculates and resets at 2:00 PM. Sending five messages between 9:00 AM and 9:15 AM and your final message at 10:30 AM means your entire allocation refreshes at 2:00 PM, exactly 5 hours from your initial 9:00 AM timestamp. When you exhaust your quota, Claude displays an on-screen countdown indicating the exact time access will resume.

A key operational factor on Claude.ai is conversation length. Claude does not treat each chat turn as an isolated prompt. To maintain context, Claude re-transmits the complete conversation history with every message you send. Turn 1 might send 200 tokens. Turn 20 re-sends the cumulative text of all 19 preceding turns, assistant responses, and code blocks. In long conversations, each new message consumes a larger portion of your 5-hour compute budget, causing the session limit warning to appear much earlier than expected.

File upload limits on Claude.ai introduce additional boundaries detailed in Anthropic's file upload guide:

  • Chat uploads: Standard chats accept individual document and image attachments up to Claude's conversation file limits.
  • Project knowledge: Claude Projects accepts files up to 30MB each with an unlimited number of files, provided the total content fits within Claude's 200,000-token context window.

While Claude Projects imposes no artificial cap on file count, the 200,000-token context window serves as a hard physical boundary. Uploading multiple technical manuals or an entire application repository quickly fills the context window to capacity. Once project knowledge occupies 150,000 tokens, every message submitted inside that project automatically begins with 150,000 tokens of overhead, exhausting your 5-hour allocation within two or three interactions.

Steps to Prevent Rate Limits Using Remote MCP Workspaces

Stuffing raw documents, full technical specifications, and entire repositories into Claude's prompt window is the primary cause of rate limit exhaustion. When an agent or user submits 80,000 tokens of raw text to answer a question that requires only three sentences from a reference guide, the interaction consumes valuable ITPM quota and accelerates session caps.

Teams typically evaluate several approaches to manage context:

  1. Local file parsing: Engineers write custom Python or shell scripts using grep or ripgrep to isolate code sections before passing them to Claude. This approach functions on an individual laptop, but it fails to scale across distributed teams or autonomous agents running in cloud containers, provides no semantic search over unstructured PDFs or scanned records, and maintains no centralized audit trail.
  2. Raw cloud object storage: Storing project assets in Amazon S3 or traditional Dropbox buckets keeps files off local machines. However, an AI agent interacting with raw storage must download entire files into memory and parse them manually, consuming bandwidth and re-creating the same token bloat when passing contents to the model.

A more practical architecture decouples file storage from model context through an intelligent workspace. Fast.io provides org-owned workspaces where files are automatically indexed for hybrid search, combining full-text lexical matching and semantic vector retrieval once Intelligence Mode is enabled.

Instead of attaching dozens of raw files to Claude, you connect your assistant or agent directly to Fast.io using the Model Context Protocol (MCP). The Fast.io remote MCP server operates at https://mcp.fast.io/mcp/tools over Streamable HTTP. Detailed setup steps are available in the Fast.io documentation.

When an agent needs information from your project repository, it does not ingest the entire file catalog. Instead, it calls the MCP search action:

{
  "server": "fastio",
  "action": "storage/search",
  "parameters": {
    "search": "authentication middleware session expiration",
    "workspace_id": "ws_dev_core"
  }
}

The workspace queries the indexed files and returns only the precise 400-token code excerpt relevant to the task. By transmitting targeted excerpts instead of an entire repository dump, the agent reduces its per-prompt token consumption to a fraction of the baseline payload. This efficiency preserves Tier 1 and Tier 2 ITPM budgets, prevents 429 throttling, and lowers response latency. Explore Fast.io storage for agents to configure persistent workspaces.

Fast.io supports multi-agent and human collaboration across these workspaces:

  • Per-file version history: Every update committed by an agent or human creates an immutable version entry, allowing instant rollback if an agentic edit introduces a regression.
  • Append-only audit log: Tracks every file access, edit, download, and permission modification across the entire workspace.
  • Collaborative Notes: Real-time co-editing documents where human developers and AI assistants review plans and document changes side by side.
  • Advisory file locks: Leases acquired and released through the MCP storage_manage tool (lock-acquire and lock-release, with lock-status on storage) to prevent parallel agents from stepping on active files.
  • Cloud Sync: Scheduled or on-demand synchronization connects existing files from Dropbox, Box, and OneDrive (Google Drive supports cloud import today, with sync coming soon).
  • Ownership transfer: An autonomous agent can initialize an organization, construct workspaces, populate documentation, and transfer ownership to a human team member while retaining administrative operational access.

Monthly plans start with a 30-day free trial that requires a credit card. Paid subscriptions include Starter at $9.99/mo (with 3 seats and 250 GB storage), Business at $49.99/mo (with 10 seats and 5 TB storage), and Enterprise at $199.99/mo (with 30 seats and 25 TB storage). Additional workspace seats and AI credits scale with team requirements. Check Fast.io pricing for complete plan terms.

How to Handle HTTP 429 Errors and Production Retries

Production systems calling Claude 3.5 Sonnet must handle rate limit errors gracefully. When your application exceeds an RPM, ITPM, or OTPM ceiling, Anthropic returns an HTTP 429 status code with a JSON payload specifying the error type:

{
  "type": "error",
  "error": {
    "type": "rate_limit_error",
    "message": "Number of request tokens has exceeded your per-minute rate limit"
  }
}

Anthropic includes informative response headers on every API transaction. Inspecting these headers allows client applications to pace requests dynamically before encountering an error:

  • anthropic-ratelimit-requests-remaining: Remaining requests available in the current bucket.
  • anthropic-ratelimit-requests-reset: The ISO 8601 timestamp when request capacity fully resets.
  • anthropic-ratelimit-tokens-remaining: Remaining tokens available before hitting ITPM or OTPM caps.
  • anthropic-ratelimit-tokens-reset: The timestamp when token capacity fully refills.
  • retry-after: On 429 responses, indicates the integer number of seconds the client must wait before retrying.

To handle transient 429 responses, implement exponential backoff with full jitter in your client code:

import time
import random
import httpx

def call_claude_with_retry(payload: dict, api_key: str, max_retries: int = 5):
    url = "https://api.anthropic.com/v1/messages"
    headers = {
        "x-api-key": api_key,
        "anthropic-version": "2023-06-01",
        "content-type": "application/json"
    }
    
    for attempt in range(max_retries):
        response = httpx.post(url, json=payload, headers=headers, timeout=60.0)
        
        if response.status_code == 200:
            return response.json()
            
        if response.status_code == 429:
            retry_after = response.headers.get("retry-after")
            if retry_after:
                sleep_seconds = float(retry_after)
            else:
                base_delay = 2 ** attempt
                sleep_seconds = base_delay + random.uniform(0.1, 1.0)
            
            time.sleep(sleep_seconds)
            continue
            
        if response.status_code == 529:
            time.sleep(2 ** attempt + random.uniform(0.5, 2.0))
            continue
            
        response.raise_for_status()
        
    raise RuntimeError("Exceeded maximum retries for Claude API call")

Distinguish HTTP 429 from HTTP 529 (overloaded_error). A 429 status code signals that your organization exceeded its allocated quota. Retrying immediately without delay will continue to fail. Conversely, an HTTP 529 status code indicates that Anthropic's backend servers are temporarily experiencing high traffic volume. When encountering a 529, back off with brief jitter and retry; the condition usually clears quickly.

Another distinct variation occurs when an organization hits its monthly spend cap. In that scenario, Anthropic returns an HTTP 429 with error.details.error_code: "enforced_spend_limit_reached". This response omits the retry-after header because requests remain blocked until the organization admin raises the spend limit in the Claude Console or the monthly billing period resets.

For large-scale, non-interactive workloads such as repository audits, batch data extraction, or synthetic dataset generation, consider the Message Batches API. The Batches API processes requests asynchronously within 24 hours under separate rate limit queues, bypassing standard real-time RPM and TPM bottlenecks while reducing token costs by half compared to standard pricing.

Sources

References used to verify factual claims in this guide.

  1. Anthropic rate limits for Claude models count only uncached input tokens toward input tokens per minute limits.

  2. Claude Projects allows an unlimited number of uploaded files up to 30MB each provided the total content fits within Claude's context window.

Frequently Asked Questions

What is the rate limit for Claude 3.5 Sonnet?

On the Anthropic API, Claude 3.5 Sonnet rate limits depend on your organization's tier. Tier 1 accounts start with 50 requests per minute (RPM), 40,000 input tokens per minute (ITPM), and 8,000 output tokens per minute (OTPM). Tier 4 accounts reach 4,000 RPM, 400,000 ITPM, and 80,000 OTPM. On Claude.ai, Pro users receive a rolling message allowance of several dozen messages every 5 hours.

How do Anthropic API rate limit tiers work?

Anthropic API tiers determine throughput capacity based on cumulative prepaid credit deposits. Accounts advance automatically through tiers as payment thresholds clear, unlocking higher request and token allowances. Enterprise tiers provide expanded monthly spend caps and custom request throughput for production deployments.

How can I avoid hitting Claude 3.5 Sonnet rate limits?

You can avoid rate limits by implementing prompt caching for static instructions, spacing API calls evenly to avoid burst throttling, using exponential backoff with jitter, and connecting Claude to remote workspaces via MCP. Storing files in an intelligent workspace lets Claude retrieve targeted excerpts rather than ingesting entire documents into prompt context.

What is the difference between a 429 error and a 529 error from Anthropic?

An HTTP 429 error (rate_limit_error) indicates that your organization exceeded its assigned requests per minute, tokens per minute, or monthly spend cap. An HTTP 529 error (overloaded_error) indicates that Anthropic's server infrastructure is temporarily at maximum capacity. Code should retry 529 errors after a short randomized delay, while 429 errors require waiting for capacity to replenish.

Do prompt-cached tokens count against Claude 3.5 Sonnet rate limits?

Tokens read from cache do not count against your input tokens per minute (ITPM) rate limit for Claude 3.5 Sonnet. Only uncached prompt tokens and newly created cache write tokens count toward your ITPM quota, allowing organizations with effective prompt caching to achieve significantly higher real-world throughput.

What happens when you hit the message limit on Claude.ai?

When you hit the message limit on Claude.ai, access pauses until the rolling 5-hour window resets. The window begins on your first interaction, and Claude displays an on-screen timer showing the exact time access will resume. Conversation history, files, and artifacts remain intact.

Related Resources

Fastio features

Keep Claude Agents Running Within API Limits

Connect your Claude agents to Fast.io workspaces through MCP. Index large file repositories for targeted retrieval instead of stuffing raw documents into context. Monthly plans start with a 30-day free trial.