AI & Agents

Claude Code Rate Limits: 5-Hour Usage Caps, 429 Errors, and Workarounds

Claude Code rate limits enforce execution thresholds across terminal requests, tokens per minute, and five-hour rolling usage windows. Autonomous tool-calling loops can trigger HTTP 429 errors within minutes as multi-file inspections compound prompt size. Diagnosing subscription caps versus API tier limits enables developers to manage session context, switch models, and connect external indexed workspaces for large codebases.

Tom Langridge 17 min read Updated
Managing Claude Code rate limits, five-hour rolling windows, and API token throughput during development.

How Claude Code Enforces Rate Limits: 5-Hour Windows and API Tiers

An autonomous coding agent in a terminal loop will happily burn through its entire five-hour token allocation in under fifteen minutes when inspecting a large repository, and it will not warn you before halting on an HTTP 429 error. The bottleneck in automated software engineering with Claude Code is almost never reasoning capacity; it is the mismatch between rapid-fire local file inspection loops and Anthropic's rolling token and request limits.

Claude Code rate limits are execution thresholds enforced by Anthropic that cap terminal requests per minute, tokens per minute, and 5-hour rolling usage windows during automated coding sessions. Unlike interacting with a chat interface where a human reads an answer before typing a response, a command-line agent runs continuous automated tool loops. Inspecting directories, reading source files, analyzing stack traces, and proposing code edits happen in rapid succession. When those operations outpace account allowances, Anthropic halts execution with rate limit errors.

Understanding how Claude Code tracks usage requires distinguishing between its two supported authentication methods: web-based subscription login and direct API keys.

Claude Pro and Team Subscriptions: The 5-Hour Rolling Window

When developers authenticate Claude Code by running the login command in their terminal, the CLI links to their Claude Pro, Max, or Team subscription account. Under this model, usage is governed by session-based conversation budgets rather than explicit per-token billing.

The core constraint on paid subscriptions is the five-hour rolling window. This limit is frequently misunderstood as a fixed daily quota that resets at midnight or on the hour. In practice, Anthropic calculates usage dynamically over the previous five hours of active operation. If you execute a sequence of token-heavy coding tasks at 9:00 AM, that consumption counts against your available budget until 2:00 PM. As individual prompts and responses age past the five-hour mark, your available capacity gradually recovers.

A critical operational factor is that subscription limits are shared across all Claude surfaces. Prompts sent through Claude Code in your terminal, conversations in the browser interface on claude.ai, and interactions in the Claude Desktop application all draw from the exact same account budget. Heavy code generation in the CLI reduces the messages available in your browser, and long discussions in the web app leave less capacity for terminal agent sessions.

Anthropic Console API Keys: RPM, ITPM, and OTPM

When developers configure Claude Code using an API key from the Anthropic Console, execution shifts to commercial API metering. Instead of a five-hour subscription window, the CLI operates under granular rate limits measured in requests per minute and tokens per minute.

According to Anthropic rate limits documentation, the Messages API evaluates throughput through three distinct metrics applied separately to each model class:

  1. Requests Per Minute (RPM): The total number of separate API calls initiated across a sixty-second window.
  2. Input Tokens Per Minute (ITPM): The volume of prompt tokens sent to the model per minute, encompassing system instructions, conversation turns, file contents, and tool schemas.
  3. Output Tokens Per Minute (OTPM): The volume of generated tokens produced by the model per minute.

Anthropic enforces these rate limits using a token bucket algorithm. Rather than resetting counters at the top of each minute, the platform continuously replenishes allowed capacity up to the account ceiling. A rate limit of 60 RPM is enforced approximately as 1 request per second. Initiating ten concurrent tool calls in two seconds can trigger rate limiting even if total traffic for the minute remains well beneath 60 requests.

Usage Tiers and Acceleration Limits

The Anthropic platform groups API accounts into named usage tiers that determine baseline throughput allowances: Start, Build, Scale, and Custom. Organizations are placed on a tier automatically based on usage history and account standing, and move up over time as they use the API. New organizations and organizations with limited usage history may begin in the Evaluation tier, with limits below the standard published limits while account history is established. Each tier also carries a monthly spend cap, and the current limits for your organization are visible on the Rate limits page in the Claude Console.

In addition to steady-state limits, Anthropic monitors traffic acceleration. If a terminal script suddenly spikes from zero activity to dozens of high-context requests within seconds, internal traffic management systems can return HTTP 429 responses to smooth consumption. Managing automated coding workflows requires understanding how multi-file tool invocations push against these ceilings.

Why Terminal Agent Loops Trigger HTTP 429 and Usage Limit Errors

The fundamental reason developers encounter Claude Code rate limits is the compounding nature of agentic tool execution. In a standard web chat, conversation growth is linear: each turn adds one human query and one assistant response. In an autonomous terminal agent, a single developer instruction initiates a multi-turn autonomous execution cycle.

When you direct Claude Code to resolve a failing test, the agent does not produce code in a single inference call. It inspects the repository structure, runs search patterns, examines referenced modules, analyzes dependencies, drafts modifications, and executes test suites. Each action represents a distinct API call that resubmits the accumulated session context back to Anthropic.

The Compounding Context Trap in CLI Loops

Every time Claude Code invokes a tool, the input prompt for the subsequent turn includes the initial prompt, the full tool call history, the tool execution results, and every file read during the session. As the agent navigates your project, the input token volume escalates dramatically with each step:

  1. Step 1 (Task Definition): The user submits an issue description. Input context contains base system instructions, the CLAUDE.md project guide, and the initial prompt. Total: approximately 3,000 tokens.
  2. Step 2 (Locating Source Files): Claude Code runs a grep command across the codebase. The command output and search results are added to the conversation. Total: approximately 5,500 tokens.
  3. Step 3 (Reading Core Logic): Claude Code opens the primary module to inspect the implementation. A 400-line source file adds 4,000 tokens. Total: approximately 9,500 tokens.
  4. Step 4 (Inspecting Imports): The module imports two internal utilities. The agent reads both files to verify function signatures. Total: approximately 18,000 tokens.
  5. Step 5 (Reviewing Test Fixtures): Claude Code reads the existing test suite and a JSON test fixture to understand current assertions. Total: approximately 28,000 tokens.
  6. Step 6 (Running the Test Suite): The agent executes the test command. The failure trace and terminal logs are captured into context. Total: approximately 38,000 tokens.
  7. Step 7 (Drafting Code Changes): The agent attempts to generate the fix. At this step, the input payload alone reaches nearly 40,000 tokens.

Across seven routine diagnostic tool steps, the CLI accumulates significant token volume. On an organization still in the Evaluation tier, or on a workspace whose own input-token limit has been set below the organization ceiling, a single diagnostic loop of this shape can exceed the per-minute token quota, causing Anthropic's gateway to return an HTTP 429 rate limit error.

Recognizing and Diagnosing HTTP 429 Errors in Claude Code

Encountering an HTTP 429 status code indicates that the Anthropic server refused a request because the client sent too many requests or consumed too many tokens in a given time period. In terminal development environments, 429 errors stem from three distinct operational bottlenecks.

Deciphering the Three Varieties of HTTP 429 Errors

A 429 error code manifests in distinct patterns depending on whether the constraint is short-term velocity, rolling subscription capacity, or financial spend caps:

  • Short-Term Concurrency Spikes (Rate Limit Exceeded): This response occurs when a rapid burst of tool executions or an unusually large multi-file prompt exceeds per-minute token or request quotas. Anthropic's API gateways return standard response headers, including retry-after, indicating how many seconds the client must wait before retrying.
  • Rolling Five-Hour Capacity Exceeded: In Claude Pro and Team subscription setups, users receive an explicit terminal notification indicating that their usage limit has been reached until a specified time (for example: "Usage limit reached. Available again at 3:15 PM"). This cap reflects total token volume accumulated across the rolling five-hour window.
  • Monthly Account Spend Ceilings: For API key users, Anthropic Console accounts enforce monthly credit spend caps. When total expenditures reach this preconfigured ceiling, all subsequent API calls are blocked with a 429 status or an explicit credit balance error until the spend limit is increased in the console dashboard.

Diagnostic Procedure for Terminal Sessions

When Claude Code halts on a rate limit, systematic triage reveals the root cause:

  1. Examine the Exact Terminal Output: Check whether the message mentions a retry countdown (seconds), a rolling window reset timestamp (hours), or a credit balance warning.
  2. Check the Anthropic Status Page: Confirm that the error is not caused by an Anthropic infrastructure incident or degraded service tier before changing local settings.
  3. Review Active Session Context Size: If the terminal agent has executed dozens of tool calls, review the cumulative context size. Large accumulated contexts trigger ITPM limits far faster than short sessions.
  4. Verify Account Spending and Tier in the Anthropic Console: For API users, check whether your organization has reached its monthly budget limit or needs tier progression to support higher concurrent throughput.

Claude Code Limits Across Plans and Tiers

Choosing the right access model for Claude Code requires comparing subscription plans against commercial API tiers. The optimal tier depends on developer workflow patterns, team size, and the degree of automation.

Comparing Claude Limits Across Plans and Access Models

To evaluate which environment suits your operational volume, compare the throughput mechanisms across account tiers:

Tier / Plan Throughput Allowance Reset Window Primary Bottleneck Recommended Mitigation
Claude Free Basic evaluation volume Rolling multi-hour window Rapid session cutoff on CLI tasks Switch to Claude Pro or API key
Claude Pro Standard subscription budget Rolling 5-hour window Multi-file agent context accumulation Compact sessions; enable pay-as-you-go credits
Claude Max Extended subscription budget Rolling 5-hour window High-volume multi-agent terminal sprints Offload repository indexing to Fast.io workspaces
API Evaluation tier Below published standard limits Continuous replenish (seconds) Single large prompt exceeding minute quota Scope file reads with permission deny rules
API Start tier Published per-model RPM, ITPM and OTPM Continuous replenish (seconds) Sustained parallel agent worker bursts Smooth request pacing and implement retry queues
API Build and Scale tiers Higher published per-model limits Continuous replenish (seconds) Monthly organizational spend caps Raise the spend limit in the Claude Console

Context Window Ceilings vs Rate Limits

Many developers conflate model context windows with rate limit allowances. Documented Claude mechanics, from Anthropic's Claude file upload documentation, establish that an individual chat accepts up to 20 files at up to 500MB each. For projects, individual files are capped at 30MB each, and file count is unlimited, provided the total content fits within Claude's context window.

However, having context window capacity to process 200,000 tokens does not grant permission to send that volume every minute. If your organization's input-token-per-minute allowance sits below the size of a single stuffed prompt, that prompt immediately violates the rate limit even though the model could theoretically process the input. One detail works in your favor here: on most Claude models only uncached input tokens count toward ITPM, so prompt caching raises effective throughput without raising your limit. Distinguishing context capacity from rate limit throughput prevents costly architecture mistakes.

Immediate Workarounds, Terminal Commands, and Context Management

When Claude Code halts on a rate limit or approaches its usage ceiling, several immediate adjustments can restore productivity and prevent subsequent disruptions. Effective session management focuses on minimizing unnecessary token payloads and directing agent attention precisely.

Pruning Working Context with Terminal Commands

The simplest way to reduce prompt token consumption is actively managing the CLI conversation history:

  • Reset Between Tasks with the slash-clear Command: When switching from one bug fix to an unrelated feature, run slash-clear. This clears accumulated conversation history and flushes previously read file contents from memory. Starting with a clean session resets the input token baseline to your initial prompt.
  • Compact Lengthy Sessions with the slash-compact Command: If you are midway through a complex debugging sequence and cannot discard active state, run slash-compact. Claude Code synthesizes prior turns into a concise operational summary, removing redundant tool outputs and discarded code attempts while preserving key decisions.
  • Select Appropriate Models with the slash-model Command: Claude Code allows dynamic model switching during sessions. For routine tasks such as generating unit test boilerplate, renaming variables, or writing documentation comments, switch to Claude 3.5 Haiku. Reserve Sonnet or Opus for architectural refactoring, subtle concurrency bugs, and complex algorithmic reasoning. Haiku requests draw down fewer conversation credits and operate under generous rate limit allowances.

Scoping File Reads and Configuring Exclusion Rules

When an agent is instructed to locate an implementation, it may read dozens of non-essential files by default. Restricting agent visibility preserves both token budgets and rate limit headroom:

  • Deny Reads on Low-Value Paths: Claude Code excludes files through permission rules rather than a dedicated ignore file. In .claude/settings.json, add Read() entries under permissions.deny to keep minified bundles, build outputs, package lockfiles, generated documentation, test fixtures, and media directories out of reach. A single unminified JavaScript bundle or lockfile can inject 50,000 tokens into prompt context in one read call.
  • Instruct Targeted Line Reads: When prompting Claude Code to examine a large file, explicitly constrain the scope: "Examine lines 40 through 120 of auth.ts to review the session validation logic." This prevents the tool from ingesting thousands of lines of unrelated helper functions.
  • Enable Pay-As-You-Go Usage Credits on Pro Plans: For developers on Claude Pro or Team subscriptions, Anthropic provides an option in account settings to enable usage credits. When your subscription reaches its five-hour session cap, Claude Code automatically transitions to consumption-based API billing without halting work. This eliminates mid-task session blocks during critical delivery windows.

Offloading Large Codebases and Documentation to Indexed Workspaces

While local terminal commands and ignore rules help manage small repositories, working across monorepos, multi-service architectures, or extensive technical documentation libraries can still strain rate limit budgets.

Reading hundreds of raw source files directly into terminal prompts forces Claude Code to re-ingest entire file contents on every tool turn. The sustainable alternative is decoupling storage and retrieval from prompt memory by storing project documents and reference files in a Fast.io shared workspace.

When project files and reference documentation reside in a Fast.io workspace, enabling Intelligence Mode automatically indexes the entire corpus for semantic and full-text search. Claude Code connects to the workspace using the remote Model Context Protocol (MCP) server at https://mcp.fast.io/mcp, authenticating with an organization API key at https://mcp.fast.io/mcp/key, as detailed in the Fast.io for agents documentation.

Instead of reading raw multi-megabyte files into prompt context, Claude Code calls Fast.io search tools to retrieve only the specific functions, schemas, or documentation paragraphs relevant to the current instruction. Files can be uploaded directly or synchronized from existing cloud storage providers like Dropbox, Box, or OneDrive, one-way or two-way, on a schedule or on demand. Google Drive supports file import today, with sync coming soon.

Fast.io leaves Anthropic's rate limits and file size allowances completely intact. What it provides is an external indexed retrieval layer where large corpuses live without saturating Claude Code's terminal token budget. Organizations can get started with a 14-day free trial, which requires a credit card. Review complete plan tiers on Fast.io pricing.

Indexing files for semantic search to prevent terminal agent token exhaustion
Fastio features

Stop Exhausting Agent Rate Limits on Raw Files

Connect Claude Code to a shared workspace where files are indexed for semantic retrieval instead of stuffing raw context into terminal prompts. Every organization starts with a 14-day free trial, which requires a credit card.

Architectural Strategies for Multi-Agent Workflows and Large Repositories

As software teams scale automated coding beyond individual developer laptops, rate limit challenges multiply. Running multiple terminal sessions concurrently, dispatching continuous integration coding bots, or combining Claude Code with tools like Cursor, Cline, or OpenClaw can exhaust organizational API allowances within minutes.

Sustainable multi-agent development requires architectural controls that pace network traffic, coordinate agent access, and prevent redundant token consumption across team members.

Rate Pacing and Exponential Backoff Implementation

When integrating Claude Code into automated CI/CD pipelines or custom developer scripts, client applications must handle HTTP 429 errors gracefully. Implementing exponential backoff with jitter prevents retrying workers from creating synchronized thundering herd spikes against the API.

A standard retry handler inspects the response headers when receiving a 429 status. If a retry-after header is present, the script pauses for the indicated duration. If the header is absent, the client applies exponential backoff based on the attempt count:

export ANTHROPIC_MAX_RETRIES=5
export ANTHROPIC_RETRY_BASE_DELAY_SECONDS=2
export ANTHROPIC_RETRY_MAX_DELAY_SECONDS=60

For custom agentic scripts interacting with Anthropic models, incorporating randomized jitter ensures that parallel workers do not retry simultaneously:

import time
import random

def execute_with_backoff(api_call, max_retries=5):
    for attempt in range(max_retries):
        try:
            return api_call()
        except Exception as e:
            if "429" in str(e) and attempt < max_retries - 1:
                base_delay = 2 ** attempt
                delay = random.uniform(1, base_delay + 1)
                time.sleep(delay)
            else:
                raise e

Establishing Shared Substrates to Avoid Duplicate Agent Reads

When multiple autonomous workers inspect the same repository in multi-agent environments, uncoordinated execution leads to severe token waste. If Agent A reads fifty documentation files to design an interface, and Agent B independently re-reads the same fifty files thirty minutes later, the team draws down double the token quota.

Teams resolve this inefficiency by establishing a shared workspace substrate. Instead of having each terminal agent independently scan and cache files in local memory, project assets, architecture decision records, and intermediate build outputs are centralized in shared workspaces.

Through versioned workspaces, granular permissions, and an append-only audit log, agents and human engineers share a consistent operational layer. Agents query pre-indexed workspace knowledge via remote MCP endpoints, inspect per-file version history to detect modifications without downloading full files, and hand finished work over to human engineers using ownership transfer. Coordinating file access through a centralized workspace layer eliminates duplicate reads, maximizes rate limit efficiency, and keeps development velocity predictable across the engineering organization.

Sources

References used to verify factual claims in this guide.

  1. Anthropic rate limits for the Messages API govern throughput through requests per minute, input tokens per minute, and output tokens per minute for each model class. New organizations and organizations with limited usage history may start in the Evaluation tier, with limits below the standard published limits.

  2. Anthropic limits chat uploads to 20 files at up to 500MB each, while project files are capped at 30MB each with no fixed file count but bounded by the context window.

Frequently Asked Questions

What is the rate limit for Claude Code?

Claude Code rate limits depend on whether you authenticate through a Claude Pro or Team subscription or use an Anthropic Console API key. Subscription users are governed by a five-hour rolling usage window shared across terminal sessions, claude.ai web chat, and Claude Desktop. API key users operate under requests per minute (RPM), input tokens per minute (ITPM), and output tokens per minute (OTPM) limits set per model and per usage tier (Start, Build, Scale, and Custom), with new organizations sometimes starting on a lower Evaluation tier. Your organization's current figures are shown on the Rate limits page in the Claude Console.

How do you fix HTTP 429 errors in Claude Code?

To resolve HTTP 429 errors in Claude Code, identify whether the cause is a short-term velocity limit, a five-hour session cap, or a monthly spend ceiling. For velocity limits, wait for the brief retry-after period, run slash-clear to flush accumulated file context, or switch to Claude 3.5 Haiku using slash-model. For five-hour session caps on Pro plans, enable usage credits in account settings to continue at standard token rates. For monthly spend cap errors on API accounts, increase the spend limit in the Claude Console.

How long is the Claude Code 5-hour reset window?

The Claude Code five-hour reset window evaluates token consumption continuously over the previous five hours of active use rather than resetting on a fixed clock schedule. As individual prompts and tool calls age past the five-hour mark, your available conversation budget gradually recovers. If you exhaust your allowance during an intensive coding sprint, full capacity returns five hours after the peak usage period ended.

Does Claude Code share rate limits with Claude chat on claude.ai?

Yes, when logged in with a Claude Pro, Max, or Team subscription, Claude Code shares its conversation budget with both the claude.ai browser interface and the Claude Desktop application. High-volume code generation and repeated multi-file reads in your terminal will draw down the capacity available for web chat, and long browser discussions will reduce your available terminal coding allowance.

Should I use an Anthropic API key or a Claude Pro subscription with Claude Code?

Use a Claude Pro subscription if you want predictable monthly billing and perform moderate coding tasks with regular breaks between sessions. Use an Anthropic Console API key if you require high-throughput automation, run continuous agentic loops across large repositories, or need explicit control over RPM and TPM throughput tiers without hitting five-hour subscription caps.

Related Resources

Fastio features

Stop Exhausting Agent Rate Limits on Raw Files

Connect Claude Code to a shared workspace where files are indexed for semantic retrieval instead of stuffing raw context into terminal prompts. Every organization starts with a 14-day free trial, which requires a credit card.