# AWS Bedrock Rate Limits: Service Quotas, ThrottlingException, and Document Retrieval

AWS Bedrock rate limits are regional, account-level quotas that cap requests and tokens per minute across foundation models. When applications send large document payloads directly into inference calls, they rapidly exhaust token allowances and trigger ThrottlingException errors. This guide explains how Bedrock quotas work, how to handle throttling with backoff and jitter, and how offloading document retrieval through an intelligent workspace prevents quota exhaustion.

Source: https://fast.io/resources/aws-bedrock-rate-limit/
Author: [Tom Langridge](https://fast.io/authors/tom-langridge/)
Last reviewed: 2026-09-16

## How AWS Bedrock Enforces Rate Limits Across Foundation Models

AWS Bedrock rate limits are regional, account-level quotas enforced by AWS Service Quotas that cap the maximum transactions per second (TPS) and processed tokens per minute for foundation model invocations. When an application streams complete multi-megabyte PDF manuals or complex compliance archives directly into model invocation payloads, token consumption spikes instantaneously, burning through per-minute token allotments and triggering immediate ThrottlingException errors. The bottleneck is not the model context window, but the velocity at which prompt tokens deplete regional throughput quotas.

AWS defines service quotas as regional limits governing the maximum operations and resources for an account, as documented in the [AWS Bedrock endpoints and quotas reference](https://docs.aws.amazon.com/general/latest/gr/bedrock.html). Amazon Bedrock divides model inference infrastructure into two core data plane dimensions: Requests Per Minute (RPM) and Tokens Per Minute (TPM). Understanding how these metrics operate across AWS regions, inference profiles, and client endpoints is essential for building resilient agent architectures and document retrieval systems.

### Requests Per Minute and Tokens Per Minute Accounting

The Requests Per Minute metric governs the frequency of calls made to Bedrock invocation APIs within a rolling sixty-second window. This includes actions such as InvokeModel, InvokeModelWithResponseStream, Converse, and ConverseStream. Amazon Bedrock uses a token bucket algorithm to enforce RPM thresholds. The service smooths traffic over short intervals rather than allowing an entire minute's allocation to arrive in a single second. If an application initiates twenty concurrent agent workers in parallel, the instantaneous burst can deplete bucket tokens, resulting in an HTTP 429 status code even if aggregate minute volume remains modest.

Tokens Per Minute accounting tracks the cumulative volume of input tokens submitted in prompt payloads and output tokens produced during model generation. Bedrock calculates TPM by evaluating the tokenized content of your system prompts, chat history, attached documents, tool definitions, and resulting model completions. Because modern foundation models support extensive context windows, developers often assume they can pass full document bodies freely. However, regional TPM allocations represent a shared velocity ceiling across your entire AWS account in that region.

### Model Invocation Endpoints: Runtime versus Mantle

Amazon Bedrock routes traffic through two primary inference endpoints, each tracked against independent service quotas:

| Metric or Endpoint | Operational Scope | Enforcement Metric | Adjustment Method |
| :--- | :--- | :--- | :--- |
| **Requests Per Minute (RPM)** | Regional, per model family | Invocation frequency over rolling window | AWS Service Quotas request |
| **Tokens Per Minute (TPM)** | Regional, per model family | Combined input and output token velocity | AWS Service Quotas request |
| **bedrock-runtime Endpoint** | Regional data plane | Standard InvokeModel and Converse actions | Service Quotas console |
| **bedrock-mantle Endpoint** | Regional data plane | Input and output token limits for chat APIs | Independent quota allocation |
| **max_tokens Parameter** | Per API invocation | Upfront token reservation during generation | Client-side request payload |

The bedrock-runtime endpoint serves standard invocations like InvokeModel and Converse. The bedrock-mantle endpoint specifically handles the OpenAI-compatible Chat Completions API and the Anthropic Messages API. Traffic dispatched to bedrock-mantle operates against its own dedicated token quotas, which evaluate input and output volume without enforcing traditional request-per-minute limits. Distributing traffic between runtime and mantle interfaces allows high-throughput systems to balance workload demands effectively across endpoints.

### The Upfront Token Reservation Mechanic

A frequent source of unexpected throttling in Amazon Bedrock is the upfront token reservation behavior. When an application dispatches a request to InvokeModel or Converse, Bedrock does not wait for text generation to finish before evaluating TPM limits. Instead, Bedrock immediately reserves quota capacity equal to the input prompt tokens plus the requested max_tokens parameter.

If a developer leaves max_tokens unconfigured or sets it to a default ceiling of 64,000 tokens, Bedrock temporarily deducts that entire quantity from the account's available TPM balance for the duration of the call. If several worker threads initiate requests concurrently, Bedrock reserves extensive capacity against your quota, even if each response only outputs 150 tokens. Once the generation completes, Bedrock releases the unused balance back to the token bucket. During the active invocation window, subsequent requests will trigger a ThrottlingException because the reserved capacity temporarily pushed the account over its regional threshold.

## Diagnosing and Resolving Bedrock ThrottlingException Errors

When an application exceeds its assigned request frequency or token velocity, Amazon Bedrock denies the API request and returns an HTTP 429 status code with the error type ThrottlingException. In automated agent pipelines, an unhandled ThrottlingException halts downstream tool execution, corrupts multi-turn conversational state, and introduces user-facing latency. Resolving these exceptions requires distinguishing between burst request limits, sustained throughput exhaustion, and artificial capacity reservations caused by oversized prompt payloads. Engineering teams must systematically identify whether failures stem from burst request pacing, context payload bloat, or excessive token reservations. Implementing structured logging, client-side retry strategies, and real-time operational alarms ensures that high-volume applications remain resilient when operating close to regional service quota boundaries. A disciplined troubleshooting methodology prevents unnecessary quota increase requests by identifying architectural inefficiencies early.

### Common Root Causes of HTTP 429 Status Codes

Investigating a ThrottlingException begins with identifying which operational boundary was crossed:

1. **Burst Concurrency Spikes:** Firing parallel agent workers simultaneously without client-side pacing depletes the token bucket before replenishment occurs.
2. **Context Payload Bloat:** Submitting full-text PDF documents, technical manuals, or dense logs in prompt context exhausts regional TPM limits within a handful of calls.
3. **Unbounded max_tokens Values:** High max_tokens values lock excessive token allowances during generation, causing false throttling during concurrent operations.
4. **Multi-Service Account Collisions:** Multiple applications, automated evaluation scripts, and internal developers sharing a single AWS account in the same region draw down the same regional quota pool.

### Exponential Backoff and Decorrelated Jitter Implementation

Standard error handling must incorporate exponential backoff with randomized jitter to manage transient throttling gracefully. Simply sleeping for a static interval causes worker threads to retry in lockstep, generating a thundering herd problem that repeatedly hammers Bedrock endpoints.

The following Python example configures a Bedrock Runtime client using the AWS SDK, combining adaptive retry modes with custom jitter logic:

```python
import time
import random
import boto3
from botocore.config import Config
from botocore.exceptions import ClientError

# configure bedrock client with adaptive backoff
bedrock_config = Config(
    retries={
        "max_attempts": 8,
        "mode": "adaptive"
    }
)

bedrock_client = boto3.client(
    "bedrock-runtime",
    region_name="us-east-1",
    config=bedrock_config
)

def invoke_bedrock_with_backoff(model_id, payload, max_retries=6):
    base_backoff = 1.0
    maximum_backoff = 32.0
    # begin retry loop
    for attempt in range(max_retries):
        try:
            response = bedrock_client.invoke_model(
                modelId=model_id,
                contentType="application/json",
                accept="application/json",
                body=payload
            )
            return response
        except ClientError as error:
            code = error.response.get("Error", {}).get("Code")
            if code == "ThrottlingException" and attempt < max_retries - 1:
                # calculate full jitter sleep duration
                calculated_cap = min(maximum_backoff, base_backoff * (2 ** attempt))
                sleep_interval = random.uniform(0.5, calculated_cap)
                time.sleep(sleep_interval)
            else:
                raise error
```

The adaptive retry mode monitors incoming server responses and dynamically adjusts client dispatch rates. Combining adaptive retries with randomized backoff intervals gives the underlying token bucket sufficient time to replenish without dropping active requests.

### Monitoring CloudWatch InvocationsThrottled Metrics

Proactive quota management requires tracking operational metrics in Amazon CloudWatch. Amazon Bedrock publishes several operational dimensions under the AWS/Bedrock namespace:

- **Invocations:** Total count of model invocation requests dispatched to the service.
- **InvocationClientErrors:** Count of 4xx responses, including ValidationException and ResourceNotFoundException.
- **InvocationServerErrors:** Count of 5xx internal server errors.
- **InvocationsThrottled:** Count of calls rejected due to rate limit or token quota exhaustion.

Engineering teams should establish CloudWatch Alarms on the InvocationsThrottled metric across every active production model. Setting an alarm threshold when throttled requests exceed a small fraction of total invocation volume provides early warning before end users experience degraded performance or service interruptions.

## The Document Context Dilemma: Why Full Payloads Break Bedrock

The most common reason engineering teams collide with Bedrock rate limits is the document context dilemma. Developers building summarization pipelines, contract analysis tools, or technical support agents often read entire multi-page files from disk and inject them directly into prompt context. While modern models boast context windows large enough to receive these files, stuffing raw documents into inference calls guarantees rapid quota exhaustion.

Passing full documents turns what should be a low-volume transactional workload into a massive token drain that overwhelms regional allocations. Understanding how document size interacts with per-minute rate limits reveals why architectural offloading is necessary for production stability.

### The Arithmetic of Context Exhaustion

Consider an enterprise agent analyzing quarterly filings, vendor agreements, and engineering specifications. A typical forty-page PDF document contains roughly twenty thousand tokens once extracted. If an application attaches two such documents to an invocation prompt to compare indemnification clauses or technical specifications, that single API call consumes forty thousand input tokens.

In an AWS account with a regional quota of two hundred thousand tokens per minute, only five concurrent agent executions will completely consume the per-minute token allotment. If a sixth user or parallel worker thread initiates a query within that same sixty-second window, Amazon Bedrock immediately throws a ThrottlingException.

The problem compounds when multi-agent frameworks run autonomous loops. An agent that queries a model four times in sequence to verify, reflect, and format findings consumes massive token volumes for a single user task. High-concurrency production deployments collapse under this payload volume, leading teams to believe they need massive quota increases when their core issue is architectural waste.

### Claude File Limits versus Bedrock Token Ceilings

Engineers transitioning from consumer chat interfaces to cloud APIs frequently carry mistaken assumptions regarding file handling. In Anthropic Claude chat, Anthropic restricts direct chat uploads to 20 files per conversation, as detailed in the [Anthropic Claude file upload guide](https://support.claude.com/en/articles/8241126-upload-files-to-claude). Claude Projects accepts individual files with no fixed file count cap, bounded solely by the total context window.

Because web applications accept multiple document uploads without immediate failure, developers often expect the underlying Bedrock APIs to handle full document ingestion with equal ease. However, cloud infrastructure evaluates requests on velocity rather than static file size. While Claude can hold an entire manual in memory, submitting that manual repeatedly through Bedrock InvokeModel calls drains your account's regional TPM allowance in seconds.

Attempting to scale document processing by feeding raw attachments directly into API prompts creates severe operational bottlenecks. Even if AWS grants a substantial quota increase, multiplying token consumption inflates operational costs and increases latency, as models spend seconds processing repetitive background text rather than focusing on the user query.

### Cross-Region Inference Profiles and Their Operational Limits

To help mitigate regional capacity constraints, AWS provides Cross-Region Inference profiles for Amazon Bedrock. By specifying a system-defined inference profile identifier (such as routing traffic across us-east-1 and us-west-2), Bedrock dynamically distributes invocation requests across multiple AWS regions based on real-time capacity and utilization.

Cross-Region Inference profiles effectively multiply your aggregate throughput envelope, providing higher burst resilience during traffic spikes. However, cross-region routing does not solve the underlying inefficiency of oversized document payloads. If every request continues to transmit tens of thousands of unindexed document tokens, the system will eventually saturate the secondary regions as well. Multi-region routing should serve as a high-availability safeguard, not a substitute for lean prompt payloads.

## Offloading Document Retrieval to Intelligent Workspaces via MCP

The definitive solution to Bedrock rate limits is decoupling document persistence and retrieval from foundation model invocations. Rather than passing entire source documents into prompt payloads, applications should store files in a dedicated workspace, index them for semantic search, and retrieve only the precise passages required to answer a specific inquiry. This architectural shift slashes prompt token volume, eliminates ThrottlingException errors, and dramatically lowers model invocation latency. Implementing an intelligent workspace allows agentic teams to collaborate across persistent records without exhausting cloud model quotas. By shifting document processing from repetitive in-prompt injection to on-demand semantic retrieval, teams preserve valuable inference capacity for synthesis and reasoning tasks, keeping operating costs predictable. Offloading document context creates an elastic retrieval tier that scales smoothly across multiple autonomous agents.

### Decoupling Storage from Foundation Model Invocations

Engineering teams have traditionally addressed document retrieval through two suboptimal paths:

- **Commodity Cloud Storage:** Storing documents in standard cloud buckets or shared consumer drives keeps files persistent, but these platforms lack built-in semantic retrieval. Agents must download whole files locally and parse them on every turn, recreating the payload bloat problem.
- **Custom Vector Infrastructure:** Provisioning standalone vector databases like OpenSearch Serverless, Pinecone, or pgvector requires building and maintaining custom chunking pipelines, embedding generation scripts, metadata sync workers, and infrastructure monitoring.

A unified alternative is an intelligent workspace platform that combines persistent, versioned storage with native retrieval interfaces designed for AI agents. Learn more about coordinating agents in shared environments with [Fast.io workspaces](/product/workspaces/).

### Fast.io Remote MCP Architecture and Intelligence Mode

Fast.io provides an intelligent cloud workspace designed for agentic teams and multi-agent systems. Instead of embedding complete files into Bedrock requests, teams upload their documents directly to Fast.io or import them via URL from Google Drive, OneDrive, Box, or Dropbox. Discover how autonomous assistants interact with persistent files on the [Fast.io for AI agents](/storage-for-agents/) hub.

When Intelligence Mode is enabled on a Fast.io workspace, the platform automatically indexes all contained documents, including PDFs, spreadsheets, technical presentations, and rich text files. Fast.io performs hybrid search across the workspace, combining full-text keyword indexing with semantic vector search.

AI agents and application backends interact with the workspace through Fast.io's remote Model Context Protocol (MCP) server:

```json
{
  "mcpServers": {
    "fastio": {
      "url": "https://mcp.fast.io/mcp",
      "headers": {
        "Authorization": "Bearer YOUR_FASTIO_API_KEY"
      }
    }
  }
}
```

Fast.io exposes Streamable HTTP at /mcp and legacy SSE at /sse. When an agent requires context to answer a prompt, it calls Fast.io's consolidated MCP toolset to query the workspace. Fast.io executes a hybrid search and returns only the relevant paragraphs, complete with precise source document citations.

By retrieving targeted context instead of transmitting a raw document archive, prompt payload volume drops dramatically. This keeps token velocity well below Bedrock regional TPM ceilings, preventing ThrottlingException errors while reducing model inference costs.

Teams can review version histories per file, inspect append-only audit logs to track changes made by autonomous agents, and use ownership transfer to hand off completed workspaces from building agents to human stakeholders. Creating a user account is free; doing real work requires an organization on a paid subscription. Monthly plans start with a trial of up to 30 days (credit card required); annual plans have no trial. Review plans and subscription options on the [Fast.io pricing page](/pricing/).

### Metadata Views for Structured Data Without Token Overhead

When applications require structured extraction from business documents like invoices, vendor agreements, and compliance records, stuffing files into multimodal vision models burns enormous token allowances.

Fast.io provides [Metadata Views](/product/document-data-extraction/) to solve this structured processing challenge. With Metadata Views, teams describe the specific fields they need in natural language. The system creates a typed schema supporting text, integers, decimals, booleans, dates, URLs, and JSON objects, automatically extracting structured values into a sortable, filterable spreadsheet.

Agents can query structured metadata directly via MCP tools without reprocessing source files or consuming Bedrock prompt tokens. This separates heavy document parsing from real-time agent decision loops, preserving Bedrock rate limits for interactive reasoning tasks.

## Requesting Quota Increases and Provisioning Bedrock Capacity

While architectural optimization through intelligent retrieval resolves the vast majority of throttling issues, growing production workloads will eventually require higher baseline throughput. AWS provides administrative pathways to scale foundation model quotas through the Service Quotas console and dedicated capacity commitments. Planning for capacity expansions requires measuring actual traffic patterns and submitting justified quota requests before launching high-concurrency features. Combining justified quota headroom with lean payload architectures provides the strongest foundation for enterprise stability, ensuring that unexpected traffic surges do not disrupt active user workflows or autonomous background tasks. Proactive capacity governance enables engineering organizations to scale production agent deployments smoothly across growing user populations while maintaining rigorous budget predictability and meeting strict customer service level agreements.

### Navigating the AWS Service Quotas Console

To request an increase for Amazon Bedrock invocation limits, follow these administrative steps:

1. **Sign in to the AWS Console:** Log in to your AWS account and navigate to the **Service Quotas** console.
2. **Locate Amazon Bedrock:** Select **AWS services** from the sidebar navigation and enter **Amazon Bedrock** into the search field.
3. **Identify Model Quotas:** Filter the quota list by your target foundation model. Bedrock displays separate entries for on-demand requests per minute and tokens per minute for each model family, such as Anthropic Claude 3.5 Sonnet.
4. **Request Quota Increase:** Select the quota you wish to modify and click **Request increase at account level**.
5. **Provide Justification:** Enter your desired limit value and include a detailed operational justification. Describe your application use case, expected concurrent user counts, client-side retry mechanisms, and context reduction strategies.

AWS reviews quota requests based on your account spend history, operational tier, and regional infrastructure capacity. Requests submitted with detailed mathematical justifications and demonstrated rate-smoothing architectures receive faster review and approval from AWS support teams.

### Calculating Production RPM and TPM Requirements

Accurately calculating required throughput prevents over-requesting quotas or falling short during peak traffic periods. Use the following formula to estimate your production requirements:

```text
Required TPM = Peak Concurrent Users * (Average Prompt Tokens + max_tokens) * Requests Per User Per Minute
```

Consider an application supporting fifty concurrent users where each user dispatches two queries per minute. If prompts carry raw documents averaging twenty thousand tokens and max_tokens is set to four thousand, the calculation yields substantial quota requirements:

```text
50 users * (20,000 + 4,000 tokens) * 2 requests/min = 2,400,000 TPM
```

Requesting multi-million TPM allowances requires extensive justification and may face regional availability limits. By contrast, if the application decouples document storage using Fast.io MCP retrieval, average prompt context shrinks to eight hundred tokens:

```text
50 users * (800 + 1,000 tokens) * 2 requests/min = 180,000 TPM
```

By eliminating the vast majority of prompt context, the application comfortably operates within standard tier quotas, keeping infrastructure costs manageable and approval timelines minimal.

### Provisioned Throughput versus Dynamic MCP Workspaces

For mission-critical enterprise applications that require guaranteed throughput without any possibility of throttling, AWS offers Provisioned Throughput. Instead of sharing multitenant on-demand pools, customers purchase dedicated Model Units (MUs) for specific models.

Provisioned Throughput guarantees dedicated processing capacity and custom model deployment options. However, Provisioned Throughput requires substantial financial commitments, often involving one-month or six-month minimum terms that cost substantial amounts per Model Unit monthly. If your application traffic fluctuates, unutilized provisioned capacity remains fully billed.

For the vast majority of engineering teams, combining on-demand Bedrock invocations with intelligent workspace retrieval via Fast.io delivers the optimal balance. By keeping prompt payloads lean, applications avoid regional throttling, maintain low response latency, and scale dynamically without paying for idle Model Units.

## Frequently asked questions

### What are the rate limits for Amazon Bedrock?

Amazon Bedrock enforces rate limits at the regional and account level through AWS Service Quotas. Limits are divided into Requests Per Minute (RPM) and Tokens Per Minute (TPM) for each foundation model family. Default quotas vary based on your AWS account history, spend tier, and regional capacity. Both the bedrock-runtime and bedrock-mantle endpoints maintain independent quota pools for supported models.

### How do I fix AWS Bedrock ThrottlingException?

To fix a Bedrock ThrottlingException (HTTP 429), implement exponential backoff with randomized jitter in your SDK client to absorb traffic bursts. Check your API payloads to ensure max_tokens is set to a realistic response size rather than high defaults, preventing excessive token reservations. Finally, reduce input prompt size by offloading large documents to an indexed workspace via MCP rather than passing raw file text in prompts.

### How to increase AWS Bedrock model quotas?

To increase your Bedrock quotas, sign in to the AWS Management Console and navigate to Service Quotas. Select Amazon Bedrock under AWS services, search for your target model, and choose either Requests Per Minute or Tokens Per Minute. Click Request increase at account level, enter your required capacity, and submit an operational justification detailing your concurrent user volume and architectural controls.

### Why does max_tokens trigger unexpected throttling in Bedrock?

When an invocation request begins, Amazon Bedrock temporarily reserves quota capacity equal to the input prompt tokens plus your max_tokens parameter before text generation completes. If max_tokens is left at a high default like 64,000, concurrent requests lock massive amounts of quota against your regional TPM limit, causing immediate ThrottlingException errors even if the model only produces a few dozen output tokens.

### How does MCP workspace retrieval prevent Bedrock rate limit errors?

Rather than attaching full multi-page documents to model prompts, an intelligent workspace indexes your files and exposes semantic search over the Model Context Protocol (MCP). Your application queries Fast.io via the tools documented on the [Fast.io for AI agents](/storage-for-agents/) hub to retrieve only the specific paragraphs relevant to the user query. This reduces prompt payloads from tens of thousands of tokens down to a few hundred, preventing TPM quota exhaustion.

## Sources

- [Anthropic: Upload files to Claude](https://support.claude.com/en/articles/8241126-upload-files-to-claude) — Anthropic restricts direct chat uploads to 20 files per conversation.
- [AWS: Amazon Bedrock endpoints and quotas](https://docs.aws.amazon.com/general/latest/gr/bedrock.html) — AWS defines service quotas as regional limits governing the maximum operations and resources for an account.

## About Fast.io

Fast.io provides shared workspaces where people and AI agents work on the same files, with built-in semantic search and citation-backed chat over what they hold. Agents reach it through a remote MCP server at https://mcp.fast.io/mcp, a REST API at https://api.fast.io/current/, and a command line client published on npm as @vividengine/fastio-cli.
