LiteLLM Rate Limits: RPM, TPM, and Upstream Gateway Workarounds
LiteLLM rate limits define the maximum requests per minute (RPM) and tokens per minute (TPM) enforced on individual users, keys, or upstream provider endpoints within a LiteLLM proxy deployment. When multi-agent systems trigger upstream 429 errors, automated router fallbacks and cooldowns maintain gateway uptime. Connecting agents to an indexed Fast.io workspace via Model Context Protocol retrieves targeted context, cutting payload sizes and preventing token exhaustion.
How LiteLLM Rate Limits Work: RPM, TPM, and Gateway Controls
In high-throughput multi-agent systems, rate limit exhaustion almost never stems from sending too many requests; it stems from sending too many tokens per request. When multiple agents concurrently re-inject full document contexts into an LLM gateway, they saturate provider token-per-minute (TPM) ceilings within seconds, turning upstream 429 Too Many Requests errors into systemic pipeline failures.
LiteLLM rate limits define the maximum requests per minute (RPM) and tokens per minute (TPM) enforced on individual users, keys, or upstream provider endpoints within a LiteLLM proxy deployment. Rather than treating an LLM gateway as a simple reverse proxy, LiteLLM functions as an active traffic coordinator. The proxy tracks consumption across both request counts and token volumes before dispatching calls to upstream model providers.
Core Metrics: RPM Versus TPM
The LiteLLM gateway meters traffic along two primary throughput axes:
- Requests Per Minute (RPM): The total count of HTTP requests completed within a rolling 60-second window. This metric protects against connection flooding, worker pool exhaustion, and upstream request frequency quotas.
- Tokens Per Minute (TPM): The combined sum of prompt input tokens sent to the model and completion output tokens generated by the model within a rolling 60-second window. Because modern frontier models bill and throttle on token consumption, TPM limits represent the primary operational bottleneck for teams running document processing or agentic coding loops.
In recent proxy releases, LiteLLM also introduces separate Input Tokens Per Minute (ITPM) and Output Tokens Per Minute (OTPM) rate limiting. This distinction is valuable for agent architectures: agents frequently transmit extensive reference documents as inputs while expecting short JSON schema responses. Separating input token metering from output token metering prevents large reading tasks from blocking output-focused workers.
Default Rate Limit Configurations
A production deployment of LiteLLM proxy balances rate limits across four distinct tiers:
- Global Proxy RPM and Budgets: Proxy-wide ceilings set in
config.yamlundergeneral_settingsorlitellm_settings. These settings define hard limits on total gateway throughput and monthly spend across all connecting teams. - User-Level and Key TPM: Granular allocations configured on virtual API keys (
/key/generate), teams (/team/new), or customer identifiers (/customer/new). These rules isolate tenants so that a runaway agent loop running on one key cannot starve other teams of API access. - Model Fallback Thresholds: Upstream reliability settings under
router_settings(num_retries,allowed_fails, andcooldown_time). When an upstream provider returns an HTTP 429 error, LiteLLM temporarily removes the failing deployment from rotation and reroutes requests to alternative model providers across 100+ LLMs. - Workspace MCP Offloading: Structural payload reduction achieved by indexing reference files in an external workspace and querying them through Model Context Protocol (MCP). By fetching concise, citation-backed excerpts instead of transmitting raw files in prompts, agents avoid consuming their per-minute token allowances.
Scope of Enforcement Across Proxy Entities
Enforcing rate limits across distributed proxy instances requires a shared persistence store. While basic model-level request throttling can run in-memory for a single container, virtual keys, user budgets, and team quotas require a connected PostgreSQL database and a Redis cache.
Understanding how these tiers interact enables platform engineers to configure precise guardrails without degrading developer productivity.
Related guides
- Gemini API Rate Limits: Tier Quotas, 429 Handling, and Large-Payload WorkflowsGemini API rate limits govern requests per minute (RPM), tokens per minute (TPM), and requests per day (RPD) across...
- Pinecone Rate Limits: Read Units, Write Units, and Vector Indexing LimitsPinecone rate limits represent throughput caps expressed in Read Units (RUs) and Write Units (WUs) that constrain how...
- Roo Code Rate Limits: Token Exhaustion, Provider Quotas, and MCP WorkspacesRoo Code rate limit refers to API rate limits (HTTP 429) hit when Roo Code's multi-step agent modes issue rapid...
- OpenRouter Rate Limits: 20 RPM Caps, 429 Errors, and Agent Storage WorkaroundsOpenRouter restricts free models to 20 requests per minute and 50 to 1,000 requests per day based on credit purchases,...
- Vertex AI Rate Limits: Understanding Quotas, TPM, RPM, and Token ConstraintsA Vertex AI rate limit is a regional Google Cloud project quota that caps API request frequency (RPM) and token...
- Azure OpenAI Rate Limits: TPM Quotas, PTU Scaling, and 429 Error ResolutionAzure OpenAI rate limits are regional and subscription-level constraints defined by Tokens Per Minute (TPM) and...
More on this subject: Agent Security and Governance (51 guides)
How to Configure Rate Limits in LiteLLM Proxy
Configuring LiteLLM rate limits involves defining model capacity in config.yaml, creating virtual keys through the administrative API, and setting up dynamic allocation algorithms to manage burst traffic.
Step 1: Defining Model-Level Limits in config.yaml
When declaring model deployments in your LiteLLM configuration file, specify the upstream provider limits directly inside litellm_params. This informs the proxy router of the capacity available on each endpoint:
model_list:
- model_name: gpt-4o
litellm_params:
model: openai/gpt-4o
api_key: os.environ/OPENAI_API_KEY
rpm: 500
tpm: 150000
- model_name: gpt-4o
litellm_params:
model: azure/gpt-4o-eastus
api_base: os.environ/AZURE_API_BASE
api_key: os.environ/AZURE_API_KEY
rpm: 1000
tpm: 300000
router_settings:
routing_strategy: "least-busy"
redis_url: os.environ/REDIS_URL
general_settings:
master_key: os.environ/LITELLM_MASTER_KEY
database_url: os.environ/DATABASE_URL
In this setup, LiteLLM routes traffic for gpt-4o across both OpenAI and Azure deployments using a least-busy strategy. The proxy router checks current utilization against the declared RPM and TPM thresholds in Redis before forwarding requests.
Step 2: Creating Teams and Virtual Keys with Rate Limits
To restrict client applications, generate virtual API keys assigned to specific teams or individual agents. The administrative endpoint accepts explicit rpm_limit and tpm_limit parameters.
The following command creates a team with shared capacity:
curl -X POST 'http://localhost:4000/team/new' \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
-H "Content-Type: application/json" \
-d '{
"team_alias": "research-agents",
"rpm_limit": 120,
"tpm_limit": 80000,
"max_budget": 500.0,
"budget_duration": "30d"
}'
The following command generates a virtual key tied to that team:
curl -X POST 'http://localhost:4000/key/generate' \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
-H "Content-Type: application/json" \
-d '{
"team_id": "YOUR_TEAM_ID",
"key_alias": "agent-worker-01",
"rpm_limit": 60,
"tpm_limit": 40000,
"models": ["gpt-4o"]
}'
When a client using this virtual key reaches the configured request frequency or token ceiling within the 60-second sliding window, LiteLLM rejects subsequent calls immediately with an HTTP 429 status code. This check occurs at the proxy gateway, preventing excess traffic from reaching upstream providers.
Step 3: Managing Reusable Rate Limit Tiers
For organizations managing hundreds of developers or automated pipelines, setting custom limits on individual keys becomes unmanageable. LiteLLM supports reusable budget and rate limit tiers via /budget/new.
The following command defines a standard tier template:
curl -X POST 'http://localhost:4000/budget/new' \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
-H "Content-Type: application/json" \
-d '{
"budget_id": "tier-standard",
"rpm_limit": 100,
"tpm_limit": 60000,
"max_budget": 200.0,
"budget_duration": "30d"
}'
Once defined, assign budget_id: "tier-standard" to newly generated keys. Updating the tier parameters in the database updates enforcement across all associated keys immediately.
LiteLLM rate limits are enforced on non-admin internal user roles, while admin users bypass proxy rate limits during testing. Always test virtual keys using standard non-admin credentials to verify that throttling behaves as planned.
Step 4: Dynamic TPM and RPM Allocation
In standard setups, static rate limits often cause artificial throttling: one key might be throttled during a spike even though other keys are idle and substantial upstream model capacity remains unused.
To address this challenge, LiteLLM provides the dynamic_rate_limiter_v3 callback. The dynamic limiter monitors model saturation in real time. While overall usage remains below a configured saturation threshold, any key can use idle capacity beyond its baseline limit. Once cluster usage crosses the saturation point, the limiter tightens enforcement, restricting each key to its reserved allocation.
Enable this behavior in config.yaml:
litellm_settings:
callbacks: ["dynamic_rate_limiter_v3"]
Dynamic allocation maintains high infrastructure utilization during normal operating conditions while preserving strict fairness when traffic surges.
What Happens When Upstream Endpoints Return HTTP 429 Errors
Even with disciplined client-side rate limits, upstream model providers will periodically return HTTP 429 Too Many Requests responses. Upstream throttling can occur due to sudden regional traffic spikes, provider infrastructure degradation, or unpredictable concurrent usage across other accounts on your organization's API credentials.
A production gateway must detect upstream 429 errors promptly and route around them without failing the client application.
The Upstream Error Lifecycle
When LiteLLM forwards an inference call to an upstream provider and receives an HTTP 429 response, it initiates an automated recovery sequence:
(Client Call) --> (LiteLLM Proxy Router) --> (Primary Deployment: OpenAI) --> HTTP 429
|
v
(Read Retry-After Header)
(Increment Failure Counter)
(Enter Deployment Cooldown)
|
v
(Evaluate Router Fallbacks)
|
v
(Alternative Deployment: Azure / Anthropic) --> HTTP 200 OK --> (Client Response)
LiteLLM parses the upstream response headers, checking for standard rate limit telemetry such as Retry-After, x-ratelimit-reset-requests, and x-ratelimit-reset-tokens.
Router Settings: Retries, Cooldowns, and Allowed Fails
To automate failover, configure resilience parameters under router_settings in config.yaml:
router_settings:
routing_strategy: "least-busy"
num_retries: 2
allowed_fails: 3
cooldown_time: 30
fallbacks:
- "gpt-4o": ["azure-gpt-4o", "claude-3-5-sonnet"]
- "claude-3-5-sonnet": ["gpt-4o"]
model_list:
- model_name: gpt-4o
litellm_params:
model: openai/gpt-4o
api_key: os.environ/OPENAI_API_KEY
rpm: 500
- model_name: azure-gpt-4o
litellm_params:
model: azure/gpt-4o-deployment
api_base: os.environ/AZURE_API_BASE
api_key: os.environ/AZURE_API_KEY
rpm: 1000
- model_name: claude-3-5-sonnet
litellm_params:
model: anthropic/claude-3-5-sonnet-20241022
api_key: os.environ/ANTHROPIC_API_KEY
rpm: 400
Here is how these parameters interact:
num_retries: 2: If a request togpt-4oreturns a 429 or network timeout, LiteLLM retries up to twice across other healthy deployments in that model group.allowed_fails: 3: If a deployment fails or returns 429 three times consecutively, LiteLLM marks that specific deployment as unhealthy.cooldown_time: 30: The unhealthy deployment is placed in a 30-second cooldown period. During cooldown, the proxy routes all traffic away from this deployment, allowing its upstream token bucket to recover.fallbacks: If all deployments forgpt-4oare in cooldown or exhaust their retries, the router automatically fails over toazure-gpt-4o, and subsequently toclaude-3-5-sonnet.
Because LiteLLM provides a standardized OpenAI-compatible interface across 100+ LLMs, client code requires zero modification to receive responses generated by fallback models.
Client-Side Headers and Observability
The proxy passes rate limit telemetry back to the client application using dedicated response headers:
x-litellm-key-remaining-requests: Remaining requests allowed on the virtual key in the current window.x-litellm-key-remaining-tokens: Remaining tokens allowed on the virtual key in the current window.x-litellm-model-api-base: The exact upstream deployment endpoint that served the request, allowing developers to verify whether a fallback occurred.
While router fallbacks and exponential backoff keep services online, relying solely on failovers creates new complications. Routing traffic to secondary models increases API billing and can introduce behavioral discrepancies in complex multi-agent workflows. The sustainable engineering solution is to fix the underlying driver of rate limit exhaustion: oversized prompt payloads.
Why Multi-Agent Workflows Exhaust Tokens and How Workspaces Fix It
Most engineering discussions around rate limits focus entirely on client-side retries, backoff algorithms, and fallback routing. While these mechanisms are necessary, they treat the symptom rather than the root cause.
In agentic systems, rate limit failures are almost always caused by prompt bloat. When autonomous agents analyze enterprise knowledge bases, codebases, or legal contracts, developers frequently re-inject tens of thousands of tokens of reference text into every single prompt turn. If five agents collaborate on an analysis, running ten iterations each, an extensive reference context creates massive volumes of input traffic in minutes. Even high-tier upstream accounts quickly exhaust their TPM quotas under this volume.
Limitations of Ad-Hoc File Storage
Engineering teams often attempt to mitigate prompt bloat using crude local workarounds:
- Local Disk Caching: Agents store reference files on local container storage and run basic string matching or keyword search scripts. This approach fails in distributed cloud environments where agent containers spin up and down ephemerally.
- Raw Object Storage: Files are uploaded to an object store like AWS S3 or Google Cloud Storage, and agents fetch entire files via pre-signed URLs. While this solves persistence, the agent must still download and parse the entire file into memory, eventually stuffing the full content back into the LLM prompt.
Neither option provides built-in semantic retrieval or structured search, forcing developers to build and maintain complex custom vector databases and chunking pipelines.
The Fast.io Architecture: Workspaces as an Intelligent Retrieval Substrate
Fast.io provides an intelligent workspace platform designed specifically for agentic teams. Rather than treating storage as passive file hosting, Fast.io transforms files into a queryable knowledge substrate accessible directly by autonomous agents.
When teams upload reference materials to an org-owned Fast.io workspace, Intelligence Mode indexes the files upon arrival. The platform provides hybrid search combining full-text matching, semantic search, and metadata filtering without requiring a separate vector database.
(Autonomous Agent)
|
v Search query: "find warranty termination clause"
(Fast.io Remote MCP Server: https://mcp.fast.io/mcp)
|
v Hybrid search: semantic and full-text index
(Fast.io Intelligent Workspace)
|
v Returns concise relevant excerpt with file citations
(Autonomous Agent)
|
v Compact prompt instead of full document
(LiteLLM Proxy Gateway) --> (Upstream Provider: OpenAI or Anthropic)
Instead of attaching an entire 80-page contract or technical manual to every LiteLLM request, the agent queries the Fast.io workspace using the Model Context Protocol (MCP). The workspace returns only the relevant excerpts accompanied by citation metadata. This reduces prompt payload sizes, replacing massive document contexts with targeted excerpts and preserving upstream token capacity.
Connecting Agents to Fast.io via Remote MCP
Fast.io hosts a remote MCP server accessible over Streamable HTTP at https://mcp.fast.io/mcp (or https://mcp.fast.io/mcp/key when using Bearer token authentication) and legacy SSE at https://mcp.fast.io/sse. Agents connect directly to this endpoint to search workspaces, inspect file structures, and retrieve document excerpts.
Here is a standard MCP configuration for agent environments:
{
"mcpServers": {
"fastio": {
"url": "https://mcp.fast.io/mcp/key",
"headers": {
"Authorization": "Bearer FASTIO_API_KEY"
}
}
}
}
Through this consolidated MCP toolset, agents execute actions such as searching indexed files, reading targeted passages, writing intermediate outputs, and organizing deliverables. For teams needing structured document extraction, Metadata Views turn unstructured PDFs and spreadsheets into queryable tabular databases, allowing agents to filter documents by typed fields before reading. Learn more on the document data extraction product page.
Versioning, Audit Logging, and Ownership Transfer
In production multi-agent environments, governance and file integrity are as critical as throughput:
- Per-File Version History: Every document modification made by an agent or human creates an immutable version record. If an agent overwrites a file or generates an imperfect edit, prior versions can be inspected and restored.
- Append-Only Audit Log: Fast.io records every file read, write, update, and share action in an append-only audit trail, providing complete operational transparency across human and agent actions.
- Ownership Transfer: When an autonomous agent provisions workspaces and generates deliverables for a human client, it can transfer primary ownership of the organization to a human colleague while maintaining scoped administrative access.
Fast.io leaves upstream LLM rate limits intact while providing the intelligent storage layer that prevents prompt bloat from triggering them in the first place.
Monthly plans start with a trial of up to 30 days (credit card required); annual plans have no trial. Explore complete feature sets and start a trial on the Fast.io pricing directory. More technical integration details are available in the agent storage overview.
Eliminate Agent Token Exhaustion with Indexed Workspaces
Stop saturating LiteLLM TPM limits with bloated prompts. Offload reference documents to a Fast.io workspace with built-in semantic search and MCP connectivity. Monthly plans start with a trial of up to 30 days (credit card required); annual plans have no trial.
How to Scale Multi-Agent Gateways Without Cascading Failures
Scaling an agentic fleet behind a LiteLLM proxy requires disciplined architectural standards. Implementing the following production practices prevents cascading rate limit failures and keeps proxy latency predictable.
1. Enforce Per-Agent Key Isolation
When configuring multi-agent systems, avoid sharing a single virtual API key across independent workers. When all agents authenticate with the same key, a single runaway agent entering an infinite reasoning loop will exhaust the entire team's RPM and TPM allocation.
Instead, provision dedicated virtual keys for each agent type using /key/generate. Apply strict rpm_limit and tpm_limit settings to individual worker instances. If a research agent gets trapped in a repetitive tool-calling loop, LiteLLM throttles only that specific worker, leaving the rest of your agent fleet operational.
2. Synchronize State with Distributed Redis
Running LiteLLM across multiple container replicas requires shared state synchronization. If each proxy container tracks rate limits locally in memory, a client can burst traffic beyond its intended limit across separate container instances.
Always configure router_settings.redis_url to point to a high-availability Redis cluster. Redis enables LiteLLM to maintain atomic sliding-window counters for RPM and TPM across all proxy instances, guaranteeing accurate quota enforcement regardless of horizontal scaling.
3. Separate Interactive Streams from Batch Workloads
Multi-agent systems execute two distinct classes of work:
- Low-Latency Interactive Tasks: User-facing chat responses, real-time code completions, and interactive decision steps.
- High-Volume Asynchronous Tasks: Bulk document ingestion, synthetic dataset generation, and automated nightly regression evaluations.
Avoid routing batch workloads through the same model deployments and virtual keys used for interactive users. Configure dedicated batch deployments with conservative RPM ceilings, or use upstream provider Batch APIs through LiteLLM. This prevents background batch processing from consuming the token burst capacity needed by customer-facing agents.
4. Cap Token Generation Proactively
When calculating TPM usage, upstream providers count both prompt input tokens and completion output tokens. Unconstrained completion generation is a frequent cause of unexpected rate limit errors.
Always specify a defensive max_tokens ceiling on all agent completion calls. If an agent task only requires a structured 200-token JSON output, setting max_tokens: 500 prevents the model from generating hundreds of hallucinated lines that consume downstream token quotas.
5. Monitor Cooldown Events and Error Spikes
Use proactive observability to monitor your proxy gateway. Configure LiteLLM's logging callbacks to export metrics to Prometheus, Datadog, or OpenTelemetry. Pay special attention to:
litellm_deployment_cooldown_total: Frequency of deployments entering temporary cooldown states.litellm_proxy_rate_limit_errors_total: Count of 429 errors returned to client applications.litellm_upstream_latency_seconds: Response latency across model deployments.
A sudden spike in cooldown events indicates that upstream provider quotas are undersized for your traffic patterns, signaling that it is time to request quota increases or expand your workspace MCP offloading strategy.
Sources
References used to verify factual claims in this guide.
-
LiteLLM rate limits are enforced on non-admin internal user roles, while admin users bypass proxy rate limits during testing.
-
LiteLLM dynamic rate limiting shares model capacity across keys and teams based on real-time saturation thresholds.
Frequently Asked Questions
How do I set rate limits in LiteLLM proxy?
You can set rate limits in LiteLLM proxy at multiple levels. Model-level limits are defined in config.yaml under model_list using the rpm and tpm parameters in litellm_params. Team and virtual key limits are configured via the administrative REST API using the /team/new and /key/generate endpoints by passing rpm_limit and tpm_limit fields. For multi-tenant applications, you can also define reusable tiers using /budget/new and assign them directly to keys or end-user customer IDs.
What happens when LiteLLM hits an upstream rate limit?
When an upstream model provider returns an HTTP 429 Too Many Requests response, LiteLLM parses the response headers for Retry-After instructions. Under router_settings, if a deployment exceeds the allowed_fails threshold, LiteLLM temporarily places it into a cooldown state for cooldown_time seconds. If retries within that deployment group fail, the proxy router automatically executes configured fallbacks, redirecting the request to an alternative healthy model deployment without interrupting the client application.
How do I prevent LLM rate limit errors in multi-agent workflows?
Preventing rate limit errors in multi-agent workflows requires reducing prompt payload sizes rather than relying solely on retries. Instead of repeatedly stuffing large reference documents into prompt context windows, store reference files in an intelligent Fast.io workspace with Intelligence Mode enabled. Agents query the workspace via Model Context Protocol (MCP) to retrieve compact, citation-backed excerpts, reducing token consumption per turn and keeping total throughput well below upstream TPM ceilings.
What is the difference between RPM and TPM in LiteLLM?
Requests Per Minute (RPM) measures the total number of individual API calls sent to the proxy within a 60-second window, protecting infrastructure from connection flooding. Tokens Per Minute (TPM) measures the aggregate volume of prompt input tokens and completion output tokens processed within that same window. In agentic workflows, TPM limits are reached much faster than RPM limits due to large prompt contexts.
How does LiteLLM dynamic rate limiting share capacity?
LiteLLM dynamic rate limiting, enabled via the dynamic_rate_limiter_v3 callback, monitors real-time model saturation across keys and teams. When total cluster usage is below a defined saturation threshold, any key can consume idle capacity beyond its baseline allocation. When cluster utilization crosses the threshold, the limiter enforces reserved share limits, ensuring fair access during heavy traffic periods.
Does LiteLLM track rate limits across multiple proxy instances?
Yes, LiteLLM tracks rate limits across multiple horizontally scaled proxy instances by connecting to a centralized Redis cluster configured via router_settings.redis_url. Redis provides atomic sliding-window counters for RPM and TPM, ensuring that virtual keys and team quotas are enforced consistently across all container replicas.
Related Resources
Eliminate Agent Token Exhaustion with Indexed Workspaces
Stop saturating LiteLLM TPM limits with bloated prompts. Offload reference documents to a Fast.io workspace with built-in semantic search and MCP connectivity. Monthly plans start with a trial of up to 30 days (credit card required); annual plans have no trial.