Gemini API Rate Limits: Tier Quotas, 429 Handling, and Large-Payload Workflows
Gemini API rate limits govern requests per minute (RPM), tokens per minute (TPM), and requests per day (RPD) across Google AI Studio tiers. While chat interactions rarely cross request thresholds, multi-agent pipelines and document workflows routinely trigger HTTP 429 errors by exceeding token quotas. Implementing exponential backoff and external workspace indexing prevents throttling.
How Gemini API Rate Limits Work Across Free and Paid Tiers
Gemini API rate limits dictate the maximum requests per minute (RPM), tokens per minute (TPM), and requests per day (RPD) an application can issue to Google AI models before receiving HTTP 429 rate limit responses. Google enforces these boundaries across all Gemini model variants, evaluating incoming traffic concurrently against every quota dimension. If a production service exceeds even one ceiling, the gateway halts subsequent calls immediately with a 429 RESOURCE_EXHAUSTED status code.
Rate limits are evaluated per Google Cloud project rather than per individual API key. Creating multiple API keys within the same project does not expand total throughput because every key draws against the identical shared quota allocation. Requests per day quotas reset daily at midnight Pacific time, while requests per minute and tokens per minute calculate on a rolling sixty-second window.
Google does not publish a static per-model quota table. The documentation states that rate limits depend on a variety of factors, including your usage tier, and directs developers to view their active limits in Google AI Studio. Limits also vary by model, and some apply only to specific models: image generation models are metered in images per minute (IPM), and some models carry a tokens-per-day (TPD) limit. Experimental and preview models are more restricted than generally available ones.
The table below outlines what each dimension governs and where the authoritative value lives:
Because Google notes that specified rate limits are not guaranteed and actual capacity may vary, production services should read their own limits programmatically or from AI Studio rather than hard-coding figures found in third-party write-ups. Engineering teams moving from prototype scripts to automated multi-agent environments often assume that request volume is their main constraint. In practice, token consumption during document-processing jobs exhausts rate limits far earlier than request counts.
Requests per Minute, Tokens per Minute, and Daily Ceilings
The three rate limit dimensions serve distinct operational functions:
- Requests Per Minute (RPM) restricts network connection velocity. This boundary guards inference servers against connection floods and high-concurrency client loops.
- Tokens Per Minute (TPM) restricts total compute density. This threshold tallies all input tokens sent in the prompt plus generated output tokens. For long-context models, input token volume dominates consumption.
- Requests Per Day (RPD) limits total cumulative requests over a twenty-four hour period. This metric provides a hard budget ceiling on unpaid exploratory usage, resetting each day at midnight Pacific time.
Because Google evaluates every request against all active boundaries, an application can fail on TPM while staying well below its RPM ceiling. A single prompt carrying a few hundred thousand tokens can consume an entire minute's token allocation on a Free tier project, causing subsequent calls to fail until the rolling minute window elapses, even though the same project has used only one of its permitted requests.
Spend Caps and Rolling Rate Limits
For paid accounts, Google imposes rolling spend-based rate limits alongside RPM and TPM quotas. These spending thresholds evaluate across rolling ten-minute intervals to protect developer billing accounts against unexpected financial runaways caused by infinite loops or compromised credentials.
Google publishes these spend rate limits per usage tier, evaluated on a rolling ten-minute window: none on Free, $10 on Tier 1, $50 on Tier 2, and $200 on Tier 3. Whether they apply to a given account depends on its billing history and standing. If burst traffic consumes inference budget faster than the allowable velocity, the platform issues an HTTP 429 RESOURCE_EXHAUSTED response even when nominal RPM and TPM limits remain unreached. Upgrading through higher billing tiers expands these rolling windows, providing smoother capacity for high-throughput production services.
Related guides
- Cohere Rate Limits: API Keys, Production Tiers, and 429 HandlingCohere rate limits are programmatic caps on the number of requests per minute (RPM) and tokens per minute (TPM) that a...
- Roo Code Rate Limits: Token Exhaustion, Provider Quotas, and MCP WorkspacesRoo Code rate limit refers to API rate limits (HTTP 429) hit when Roo Code's multi-step agent modes issue rapid...
- Google AI Studio Rate Limits: Free Tier Quotas, TPM, and Handling 429 ErrorsGoogle AI Studio rate limits enforce operational caps across requests per minute, tokens per minute, and daily request...
- Groq API Rate Limits: LPU Tier Quotas, TPM Ceilings, and Document HandlingGroq rate limit policies govern API throughput across GroqCloud LPUs through concurrent requests, requests per minute,...
- LiteLLM Rate Limits: RPM, TPM, and Upstream Gateway WorkaroundsLiteLLM rate limits define the maximum requests per minute (RPM) and tokens per minute (TPM) enforced on individual...
- Pinecone Rate Limits: Read Units, Write Units, and Vector Indexing LimitsPinecone rate limits represent throughput caps expressed in Read Units (RUs) and Write Units (WUs) that constrain how...
More on this subject: Agent Security and Governance (51 guides)
What You Need to Know About Gemini API Quota Tiers
Understanding Gemini API quota limits across project environments requires examining how Google enforces both concurrency and daily budget thresholds. Access to the Gemini API is organized into structured usage tiers that determine model availability, concurrency, and maximum rate limits. Google provides initial access to the Gemini API free of charge with baseline quotas before scaling into pay-as-you-go tiers for production applications. Moving between tiers occurs automatically as an organization links billing credentials and establishes verifiable payment history.
The progression path across account tiers follows specific billing qualification milestones:
Qualification for Tiers 2 and 3 is based on total cumulative spending on Google Cloud services for the billing account linked to your project, not on Gemini API usage alone. Transitioning from the Free tier to Tier 1 takes effect almost immediately upon attaching a valid billing instrument in Google AI Studio or the Google Cloud console. Subsequent tier advancements take effect within ten minutes of meeting the criteria, although Google notes that an upgrade request can still be denied in rare cases after review.
Free Tier Operational Constraints
The Free Tier provides a sandbox for evaluation, syntax experimentation, and small-scale automation. However, teams evaluating the Free Tier must account for specific operational trade-offs:
- Stricter RPM and TPM Boundaries. Low token allocations prevent concurrent agent processing and large-file analysis.
- Training Data Usage. Under Google terms of service for the Free Tier, prompt inputs and model outputs may be reviewed by human annotators and used to train Google products.
- No Service Level Agreement. Production availability is not guaranteed on unpaid endpoints, and capacity can be dynamically throttled during regional peak hours.
For teams building commercial applications or handling private company files, linking a billing account to enter Tier 1 is necessary to ensure data privacy and establish reliable throughput.
Interactive Inference, Priority Inference, and Batch Processing
Google segments Gemini API traffic into distinct processing pathways, each carrying different rate limit behaviors:
- Interactive Traffic: Standard synchronous requests sent to the
generateContentendpoint. These calls are metered directly against standard RPM and TPM quotas. - Priority Inference: A dedicated routing option designed for mission-critical workloads requiring minimal latency variance. Priority consumption holds its own rate limits, documented as 0.3 times the standard rate limit for each model and tier, while still counting toward overall interactive traffic.
- Batch API: Asynchronous offline processing designed for bulk data jobs. Batch requests run asynchronously for bulk offline processing at a discounted rate, handling high-volume workloads without drawing down real-time interactive TPM allocations.
How to Diagnose and Resolve HTTP 429 Rate Limit Errors
When an application exceeds an active rate limit, the Gemini API returns an HTTP 429 response accompanied by the status string RESOURCE_EXHAUSTED. Resolving these errors requires inspecting the specific error message to determine which quota boundary was breached.
A standard Gemini API 429 error payload typically provides diagnostic context:
{
"error": {
"code": 429,
"message": "Resource has been exhausted (e.g. check quota).",
"status": "RESOURCE_EXHAUSTED",
"details": [
{
"@type": "type.googleapis.com/google.rpc.QuotaFailure",
"violations": [
{
"subject": "project:1029384756",
"description": "Quota exceeded for quota metric 'GenerateContent requests per minute' and limit 'GenerateContent requests per minute per project'."
}
]
}
]
}
}
The description field reveals whether the failure stems from requests per minute, tokens per minute, requests per day, or short-term spend rate limits.
Implementing Exponential Backoff with Jitter
The first line of defense against transient 429 errors is an automated retry policy incorporating truncated exponential backoff and randomized jitter. Retrying immediately upon receiving an HTTP 429 error creates a cascading thundering herd condition that compounds server congestion.
Here is a resilient retry implementation in Python using standard libraries:
import time
import random
import requests
def call_gemini_with_backoff(url: str, payload: dict, headers: dict, max_retries: int = 5) -> dict:
base_delay = 1.0
max_delay = 32.0
# Retry loop with exponential backoff
for attempt in range(max_retries):
response = requests.post(url, json=payload, headers=headers)
# Check success
if response.status_code == 200:
return response.json()
# Handle rate limiting
if response.status_code == 429:
retry_after = response.headers.get("Retry-After")
if retry_after:
delay = float(retry_after)
else:
calculated_delay = min(max_delay, base_delay * (2 ** attempt))
delay = random.uniform(0, calculated_delay)
time.sleep(delay)
continue
# Terminate on non-rate-limit errors
response.raise_for_status()
# Fallback exception
raise RuntimeError("Exceeded maximum retry attempts due to rate limits.")
Adding randomized jitter prevents synchronized worker threads from striking the API endpoint simultaneously on repeated intervals.
Client-Side Rate Limiting with Token Buckets
Reactive retries prevent crashes, but proactive rate limiting prevents 429 errors from occurring in the first place. In Node.js or TypeScript agent services, implementing a client-side token bucket or leaky bucket queue coordinates concurrent worker tasks before network transmission.
Here is a practical client-side rate limiter in TypeScript:
export class TokenBucketRateLimiter {
private tokens: number;
private lastRefill: number;
private readonly capacity: number;
private readonly refillRatePerSecond: number;
// Initialize limiter
constructor(capacity: number, refillRatePerSecond: number) {
this.capacity = capacity;
this.tokens = capacity;
this.refillRatePerSecond = refillRatePerSecond;
this.lastRefill = Date.now();
}
// Refill tokens
private refill(): void {
const now = Date.now();
const elapsedSeconds = (now - this.lastRefill) / 1000;
this.tokens = Math.min(this.capacity, this.tokens + elapsedSeconds * this.refillRatePerSecond);
this.lastRefill = now;
}
// Acquire execution tokens
async acquire(cost: number = 1): Promise<void> {
while (true) {
this.refill();
if (this.tokens >= cost) {
this.tokens -= cost;
return;
}
const deficit = cost - this.tokens;
const waitMs = Math.ceil((deficit / this.refillRatePerSecond) * 1000);
await new Promise((resolve) => setTimeout(resolve, Math.max(50, waitMs)));
}
}
}
Integrating a token bucket limiter directly into your agent dispatch pipeline guarantees that outbound requests respect defined RPM constraints regardless of how many parallel agent tasks trigger simultaneously.
Decouple Gemini Pipelines from Token Rate Limits
Store and index document archives in shared, persistent workspaces. Query via MCP tools to send compact prompts and protect your API quotas. Monthly plans start with a trial of up to 30 days (credit card required); annual plans have no trial.
Why Large Document Payloads Trigger TPM Limits First
A critical gap in standard developer documentation is the failure to explain how document processing destabilizes token quotas. Understanding how Gemini RPM TPM limits interact during document-processing jobs is essential for preventing unexpected downtime. Developers review rate limit charts showing 1,000 RPM on Tier 1 and conclude that their system has ample headroom. However, when deploying autonomous agents that analyze real-world business documents, applications hit tokens per minute (TPM) limits long before approaching connection limits.
Consider a practical document analysis pipeline. A single 200-page scanned legal agreement, insurance policy package, or technical handbook routinely translates into 250,000 to 400,000 tokens when converted to text and attached to a prompt.
If an orchestrator launches three parallel agent tasks to cross-examine clauses across that file, a single synchronized execution batch emits hundreds of thousands of tokens per batch. On a Free tier project, the very first request can exceed the minute's token capacity outright. Even on a paid tier with a far larger allocation, multiple concurrent agent passes over the same document will exhaust the per-minute budget, because each pass re-transmits the whole file.
Naive document handling introduces three systemic failure points:
- Silent Pipeline Stalls: When TPM exhaustion occurs, all background agent tasks halt simultaneously, waiting out the sixty-second sliding window.
- Context Degradation: Flooding an LLM prompt with hundreds of pages of unindexed text introduces retrieval noise, increasing the probability of hallucinated answers.
- Compounded Financial Costs: Repeatedly re-transmitting entire document files in every conversational turn burns inference budget rapidly, triggering short-term spend rate limits.
Why File Stuffing Fails at Scale
Many developer tutorials recommend uploading raw files directly into the Gemini Files API or converting documents into base64 strings embedded inside prompt payloads. While convenient for one-off scripts, this approach breaks down in autonomous workflows.
When multiple autonomous agents collaborate on a project, each agent requires access to specific factual excerpts. If every agent loads the complete source corpus into its context window, token consumption scales linearly with both corpus size and agent count. An agency running five specialist agents against a 500-page technical archive will consume millions of tokens per task cycle, producing constant 429 bottlenecks and unnecessary API expense.
Trade-Offs of Traditional Local Retrieval Systems
To mitigate context stuffing, engineering teams frequently attempt to build local vector database pipelines using open-source embedding models and local vector stores. While this architecture reduces prompt payload sizes, it introduces substantial infrastructure complexity:
- Document Chunking and Parsing: Engineering teams must build and maintain custom parsers for PDF layouts, slide decks, spreadsheet grids, and scanned images.
- Vector Database Operations: Managing standalone vector indexes requires operational overhead, cluster monitoring, memory tuning, and embedding synchronization.
- Synchronization Lag: When source files are updated by human collaborators, local vector databases fall out of sync unless custom polling daemons or webhooks are constructed.
Instead of building custom chunking and retrieval infrastructure from scratch, modern agent architectures decouple file persistence and semantic search into a dedicated workspace layer.
How to Decouple Document Storage with External Workspace Indexing
The reliable pattern for scaling Gemini agent pipelines is to separate document storage and semantic indexing from model inference. Rather than stuffing large files into prompt contexts or managing custom vector databases, development teams place source corpora into persistent, intelligent workspaces.
In an intelligent workspace, files are indexed automatically upon arrival for full-text and semantic retrieval. Instead of transmitting a 300-page operational manual into a Gemini prompt, an agent connects to the workspace, executes a targeted search query, and extracts only the relevant text fragments. Passing three focused paragraphs into the prompt consumes 800 tokens instead of 300,000 tokens, preserving your Gemini TPM rate limit allowance for core application traffic.
Fast.io delivers this workspace architecture for agentic teams. Workspaces provide shared, organization-owned file storage where humans and AI agents interact seamlessly. Files can be uploaded directly or imported through cloud storage integrations. Cloud Sync ships for Dropbox, Box, and OneDrive, while Google Drive is import today, with sync coming soon.
Once stored in a workspace, documents are processed by Intelligence Mode, which builds a hybrid search index combining semantic meaning and keyword matching. Autonomous agents access the workspace using standard protocols through the remote Fast.io Model Context Protocol (MCP) server at https://mcp.fast.io/mcp (or https://mcp.fast.io/mcp/key with Bearer token authentication).
By retrieving only indexed excerpts through MCP tooling, multi-agent systems process thousands of document queries daily without approaching Google API rate limits. Teams working with structured file extraction can also use Metadata Views to turn document collections into typed, queryable datasets without manual OCR rules.
Collaborative integrity is preserved through built-in operational features:
- Per-File Version History: Concurrent agent writes and updates never overwrite prior states destructively. Every revision is preserved with rollback capability.
- Granular Access Permissions: Permissions configure cleanly across organization, workspace, folder, and file tiers.
- Append-Only Audit Log: Every file access, search query, and metadata modification is tracked permanently for administrative oversight.
- Ownership Transfer: Autonomous agents can provision workspaces, index client records, and hand off ownership cleanly to human stakeholders.
Creating an account is free; doing real work requires an organization on a paid subscription. Monthly plans start with a trial of up to 30 days (credit card required); annual plans have no trial. Subscription plans are Starter at $9.99/mo, Business at $49.99/mo, and Enterprise at $199.99/mo on Fast.io pricing. For developers building agent integrations, documentation and tool references are available in the Fast.io for Agents documentation, with technical schema details at https://mcp.fast.io/skill.md and onboarding guides at https://fast.io/llms.txt.
Configuring Gemini Agents with Fast.io MCP
Connecting an AI agent or custom Python script to Fast.io requires only configuring the remote MCP endpoint. Agents connect over Streamable HTTP at /mcp or legacy Server-Sent Events at /sse.
Here is an example MCP client configuration snippet:
{
"mcpServers": {
"fastio": {
"url": "https://mcp.fast.io/mcp/key",
"headers": {
"Authorization": "Bearer YOUR_FASTIO_API_KEY"
}
}
}
}
When an agent needs to answer an inquiry regarding a company policy or technical specification, it issues a search query against the workspace. The MCP server returns precise, citation-backed snippets that the agent inserts into its Gemini prompt. The resulting payload stays exceptionally small, eliminating rate limit throttling.
Sources
References used to verify factual claims in this guide.
-
Gemini API rate limits are applied per Google Cloud project rather than per API key, with daily request limits resetting at midnight Pacific time. Google does not publish a fixed per-model Gemini rate limit table and directs developers to view their active limits in Google AI Studio. Gemini spend-based rate limits are evaluated on a rolling 10-minute window at $10 on Tier 1, $50 on Tier 2, and $200 on Tier 3.
-
Google provides initial access to the Gemini API free of charge with baseline quotas before scaling into pay-as-you-go tiers for production applications.
Frequently Asked Questions
What is the rate limit for Gemini API?
Gemini API rate limits depend on the model you call and your project's usage tier, and Google does not publish a fixed per-model table. Limits are measured across requests per minute (RPM), input tokens per minute (TPM), and requests per day (RPD), applied per project rather than per API key, with RPD resetting at midnight Pacific time. To see the figures that apply to you, open the rate limit page in Google AI Studio or read them with the Rate Limits API.
How do I fix Gemini API 429 rate limit exceeded?
Fixing an HTTP 429 error requires diagnosing whether you breached requests per minute, tokens per minute, requests per day, or rolling spend limits. Implement truncated exponential backoff with randomized jitter in your client code, establish client-side token bucket rate limiters, and link a billing account to advance to Tier 1. To resolve token rate limit exhaustion, decouple large documents into an indexed external workspace instead of attaching full files directly to prompts.
What is the Gemini API free tier request limit?
The Gemini API Free tier enforces model-specific request limits rather than one shared number, and Google surfaces the current values in Google AI Studio instead of publishing them in the rate limits documentation. Faster Flash-class models carry the highest free allowances, while Pro-class reasoning models and any experimental or preview model are restricted much more tightly. Daily quotas reset at midnight Pacific time.
Why do document uploads trigger Gemini TPM limits before RPM limits?
A standard multi-page PDF or technical manual contains hundreds of thousands of tokens. Submitting large document payloads in a prompt or across parallel agent calls consumes the tokens per minute quota immediately, even when total request velocity remains well below the allowable requests per minute threshold.
How does external workspace retrieval reduce Gemini API token consumption?
External workspaces index file corpora on arrival using semantic and keyword search. Instead of passing an entire document into Gemini, agents query the workspace via MCP tools and retrieve only the relevant passages. This approach reduces prompt size from hundreds of thousands of tokens to fewer than one thousand tokens, preserving rate limit headroom.
Related Resources
Decouple Gemini Pipelines from Token Rate Limits
Store and index document archives in shared, persistent workspaces. Query via MCP tools to send compact prompts and protect your API quotas. Monthly plans start with a trial of up to 30 days (credit card required); annual plans have no trial.