Google Gemini Token Limits: 1M Context Windows, 65K Output Caps, and API Quotas
Google Gemini enforces an input token limit of 1,048,576 tokens and an output ceiling of 65,536 tokens on current Gemini 3 and 2.5 models. The Gemini API also imposes rolling tokens per minute (TPM) caps that vary by model and usage tier and are published only in Google AI Studio. Heavy document workflows routinely trigger HTTP 429 errors by exceeding per-minute token throughput. Storing files in Fast.io workspaces lets assistants retrieve indexed excerpts via MCP without prompt bloat.
What Are the Exact Token Limits for Google Gemini Models?
Google Gemini enforces an input limit of 1,048,576 tokens and a maximum completion output of 65,536 tokens per request on current Gemini 3 and 2.5 models, alongside rolling per-minute throughput quotas that Google sets per model and per usage tier and exposes in Google AI Studio rather than in its published documentation. A Gemini token limit defines the maximum input tokens, completion output tokens, and per-minute token throughput (TPM) enforced by the Google Gemini API across developer tiers.
Many engineering teams confuse overall context window capacity with rolling rate limits. While technical discussions highlight the one-million-token input capacity of current Gemini models, an application cannot continuously transmit prompts of that magnitude without colliding with operational rate limits. Google evaluates incoming API requests across multiple distinct token dimensions simultaneously:
- Input Context Window: The cumulative volume of system instructions, conversation history, structured tool declarations, and attached files accepted in a single request.
- Output Generation Limit: The absolute ceiling on completion tokens generated in a single response turn.
- Tokens Per Minute (TPM): The rate-limited throughput ceiling evaluated on a rolling sixty-second window across an entire Google Cloud project.
- Spend-Based Velocity Caps: Monetary rate limits evaluated over rolling ten-minute intervals to protect billing accounts from accidental consumption bursts.
The table below outlines verified context capacities and output limits across Google Gemini developer models as of September 2026:
Earlier Gemini 1.5 models capped completion generations at 8,192 tokens. Current Gemini 2.5 and 3.8 Flash models expand this output ceiling eightfold to 65,536 tokens, enabling extensive code generation, complete document drafting, and long-form structured data transforms.
Input Context Windows Versus Output Generation Limits
The fundamental token boundary of any large language model is its context window. Google documentation states that the context window defines the combined limit of input and output tokens across an interaction. However, the input ceiling and the output cap operate under asymmetric constraints.
On current Gemini 3 and 2.5 models, the input context window reaches 1,048,576 tokens, accommodating extensive multi-file repositories or media recordings. In contrast, completion output is restricted to 65,536 tokens. A model cannot generate a response that matches the size of its maximum input.
This asymmetry shapes software architecture:
- Asymmetric Payloads: You can supply a 900,000-token codebase to Gemini 3.8 Flash, but the model completion output reaches a maximum ceiling of 65,536 tokens of analysis or refactored code in a single API round-trip.
- Context Exhaustion: In a multi-turn session, output tokens from turn one become input tokens for turn two, consuming both the 1,048,576 context window and rolling TPM allowances.
- Finish Reasons: If a generation reaches the 65,536-token ceiling before the model concludes its reasoning, the API terminates the stream with a
finish_reasonofMAX_TOKENS, resulting in truncated text or malformed JSON payloads.
How Multimodal Data Converts to Gemini Tokens
In Google Gemini, non-text media files are converted directly into token representations within the neural network:
- Text and Source Code: For standard English text, Google documents that one token corresponds to approximately four characters, meaning 100 tokens equals roughly 60 to 80 English words. Source code and structured JSON exhibit higher token density due to indentation and symbols.
- Images: Images with dimensions of 384 pixels or smaller consume exactly 258 tokens. Larger images are divided into standardized 768 by 768 pixel tiles, with each tile consuming 258 tokens. An architectural blueprint with complex dimensions can consume numerous token tiles.
- Video Footage: The Gemini API samples video at one frame per second during static processing, with each second converting to approximately 263 tokens. A ten-minute video consumes roughly 157,800 input tokens.
- Audio Streams: Audio tracks tokenize at 32 tokens per second, which equals 1,920 tokens for every minute of recorded speech.
- PDF Documents: Text pages parse as text tokens, while embedded charts and scanned pages convert into image tiles at 258 tokens per tile.
Related guides
- Gemini Message Limits in 2026: Quotas, Cooldowns, and WorkaroundsA Gemini message limit restricts how many prompts or API requests you can submit within a rolling window. Learn the...
- DeepSeek Token Limit: Context Windows, Output Caps, and Token WorkaroundsThe DeepSeek token limit consists of a 1M-token context window for prompt ingestion and a 384K-token output ceiling per...
- Continue.dev Token Limit: Context Window Configuration and Codebase IndexingThe Continue.dev token limit is the maximum context length configured in Continue's config.json or config.yaml file...
- How to Manage LangChain Token Limits: Memory, Pruning, and MCP WorkspacesUnderstanding the LangChain token limit helps developers prevent context overflow errors in complex LLM pipelines. This...
- ChatGPT Character Limits: The 25,000-Character Paste Limit and SolutionsThe ChatGPT web interface enforces a frontend paste limit that triggers truncation warnings or forces file attachment...
- Custom GPT Token Limit: Instruction Limits, Context Windows, and RetrievalThe Custom GPT token limit refers to the 8,000-character constraint on configuration instructions and the dynamic...
More on this subject: AI Agents: General Guides (99 guides)
How Many Tokens per Minute Does the Gemini API Allow Across Tiers?
Rate limits regulate the velocity at which an application can dispatch requests and consume tokens. Google enforces rate limits on a per-project basis rather than per API key. If a developer creates ten separate API keys within the same Google Cloud project, all ten keys share the exact same underlying quota pool.
Google evaluates API traffic across three primary quota dimensions:
- Requests Per Minute (RPM): Restricts API calls dispatched within a rolling sixty-second window.
- Tokens Per Minute (TPM): Restricts total input and output tokens processed across all requests in a rolling sixty-second window.
- Requests Per Day (RPD): Imposes a daily request volume cap that resets each night at midnight Pacific time.
Google no longer publishes a fixed per-model quota table. Its rate limits documentation states that limits depend on factors such as your usage tier and can be viewed in Google AI Studio, and adds that specified rate limits are not guaranteed and actual capacity may vary. Limits differ by model, and some apply only to certain models: image-capable models are metered in images per minute (IPM), and some models carry a tokens-per-day (TPD) ceiling. Experimental and preview models are more restricted than generally available ones.
What Google does publish per tier is the spend-based rate limit, evaluated on a rolling ten-minute window:
The practical consequence is unchanged by the missing table. While Gemini 3.8 Flash accepts 1,048,576 tokens in a single request, no tier lets you sustain that volume every minute. A service that sends an 800,000-token prompt consumes the bulk of its minute's token allocation in one call, and a second concurrent worker dispatching a similar prompt inside the same window fails immediately with an HTTP 429 rate limit error.
Free Tier Limits Versus Paid Tier Quotas
The progression from evaluation sandbox to production deployment requires transitioning across Google AI Studio usage tiers:
- Free Tier: Available immediately upon creating a project in Google AI Studio. Flash-class models carry the highest free allowances, while Pro-class reasoning models and any experimental or preview model are restricted considerably more tightly. Free Tier inputs and outputs may be reviewed by human evaluators and used to train Google models.
- Tier 1 (Pay-as-You-Go): Activated immediately upon linking a valid Google Cloud billing account. Tier 1 raises per-minute request and token allowances across every model class and carries a $250 billing tier cap with a $10 spend rate limit per rolling ten minutes. Customer data is exempt from model training.
- Tier 2: Unlocked automatically after meeting initial cumulative spending thresholds on Google Cloud services and maintaining an active account following the first successful payment.
- Tier 3: Unlocked after meeting higher cumulative spending thresholds and maintaining an established operational history, providing maximum throughput for enterprise deployments.
Moving from the Free Tier to Tier 1 raises every per-minute ceiling, but any ceiling is easily exhausted when running automated agents across large document sets, because each agent turn re-transmits the whole corpus.
Rolling Spend Limits and Batch API Enqueued Tokens
In addition to RPM and TPM limits, Google enforces rolling spend-based rate limits across rolling ten-minute windows:
- Free Tier: Not applicable (unpaid evaluation).
- Tier 1: Entry spend rate limit per rolling ten-minute window.
- Tier 2: Intermediate spend rate limit per rolling ten-minute window.
- Tier 3: Highest standard spend rate limit per rolling ten-minute window for enterprise deployments.
If your application sends rapid bursts of high-token prompts, you can trigger an HTTP 429 response from spend velocity limits even when nominal RPM and TPM metrics appear healthy.
For high-volume offline workloads, Google provides the Batch API. The Batch API processes asynchronous jobs at a discounted rate compared to standard interactive pricing without drawing down real-time TPM allocations:
- Tier 1: Enqueues millions of tokens for Flash models and Pro reasoning variants.
- Tier 2: Significantly expands active enqueued token pools for production scale.
- Tier 3: Allocates maximum enqueued capacity for high-volume enterprise pipelines.
Batch API jobs accommodate large individual input files and extensive cumulative file storage across concurrent batch requests.
What Happens When You Exceed the Gemini Token Limit?
When your application breaches an active quota or token boundary, the Gemini API halts execution and returns a structured error payload:
- HTTP 429 RESOURCE_EXHAUSTED: Returned when you exceed RPM, TPM, daily RPD, or rolling spend rate limits. The request is rejected without processing, and no billing charges are incurred.
- HTTP 400 INVALID_ARGUMENT: Returned when the input prompt exceeds the model physical context window (for example, attempting to send 1,200,000 tokens into Gemini 3.8 Flash, which accepts a maximum of 1,048,576 tokens).
- Generation Truncation (finish_reason: MAX_TOKENS): Returned with an HTTP 200 status code when the prompt succeeds, but the generated response hits the 65,536 output token ceiling. The output stops abruptly mid-sentence or mid-object.
Understanding these failure states determines whether your service crashes or recovers smoothly.
Decoding HTTP 429 and RESOURCE_EXHAUSTED Errors
When an HTTP 429 occurs, the response body provides critical diagnostic details that distinguish between request count exhaustion and token volume saturation:
{
"error": {
"code": 429,
"message": "Resource has been exhausted (e.g. check quota).",
"status": "RESOURCE_EXHAUSTED",
"details": [
{
"@type": "type.googleapis.com/google.rpc.QuotaFailure",
"violations": [
{
"subject": "project:987654321098",
"description": "Quota exceeded for quota metric 'Tokens per minute' and limit 'Tokens per minute per project' of service 'generativelanguage.googleapis.com'"
}
]
},
{
"@type": "type.googleapis.com/google.rpc.ErrorInfo",
"reason": "RATE_LIMIT_EXHAUSTED",
"domain": "googleapis.com",
"metadata": {
"consumer": "projects/987654321098",
"quota_location": "global",
"quota_metric": "generativelanguage.googleapis.com/generate_content_input_tokens",
"quota_value": "1000000"
}
}
]
}
}
In the payload above, inspecting quota_metric reveals that generate_content_input_tokens triggered the failure, and quota_value reports the exact per-minute input token ceiling that applied to the request.
Calculating Request Tokens and Implementing Exponential Backoff
To avoid unexpected 429 errors and context overflows, production applications must calculate prompt token counts before dispatching requests. In Python using the Google GenAI SDK:
import os
from google import genai
client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])
prompt_text = "Analyze this quarterly report dataset..."
# count tokens prior to generation
token_count = client.models.count_tokens(
model="gemini-3.8-flash",
contents=prompt_text
)
print(f"Total input tokens: {token_count.total_tokens}")
if token_count.total_tokens > 950000:
print("Warning: Approaching the per-minute token ceiling. Route to external retrieval.")
When unexpected HTTP 429 errors occur due to concurrent traffic spikes, client applications must implement exponential backoff with randomized jitter rather than retrying immediately:
import time
import random
from google.genai.errors import APIError
def execute_with_backoff(client, model, contents, max_retries=5):
base_delay = 2.0
for attempt in range(max_retries):
try:
return client.models.generate_content(
model=model,
contents=contents
)
except APIError as e:
if e.code == 429 and attempt < max_retries - 1:
# calculate backoff with full jitter
sleep_time = (base_delay * (2 ** attempt)) + random.uniform(0.5, 1.5)
time.sleep(sleep_time)
else:
raise e
While exponential backoff resolves transient RPM spikes, it cannot solve systemic TPM exhaustion caused by repetitive, oversized prompt payloads.
Keep Gemini prompt context lean with Fast.io workspaces
Store, index, and query massive document archives through our remote MCP server instead of burning through rolling TPM limits. Every organization starts with a 14-day free trial, which requires a credit card. Subscription plans are Starter at `$9.99/mo`, Business at `$49.99/mo`, and Enterprise at `$199.99/mo` on [Fast.io pricing](/pricing/).
Why Does Stuffing Files into Prompts Break Gemini Token Budgets?
The availability of a 1,048,576-token context window tempts developers into using prompt payloads as a replacement for databases and search indexes. In naive implementations, developers dump entire folder trees, technical manuals, and conversation transcripts into every single API request.
While this approach works for isolated one-off scripts, it breaks down rapidly in production pipelines. Treating a large context window as a persistent document repository introduces three compounding problems:
- Token Throughput Saturation: Massive prompts burn through your project rolling TPM allowance on every call.
- Latency Explosions: Time to First Token (TTFT) scales with prompt token volume, turning real-time interactions into sluggish waiting games.
- Attention Degradation: Language models experience recall drift and reasoning inconsistencies when navigating dense, unstructured text blocks.
The Compounding Math of Multi-Turn Context Expansion
The primary trap of in-context document processing is multi-turn conversational expansion. In stateless API protocols, language models have no internal memory across independent calls. To maintain conversation context, client applications must append the entire conversation history, including all prior documents, instructions, and model replies, into every subsequent request.
Consider an autonomous document review assistant examining a 250,000-token legal agreement across four conversational turns:
- Turn 1: Client submits instructions (2,000 tokens), the document (250,000 tokens), and question 1 (500 tokens). Gemini processes 252,500 input tokens and generates 1,500 output tokens. Total turn tokens: 254,000 tokens.
- Turn 2: Client appends the document, question 1, answer 1, and question 2. Gemini processes 254,500 input tokens and generates 2,000 output tokens. Total turn tokens: 256,500 tokens. Cumulative tokens: 510,500 tokens.
- Turn 3: Client resends the full history and question 3. Gemini processes 257,000 input tokens and generates 1,800 output tokens. Total turn tokens: 258,800 tokens. Cumulative tokens: 769,300 tokens.
- Turn 4: Client resends the full history and question 4. Gemini processes 259,300 input tokens and generates 2,200 output tokens. Total turn tokens: 261,500 tokens. Cumulative tokens: 1,030,800 tokens.
Over four questions, that single session processed over one million tokens. On a Free tier project, the later turns fail with an HTTP 429 error. On a paid tier, running several concurrent users through this same workflow still exhausts the project's throughput allocation. Either way you are paying for and reprocessing the exact same 250,000 tokens four times in succession.
Latency Spikes, Attention Drift, and Context Caching Costs
Beyond throughput caps, massive prompts inflict severe performance penalties on application responsiveness and output quality:
Generation Latency Spikes
Processing prompt tokens requires computing the initial key-value (KV) attention cache across inference hardware before emitting the first response token. Time to First Token (TTFT) scales directly with input size. A prompt containing 5,000 tokens typically begins generating output within 800 milliseconds on Gemini 3.8 Flash. If you transmit an 800,000-token document collection, TTFT surges to 15 to 40 seconds. For autonomous agents running iterative tool loops, waiting half a minute for every turn paralyzes execution.
Attention Diffusion Across Dense Context
While synthetic benchmark evaluations demonstrate clean retrieval in isolated tests, real-world technical and legal documents behave differently. Transformer attention mechanisms tend to prioritize tokens situated near the very beginning of the prompt (system instructions) and the very end (the latest user prompt). When hundreds of pages of complex specifications are stuffed into the middle of a 700,000-token context payload, models experience attention diffusion. Minor qualification clauses, cross-file variable references, and conflicting policy exceptions buried in the middle region are frequently overlooked.
Limitations of Context Caching
Google offers context caching to reduce input token billing on prompts that exceed 32,768 tokens. When caching is active, Gemini caches the precomputed KV state for a minimum of one hour. However, context caching introduces rigid operational constraints:
- Minimum Threshold: Caching requires at least 32,768 tokens; smaller prompt segments cannot benefit.
- Storage Time-to-Live Fees: Cached tokens incur an hourly storage charge per million tokens in addition to invocation costs.
- Prefix Sensitivity: The cached tokens must sit at the exact beginning of the prompt. Any modification to system instructions or early text invalidates the entire cache, forcing a full re-computation.
- Volatile File Sets: In collaborative team environments where files are actively edited or versioned throughout the day, maintaining valid static caches becomes impractical.
How Can External Workspaces Solve Gemini File and Token Limits?
The architectural solution to Gemini token limits and TPM bottlenecks is decoupling persistent document storage from LLM prompt execution. Instead of uploading entire multi-megabyte file binders into every prompt turn, production teams store their corpus in an external workspace, index documents automatically for hybrid search, and retrieve only the relevant passages required to answer the immediate question.
Organizations often evaluate local disk directories, AWS S3 buckets, or general cloud drives like Google Drive and Dropbox for file persistence. While basic object storage holds files securely, raw cloud storage lacks native indexing for artificial intelligence, leaving developers to assemble separate embedding scripts, vector databases, and retrieval pipelines by hand.
Fast.io provides shared cloud workspaces designed specifically for agentic teams and AI pipelines. Fast.io leaves every vendor's own upload limit exactly where it is; what it adds is a searchable place for the files that do not fit.
In consumer chat interfaces, file handling constraints restrict document volume. For example, in Claude Projects, project knowledge is limited by the context window, and individual files face strict size boundaries according to Anthropic documentation. Gemini web interfaces face similar browser upload limits, file parsing timeouts, and memory constraints. Moving your repository to an external workspace eliminates these client-side restrictions.
Decoupling File Ingestion from Prompt Context with Fast.io MCP
Fast.io provides a consolidated Model Context Protocol (MCP) server that connects Gemini agents directly to shared workspace files. The remote MCP server operates over Streamable HTTP at https://mcp.fast.io/mcp (or https://mcp.fast.io/mcp/key when using bearer token authentication), with a legacy SSE transport available at https://mcp.fast.io/sse.
Setting up external file retrieval follows four straightforward steps:
- Centralize Documents in an Org-Owned Workspace: Create a shared workspace in Fast.io. You can upload large files directly using chunked uploads or import existing archives from Google Drive, Dropbox, Box, or OneDrive. Fast.io supports one-time cloud import for Google Drive today, with active sync coming soon.
- Enable Intelligence Mode for Automatic Indexing: Turn on Intelligence Mode in the workspace settings. Fast.io automatically indexes all incoming documents for hybrid search, combining full-text lexical search with semantic vector search. No external vector database, embedding model, or chunking script is required.
- Connect Gemini via the Remote MCP Server: Register Fast.io's remote MCP endpoint in your agent configuration, IDE, or desktop assistant.
- Retrieve Precise Passages on Demand: When a user asks a question, the assistant calls the Fast.io MCP search tool to query the workspace. Instead of transmitting 400,000 document tokens, the assistant retrieves only the top three relevant paragraphs with source citations, inserting just 1,500 tokens into Gemini's context window.
Below is an MCP client configuration example for connecting an assistant runtime to Fast.io:
{
"mcpServers": {
"fastio": {
"url": "https://mcp.fast.io/mcp/key",
"headers": {
"Authorization": "Bearer YOUR_FASTIO_API_KEY"
}
}
}
}
By retrieving only 1,500 to 3,000 tokens per turn instead of 300,000 tokens, your application preserves nearly all of its rolling TPM quota, slashes generation latency from thirty seconds to sub-second responses, and eliminates HTTP 429 throttling.
Extracting Structured Document Records with Metadata Views
Unstructured conversational search solves half the file problem. The other half involves extracting structured business intelligence from documents, such as contract renewal dates, indemnification limits, invoice totals, or technical parameters.
Rather than writing custom prompt-engineering scripts that burn thousands of tokens asking Gemini to output JSON schemas, teams use Metadata Views. Metadata Views turn documents into a live, queryable database without requiring OCR rules or rigid templates.
Users describe the fields they want extracted in natural language, and Fast.io builds a typed schema matching files in the workspace:
- Text: Vendor names, jurisdiction clauses, software package identifiers.
- Integer and Decimal: Line item quantities, invoice totals, rate limits, latency metrics.
- Boolean: Exclusivity clauses, compliance indicators, active status flags.
- URL and JSON: API endpoints, structured configuration objects, reference links.
- Date and Time: Effective dates, expiration deadlines, milestone timestamps.
Metadata Views work across PDFs, spreadsheets, presentations, scanned pages, and images. When new columns are added to a view, Fast.io updates the schema without requiring manual reprocessing. AI assistants can create views, trigger extraction, and query structured records directly through Fast.io MCP tools.
While Intelligence Mode handles semantic search and summarization across unstructured text, Metadata Views serve as the structured data extraction layer. Combined with per-file version history, an append-only audit log, Collaborative Notes, and agent-to-human ownership transfer, Fast.io provides the persistent coordination substrate for multi-agent workflows.
Every organization starts with a 14-day free trial, which requires a credit card. Subscription plans are Starter at $9.99/mo, Business at $49.99/mo, and Enterprise at $199.99/mo on Fast.io pricing. For developers building agent integrations, documentation and tool references are available in Fast.io storage for agents, with technical schema details at https://mcp.fast.io/skill.md and onboarding guides at https://fast.io/llms.txt.
Sources
References used to verify factual claims in this guide.
-
Rate limits regulate the number of requests made to the Gemini API across requests per minute, tokens per minute, and requests per day. Google publishes Gemini rate limits in Google AI Studio rather than as a fixed per-model table, and states that specified limits are not guaranteed.
-
The Gemini context window defines the combined limit of input and output tokens across an interaction.
Frequently Asked Questions
What is the token limit for the Gemini API?
Current Google Gemini 3 and 2.5 models support an input context window of 1,048,576 tokens and an output generation ceiling of 65,536 tokens per request. The retired Gemini 1.5 Pro previously supported up to 2,097,152 input tokens but capped output at 8,192 tokens.
How many tokens per minute does the Gemini API allow?
Google sets tokens-per-minute allowances per model and per usage tier and publishes them in Google AI Studio rather than in its rate limits documentation, noting that specified limits are not guaranteed. Flash-class models carry the highest allowances on the Free tier, Pro-class and preview models considerably lower ones, and every allowance rises as a project moves from Free to Tier 1, 2 and 3. Read your project's current figures in AI Studio or through the Rate Limits API.
What happens when you exceed the Gemini token limit?
If an individual request exceeds the 1,048,576-token input window, the API returns an HTTP 400 Invalid Argument error. If generation reaches the 65,536-token output cap, the model halts with finish_reason MAX_TOKENS. If your application exceeds rolling TPM or RPM limits, the API returns an HTTP 429 RESOURCE_EXHAUSTED error.
Are Gemini rate limits applied per API key or per project?
Gemini API rate limits are enforced per Google Cloud project, not per API key. Generating multiple API keys within the same project shares the exact same RPM, TPM, and daily RPD quota allocation.
How does Fast.io help with Gemini token limits?
Fast.io provides shared workspaces that index large document collections for hybrid search. Instead of passing massive multi-megabyte files into Gemini prompts and consuming millions of tokens, an assistant connects via the remote Fast.io MCP server to retrieve only the relevant passages with citations on demand as described in [Fast.io storage for agents](/storage-for-agents/).
Related Resources
Keep Gemini prompt context lean with Fast.io workspaces
Store, index, and query massive document archives through our remote MCP server instead of burning through rolling TPM limits. Every organization starts with a 14-day free trial, which requires a credit card. Subscription plans are Starter at `$9.99/mo`, Business at `$49.99/mo`, and Enterprise at `$199.99/mo` on [Fast.io pricing](/pricing/).