Groq API Rate Limits: LPU Tier Quotas, TPM Ceilings, and Document Handling
Groq rate limit policies govern API throughput across GroqCloud LPUs through concurrent requests, requests per minute, and tokens per minute. While Groq lists generation speeds up to 1,000 tokens per second, strict token-per-minute ceilings create unexpected bottlenecks for document-heavy agent prompts. Managing production throughput requires decoding rate limit headers, configuring exponential backoff, and decoupling file storage from prompt context.
How Groq Calculates LPU Rate Limits and Tier Quotas
Groq publishes a generation speed alongside every model it lists, and across its production models those figures run from 280 tokens per second on Llama 3.3 70B up to 1,000 on GPT OSS 20B. That raw generation speed changes how engineering teams encounter infrastructure limits: on GroqCloud, rate limits rather than token costs or compute availability become the primary operational constraint.
Groq rate limits govern the throughput of API calls to GroqCloud's LPU inference engine, constrained by concurrent requests, requests per minute (RPM), and tokens per minute (TPM) per model. Because Language Processing Units process tokens at hardware speeds, an application can consume thousands of tokens in fractions of a second. Understanding how Groq calculates and enforces these quotas is essential for preventing production disruptions.
How Groq Measures API Consumption
As outlined in the Groq rate limits documentation, GroqCloud tracks API traffic across six distinct dimensions:
- Requests Per Minute (RPM): The total number of completed API calls permitted within a rolling 60-second window.
- Requests Per Day (RPD): The aggregate number of requests allowed across a 24-hour UTC window.
- Tokens Per Minute (TPM): The total token count, combining input prompt tokens and generated output tokens, processed within any rolling 60-second window.
- Tokens Per Day (TPD): The cumulative daily token ceiling for a specific model family.
- Audio Seconds Per Hour (ASH) and Day (ASD): Specific execution caps applied exclusively to Whisper transcription models.
- Input and Output Token Ceilings (ITPM and OTPM): Granular sub-quotas configured on select enterprise accounts to regulate prompt ingestion separately from output generation.
Rate limits apply at the organization level, not individual users. You can hit any limit type depending on which threshold you reach first. If an organization runs multiple automated worker processes, agentic pipelines, and developer environments under a single account, every process shares the same organization-level pool. When one worker spikes in volume, other services in the organization encounter immediate throttling.
The LPU Generation Speed Dynamic
Traditional inference platforms throttle requests primarily around compute capacity. If an inference cluster runs out of tensor cores, user requests queue up, latency increases, and response times degrade gradually.
Groq's deterministically scheduled LPU architecture behaves differently. Because execution latency remains low even under heavy generation, requests never linger in a processing queue. An agent sending three consecutive 3,500-token document prompts will process them nearly instantaneously, consuming 10,500 tokens in two seconds. On a tier configured with a 6,000 or 12,000 TPM limit, that single short burst consumes the entire minute's quota instantly, triggering HTTP 429 errors for all subsequent requests until the minute resets.
Related guides
- Azure OpenAI Rate Limits: TPM Quotas, PTU Scaling, and 429 Error ResolutionAzure OpenAI rate limits are regional and subscription-level constraints defined by Tokens Per Minute (TPM) and...
- LlamaIndex Rate Limits: Ingestion Batching, Embedding Quotas, and Offloaded IndexingLlamaIndex rate limits are API request and token bottlenecks triggered while parsing, chunking, and embedding large...
- Gemini API Rate Limits: Tier Quotas, 429 Handling, and Large-Payload WorkflowsGemini API rate limits govern requests per minute (RPM), tokens per minute (TPM), and requests per day (RPD) across...
- Google AI Studio Rate Limits: Free Tier Quotas, TPM, and Handling 429 ErrorsGoogle AI Studio rate limits enforce operational caps across requests per minute, tokens per minute, and daily request...
- AWS Bedrock Rate Limits: Service Quotas, ThrottlingException, and Document RetrievalAWS Bedrock rate limits are regional, account-level quotas that cap requests and tokens per minute across foundation...
- Open WebUI Rate Limits: User Throttling, Model API Quotas, and RAG ScalingOpen WebUI rate limits encompass both administrative per-user request constraints configured in the web interface and...
More on this subject: Agent Security and Governance (51 guides)
Comparing Rate Limits Across Groq Models and Tiers
Groq organizes account quotas into two primary service levels: the Developer Free tier and the Developer Plan. While the Free tier enables rapid testing and prototyping without requiring upfront commitments, production systems quickly require the higher throughput ceilings of paid tiers.
The following table summarizes the baseline rate limits across Groq's primary model catalog as documented in the Groq Console:
Key Differences Between Free and Paid Quotas
On the Free Plan, rate limits are calibrated for single-user experimentation. A ceiling of 30 RPM and 6,000 to 12,000 TPM means that any agent executing continuous multi-turn tool loops will reach quota exhaustion within three to four turns.
Upgrading to the Developer Plan provides substantial headroom. By providing billing details, an organization unlocks base limits of 1,000 RPM and 300,000 TPM across standard Llama models. For enterprise workloads requiring dedicated capacity, Groq offers custom Performance Tiers with negotiated throughput agreements.
Rate Limit Reset Mechanics
Groq enforces token limits on a sliding minute window rather than a fixed clock boundary. If your application consumes 5,000 tokens at 10:00:15, those tokens do not reset at 10:01:00. Instead, they clear 60 seconds after execution at 10:01:15.
Daily quotas (RPD and TPD) follow UTC midnight resets. If an automated script consumes the 100,000 daily token allowance on Llama 3.3 70B during morning testing, the account remains blocked from making further calls to that model until 00:00 UTC.
How to Inspect Rate Limit Headers and Handle 429 Errors
When an application exceeds an RPM, TPM, or RPD threshold, GroqCloud returns an HTTP 429 (Too Many Requests) status code. Rather than treating 429 responses as fatal crashes, resilient applications parse response headers to pace traffic dynamically.
Rate Limit Header Specifications
Every HTTP response from Groq includes detailed diagnostic headers that disclose the current status of your organization's quotas:
x-ratelimit-limit-requests: The maximum number of requests allowed per day (RPD).x-ratelimit-limit-tokens: The maximum tokens per minute allowed (TPM).x-ratelimit-remaining-requests: The number of requests remaining in your daily allowance.x-ratelimit-remaining-tokens: The number of tokens remaining in the rolling minute window.x-ratelimit-reset-requests: Time remaining until the daily request quota resets (formatted like2m59.56sor14h12m).x-ratelimit-reset-tokens: Time remaining until the per-minute token quota replenishes (formatted like7.66s).retry-after: Included exclusively on HTTP 429 responses. Specifies the integer number of seconds the client must pause before retrying.
Resilient Python Request Handler
To handle burst limits reliably, wrap API calls in a retry handler that inspects headers and implements exponential backoff with jitter:
import json
import random
import time
import urllib.error
import urllib.request
GROQ_API_URL = "https://api.groq.com/openai/v1/chat/completions"
def call_groq_resilient(api_key, payload, max_retries=5):
headers = {"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"}
data = json.dumps(payload).encode("utf-8")
for attempt in range(max_retries):
req = urllib.request.Request(GROQ_API_URL, data=data, headers=headers)
try:
with urllib.request.urlopen(req) as resp:
remaining_tokens = resp.headers.get("x-ratelimit-remaining-tokens")
if remaining_tokens and int(remaining_tokens) < 1000:
time.sleep(1.0)
return json.loads(resp.read().decode("utf-8"))
except urllib.error.HTTPError as e:
if e.code == 429:
retry_after = e.headers.get("retry-after")
if retry_after:
sleep_time = float(retry_after)
else:
sleep_time = (2 ** attempt) + random.uniform(0.1, 1.0)
if attempt == max_retries - 1:
raise RuntimeError(f"Rate limit exceeded after {max_retries} attempts: {e}")
time.sleep(sleep_time)
else:
raise e
raise RuntimeError("Maximum retries exhausted")
This implementation inspects remaining token capacity after each successful call, introduces intentional spacing before quotas are breached, and respects vendor-dictated retry-after values during unexpected bursts.
Eliminate Groq TPM Bottlenecks with Intelligent Workspaces
Stop stuffing full documents into prompt payloads. Connect your agents to Fast.io workspaces via MCP for automatic indexing, semantic retrieval, and structured document extraction. Monthly plans start with a trial of up to 30 days (credit card required); annual plans have no trial. Plans are Starter at $9.99/mo, Business at $49.99/mo, and Enterprise at $199.99/mo.
Why Context Stuffing Fails on Groq LPU Models
The most common trigger of unexpected Groq rate limits is passing raw document content directly inside prompt context.
Llama 3.3 70B and Llama 3.1 8B offer a 128,000-token context window. This large window often leads developers to believe they can paste complete PDF documents, lengthy contract agreements, quarterly filings, or source code repositories directly into the messages array of a single request.
The Arithmetic of TPM Failures
While the model architecture can process 128,000 tokens in a single request, the API rate limit does not permit it.
Consider an agent analyzing a 40-page operational manual containing 20,000 tokens. When the agent sends this manual to Llama 3.3 70B on the Developer Free tier (12,000 TPM ceiling), the request fails immediately:
- Document size: 20,000 tokens
- Free tier allowance: 12,000 tokens per minute
- Result: Immediate HTTP 429 rejection before a single word is generated
Even on the Developer Plan with a 300,000 TPM limit, running fifteen concurrent agent threads evaluating 20,000-token files simultaneously demands 300,000 tokens in one burst. The initial calls succeed, and subsequent concurrent calls fail with rate limit errors.
Prompt Caching Constraints
Groq supports prompt caching, which reduces input processing latency and discounts cached tokens. Crucially, cached tokens do not count toward your organization's rate limits on subsequent requests.
However, prompt caching does not solve the initial document ingestion problem:
- The initial request that populates the cache must still pass all RPM and TPM rate limit gates.
- If the prompt has slight variations, dynamic system timestamps, or unique per-user prefixes, cache misses occur.
- In dynamic multi-agent workflows where documents are edited, versioned, or updated continuously, cache invalidation forces repeated full-context re-evaluations that trigger 429 limits.
Decoupling Document Storage and Retrieval with Fast.io Workspaces
To prevent document-heavy agent workflows from hitting Groq's TPM ceilings, production architectures decouple persistent file storage and indexing from the LLM prompt payload.
Instead of passing entire files to the inference engine, files reside in shared workspaces where they are indexed once and retrieved selectively.
Shared Workspaces with Intelligence Mode
In an intelligent workspace architecture, files are stored centrally in an organization-owned workspace. Fast.io workspaces accept file uploads, imports from cloud storage (Google Drive, Dropbox, Box, and OneDrive), and direct writes from automated agents. When setting up storage for agents, persistent workspaces provide the grounding layer for shared files.
When Intelligence Mode is enabled on a workspace, files are indexed automatically for hybrid search, combining full-text lexical search and semantic retrieval. Documents are chunked and embedded at rest without requiring a separate external vector database or custom chunking pipeline.
Structured Extraction via Metadata Views
When workflows require structured data from incoming files, stuffing entire documents into Groq to extract JSON fields consumes massive token budgets.
Fast.io Metadata Views provide an alternative. Users and agents define a typed schema using natural language, specifying target fields such as text strings, integers, decimals, booleans, dates, URLs, or structured JSON.
Metadata Views automatically process PDFs, Word documents, scanned images, and spreadsheets across the workspace, extracting the designated fields into a structured, queryable view. Agents query the extracted structured data directly rather than paying raw token costs to re-parse entire documents through the inference API.
Connecting Agents via Model Context Protocol
Autonomous agents connect to Fast.io workspaces through the remote Model Context Protocol (MCP) server over Streamable HTTP (https://mcp.fast.io/mcp or https://mcp.fast.io/mcp/key with Bearer authentication). The Fast.io MCP integration gives agents action-based tools to query workspace files directly.
Rather than sending 30,000 tokens of raw file data to Groq, the agent executes an MCP search tool against the workspace:
- The agent queries the workspace using natural language or metadata filters.
- The Fast.io MCP server returns only the top relevant text passages, typically spanning 300 to 600 tokens, accompanied by document citations.
- The agent formats a lightweight prompt containing only the retrieved passages and submits it to Groq.
By reducing prompt payloads from 30,000 tokens down to 600 tokens, a 300,000 TPM limit on Groq can sustain 500 requests per minute rather than failing on request number ten. Reviewing Fast.io pricing shows that organizations start on paid subscriptions with predictable limits.
Because Fast.io maintains full per-file version history and an append-only audit log, concurrent agents can read and write workspace documents simultaneously without conflicting updates or lost context.
Production Best Practices for Groq Inference Workflows
Scaling applications on GroqCloud requires designing systems that operate comfortably within token ceilings. Implement these operational patterns across your inference pipelines:
1. Multi-Model Tiering and Intelligent Routing
Avoid routing every task to large 70B parameter models. Llama 3 8B models provide higher rate limit headroom and execute faster on LPUs:
- Use
llama-3.1-8b-instantwith its higher per-minute token ceiling for initial intent classification, content routing, query decomposition, and output filtering. - Route complex synthesis, legal reasoning, and final decision-making to
llama-3.3-70b-versatile. - By filtering routine classification and query decomposition through smaller models, you preserve high-value TPM allowances on larger models for tasks that genuinely require deep reasoning.
2. Use Asynchronous Batch Processing
For non-real-time tasks such as nightly evaluations, batch classification, or bulk data enrichment, route requests through Groq's Batch API.
Batch requests operate with separate quota pools and receive substantial pricing discounts compared to on-demand execution. By queuing offline processing into batch jobs, real-time user-facing features never contend with background tasks for RPM and TPM allocations.
3. Implement Client-Side Token Budgeting
Do not rely on the API to reject requests after the fact. In high-throughput distributed systems, maintain a client-side token bucket:
- Estimate prompt token counts before sending using standard token estimation heuristics (approximately 4 characters per token for English text).
- Track rolling token consumption across your application workers in a shared memory store.
- If scheduled requests approach your organization's TPM ceiling, queue requests locally rather than triggering API-level 429 rejections.
4. Separate Environments by Organization
Because Groq rate limits apply at the organization level, never share a single organization across production, staging, and local developer environments. A developer testing a recursive agent loop locally can consume the organization's entire TPM ceiling, causing immediate outages for production users. Establish isolated organizations with distinct API keys for each operational stage.
Sources
References used to verify factual claims in this guide.
-
Groq lists generation speeds across its production models running from 280 tokens per second on Llama 3.3 70B up to 1,000 on GPT OSS 20B.
-
Groq returns an HTTP 429 Too Many Requests status when an organization exceeds its rate limits. Groq API rate limits apply at the organization level across all users and trigger on whichever quota threshold is reached first.
Frequently Asked Questions
What are the rate limits for Groq API?
Groq API rate limits are applied at the organization level across requests per minute (RPM), requests per day (RPD), tokens per minute (TPM), and tokens per day (TPD). On the Developer Free tier, top models such as Llama 3.3 70B are capped at 30 RPM and 12,000 TPM, while smaller models like Llama 3.1 8B have limits between 6,000 and 30,000 TPM. The Developer Plan raises baseline limits to 1,000 RPM and 300,000 TPM across standard models.
How to avoid Groq 429 rate limit errors?
To avoid HTTP 429 rate limit errors on Groq, implement exponential backoff with jitter and respect the retry-after header returned in response payloads. Monitor x-ratelimit-remaining-tokens to pause requests proactively before limits are breached. In document processing workflows, avoid passing whole files into prompts; instead, index documents in an external workspace like Fast.io and retrieve only relevant passages via MCP to keep prompt payloads around 600 tokens.
Can you increase Groq TPM limits?
Yes. You can increase Groq TPM limits by upgrading from the Developer Free tier to the paid Developer Plan, which elevates base limits from 12,000 TPM to a baseline of 300,000 TPM. For high-volume production applications requiring even greater throughput, organizations can contact Groq sales for enterprise plans that provide custom rate limit tiers and dedicated LPU capacity.
How does Groq calculate prompt caching against rate limits?
Groq supports prompt caching for repeated identical prompt prefixes. When a request hits a cached prompt prefix, those cached input tokens do not count toward your organization's tokens per minute (TPM) rate limits. However, the initial request that establishes the cache must fit entirely within your available TPM allowance, and cache misses caused by dynamic prompt modifications will consume full token quotas.
Related Resources
Eliminate Groq TPM Bottlenecks with Intelligent Workspaces
Stop stuffing full documents into prompt payloads. Connect your agents to Fast.io workspaces via MCP for automatic indexing, semantic retrieval, and structured document extraction. Monthly plans start with a trial of up to 30 days (credit card required); annual plans have no trial. Plans are Starter at $9.99/mo, Business at $49.99/mo, and Enterprise at $199.99/mo.