# Grok Limits: Rate Caps, File Sizes, and Context Window Constraints

Grok limits encompass the query frequency caps, context length windows, and file attachment thresholds enforced by xAI across Free, Premium, SuperGrok, and API tiers. While free web users operate under rolling query restrictions, Grok web and mobile apps enforce a file upload limit of up to 150MB per file. Managing large document collections requires decoupling file storage from context windows.

Source: https://fast.io/resources/grok-limits/
Author: [Derek Labian](https://fast.io/authors/derek-labian/)
Last reviewed: 2026-09-25

## What Are the Current Rate Limits and Usage Caps on Grok?

As of September 2026, free tier users of Grok on the X platform face rolling caps of 10 to 20 queries every 2 hours, while Grok web and mobile apps enforce a file upload limit of up to 150MB per file. Understanding these thresholds is essential for individual users, researchers, and engineering teams integrating xAI models into production systems.

"Grok limits encompass the query frequency caps, context length windows, and file attachment thresholds enforced by xAI across Free, Premium, and SuperGrok accounts." When you exceed any of these operational constraints, the platform halts interactions with temporary lockouts, UI throttling notices, or HTTP 429 Too Many Requests errors.

### The Dynamics of Rolling Usage Windows

Unlike platforms that reset allowances at a fixed local time each morning, consumer Grok on the X platform measures usage across a dynamic 2-hour rolling window. Every query you submit registers a timestamp in xAI's metering cluster. The system continually evaluates how many requests occurred in the preceding 120 minutes.

During periods of standard platform traffic, a free account can typically submit between 10 and 20 prompt turns before the interface blocks further input. Once you hit that threshold, the system displays a countdown indicating when the oldest request in the current window expires. If your usage was concentrated within a rapid 10-minute span, you must wait nearly the full 2 hours before your quota clears.

When traffic surges across xAI's Colossus GPU infrastructure, xAI dynamically constrains these rolling quotas. Free tier users frequently report allowances dropping toward 10 messages, and response latency increasing as compute clusters prioritize paid tiers.

### Paid Consumer Tiers: X Premium, SuperGrok, and SuperGrok Heavy

To access higher frequency caps, users choose between social subscriptions on the X platform and standalone SuperGrok tiers on grok.com:

* **X Premium and Premium+.** Subscribing through X raises the rolling cap to approximately 50 queries every 2 hours on standard models like Grok 2 and Grok 3. However, these plans remain tied to rolling time constraints and offer limited access to xAI's advanced reasoning models.
* **SuperGrok.** Hosted on grok.com, SuperGrok moves away from simple message counts to a shared compute pool. Users receive conversational queries along with image generations via Grok Imagine under current throttled caps. Complex operations like Think Mode and DeepSearch consume more compute per turn, reducing total message volume faster than standard text generation.
* **SuperGrok Heavy.** Built for power researchers and creators, SuperGrok Heavy scales daily allowances, provides priority queue routing, and unlocks xAI's video generation models. Even on this tier, video generation remains throttled based on real-time server load.

### Grok Tier Limits and Constraints Compared

The following comparison details the operational limits across consumer tiers and developer endpoints as checked on September 25, 2026:

| Platform Tier or Interface | Query Rate Limit | Context Window | File Upload Cap | Verified Date |
| :--- | :--- | :--- | :--- | :--- |
| Grok Free (X Platform) | 10 to 20 queries per 2 hours (rolling) | 128,000 tokens | 150 MB (web) | 2026-09-25 |
| X Premium ($8/mo) | Approximately 50 queries per 2 hours | 128,000 tokens | 150 MB (web) | 2026-09-25 |
| SuperGrok ($30/mo) | ~100 queries/day (compute pool) | 500,000 tokens | 150 MB (web) | 2026-09-25 |
| SuperGrok Heavy ($300/mo) | ~500 queries/day (priority queue) | Up to 1,000,000 tokens | 150 MB (web) | 2026-09-25 |
| xAI API (Tier 0, $0 spend) | 37 to 150 RPS, 10M to 50M TPM | 256,000 to 1,000,000 tokens | 48 MB (Files API) | 2026-09-25 |
| xAI API (Tier 4, $5,000 spend) | 208 to 500 RPS, 85M to 100M TPM | 256,000 to 1,000,000 tokens | 48 MB (Files API) | 2026-09-25 |

For both free and paid accounts, running into these limits is not merely an inconvenience; it halts active research sessions and breaks automated developer scripts that lack backoff handling.

## What Are the File Upload Size and Attachment Limits for Grok?

Direct file uploads in Grok allow users to analyze spreadsheets, inspect source code, parse PDFs, and process images. However, the size and quantity of files you can attach vary sharply depending on whether you access Grok through a browser, a mobile app, or the programmatic API.

### Web and Mobile Interface Ceilings

When using grok.com or the Grok panel on X, Grok web and mobile apps enforce a file upload limit of up to 150MB per file. This ceiling applies to individual documents, spreadsheets, presentations, and archives uploaded via the attachment icon.

In addition to the per-file size cap, interface clients enforce batch attachment constraints:

* **Desktop Web Browsers:** You can attach multiple individual files in a single prompt session, provided the aggregate volume does not trigger browser memory exhaustion or exceed network request timeouts.
* **Mobile Applications (iOS and Android):** The native mobile apps restrict attachments to 20 files per message. This limit prevents mobile memory crashes during file preprocessing and local image compression.
* **Supported File Types:** Grok natively parses text formats (.txt, .md, .csv, .json), code files (.py, .ts, .js, .go, .rs, .c, .cpp, .html, .css), documents (.pdf, .docx, .xlsx, .pptx), and image files (.png, .jpg, .webp).

While this ceiling accommodates standard business documents, uploading large files into chat interfaces introduces performance tradeoffs. The browser must read, parse, and upload the entire document before the model processes your query. On high-latency connections, uploading multiple large PDFs can cause timeouts before Grok generates its first token.

### Developer Ceilings on the xAI Files API

For engineers building automated pipelines, direct file uploads must pass through the xAI Files API. Here, the developer xAI Files API enforces a file size limit of 48MB per file for attachments.

Key developer constraints include:

* **Attachment Size Boundary:** The developer xAI Files API enforces a file size limit of 48MB per file for attachments across all standard document formats.
* **Vision and Image Understanding:** Standalone images submitted for visual reasoning have a separate 20-megabyte ceiling per image payload.
* **Storage and Persistence:** Files uploaded via the Files API remain private to your team organization. They can be referenced by unique file identifiers across multiple chat completions or grouped into persistent Collections for semantic search.

Attempting to upload a document exceeding accepted boundaries to the Files API returns an immediate client error indicating payload rejection. Furthermore, when files are transmitted as base64-encoded strings within JSON request bodies, the encoding process introduces substantial byte expansion. Binary data expands when converted to base64 text, causing programmatic API requests to fail unexpectedly.

### The Hidden Bottleneck of Direct File Attachments

Attaching large files directly to a conversational prompt creates a compounding computational bottleneck. LLMs are stateless by nature. When you attach a document to a conversation and ask five sequential follow-up questions, Grok does not store a persistent mental index of that file.

Instead, the chat client resubmits the entire text extraction of that document alongside your message history on every single turn. This architecture causes three major problems:

1. **Rapid Quota Depletion:** On SuperGrok, resending hundreds of thousands of tokens of document text on every turn quickly drains your shared compute pool. On the API, it burns through your rate limits within seconds.
2. **Context Fragmentation:** Pushing hundreds of pages into the prompt crowds out system instructions and prior conversational turns, increasing the risk of the model forgetting critical constraints.
3. **Severe Response Latency:** Generating a response requires processing every document token in the prompt before generating the first completion token. Time-to-first-token metrics spike when large files remain directly attached to the session.

## How Large Are Grok Context Windows Across Models?

A model's context window defines the maximum number of tokens it can hold in active memory across the system prompt, conversational history, attached documents, and generated output. Over recent releases, xAI has expanded Grok's context window capabilities, offering distinct models tailored for reasoning, high-speed interaction, and coding.

### Context Window Capacity Across the Grok Model Family

The context window is not identical across all Grok variants. As documented in xAI's model registry as of September 2026, context capacity spans from 256,000 tokens through 1,000,000 tokens:

* **Grok 4.7 (Flagship Multimodal Reasoning):** Features a 500,000-token context window (`maxPromptLength: 500,000`). Grok 4.7 handles complex logic, code synthesis, and image analysis while maintaining coherent long-range reasoning.
* **Grok 4.6 and Grok 4.5:** Both earlier flagship iterations also feature a 500,000-token context window, balancing reasoning depth with high throughput.
* **Grok 4.3 and Grok 4.20 (Standard and Reasoning):** Support an expansive 1,000,000-token context window (`maxPromptLength: 1,000,000`). These models are designed for extensive text synthesis, multi-document comparison, and large log ingestion.
* **Grok-build-0.1 (Specialized Code Model):** Operates with a 256,000-token context window (`maxPromptLength: 256,000`). This window is optimized for repository inspections, pull request evaluations, and targeted test generation.

To put these figures into perspective, 500,000 tokens corresponds to roughly 375,000 words of English prose, equivalent to an extensive technical manual. A 1,000,000-token window can theoretically ingest multiple full-length books or an entire codebase in a single prompt.

### How Context Tokens Accumulate in Practice

While a 500,000-token window sounds virtually limitless, tokens accumulate far faster than most users expect. Every conversational interaction includes multiple overhead sources:

* **System Instructions and Tool Schemas:** If you invoke function calling or web search tools, their JSON definitions consume thousands of tokens before your user prompt is even evaluated.
* **Conversational History:** Every user question and model response is retained in memory. In an extended conversation, earlier responses are continuously reparsed by the model on every subsequent submission.
* **Reasoning Tokens:** On reasoning models such as Grok 4.7 and Grok 4.20 Reasoning, the model generates internal chain-of-thought tokens. These reasoning tokens count directly against your context window budget and rate limits, even though they may be hidden or collapsed in the user interface.
* **Document Ingestion:** A dense financial statement or legal filing can easily consume tens of thousands of tokens once parsed into plain text.

### The Limits of In-Context Retrieval

Having an expansive context window does not eliminate retrieval problems. Academic evaluations and practical engineering tests consistently observe the "needle in a haystack" degradation curve:

As the active context expands into hundreds of thousands of tokens, transformer models suffer from attention dispersion. Information placed in the middle of a massive prompt is frequently overlooked in favor of facts stated at the very beginning or end of the document. If your goal is extracting a specific contract clause, auditing an invoice line item, or retrieving an obscure function definition, stuffing raw text into Grok often yields lower factual accuracy than retrieving only the relevant excerpt through structured search.

Furthermore, processing hundreds of thousands of tokens on every conversational turn is economically inefficient. Token pricing and rate limits make brute-force prompt stuffing unsustainable for production applications.

## What Are the xAI API Rate Limits and Spend Tiers?

When transitioning from consumer interfaces to automated software architectures, teams interact with xAI's developer API. The API replaces rolling 2-hour user caps with mathematically defined infrastructure quotas designed to protect xAI's inference clusters from sudden traffic bursts.

### Two-Dimensional Rate Limiting: RPS and TPM

According to official developer documentation, xAI enforces per-model programmatic rate limits on two dimensions including requests per second and tokens per minute:

1. **Requests Per Second (RPS):** Governs how many individual API calls your team can initiate within a single second. Under xAI programmatic rate limits, requests per second protect server worker nodes from micro-burst overloads.
2. **Tokens Per Minute (TPM):** Dictates the total volume of tokens processed across all requests in a rolling 60-second window. Under xAI programmatic rate limits, tokens per minute scale with your account tier.

Every token consumed by your API requests draws down your allocation. Under xAI programmatic rate limits, tokens per minute include prompt tokens, completion tokens, reasoning tokens, and cached tokens.

### Spend-Based Rate Limit Tiers

Under xAI programmatic rate limits, requests per second and tokens per minute scale across tiers based on cumulative billing spend since January 1, 2026. As your organization purchases prepaid credits or fulfills invoices, tiers unlock permanently without manual support requests:

* **Tier 0:** Baseline sandbox tier for prototyping before establishing a billing history. Grok 4.7 provides 150 requests per second and 50 million tokens per minute, while Grok 4.3 and Grok 4.20 provide 37 requests per second and 10 million tokens per minute. Grok-build-0.1 offers 37 requests per second.
* **Tier 1:** Unlocks production throughput. Grok 4.7 scales to 172 requests per second and 53 million tokens per minute; Grok 4.3 scales to 50 requests per second.
* **Tier 2:** Intermediate scaling tier. Grok 4.7 allows 208 requests per second and 60 million tokens per minute; Grok 4.3 allows 75 requests per second.
* **Tier 3:** High-volume deployment tier. Grok 4.7 provides 312 requests per second and 74 million tokens per minute; Grok 4.3 provides 125 requests per second.
* **Tier 4:** Enterprise-grade quota. Grok 4.7 reaches 500 requests per second and 100 million tokens per minute; Grok 4.3 reaches 208 requests per second.
* **Enterprise:** Custom allocations available upon direct agreement with xAI engineering sales.

Specialized endpoints carry distinct concurrency limits. Grok Imagine image generation operates at 6 requests per second on Tier 0, scaling to 100 requests per second on Tier 4. Video generation operates at 10 requests per second on Tier 0, scaling to 158 requests per second on Tier 4. Voice endpoints, including Speech to Speech, are constrained by Concurrent Sessions (CST), offering 10 concurrent streams on Tier 0 through 200 on Tier 4.

### Handling HTTP 429 Errors with Exponential Backoff

Exceeding either your requests per second ceiling or your tokens per minute allowance causes the xAI API to reject requests with an HTTP 429 Too Many Requests status code. Production systems must implement exponential backoff with jitter to handle these rate events gracefully.

The following Python example demonstrates a resilient request wrapper using the standard OpenAI client pointed at the [xAI API endpoint](https://docs.x.ai/developers/rate-limits):

```python
import os
import time
import random
from openai import OpenAI, RateLimitError

client = OpenAI(
    base_url="https://api.x.ai/v1",
    api_key=os.environ.get("XAI_API_KEY")
)

def query_grok_with_backoff(prompt, max_retries=5, base_delay=1.0):
    for attempt in range(max_retries):
        try:
            response = client.chat.completions.create(
                model="grok-4.7",
                messages=[
                    {"role": "system", "content": "You are a precise technical analyst."},
                    {"role": "user", "content": prompt}
                ]
            )
            return response.choices[0].message.content
        except RateLimitError as err:
            if attempt == max_retries - 1:
                raise err
            jitter = random.uniform(0.5, 1.5)
            delay = (base_delay * (2 ** attempt)) * jitter
            time.sleep(delay)

output = query_grok_with_backoff("Analyze system logs for anomalies.")
print(output)
```

By decoupling sudden spikes in request volume from your core application logic, exponential backoff ensures that temporary 429 errors do not corrupt downstream data pipelines.

## How Can You Analyze Files That Exceed Grok Upload Limits?

When building enterprise workflows or multi-agent research pipelines, teams frequently need to analyze files that exceed Grok's operational limits. A team analyzing an extensive legal discovery archive, a large software repository, or thousands of client invoices cannot upload those files directly to [grok.com](https://grok.com) or push them through the xAI Files API.

Attempting to force massive corpora through direct model uploads leads to broken workflows:

* You hit the upload boundary on the Files API or web interface.
* You exhaust your rate limits and tokens per minute allocation on the first conversational turn.
* You trigger attention degradation and hallucinations across bloated multi-hundred-thousand-token prompts.
* You incur massive billing costs by re-sending the same static context with every follow-up question.

To analyze large document collections reliably, engineering teams must separate document storage and retrieval from active model context.

### Traditional Workarounds and Their Tradeoffs

Teams seeking to bypass Grok file upload limits historically turn to two architectural patterns, each carrying significant engineering overhead:

1. **Custom Vector Databases and Local Chunking Scripts.** Developers write custom Python scripts to extract text from PDFs, split documents into smaller chunks, compute vector embeddings using an external model, and store the embeddings in a database like Pinecone, Qdrant, or Chroma. While functional, this approach requires building and maintaining custom parsing infrastructure, handling extraction failures, managing embedding model versions, and writing complex retrieval logic.
2. **Commodity Cloud Storage.** Storing files in basic cloud buckets solves the file size problem, allowing uploads of large datasets without restriction. However, commodity object storage lacks intelligence. Cloud buckets cannot index document contents, generate citations, extract structured data, or expose an interface that AI models can query autonomously.

### The Intelligent Workspace Solution with Fast.io

Fast.io provides a dedicated intelligent workspace layer designed specifically for agentic teams and large document corpora. Instead of attempting to attach files directly to Grok, you store your complete document archive in a shared, organization-owned [Fast.io workspace](/product/workspaces/).

Fast.io solves the large-corpus problem through core workspace capabilities:

* **Ingest Data Without Bandwidth Bottlenecks:** Upload extensive document libraries directly through chunked uploads, or bring files in via one-time [cloud import](/product/cloud-import/) from Google Drive, OneDrive, Dropbox, or Box over secure OAuth without consuming local bandwidth.
* **Automatic Intelligence and Hybrid Indexing:** Once files arrive in a workspace, enabling [Intelligence Mode](/product/ai/) triggers automatic indexing. Fast.io indexes content using hybrid search, combining full-text keyword indexing with semantic vector embeddings. Documents become immediately searchable by concept, exact phrasing, or specific metadata values.
* **Remote Model Context Protocol (MCP) Integration:** Fast.io exposes a consolidated MCP toolset at `https://mcp.fast.io/mcp` over Streamable HTTP (with legacy SSE available at `https://mcp.fast.io/sse`). AI assistants connect directly using an API key from `https://mcp.fast.io/mcp/key` as detailed in [Fast.io for Agents](/storage-for-agents/). Instead of attaching raw files to Grok's prompt, the assistant calls Fast.io search tools to retrieve only the relevant paragraphs and exact document citations needed to answer the user's prompt.
* **Structured Data Extraction via Metadata Views:** For teams handling contracts, invoices, or research records, [Metadata Views](/product/document-data-extraction/) turn unstructured documents into live, queryable databases. Users define extraction fields in natural language, and the system automatically extracts typed values across all matching files in the workspace. Agents can query these structured tables via MCP without reading through raw document text.
* **Shared Substrate for Humans and Agents:** Fast.io workspaces are organization-owned, featuring granular permissions, Collaborative Notes coordinated through Agent Intents, per-file version history, and an append-only audit log. An AI agent can research documents, extract findings, and write summaries into Collaborative Notes while human colleagues review and refine the output.

Every organization starts with a 14-day free trial, which requires a credit card. Paid subscription tiers on the [Fast.io pricing page](/pricing/) include Starter, Business, and Enterprise plans.

By pairing Grok's reasoning power with an intelligent workspace, teams bypass arbitrary file upload limits, eliminate context window token waste, and maintain a verifiable, citation-backed system of record for all business data.

## Frequently asked questions

### What is the limit on Grok?

Grok limits depend on your access tier. Free users on X face rolling rate caps of 10 to 20 messages every 2 hours and a 128,000-token context window. SuperGrok provides roughly 100 queries per day from a shared compute pool and a 500,000-token context window. Under xAI programmatic rate limits, developer requests per second and tokens per minute scale across tiers based on spend.

### How many messages can you send on Grok in 2 hours?

Free tier accounts on the X platform can typically send between 10 and 20 messages in a rolling 2-hour window before being throttled. During periods of peak traffic on xAI infrastructure, this limit may dynamically tighten toward 10 messages. X Premium subscribers receive higher rolling caps of approximately 50 messages every 2 hours.

### Does Grok have a file upload limit?

Yes. Grok web and mobile apps enforce a file upload limit of up to 150MB per file for attachments. For programmatic developers, the developer xAI Files API enforces a file size limit of 48MB per file for attachments and a dedicated ceiling for vision images.

### What happens when you hit a 429 Too Many Requests error on the xAI API?

An HTTP 429 error indicates that your application exceeded either its requests per second ceiling or tokens per minute budget for that model. Applications should catch rate limit exceptions and apply exponential backoff with randomized jitter to retry requests without overwhelming the API.

### When do Grok usage limits reset?

Consumer limits on the X platform reset continuously on a rolling 120-minute window measured from the timestamp of each submitted query. SuperGrok daily quotas reset at midnight UTC. Developer API rate limits reset continuously across 1-second and 60-second sliding windows for requests per second and tokens per minute respectively.

### How can teams analyze large document collections without exceeding Grok context windows?

Rather than uploading large files directly to Grok chat sessions, teams store their documents in a shared Fast.io workspace. Fast.io automatically indexes files with hybrid search once Intelligence Mode is enabled, allowing AI models to retrieve targeted excerpts and citations via the Model Context Protocol (MCP) without exhausting context windows.

## Sources

- [SpaceXAI Docs: Rate Limits](https://docs.x.ai/developers/rate-limits) — xAI enforces per-model programmatic rate limits on two dimensions including requests per second and tokens per minute.
- [xAI Documentation: Grok FAQ](https://docs.x.ai/grok/faq) — Grok web and mobile apps enforce a file upload limit of up to 150MB per file, while the developer xAI Files API enforces a file size limit of 48MB per file for attachments.

## About Fast.io

Fast.io provides shared workspaces where people and AI agents work on the same files, with built-in semantic search and citation-backed chat over what they hold. Agents reach it through a remote MCP server at https://mcp.fast.io/mcp, a REST API at https://api.fast.io/current/, and a command line client published on npm as @vividengine/fastio-cli.
