# Cerebras Rate Limits: Wafer-Scale Inference Tiers, TPM Caps, and Payload Optimization

Cerebras rate limits govern the frequency of requests and token throughput permissible per minute when querying Cerebras Wafer-Scale clusters. Enforced through a dual-bucket architecture of uncached and total tokens per minute, limits can be exhausted in seconds at 1,500 tokens per second. Instead of sending multi-file contexts directly into API payloads, engineering teams index documents in shared workspaces for targeted semantic retrieval.

Source: https://fast.io/resources/cerebras-rate-limit/
Author: [Derek Labian](https://fast.io/authors/derek-labian/)
Last reviewed: 2026-09-17

## How Cerebras Enforces Rate Limits: Dual-Bucket Architecture and Quota Replenishment

Cerebras rate limits govern the frequency of requests and token throughput permissible per minute when querying Cerebras Wafer-Scale clusters. According to [Cerebras Inference rate limits documentation](https://inference-docs.cerebras.ai/support/rate-limits), Cerebras enforces throughput using a dual-bucket architecture: every organization receives both an uncached token limit and a total token limit. When inference speeds reach 1,500 to 1,800 tokens per second on Wafer-Scale Engine hardware, unthrottled agent loops can exhaust token allowances in seconds. Understanding how these rate limits operate is essential for engineering teams deploying high-velocity applications on Cerebras infrastructure.

Rate limits are designed to preserve cluster stability and ensure predictable performance across organizations. Unlike GPU clusters that batch requests across distributed nodes with varying network latency, Cerebras Wafer-Scale systems execute inference on a single massive wafer of silicon. The speed advantages are substantial, but high generation velocity means that applications consume their token allowances far more rapidly than on traditional hardware.

### Uncached TPM vs Total TPM: The Dual-Bucket Model

Cerebras separates token consumption into two independent accounting buckets, evaluated concurrently on every request:

* **Uncached Tokens per Minute (Uncached TPM):** Measures tokens that require full model compute, representing cache misses. Uncached TPM is the primary system constraint because it directly measures raw compute operations on the Wafer-Scale Engine.
* **Total Tokens per Minute (Total TPM):** Measures combined token volume, including both fresh uncached tokens and reused prompt cache hits. By default, Cerebras sets the Total TPM ceiling to three times your uncached TPM limit.

Both buckets are enforced independently. If an application stays below its uncached limit but submits massive prompt payloads that exceed the Total TPM ceiling, the request will be rejected. Conversely, if an application exceeds its uncached compute allowance even while total tokens remain modest, the request is throttled. When Cerebras returns an HTTP 429 status code, the response body explicitly indicates which bucket was breached, allowing developers to diagnose whether the bottleneck stems from cache misses or aggregate payload volume.

Because cached tokens do not consume uncached quota, maintaining high prompt cache hit rates allows applications to process substantially higher total throughput within existing tier limits.

### Quota Replenishment via Continuous Token Bucketing

Cerebras enforces limits using a continuous token bucket algorithm rather than resetting counters at the top of an hour or minute. In a fixed-window system, an application can consume its entire quota in the first two seconds of a minute and remain completely starved for the remaining 58 seconds.

Under the Cerebras token bucket implementation, available capacity replenishes continuously across time:

```
Available quota = min(Rate limit, Rate limit + replenished tokens by time - current usage)
```

As an application pauses between tool calls or processes intermediate operations, its request bucket and token buckets refill automatically up to the tier ceiling. This design accommodates brief bursts of concurrent requests while smoothing outbound traffic over the rolling window.

### Pre-Request Token Estimation and Max Completion Tokens

A critical operational detail of Cerebras inference is that rate limits are checked before text generation begins. When an API call arrives, Cerebras estimates the maximum token volume the request could consume:

1. The system counts the prompt input tokens.
2. The system adds either the explicit `max_completion_tokens` parameter or an internal upper-bound estimate of anticipated output tokens.

If this projected token total exceeds the organization's remaining quota in either the uncached or total bucket, Cerebras rejects the request immediately with an HTTP 429 error before processing a single token. Once generation completes, the system reconciles the quota to reflect actual tokens consumed.

Developers who omit `max_completion_tokens` or assign arbitrarily large values, such as setting an 8,192 token output ceiling for a single-sentence classification task, frequently trigger false rate limit errors. Setting `max_completion_tokens` to an accurate, conservative ceiling for each prompt is an immediate operational fix for unexpected throttling.

## Comparing Cerebras Inference Tiers: Free Trial, Developer, and Enterprise Quotas

Cerebras provisions rate limits at the organization level rather than per user or per API key. Organizational limits vary depending on the chosen model and account tier: Free Trial, Developer (Pay as You Go), and Enterprise. Choosing the right tier determines whether an application can support concurrent user traffic or autonomous agent execution.

### Free Trial Tier Limitations and Windows

New accounts receive starter trial credits after verifying a payment method. These credits remain valid for 30 days and provide access to public models for evaluation. However, the Free Trial tier applies strict throughput boundaries across requests and tokens:

* **Short-Term Velocity Caps:** Free Trial accounts are restricted to 5 requests per minute (RPM) on flagship models such as `gpt-oss-120b` and `qwen-3.8-27b`. Uncached throughput is restricted to 30,000 tokens per minute, with a Total TPM ceiling of 90,000 tokens per minute.
* **Hourly and Daily Windows:** In addition to per-minute limits, the Free Trial enforces aggregate window caps of 1,000,000 tokens per hour and 1,000,000 tokens per day. Once an account consumes 1,000,000 tokens within a day, API calls stop until the window rolls over.
* **Payload Constraints:** Multimodal models like `qwen-3.8-27b` on the Free Trial tier limit image attachments to 2 images per request and enforce a 10 MiB total request payload size.
* **Legacy Evaluation Profiles:** Legacy model configurations, such as `llama-3.1-8b`, historically operated under an allowance of 30 RPM and 60,000 TPM for basic developer testing.

Once Free Trial credits expire or are depleted, API access pauses until the organization adds billing details.

### Developer Tier Throughput and Budget Flexibility

Purchasing API credits in the Cerebras Cloud Console immediately elevates an account to the Developer tier. The Developer tier removes hourly and daily token limits completely, allowing organizations to process as many tokens as their budget allows.

Throughput limits on the Developer tier expand to support production workloads:

* **`gpt-oss-120b`:** Provides 1,000 RPM, 1,000,000 Uncached TPM, and 3,000,000 Total TPM. This generous ceiling supports high-concurrency chat applications and parallel worker pools.
* **`qwen-3.8-27b`:** Provides 300 RPM, 150,000 Uncached TPM, and 450,000 Total TPM. Image input limits expand to 10 images per request while maintaining the 10 MiB payload cap.

While 150,000 Uncached TPM represents a solid ceiling for interactive human chat, it creates a narrow operational corridor for autonomous coding agents and document-heavy workflows that re-send context repeatedly.

### Enterprise Tier and Dedicated Endpoints

For enterprise workloads requiring guaranteed capacity and tailored service agreements, Cerebras offers Enterprise provisions:

* **Custom Quota Allocation:** Enterprise accounts receive custom RPM and TPM allocations negotiated around specific application concurrency profiles.
* **Dedicated Endpoints:** Organizations can reserve dedicated Wafer-Scale Engine hardware. Dedicated Endpoints eliminate multi-tenant contention, provide deterministic sub-second latencies, and bypass shared rate limit pools entirely.

### Rate Limit Comparison Across Tiers

The following comparison summarizes rate limits and window constraints across Cerebras inference tiers as documented in platform specifications:

| Plan Tier | Model Name | Requests / Min (RPM) | Uncached Tokens / Min | Total Tokens / Min | Window Constraints |
| :--- | :--- | :--- | :--- | :--- | :--- |
| **Free Trial** | `gpt-oss-120b` | 5 RPM | 30,000 Uncached TPM | 90,000 Total TPM | 1,000,000 TPH, 1,000,000 TPD (30-day credit expiry) |
| **Free Trial** | `qwen-3.8-27b` | 5 RPM | 30,000 Uncached TPM | 90,000 Total TPM | 1,000,000 TPH, 1,000,000 TPD (2 images per request) |
| **Free Trial (Legacy)** | `llama-3.1-8b` | 30 RPM | 60,000 Uncached TPM | 180,000 Total TPM | Standard prototype allowance |
| **Developer Tier** | `gpt-oss-120b` | 1,000 RPM | 1,000,000 Uncached TPM | 3,000,000 Total TPM | No hourly or daily caps; pay as you go |
| **Developer Tier** | `qwen-3.8-27b` | 300 RPM | 150,000 Uncached TPM | 450,000 Total TPM | No hourly or daily caps (10 images per request) |
| **Enterprise** | Custom Profiles | Custom | Custom Allocation | Custom Allocation | Dedicated Endpoints and SLA agreements |

## Why Agentic Coding and Document Bloat Trigger Immediate 429 Errors

The primary reason engineering teams encounter HTTP 429 errors on Cerebras is not high user concurrency. It is the interaction between wafer-scale inference velocity and autonomous agent payload bloat.

Cerebras hardware delivers extreme generation speeds, reaching 1,500 to 1,800 tokens per second on models like `qwen-3.8-27b`. In standard human conversational interfaces, generating a 500-token answer takes approximately 300 milliseconds. The human user reads the response during a natural pause before typing their next prompt. This natural pause allows the token bucket to replenish completely between turns.

In autonomous agent systems, however, there is no human pause.

### The Arithmetic of Autonomous Token Exhaustion

When an autonomous coding assistant or multi-file research agent operates, it dispatches tool calls in rapid succession. A code agent might inspect a directory, read five source files, run a linter, inspect test output, and rewrite a module.

Each iteration through the agent loop re-submits the entire conversation history, tool declarations, intermediate tool results, and file contents. The prompt grows larger with every step:

1. **Step 1:** Initial prompt and codebase overview (12,000 tokens).
2. **Step 2:** Step 1 plus file read outputs for three modules (28,000 tokens).
3. **Step 3:** Step 2 plus linter errors and test traces (35,000 tokens).
4. **Step 4:** Step 3 plus dependency specifications and documentation files (42,000 tokens).

On the Developer tier for `qwen-3.8-27b`, the uncached limit is 150,000 tokens per minute. Because Cerebras streams responses at 1,500 tokens per second, each step completes in under half a second. If the agent executes five tool iterations within 40 seconds, the cumulative input tokens reach 145,000. On subsequent iterations, the accumulating payload pushes cumulative consumption past the 150,000-token per-minute threshold, and Cerebras halts execution with an HTTP 429 rate limit error.

The agent stalls mid-task, not because the model was incapable of solving the problem, but because the architecture treated the LLM prompt as an arbitrary file transport layer.

### Why Prompt Caching Only Partially Mitigates the Problem

Prompt caching allows unchanged prompt prefixes to bypass the uncached token bucket, drawing only from the higher Total TPM bucket (which defaults to three times uncached capacity). While prompt caching is helpful, relying on it as your sole mitigation strategy introduces several vulnerabilities:

* **Prefix Invalidation on Dynamic Tool Ingestion:** Prompt caching requires byte-exact prefix continuity. If an agent inserts dynamic context (such as timestamps, variable tool outputs, or files read in changing order) early in the prompt, the cache invalidates from that point onward, converting subsequent tokens into uncached cache misses.
* **Initial Ingestion Still Costs Uncached TPM:** The first time a document set or codebase is submitted, it must be written into the cache. Cache creation requires full compute and counts against your uncached TPM allowance. Submitting three large documents simultaneously can exhaust your uncached limit during the initial ingestion call alone.
* **No Price Discount on Cached Input:** In Cerebras inference tier billing, cached input tokens are billed at standard input rates rather than receiving the steep pricing discounts commonly offered on traditional GPU platforms. Sending massive file payloads repeatedly remains costly even when cache hit rates are high.

To maintain continuous agent operations without hitting 429 boundaries, teams must stop stuffing raw files into prompts and decouple persistent file storage from inference execution.

## Optimizing Payloads with Remote MCP Workspaces and Hybrid Search

The most effective strategy for staying under Cerebras rate limits is decoupling document storage from prompt payloads. Instead of attaching multi-megabyte PDFs, technical manuals, or full repositories directly to API calls, engineering teams store their corpus in external workspaces and retrieve targeted context on demand.

### Decoupling Document Storage from Inference Context

Rather than forcing the LLM to ingest entire documents on every turn, production systems implement a retrieval-first architecture:

1. **Persistent Cloud Storage:** Documents and data files live in a shared cloud repository.
2. **Automated Document Indexing:** Files are indexed automatically upon arrival for both full-text keywords and semantic vector representations.
3. **Targeted Excerpt Retrieval:** When an agent requires factual reference data, it queries the index and receives only the relevant paragraphs or data points.
4. **Lean Prompt Construction:** The prompt contains only the targeted excerpts (typically 300 to 800 tokens) rather than complete 40,000-token file bundles.

By shifting from whole-document ingestion to precision retrieval, prompt token volume drops by orders of magnitude, allowing autonomous agents to run dozens of tool iterations without approaching Cerebras uncached TPM caps.

### Managing Project Corpuses with Fast.io Workspaces

[Fast.io workspaces](/product/workspaces/) provide the storage and coordination substrate for agentic teams and multi-agent systems. When teams upload files or import them from external sources, Fast.io's Intelligence Mode indexes document contents automatically in the cloud.

The platform parses diverse file formats, including PDFs, Microsoft Word documents, presentations, spreadsheets, and source code. Fast.io creates a hybrid search index combining keyword search and semantic vector embeddings, eliminating the need to deploy and maintain external vector databases, embedding pipelines, or custom chunking scripts.

For workflows that require structured data extraction from invoices, contracts, or technical reports, [Metadata Views](/product/document-data-extraction/) convert unstructured documents into queryable tables. Users describe target attributes in natural language (such as contract renewal dates, indemnification caps, policy numbers, or line-item totals). Fast.io designs a typed schema (supporting Text, Integer, Decimal, Boolean, URL, JSON, and Date & Time) and extracts structured values automatically without requiring manual OCR templates.

Monthly plans start with a trial of up to 30 days (credit card required); annual plans have no trial. Subscription plans on [Fast.io pricing](/pricing/) include Starter, Business, and Enterprise options tailored to different storage sizes and team seats.

### Connecting Cerebras Agents to Fast.io via Remote MCP

Fast.io provides native integration with agent ecosystems through a consolidated remote Model Context Protocol (MCP) server. The server operates over Streamable HTTP at `https://mcp.fast.io/mcp` (with legacy SSE support at `https://mcp.fast.io/sse`). For secure programmatic agent access, clients connect to `https://mcp.fast.io/mcp/key` using Bearer token authentication.

For developers configuring agent environments, review the [storage for agents guide](/storage-for-agents/) and onboarding specifications at [fast.io/llms.txt](https://fast.io/llms.txt).

Configure your MCP client configuration file (such as `claude_desktop_config.json` or custom agent settings) to connect to the remote endpoint:

```json
{
  "mcpServers": {
    "fastio": {
      "url": "https://mcp.fast.io/mcp/key",
      "headers": {
        "Authorization": "Bearer YOUR_FASTIO_API_KEY"
      }
    }
  }
}
```

### Querying Cerebras with Workspace Retrieval

When querying Cerebras models through the OpenAI-compatible Python client, the agent queries the Fast.io workspace for targeted context before dispatching the inference request:

```python
from openai import OpenAI

client = OpenAI(
    base_url="https://api.cerebras.ai/v1",
    api_key="YOUR_CEREBRAS_API_KEY"
)

focused_context = (
    "Section 4.2: Maximum liability cap is established in accordance with Delaware law. "
    "Governing law is specified as the State of Delaware."
)

prompt_message = f"Context: {focused_context} Question: What is the liability cap?"

response = client.chat.completions.create(
    model="qwen-3.8-27b",
    messages=[
        {
            "role": "system",
            "content": "You are a legal review assistant. Rely strictly on provided context."
        },
        {
            "role": "user",
            "content": prompt_message
        }
    ],
    max_completion_tokens=150,
    temperature=0.2
)

print(response.choices[0].message.content)
```

By substituting raw document attachments with focused excerpt retrieval, the request consumes approximately 200 input tokens instead of 35,000 tokens. This reduction keeps your application running at peak wafer-scale speed while remaining safely below Cerebras rate limits.

## Handling HTTP 429 Errors: Headers, Exponential Backoff, and Payload Hygiene

Even in optimized environments, traffic spikes and concurrent worker pools can occasionally encounter rate limits. Handling HTTP 429 errors cleanly requires monitoring response headers, implementing jittered exponential backoff, and enforcing strict payload hygiene.

### Inspecting Cerebras Rate Limit Response Headers

When Cerebras throttles an API request, it returns an HTTP 429 Too Many Requests status code. The response payload and HTTP headers provide diagnostic information to guide retry timing:

* `retry-after`: Specifies the duration in seconds that your client must wait before retrying the request. When present, this header provides the definitive pause interval dictated by the server's token bucket replenishment rate.
* `x-ratelimit-limit-requests-minute`: The organization's total RPM limit for the target model.
* `x-ratelimit-remaining-requests-minute`: Remaining request capacity in the current minute window.
* `x-ratelimit-limit-tokens-minute`: Total tokens per minute allowed for the organization.
* `x-ratelimit-remaining-tokens-minute`: Estimated tokens remaining in the replenishment bucket.

Monitoring these headers in client middleware allows applications to track quota headroom dynamically and throttle outbound dispatches proactively before errors occur.

### Implementing Exponential Backoff with Randomized Jitter

When an HTTP 429 error occurs, client code must never retry immediately. Immediate retries consume depleting request capacity and extend the throttling penalty.

Instead, inspect the `retry-after` header. If present, pause for the specified seconds plus a small randomized jitter (such as 100 to 500 milliseconds) to prevent synchronized retry waves across worker threads. If the header is absent, apply exponential backoff:

```typescript
async function executeWithBackoff<T>(
  apiCall: () => Promise<T>,
  maxRetries = 5,
  baseDelayMs = 1000
): Promise<T> {
  for (let attempt = 0; attempt < maxRetries; attempt++) {
    try {
      return await apiCall();
    } catch (error: any) {
      const isRateLimited = error?.status === 429 || error?.statusCode === 429;
      if (!isRateLimited || attempt === maxRetries - 1) {
        throw error;
      }
      const retryAfterHeader = error?.headers?.get?.("retry-after");
      let delayMs: number;
      if (retryAfterHeader) {
        delayMs = parseFloat(retryAfterHeader) * 1000 + Math.random() * 500;
      } else {
        const backoff = baseDelayMs * Math.pow(2, attempt);
        delayMs = backoff * (0.5 + Math.random() * 0.5);
      }
      await new Promise((resolve) => setTimeout(resolve, delayMs));
    }
  }
  throw new Error("Maximum retry attempts exceeded");
}
```

### Production Checklist for Staying Under Cerebras Quotas

Follow this engineering checklist to maintain reliable throughput on Cerebras inference endpoints:

* **1. Set `max_completion_tokens` Conservatively:** Always define an explicit, realistic output cap. Because Cerebras evaluates rate limits pre-request by combining input tokens with `max_completion_tokens`, excessive output limits cause false 429 rejections.
* **2. Decouple File Corpuses from Prompts:** Never embed multi-page documents or complete repositories in system instructions. Index archives in Fast.io workspaces and use MCP tools to retrieve focused excerpts.
* **3. Structure Fixed System Prompts for Caching:** Place unchanging system instructions and tool definitions at the very beginning of the prompt. Avoid placing dynamic variables, user IDs, or timestamps ahead of static instructions.
* **4. Prune Conversational Turn History:** Implement a rolling context window for multi-turn agent chats. Summarize older conversation turns or discard intermediate tool noise to prevent unbounded token growth.
* **5. Distribute Workloads Across Models:** Cerebras meters rate limits independently per model class. If an application performs both complex reasoning and simple formatting, route formatting tasks to lighter models to preserve your primary token allocation.
* **6. Pace Concurrent Worker Pools:** When deploying parallel background threads, use token-aware rate limiters (such as leaky-bucket dispatchers) to smooth outbound requests across the rolling minute rather than firing synchronized batches.

## Frequently asked questions

### What is the rate limit for Cerebras inference?

Cerebras rate limits govern API request frequency and token velocity, evaluated independently per model and tier. On the Free Trial tier, limits are typically 5 RPM and 30,000 Uncached TPM (with 90,000 Total TPM) on flagship models like gpt-oss-120b and qwen-3.8-27b, alongside daily caps of 1,000,000 tokens. On the Developer tier (Pay as You Go), hourly and daily caps are removed, providing 1,000 RPM and 1,000,000 Uncached TPM on gpt-oss-120b, and 300 RPM with 150,000 Uncached TPM on qwen-3.8-27b.

### How do I handle Cerebras TPM rate limit errors?

To resolve Cerebras TPM rate limit errors, inspect the HTTP 429 response to see whether you exceeded the uncached or total token limit. Implement exponential backoff with randomized jitter using the retry-after header. To permanently prevent errors, configure max_completion_tokens conservatively to stop pre-request overestimation, prune chat history, and index large document files in a shared workspace like Fast.io so agents retrieve only targeted excerpts via MCP.

### How can I reduce token usage with Cerebras?

You can reduce Cerebras token usage by decoupling file storage from prompt construction. Instead of sending full PDFs, code files, or spreadsheets in API calls, store files in a Fast.io workspace with Intelligence Mode enabled. An agent querying Fast.io via remote MCP retrieves only the specific paragraphs needed for its task, dramatically reducing prompt token volume while keeping requests well below Cerebras TPM limits.

### What is the difference between uncached TPM and total TPM on Cerebras?

Uncached TPM measures tokens that require fresh compute on the Wafer-Scale Engine, representing cache misses. Total TPM measures combined token volume, including both uncached tokens and prompt cache hits. By default, Total TPM is set to three times your uncached TPM limit. Both buckets are enforced independently, and a 429 error message indicates which limit was exceeded.

### Does Cerebras support prompt caching to bypass rate limits?

Prompt caching allows repeated prompt prefixes to draw from the higher Total TPM bucket rather than the primary Uncached TPM bucket. However, cached tokens still count toward your Total TPM ceiling and initial cache writes consume uncached tokens. Furthermore, Cerebras bills cached input tokens at standard input rates, so caching alone does not eliminate the cost or throughput risks of oversized document prompts.

### Why does Cerebras rate limit requests before they finish processing?

Cerebras evaluates rate limits pre-request by adding prompt input tokens to either the max_completion_tokens parameter or an internal estimate of output length. If this combined estimate exceeds your available quota, Cerebras rejects the request before inference starts. Setting max_completion_tokens to a realistic value prevents premature rate limiting.

### How do I upgrade from the Cerebras Free Trial to the Developer tier?

You upgrade to the Developer tier by purchasing API credits under the Billing tab in the Cerebras Cloud Console. Your first credit purchase removes daily token caps and expands per-minute allowances across supported models, reaching 1,000 RPM and 1,000,000 Uncached TPM on gpt-oss-120b.

## Sources

- [Cerebras Inference Documentation: Rate Limits](https://inference-docs.cerebras.ai/support/rate-limits) — Cerebras enforces API rate limits using a dual-bucket model separating uncached tokens from total tokens.
- [Cerebras Inference Documentation: Rate Limits](https://inference-docs.cerebras.ai/support/rate-limits) — Cerebras employs a continuous token bucketing algorithm to replenish capacity rather than resetting quotas at fixed clock intervals.

## About Fast.io

Fast.io provides shared workspaces where people and AI agents work on the same files, with built-in semantic search and citation-backed chat over what they hold. Agents reach it through a remote MCP server at https://mcp.fast.io/mcp, a REST API at https://api.fast.io/current/, and a command line client published on npm as @vividengine/fastio-cli.
