AI & Agents

Cohere Rate Limits: API Keys, Production Tiers, and 429 Handling

Cohere rate limits are programmatic caps on the number of requests per minute (RPM) and tokens per minute (TPM) that a client application can submit to Cohere API endpoints. Free trial keys are limited to 1,000 calls per month with strict per-minute caps, while production keys provide 500 requests per minute on Command models. When agents submit large multi-file contexts, token exhaustion triggers HTTP 429 errors. Indexing files in shared workspaces prevents quota exhaustion.

Derek Labian 18 min read Updated
Monitor Cohere inference and embedding rate limits across evaluation and production keys while querying indexed workspaces.

How Cohere Enforces Rate Limits: Key Tiers, Architecture, and Quota Replenishment

When an autonomous agent pipeline repeatedly resends entire multi-file collections to evaluate context, it exhausts Cohere per-minute API quotas in seconds. The failure is rarely a defect in the prompt; it is an architectural flaw where the language model API is forced to act as an unindexed file transport layer. Cohere rate limits are programmatic caps on the number of requests per minute (RPM) and tokens per minute (TPM) that a client application can submit to Cohere's API endpoints. According to Cohere documentation on rate limits, the platform provisions access through two distinct categories of credentials: evaluation keys and production keys.

Cohere differentiates API access between free evaluation keys and paid production keys with higher usage limits. Evaluation keys, also referred to as trial keys, are designed for experimentation, functional testing, and initial development. They provide free access to Cohere models without requiring an upfront credit card, but they impose strict throughput ceilings. Specifically, trial keys (and prod keys on newer Chat model variants) are limited to 1,000 API calls a month. On a per-minute basis, trial keys are restricted to 20 requests per minute on Command generation models, 10 requests per minute on Rerank, and 5 requests per minute on specialized endpoints like batch EmbedJob and Audio Transcriptions.

Production keys require the organization to configure active billing in the Cohere Dashboard. Once upgraded, the monthly call ceiling is removed, and per-minute throughput expands to handle commercial production traffic. For example, production keys support 500 requests per minute on standard generation models including Command A, Command R+, Command R, Command R7B, and North Mini Code, alongside 1,000 requests per minute on Rerank and 2,000 inputs per minute on Embed.

Rolling Window Velocity and Continuous Replenishment

Cohere calculates request velocity using a rolling 60-second window rather than resetting usage counters at the top of a clock minute. In a fixed-window architecture, an application could fire 500 requests in the final two seconds of a minute, followed immediately by 500 requests in the first two seconds of the next minute, resulting in an unbuffered burst of 1,000 requests over four seconds. A rolling window prevents this by evaluating request timestamps continuously.

Every incoming API request adds an event to the rolling counter. If the number of requests within the trailing 60 seconds reaches your tier limit, any additional incoming call is rejected with an HTTP 429 status code. Capacity replenishes continuously as older requests pass beyond the 60-second horizon. Understanding this rolling mechanism is critical for background processing loops: firing a burst of unthrottled worker threads at once will saturate your quota instantly, stalling downstream operations until the rolling minute clears.

Requests Per Minute Versus Inputs Per Minute

Rate limits on Cohere depend on the nature of the endpoint. Generation models are governed primarily by requests per minute (RPM). In contrast, embedding models evaluate throughput based on inputs per minute (IPM).

When invoking the Embed API, a single HTTP request can bundle hundreds of discrete text snippets in its input array parameter. Cohere meters this endpoint by counting the total number of text items submitted across all requests within the rolling minute. On both trial and production keys, the standard Embed endpoint processes 2,000 inputs per minute. If you submit a single API request containing 2,001 strings, the call fails immediately, even though your application only dispatched one HTTP request. For multimodal embeddings processing images, the production ceiling is set to 400 inputs per minute, while trial keys permit 5 image inputs per minute.

Comparing Cohere Rate Limits Across Models and Endpoints

Cohere applies distinct rate limits depending on the target model family, task type, and key tier. While high-volume classification and semantic search operations often rely on Embed and Rerank endpoints, interactive reasoning and multi-turn workflows depend on the Command model family.

The table below details current default rate limits across key tiers for Cohere API endpoints as published in official documentation:

Model or Endpoint Trial Key Rate Limit Production Key Rate Limit
Command A 20 req / min (1,000 calls / month) 500 req / min
Command R+ 20 req / min (1,000 calls / month) 500 req / min
Command R 20 req / min (1,000 calls / month) 500 req / min
Command R7B 20 req / min (1,000 calls / month) 500 req / min
North Mini Code 20 req / min (1,000 calls / month) 500 req / min
Command A+ 20 req / min (1,000 calls / month) Custom quota (contact sales)
Command A Reasoning 20 req / min (1,000 calls / month) Custom quota (contact sales)
Command A Translate 20 req / min (1,000 calls / month) Custom quota (contact sales)
Command A Vision 20 req / min (1,000 calls / month) Custom quota (contact sales)
Embed (Text) 2,000 inputs / min 2,000 inputs / min
Embed (Images) 5 inputs / min 400 inputs / min
EmbedJob (Batch) 5 req / min 50 req / min
Rerank 10 req / min 1,000 req / min
Tokenize 100 req / min 2,000 req / min
Parse 500 req / min 500 req / min
Default Endpoints 500 req / min 500 req / min

Command Generation Models and Reasoning Variants

Standard Command models (Command A, Command R+, Command R, and Command R7B) provide a reliable production allocation of 500 requests per minute. This volume accommodates parallelized web applications, interactive customer support agents, and internal analysis workflows.

However, Cohere's newer model variants, including Command A+, Command A Reasoning, Command A Translate, and Command A Vision, operate under special access rules. On trial keys, these models remain capped at 20 requests per minute and are subject to the monthly 1,000 call ceiling. On production keys, these models do not automatically inherit the standard 500 RPM allocation. Production keys function like trial keys on these variants unless the organization contacts Cohere sales to negotiate a dedicated production quota. Teams planning to deploy reasoning or vision agents in customer-facing production must verify their quota allocation beforehand to avoid unexpected throttling in production environments.

Search, Rerank, and Batch Embeddings

Vector search architectures frequently pair Cohere Embed with Cohere Rerank. Understanding how rate limits intersect across these two stages is critical for retrieval-augmented generation (RAG):

  • Embedding Vector Ingestion: The Embed endpoint accepts batches of text strings reaching 2,000 inputs per minute. For offline document corpus ingestion that exceeds this throughput, developers should use the EmbedJob endpoint. EmbedJob processes large datasets asynchronously in the background, with a dispatch limit of 50 requests per minute on production keys.
  • Two-Stage Retrieval with Rerank: After retrieving candidate documents from a vector index or lexical search engine, applications pass top candidates through the Rerank endpoint. While trial keys allow 10 requests per minute, production keys support 1,000 requests per minute. This allows production systems to execute real-time reranking on every user query without bottlenecking search latency.
  • Tokenization Overhead: The Tokenize endpoint supports 2,000 requests per minute on production keys (100 on trial keys), allowing client applications to pre-calculate token boundaries and enforce payload hygiene before submitting generation requests.

Why Autonomous Agents and Payload Bloat Trigger Immediate HTTP 429 Errors

When engineering teams encounter HTTP 429 errors from Cohere, the immediate assumption is often that concurrent user traffic has exceeded capacity. In modern agentic architectures, however, rate limit exhaustion is rarely caused by user concurrency. It is driven by autonomous execution loops and unchecked payload bloat.

Unlike human chat interfaces where users pause between questions, autonomous agents execute in rapid automated cycles. A research or coding agent might inspect a file, issue a tool call, inspect output, query another file, and refine its conclusions. Each turn executes within hundreds of milliseconds.

The Compounding Overhead of Multi-Turn Context

Consider an autonomous agent assigned to audit a collection of legal agreements or technical manuals. The agent loop operates sequentially:

  1. Step 1 (Initialization): System instructions, persona guidelines, and initial objective (3,000 tokens).
  2. Step 2 (First File Read): Step 1 context plus full contents of contract agreement A (18,000 tokens).
  3. Step 3 (Tool Invocation): Step 2 context plus extraction output, error logs, and contract agreement B (34,000 tokens).
  4. Step 4 (Cross-Referencing): Step 3 context plus contract agreement C and compliance checklist (52,000 tokens).

With every iteration, the prompt payload expands. If the agent makes five tool calls within thirty seconds, it submits hundreds of thousands of cumulative tokens across successive requests. On a trial key, firing five requests within 15 seconds consumes a quarter of the entire per-minute allowance (20 RPM), and running four such evaluation runs exhausts the 1,000 call monthly allocation entirely.

Even on production keys with 500 RPM, retransmitting megabytes of repetitive context across every turn causes network latency, spikes token billing, and increases the probability of encountering concurrency throttles during peak traffic periods.

The Architectural Flaw of Treating Prompts as Storage

The underlying issue is that many agent frameworks treat the model prompt as an unindexed file transport layer. When an agent needs information from a 100-page operational manual or a 50-file software repository, dumping the entire raw corpus into the prompt forces the model to reprocess identical tokens on every cycle.

Most competitive advice recommends adding arbitrary pauses or purchasing enterprise tier expansions. While backoff logic is necessary for resilience, it does not fix the root cause. The proper solution is decoupling file persistence from model inference. When documents live in an intelligent storage layer, the agent retrieves only the specific sentences or data points needed for the immediate step, keeping prompt payloads compact and request rates stable.

Diagram showing prompt token bloat and rate limit exhaustion in agent workflows
Fastio features

Query document collections without exhausting Cohere rate limits

Connect Cohere agents to indexed workspaces over remote MCP to search multi-gigabyte files without payload bloat. Monthly plans start with a trial of up to 30 days (credit card required); annual plans have no trial.

Decoupling Document Storage from Inference Context with Remote Workspaces

To eliminate payload bloat and prevent rate limit exhaustion, production architectures decouple document storage from the LLM context window. Instead of attaching full PDF files, spreadsheets, or code repositories directly to Cohere API requests, engineering teams store their corpus in external workspaces and query indexed content on demand.

Architecture of a Retrieval-First Agent Pipeline

A retrieval-first architecture replaces raw file stuffing with targeted precision extraction:

  1. Persistent Cloud Storage: Project files and reference documentation live in a shared cloud repository rather than in transient local containers.
  2. Automated Document Indexing: When files are uploaded or imported, an automated indexing engine extracts text, generates vector embeddings, and builds full-text search indexes.
  3. Targeted Excerpt Retrieval: When an agent needs factual evidence, it submits a targeted search query and receives only the relevant 300 to 600 token excerpts.
  4. Lean Context Construction: The agent passes only the retrieved excerpts into Cohere's Chat API, executing reasoning over focused text rather than megabytes of raw files.

This retrieval pattern reduces prompt token volume by orders of magnitude, allowing agents to run dozens of tool iterations without approaching rate limits.

Indexing Corpuses with Fast.io Workspaces

Fast.io workspaces provide the persistent storage and coordination substrate for agentic teams. When documents are placed in an org-owned workspace, Fast.io's Intelligence Mode indexes file contents automatically in the cloud.

The platform handles diverse document formats, including PDFs, Microsoft Office documents, spreadsheets, text files, and source code. Fast.io creates a hybrid search index combining full-text keyword search and semantic vector embeddings, eliminating the need for teams to deploy, configure, and maintain separate vector databases or embedding pipelines.

For workflows that require structured field extraction from financial records, legal contracts, or inventory spreadsheets, Metadata Views convert unstructured documents into queryable tables. Teams describe the fields they want extracted in natural language, such as contract effective dates, counterparties, invoice totals, or insurance policy limits. Fast.io generates a typed schema (supporting Text, Integer, Decimal, Boolean, URL, JSON, and Date & Time) and extracts structured values directly from files without manual OCR templates.

Monthly plans start with a trial of up to 30 days (credit card required); annual plans have no trial. Subscription plans on Fast.io pricing include Starter, Business, and Enterprise options tailored to different storage sizes and team seats.

Connecting Cohere Agents to Fast.io via Remote MCP

Fast.io exposes its storage and intelligence capabilities through a consolidated remote Model Context Protocol (MCP) server. The server operates over Streamable HTTP at https://mcp.fast.io/mcp (with legacy SSE available at https://mcp.fast.io/sse). Programmatic agents authenticate securely against https://mcp.fast.io/mcp/key using a Bearer token.

For developers setting up agent environments, review the storage for agents guide and onboarding documentation at fast.io/llms.txt.

Configure your MCP client settings to connect to the remote Fast.io endpoint:

{
  "mcpServers": {
    "fastio": {
      "url": "https://mcp.fast.io/mcp/key",
      "headers": {
        "Authorization": "Bearer YOUR_FASTIO_API_KEY"
      }
    }
  }
}

Lean Context Querying with Cohere Client

When querying Cohere models in Python, the agent retrieves targeted excerpts from Fast.io through MCP before dispatching the inference request:

import cohere

client = cohere.ClientV2(api_key="YOUR_COHERE_API_KEY")

targeted_context = (
    "Section 8.1: Service Level Agreement. Monthly uptime commitment is maintained "
    "at standard enterprise availability tiers. Credits are issued upon verified request."
)

user_query = "What is the service commitment stated in the SLA?"

prompt = f"Context: {targeted_context} Question: {user_query}"

response = client.chat(
    model="command-r-plus",
    messages=[
        {
            "role": "system",
            "content": "You are a precise technical analyst. Answer solely based on the provided context."
        },
        {
            "role": "user",
            "content": prompt
        }
    ]
)

print(response.message.content[0].text)

By passing targeted excerpts rather than uploading the entire service agreement on every turn, prompt input overhead drops from tens of thousands of tokens to a few concise paragraphs, protecting your application from rate limits while maintaining fast response times.

Handling HTTP 429 Errors: Response Headers, Exponential Backoff, and SDK Integration

Even in optimized production systems, bursty traffic and parallel worker pools can occasionally exceed rate limits. Building resilient applications requires inspecting HTTP 429 response headers, implementing jittered exponential backoff, and catching typed SDK exceptions.

Inspecting Cohere Rate Limit Response Headers

When Cohere rejects an API request due to rate limiting, it returns an HTTP 429 Too Many Requests status code. The response headers provide essential diagnostic metadata:

  • retry-after: Indicates the exact duration in seconds that the client must wait before resubmitting the request. Respecting this header ensures that client retries align precisely with Cohere's rolling window replenishment.
  • x-ratelimit-limit-requests-minute: The total number of requests per minute permitted for the credential on the requested endpoint.
  • x-ratelimit-remaining-requests-minute: The remaining request capacity available in the current rolling minute window.

Monitoring these headers in client middleware allows applications to track quota depletion dynamically and introduce proactive throttling before errors occur.

Implementing Exponential Backoff with Randomized Jitter in Python

When an application encounters an HTTP 429 error, retrying immediately is counterproductive. Immediate retries consume depleting request capacity and extend the throttling penalty. If multiple distributed workers retry simultaneously, they create a synchronized wave of requests, known as a thundering herd, that repeatedly triggers rate limits.

The recommended pattern combines exponential backoff with randomized jitter. The delay expands exponentially on each failed attempt, while a random jitter factor decorrelates retry timing across worker threads:

import time
import random
import cohere

client = cohere.ClientV2(api_key="YOUR_COHERE_API_KEY")

def execute_chat_with_retry(messages, model="command-r-plus", max_retries=5, base_delay=1.0):
    for attempt in range(max_retries):
        try:
            return client.chat(model=model, messages=messages)
        except cohere.errors.TooManyRequestsError as error:
            if attempt == max_retries - 1:
                raise error
            retry_after = getattr(error, "retry_after", None)
            if retry_after is not None:
                sleep_time = float(retry_after) + random.uniform(0.1, 0.5)
            else:
                backoff = base_delay * (2 ** attempt)
                sleep_time = backoff * random.uniform(0.8, 1.2)
            time.sleep(sleep_time)
    raise RuntimeError("Failed after maximum retries")

TypeScript Client Implementation with Header Parsing

In Node.js or browser environments, client applications can implement exponential backoff by wrapping the fetch call or SDK client:

interface RequestOptions {
  maxRetries?: number;
  baseDelayMs?: number;
}

async function fetchWithBackoff<T>(
  apiCall: () => Promise<T>,
  options: RequestOptions = {}
): Promise<T> {
  const maxRetries = options.maxRetries ?? 5;
  const baseDelayMs = options.baseDelayMs ?? 1000;
  for (let attempt = 0; attempt < maxRetries; attempt++) {
    try {
      return await apiCall();
    } catch (error: any) {
      const isRateLimited = error?.status === 429 || error?.statusCode === 429;
      if (!isRateLimited || attempt === maxRetries - 1) {
        throw error;
      }
      const retryAfterHeader = error?.headers?.get?.("retry-after");
      let delayMs: number;
      if (retryAfterHeader) {
        delayMs = parseFloat(retryAfterHeader) * 1000 + Math.random() * 400;
      } else {
        const exponentialDelay = baseDelayMs * Math.pow(2, attempt);
        delayMs = exponentialDelay * (0.8 + Math.random() * 0.4);
      }
      await new Promise((resolve) => setTimeout(resolve, delayMs));
    }
  }
  throw new Error("Maximum retry limit exceeded");
}

Production Checklist for Cohere Rate Limit Management

Follow this engineering checklist to maintain reliable throughput on Cohere endpoints:

  • 1. Upgrade to Production Keys for Commercial Deployments: Do not run customer-facing applications on free trial keys. Adding billing details removes the 1,000 monthly call ceiling and expands RPM allowances.
  • 2. Confirm Custom Allocations for Advanced Models: If using Command A Reasoning, Command A Vision, or Command A+, contact Cohere sales to verify production quotas before launching.
  • 3. Decouple Corpuses from Prompt Payloads: Avoid sending raw files into model prompts. Store reference corpuses in Fast.io workspaces with Intelligence Mode enabled and query targeted excerpts via remote MCP.
  • 4. Respect the Retry-After Header: Configure retry handlers to honor the retry-after header value rather than using static sleep intervals.
  • 5. Apply Randomized Jitter to All Retries: Decorrelate retry attempts across parallel workers to prevent synchronized request spikes.
  • 6. Batch Vector Ingestion with EmbedJob: When embedding thousands of documents, use asynchronous EmbedJob batches rather than firing rapid calls against synchronous Embed endpoints.
  • 7. Monitor Quota Metrics in the Cohere Dashboard: Regularly review usage graphs and request velocity under the Billing and Usage tabs to anticipate capacity needs before scaling.

Sources

References used to verify factual claims in this guide.

  1. Cohere limits trial API keys and production keys on newer chat variants to 1,000 API calls a month. Cohere differentiates API access between free evaluation keys and paid production keys with higher usage limits.

Frequently Asked Questions

What is the rate limit on Cohere free trial keys?

Cohere free trial evaluation keys are restricted to a total ceiling of 1,000 API calls per month across all endpoints. On a per-minute basis, trial keys allow 20 requests per minute on Command generation models, 10 requests per minute on Rerank, 5 requests per minute on Audio Transcriptions and batch EmbedJobs, 100 requests per minute on Tokenize, and 2,000 inputs per minute on text Embed.

How do I increase my Cohere rate limit?

You increase your Cohere rate limit by adding billing information in the Cohere Dashboard to upgrade from a trial key to a production key. Upgrading removes the 1,000 monthly call ceiling and increases generation limits to 500 requests per minute on Command A, Command R+, Command R, Command R7B, and North Mini Code. For higher custom volume or production access to advanced reasoning and vision variants, contact Cohere sales at sales@cohere.com.

How do I handle Cohere 429 rate limit errors?

To handle Cohere HTTP 429 errors, catch the cohere.errors.TooManyRequestsError exception in Python or check for status code 429 in HTTP clients. Inspect the retry-after header in the response and pause for the indicated number of seconds plus a small randomized jitter. If the header is missing, implement exponential backoff starting at one second. To permanently prevent rate limits in multi-turn agent workflows, index large document corpuses in a shared workspace like Fast.io and retrieve targeted excerpts via remote MCP rather than re-sending raw files.

What is the difference between requests per minute and inputs per minute on Cohere?

Requests per minute (RPM) measures the number of distinct HTTP requests sent to an endpoint within a rolling 60-second window, which is how Command generation and Rerank endpoints are governed. Inputs per minute (IPM) measures the total count of text snippets or images included in the payload across all requests. Cohere's Embed endpoint permits 2,000 text inputs per minute regardless of whether they are sent across ten requests or a single batch request.

Does Cohere enforce daily or monthly token limits?

Cohere does not enforce a daily token cap on standard production keys; usage is billed per token on a pay-as-you-go basis up to your configured monthly billing threshold. However, free trial keys are governed by a strict monthly ceiling of 1,000 API calls. Newer model variants like Command A Reasoning also operate under trial call limits until production quota is explicitly provisioned by Cohere.

How does workspace retrieval reduce Cohere rate limit errors in agent pipelines?

Autonomous agents frequently trigger rate limits by re-submitting complete documents and conversational history on every turn of a loop. By storing documents in a Fast.io workspace with Intelligence Mode enabled, files are indexed for semantic and full-text search. The agent queries Fast.io over remote MCP to fetch only the relevant 300 to 600 token passages needed for its immediate step. This reduces prompt payload size substantially and stops redundant API calls, keeping agent pipelines safely within Cohere rate limits.

Related Resources

Fastio features

Query document collections without exhausting Cohere rate limits

Connect Cohere agents to indexed workspaces over remote MCP to search multi-gigabyte files without payload bloat. Monthly plans start with a trial of up to 30 days (credit card required); annual plans have no trial.