AI & Agents

LlamaIndex Rate Limits: Ingestion Batching, Embedding Quotas, and Offloaded Indexing

LlamaIndex rate limits are API request and token bottlenecks triggered while parsing, chunking, and embedding large file corpora through third-party embedding models. When unthrottled ingestion pipelines process hundreds of documents simultaneously, upstream providers enforce per-minute caps that return HTTP 429 errors. Tuning embedding batch sizes and offloading indexing to intelligent workspaces via remote MCP prevents token exhaustion and stabilizes agent pipelines.

Tom Langridge 16 min read Updated
Manage LlamaIndex ingestion rate limits by tuning embedding batch sizes and offloading vector indexing to remote workspaces.

What Causes LlamaIndex Rate Limits During Document Ingestion

Ingesting a corpus of 1,000 PDFs through local pipelines frequently triggers OpenAI Tier 2 embedding rate limits within minutes. When an automated ingestion pipeline parses and chunks hundreds of documents simultaneously, unthrottled worker processes flood third-party embedding endpoints with heavy token volumes, triggering HTTP 429 errors that halt data indexing.

LlamaIndex rate limits are API request and token bottlenecks triggered while parsing, chunking, and embedding large file corpora through third-party embedding models. Understanding how these rate limits operate requires examining the intersection between client-side ingestion architecture and upstream provider token quotas.

LlamaIndex itself is an open-source data framework that orchestrates document ingestion, chunking, embedding, and retrieval. It does not enforce rate limits on its own core software. Instead, rate limits originate from external foundation model APIs, such as OpenAI, Cohere, Voyage AI, or Anthropic, that LlamaIndex calls to generate vector representations or extract metadata. When engineers build retrieval-augmented generation (RAG) pipelines, they often treat document ingestion as a simple file reading loop, only to discover that embedding thousands of chunks overruns API velocity thresholds.

Upstream Quotas: RPM, TPM, and Tier Ceilings

The foundation model providers enforce rate limits to protect infrastructure from overload and distribute capacity fairly among developers. According to the official OpenAI rate limits guide, rate limits are defined at the organization level and at the project level, not user level. These constraints operate across several distinct dimensions:

  • Requests Per Minute (RPM): The total number of individual HTTP calls allowed within a rolling sixty-second window.
  • Tokens Per Minute (TPM): The total volume of input text tokens processed across all requests in a rolling sixty-second window.
  • Monthly Usage Limits: The maximum dollar spend permitted per calendar month, determined by account qualification tier.

The table below outlines the usage tier pricing and monthly plan limits across developer accounts on the OpenAI platform:

Usage Tier Qualification Threshold Monthly Plan Limit Typical Embedding Rate Ceiling
Free Tier Verified account in supported geography $100 monthly cap 3 RPM, 10,000 TPM
Tier 1 $5 paid $100 monthly cap 500 RPM, 1,000,000 TPM
Tier 2 $50 paid and 7 days elapsed $500 monthly cap 5,000 RPM, 5,000,000 TPM
Tier 3 $100 paid and 7 days elapsed $1,000 monthly cap 5,000 RPM, 10,000,000 TPM
Tier 4 $250 paid and 14 days elapsed $5,000 monthly cap 10,000 RPM, 20,000,000 TPM
Tier 5 $1,000 paid and 14 days elapsed $200,000 monthly cap 10,000 RPM, 50,000,000 TPM

When an application breaches either the RPM or TPM limit, the upstream API immediately responds with an HTTP 429 status code and a structured rate_limit_exceeded error message. In document ingestion pipelines, TPM limits represent the primary bottleneck because embedding models process substantial text volumes in single API calls.

The Ingestion Math: Why Multi-Document Collections Exhaust Quotas

To understand why rate limits trigger during local batch processing, consider what happens when ingesting a collection of hundreds of technical documents, legal agreements, or research files.

When large documents are ingested through a standard node parser with chunking enabled, each document produces numerous discrete text nodes. A collection containing thousands of pages easily generates tens of thousands of chunks.

If an unthrottled ingestion script runs with multiple concurrent threads or async workers, it attempts to push these nodes to the embedding endpoint as fast as the network permits. On introductory tiers, transmitting these chunks saturates the rolling minute token allowance in seconds. Even on higher developer tiers, a burst of parallel requests breaches the rolling window quickly, causing the ingestion job to halt with repeated 429 exceptions.

How to Configure Ingestion Pipelines for Optimal Batch Sizes and Concurrency

The primary cause of LlamaIndex ingestion rate limits is unthrottled concurrency combined with oversized embedding requests. The optimal embed_batch_size setting for OpenAI embedding models in LlamaIndex is typically 20 to 50 nodes per batch, paired with num_workers=1 or num_workers=2 in IngestionPipeline transformations. This configuration balances API request latency with token throughput while preventing bursts that trigger HTTP 429 errors.

Managing throughput requires adjusting parameters across the embedding model and the ingestion pipeline itself. Developers often look for batch size settings on the IngestionPipeline class, but in LlamaIndex, batching is configured on the underlying embedding component.

Configuring embed_batch_size on the Embedding Model

The embed_batch_size parameter belongs to the BaseEmbedding class from which all provider-specific embedding integrations inherit. When you instantiate an embedding model, passing embed_batch_size determines how many text nodes LlamaIndex bundles into a single API payload.

Finding the correct batch size requires balancing two failure modes:

  • Oversized Batches: Large batches reduce the total count of HTTP requests, which protects your RPM quota. However, sending hundreds of chunks in a single call creates a heavy token payload. A single call can exhaust the per-request token ceiling or immediately consume your organization's rolling TPM budget.
  • Undersized Batches: Very small batches prevent TPM spikes because each call carries only a few hundred tokens. However, processing a large document collection with tiny batches requires thousands of individual HTTP requests. At high concurrency, this easily exceeds RPM ceilings on lower tiers and introduces substantial network round-trip overhead.

For models like text-embedding-3-small and text-embedding-3-large, an embed_batch_size of 20 to 40 nodes provides a reliable balance. A batch of 30 nodes with standard chunk sizes produces a predictable token payload per request, well within per-call payload allowances and straightforward to pace within rolling rate windows.

Managing Worker Threads with num_workers

The IngestionPipeline class accepts a num_workers argument that controls execution parallelism. While increasing num_workers speeds up CPU-bound transformations such as regular-expression text cleaning and sentence splitting, it multiplies API traffic when applied to remote embedding steps.

When worker concurrency is configured aggressively alongside standard batch sizes, the pipeline initiates numerous concurrent requests simultaneously, pushing hundreds of nodes and heavy token volumes across the network at once. Setting num_workers=1 enforces strict serial execution, where each batch must complete before the next begins. For production ingestion runs where API quotas are tight, num_workers=1 or num_workers=2 prevents concurrent request clustering.

Building a Rate-Aware Ingestion Pipeline

The following implementation demonstrates how to construct a rate-conscious ingestion pipeline in Python using LlamaIndex. It configures the embedding model with an explicit batch size and runs transformations with controlled worker concurrency:

from llama_index.core import SimpleDirectoryReader
from llama_index.core.ingestion import IngestionPipeline
from llama_index.core.node_parser import SentenceSplitter
from llama_index.embeddings.openai import OpenAIEmbedding

embed_model = OpenAIEmbedding(
    model="text-embedding-3-small",
    embed_batch_size=30,
    max_retries=5,
    timeout=60.0,
)

node_parser = SentenceSplitter(
    chunk_size=512,
    chunk_overlap=50,
)

pipeline = IngestionPipeline(
    transformations=[
        node_parser,
        embed_model,
    ],
)

documents = SimpleDirectoryReader("./source_documents").load_data()
nodes = pipeline.run(
    documents=documents,
    num_workers=1,
    show_progress=True,
)

print(f"Successfully processed and embedded {len(nodes)} nodes.")

This configuration maintains manageable payload sizes on each request and ensures requests execute sequentially, eliminating the sudden traffic spikes that trigger 429 errors.

How to Implement Client-Side Throttling and Backoff Retries

Pacing batch sizes and limiting worker concurrency lowers the likelihood of hitting rate limits. However, when ingesting large document collections, external factors such as network hiccups and upstream provider latency fluctuations can still trigger sporadic HTTP 429 status codes. Production systems must implement resilient retry strategies and document caching.

Relying on naive sleep intervals (such as calling a fixed sleep delay after an error) is insufficient for production pipelines. If multiple parallel workers encounter rate limits simultaneously and each sleeps for a fixed duration, all workers wake up at the exact same moment and retry in unison. This synchronized wave immediately triggers another rate limit error, creating a thundering herd problem.

Exponential Backoff and Randomized Jitter

When API rate limits trigger, resilient pipelines apply exponential backoff combined with randomized jitter. When an API call fails with a rate limit error, the application calculates a wait duration that doubles with each consecutive retry attempt, while adding random variance to desynchronize concurrent threads.

Libraries like tenacity provide battle-tested retry decorators that can wrap custom embedding functions or API requests:

import time
from tenacity import (
    retry,
    stop_after_attempt,
    wait_random_exponential,
    retry_if_exception_type,
)
from openai import RateLimitError

@retry(
    wait=wait_random_exponential(min=4, max=60),
    stop=stop_after_attempt(6),
    retry=retry_if_exception_type(RateLimitError),
    reraise=True,
)
def execute_embedding_with_backoff(embed_model, text_batch):
    return embed_model.get_text_embedding_batch(text_batch)

Using randomized exponential backoff ensures that transient provider load spikes resolve naturally without crashing the ingestion job.

Pacing API Calls with RateLimiter

To pace API requests proactively, LlamaIndex supports explicit rate limiting via the rate_limiter parameter on BaseEmbedding classes. By injecting a rate limiter that enforces a token-bucket or leaky-bucket algorithm, the pipeline pauses execution proactively before making an API call, rather than waiting for an HTTP 429 failure.

Configuring a client-side rate limiter to pace requests well below nominal account ceilings provides a valuable safety buffer that absorbs estimation discrepancies and prevents the pipeline from bumping into provider limits.

Avoiding Re-Embedding with DocstoreStrategy.UPSERTS

The most efficient way to avoid embedding rate limits is to never embed the same document twice. In standard development scripts, running an ingestion pipeline repeatedly re-chunks and re-embeds the entire document collection from scratch, consuming API quota redundantly.

LlamaIndex solves this through the docstore_strategy parameter on IngestionPipeline. By pairing the pipeline with a persistent document store, LlamaIndex computes a unique hash for each document and node.

When configured with DocstoreStrategy.UPSERTS, the pipeline checks whether a document already exists in the document store:

  • If the document is absent, the pipeline processes and embeds it.
  • If the document exists and its hash is identical, the pipeline skips embedding generation entirely and reuses the cached node vectors.
  • If the document exists but its hash has changed, the pipeline updates the document and generates new embeddings only for modified sections.
from llama_index.core.ingestion import IngestionPipeline, DocstoreStrategy
from llama_index.core.storage.docstore import SimpleDocumentStore
from llama_index.core.node_parser import SentenceSplitter
from llama_index.embeddings.openai import OpenAIEmbedding
import os

docstore_path = "./pipeline_cache/docstore.json"
if os.path.exists(docstore_path):
    docstore = SimpleDocumentStore.from_persist_path(docstore_path)
else:
    docstore = SimpleDocumentStore()

embed_model = OpenAIEmbedding(
    model="text-embedding-3-small",
    embed_batch_size=30,
)

pipeline = IngestionPipeline(
    transformations=[
        SentenceSplitter(chunk_size=512, chunk_overlap=50),
        embed_model,
    ],
    docstore=docstore,
    docstore_strategy=DocstoreStrategy.UPSERTS,
)

nodes = pipeline.run(documents=documents, num_workers=1)
pipeline.persist(persist_dir="./pipeline_cache")

Persisting document state eliminates repetitive embedding costs and prevents pipeline re-runs from triggering unnecessary rate limit exceptions.

Client-side retry logic and backoff pacing in document ingestion pipelines
Fastio features

Stop hitting embedding rate limits in agent workflows

Move document indexing out of brittle local retry loops and into intelligent workspaces with semantic search and remote MCP integration. Starts with a 30-day free trial.

Why Client-Side Retry Loops Create Bottlenecks in Agent Workflows

Most technical tutorials and vendor documentation approach rate limits purely as a client-side tuning challenge, instructing engineers to write increasingly complex retry loops and sleep decorators. In simple batch scripts running overnight, adding backoff delays may seem acceptable. In production agentic systems, however, relying on client-side retry loops creates severe architectural liabilities.

Guides only teach client-side retry sleep loops instead of demonstrating remote indexing via MCP. While a backoff loop prevents an immediate script crash, it does not solve the fundamental constraint: local pipelines are still responsible for serializing, chunking, transmitting, and vectorizing massive document collections over external HTTP connections.

Latency Spikes and Agent Thread Blocking

When an autonomous agent processes incoming documents to answer a user question or complete a multi-step research task, entering an extended backoff sleep loop blocks the agent execution thread. Modern agent frameworks operate under strict timeout constraints. When tool calls stall during prolonged embedding backoff delays, surrounding agent orchestrators treat the operation as timed out, terminating the user session.

In interactive applications, users expect prompt responses. Forcing an agent to wait out provider rate limits turns an automated workflow into a sluggish, unresponsive experience.

Resource Exhaustion and Pipeline Fragility

The practice of holding un-embedded document chunks in memory across multiple backoff attempts inflates process memory consumption. If an ingestion worker crashes or restarts during an extended sleep cycle, the intermediate pipeline state is lost unless comprehensive distributed checkpointing is implemented.

Local embedding pipelines also require embedding API keys to be distributed across all client environments, worker nodes, and developer laptops. Managing API key rotation, spending limits, and organizational permissions across distributed execution environments introduces administrative complexity and increases the risk of key leakage.

Upstream Provider Degradation and Cascading Failure

When upstream model providers experience service degradation or regional traffic surges, provider rate limits become dynamic, rejecting requests well below nominal account tier ceilings.

Client-side retry loops respond to provider slowdowns by retrying repeatedly, worsening provider congestion and trapping the client in an infinite backoff cycle until retry budgets expire. When multiple agents or pipeline workers compete for the same organization-wide token pool, a single runaway batch script can exhaust quotas for all other applications and developer sessions across the company.

How to Offload Indexing to Intelligent Workspaces via Remote MCP

Offloading indexing to a managed workspace server eliminates local embedding API calls. Instead of maintaining local embedding pipelines, tuning embed_batch_size, and managing API keys for third-party vector databases, teams can store their document corpora directly in an intelligent cloud workspace.

This architectural shift separates file storage and document indexing from agent execution. Rather than pulling raw files to a local machine, parsing text, and sending chunked requests to OpenAI or Cohere, files live in an organization-owned workspace where indexing happens automatically in the cloud.

How Fast.io Intelligent Workspaces Index Knowledge

The Fast.io workspaces platform stores documents in shared cloud storage designed for agentic teams. Users and automated agents can upload files directly or import entire document repositories from cloud storage providers, including Google Drive, Dropbox, Box, and OneDrive.

When Intelligence Mode is enabled on a workspace, Fast.io automatically indexes documents on arrival. The workspace platform extracts text from PDFs, spreadsheets, presentations, and scans, generating full-text, semantic, and metadata indexes in the background. Because document processing occurs directly within the workspace infrastructure:

  • Local scripts and autonomous agents make zero calls to third-party embedding APIs.
  • Ingestion runs never trigger client-side TPM or RPM rate limit errors.
  • Developers do not need to configure embedding batch sizes, worker counts, or backoff decorators.
  • Document indexing scales across cloud workers without drawing down organizational LLM API allowances.

Querying Pre-Indexed Knowledge Through Remote MCP

When autonomous agents connect to Fast.io workspaces through the remote Model Context Protocol (MCP) server hosted at https://mcp.fast.io/mcp using Streamable HTTP, or https://mcp.fast.io/mcp/key with API key authentication, they gain access to a consolidated MCP toolset. Agents can also connect via legacy SSE at https://mcp.fast.io/sse.

Instead of stuffing large files into prompt context or running local LlamaIndex ingestion pipelines, the agent calls Fast.io's consolidated MCP toolset to search the workspace. The workspace executes hybrid search (combining full-text search, semantic relevance, and metadata filtering) and returns ranked, citation-backed text passages directly to the agent.

The following example illustrates how an autonomous agent or developer configures a remote MCP connection to query an intelligent workspace:

{
  "mcpServers": {
    "fastio": {
      "url": "https://mcp.fast.io/mcp/key",
      "headers": {
        "Authorization": "Bearer YOUR_FASTIO_API_KEY"
      }
    }
  }
}

Once connected, the agent invokes the consolidated storage tool with the search action:

{
  "name": "storage",
  "arguments": {
    "action": "search",
    "workspace_id": "ws_enterprise_contracts",
    "query": "What are the indemnification terms in the vendor master agreement?",
    "semantic": true
  }
}

The agent receives verified, citation-backed results extracted from thousands of documents in milliseconds, without performing any client-side vector operations or managing embedding rate limits.

Collaboration, Governance, and Structured Extraction

The shift toward shared workspaces provides governance and collaboration capabilities that local vector pipelines lack:

  • Per-File Version History: Every document in the workspace maintains full version history. When updated contracts or technical manuals arrive, the workspace updates indexes while preserving prior versions for complete traceability.
  • Granular Permissions: Permissions are enforced across organizations, workspaces, folders, and files. Agents access only the specific workspaces they are granted, preventing unauthorized data exposure across multi-agent environments.
  • Append-Only Audit Log: Every file read, search query, upload, and permission change is recorded in an immutable audit log, providing complete transparency into human and agent actions.
  • Structured Data Extraction: In addition to unstructured semantic search, Fast.io provides Metadata Views to convert document repositories into queryable tables. Metadata Views extract typed schema fields (such as contract renewal dates, policy limits, or invoice totals) from PDFs, images, and documents without manual template configuration.

Subscription Plans and Trial Access

Monthly plans start with a trial of up to 30 days (credit card required); annual plans have no trial. Paid plans include:

Plan Tier Monthly Price Best For
Starter $9.99/mo Developers deploying initial agent workspaces and project storage
Business $49.99/mo Growing teams requiring expanded storage capacity and workspace intelligence
Enterprise $199.99/mo Multi-agent production workloads and high-volume organizational repositories

Autonomous agents can register an organization, configure workspaces and shares, and transfer ownership to human administrators while retaining scoped administrative access.

Sources

References used to verify factual claims in this guide.

  1. 1 OpenAI: Rate Limits Guide Accessed

    OpenAI rate limits are enforced across developer accounts at both the organization and project levels rather than per user.

Frequently Asked Questions

How do I avoid rate limits in LlamaIndex?

To avoid rate limits in LlamaIndex, set embed_batch_size between 20 and 50 on your embedding model, limit concurrency by setting num_workers=1 or 2 in IngestionPipeline, and enable persistent docstore caching with DocstoreStrategy.UPSERTS to prevent re-embedding unchanged files. For large corpora, offload document indexing to an intelligent workspace like Fast.io via remote MCP, eliminating local embedding API calls entirely.

How do I throttle embeddings in LlamaIndex?

You can throttle embeddings in LlamaIndex by configuring the optional rate_limiter parameter on BaseEmbedding classes or by wrapping embedding calls with exponential backoff retry decorators using libraries like tenacity. This ensures that when requests approach provider velocity boundaries, the client pauses and retries with randomized jitter rather than failing with an unhandled exception.

What is the best batch size for LlamaIndex ingestion?

For OpenAI embedding models like text-embedding-3-small and text-embedding-3-large, the optimal batch size is typically 20 to 50 nodes per batch. This range minimizes HTTP request overhead while keeping per-request token payloads safely below per-minute token throughput limits.

Why do I get HTTP 429 errors when using IngestionPipeline with OpenAI?

HTTP 429 errors occur when your ingestion pipeline exceeds OpenAI's Requests Per Minute (RPM) or Tokens Per Minute (TPM) limits for your account's usage tier. High concurrency settings or processing hundreds of document chunks in a short burst exhausts the rolling sixty-second token window, prompting the API to reject subsequent requests.

How does offloading indexing to a workspace server eliminate embedding rate limits?

Offloading indexing moves document parsing, chunking, and embedding from local client machines to cloud workspace servers. When documents are placed in an intelligent workspace with Intelligence Mode enabled, the workspace automatically indexes the content on arrival. Autonomous agents then search the pre-indexed corpus via remote MCP without making local embedding API calls.

Related Resources

Fastio features

Stop hitting embedding rate limits in agent workflows

Move document indexing out of brittle local retry loops and into intelligent workspaces with semantic search and remote MCP integration. Starts with a 30-day free trial.