# Managing LangGraph Rate Limits, Node Throttling, and State Stores

LangGraph workflows trigger HTTP 429 rate limits when parallel branches and multi-agent graphs exceed upstream provider throughput ceilings. Resolving these bottlenecks requires distinguishing graph recursion bounds from API rate limits, configuring retry policies with backoff, and throttling concurrent node execution. Decoupling document storage from the graph state store prevents checkpoint bloat and token-heavy payload transfers across execution cycles.

Source: https://fast.io/resources/langgraph-rate-limit/
Author: [Derek Labian](https://fast.io/authors/derek-labian/)
Last reviewed: 2026-09-24

## How LangGraph Rate Limits Differ from Graph Recursion Bounds

When an automated LangGraph workflow halts with an execution error, developers frequently confuse internal graph recursion ceilings with upstream API rate limit throttling. LangGraph enforces an internal `recursion_limit` of 25 super-steps by default to catch runaway infinite loops and cyclic execution graphs. An HTTP 429 error, by contrast, indicates that outbound model calls have breached an external provider's requests-per-minute (RPM) or tokens-per-minute (TPM) quota. Because LangGraph's default retry mechanism excludes client-side 4xx errors, an unhandled 429 immediately crashes the entire graph run.

Understanding this boundary is essential for diagnosing failures in production agent graphs:

LangGraph rate limits occur when branching nodes or multi-agent state graphs issue concurrent LLM calls that exceed upstream provider RPM/TPM thresholds.

While both failure modes halt execution, their root causes, system boundaries, and diagnostic traces share nothing in common:

| Dimension | Graph Recursion Limit | Provider Rate Limit (429) |
| :--- | :--- | :--- |
| **System Boundary** | Internal to LangGraph runtime | External upstream provider API |
| **Root Mechanism** | Super-step count exceeds configured limit | Token or request rate breaches quota |
| **Default Threshold** | 25 super-steps per invocation | Provider-specific tier limits |
| **Raised Exception** | `GraphRecursionError` | `openai.RateLimitError` or HTTP 429 |
| **Primary Resolution** | Graph topology review or `recursion_limit` config | Jittered retries, rate limiters, storage decoupling |

A standard LangChain chain executes sequentially, making it relatively straightforward to pace outbound requests. LangGraph, however, is built for cyclic, branching, and multi-agent topologies. When a graph branches into parallel paths or dispatches autonomous sub-agents, request frequency spikes unpredictably. Developers troubleshooting these crashes often check the [LangGraph fault tolerance documentation](https://docs.langchain.com/oss/python/langgraph/fault-tolerance) and the [LangSmith rate limiting guide](https://docs.langchain.com/langsmith/handle-model-rate-limiting) to isolate whether the failure stems from runaway graph recursion or third-party API throttling. Treating a 429 rate limit as a graph logic error leads teams to inflate recursion limits unnecessarily, which leaves the true bottleneck unaddressed.

## Why Parallel Branching and Multi-Agent Fan-Out Exhaust Provider Quotas

In production LangGraph applications, rate limit exhaustion rarely happens during single-node sequential reasoning. It happens when graphs fan out.

LangGraph parallel fan-out patterns multiply token throughput by the branch factor. When a routing node uses the `Send` API or multiple conditional edges to dispatch work across parallel worker nodes, every branch executes within the exact same super-step. If a document router identifies ten distinct sections in an intake file and spawns ten analyzer nodes simultaneously, ten concurrent LLM completion requests hit the model provider in the exact same millisecond window.

The mathematical consequence on token allowances is immediate. Consider an intake node processing a technical contract:

1. Each analyzer worker receives a prompt containing extracted document context.
2. Each worker targets a structured completion extraction.
3. A fan-out across ten workers dispatches ten concurrent requests simultaneously.

If your organization operates on an entry or intermediate provider tier, a single fan-out super-step instantly triggers an HTTP 429 response. Even if your daily quota is generous, short-window burst limits trip before the graph completes its first step.

State reducers present an additional operational hazard during fan-outs. LangGraph allows parallel nodes to return state updates that combine through reducers, such as `Annotated[list, operator.add]`. While reducers prevent concurrent write collisions within the internal state dictionary, they do nothing to pace the outbound network calls that produce those updates. The graph engine eagerly fires every node ready for execution, placing the full burden of traffic shaping on your API client configuration.

## Steps to Configure Node-Level Retries with Exponential Backoff

To build fault-tolerant agent graphs, teams must configure explicit retry policies on nodes that make external network calls. LangGraph provides a declarative `RetryPolicy` primitive that can be attached directly to individual nodes or applied globally across the entire graph.

However, relying on default retry settings will not resolve HTTP 429 errors. LangGraph's default retry policy (`default_retry_on`) automatically retries transient network interruptions and server errors, but for exceptions originating from standard HTTP libraries like `requests` and `httpx`, it only retries on 5xx status codes. Because an HTTP 429 is a 4xx client status code, the default handler allows the exception to bubble up and crash the graph immediately.

To handle rate limit throttling, you must define a custom `retry_on` callable that inspects the exception and explicitly approves 429 errors for retrying with jittered exponential backoff. Following structured configuration steps ensures that each node gracefully waits for quotas to reset instead of failing during traffic spikes.

### Declaring a Jittered Retry Policy for Graph Nodes

The following implementation configures a `RetryPolicy` that intercepts rate limit exceptions from upstream providers, applies exponential backoff with randomized jitter to prevent thundering herd collisions, and attaches to worker nodes:

```python
import httpx
from typing import TypedDict, Annotated
import operator
from langgraph.graph import StateGraph, START, END
from langgraph.types import RetryPolicy, default_retry_on
from langchain_openai import ChatOpenAI

class AgentState(TypedDict):
    query: str
    documents: list[str]
    analyses: Annotated[list[str], operator.add]

def is_retryable_rate_limit(exc: BaseException) -> bool:
    # check for standard HTTP 429 status codes from httpx
    if isinstance(exc, httpx.HTTPStatusError) and exc.response.status_code == 429:
        return True
    # check for provider-specific rate limit exceptions
    exc_type_name = type(exc).__name__
    if "RateLimitError" in exc_type_name:
        return True
    # fall back to LangGraph default retry rules for network and 5xx errors
    return default_retry_on(exc)

# configure exponential backoff with jitter
rate_limit_policy = RetryPolicy(
    max_attempts=5,
    initial_interval=1.0,
    backoff_factor=2.0,
    max_interval=32.0,
    jitter=True,
    retry_on=is_retryable_rate_limit,
)

model = ChatOpenAI(model="gpt-4o-mini", max_retries=0)

def analyze_document(state: AgentState) -> dict:
    doc = state["documents"][0]
    response = model.invoke(f"Analyze the following excerpt: {doc}")
    return {"analyses": [response.content]}

builder = StateGraph(AgentState)
builder.add_node("analyzer", analyze_document, retry_policy=rate_limit_policy)
builder.add_edge(START, "analyzer")
builder.add_edge("analyzer", END)
graph = builder.compile()
```

Notice that `max_retries=0` is passed to the underlying chat model client when delegating retry responsibility to LangGraph. This prevents nested retry loops, where the model client retries five times internally before LangGraph retries the entire node another five times.

### Setting Global Defaults and Handling Retry Exhaustion

Rather than attaching `retry_policy` manually to every node in a complex graph, use `builder.set_node_defaults(retry_policy=rate_limit_policy)` to apply baseline fault tolerance across all nodes simultaneously.

When rate limits persist beyond the maximum attempt threshold, LangGraph allows you to define an `error_handler` callback on the node. The error handler receives a `NodeError` object containing the failure trace, allowing you to update state cleanly, write a partial failure flag, or route execution to a fallback node instead of terminating the overall pipeline.

## How to Throttle Concurrency and Pace Request Frequency

While retry policies repair transient failures after they occur, proactive traffic management prevents rate limit spikes before requests leave your application. High-volume workflows require two complementary throttling controls: concurrency caps on the graph runtime and client-side token bucket rate limiters on model instances.

Relying solely on reactive retries can still lead to cascading failures if every worker in a high-concurrency graph retries at similar intervals. When dozens of parallel branches simultaneously experience throttling, repeated retry storms saturate upstream connection pools and exhaust daily quota budgets. Establishing proactive controls bounds the operational surface of your graph and ensures steady-state throughput.

By combining invocation concurrency limits with in-memory request pacing, you establish predictable throughput that remains safely within provider quotas.

### Restricting Parallel Branching with max_concurrency

LangGraph allows you to pass runtime configuration options directly to `graph.invoke()`, `graph.ainvoke()`, or `graph.stream()`. The `max_concurrency` setting controls the maximum number of worker threads or coroutines that LangGraph executes simultaneously during a super-step.

```python
# enforce a hard ceiling on concurrent node executions
config = {
    "configurable": {"thread_id": "session-101"},
    "max_concurrency": 3,
}

result = graph.invoke(
    {"query": "Review quarterly statements", "documents": document_batch},
    config=config,
)
```

When an intake node uses the `Send` API to spawn ten tasks, a `max_concurrency` value of 3 ensures that LangGraph executes only three nodes concurrently. The remaining tasks wait in the execution queue until active branches finish, flattening a high-token burst into a manageable sequential cadence.

### Implementing Client-Side Token Bucket Rate Limiters

Concurrency limits restrict how many nodes run at once, but they do not account for request frequency over rolling time windows. If three concurrent nodes finish in 500 milliseconds and immediately trigger the next three, you can still breach requests-per-minute (RPM) ceilings.

To control request frequency, attach an `InMemoryRateLimiter` from `langchain_core.rate_limiters` directly to your model definition:

```python
from langchain_core.rate_limiters import InMemoryRateLimiter
from langchain_openai import ChatOpenAI

# pace model calls to 2 requests per second with a burst bucket of 5
rate_limiter = InMemoryRateLimiter(
    requests_per_second=2.0,
    check_every_n_seconds=0.1,
    max_bucket_size=5,
)

throttled_model = ChatOpenAI(
    model="gpt-4o",
    rate_limiter=rate_limiter,
)
```

The rate limiter maintains an in-memory token bucket. Even if multiple asynchronous nodes invoke the model simultaneously, requests block cooperatively until bucket capacity becomes available. This converts bursty fan-out spikes into a smooth, steady stream of requests that avoids triggering an HTTP 429 response. For distributed agent deployments where tasks run across multiple worker hosts, pairing client rate limiters with centralized [storage for agents](/storage-for-agents/) keeps coordination clean and predictable.

## Why Decoupling Document Storage Eliminates Token and Checkpoint Bloat

Many engineering teams spend days tuning backoff intervals and concurrency parameters without realizing that their rate limits are driven by how they store documents in graph state.

In a naive LangGraph implementation, developers place entire document bodies, parsed PDF texts, and raw OCR outputs directly into the graph's `TypedDict` state. This creates two severe architectural bottlenecks:

First, LangGraph checkpointers, such as `PostgresSaver` or `SqliteSaver`, write the entire state dictionary to persistent storage at the end of every super-step. Storing megabytes of raw text across dozens of graph execution steps causes database bloat and serialization latency.

Second, when downstream nodes construct LLM prompts, they repeatedly format these massive text blobs into messages. Passing thousands of redundant tokens into the model across multi-turn agent loops rapidly exhausts your tokens-per-minute quota, triggering persistent 429 rate limit throttling.

Decoupling document state from graph checkpoints cuts token transfer overhead by storing document bodies externally in persistent workspaces and passing only lightweight file identifiers across graph steps.

### The Externalized State Pattern

Instead of serializing raw document payloads into graph memory, production systems store the authoritative files in an external workspace layer. The LangGraph state maintains only lightweight pointers, such as file IDs, storage paths, or extraction record keys:

```python
class ScalableState(TypedDict):
    workspace_id: str
    target_document_ids: list[str]
    extracted_summary_keys: Annotated[list[str], operator.add]
```

When an agent node needs information, it queries the external storage layer for targeted passages rather than dumping the full document corpus into the prompt.

Traditional alternatives for external persistence include local disk storage or raw cloud object stores like Amazon S3. Local filesystems fail when graphs run across distributed worker pods or serverless execution environments. Raw object stores solve persistence but require engineering teams to build custom text extraction, chunking, indexing, and vector retrieval pipelines from scratch.

### Intelligent Workspaces for Agentic Teams

Fast.io provides an intelligent cloud workspace designed specifically for agentic workflows and multi-agent coordination. Instead of managing standalone object storage buckets and disconnected vector databases, teams organize documents in shared org-owned workspaces.

When you enable Intelligence Mode on a Fast.io workspace, uploaded documents are automatically indexed for hybrid search, combining exact full-text matching with semantic retrieval. Agents connecting through the remote [Fast.io MCP server](/storage-for-agents/) at `https://mcp.fast.io/mcp` or the REST API at `https://api.fast.io/current/` can search indexed files and retrieve specific, citation-backed snippets on demand.

For structured workflows, [Metadata Views](/product/document-data-extraction/) turn workspace files into typed, queryable tabular schemas. Rather than having a LangGraph node process an entire invoice or contract to find critical terms, Metadata Views extracts fields such as renewal dates, payment terms, and counterparties directly into structured records that agents query via MCP.

By offloading document indexing and structured extraction to Fast.io workspaces, LangGraph nodes exchange compact metadata references instead of full text corpora. This architectural separation preserves your token quota for genuine reasoning steps, prevents checkpoint serialization bloat, and provides an append-only audit log tracking every document change.

Monthly plans start with a trial of up to 30 days (credit card required); annual plans have no trial. Review the Fast.io [pricing](/pricing/) options across Starter, Business, and Enterprise tiers to select appropriate storage volumes and seat allotments.

## Checklist for Production Hardening: Dead-Letter Queues and Observability

Operating LangGraph in mission-critical production environments requires monitoring and failure recovery strategies that protect upstream pipelines from cascading collapse. In multi-tenant systems or automated data ingestion pipelines, handling rate limits extends beyond simple backoff logic to encompass comprehensive dead-letter routing, telemetry tracing, and model fallback chains.

When rate limits strike despite proactive throttling, well-architected systems degrade gracefully rather than halting completely. Equipping your graph with structured error boundaries ensures that transient rate limits on a single item do not compromise large batch operations. Furthermore, systematic observability allows operations teams to measure real-time throughput margins, identify capacity bottlenecks, and tune concurrency caps before users encounter degraded performance.

### Tracing Token Consumption and Bottlenecks with LangSmith

Diagnosing which nodes cause rate limit spikes requires deep visibility into graph execution. Enabling LangSmith tracing provides granular telemetry on latency, error traces, and exact token consumption for every node attempt:

```bash
export LANGSMITH_TRACING="true"
export LANGSMITH_API_KEY="your-langsmith-key"
export LANGSMITH_PROJECT="langgraph-production"
```

Within the LangSmith trace dashboard, look for execution clusters where multiple model calls overlap on the timeline. If multiple nodes execute concurrently and each consumes thousands of prompt tokens, you have identified the precise fan-out step responsible for tripping upstream TPM thresholds.

### Implementing Dead-Letter Queues in Graph State

In batch document processing pipelines, a rate limit failure on a single document should not abort the processing of other valid files. You can implement dead-letter fault tolerance by capturing exhausted errors into an error partition in state:

```python
class BatchProcessingState(TypedDict):
    active_file_ids: list[str]
    completed_records: Annotated[list[dict], operator.add]
    failed_records: Annotated[list[dict], operator.add]

def resilient_worker(state: BatchProcessingState) -> dict:
    file_id = state["active_file_ids"][0]
    try:
        result = execute_model_extraction(file_id)
        return {"completed_records": [result]}
    except Exception as exc:
        return {
            "failed_records": [{
                "file_id": file_id,
                "error": str(exc),
                "timestamp": "2026-09-24",
            }]
        }
```

By isolating failing items into `failed_records`, downstream aggregator nodes can synthesize all successful outputs while emitting a clean report of items requiring reprocessing.

### Production Reliability Checklist

Before deploying LangGraph workflows to production, verify your configuration against this operational checklist:

* **Configure Jittered Backoff**: Attach an explicit `RetryPolicy` to every node interacting with external APIs, ensuring `retry_on` catches HTTP 429 and rate limit errors.
* **Enforce Invocation Concurrency**: Set `max_concurrency` in graph execution configs to bound the maximum number of simultaneous worker branches.
* **Apply Client-Side Rate Limiters**: Wrap model clients with `InMemoryRateLimiter` to pace request bursts into consistent per-second intervals.
* **Decouple Document Payloads**: Store file corpora in persistent external workspaces like Fast.io, passing compact file IDs in graph state instead of raw text.
* **Disable Redundant Client Retries**: Set `max_retries=0` on underlying chat models when LangGraph's node-level `RetryPolicy` manages backoff.
* **Monitor Telemetry**: Track per-node token spikes in LangSmith to adjust concurrency thresholds before reaching provider limits.

## Frequently asked questions

### How do I handle rate limits in LangGraph?

Handle rate limits in LangGraph by applying a RetryPolicy with exponential backoff and jitter to network-facing nodes, setting max_concurrency on graph invocations to cap parallel branches, and wrapping chat models with an InMemoryRateLimiter. Additionally, externalize document storage so that large file payloads do not inflate token counts across iterative graph cycles.

### What causes 429 errors in LangGraph workflows?

HTTP 429 errors in LangGraph are caused by workflows exceeding model provider requests-per-minute (RPM) or tokens-per-minute (TPM) limits. This frequently occurs during parallel fan-out patterns, map-reduce operations using the Send API, or multi-agent loops where concurrent nodes dispatch simultaneous LLM completion requests.

### How do I limit concurrent node execution in LangGraph?

Limit concurrent node execution in LangGraph by passing max_concurrency in the configuration dictionary during graph invocation, such as graph.invoke(inputs, config={'max_concurrency': 3}). This restricts the number of parallel node tasks that LangGraph executes simultaneously across branching paths.

### What is the difference between recursion_limit and an API rate limit?

The recursion_limit is an internal LangGraph safety setting that halts execution when a graph exceeds a specified number of super-steps (25 by default) to stop infinite loops. An API rate limit is an external threshold enforced by model providers on request frequency or token volume, resulting in an HTTP 429 status code.

### Why does storing raw document text in LangGraph state worsen rate limits?

Storing full document text in LangGraph state causes prompt bloat, as downstream nodes repeatedly format massive strings into model messages. This rapidly drains tokens-per-minute quotas and triggers 429 throttling, while also causing checkpoint database bloat during state serialization.

## Sources

- [Docs by LangChain: How to handle model rate limits](https://docs.langchain.com/langsmith/handle-model-rate-limiting) — LangChain components provide exponential backoff retry methods to handle third-party model provider rate limits.
- [LangGraph Documentation: Fault tolerance](https://docs.langchain.com/oss/python/langgraph/fault-tolerance) — LangGraph's default node retry handler excludes 4xx client errors from automated retries for popular HTTP libraries.

## About Fast.io

Fast.io provides shared workspaces where people and AI agents work on the same files, with built-in semantic search and citation-backed chat over what they hold. Agents reach it through a remote MCP server at https://mcp.fast.io/mcp, a REST API at https://api.fast.io/current/, and a command line client published on npm as @vividengine/fastio-cli.
