LangGraph Recursion Limit: GraphRecursionError and State Fixes
A LangGraph GraphRecursionError occurs when a compiled graph exceeds its maximum allowed execution steps before hitting a stop condition. While increasing the recursion_limit in the invoke configuration provides a quick override, permanent resolution requires fixing cyclic routing, tracking iteration counters in state, and offloading repetitive document retrieval to an external indexed workspace.
What Causes GraphRecursionError and Execution Step Failures
When an autonomous agent enters an unconstrained cycle inspecting files or retrying invalid tool arguments, it does not stop on its own until an external guard halts execution. In LangGraph, that safety mechanism is the recursion limit, which terminates execution and raises GraphRecursionError when execution steps exceed the configured threshold before reaching an end state.
The LangGraph recursion limit is a safety mechanism that halts graph execution and raises a GraphRecursionError if an agent workflow exceeds a configured step count (defaulting to 25 steps) before reaching an end state.
Rather than measuring wall-clock latency or model token counts, LangGraph tracks discrete execution steps managed by its Pregel runtime. In this execution model, the runtime evaluates graph state in discrete super-steps. A single step corresponds to executing all active nodes scheduled for that turn, writing their state updates, and evaluating conditional edges to determine the next active nodes. As documented in the LangChain GraphRecursionError guide, reaching this threshold indicates that the graph reached its maximum allowed step count prior to finding a stop condition.
In a standard agent architecture, such as a ReAct loop, a single task iteration consumes two graph steps:
- The Model Step: The agent node invokes the language model with the current conversation history and receives a tool call request.
- The Tool Execution Step: The tool node executes the requested tool and writes the tool output message back to the graph state.
Because every tool interaction consumes two full graph steps, an agent operating under the default recursion limit of 25 steps raises a GraphRecursionError after 12 iterations. When graphs incorporate additional nodes, such as input guardrails, validation checks, or reflection steps, a single logical cycle can consume four or five graph steps. In those architectures, the agent exhausts its 25-step allowance within five iterations.
Consider this minimal reproducible example of an unconstrained loop:
from typing import TypedDict
from langgraph.graph import StateGraph, START, END
class AgentState(TypedDict):
query: str
attempts: int
def analyze_query(state: AgentState) -> AgentState:
print(f"Analyzing query: {state['query']}")
return {"query": state["query"], "attempts": state.get("attempts", 0) + 1}
def evaluate_results(state: AgentState) -> AgentState:
print(f"Evaluating attempt: {state['attempts']}")
return state
builder = StateGraph(AgentState)
builder.add_node("analyze", analyze_query)
builder.add_node("evaluate", evaluate_results)
builder.add_edge(START, "analyze")
builder.add_edge("analyze", "evaluate")
builder.add_edge("evaluate", "analyze") # intentional cyclic edge without exit
graph = builder.compile()
### Executing this graph triggers GraphRecursionError on step 25
try:
graph.invoke({"query": "verify system state", "attempts": 0})
except Exception as err:
print(f"Caught expected error: {type(err).__name__} - {err}")
In this script, the unconditional edge from evaluate back to analyze forms an infinite cycle. Once the Pregel runtime records the 25th step, it raises GraphRecursionError and halts immediately, preventing uncontrolled execution.
Related guides
- How to Connect LangGraph Agent Workflows to OneDrive DocumentsConnecting LangGraph to OneDrive enables cyclic agent workflows and state machines to retrieve, synthesize, and ground...
- How to Manage AI Agent State: Patterns for PersistenceState management is how agents save, retrieve, and sync their work, memory, and files across sessions. Without it,...
- ChatGPT Canvas Limits: File Sizes, Document Length, and WorkaroundsChatGPT Canvas limits documents to practical thresholds of roughly 4,000 lines of code or 25,000 words, running in an...
- ChatGPT Character Limits: The 25,000-Character Paste Limit and SolutionsThe ChatGPT web interface enforces a frontend paste limit that triggers truncation warnings or forces file attachment...
- ChatGPT Message Limits: Quotas, Cooldowns, and Large File WorkaroundsText chat in ChatGPT is unlimited; uploads, images, and voice are capped within rolling windows. When conducting...
- Continue.dev Token Limit: Context Window Configuration and Codebase IndexingThe Continue.dev token limit is the maximum context length configured in Continue's config.json or config.yaml file...
More on this subject: AI Agents: General Guides (99 guides)
How to Increase and Configure the LangGraph Recursion Limit
When a workflow legitimately requires more execution steps to complete a complex plan, you can override the default limit at runtime. You configure the recursion limit by passing a recursion_limit parameter inside the config dictionary when calling .invoke(), .ainvoke(), .stream(), or .astream().
The recursion_limit parameter must sit as a top-level key in the RunnableConfig dictionary. A common implementation bug occurs when developers nest recursion_limit inside the configurable dictionary alongside parameters like thread_id. When placed inside configurable, LangGraph ignores the key and falls back to the default limit of 25 steps.
The following script demonstrates the correct configuration syntax:
from langgraph.graph import StateGraph, START, END
from langgraph.errors import GraphRecursionError
from typing import TypedDict
class PipelineState(TypedDict):
step_count: int
def process_step(state: PipelineState) -> PipelineState:
current = state.get("step_count", 0) + 1
return {"step_count": current}
def check_completion(state: PipelineState) -> str:
if state["step_count"] >= 40:
return END
return "process"
builder = StateGraph(PipelineState)
builder.add_node("process", process_step)
builder.add_edge(START, "process")
builder.add_conditional_edges("process", check_completion, {"process": "process", END: END})
graph = builder.compile()
### Correct configuration: recursion_limit is a top-level config key
runtime_config = {
"recursion_limit": 100,
"configurable": {
"thread_id": "session-prod-001"
}
}
try:
final_state = graph.invoke({"step_count": 0}, config=runtime_config)
print(f"Graph completed successfully after {final_state['step_count']} steps.")
except GraphRecursionError:
print("Execution halted: recursion limit reached before completion.")
Handling GraphRecursionError in Production
Production systems should catch GraphRecursionError explicitly rather than letting uncaught exceptions crash application threads. When caught, your application can inspect the latest checkpoint state, summarize partial results, or alert an operator:
from langgraph.errors import GraphRecursionError
def run_agent_safely(compiled_graph, initial_payload, config):
try:
return compiled_graph.invoke(initial_payload, config=config)
except GraphRecursionError as error:
### Retrieve the latest checkpoint from your checkpointer if configured
checkpoint = compiled_graph.get_state(config)
latest_values = checkpoint.values if checkpoint else {}
return {
"status": "halted_at_limit",
"partial_state": latest_values,
"error_detail": str(error)
}
The Risk of Premature Limit Increases
Increasing recursion_limit from 25 to 100 or 1000 is often treated as a universal quick fix in forum discussions. While appropriate for deep search trees, recursive theorem proving, or complex code generation with extensive test suites, increasing the limit without auditing graph logic introduces severe operational risks:
- Compounded API Expenses: An agent trapped in a repetitive tool loop calling model endpoints at each turn will execute 100 model calls instead of 25 before failing, quadrupling your inference bill.
- Latency Spikes: An unresponsive loop running for 100 turns blocks client connections for several minutes, degrading user experience.
- State Pollution: Looping nodes often append repetitive error messages to state, consuming token context and degrading the quality of subsequent turns.
Four Architecture Steps to Stop Cyclic Graph Loops
Fixing GraphRecursionError permanently requires addressing root causes in your graph state and edge transitions. The four patterns below provide deterministic stop conditions, preventing agents from entering infinite cycles.
1. Explicit Iteration Counters in Graph State
Relying solely on the language model to decide when to finish invites failure when prompts are ambiguous or tool outputs unexpected. Track iteration counts directly in your state schema and enforce a hard boundary in conditional edges:
from typing import TypedDict, Literal
from langgraph.graph import StateGraph, START, END
class GuardedState(TypedDict):
task: str
iteration: int
is_complete: bool
def execute_task(state: GuardedState) -> GuardedState:
current_iter = state.get("iteration", 0) + 1
complete = current_iter >= 3 # simulated completion
return {"iteration": current_iter, "is_complete": complete}
def route_next(state: GuardedState) -> Literal["execute", "fallback_summary", "__end__"]:
if state.get("is_complete", False):
return END
### Hard guard: bail out before reaching the Pregel recursion limit
if state.get("iteration", 0) >= 10:
return "fallback_summary"
return "execute"
def fallback_summary(state: GuardedState) -> GuardedState:
return {"task": "Task reached iteration ceiling; returning partial findings."}
builder = StateGraph(GuardedState)
builder.add_node("execute", execute_task)
builder.add_node("fallback_summary", fallback_summary)
builder.add_edge(START, "execute")
builder.add_conditional_edges(
"execute",
route_next,
{"execute": "execute", "fallback_summary": "fallback_summary", END: END}
)
builder.add_edge("fallback_summary", END)
guarded_graph = builder.compile()
2. Guarding Conditional Edges Against Repeated Tool Invocations
When an agent invokes a tool that fails, language models frequently retry the exact same tool call with identical arguments. If the router sends execution back to the tool node without state changes, the agent repeats the error until hitting the recursion limit.
To prevent this pattern, inspect message history in your router. If the last two tool outputs report identical errors, route the agent to a dedicated diagnosis node or prompt the model with an explicit correction directive rather than executing the failing tool again.
3. Message Trimming and Context Window Stabilization
As execution steps accumulate, message arrays stored in state expand. Models presented with lengthy histories containing earlier failed tool attempts often fixate on previous mistakes, repeating patterns that led to failure.
Use functional message trimming to keep conversation context concise across long-running graphs:
from langchain_core.messages import trim_messages
from langchain_openai import ChatOpenAI
model = ChatOpenAI(model="gpt-4o", temperature=0)
trimmer = trim_messages(
max_tokens=4000,
strategy="last",
token_counter=model,
include_system=True,
allow_partial=False,
start_on="human"
)
def agent_node(state):
trimmed = trimmer.invoke(state["messages"])
response = model.invoke(trimmed)
return {"messages": [response]}
4. Hierarchical Subgraph Isolation
When workflows include multi-step subtasks, such as researching multiple documentation sources or validating code syntax, structure those subtasks as isolated subgraphs.
A child subgraph maintains its own state and its own recursion boundary. If a child subgraph encounters an edge cycle, the parent graph catches the boundary, logs the failure, and moves forward with alternative execution steps rather than failing the entire pipeline.
Stop agent execution loops with indexed workspace search
Decouple multi-file document retrieval from your LangGraph state. Fast.io indexes team files automatically so your agents query context through MCP in a single step instead of cycling through file reads. Every organization starts with a 14-day free trial, which requires a credit card.
Why Multi-File Retrieval Causes Graph Step Loops
One of the most frequent triggers of GraphRecursionError in coding and document agents is iterative file exploration. When an agent attempts to answer questions across a large code repository or document archive by cycling through local file tools (such as list_files, read_file, and grep_search), each file inspection consumes two full graph steps.
Scanning a repository containing dozens of files will trigger a GraphRecursionError within 10 iterations, long before the agent locates relevant content. Bumping recursion_limit to 200 does not solve the underlying problem; it merely lets the agent wander through local folders longer while burning tokens.
The Tradeoff Between Local Disk Traversal and Indexed Workspaces
Teams typically choose between three retrieval architectures when designing document agents:
- Local Disk Traversal in Graph Loops: The agent issues sequential tool calls to list directories and read individual files. This approach causes rapid step exhaustion, high token consumption, and fragility when repository structures change.
- Custom Local Vector Stores: The developer scripts chunking, embeddings, and vector database synchronization locally. While this reduces step counts, it requires ongoing maintenance, dedicated infrastructure, and complex metadata synchronization.
- Remote Indexed Workspaces: Files reside in an organization-managed cloud workspace that handles document indexing and hybrid retrieval automatically. The agent queries the workspace over standard protocols, collapsing complex multi-step traversals into single-turn answers.
The Fast.io Remote MCP Workspace Architecture
Fast.io provides shared workspaces designed specifically for agentic teams and multi-agent coordination. Instead of forcing a LangGraph agent to navigate directory structures node by node, team files are stored in Fast.io workspaces with Intelligence Mode enabled.
When Intelligence Mode is active, files (including PDFs, Word documents, spreadsheets, markdown files, and code) are automatically indexed for hybrid search combining full-text keyword matching and semantic vector retrieval.
Fast.io exposes this indexing layer through its remote MCP server at https://mcp.fast.io/mcp. You can review agent onboarding protocols at Fast.io agent onboarding and explore setup details on the Fast.io workspace storage for agents page. Rather than executing ten consecutive file inspection steps inside LangGraph, the agent calls the Fast.io MCP endpoint once. The workspace returns relevant excerpts and citations directly into the model context.
Here is how to connect a LangGraph agent to a Fast.io MCP workspace:
import asyncio
from langchain_mcp_adapters.client import MultiServerMCPClient
from langchain_openai import ChatOpenAI
from langgraph.prebuilt import create_react_agent
async def run_workspace_agent():
model = ChatOpenAI(model="gpt-4o", temperature=0)
### Connect to the remote Fast.io MCP endpoint
connections = {
"fastio": {
"transport": "http",
"url": "https://mcp.fast.io/mcp/key",
"headers": {
"Authorization": "Bearer YOUR_FASTIO_API_KEY"
}
}
}
client = MultiServerMCPClient(connections)
tools = await client.get_tools()
### Create the agent with Fast.io search tools
agent = create_react_agent(model, tools)
### A single query retrieves indexed content without iterative disk loops
query = "Find the deployment policy for staging environments and summarize the rollback steps."
result = await agent.ainvoke(
{"messages": [("user", query)]},
config={"recursion_limit": 25}
)
print(result["messages"][-1].content)
if __name__ == "__main__":
asyncio.run(run_workspace_agent())
Keeping State Clean and Auditable
Decoupling document retrieval to an indexed Fast.io workspace provides several operational advantages:
- Step Reduction: The agent resolves document queries in a single model-tool step pair rather than cycling through multiple file-reading iterations.
- Per-File Version History: Every document in Fast.io maintains complete revision history. If an agent writes or updates a file in the workspace, prior versions remain intact and can be restored.
- Append-Only Audit Log: Every file read, write, share, and metadata query is recorded in the workspace audit log, providing complete traceability for human operators.
- Metadata Views: Teams can define structured schemas across workspace documents without writing custom parsing scripts. Agents query these views directly via MCP.
Explore how engineering teams configure these environments on the Fast.io workspace storage for agents page, or examine subscription options on the Fast.io pricing page.
Production Diagnostics and Step Tracing Checklist
Diagnosing the root cause of a GraphRecursionError requires systematic inspection of graph state transitions. Use this diagnostic checklist when troubleshooting graphs that fail in production:
Diagnostic Checklist
- Verify Config Parameter Placement: Confirm that
recursion_limitis placed at the top level of theRunnableConfigdictionary and not nested insideconfigurable. - Inspect Checkpoint State History: If your graph compiles with a checkpointer (such as MemorySaver or a database checkpointer), inspect the state history using
graph.get_state_history(config). Identify the exact sequence of nodes that repeated prior to failure. - Audit Conditional Edge Routers: Verify that conditional edge functions evaluate all possible model responses, including tool calls, empty responses, and unexpected strings. Ensure every branch path has a guaranteed route to
END. - Enforce Iteration Counters: Add an explicit integer counter in state that increments at every node execution. If the counter passes a predefined threshold, route to a graceful fallback node instead of allowing the runtime to crash.
- Inspect Tool Error Payloads: Ensure tools return actionable, structured error messages. When tools return generic failure messages, models tend to retry identical calls in an infinite loop.
- Decouple Document and Repository Search: Move large file collections out of prompt memory and local filesystem iteration into an indexed workspace accessible via MCP.
Monitoring Step Progression in Real Time
In production deployments, you can monitor graph step progression in real time using LangGraph event streaming. Inspecting the metadata dictionary reveals the current step index:
async def stream_and_monitor(graph, inputs, config):
async for event in graph.astream_events(inputs, config=config, version="v2"):
kind = event.get("event")
metadata = event.get("metadata", {})
current_step = metadata.get("langgraph_step", None)
if kind == "on_chain_start" and current_step is not None:
node_name = event.get("name", "unknown")
print(f"Step: {current_step} - Entering node: {node_name}")
### Proactive alerting if execution approaches the limit
configured_limit = config.get("recursion_limit", 25)
if current_step >= configured_limit - 3:
print(f"WARNING: Graph step {current_step} is near limit {configured_limit}!")
Summary of Symptoms and Architectural Remediations
By combining explicit state guards, message pruning, and indexed workspace retrieval, developers build resilient agent graphs that complete complex tasks without hitting execution limits.
Sources
References used to verify factual claims in this guide.
-
LangGraph StateGraph execution terminates and raises an error when reaching the configured step limit before reaching a stop condition.
Frequently Asked Questions
What is the default recursion limit in LangGraph?
LangGraph enforces a default recursion limit of 25 execution steps per invoke or stream call in its Pregel runtime. A single step represents one super-step where all scheduled nodes execute. In standard tool-calling agent loops where each iteration requires a model step and a tool execution step, the default limit is reached after 12 iterations.
How do I fix GraphRecursionError in LangGraph?
To fix GraphRecursionError, inspect your graph for infinite cycles in conditional edges, add an explicit iteration counter to state with a hard stop condition, and ensure tools return descriptive error messages so the model does not repeat identical failed calls. For document-heavy agents, offload file exploration to an indexed MCP workspace instead of iterating through files in graph loops.
How do you increase the recursion limit in LangGraph?
You increase the recursion limit by passing recursion_limit as a top-level key in the config dictionary when invoking your graph, such as graph.invoke(inputs, config={'recursion_limit': 100}). Ensure recursion_limit is placed at the root of the configuration dictionary rather than inside the nested configurable object.
What is the difference between recursion_limit and max_iterations in LangGraph?
The recursion_limit parameter is a LangGraph runtime guard that counts graph super-steps across all nodes in a compiled graph. In contrast, max_iterations is typically a property of individual agent executors or custom loops that counts higher-level cognitive turns. Because one cognitive turn often requires multiple graph steps, a graph reaches its recursion_limit much sooner than an equivalent iteration count.
Does increasing the recursion limit increase API token costs?
Increasing the recursion limit does not inherently increase costs if the graph reaches an end state efficiently. However, if the limit was reached due to an infinite prompt loop or repeating tool failure, increasing the limit permits the graph to continue calling model endpoints repeatedly, multiplying inference costs before failing.
Related Resources
Stop agent execution loops with indexed workspace search
Decouple multi-file document retrieval from your LangGraph state. Fast.io indexes team files automatically so your agents query context through MCP in a single step instead of cycling through file reads. Every organization starts with a 14-day free trial, which requires a credit card.