Claude Compaction Failed Unexpectedly: Causes, Fixes, and Context Offloading
Claude Code triggers automated compaction when session context approaches the 200,000-token boundary, summarizing prior conversation turns to preserve working memory. When oversized file reads, sprawling tool outputs, or heavy startup prompts exceed compression thresholds, the CLI halts with unexpected compaction failures. Resolving these errors requires targeted session recovery, scoped manual compaction, and offloading heavy documentation to indexed workspaces via Model Context Protocol.
What Is Claude Code Compaction and Why Does It Fail?
In Claude chats, Anthropic documents that chat uploads allow up to 20 files at up to 500 MB per file, while project files are capped at 30 MB per file with an unlimited file count bounded only by the context window. In the Claude Code command line interface, conversations do not stop at upload dialogs. Claude Code runs directly against your local workspace, executing terminal commands, evaluating test suites, and reading source files until memory reaches the model's 200,000-token context ceiling.
Claude Code compaction is an automated background process that summarizes prior conversation turns and tool results to free up context window capacity before the next LLM call.
When working on extended coding tasks, an AI assistant accumulates substantial working state. Every prompt sent to Anthropic's API includes system instructions, repository guidelines from CLAUDE.md, active Model Context Protocol (MCP) tool schemas, the running conversation history, and the full text of any files read during previous turns. In a fast-moving terminal session, running test runners, inspecting build logs, and checking repository files can consume the vast majority of working memory within an hour of active development.
To keep the session operational without abruptly truncating history, Claude Code triggers an automatic compaction routine when the message buffer approaches the model's context ceiling. During a healthy compaction cycle, Claude Code generates a concise summary of earlier conversational steps, preserves active plans and pending code modifications, discards ephemeral command outputs, and rebuilds the context window.
However, developers frequently encounter the error: claude compaction failed unexpectedly (or automatic compaction failed). When this occurs, the CLI is unable to compress the conversation history into a workable prompt, rendering the active session unresponsive or repeatedly aborting subsequent requests.
The table below contrasts documented file and context limits across Claude operating environments:
Understanding why compaction breaks requires looking closely at how context accumulates inside the CLI runtime and identifying which operations overwhelm the summarizer.
Related guides
- Claude Code Session Limit Reached: Causes, Compaction, and WorkaroundsClaude Code session limits occur when cumulative terminal output, command history, and file reads saturate the agent...
- Claude Code Message Limit: Quotas, Compaction, and Large-Context WorkaroundsClaude Code message limits define the maximum prompt token size and message frequency allowed in a CLI session before...
- Claude 3.5 Sonnet Context Window: 200,000 Token Limit and Output BudgetsThe Claude 3.5 Sonnet context window is 200,000 input tokens with a maximum output limit of 8,192 tokens per request....
- Claude 3.7 Sonnet Context Window: 200K Tokens, Extended Thinking, and PricingAnthropic's Claude 3.7 Sonnet pairs a 200,000-token input context window with a dynamic thinking budget capable of...
- Claude Desktop Context Window: Token Limits, MCP Overhead, and Document RetrievalClaude Desktop operates with a standard 200,000-token input context window (with options up to 1,000,000 tokens on...
- Claude Haiku Context Window: Token Limits, Latency, and WorkaroundsThe Claude Haiku context window provides a 200,000-token input memory buffer for high-speed processing across...
More on this subject: Claude and Claude Code (249 guides)
The Four Root Causes of Claude Compaction Errors
Compaction failures in Claude Code rarely stem from random infrastructure outages. Instead, they occur when the conversational state violates mathematical or memory assumptions required by the LLM summarization pipeline.
1. Oversized Single-Turn Outputs and Monolithic File Reads
The most common trigger for unexpected compaction crashes is a massive tool output returned in a single turn. Claude Code requires conversational headroom to construct and execute a compaction prompt. When an agent executes a terminal command that dumps minified JavaScript bundles, unpaginated SQL database exports, or voluminous build logs into standard output, vast numbers of tokens enter the message buffer instantaneously.
When this single-turn payload approaches or exceeds the remaining context budget, the summarizer cannot construct a valid API request. Claude Code's error reference notes a related edge condition: single-exchange conversations cannot be compacted. If a session consists of one prompt followed by an enormous tool execution, the engine cannot split the history into past turns to summarize and active turns to preserve, triggering an unrecoverable failure.
2. Auto-Compaction Thrashing Loops
A second common failure mode is context thrashing. Claude Code documentation documents this behavior under the alert: Autocompact is thrashing: the context refilled to the limit....
Thrashing happens when automatic compaction succeeds in compressing the prior conversation, freeing up a modest slice of context, but the model's immediate next planned action re-reads the exact same large source file or re-runs the exact same verbose build command. Within one turn, the context window fills back to maximum capacity. When this cycle repeats multiple times consecutively, Claude Code deliberately aborts the loop to prevent wasting API credits on an unproductive execution cycle.
Autocompact is thrashing: the context refilled to the limit immediately after compaction.
Claude Code stopped retrying to avoid wasting API calls on a loop that is not making progress.
3. Startup Overhead Inflation from CLAUDE.md and MCP Schemas
Every turn sent to Anthropic's models includes a base layer of system context that precedes user messages. This base layer contains Anthropic's system prompt, the contents of your local CLAUDE.md file, active skill documentation, and the full JSON Schema definitions for every configured MCP tool.
When teams include extensive architecture manuals, code style encyclopedias, and numerous external MCP servers, the base context can consume tens of thousands of tokens on turn one. This static overhead reduces the dynamic space available for code generation. When the total context reaches the auto-compact window, the engine attempts to compress the session, but the static overhead cannot be removed by compaction. As a result, the summarizer has too little room to emit its summary, resulting in an immediate compaction error.
4. Memory Cache Corruption and Orphaned Terminal Processes
During aggressive auto-compaction routines, Claude Code interacts with prompt caching mechanisms in Anthropic's API. If network connectivity drops or the API returns an unexpected HTTP 500 error while serializing the conversation tree, the local SQLite state or memory cache can enter an inconsistent state.
In these situations, the CLI may hang indefinitely on the notification Compacting conversation.... Terminal keystrokes like Ctrl+C may fail to interrupt the process if the event loop is blocked waiting on an unhandled child process, leaving orphaned Node.js processes running in the background while the developer's terminal remains locked.
How to Recover From Compaction Failures in Claude Code
When Claude Code halts with an unexpected compaction failure or becomes trapped in a thrashing loop, developers need an orderly recovery process that restores terminal responsiveness without discarding uncommitted code modifications.
Step 1: Step Back from Overloaded Turns
If a session fails immediately after running a command that produced an oversized terminal output, the fastest recovery method is stepping back to the state prior to that command. Press Esc twice in rapid succession (or execute the /rewind command). This action rolls back the most recent conversational turn and its associated tool output, instantly dropping the token count back below the critical threshold.
Once the oversized payload is purged, you can manually trigger compaction with explicit instructions:
/compact focus on the architectural plan and current git diff
Providing explicit focus arguments directs the summarizer to discard noisy shell outputs while preserving the technical decisions and file modifications needed for your task.
Step 2: Inspect Context Allocation with the /context Command
To determine whether conversational history or startup overhead caused the failure, run the /context command in your terminal. Claude Code renders a breakdown of token consumption across system components:
- Messages Row: Reflects dynamic user prompts, assistant replies, file reads, and bash tool results. If this row represents the vast majority of consumed tokens, conversational bloat is the primary culprit.
- CLAUDE.md and System Rules: Reflects static instructions loaded from disk. When this category consumes substantial context space, static documentation constricts working headroom for active development.
- MCP Tools: Reflects JSON schemas injected by connected MCP servers. Loading multiple expansive servers injects thousands of tokens into every single turn.
If the startup rows consume a large fraction of the context window, trimming those files provides immediate relief.
Step 3: Scope File Reads to Explicit Line Ranges
If the CLI fails with auto-compaction thrashing, the root cause is almost always reading full source files that span thousands of lines. Rather than allowing Claude Code to read entire files into context, instruct the assistant to read targeted sections:
claude "read lines 120-220 of src/services/billing-engine.ts"
Restricting file reads to targeted line ranges, specific function signatures, or AST exports prevents the message buffer from refilling immediately after a compaction cycle completes.
Step 4: Clear Session Cache and Kill Orphaned Processes
If Claude Code remains frozen on Compacting conversation..., signal interruption might fail. In this scenario, take the following operational steps:
- Close the affected terminal pane or window.
- Search for and terminate orphaned Claude CLI processes:
pkill -f claude
- Re-open your terminal in the project directory and resume the session using the resume flag:
claude --resume
- If resuming the session immediately re-triggers the compaction failure, the conversation graph is corrupted. Execute
/clearto start a fresh session. Because Claude Code writes code directly to files on disk, your source code modifications remain completely intact. You can summarize the pending task in a fresh prompt and resume work with an empty context window.
Stop Claude Code context crashes with indexed workspace storage
Keep documentation, architecture schemas, and API specs in an intelligent workspace. Let Claude Code query indexed files via MCP without overloading session context. Starts with a 30-day free trial requiring a credit card.
Preventing Context Bloat With Indexed Remote Workspaces
While tactical fixes like /compact and /clear rescue broken sessions, they do not resolve the underlying architectural limitation: inlining large technical files into an LLM's active prompt is inherently fragile.
Engineering teams frequently maintain extensive documentation corpuses, including OpenAPI specifications, database schema dumps, architectural decision records (ADRs), and compliance policies. When developers ask Claude Code to write a new API endpoint, the assistant often reads several hefty documentation files to understand data contracts. Inlining these files floods the context window with immense token payloads before a single line of code is produced, frequently inducing compaction errors and auto-compact thrashing.
Comparing File Persistence and Retrieval Patterns
Development teams address context limits using several common architectures:
- Manual markdown summaries: Developers write abridged cheat sheets for the agent. While memory-efficient, these summaries require constant manual updates and quickly drift out of date as code evolves.
- Amazon S3 and raw object storage: Teams store raw documentation in cloud buckets. However, object storage lacks built-in semantic search, requiring engineers to build, host, and maintain separate chunking and vector indexing infrastructure.
- Consumer cloud storage (Google Drive, Dropbox): Traditional file sync services synchronize desktop folders, but their APIs enforce strict rate limits when accessed concurrently by autonomous coding tools, and they do not index file contents for semantic RAG queries.
The Fast.io Solution: Persistent Intelligent Workspaces
Intelligent workspaces eliminate context bloat by serving as an indexed knowledge coordination layer. Instead of forcing Claude Code to ingest entire documentation libraries, organizations store technical specifications in Fast.io shared workspaces.
Fast.io provides persistent storage for AI agents, allowing human developers and autonomous coding assistants to work against a unified, version-controlled source of truth. Teams upload schema files, API contracts, and architecture manuals once. Through Fast.io Workspace Intelligence, files are automatically indexed for full-text and semantic search upon arrival, completely removing the need to manage external vector databases or embedding pipelines.
Connecting Claude Code to Fast.io via Streamable HTTP
Coding assistants connect directly to Fast.io workspaces over Streamable HTTP using the remote Model Context Protocol (MCP) server. Configure the connection directly in your terminal:
claude mcp add --transport http fast-io https://mcp.fast.io/mcp/code
After adding the server, run /mcp inside Claude Code to complete the browser-based OAuth authentication flow. For general desktop and web Claude applications, connect to the endpoint https://mcp.fast.io/mcp/tools. Detailed setup instructions and protocol documentation are available at Fast.io MCP Documentation.
Using the MCP toolset, Claude Code queries workspace files on demand using search actions. When building a new integration, the assistant searches the workspace for specific schema definitions, retrieves the exact thirty lines required, and injects only those lines into its context window. Local prompt usage drops from hefty file dumps to compact excerpts, eliminating compaction failures at the source.
For engineering teams running multiple autonomous agents, Fast.io workspaces include per-file version history and a detailed activity log, ensuring full visibility into file changes. Advisory file locks (lock-acquire and lock-release on the storage_manage tool) coordinate concurrent access between multiple agents, while ownership transfer capabilities allow autonomous tools to construct workspace assets and hand them off directly to human project leads.
Organizations can test Fast.io on paid monthly subscriptions that begin with a 30-day free trial requiring a credit card. Subscriptions scale across transparent organizational tiers on Fast.io pricing:
Additional AI credits can be purchased as team usage expands.
Configuration Best Practices for Long-Running Agent Sessions
To maintain high velocity during multi-hour coding sessions, developers should implement defensive configuration practices that prevent context exhaustion before compaction errors occur.
1. Configure the Auto-Compact Threshold
Claude Code provides fine-grained control over when automatic compaction triggers. If you work through an LLM gateway or proxy with lower context caps than native Anthropic endpoints, you can tune the auto-compact threshold using the CLAUDE_CODE_AUTO_COMPACT_WINDOW environment variable:
export CLAUDE_CODE_AUTO_COMPACT_WINDOW="160000"
The environment variable accepts plain integers between 100,000 and 1,000,000 tokens. Setting this value lower than your gateway's hard timeout ensures Claude Code compresses history before external proxies reject your requests. You can also specify this setting at launch using claude --autocompact 160000.
2. Maintain a Lean CLAUDE.md File
Because CLAUDE.md is re-read and injected into every prompt turn, bloating this file directly reduces your operational context. Structure CLAUDE.md to include only essential developer instructions:
- Build, lint, and test commands.
- Key code formatting and architectural conventions.
- Critical invariants that must never be broken.
Avoid pasting full API references, changelogs, or code examples into CLAUDE.md. Keep reference material in Fast.io workspaces where Claude Code can search it dynamically.
3. Curate Active MCP Servers
Each configured MCP server contributes its tool schemas to every single conversation turn. If your configuration includes database managers, deployment tools, browser automation drivers, and cloud consoles simultaneously, tool definitions alone can consume vast amounts of working memory.
Review your .mcp.json or user configuration regularly. Keep only the tools required for your active task enabled, and disable dormant servers to preserve working memory.
4. Delegate Exploratory Work to Subagents
When investigating bugs that require searching across dozens of source files or parsing extensive log directories, delegate the research to a subagent. Subagents execute in isolated context windows, preventing massive scratchpad exploration from polluting your main session. Once the subagent identifies the relevant file and line numbers, it reports a concise summary back to the parent session, leaving your primary context window clean and free from compaction pressure.
Sources
References used to verify factual claims in this guide.
-
In Claude chats, uploads are capped at 20 files per chat at up to 500 MB per file, while project files are capped at 30 MB per file with an unlimited file count bounded by the context window.
-
When automatic compaction succeeds but subsequent tool outputs or file reads immediately refill the context window repeatedly, Claude Code stops retrying to avoid wasting API calls on loops that make no progress.
Frequently Asked Questions
Why does Claude Code say compaction failed unexpectedly?
Claude Code displays compaction failed unexpectedly when the background summarization routine cannot compress the conversation history into the remaining context window budget. Common triggers include single-turn tool outputs that exceed token limits, unpaginated file reads, corrupted prompt caches, or having a single-exchange session where past turns cannot be partitioned from active state.
How do I fix Claude Code stuck on compacting?
If Claude Code is stuck on compacting, first attempt to interrupt the process using Ctrl+C. If the terminal remains unresponsive, terminate the orphaned CLI process using pkill -f claude, reopen your terminal, and resume the session with claude --resume. If the compaction loop recurs immediately, execute /clear to start a fresh session while keeping your edited source files safe on disk.
What causes autocompact is thrashing errors in Claude Code?
Autocompact thrashing occurs when automatic compaction successfully summarizes earlier messages, but subsequent actions immediately refill the context window to capacity several times consecutively. This typically happens when Claude Code repeatedly reads large source files whole. Claude Code halts execution to avoid burning API calls in an unproductive loop.
What is the difference between /compact and /clear in Claude Code?
/compact summarizes prior messages and tool results to free context capacity while retaining active plans, recent edits, and project guidelines. In contrast, /clear wipes the entire conversation history, resetting token usage to the baseline startup state. Neither command modifies files saved to disk.
How do I adjust the auto-compact threshold in Claude Code?
You can adjust the auto-compact threshold by setting the CLAUDE_CODE_AUTO_COMPACT_WINDOW environment variable to a plain integer between 100000 and 1000000 tokens (for example, export CLAUDE_CODE_AUTO_COMPACT_WINDOW="160000"). You can also run /autocompact <tokens> inside an active session or launch the CLI with claude --autocompact <tokens>.
Can Claude Code access external documentation without inlining files into context?
Yes. Instead of inlining extensive files into local context, teams store documentation in Fast.io intelligent workspaces. Claude Code connects over Streamable HTTP via the remote MCP server for coding agents, as detailed on Fast.io for Agents (/storage-for-agents/). The assistant searches indexed documentation on demand, retrieving only relevant snippets and keeping local context usage minimal.
Related Resources
Stop Claude Code context crashes with indexed workspace storage
Keep documentation, architecture schemas, and API specs in an intelligent workspace. Let Claude Code query indexed files via MCP without overloading session context. Starts with a 30-day free trial requiring a credit card.