# Claude Compaction Failed Unexpectedly: Causes, Fixes, and Context Offloading

Claude Code triggers automated compaction when session context approaches the 200,000-token boundary, summarizing prior conversation turns to preserve working memory. When oversized file reads, sprawling tool outputs, or heavy startup prompts exceed compression thresholds, the CLI halts with unexpected compaction failures. Resolving these errors requires targeted session recovery, scoped manual compaction, and offloading heavy documentation to indexed workspaces via Model Context Protocol.

Source: https://fast.io/resources/claude-compaction-failed-unexpectedly/
Author: [Derek Labian](https://fast.io/authors/derek-labian/)
Last reviewed: 2026-10-08

## What Is Claude Code Compaction and Why Does It Fail?

In Claude chats, Anthropic documents that chat uploads allow up to 20 files at up to 500 MB per file, while project files are capped at 30 MB per file with an unlimited file count bounded only by the context window. In the Claude Code command line interface, conversations do not stop at upload dialogs. Claude Code runs directly against your local workspace, executing terminal commands, evaluating test suites, and reading source files until memory reaches the model's 200,000-token context ceiling.

Claude Code compaction is an automated background process that summarizes prior conversation turns and tool results to free up context window capacity before the next LLM call.

When working on extended coding tasks, an AI assistant accumulates substantial working state. Every prompt sent to Anthropic's API includes system instructions, repository guidelines from `CLAUDE.md`, active Model Context Protocol (MCP) tool schemas, the running conversation history, and the full text of any files read during previous turns. In a fast-moving terminal session, running test runners, inspecting build logs, and checking repository files can consume the vast majority of working memory within an hour of active development.

To keep the session operational without abruptly truncating history, Claude Code triggers an automatic compaction routine when the message buffer approaches the model's context ceiling. During a healthy compaction cycle, Claude Code generates a concise summary of earlier conversational steps, preserves active plans and pending code modifications, discards ephemeral command outputs, and rebuilds the context window.

However, developers frequently encounter the error: `claude compaction failed unexpectedly` (or `automatic compaction failed`). When this occurs, the CLI is unable to compress the conversation history into a workable prompt, rendering the active session unresponsive or repeatedly aborting subsequent requests.

The table below contrasts documented file and context limits across Claude operating environments:

| Claude Environment | Input or Upload Ceiling | Retention and Compaction Behavior | Primary Operational Constraint | Verified Source Date |
|---|---|---|---|---|
| Claude Web and Desktop Chat | 20 files per chat, 500 MB per file | Retains message history until conversational limits prompt a new chat | Hard file count cap and 1,000-page PDF ceiling | 2026-09-14 |
| Claude Projects | 30 MB per file, unlimited file count | Persists files across conversations within the model context window | Total project content bounded by model context capacity | 2026-09-14 |
| Claude Code CLI (Default) | Inlined local files and bash outputs | Auto-compaction triggers as context approaches 200,000 tokens | Runaway tool outputs and thrashing loops halt execution | 2026-10-01 |
| Claude Code CLI (Overridden) | Configured via `CLAUDE_CODE_AUTO_COMPACT_WINDOW` | Compaction threshold customizable between 100,000 and 1,000,000 tokens | Effective window bounded by model context capacity or API gateway limits | 2026-10-01 |

Understanding why compaction breaks requires looking closely at how context accumulates inside the CLI runtime and identifying which operations overwhelm the summarizer.

## The Four Root Causes of Claude Compaction Errors

Compaction failures in Claude Code rarely stem from random infrastructure outages. Instead, they occur when the conversational state violates mathematical or memory assumptions required by the LLM summarization pipeline.

### 1. Oversized Single-Turn Outputs and Monolithic File Reads

The most common trigger for unexpected compaction crashes is a massive tool output returned in a single turn. Claude Code requires conversational headroom to construct and execute a compaction prompt. When an agent executes a terminal command that dumps minified JavaScript bundles, unpaginated SQL database exports, or voluminous build logs into standard output, vast numbers of tokens enter the message buffer instantaneously.

When this single-turn payload approaches or exceeds the remaining context budget, the summarizer cannot construct a valid API request. Claude Code's error reference notes a related edge condition: single-exchange conversations cannot be compacted. If a session consists of one prompt followed by an enormous tool execution, the engine cannot split the history into past turns to summarize and active turns to preserve, triggering an unrecoverable failure.

### 2. Auto-Compaction Thrashing Loops

A second common failure mode is context thrashing. Claude Code documentation documents this behavior under the alert: `Autocompact is thrashing: the context refilled to the limit...`.

Thrashing happens when automatic compaction succeeds in compressing the prior conversation, freeing up a modest slice of context, but the model's immediate next planned action re-reads the exact same large source file or re-runs the exact same verbose build command. Within one turn, the context window fills back to maximum capacity. When this cycle repeats multiple times consecutively, Claude Code deliberately aborts the loop to prevent wasting API credits on an unproductive execution cycle.

```text
Autocompact is thrashing: the context refilled to the limit immediately after compaction.
Claude Code stopped retrying to avoid wasting API calls on a loop that is not making progress.
```

### 3. Startup Overhead Inflation from CLAUDE.md and MCP Schemas

Every turn sent to Anthropic's models includes a base layer of system context that precedes user messages. This base layer contains Anthropic's system prompt, the contents of your local `CLAUDE.md` file, active skill documentation, and the full JSON Schema definitions for every configured MCP tool.

When teams include extensive architecture manuals, code style encyclopedias, and numerous external MCP servers, the base context can consume tens of thousands of tokens on turn one. This static overhead reduces the dynamic space available for code generation. When the total context reaches the auto-compact window, the engine attempts to compress the session, but the static overhead cannot be removed by compaction. As a result, the summarizer has too little room to emit its summary, resulting in an immediate compaction error.

### 4. Memory Cache Corruption and Orphaned Terminal Processes

During aggressive auto-compaction routines, Claude Code interacts with prompt caching mechanisms in Anthropic's API. If network connectivity drops or the API returns an unexpected HTTP 500 error while serializing the conversation tree, the local SQLite state or memory cache can enter an inconsistent state.

In these situations, the CLI may hang indefinitely on the notification `Compacting conversation...`. Terminal keystrokes like `Ctrl+C` may fail to interrupt the process if the event loop is blocked waiting on an unhandled child process, leaving orphaned Node.js processes running in the background while the developer's terminal remains locked.

## How to Recover From Compaction Failures in Claude Code

When Claude Code halts with an unexpected compaction failure or becomes trapped in a thrashing loop, developers need an orderly recovery process that restores terminal responsiveness without discarding uncommitted code modifications.

### Step 1: Step Back from Overloaded Turns

If a session fails immediately after running a command that produced an oversized terminal output, the fastest recovery method is stepping back to the state prior to that command. Press `Esc` twice in rapid succession (or execute the `/rewind` command). This action rolls back the most recent conversational turn and its associated tool output, instantly dropping the token count back below the critical threshold.

Once the oversized payload is purged, you can manually trigger compaction with explicit instructions:

```bash
/compact focus on the architectural plan and current git diff
```

Providing explicit focus arguments directs the summarizer to discard noisy shell outputs while preserving the technical decisions and file modifications needed for your task.

### Step 2: Inspect Context Allocation with the /context Command

To determine whether conversational history or startup overhead caused the failure, run the `/context` command in your terminal. Claude Code renders a breakdown of token consumption across system components:

*   **Messages Row:** Reflects dynamic user prompts, assistant replies, file reads, and bash tool results. If this row represents the vast majority of consumed tokens, conversational bloat is the primary culprit.
*   **CLAUDE.md and System Rules:** Reflects static instructions loaded from disk. When this category consumes substantial context space, static documentation constricts working headroom for active development.
*   **MCP Tools:** Reflects JSON schemas injected by connected MCP servers. Loading multiple expansive servers injects thousands of tokens into every single turn.

If the startup rows consume a large fraction of the context window, trimming those files provides immediate relief.

### Step 3: Scope File Reads to Explicit Line Ranges

If the CLI fails with auto-compaction thrashing, the root cause is almost always reading full source files that span thousands of lines. Rather than allowing Claude Code to read entire files into context, instruct the assistant to read targeted sections:

```bash
claude "read lines 120-220 of src/services/billing-engine.ts"
```

Restricting file reads to targeted line ranges, specific function signatures, or AST exports prevents the message buffer from refilling immediately after a compaction cycle completes.

### Step 4: Clear Session Cache and Kill Orphaned Processes

If Claude Code remains frozen on `Compacting conversation...`, signal interruption might fail. In this scenario, take the following operational steps:

1. Close the affected terminal pane or window.
2. Search for and terminate orphaned Claude CLI processes:
  

```bash
   pkill -f claude
  

```
3. Re-open your terminal in the project directory and resume the session using the resume flag:
  

```bash
   claude --resume
  

```
4. If resuming the session immediately re-triggers the compaction failure, the conversation graph is corrupted. Execute `/clear` to start a fresh session. Because Claude Code writes code directly to files on disk, your source code modifications remain completely intact. You can summarize the pending task in a fresh prompt and resume work with an empty context window.

## Preventing Context Bloat With Indexed Remote Workspaces

While tactical fixes like `/compact` and `/clear` rescue broken sessions, they do not resolve the underlying architectural limitation: inlining large technical files into an LLM's active prompt is inherently fragile.

Engineering teams frequently maintain extensive documentation corpuses, including OpenAPI specifications, database schema dumps, architectural decision records (ADRs), and compliance policies. When developers ask Claude Code to write a new API endpoint, the assistant often reads several hefty documentation files to understand data contracts. Inlining these files floods the context window with immense token payloads before a single line of code is produced, frequently inducing compaction errors and auto-compact thrashing.

### Comparing File Persistence and Retrieval Patterns

Development teams address context limits using several common architectures:

*   **Manual markdown summaries:** Developers write abridged cheat sheets for the agent. While memory-efficient, these summaries require constant manual updates and quickly drift out of date as code evolves.
*   **Amazon S3 and raw object storage:** Teams store raw documentation in cloud buckets. However, object storage lacks built-in semantic search, requiring engineers to build, host, and maintain separate chunking and vector indexing infrastructure.
*   **Consumer cloud storage (Google Drive, Dropbox):** Traditional file sync services synchronize desktop folders, but their APIs enforce strict rate limits when accessed concurrently by autonomous coding tools, and they do not index file contents for semantic RAG queries.

### The Fast.io Solution: Persistent Intelligent Workspaces

Intelligent workspaces eliminate context bloat by serving as an indexed knowledge coordination layer. Instead of forcing Claude Code to ingest entire documentation libraries, organizations store technical specifications in [Fast.io shared workspaces](/product/workspaces/).

Fast.io provides [persistent storage for AI agents](/storage-for-agents/), allowing human developers and autonomous coding assistants to work against a unified, version-controlled source of truth. Teams upload schema files, API contracts, and architecture manuals once. Through [Fast.io Workspace Intelligence](/product/ai/), files are automatically indexed for full-text and semantic search upon arrival, completely removing the need to manage external vector databases or embedding pipelines.

### Connecting Claude Code to Fast.io via Streamable HTTP

Coding assistants connect directly to Fast.io workspaces over Streamable HTTP using the remote Model Context Protocol (MCP) server. Configure the connection directly in your terminal:

```bash
claude mcp add --transport http fast-io https://mcp.fast.io/mcp/code
```

After adding the server, run `/mcp` inside Claude Code to complete the browser-based OAuth authentication flow. For general desktop and web Claude applications, connect to the endpoint `https://mcp.fast.io/mcp/tools`. Detailed setup instructions and protocol documentation are available at [Fast.io MCP Documentation](https://mcp.fast.io/docs).

Using the MCP toolset, Claude Code queries workspace files on demand using search actions. When building a new integration, the assistant searches the workspace for specific schema definitions, retrieves the exact thirty lines required, and injects only those lines into its context window. Local prompt usage drops from hefty file dumps to compact excerpts, eliminating compaction failures at the source.

For engineering teams running multiple autonomous agents, Fast.io workspaces include per-file version history and a detailed activity log, ensuring full visibility into file changes. Advisory file locks (`lock-acquire` and `lock-release` on the `storage_manage` tool) coordinate concurrent access between multiple agents, while ownership transfer capabilities allow autonomous tools to construct workspace assets and hand them off directly to human project leads.

Organizations can test Fast.io on paid monthly subscriptions that begin with a 30-day free trial requiring a credit card. Subscriptions scale across transparent organizational tiers on [Fast.io pricing](/pricing/):

| Plan Tier | Monthly Billing | Workspace Allowance | Included AI Work Credits |
|---|---|---|---|
| Starter | $9.99 per month | 5 workspaces | 100,000 credits |
| Business | $49.99 per month | 50 workspaces | 600,000 credits |
| Enterprise | $199.99 per month | 200 workspaces | 3,000,000 credits |

Additional AI credits can be purchased as team usage expands.

## Configuration Best Practices for Long-Running Agent Sessions

To maintain high velocity during multi-hour coding sessions, developers should implement defensive configuration practices that prevent context exhaustion before compaction errors occur.

### 1. Configure the Auto-Compact Threshold

Claude Code provides fine-grained control over when automatic compaction triggers. If you work through an LLM gateway or proxy with lower context caps than native Anthropic endpoints, you can tune the auto-compact threshold using the `CLAUDE_CODE_AUTO_COMPACT_WINDOW` environment variable:

```bash
export CLAUDE_CODE_AUTO_COMPACT_WINDOW="160000"
```

The environment variable accepts plain integers between 100,000 and 1,000,000 tokens. Setting this value lower than your gateway's hard timeout ensures Claude Code compresses history before external proxies reject your requests. You can also specify this setting at launch using `claude --autocompact 160000`.

### 2. Maintain a Lean CLAUDE.md File

Because `CLAUDE.md` is re-read and injected into every prompt turn, bloating this file directly reduces your operational context. Structure `CLAUDE.md` to include only essential developer instructions:

*   Build, lint, and test commands.
*   Key code formatting and architectural conventions.
*   Critical invariants that must never be broken.

Avoid pasting full API references, changelogs, or code examples into `CLAUDE.md`. Keep reference material in Fast.io workspaces where Claude Code can search it dynamically.

### 3. Curate Active MCP Servers

Each configured MCP server contributes its tool schemas to every single conversation turn. If your configuration includes database managers, deployment tools, browser automation drivers, and cloud consoles simultaneously, tool definitions alone can consume vast amounts of working memory.

Review your `.mcp.json` or user configuration regularly. Keep only the tools required for your active task enabled, and disable dormant servers to preserve working memory.

### 4. Delegate Exploratory Work to Subagents

When investigating bugs that require searching across dozens of source files or parsing extensive log directories, delegate the research to a subagent. Subagents execute in isolated context windows, preventing massive scratchpad exploration from polluting your main session. Once the subagent identifies the relevant file and line numbers, it reports a concise summary back to the parent session, leaving your primary context window clean and free from compaction pressure.

## Frequently asked questions

### Why does Claude Code say compaction failed unexpectedly?

Claude Code displays compaction failed unexpectedly when the background summarization routine cannot compress the conversation history into the remaining context window budget. Common triggers include single-turn tool outputs that exceed token limits, unpaginated file reads, corrupted prompt caches, or having a single-exchange session where past turns cannot be partitioned from active state.

### How do I fix Claude Code stuck on compacting?

If Claude Code is stuck on compacting, first attempt to interrupt the process using Ctrl+C. If the terminal remains unresponsive, terminate the orphaned CLI process using pkill -f claude, reopen your terminal, and resume the session with claude --resume. If the compaction loop recurs immediately, execute /clear to start a fresh session while keeping your edited source files safe on disk.

### What causes autocompact is thrashing errors in Claude Code?

Autocompact thrashing occurs when automatic compaction successfully summarizes earlier messages, but subsequent actions immediately refill the context window to capacity several times consecutively. This typically happens when Claude Code repeatedly reads large source files whole. Claude Code halts execution to avoid burning API calls in an unproductive loop.

### What is the difference between /compact and /clear in Claude Code?

/compact summarizes prior messages and tool results to free context capacity while retaining active plans, recent edits, and project guidelines. In contrast, /clear wipes the entire conversation history, resetting token usage to the baseline startup state. Neither command modifies files saved to disk.

### How do I adjust the auto-compact threshold in Claude Code?

You can adjust the auto-compact threshold by setting the CLAUDE_CODE_AUTO_COMPACT_WINDOW environment variable to a plain integer between 100000 and 1000000 tokens (for example, export CLAUDE_CODE_AUTO_COMPACT_WINDOW="160000"). You can also run /autocompact <tokens> inside an active session or launch the CLI with claude --autocompact <tokens>.

### Can Claude Code access external documentation without inlining files into context?

Yes. Instead of inlining extensive files into local context, teams store documentation in Fast.io intelligent workspaces. Claude Code connects over Streamable HTTP via the remote MCP server for coding agents, as detailed on Fast.io for Agents (/storage-for-agents/). The assistant searches indexed documentation on demand, retrieving only relevant snippets and keeping local context usage minimal.

## Sources

- [Anthropic: Upload files to Claude](https://support.claude.com/en/articles/8241126-upload-files-to-claude): In Claude chats, uploads are capped at 20 files per chat at up to 500 MB per file, while project files are capped at 30 MB per file with an unlimited file count bounded by the context window.
- [Claude Code Docs: Troubleshooting](https://code.claude.com/docs/en/troubleshooting): When automatic compaction succeeds but subsequent tool outputs or file reads immediately refill the context window repeatedly, Claude Code stops retrying to avoid wasting API calls on loops that make no progress.

## About Fast.io

Fast.io provides shared workspaces where people and AI agents work on the same files, with built-in semantic search and citation-backed chat over what they hold. Agents reach it through a remote MCP server, a REST API at https://api.fast.io/current/, and a command line client published on npm as @vividengine/fastio-cli. MCP setup is at https://mcp.fast.io/docs: Claude and most MCP clients connect to https://mcp.fast.io/mcp/tools, ChatGPT to https://mcp.fast.io/mcp/operations, and coding agents to https://mcp.fast.io/mcp/code.
