# Claude Code Degradation: Why Reasoning Declines in Long Sessions and How to Prevent It

Claude Code degradation occurs when extended agent sessions accumulate excessive tool outputs, bash logs, and lossy compactions, causing instruction drift and file state errors. While Claude supports large context capacities, attention dilution erodes reasoning precision long before reaching token limits. Understanding these context mechanics and structuring sessions with modular persistence keeps agent execution reliable.

Source: https://fast.io/resources/claude-code-degradation/
Author: [Tom Langridge](https://fast.io/authors/tom-langridge/)
Last reviewed: 2026-10-08

## Why Claude Code Reasoning Degrades in Long Sessions

Anthropic's documented Claude limits specify that a standard Claude chat accepts up to 20 files at up to 500MB each, while Claude Projects accepts files up to 30MB each with unlimited file count provided the total content fits within Claude's context window, as detailed in [Anthropic's documented Claude limits](https://support.claude.com/en/articles/8241126-upload-files-to-claude). For Claude Code, the command-line coding agent, the runtime operates against a 200,000 token context window. In extended terminal sessions, software engineers frequently watch code quality deteriorate long before reaching that token ceiling.

Claude Code degradation describes the progressive loss of reasoning precision, adherence to system instructions, and file state awareness that occurs as an agent session accumulates excessive context and repeated compactions.

At the start of a terminal session, Claude Code reasons with precision. It inspects repository structures, adheres to style rules, and produces clean git diffs. As the session accumulates bash tool outputs, compiler logs, and full-file inspections, that reliability breaks down. The agent repeats failed shell commands, forgets negative constraints, misidentifies line numbers in active buffers, and hallucinates edits. Developers troubleshooting this behavior often suspect a claude code memory leak in the CLI process. In reality, the issue is token accumulation inside the context window rather than an operating system process memory leak.

### The Three Primary Degradation Mechanisms

This operational decline is driven by three architectural mechanisms:

1. **Context Window Saturation:** As conversational history, loaded files, and shell outputs accumulate, the active token count approaches maximum capacity. This leaves reduced token headroom for multi-step reasoning, truncating complex logic.
2. **Attention Dilution from Tool-Call Outputs:** Modern transformer architectures distribute attention weights across every token in the active prompt prefix. When a session accumulates thousands of tokens of compiler outputs, test traces, and whole-file reads, mathematical attention allocated to core system constraints thins out. Irrelevant terminal outputs crowd out critical instructions.
3. **Compaction Lossiness:** To prevent out-of-memory failures, Claude Code periodically summarizes conversational history using the `/compact` command or automated compaction routines. While primary objectives survive, subtle debugging context, negative constraints, and exact file state representations are permanently discarded.

Understanding how to prevent claude code performance degradation requires examining how the agent structures its context window on every turn.

## Inside the Agentic Context Stack: How Tokens Accumulate

Large language models retain no working memory between API calls. Each time Claude Code executes a turn, it constructs a complete prompt payload sent to the Anthropic API. The agent recreates session state on every interaction by stacking multiple layers into an ordered prompt prefix.

### The Internal Context Window Stack

The Claude Code context window is organized into distinct functional layers:

* **System Prompt (~4,200 tokens):** Core behavioral instructions defined by the agent runtime, establishing tool schemas, execution rules, and formatting.
* **Auto Memory (`MEMORY.md`):** Persistent memory across sessions where Claude records build commands, codebase patterns, and user preferences.
* **Environment Information:** Shell metadata including working directory, operating system, default shell, git branch, and commit hash.
* **MCP Tool Definitions:** Schemas for external tools. Claude Code defers tool schemas and uses tool search to retrieve specific definitions on demand.
* **Skill Index:** Descriptions of loaded skills and slash commands, loaded into active context only when invoked.
* **Project Configuration (`CLAUDE.md`):** Global rules from `~/.claude/CLAUDE.md` and repository guidelines from the root `CLAUDE.md`.
* **Path-Scoped Rules:** Targeted rule files in `.claude/rules/*.md` loaded dynamically when matching paths are accessed.
* **Conversational Turns:** Alternating user prompts, assistant reasoning blocks, and tool invocation tags.
* **Raw Tool Results:** Direct output returned by tools, including whole-file contents, edit diffs, and bash stdout and stderr.

### Why Attention Dilution Degrades Reasoning

Attention dilution is a structural property of self-attention in transformer models. When Claude processes a turn, every token attends to every other token. In a fresh session with 8,000 tokens of system instructions and conventions, a rule such as "never edit generated database migration files" commands a prominent share of attention weights.

```text
Fresh Session (10,000 tokens):
Core System Prompt and CLAUDE.md occupy active memory.
User instructions command primary attention.

Saturated Session (160,000 tokens):
Terminal logs, test traces, and whole-file reads dominate active memory.
Core instructions represent a tiny fraction of the prompt prefix.
```

As the developer runs tests, inspects dependencies, and reads source files, thousands of lines of output enter context. When active tokens reach 150,000, original rules represent a tiny fraction of the prompt prefix. The mathematical attention available for instruction adherence is diluted across verbose test logs and build artifacts. This causes claude context degradation, where the agent begins violating rules established at the start of the conversation.

| Context Layer | Typical Token Footprint | Invalidation Trigger | Degradation Impact |
| :--- | :--- | :--- | :--- |
| System Prompt | 4,000 to 4,500 tokens | Tool set modifications | Minimal; static and cached |
| Project `CLAUDE.md` | 1,000 to 3,000 tokens | Session start or `/compact` | High if bloated; keep concise |
| Path-Scoped Rules | 300 to 1,500 tokens per rule | File path match | Low; loads only when relevant |
| Bash Command Logs | 1,000 to 25,000 tokens per call | Executed commands | Severe; primary driver of dilution |
| Whole-File Reads | 2,000 to 40,000 tokens per file | File viewing tool | Severe; crowds out working memory |
| Conversational History | 5,000 to 50,000 tokens | Each conversational turn | Moderate; accumulates progressively |

## How Compaction Lossiness and Thrashing Break File Awareness

When conversational context approaches the operational limit of the context window, Claude Code intervenes to prevent an out-of-memory failure. It does this through compaction, either triggered manually via the `/compact` command or automatically by the runtime.

### How Compaction Operates Internally

Compaction executes a structured multi-phase reduction:

1. **Tool Output Pruning:** Older tool results, particularly lengthy stdout and stderr streams from bash commands, are stripped or truncated first.
2. **Context Summarization:** Claude Code issues an out-of-band request containing the conversation history. The model produces a structured markdown summary capturing core goals, decisions, modified files, and outstanding tasks.
3. **State Reassembly:** Claude Code clears the message history and reconstructs a new conversation prefix. It reloads the system prompt, `CLAUDE.md`, and `MEMORY.md`, injecting the summary as baseline context.

### What Survives Compaction Versus What Is Lost

Compaction allows the session to proceed, but it alters the agent's internal knowledge state:

* **What Survives:** High-level project objectives, major architectural choices cited in the summary, active file paths, and general task status.
* **What Is Lost:** Fine-grained negative constraints ("do not modify helper functions in auth.ts"), intermediate hypotheses evaluated and rejected, subtle line-by-line syntax choices, and exact syntax tree representations of files read prior to compaction.

This information loss causes claude code hallucinations long sessions. After compaction, Claude remembers that it modified `auth.ts`, but it no longer holds the exact text of `auth.ts` in its prompt prefix. If asked to make a follow-up edit, it generates code based on an assumed mental model of the file, resulting in failed patches, duplicate function declarations, or syntax errors.

### The Auto-Compaction Thrashing Loop

A severe failure mode in long-running agent workflows is auto-compaction thrashing. This occurs when a repository contains massive individual files, generated bundles, or voluminous test suites. If a single file read or bash output exceeds 30,000 tokens, the context window refills almost immediately after a compaction cycle completes.

```text
Context Saturated (Approaching 180,000 tokens)
               |
               v
       Auto-Compaction Runs
               |
               v
Context Reduced (Summary generated)
               |
               v
Claude Re-Reads Massive File or Verbose Test Suite
               |
               v
Context Refills Immediately
               |
               v
Auto-Compaction Triggers Again (Thrashing Loop)
```

When this cycle repeats multiple times in rapid succession, Claude Code detects that compaction is failing to maintain usable headroom. The runtime halts execution and outputs an auto-compaction thrashing error.

### Compaction and Prompt Caching Invalidation

Compaction also carries a performance and cost penalty related to prompt caching. Claude Code uses prompt caching to accelerate response times and reduce API billing based on exact prefix matching.

Because compaction rewrites conversational history into a summary, the message prefix changes completely, invalidating the entire cached conversation. The first turn following compaction experiences higher latency as the API processes the new prompt prefix from scratch.

## Steps to Prevent Context Saturation and Agent Drift

Preventing Claude Code degradation requires moving away from the pattern of treating an agent session as an infinite, all-knowing stream. Production developers use explicit operational discipline, session modularization, and file-based state tracking.

### 1. Enforce Task-Scoped Sessions

The most effective protection against context rot is limiting each Claude Code session to a single, tightly defined objective. Divide large features into modular units:

* **Session 1:** Define database schemas and generate migration scripts. Run tests, verify output, commit changes to git, and terminate.
* **Session 2:** Launch a fresh session with `claude`. Implement API authentication routes against the committed schema, commit, and exit.
* **Session 3:** Launch a fresh session to write unit and integration tests for the authentication routes.

Use the `/clear` command between related sub-tasks in the same terminal instance. The `/clear` command wipes conversation history and tool outputs, reloading a fresh context window while preserving your active working directory.

### 2. Offload State to Repository Markdown Files

Do not rely on conversational history to remember project plans or technical decisions. Conversational memory is volatile and subject to compaction loss. Store operational state directly in repository markdown files:

* `PLAN.md`: A structured checklist of implementation steps, architectural choices, and technical requirements.
* `SCRATCHPAD.md`: A temporary working file where the agent records discovered interface contracts, curl responses, and debug notes.

Point Claude directly to the plan file: "Read `PLAN.md` and implement step 3." This injects structured context in roughly 500 tokens, bypassing hundreds of turns of stale debugging dialogue.

### 3. Configure Compact Instructions in CLAUDE.md

Add a dedicated Compact Instructions section to your repository's `CLAUDE.md` file:

```markdown
### Compact Instructions
When compacting conversation history, you must always preserve:
- The exact list of modified files and pending Git commits
- Unresolved bug hypotheses and failed test edge cases
- Explicit architectural constraints: never edit files in `src/generated/`
- Current API endpoint contracts and payload schemas under active development
```

Claude Code incorporates these instructions into its summarization prompt, ensuring critical negative constraints survive summarization.

### 4. Delegate Heavy Investigation to Subagents

When investigating a sprawling codebase or running experimental benchmarks, do not perform that work in your primary session. Claude Code supports custom subagents that execute tasks inside isolated context windows:

```text
Primary Session (Clean Context: ~15,000 tokens)
       |
       +---> Spawns Research Subagent
       |         |
       |         +-- Reads 12 files (35,000 tokens)
       |         +-- Runs grep sweeps across codebase (8,000 tokens)
       |         +-- Synthesizes findings
       |
       +---< Returns 400-token structured summary
Primary Session Context Remains Crisp (~15,400 tokens)
```

The tokens generated by exploratory file reads, syntax parsing, and grep outputs remain quarantined inside the subagent's temporary context window. When the subagent finishes, it returns a concise summary to the primary session, keeping working context focused on implementation.

### 5. Monitor Context Usage with `/context`

Run the `/context` command in your terminal periodically to view a visual breakdown of consumed tokens across system prompts, `CLAUDE.md` files, tool calls, and conversation history. If tool outputs occupy the bulk of your context budget, run a focused compaction with `/compact focus on <target>`, or commit your progress and run `/clear` to start fresh with a clean slate.

## Offloading Large Corpora to Persistent Workspaces via Fast.io MCP

In enterprise engineering environments, project knowledge frequently outgrows local markdown files and single-repository boundaries. Complex microservice architectures, OpenAPI specifications, legacy documentation, compliance matrices, and cross-team dependencies easily span tens of thousands of tokens.

Attaching these massive assets directly to Claude chats or dumping them into `CLAUDE.md` guarantees rapid context saturation and performance degradation. Storing reference documentation in local git repositories also bloats clone times and forces coding agents to perform expensive local grep sweeps that consume session tokens.

### Externalizing Reference Knowledge to Fast.io Workspaces

The sustainable architectural pattern for large-corpus development is decoupling working memory from reference knowledge. Developers achieve this by offloading reference corpora into persistent, cloud-hosted [Fast.io intelligent workspaces](/product/workspaces/) accessible via the Model Context Protocol (MCP).

```text
                      +------------------------------------------+
                      |         Fast.io Cloud Workspace          |
                      |  - API Specs, Schemas, Architecture Docs |
                      |  - Intelligence Mode (Auto-Indexed RAG)   |
                      |  - Version History & Audit Log           |
                      +------------------------------------------+
                                           ^
                                           | Streamable HTTP
                                           | Semantic Queries & Citations
                                           v
+------------------------------------------------------------------------+
|                          Claude Code CLI                               |
|  - Active Terminal Session                                             |
|  - MCP Tool Search (queries Fast.io on demand via /mcp/code)          |
|  - Working Context Stays Lean (around 25,000 tokens)                   |
+------------------------------------------------------------------------+
```

Fast.io provides shared, organization-owned cloud workspaces designed for teams of humans and AI agents, delivering [storage for AI agents](/storage-for-agents/). When Intelligence Mode is enabled, Fast.io automatically indexes uploaded documents for hybrid retrieval, combining full-text keyword indexing, semantic vector search, and metadata filtering via [workspace intelligence](/product/ai/).

When Claude Code requires information about an internal API schema or database policy, it executes a targeted search query through the MCP server. Fast.io returns only the relevant paragraphs with document citations, injecting 300 tokens of high-precision context instead of 40,000 tokens of raw file data.

### Connecting Claude Code to Fast.io MCP

Claude Code connects to Fast.io using the remote MCP server over Streamable HTTP, as documented in the [Fast.io MCP setup guide](https://mcp.fast.io/docs).

To add the Fast.io MCP server to your Claude Code environment, execute the following command:

```bash
claude mcp add --transport http fast-io https://mcp.fast.io/mcp/code
```

After adding the server, authenticate by running `/mcp` inside your Claude Code session to initiate an OAuth 2.0 authorization flow in your browser. For project-level configuration shared via version control, add the Fast.io server entry to your repository's `.mcp.json` file:

```json
{
  "mcpServers": {
    "fast-io": {
      "type": "http",
      "url": "https://mcp.fast.io/mcp/code"
    }
  }
}
```

The `/mcp/code` endpoint exposes consolidated tools tailored for coding agents, detailed in the [MCP tooling reference](https://mcp.fast.io/skill.md):
* `search`: Performs semantic and keyword queries across workspace documents, returning matched snippets with exact file paths and citations.
* `execute`: Retrieves targeted file contents and executes workspace inspection commands.
* `upload` and `upload_manage`: Persists generated artifacts, benchmark summaries, or architecture notes back into the shared workspace.

### Collaborative Multi-Agent Persistence

Offloading reference corpora to Fast.io workspaces provides structural advantages beyond token savings:

* **Durable Version History:** Every document and artifact written to a Fast.io workspace retains full per-file version history via [persistent cloud storage](/product/ai-storage/). Prior versions remain auditable and restorable if an agent overwrites an API contract.
* **Advisory File Locks:** When multiple agents or developers work against the same workspace assets, agents coordinate using advisory file leases. The `lock-acquire` and `lock-release` actions on the `storage_manage` tool allow agents to signal active edits.
* **Cross-Cloud Synchronization:** Fast.io provides Cloud Sync for Dropbox, Box, and OneDrive, supporting one-way or two-way synchronization on a schedule or on demand. Google Drive files can be imported directly today, with automated sync scheduled for future release.
* **Transparent Pricing Structure:** Creating an account is free; performing operational work requires an organization on a paid subscription. Monthly plans start with a 30-day free trial that requires a credit card and ends early if trial credits are exhausted. Annual plans start paid with no trial. Full details are available on the [Fast.io pricing page](/pricing/).

| Fast.io Plan | Monthly Price | Included Storage | Included Seats | Monthly AI Credits |
| :--- | :--- | :--- | :--- | :--- |
| Starter | $9.99/mo | 250 GB | 3 seats | 100,000 credits |
| Business | $49.99/mo | 5 TB | 10 seats | 600,000 credits |
| Enterprise | $199.99/mo | 25 TB | 30 seats | 3,000,000 credits |

AI credits meter workspace semantic indexing and intelligence queries across connected agent sessions.

## Diagnostic Checklist for Triage and Session Recovery

When Claude Code begins displaying signs of degradation during an active development session, continuing to argue with the model in natural language only compounds the problem. Each corrective prompt adds more conversational turns, further diluting attention.

Follow this systematic checklist to diagnose, triage, and recover degraded agent sessions:

### Step 1: Diagnose Context Allocation with `/context`

Pause the active workflow and inspect token distribution:

```text
/context
```

Examine the output breakdown:
* If tool results and bash command output account for the bulk of total context, your session is suffering from attention dilution.
* If conversation history reaches 40,000 tokens, intermediate instructions have begun to drift.
* Check whether large whole files were read into context unnecessarily.

### Step 2: Choose the Correct Recovery Command

Select the triage action that matches your workflow state rather than defaulting to passive continuation:

| Operational State | Recommended Command | Mechanism and Tradeoff |
| :--- | :--- | :--- |
| Agent went down a flawed logic path or broke code | Press `Esc` twice or run `/rewind` | Reverts conversation and file system to a prior checkpoint. Preserves warm prompt cache. |
| Current sub-task is complete, but session history is noisy | `/compact focus on <next objective>` | Condenses history into a focused summary. Clears raw tool outputs but incurs a cache rebuild turn. |
| Switching to a completely new feature or module | Commit changes to git and run `/clear` | Completely wipes working context. Re-initializes clean context window with zero baggage. |
| Tool thrashing error or unrecoverable drift | Terminate CLI process and restart | Fresh session startup. Re-reads project `CLAUDE.md` and initializes clean token space. |

### Step 3: Validate File System State Against Git

Never trust an agent's conversational claim that an edit was applied successfully after a long session. Always verify actual disk state against git:

```bash
git status
git diff
```

Look for common degradation artifacts: duplicate definitions from stale buffer models, reverted edge-case fixes overwritten during later turns, and stray debug logs.

### Step 4: Codify Missing Constraints in CLAUDE.md

If Claude repeatedly violated an architectural rule, do not simply re-state it in the chat prompt. Chat prompts vanish when the session ends. Open your repository's `CLAUDE.md` or create a path-scoped rule in `.claude/rules/` and write the constraint explicitly:

```markdown
### Rule: Prevent Direct Database Modifications
Never generate raw SQL migrations inside application service layers.
Always place schema changes in `packages/database/migrations/`.
```

By encoding rules in persistent repository files, you ensure that future sessions, subagents, and post-compaction states enforce the constraint automatically.

### Diagnostic Matrix: Symptoms and Remedies

| Observed Symptom | Primary Root Cause | Immediate Triage Action | Permanent Architectural Fix |
| :--- | :--- | :--- | :--- |
| Agent repeats previously failed bash commands | Attention dilution from verbose error logs | Run `/rewind` to pre-error state | Filter bash commands with quiet flags |
| Hallucinates edits against outdated function signatures | Compaction lossiness; stale syntax tree | Run `git diff` to inspect state, then `/clear` | Keep sessions task-scoped; offload plans to `PLAN.md` |
| Ignores negative rules established in opening prompt | Low attention weight on distant conversational tokens | Re-state rule in a fresh prompt or `/compact` | Move negative constraints to root `CLAUDE.md` |
| Auto-compaction halts with thrashing error | Single massive file read or command output refilling window | Kill session; delete large output logs | Store large corpora in Fast.io workspace via MCP |

## Frequently asked questions

### Why does Claude Code get worse during long sessions?

Claude Code performance degrades during extended sessions due to attention dilution, context window saturation, and compaction lossiness. As a session progresses, verbose tool outputs from bash commands and whole-file reads consume thousands of tokens, dispersing the transformer model's attention weights away from core system instructions. When automatic compaction summarizes the conversation to free up space, detailed negative constraints and exact file representations are compressed away, leading to hallucinated edits and repeated errors.

### How do you prevent Claude Code from hallucinating in large codebases?

To prevent hallucinations in large codebases, keep sessions task-scoped and limit each session to a single feature or bug fix. Offload project plans and task tracking to repository markdown files like PLAN.md instead of relying on chat memory. Delegate codebase exploration and heavy file searches to isolated subagents so that large file reads do not bloat the primary context window. Finally, store large reference documentation and specifications in external workspaces connected via the Fast.io MCP server so the agent retrieves targeted snippets rather than ingesting whole files.

### What is context degradation in AI coding agents?

Context degradation is the progressive decline in an AI coding agent's reasoning precision, instruction following, and environment tracking as its active context window fills with tokens. In terminal agents like Claude Code, raw terminal logs, compiler errors, and file contents accumulate in the prompt prefix on every turn. This attention dilution causes the model to overlook critical rules established earlier in the conversation, resulting in faulty code generation and circular debugging attempts.

### How does the /compact command affect Claude Code prompt caching?

The /compact command summarizes conversation history to reclaim context window space, but doing so invalidates the conversation layer of Anthropic's prompt cache. Because prompt caching requires an exact prefix match from the start of the prompt, replacing conversation history with a summary breaks the cached prefix. The turn immediately following compaction requires a full uncached reprocessing of the new prompt, resulting in higher latency and standard token processing billing.

### When should developers use subagents instead of a single Claude Code session?

Developers should use subagents for exploratory, high-token tasks such as searching across large directory trees, reading multiple candidate files, or running verbose diagnostic suites. Subagents operate in an isolated context window with their own copy of CLAUDE.md. When the subagent completes its task, it returns a concise summary to the primary session, preventing tens of thousands of exploratory tokens from diluting the main session's working memory.

## Sources

- [Claude Help Center: Upload files to Claude](https://support.claude.com/en/articles/8241126-upload-files-to-claude): Anthropic's documented Claude limits specify that a standard Claude chat accepts up to 20 files at up to 500MB each, while Claude Projects accepts files up to 30MB each with unlimited file count provided the total content fits within Claude's context window, as detailed in Anthropic's documented Claude limits.

## About Fast.io

Fast.io provides shared workspaces where people and AI agents work on the same files, with built-in semantic search and citation-backed chat over what they hold. Agents reach it through a remote MCP server, a REST API at https://api.fast.io/current/, and a command line client published on npm as @vividengine/fastio-cli. MCP setup is at https://mcp.fast.io/docs: Claude and most MCP clients connect to https://mcp.fast.io/mcp/tools, ChatGPT to https://mcp.fast.io/mcp/operations, and coding agents to https://mcp.fast.io/mcp/code.
