# Anthropic Context Window: Claude Token Limits, Prompt Caching, and MCP Workspaces

The Anthropic context window defines the active token capacity available across Claude models for system instructions, conversations, and attached documents. While standard models support a baseline 200,000-token window, managing large files requires understanding prompt caching breakpoints and project constraints. Connecting Claude to an indexed workspace via the Model Context Protocol allows teams to query expansive corpora through targeted search without exhausting active tokens.

Source: https://fast.io/resources/anthropic-context-window/
Author: [Tom Langridge](https://fast.io/authors/tom-langridge/)
Last reviewed: 2026-09-22

## The Anthropic Context Window: Token Limits Across Claude Models

The Anthropic context window is the 200,000-token capacity supported across all modern Claude 3 and 3.5/3.7 models, allowing roughly 150,000 words or 500 pages of input text per prompt. In production deployments, this parameter determines how much raw information an engineer can place in front of the model before conversational turns degrade or system instructions are pushed out of memory. While Anthropic has introduced extended 500,000-token and 1,000,000-token windows for specialized workloads on newer models like Claude Sonnet and Opus, the standard 200,000-token tier remains the default operational baseline across the Claude API and standard developer subscriptions.

Understanding how tokens map to operational artifacts is essential for designing resilient AI agent workflows. In Anthropic's tokenizer, one token corresponds to approximately 0.75 English words, or about four characters. A 200,000-token buffer accommodates roughly 150,000 words of conversational English prose. In technical workflows, however, token density increases sharply. Code snippets, JSON schemas, XML wrappers, API payloads, and stack traces contain frequent punctuation, unique variable identifiers, and structured indentation. These elements consume tokens at a noticeably higher rate than standard natural language sentences. A codebase containing 10,000 lines of source code can consume anywhere from 60,000 to 120,000 tokens depending on comment density and syntax complexity.

The table below outlines the documented context windows, prompt caching minimum breakpoints, write rates, and read discounts across modern Claude models as of September 2026:

| Claude Model | Standard Context Window | Prompt Caching Minimum | 5-Minute Cache Write Rate | Cache Read Discount |
| --- | --- | --- | --- | --- |
| Claude 3.7 Sonnet | 200,000 tokens | 1,024 tokens | 1.25 times base input | Up to 90% |
| Claude 3.5 Sonnet | 200,000 tokens | 1,024 tokens | 1.25 times base input | Up to 90% |
| Claude 3.5 Haiku | 200,000 tokens | 2,048 tokens | 1.25 times base input | Up to 90% |
| Claude 3.0 Opus | 200,000 tokens | 1,024 tokens | 1.25 times base input | Up to 90% |
| Claude 3.0 Haiku | 200,000 tokens | 2,048 tokens | 1.25 times base input | Up to 90% |

In practice, an agent prompt never dedicates its entire context window to user documents. Every call must partition its token budget across several competing operational components:

* **System instructions and role framing:** Technical system prompts, security boundaries, style rules, and behavioral constraints typically consume between 1,000 and 5,000 tokens.
* **Tool definitions and schema declarations:** Declaring external functions, Model Context Protocol schemas, and API parameters consumes between 2,000 and 8,000 tokens before any action is taken.
* **Conversational history and tool results:** Multi-turn exchanges, previous tool inputs, returned outputs, and intermediate reasoning steps accumulate cumulatively with each interaction cycle.
* **Working document context:** The remaining balance represents the headroom available for active code files, reference manuals, and user inputs.

When an application loads 160,000 tokens of reference material into an active session, it leaves only 40,000 tokens for system rules, tool declarations, user prompts, and model output generation. If conversational turns continue inside that saturated buffer, the model quickly reaches its operational ceiling.

## How Prompt Caching Interacts With the Context Window

Repeatedly transmitting large context payloads over network APIs introduces steep computational costs and noticeable latency. When an agent passes 150,000 tokens of API documentation to Claude on every conversational turn, the model must process that static background text repeatedly. Anthropic prompt caching solves this inefficiency by allowing the API to store static prefixes in memory, dramatically accelerating response times and lowering billing charges.

Prompt caching operates on an exact prefix-matching architecture. When a developer flags a static segment of a prompt using ephemeral cache control parameters, Anthropic caches the computational representations of those tokens. On subsequent API calls containing the exact same prefix, Claude reuses the cached state rather than reprocessing the text from scratch.

To apply prompt caching effectively, developers must respect four technical constraints:

* **Minimum token thresholds:** Prompts must satisfy strict minimum token lengths to trigger caching. Claude Sonnet and Claude Opus require a minimum prefix of 1,024 tokens. Claude Haiku requires a minimum prefix of 2,048 tokens. Prompts below these breakpoints bypass the cache entirely and are processed as standard input tokens.
* **Prefix ordering requirements:** Anthropic constructs cache keys hierarchically from the top of the prompt downward in a strict sequence: tools first, followed by system instructions, followed by the message sequence. If a single character changes in an earlier block, every cached block that follows it is invalidated immediately.
* **Cache control breakpoints:** Developers can define up to four explicit cache breakpoints per prompt using the `cache_control: {"type": "ephemeral"}` parameter on tools, system messages, or specific conversational turns.
* **Time-to-Live (TTL) expiration:** By default, cached prefixes remain in memory for a minimum lifetime of five minutes. Each time a request reads from the cache, the five-minute timer refreshes at no additional cost. Anthropic also provides an optional one-hour cache duration for workflows with longer intervals between prompts, billed at two times the base input price instead of the standard 1.25 times write rate.

The financial and latency benefits of prompt caching are substantial. Anthropic prompt caching reduces input token costs by up to 90% and latency by up to 85% for long prompts. For an engineering team querying a 100,000-token API reference across dozens of development questions, caching the static specification reduces input costs from full rates down to tenth-rate reads on all subsequent calls within the TTL window.

The following Python example illustrates how to configure explicit cache breakpoints on tool definitions and system instructions using the Anthropic API:

```python
import anthropic

client = anthropic.Anthropic()

response = client.messages.create(
    model="claude-3-7-sonnet-20250219",
    max_tokens=2048,
    system=[
        {
            "type": "text",
            "text": "You are a senior software architect analyzing system designs.",
        },
        {
            "type": "text",
            "text": "ARCHITECTURAL GUIDELINES: Apply modular domain boundaries and verify data contracts.",
            "cache_control": {"type": "ephemeral"},
        },
    ],
    messages=[
        {
            "role": "user",
            "content": "Evaluate our microservices decomposition strategy.",
        }
    ],
)

print(f"Cache write tokens: {response.usage.cache_creation_input_tokens}")
print(f"Cache read tokens: {response.usage.cache_read_input_tokens}")
print(f"Standard input tokens: {response.usage.input_tokens}")
```

While prompt caching significantly reduces latency and financial overhead, it does not expand the physical dimensions of the 200,000-token context window. A 150,000-token cached document still consumes 150,000 tokens of the model's total capacity, leaving the remaining 50,000 tokens for generation, tool definitions, and user interaction.

## File Attachment Constraints in Claude Chats and Projects

Many developers first encounter token boundaries when attaching files directly to Claude conversations or configuring Claude Projects. Anthropic enforces distinct file ingestion rules depending on whether content is uploaded into a direct chat session or stored inside a project workspace.

In standard chat conversations, Claude accepts a ceiling of twenty attachments per message, with an individual file size ceiling of 500 megabytes per file. Claude accepts a wide range of document types, including PDF, DOCX, TXT, CSV, HTML, and code files. PDF processing behavior depends directly on document length:

* **PDFs of 100 pages or fewer:** Claude applies multimodal visual parsing, inspecting text alongside diagrams, embedded charts, tables, and graphic elements.
* **PDFs between 101 and 1,000 pages:** Claude falls back to text-only extraction, stripping all visual figures and evaluating raw character streams.
* **PDFs exceeding 1,000 pages:** The interface rejects the file entirely, triggering an explicit file size error.

Claude Projects introduces a completely different mechanism for persistent reference context. According to Anthropic's documented file upload specifications, project files have an individual file size limit of 30 megabytes per file. The number of uploaded files in a project is completely unlimited, but the total extracted content across all files must fit within Claude's context window.

This constraint debunks the common myth that Claude Projects enforces an arbitrary file-count cap, such as five or ten files. A team can successfully upload 50 concise configuration files of 2KB each without issue. Conversely, uploading a single 25 megabyte markdown repository dump or a dense legal contract corpus will immediately exceed the 200,000-token ceiling. In project environments, file count is irrelevant; cumulative token volume is the only binding limit.

When project files or direct chat attachments exhaust the context window, applications experience several predictable failure modes:

* **Immediate prompt rejection:** If extracted document tokens combined with system instructions exceed 200,000 tokens, the API or web interface returns a prompt length error and refuses execution.
* **Attention dilution (lost in the middle):** When models process context windows packed near their physical limit, attention mechanisms distribute weights across vast token distances. Information placed in the middle portions of an expansive prompt suffers noticeably higher retrieval error rates compared to content positioned near the beginning or end of the document.
* **Context thrashing under auto-compaction:** In Claude environments where automatic context management is active, the system compresses earlier conversational turns into brief narrative summaries when total tokens approach the limit. While this compression prevents hard runtime crashes, it destroys granular technical details. If database constraints, API schemas, or variable naming conventions were agreed upon in early messages, lossy summarization frequently flattens them into generic notes like 'reviewed backend requirements.' The model then hallucinates parameters or reintroduces previously rejected bugs.

## Why Context Window Ceilings Drive the Shift to MCP Retrieval

Encountering context saturation forces engineering teams to reconsider their architectural strategy for handling large document corpora. The instinct to solve context limits by waiting for larger model windows overlooks the fundamental economics and performance dynamics of large language models. Evaluating 500,000 or 1,000,000 tokens on every prompt incurs significant computational latency, high token costs, and increased susceptibility to hallucination.

Instead of stuffing complete documentation libraries directly into the prompt context, high-performance agent architectures decouple storage from reasoning. Rather than making the model read the entire library on every turn, the agent queries an indexed repository on demand and ingests only the relevant paragraphs.

Teams handling expansive context generally evaluate three operational approaches:

* **Local filesystem inspection:** Developers configure scripts or local language servers to search local directories using text grep or Abstract Syntax Tree parsing. While functional on a single developer workstation, local filesystem tools cannot synchronize state across distributed teams, struggle with complex non-text formats like scanned PDFs, and risk process hangs when interacting with virtual cloud-sync placeholders.
* **Self-hosted vector databases:** Teams deploy specialized vector databases like Pinecone, Weaviate, or Qdrant. While vector retrieval supports semantic similarity search, maintaining standalone database clusters requires dedicated embedding pipelines, custom chunking logic, ongoing index synchronization, and complex access control maintenance.
* **Remote Model Context Protocol (MCP) servers:** Anthropic established the Model Context Protocol as an open standard for connecting AI models to external tools, databases, and document repositories. By exposing indexed knowledge bases over standardized MCP endpoints, models can discover available tools dynamically, execute targeted queries, and retrieve precise information snippets without loading full documents into prompt memory.

The architectural contrast between in-context document stuffing and MCP-based retrieval highlights why external workspaces have become the standard for agent engineering:

| Operational Dimension | Direct Context Ingestion | Remote MCP Workspace Retrieval |
| --- | --- | --- |
| Context Window Overhead | Consumes 50,000 to 180,000 active tokens permanently | Consumes 300 to 1,000 tokens per retrieved excerpt |
| Corpus Scale Limit | Hard ceiling at model limit (200,000 tokens) | Virtually unbounded (gigabytes of indexed files) |
| Reasoning Headroom | Severely restricted by static document volume | Preserves 190,000+ tokens for multi-step agent reasoning |
| Retrieval Accuracy | Suffers from attention dilution across dense text | Delivers targeted, pre-ranked excerpts with file citations |
| Multi-Turn Economics | Accumulates full input costs on every conversational turn | Reuses minimal tokens, maximizing prompt caching efficiency |

By shifting document storage and search out of the active context window and into an external MCP retrieval tier, Claude maintains full reasoning capacity. The model focuses its attention on the developer's immediate objective rather than carrying megabytes of passive reference text through every conversational step.

## Structuring External Knowledge Bases and Workspaces in Fast.io

Solving the context window challenge for enterprise documentation requires an intelligent storage layer designed for AI agents and human teams alike. Storing technical specifications, customer records, and architectural blueprints in an external Fast.io workspace allows Claude to access gigabytes of reference material through targeted retrieval without consuming precious prompt tokens. Fast.io leaves Anthropic's context window exactly where it is, providing an indexed repository for the files that do not fit inside the prompt.

Teams can populate Fast.io workspaces through direct file uploads or automated cloud import. Fast.io supports importing files directly from Dropbox, Box, and OneDrive, while Google Drive imports today with sync coming soon. This enables organizations to centralize legacy documentation repositories, technical manuals, and multi-gigabyte project assets into dedicated workspaces without writing custom synchronization scripts.

Once files enter a workspace, enabling Intelligence Mode activates built-in retrieval-augmented generation. Fast.io automatically constructs a hybrid index combining exact full-text keyword matching, semantic vector embeddings, and search-by-metadata-value without requiring external vector databases or chunking pipelines. When structured documents like invoices, policies, or contracts require schema extraction, Metadata Views at [/product/document-data-extraction/](/product/document-data-extraction/) turn unstructured files into queryable data grids with typed columns.

Claude connects to these indexed workspaces using Anthropic's Model Context Protocol. Fast.io provides a remote MCP server supporting modern Streamable HTTP at `https://mcp.fast.io/mcp` (or `https://mcp.fast.io/mcp/key` with Bearer token authentication) alongside legacy Server-Sent Events at `https://mcp.fast.io/sse`.

Developers using Claude Desktop or Claude Code can connect Fast.io by adding the server configuration to their MCP settings file:

```json
{
  "mcpServers": {
    "fastio": {
      "url": "https://mcp.fast.io/mcp/key",
      "headers": {
        "Authorization": "Bearer YOUR_FASTIO_API_KEY"
      }
    }
  }
}
```

Once configured, Claude automatically discovers Fast.io's consolidated MCP tools. The interaction workflow proceeds through five clean operational phases:

1. **User request:** A developer asks Claude: 'What are our data retention requirements for EU customer records?'
2. **Context evaluation:** Claude recognizes that data retention policies are not loaded into its local prompt context.
3. **Targeted MCP search:** Claude invokes the Fast.io `storage` tool using the `search` action, querying the workspace with the terms 'EU customer data retention compliance.'
4. **Focused excerpt injection:** Fast.io executes hybrid keyword and semantic retrieval, returning the exact two relevant paragraphs accompanied by source file citations.
5. **Accurate synthesis:** Claude incorporates the 400-token excerpt into its active prompt and produces a precise answer, preserving over 195,000 tokens of context headroom for subsequent coding and reasoning tasks.

Beyond offloading context tokens, centralizing project knowledge in Fast.io provides production governance features that direct file attachments cannot match:

* **Per-file version history:** Every document retains complete revision tracking, ensuring agents and human collaborators can audit changes, compare diffs, and restore previous versions.
* **Append-only audit log:** Every file read, search query, document update, and permission change is recorded in an immutable audit trail for complete operational traceability.
* **Collaborative Notes:** Fast.io Notes provides real-time co-editing surfaces where human developers and autonomous AI agents collaborate as first-class multiplayer editors.
* **Granular permissions:** Organizations enforce strict access controls across organization, workspace, folder, and file tiers, ensuring AI assistants only query documentation authorized for their operational scope.

Engineering teams can evaluate external workspace retrieval during onboarding. Creating an account is free; doing real work requires an organization on a paid subscription. Monthly plans start with a trial of up to 30 days on [Fast.io pricing](/pricing/).

## Frequently asked questions

### What is the Anthropic context window limit?

The standard Anthropic context window is 200,000 tokens across all modern Claude 3 and 3.5/3.7 models, representing approximately 150,000 words or 500 pages of text. While select models support extended 500,000-token or 1,000,000-token windows in specific enterprise environments, 200,000 tokens remains the standard baseline across the public API and developer plans.

### How does prompt caching work with Anthropic's context window?

Anthropic prompt caching allows customers to reduce costs by up to 90% and latency by up to 85% for long prompts, with a default five-minute time-to-live that refreshes upon every read. Prompts must satisfy minimum thresholds of 1,024 tokens on Claude Sonnet and Opus, or 2,048 tokens on Claude Haiku to trigger caching.

### How many pages of text is 200k tokens in Claude?

In standard English prose, 200,000 tokens equates to roughly 150,000 words, which corresponds to approximately 500 single-spaced printed pages. In technical workflows containing source code, JSON schemas, or configuration files, token density is higher, yielding closer to 300 to 400 pages of content.

### Can you expand Anthropic's context window beyond 200,000 tokens?

While Anthropic offers 500,000-token and 1,000,000-token context windows on select newer models under specialized tiers, expanding raw context increases latency and costs. A more scalable approach is connecting Claude to an external workspace like Fast.io via the Model Context Protocol, enabling semantic queries across gigabytes of documents while loading only relevant passages into active memory.

### What are the file upload limits in Claude chats versus Claude Projects?

Direct chat sessions accept a ceiling of twenty files at 500 megabytes each, with PDFs capped at 1,000 pages. Claude Projects accepts individual files up to 30 megabytes each with no fixed limit on total file count; however, the cumulative extracted text across all project files must fit entirely within Claude's 200,000-token context window.

### Does prompt caching increase the available context window size?

Prompt caching does not increase the physical context window ceiling. A 150,000-token document cached in memory still occupies 150,000 tokens of the model's 200,000-token capacity. Prompt caching optimizes billing costs and response latency rather than expanding total token volume.

## Sources

- [Claude Help Center: How large is the context window on paid Claude plans?](https://support.claude.com/en/articles/8606394-how-large-is-the-context-window-on-paid-claude-plans) — Outside of newer extended models, Claude supports a 200,000-token context window that can ingest roughly 500 pages of text on paid plans.
- [Anthropic News: Prompt Caching](https://www.anthropic.com/news/prompt-caching) — Anthropic prompt caching allows customers to reduce costs by up to 90% and latency by up to 85% for long prompts.

## About Fast.io

Fast.io provides shared workspaces where people and AI agents work on the same files, with built-in semantic search and citation-backed chat over what they hold. Agents reach it through a remote MCP server at https://mcp.fast.io/mcp, a REST API at https://api.fast.io/current/, and a command line client published on npm as @vividengine/fastio-cli.
