# Copilot Context Window: Limits in Chat, CLI, and Agent Mode

The Copilot context window determines how many tokens of code, conversation history, and system instructions GitHub Copilot or Microsoft Copilot can evaluate at once. Standard IDE chat interfaces cap prompt context between 8,000 and 32,000 tokens, while GitHub Copilot CLI and agent mode reach 128,000 tokens with background compaction. Indexing extensive technical documentation in an external workspace over MCP prevents context exhaustion.

Source: https://fast.io/resources/copilot-context-window/
Author: [Derek Labian](https://fast.io/authors/derek-labian/)
Last reviewed: 2026-09-28

## What Is the Copilot Context Window Across Chat, CLI, and Agent Mode?

The Copilot context window is the maximum number of tokens, spanning conversation history, system instructions, active files, and editor selections, that GitHub Copilot or Microsoft Copilot can evaluate in a single interaction. While standard IDE chat interfaces allocate between 8,000 and 32,000 tokens of context depending on the underlying model, GitHub Copilot CLI and agent mode reach 128,000 tokens with GPT-4o and `Claude 3.5 Sonnet`. Understanding these distinct tiers prevents silent truncation errors when engineers attempt to feed entire repositories, complex multi-file debugging logs, or long architectural documents into an interactive session.

A primary source of developer confusion stems from conflating foundation model API context ceilings with product interface prompt allowances. When OpenAI releases a model like GPT-4o with a 128,000-token context window, engineers frequently assume their Visual Studio Code chat panel can ingest an entire 100,000-token project in a single prompt. In practice, developer tools enforce strict intermediate boundaries. Running maximum-length context windows on every keystroke introduces unacceptable inference latency and excessive GPU compute costs. To keep code completions sub-second and chat responses responsive, GitHub Copilot partitions context across specialized operational modes.

The following comparison details how token boundaries, usable prompt allowances, output limits, and compaction behaviors vary across Copilot surfaces:

| Operational Mode | Model Context Capacity | Usable Prompt Allowance | Output Generation Ceiling | Context Eviction / Compaction Behavior |
| :--- | :--- | :--- | :--- | :--- |
| **IDE Inline Completion** | Active file buffer | Cursor proximity window | Single statement or block | Local editor heuristics prioritize neighboring lines and recent edits |
| **IDE Copilot Chat (VS Code, JetBrains)** | 8,000 to 32,000 tokens | Approximately 4,000 to 18,000 tokens | Up to 4,096 tokens | Drops oldest conversational turns silently when limit is approached |
| **GitHub Copilot CLI** | Up to 128,000 tokens | Dynamic working set | Up to 8,192 tokens | Background compaction triggers at 80% capacity with 20% reserved buffer |
| **Copilot Agent Mode (VS Code)** | Up to 128,000 tokens | Dynamic workspace context | Model-dependent (4,096 to 8,192 tokens) | Tool execution outputs buffered; files referenced via index or truncated |
| **Microsoft 365 Copilot** | Model-dependent (up to 128,000 tokens) | 2,000 to 4,000 character prompt box | 30 turns per topic | Grounded via Microsoft Graph; references up to 20 files per agent query |
| **Copilot Studio Custom Agents** | 128,000 to 400,000 tokens | Configured prompt budget | Model-dependent | Model selection governs prompt ceiling (128k general, 400k reasoning) |

These operational boundaries mean that your choice of interface directly dictates how much project knowledge the model can inspect simultaneously. While quick inline completions require minimal peripheral context, multi-file refactoring and command-line automation demand deliberate context management to avoid losing critical architectural constraints.

## How Context Budgets Work: System Prompts, Tool Schemas, and the 40% Reserve

Even when an interface advertises a 32,000-token or 128,000-token ceiling, the space available for your actual code is substantially smaller than the headline number suggests. In GitHub Copilot and Microsoft Copilot, context windows function as unified memory buffers that must accommodate multiple competing requirements before processing your first keystroke.

Every request sent through the Copilot pipeline consumes tokens across several mandatory categories:

* **System Prompts:** The baseline persona instructions, guardrails, and behavioral rules that define how the assistant formats code, cites sources, and handles security restrictions.
* **Custom Instructions:** Repository-specific guidance provided in configuration files such as `.github/copilot-instructions.md`, which inject coding standards and project rules into every turn.
* **Built-in Tool Schemas:** The JSON schema definitions for local IDE tools, file search utilities, symbol finders, and terminal execution commands.
* **External MCP Tool Definitions:** When developers connect external tools through the Model Context Protocol, the complete schema for each exposed tool action is permanently loaded into the active context window.
* **Conversation Turns:** The accumulated history of user prompts, assistant answers, and intermediate reasoning steps across the current thread.
* **Active File Snippets:** Content explicitly referenced using `@file` or gathered through editor heuristics based on open tabs and cursor position.
* **Reserved Output Space:** A preallocated token buffer reserved exclusively for model response generation.

### Why Empty Chats Show Up to 40% Context Usage

Developers monitoring context consumption in Visual Studio Code often encounter an unexpected phenomenon: starting a fresh chat session and typing a minimal greeting like "hello" can immediately register up to 40% context usage. This behavior is neither a bug nor a memory leak.

As highlighted in community discussions regarding Copilot context consumption, GitHub Copilot chat preallocates system instructions, tool schemas, and reserved output space that can consume up to 40% of the context capacity even on minimal prompts. The reserved output allocation ensures that when you ask Copilot to generate a complex refactoring patch or write unit tests, the model does not run out of generation headroom halfway through emitting code. Because output capacity cannot be dynamically borrowed from input during active token generation, the system protects that buffer from being overwritten by user attachments.

### Tool Output Limits and the 20 KiB Threshold

In agentic workflows, tool executions represent the fastest way to exhaust an active context window. Running a test suite that produces verbose stack traces, executing a package manager build command, or reading an entire generated JSON schema can easily generate tens of thousands of tokens of output in a few milliseconds.

To prevent command output from flooding the model context, GitHub Copilot CLI implements an automatic output cutoff. Tool responses exceeding `20 KiB` are redirected by default to temporary files on disk. Instead of pasting thousands of raw lines into the model prompt, Copilot passes the file path and a concise preview to the assistant. Developers who need larger terminal logs injected directly into the conversation can adjust this behavior by setting the `COPILOT_LARGE_OUTPUT_THRESHOLD_BYTES` environment variable, though increasing this threshold reduces the remaining token space available for subsequent reasoning and tool calls.

## Context Compaction, Checkpoints, and Multi-Turn Retention in Copilot CLI

When running complex terminal tasks or multi-stage code refactoring in GitHub Copilot CLI, sessions can span dozens of command executions, code reviews, and file modifications. Because raw message logs would rapidly exceed the 128,000-token boundary, Copilot CLI relies on automatic background compaction to sustain long-running sessions without dropping critical project continuity.

### Automatic Compaction at 80% Capacity

In GitHub Copilot CLI, the system actively monitors the token consumption of the active working set. GitHub Copilot CLI starts automatic context compaction when conversation history reaches approximately 80% of context window capacity, maintaining a 20% headroom buffer for active tool execution.

This 20% headroom buffer prevents terminal execution commands and tool calls from failing abruptly while compaction occurs in the background. If a series of heavy tool calls drives context consumption toward full capacity before the background summarization completes, the CLI pauses execution briefly to allow the compaction pipeline to finish. Developers can also trigger this process manually at any time by issuing the `/compact` command, clearing obsolete conversational fluff before beginning a distinct phase of work.

### Structured Summaries and Historical Checkpoints

The compaction mechanism does not simply delete older messages from the top of the buffer. Instead, the CLI executes an internal summarization routine that extracts essential session state:

1. The assistant takes a complete snapshot of the conversation history.
2. The model generates a structured technical summary capturing high-level goals, decisions made, files inspected or modified, and planned next steps.
3. The detailed conversational history is replaced with the structured summary, preserving original user instructions and active plan items.
4. Messages received while the background compaction ran are merged into the clean working set.

Every time compaction occurs, the CLI creates a permanent checkpoint. Checkpoints represent serialized snapshots of the compaction summary saved directly in the session workspace. Developers inspect saved checkpoints by entering the following command:

```text
/session checkpoints
```

Entering `/session checkpoints 1` displays the exact technical state captured during the first compaction pass. Reviewing checkpoints is invaluable when an assistant appears confused about an earlier decision: it allows the developer to determine whether a subtle constraint was omitted during summarization or whether the task needs to be re-anchored with an explicit prompt update.

## Claude Projects vs Copilot: Why Large Corpora Exhaust Model Windows

Understanding the limits of coding assistants requires examining how context boundaries operate across different vendor architectures. Anthropic file upload rules provide an instructive contrast to Copilot in-editor boundaries.

As documented in Anthropic platform guidance, standard Claude chat accepts individual files up to `500 MB` each, while a Claude Project accepts files up to `30 MB` with no fixed cap on overall project file count. However, the total content across all uploaded files must fit entirely within Claude context window. Claude Projects has no fixed file-count cap, meaning the practical ceiling on any project is the context window itself. Once an engineering team uploads several extensive specification documents, schema definitions, and API guides, the cumulative tokens reach that ceiling, forcing developers to look for alternative context architectures.

Copilot environments face an identical architectural ceiling. Whether working within the 32,000-token boundary of an IDE chat or the 128,000-token boundary of Copilot CLI, stuffing raw files directly into the prompt buffer triggers severe operational tradeoffs:

* **Context Rot and Attention Degradation:** Large language models experience reduced retrieval accuracy when evaluated across tens of thousands of tokens, commonly known as the "lost in the middle" problem. When critical business logic rules or database constraints are buried in an oversized prompt, the model frequently ignores instructions or hallucinates non-existent function arguments.
* **High Latency Overhead:** Passing 100,000 tokens of file context into an LLM on every conversational turn causes significant inference delay, turning real-time interactive development into a sluggish waiting game.
* **Non-Code Documentation Exclusion:** Modern software engineering relies heavily on non-code assets, including architectural decision records, product requirement documents, OpenAPI specifications, and deployment runbooks. Storing these large assets inside Git repositories to satisfy local editor extensions inflates repository clone times and consumes expensive enterprise code index quotas.

Attempting to solve context limitations by requesting ever-larger model windows ignores the economic and cognitive reality of language models. For engineering teams managing extensive documentation sets, the answer is not stuffing larger files into the chat window, but decoupling persistent file storage from the prompt evaluation buffer.

## How to Extend Assistant Context with External Workspaces and Remote MCP

Engineering teams running sophisticated development workflows need persistent, searchable access to project context that extends beyond local editor buffers and CLI compaction summaries. When coordinating multi-agent coding sessions, microservice architectures, or distributed team specifications, relying on transient prompt memory creates isolated information silos.

Teams historically attempted to solve this gap with temporary workarounds:

* **Local Vector Databases:** Running local Chroma or Qdrant instances allows engineers to index documentation locally. However, local vector stores cannot be shared easily across a distributed team, require ongoing database administration, and fail to track upstream repository changes.
* **Git Submodule Bloat:** Committing large specification PDFs, database schema dumps, and customer requirements directly into application repositories slows down version control operations and wastes enterprise repository index allowances.
* **Unstructured Object Buckets:** Storing documentation in Amazon S3 or Google Cloud Storage provides persistence, but raw cloud storage lacks native semantic indexing and real-time agent search interfaces.

### Intelligent Workspaces for Human-Agent Teams

Fast.io provides an alternative architecture by serving as an intelligent workspace platform where technical documentation, architectural specifications, and agent outputs live alongside development teams. For engineering groups evaluating dedicated infrastructure, [storage for AI agents](/storage-for-agents/) unifies persistent team files and agent scratchpads.

Rather than attempting to pack comprehensive system specifications into an IDE prompt buffer or relying on periodic git indexers, teams place their documentation corpus into an org-owned [Fast.io workspace](/product/workspaces/). Files can be uploaded directly or synchronized from existing cloud storage: Cloud Sync connects Dropbox, Box, and OneDrive on a scheduled or on-demand basis, while Google Drive operates as an import connector today with sync coming soon.

Once assets reside in a workspace, [Fast.io AI features](/product/ai/) enable Intelligence Mode to index files automatically for hybrid search, combining full-text keyword indexing with semantic vector retrieval. Assistants and human developers query the exact same knowledge base, retrieving concise, highly relevant excerpts with verified document citations rather than stuffing massive source files into conversational memory. For structured records like API catalogs or configuration tables, [Metadata Views](/product/document-data-extraction/) extract typed attributes without manual parsing.

```json
{
  "mcpServers": {
    "fastio": {
      "url": "https://mcp.fast.io/mcp/code"
    }
  }
}
```

### Remote Model Context Protocol and Concurrency Controls

Coding assistants connect to Fast.io through the remote Model Context Protocol endpoint at `https://mcp.fast.io/mcp/code`, as documented in the [storage for AI agents](/storage-for-agents/) guide and the [Fast.io MCP documentation](https://mcp.fast.io/docs). Through a consolidated MCP toolset, agents search project documentation on demand, read structured Metadata Views, and coordinate task handoffs with human teammates using Collaborative Notes.

When multiple developers and autonomous agents interact within the same project workspace, built-in governance controls keep changes safe and auditable:

* **Per-File Version History:** Every write operation generates a permanent version record, allowing engineers to audit agent modifications, inspect code diffs, and restore prior file states.
* **Advisory File Locks:** Agents acquire, heartbeat, and release file leases through the MCP storage tool (`lock-acquire`, `lock-status`, `lock-release`), signaling active edits to other participants without blocking urgent human updates.
* **Activity Events Subscriptions:** The workspace events feed allows services to subscribe to file updates over WebSockets or long-polling, automatically triggering downstream testing or agent execution when a new schema version is published.

Every organization starts with a 30-day trial, which requires a credit card. Subscriptions are available on the Starter plan at `$29/mo` (5 seats, `1 TB` storage, 300,000 credits a month), Business at `$99/mo` (20 seats, `10 TB` storage, 1,200,000 credits a month), and Enterprise at `$299/mo` (50 seats, `50 TB` storage, 4,500,000 credits a month) on [Fast.io pricing](/pricing/). Credits meter AI work. Storage and seat allowances are included with each subscription. Additional credit packs cost `$10` per 100,000 credits.

## Frequently asked questions

### What is the context window for GitHub Copilot?

GitHub Copilot context window size depends on the interface and model. In Visual Studio Code chat, the context window allocates between 8,000 and 32,000 tokens for chat history and active editor buffers. In GitHub Copilot CLI and agent mode, the context window supports 128,000 tokens with GPT-4o and `Claude 3.5 Sonnet`, with extended `1M-token` modes available in select preview environments.

### Does GitHub Copilot support 128k context?

GitHub Copilot supports a 128,000-token context window when using GitHub Copilot CLI and agentic modes powered by models like GPT-4o and `Claude 3.5 Sonnet`. Standard in-editor chat panels in Visual Studio Code and JetBrains IDEs operate with tighter prompt allocations between 8,000 and 32,000 tokens to ensure low latency and responsive completions.

### How do I fix Copilot running out of context?

To prevent Copilot from running out of context, clear long conversation threads periodically, use the /compact command in Copilot CLI, and reference only relevant files using @file mentions rather than loading broad folders. For projects with large documentation sets, connect an external intelligent workspace over the Model Context Protocol to retrieve relevant sections dynamically with citations.

### Why does GitHub Copilot show 40% context usage on an empty chat?

GitHub Copilot displays up to 40% context usage on empty chats or minimal prompts because the interface preallocates system prompts, built-in tool definitions, custom repository instructions, and a reserved output buffer. This reserved buffer ensures the model always has sufficient generation capacity to return complete code blocks without truncating mid-response.

### What is the difference between Copilot Chat and Copilot CLI context limits?

Copilot Chat runs inside integrated development environments with an 8,000 to 32,000 token context window tailored for immediate editor assistance and diff previews. Copilot CLI runs in the terminal with 128,000 tokens, supporting autonomous tool executions, automatic background compaction at 80% usage, and persistent session checkpoints.

### How does GitHub Copilot CLI compaction work?

GitHub Copilot CLI compaction triggers automatically when conversation history reaches approximately 80% of context window capacity. The CLI snapshots the session, generates a structured summary of technical decisions, files touched, and next steps, and replaces detailed conversation history with the summary while keeping custom instructions and plan state intact.

## Sources

- [GitHub Docs: Managing context in GitHub Copilot CLI](https://docs.github.com/en/copilot/concepts/agents/copilot-cli/context-management): GitHub Copilot CLI starts automatic context compaction when conversation history reaches approximately 80% of context window capacity, maintaining a 20% headroom buffer for active tool execution.
- [GitHub Community: Discussion #188691](https://github.com/orgs/community/discussions/188691): GitHub Copilot chat preallocates system instructions, tool schemas, and reserved output space that can consume up to 40% of the context capacity even on minimal prompts.

## About Fast.io

Fast.io provides shared workspaces where people and AI agents work on the same files, with built-in semantic search and citation-backed chat over what they hold. Agents reach it through a remote MCP server, a REST API at https://api.fast.io/current/, and a command line client published on npm as @vividengine/fastio-cli. MCP setup is at https://mcp.fast.io/docs: Claude and most MCP clients connect to https://mcp.fast.io/mcp/tools, ChatGPT to https://mcp.fast.io/mcp/operations, and coding agents to https://mcp.fast.io/mcp/code.
