# Continue.dev Token Limit: Context Window Configuration and Codebase Indexing

The Continue.dev token limit is the maximum context length configured in Continue's config.json or config.yaml file that determines how much code, chat history, and codebase context can be passed to the underlying model. Configuring contextLength, setting provider parameters like Ollama num_ctx, and offloading repository documentation to an intelligent workspace prevents editor freezing and prompt truncation.

Source: https://fast.io/resources/continue-dev-token-limit/
Author: [Tom Langridge](https://fast.io/authors/tom-langridge/)
Last reviewed: 2026-09-23

## How Continue.dev Enforces Token Limits and Context Window Boundaries

When a developer pastes a multi-file refactoring prompt or an architectural review into Continue within Visual Studio Code or JetBrains, the assistant frequently stops mid-generation, drops function definitions, or throws a context window error. This breakdown occurs because Continue functions as an orchestration layer between the code editor and an underlying language model, applying its own context window ceilings before prompt tokens ever reach the provider API.

The Continue.dev token limit is the maximum context length configured in Continue's config.json or config.ts file that determines how much code, chat history, and codebase context can be passed to the underlying model. Modern releases also support config.yaml as the standard configuration format. Continue.dev defaults contextLength to 4,096 tokens if unspecified, which truncates large file diffs and multi-file code reviews.

To understand why token limits cause unexpected behavior in development environments, software engineers must recognize that Continue constructs prompts from several competing sources. Every request sent to the model consumes a portion of a single, shared context budget:

* System instructions: Continue injects foundational instructions, model personality rules, formatting directives, and custom rule files defined in `.continue/rules`.
* Conversation history: Previous messages, user queries, code responses, and tool invocation outputs accumulate as a conversation progresses.
* Active file buffers: Code from currently opened tabs or active editor selections is packaged into the prompt to provide immediate situational awareness.
* Pinned context modifiers: When a developer explicitly references files or terminal output using `@file`, `@folder`, `@diff`, or `@terminal`, Continue reads the source content into the prompt buffer.
* Codebase retrieval chunks: Semantic search queries and repository map summaries pull relevant snippets into the prompt payload.
* Generation reserve: The model requires an unallocated slice of the context window to generate its completion, defined by maxTokens.

Continue enforces distinct token thresholds across different subsystems rather than applying a single universal ceiling. The primary chat interface handles conversational reasoning and multi-step tasks, governed by the contextLength property assigned to the active model. Tab autocomplete operates under strict latency constraints, typically requiring sub-second response times and dedicated smaller context windows. Inline code editing (/edit) passes the target file alongside the prompt instructions; if the file combined with the instructions exceeds the context ceiling, diff generation fails or outputs broken syntax trees. Codebase indexing segments source files into discrete token chunks using embedding parameters like maxEmbeddingChunkSize.

The table below outlines how Continue handles token allocations across its core subsystems:

| Subsystem | Configuration Parameter | Default Token Budget | Typical Practical Ceiling | Behavior When Exceeded |
| :--- | :--- | :--- | :--- | :--- |
| Chat and Agent Mode | contextLength | 4,096 | Model native (up to 200,000) | Context truncation, instruction drift, dropped code |
| Tab Autocomplete | contextLength | 2,048 | 2,048 to 4,096 | High completion latency, discarded file context |
| Inline Edits (/edit) | contextLength | 4,096 | Model native (up to 128,000) | Broken syntax trees, partial diff applications |
| Embedding Chunks | maxEmbeddingChunkSize | 512 | 1,024 | Embedding API rejections, chunk truncation |
| Model Output | maxTokens | 2,048 | 4,096 to 16,384 | Truncated responses, cut-off functions |

Understanding these subsystem boundaries is essential for configuring Continue correctly. When a model's declared context window is smaller than the active prompt payload, Continue truncates context, dropping earlier conversation turns and reference snippets to fit the remaining content into the model's memory buffer.

## How to Configure contextLength and maxTokens Across LLM Providers

Configuring token limits in Continue requires modifying the global or workspace configuration file. Continue supports configuration through YAML or JSON files located in user profile directories or workspace roots.

Configuration files are located in standard paths across operating systems:

* macOS and Linux: `~/.continue/config.yaml` (modern) or `~/.continue/config.json` (legacy)
* Windows: `%USERPROFILE%\.continue\config.yaml` or `%USERPROFILE%\.continue\config.json`
* Workspace Root: `.continue/config.yaml` or `.continue/config.json` in the repository root for project-specific overrides

Developers can also access the active configuration directly from the editor. Clicking the gear icon in the lower-right corner of the Continue sidebar in Visual Studio Code or JetBrains opens the active configuration file in an editor tab.

To adjust token limits, developers configure three primary properties: contextLength, maxTokens, and requestOptions. The JSON snippet below demonstrates the syntax for config.json:

```json
{
  "models": [
    {
      "title": "Anthropic Claude 3.5 Sonnet",
      "provider": "anthropic",
      "model": "claude-3-5-sonnet-latest",
      "apiKey": "YOUR_ANTHROPIC_API_KEY",
      "contextLength": 200000,
      "defaultCompletionOptions": {
        "maxTokens": 8192
      },
      "requestOptions": {
        "timeout": 60000
      }
    },
    {
      "title": "Ollama Llama 3.1",
      "provider": "ollama",
      "model": "llama3.1:latest",
      "contextLength": 8192,
      "requestOptions": {
        "timeout": 120000,
        "extraBodyProperties": {
          "num_ctx": 8192
        }
      }
    }
  ]
}
```

In modern Continue releases using config.yaml, the equivalent declaration defines contextLength and maxTokens under defaultCompletionOptions:

```yaml
models:
  - name: Anthropic Claude 3.5 Sonnet
    provider: anthropic
    model: claude-3-5-sonnet-latest
    apiKey: YOUR_ANTHROPIC_API_KEY
    defaultCompletionOptions:
      contextLength: 200000
      maxTokens: 8192
    requestOptions:
      timeout: 60000
  - name: Ollama Llama 3.1
    provider: ollama
    model: llama3.1:latest
    defaultCompletionOptions:
      contextLength: 8192
    requestOptions:
      timeout: 120000
      extraBodyProperties:
        num_ctx: 8192
```

Each model provider interacts with contextLength differently:

* Ollama: Setting contextLength in Continue is not always sufficient for local Ollama instances. Continue controls the client-side truncation window, but Ollama allocates GPU VRAM and CPU memory based on its internal num_ctx parameter. If you declare contextLength as 8,192 in Continue but Ollama defaults to 2,048 or 4,096 tokens, Ollama truncates incoming prompts on the server side. Developers must explicitly pass num_ctx inside requestOptions.extraBodyProperties to synchronize both sides of the connection.
* Anthropic: The `Claude 3.5 Sonnet` model features a 200,000 token context window. Claude Projects accepts files up to 30MB each with an unlimited file count provided total content fits within the context window. If contextLength is omitted from Continue's configuration, the extension clamps requests to its default 4,096 tokens, discarding the vast majority of Claude's capacity. Setting contextLength to 200,000 and maxTokens to 8,192 unlocks full multi-file analysis.
* OpenAI: For `GPT-4o`, Continue supports a `128000` token input context alongside a `16384` token generation reserve. Setting contextLength to 128,000 and maxTokens to 16,384 ensures Continue does not prematurely cut off deep refactoring routines.
* Request options and timeouts: When sending prompts containing tens of thousands of tokens, inference latency increases. Setting requestOptions.timeout to 60,000 or 120,000 milliseconds prevents the IDE from dropping the connection before the provider begins streaming tokens.

## Why Codebase Indexing Causes IDE Freezing and Memory Bottlenecks

Continue includes a codebase indexing subsystem that scans repository files to build an index for semantic search and codebase awareness. While helpful for small utilities, local LanceDB vector indexing in Continue.dev can cause memory bottlenecks when scanning repositories with thousands of files without external storage boundaries.

The indexing lifecycle operates through a multi-step pipeline inside the editor extension host:

1. Repository traversal: Continue scans the project directory tree, reading file contents from disk into editor process memory.
2. Document chunking: Source files are broken into discrete text chunks according to maxEmbeddingChunkSize, typically 256 to 512 tokens per chunk.
3. Vector embedding generation: Chunks are transmitted to an embedding provider, such as Voyage AI, Ollama nomic-embed-text, or OpenAI text-embedding-3-small, to generate high-dimensional vector representations.
4. Database persistence: Vector embeddings are stored locally in LanceDB format at `~/.continue/index/lancedb`, while file hashes and indexing metadata are written to `~/.continue/index/index.sqlite`.

When applied to large codebases, this architecture encounters severe resource constraints. Large codebases frequently contain generated artifacts, third-party libraries in node_modules, Python virtual environments, compiled binaries, build caches, and large JSON fixtures. If Continue attempts to parse and embed these files, the VS Code or JetBrains extension host experiences severe memory pressure.

LanceDB uses an append-only, versioned data format. As files change and re-indexing runs occur, LanceDB retains historical versions to allow rollback capabilities. In active repositories with frequent Git branch switches, the local index directory can expand to tens of gigabytes, consuming disk space and causing disk I/O bottlenecks. When the extension host memory exceeds Node.js heap allocations, the editor freezes, completions stop responding, and background indexing halts.

Developers can mitigate local indexing degradation through targeted maintenance steps:

* Create a workspace .continueignore file: Place a `.continueignore` file in the root of your project using standard `.gitignore` syntax. Add entries for build directories, minified bundles, lockfiles, data dumps, and media assets.
* Configure global ignore rules: Add global ignore patterns in `~/.continue/.continueignore` (macOS/Linux) or `%USERPROFILE%\.continue\.continueignore` (Windows) to prevent indexing standard cache folders across all projects.
* Purge corrupted index caches: If the index becomes corrupted or the IDE persistently hangs during startup, close the editor completely and delete the `~/.continue/index` directory. Restarting the editor triggers a clean re-index.

While ignore rules help, local vector indexing faces an inherent architectural limit. Code editors are designed for responsive, low-latency text manipulation, not for hosting resource-intensive vector databases across enterprise-scale repositories.

## How to Offload Large Repositories and Architecture Context to Fast.io MCP

When engineering teams work with monorepos, microservice architectures, or extensive technical documentation, keeping all contextual material inside local editor memory or within prompt tokens becomes unsustainable. Rather than constantly tweaking local LanceDB chunk sizes or fighting context window limits, teams can offload their technical corpus to an external intelligent workspace.

Traditional storage tools like Google Drive, Dropbox, or raw cloud buckets allow file uploads, but they function as passive file repositories. They do not index codebases and technical documentation for semantic agent retrieval. When a developer asks an AI assistant about system architecture, commodity drives cannot deliver relevant technical passages directly into the assistant's reasoning loop.

[Fast.io workspaces](/product/workspaces/) provide an intelligent coordination layer designed for developers and AI agents. Teams upload or import source repositories, technical specifications, API contracts, database schemas, and architectural decision records into a shared workspace. Once uploaded, Intelligence Mode automatically indexes all files for hybrid search, combining full-text keyword matching, semantic meaning search, and metadata filtering without requiring a separate vector database.

Continue connects to Fast.io through its remote Model Context Protocol (MCP) server. Instead of forcing Continue to scan thousands of local files or paste large documentation files into chat prompts, the assistant queries Fast.io over Streamable HTTP at `https://mcp.fast.io/mcp` (or `https://mcp.fast.io/mcp/key` with Bearer token authentication).

To connect Continue to Fast.io via MCP, add the server definition to your Continue configuration file. In `config.yaml`, add the server under `mcpServers`:

```yaml
mcpServers:
  - name: fastio
    url: https://mcp.fast.io/mcp/key
    headers:
      Authorization: Bearer YOUR_FASTIO_API_KEY
```

In legacy `config.json`, configure the remote transport under the experimental block:

```json
{
  "experimental": {
    "modelContextProtocolServers": [
      {
        "transport": {
          "type": "streamable-http",
          "url": "https://mcp.fast.io/mcp/key",
          "headers": {
            "Authorization": "Bearer YOUR_FASTIO_API_KEY"
          }
        }
      }
    ]
  }
}
```

When configured, Continue gains access to consolidated MCP tools for workspace exploration, document retrieval, and semantic search. When a developer prompts the assistant to implement a new service according to enterprise specifications, Continue queries the workspace, retrieves only the relevant paragraphs and schema definitions, and injects them into the prompt. This leaves the model's native context window completely free for code generation, syntax validation, and diff creation.

Beyond saving local token capacity, storing team context in Fast.io introduces critical collaborative controls:

* Per-file version history: Every document and code asset maintains a complete, restorable version history, keeping modifications between teammates and automated agents auditable.
* Append-only audit logs: Teams track every read, write, and export event across workspaces, providing accountability for automated agent actions.
* Granular permissions: Access can be scoped across organization, workspace, folder, and file tiers, ensuring coding agents only reach authorized repositories.
* Ownership transfer: An agent can create an organization, configure workspaces, index repositories, and transfer administrative ownership to a human team lead via an invite link while retaining operational access.

Creating an account is free; doing real work requires an organization on a paid subscription. Every organization starts with a 14-day free trial, which requires a credit card. Plans are Starter at `$9.99/mo`, Business at `$49.99/mo`, and Enterprise at `$199.99/mo` on [Fast.io pricing](/pricing/). Teams exploring agent infrastructure can review [Fast.io storage for agents](/storage-for-agents/) to integrate remote MCP endpoints into their daily development workflows.

## Best Practices for Managing Context Windows in Team Development

Managing token limits and context boundaries effectively requires intentional operational patterns across an engineering team. Following these five best practices prevents IDE freezing, controls API costs, and maintains high generation accuracy:

* Use targeted context referencing: Avoid broad folder queries like `@folder` on massive source trees. Instead, pinpoint specific files using `@file` or focus on active modifications with `@diff`. Directing the model to exact code segments keeps prompt token consumption well below default context ceilings.
* Maintain committed ignore files: Commit a comprehensive `.continueignore` file to your repository root alongside `.gitignore`. Ensure all team members exclude build artifacts, package managers, database dumps, and generated documentation from local indexing passes.
* Segment models by development task: Do not use a single model configuration for all editor operations. Configure fast, cost-effective models with modest context limits (4,096 to 8,192 tokens) for tab autocomplete and quick single-function edits. Reserve frontier models with large context windows (128,000 to 200,000 tokens) for multi-file refactoring, code reviews, and architectural planning.
* Offload team knowledge to remote workspaces: Rather than forcing each developer's laptop to index multi-gigabyte repositories, store shared architecture documents, API schemas, and technical specs in an intelligent workspace. Let developers query that knowledge on demand through remote MCP connections, preserving local editor responsiveness.
* Monitor token usage and latency: Regularly check Continue's token counters and request latencies. If generation latency spikes or responses cut off unexpectedly, review whether active editor tabs or broad context providers are flooding the prompt budget. Adjust contextLength and maxTokens settings to restore balanced performance.

## Frequently asked questions

### How do I increase the context length in Continue.dev?

To increase context length in Continue.dev, open your configuration file (`~/.continue/config.yaml` or `~/.continue/config.json`) and set the `contextLength` parameter under your model definition. In modern YAML configs, declare `contextLength: 200000` (or your model's maximum supported limit) inside `defaultCompletionOptions`. Save the file, and Continue will automatically reload the new token ceiling.

### Where is the Continue.dev config.json or config.yaml file located?

The global configuration file is located at `~/.continue/config.yaml` (or `config.json`) on macOS and Linux, and `%USERPROFILE%\.continue\config.yaml` on Windows. You can also place a `.continue/config.yaml` file in the root of your workspace for repository-specific settings, or click the gear icon in the lower-right corner of the Continue sidebar inside your IDE.

### How does Continue.dev handle large codebase indexing?

Continue indexes codebases by scanning project files, chunking them into discrete token segments (governed by `maxEmbeddingChunkSize`), and generating vector embeddings stored locally in LanceDB (`~/.continue/index/lancedb`) with metadata in SQLite (`~/.continue/index/index.sqlite`). For large repositories, you should configure a `.continueignore` file to prevent indexing build caches and large data assets that cause IDE memory exhaustion.

### Why does Ollama return context length errors in Continue.dev?

Ollama manages its memory and context window through its own server-side parameter called `num_ctx`. If you set `contextLength` in Continue but do not configure `num_ctx` in Ollama, Ollama uses its internal default (often 2,048 or 4,096 tokens) or fails if system VRAM is exceeded. To resolve this, specify `num_ctx` inside `requestOptions.extraBodyProperties` in Continue's model configuration.

### What is the difference between contextLength and maxTokens in Continue.dev?

`contextLength` defines the total working memory budget for both input prompt tokens and output generation tokens combined. `maxTokens` specifies the maximum number of tokens the model is allowed to produce in its generated response. The `maxTokens` value must always be smaller than `contextLength` to ensure adequate space remains for system instructions, conversation history, and code context.

### How can teams share codebase context without exhausting local editor memory?

Teams can store architectural documentation, technical specifications, and repository assets in an intelligent workspace like Fast.io. With Intelligence Mode enabled, files are indexed automatically for hybrid search. Developers connect Continue to Fast.io via the remote MCP server at `https://mcp.fast.io/mcp`, allowing the assistant to query documents on demand without consuming local editor memory or filling prompt budgets.

## Sources

- [Anthropic Help Center: Upload files to Claude](https://support.claude.com/en/articles/8241126-upload-files-to-claude) — Claude Projects accepts files up to 30MB each with an unlimited file count provided total content fits within the context window.

## About Fast.io

Fast.io provides shared workspaces where people and AI agents work on the same files, with built-in semantic search and citation-backed chat over what they hold. Agents reach it through a remote MCP server at https://mcp.fast.io/mcp, a REST API at https://api.fast.io/current/, and a command line client published on npm as @vividengine/fastio-cli.
