AI & Agents

Cursor Token Usage: How to Track, Monitor, and Reduce AI Spend

Cursor token usage measures the volume of prompt and completion tokens processed during chat, codebase indexing, and Composer agent loops. Because multi-turn agent sessions repeatedly re-send workspace context, unmanaged repositories can burn through monthly compute credits in days. Teams can audit consumption through the Cursor dashboard, prune local context with ignore files, and offload static reference corpora to an indexed workspace via remote MCP.

Derek Labian 16 min read Updated
Auditing token consumption and context indexing prevents runaway AI spend across multi-turn agent sessions.

How Cursor Token Usage and Request Billing Work

A multi-turn Composer session rapidly accumulates tens of thousands of tokens when an agent ingests whole directory trees, inspects dependencies, and generates multi-file diffs. Many developers assume the editor charges a flat subscription fee for unlimited edits, but Cursor meters frontier models against actual token volume and compute credit pools.

Cursor token usage refers to the total volume of input prompt and output completion tokens consumed when executing chat queries, @codebase indexing, and multi-turn Composer agent loops in the Cursor editor.

Every interaction with an AI model breaks down into distinct token categories that carry different computational and financial costs:

  • Input Tokens. The text, system instructions, and file content sent to the model. Input tokens include your prompt, conversation history from previous turns, active rules files, files referenced with @file or @folder, and snippets retrieved during codebase searches.
  • Cache Write Tokens. Tokens processed when new context enters the model's prompt cache. Modern frontier models support prompt caching to store repeated system instructions and code context in memory across consecutive turns. Writing new context into the cache costs slightly more than standard input processing on the initial turn.
  • Cache Read Tokens. Tokens retrieved directly from the prompt cache during subsequent turns. When context remains static, cache reads cost a small fraction of standard input tokens, making long chat sessions more economical.
  • Output Tokens. The tokens generated by the model in response. These include generated code, diff blocks, terminal command proposals, and reasoning traces. Output tokens cost three to five times more per million units than input tokens across all major model providers.

Does Cursor Charge per Token or per Request?

Developers frequently ask whether Cursor charges per token or per request. The answer depends on which model tier and billing system you select.

Historically, Cursor marketed subscriptions around counts of fast requests, such as monthly fast request allowances on the Pro plan. However, modern frontier models vary widely in compute requirements. A single request using an extended reasoning model on a deep codebase context costs far more to serve than a single completion on a concise helper function.

To reflect these operational realities, Cursor transitioned to a dollar-denominated compute credit pool, as detailed in the official Cursor pricing documentation. Each paid plan provides a specific monthly allowance of model usage. First-party models (such as Composer and Grok) run within your base subscription allowance without extra token surcharges. When you manually select third-party frontier models (such as Claude Sonnet, Claude Opus, or specialized reasoning engines), Cursor meters your usage based on the actual input and output tokens consumed.

On Teams and Enterprise plans, Cursor applies a dedicated token surcharge per million tokens for third-party models. This platform fee applies even when you configure your own API keys through Bring Your Own Key (BYOK) settings. Once your monthly credit allowance is exhausted, Cursor stops processing requests unless you enable on-demand usage, which bills additional consumption in arrears.

Model Tier Primary Use Case Context Overhead Billing Mechanism
First-Party Models (Composer, Grok) Rapid inline generation and agent tasks Moderate Included in base subscription allowance
Frontier Third-Party Models Complex multi-file architecture High Deducted from monthly compute credits based on tokens
Extended Reasoning Models Deep debugging and multi-step verification Highest Rapidly depletes compute credits via internal reasoning tokens
Bring Your Own Key (Teams and Enterprise) Direct provider billing Full API rate Provider token cost plus platform token surcharge

Understanding this distinction clarifies why two developers on the same Pro tier plan can experience wildly different monthly runtimes. A developer executing targeted edits with first-party models rarely hits plan thresholds. A developer running continuous multi-file Composer loops against unindexed repositories can exhaust their credit pool in days. Teams managing multi-agent workflows can coordinate shared documentation using Fast.io Workspaces and connect models to shared assets through Fast.io AI.

Tracking and Auditing Token Consumption in Cursor

Controlling AI spend starts with knowing exactly where tokens go. Cursor provides two separate administrative surfaces to inspect consumption: the in-editor settings panel and the web management dashboard.

Checking Usage in the Editor Settings

To inspect immediate status without leaving your code:

  1. Open settings using Cmd+, on macOS or Ctrl+, on Windows and Linux.
  2. Select Models from the sidebar.
  3. Review your active model list and verify which providers are enabled.
  4. Check your account connection status to confirm whether requests route through Cursor's managed pool or your own API credentials.

The editor interface displays operational toggles and immediate model availability. For granular historical consumption, request counts, and dollar balances, you must inspect the centralized web dashboard.

Auditing the Web Management Dashboard

The primary hub for monitoring token usage is cursor.com/dashboard. Log in with your account credentials to access detailed billing metrics organized across two primary tabs:

  • The Usage Tab. This view breaks down every request processed during your active billing cycle. It separates usage by model family (for example, first-party Composer runs versus third-party frontier models) and shows the total volume of requests submitted. The dashboard visualizes daily consumption spikes, allowing you to trace large credit drains back to specific coding sessions or complex refactoring operations.
  • The Spending Tab. This view tracks financial accruals, on-demand compute charges, and remaining plan allowances. For individual accounts, it confirms whether on-demand usage is active. For team administrators, it provides seat-level breakdowns showing which team members consume the largest share of pooled compute.

Setting Hard Spend Limits

Uncapped on-demand billing exposes developers to unexpected monthly charges. A runaway background agent loop or an accidental prompt that repeatedly reads large database dumps can generate unexpected fees overnight.

To establish financial guardrails:

  1. Open cursor.com/dashboard and select the Spending tab.
  2. Locate the Spend Limit configuration control.
  3. Enter a hard monthly spending ceiling beyond your included plan tier.
  4. Save the threshold.

Once your total consumption reaches this figure, Cursor automatically halts on-demand requests. The editor displays a credit exhaustion alert rather than continuing to bill your credit card in arrears. Teams can configure both organization-wide spending limits and per-seat caps to prevent individual automated experiments from draining the department budget.

High-Burn Symptom Root Operational Cause Corrective Action
Monthly credits exhausted within days Unscoped @codebase prompts pulling entire repos Switch to precise @file and @symbol references
Sluggish response generation in Composer Context window bloated past reasonable limits Clear session history and start a fresh Composer thread
High on-demand charges on Teams plan Third-party frontier models selected by default Standardize team defaults on first-party models
Frequent cache invalidation Changing files early in long conversational threads Move static reference files out of the active edit directory

Why Composer Agent Sessions Consume Massive Token Volumes

Developers transitioning from standard inline autocomplete to Composer agent mode often notice their token burn rate multiplying rapidly. This increase is not a billing glitch. It is the direct consequence of how autonomous agent loops manage context.

In standard chat, an interaction is single-turn: you provide a prompt, the editor attaches the open file, the model responds, and the transaction ends. Composer functions differently. It executes multi-turn iterative loops where the model inspects files, runs terminal checks, generates diffs, analyzes errors, and applies revisions across multiple files.

The Mechanics of Context Compounding

The primary driver of token bloat in Composer is conversational compounding. In an agent loop, each new turn re-transmits the complete conversation history from every preceding turn. The model must receive the entire context history to maintain continuity and track which files it has already modified.

Consider a typical multi-turn refactoring task on a modest codebase where each prompt re-submits the full conversation history. Turn one might ingest your prompt and a single target file, consuming five thousand tokens. In turn two, Composer re-sends the original prompt, the first file, the model's generated diff, and the new follow-up instruction, climbing past twelve thousand tokens. By turn five, the cumulative payload sent to the model frequently exceeds thirty-five thousand tokens. Across a complete ten-turn debugging loop, cumulative input token volume easily compounds past one hundred thousand tokens for a single task.

The Codebase Indexing Overhead

Another major source of token waste is broad @codebase prompting. When you type @codebase in a prompt, Cursor executes a local vector search across your repository embeddings to retrieve relevant code snippets.

While vector search works well for finding specific method signatures, broad prompts (such as "How does authentication work across this app?") force the indexer to pull dozens of code chunks into the prompt context. If your repository contains generated files, minified bundles, build outputs, or mock data, the search index injects thousands of tokens of irrelevant text directly into the model's working window.

Upload Limits and Context Ceilings Across AI Assistants

Every AI development tool imposes technical boundaries on file intake and context size. According to Anthropic's documented Claude limits, chat uploads accept up to 20 files at up to 500MB each while project files accept up to 30MB each with no fixed file-count cap as long as content fits within the context window.

Whether working in Claude Projects or Cursor Composer, the operational bottleneck remains identical. When you pack static reference manuals, API documentation, and large schemas directly into prompt context, you quickly exhaust the context window and trigger rapid credit depletion.

Autonomous agent execution loop managing files and context in a workspace
Fastio features

Control Cursor Token Spend With an Indexed MCP Workspace

Connect Cursor to an indexed Fast.io workspace via remote MCP to search extensive reference documentation without exhausting prompt context. Monthly plans start with a 30-day free trial (credit card required).

Pruning Local Context With Ignore Files and Prompt Scoping

Developers can cut Cursor token consumption by enforcing strict local hygiene. Following a systematic optimization checklist prevents unnecessary files from entering the context window while keeping agent sessions focused.

The Four-Step Context Optimization Checklist

To keep token usage under control across daily development:

  1. Check Usage in Cursor Settings. Regularly inspect your active model choices and verify your spending limits on cursor.com/dashboard.
  2. Audit Composer Context Ingestion. Monitor active Composer sessions to identify when conversational turns compound past reasonable limits.
  3. Configure Ignore Files for Local Bloat. Use .cursorignore and .cursorindexingignore to block build outputs, dependencies, and lock files.
  4. Offload Reference Corpora to an Indexed Workspace. Move heavy external documentation, API schemas, and architecture guides into an external workspace accessible via remote MCP.

Configuring .cursorignore and .cursorindexingignore

Cursor respects your project's .gitignore by default. However, standard git repositories frequently track files that developers need in version control but should never expose to an AI model.

Cursor provides two distinct ignore mechanisms with different operational scopes:

  • .cursorignore. This file provides a complete block. Files matching patterns in .cursorignore are excluded from both AI indexing and runtime tool calls. The model cannot read them, and manual @file mentions will not resolve them.
  • .cursorindexingignore. This file excludes matching files from background vector indexing while keeping them accessible for explicit @file mentions in chat. Use this configuration for auxiliary utility code that you occasionally need to inspect manually but do not want cluttering semantic search results.

Create a .cursorignore file in the root directory of your project to prune build directories, dependencies, lock files, and large data fixtures:

#--- Dependencies and packages ---
node_modules/
vendor/
.venv/
target/

#--- Build artifacts and compiled files ---
dist/
build/
out/
*.bundle.js
*.min.js

#--- Lock files and dependency manifests ---
package-lock.json
pnpm-lock.yaml
yarn.lock
Cargo.lock
poetry.lock

#--- Logs, test reports, and temporary caches ---
*.log
coverage/
.next/
.turbo/
.cache/

#--- Data fixtures and mock datasets ---
fixtures/
data/*.csv
data/*.json

Scoping Prompts and Resetting Composer Threads

Pruning files from the repository index is only half the battle. How you prompt Composer determines how many tokens get transmitted on each turn.

  • Use Explicit References. Avoid generic prompts that trigger broad repository scans. Instead of asking "Where is customer billing handled?", use @file src/billing/stripe.ts or reference a specific interface with @symbol BillingService. Scoped prompts tell Cursor precisely which code blocks to attach, bypassing vector retrieval overhead entirely.
  • Clear Threads Promptly. When an agent finishes a specific refactoring task or debugging session, close the Composer thread and start a new session. Lingering conversational history from an earlier bug fix forces every subsequent turn to re-transmit obsolete diffs and error traces.
  • Branch Complex Tasks. For multi-file architecture changes, break the task into discrete phases. Have Composer generate the database migrations in thread one, verify the diff, and commit the code. Open thread two to build the service layer, attaching only the newly created migration schema.

Offloading Reference Corpora to an Indexed Workspace via Remote MCP

Adding files to .cursorignore solves the problem of local context bloat, but it introduces a severe operational gap: the AI cannot see those files at all.

Engineering teams frequently need their coding agents to consult extensive technical documentation, such as internal API specifications, third-party vendor SDK manuals, database schemas, and compliance frameworks. If you keep these reference assets inside your repository, they inflate background vector embeddings and consume thousands of prompt tokens every time they are mentioned. If you add them to .cursorignore, the model is forced to guess API signatures and hallucinates outdated methods.

The External Indexing Architecture

The solution is to decouple static reference documentation from your local source tree. Instead of cluttering your code repository with multi-megabyte PDFs and Markdown dumps, offload the reference corpus to an external cloud workspace on Fast.io.

When you place documents into a Fast.io workspace and enable Intelligence Mode, the platform automatically indexes the files for hybrid search combining full-text keyword indexing and semantic embeddings. Cursor connects to the workspace through the remote Model Context Protocol (MCP) server.

Rather than forcing Cursor to ingest a 300-page API manual into its active prompt buffer, the agent issues targeted search queries over MCP. The remote server performs semantic retrieval across your indexed files and returns only the precise paragraph or code snippet needed to answer the prompt. This retrieval model keeps active context lean, avoiding the heavy token compounding that drains compute credits in Composer.

Connecting Cursor to Fast.io via Remote MCP

Cursor supports remote MCP servers over Streamable HTTP. To link your project or user environment to an indexed documentation workspace, configure ~/.cursor/mcp.json (or .cursor/mcp.json in a project):

{
  "mcpServers": {
    "fastio": {
      "url": "https://mcp.fast.io/mcp/code"
    }
  }
}

Sign in with OAuth in the browser when Cursor connects. The Review Permissions screen lets you select Read Only or Read & Write access and choose which organizations and workspaces the connection can reach. For complete setup details, see the Fastio MCP documentation. Once configured, Cursor registers the Fast.io toolset. When your prompt requires reference documentation, the model calls the search tool, retrieves verified excerpts with source citations, and generates code without stuffing raw documentation into your prompt window.

Synchronizing Team Knowledge and Workspace Plans

Keeping reference documentation current across engineering teams requires continuous coordination. Fast.io supports Cloud Sync for Dropbox, Box, and OneDrive, allowing your team to synchronize reference folders on a schedule or on demand. Teams can also import documentation directly from Google Drive.

For structured extraction across technical assets, Metadata Views allow teams to define typed schemas that automatically extract properties like API endpoints, SDK versions, and deprecation dates without writing manual OCR rules. Team members collaborate on shared files through Fast.io Collaboration.

Plan Tier Monthly Price Included Seats Storage Allowance Included Monthly Credits Max Single Upload
Starter $9.99/mo 3 seats 250 GB 100,000 credits 25 GB
Business $49.99/mo 10 seats 5 TB 600,000 credits 50 GB
Enterprise $199.99/mo 30 seats 25 TB 3,000,000 credits 100 GB

Monthly plans start with a 30-day free trial requiring a credit card, allowing engineering teams to test workspace indexing with their existing Cursor workflows. Additional compute credits and storage can be added directly from organization billing when scaling automated agent workflows. To explore architecture patterns for persistent agent storage, read about Fast.io Storage for Agents.

Sources

References used to verify factual claims in this guide.

  1. Anthropic limits Claude chat uploads to 20 files at up to 500MB each and project files to 30MB each with no fixed file-count cap.

Frequently Asked Questions

How do I check token usage in Cursor?

You can inspect active model availability inside the editor under the model settings panel. For granular historical consumption, request counts, and dollar balances, log in to cursor.com/dashboard. The Usage tab details request volumes broken down by model family across your billing cycle, while the Spending tab tracks compute credit balances, on-demand accruals, and hard spend caps.

Does Cursor charge per token or per request?

Cursor uses a blended approach depending on your plan tier and selected models. Standard subscriptions provide an included allowance of model requests and compute credits for first-party models like Composer. When you select third-party frontier models, Cursor meters usage based on the actual input and output tokens consumed, deducting compute credits from your monthly pool or billing on-demand token surcharges once plan allowances are exhausted.

Why is Cursor using so many tokens during Composer sessions?

Multi-turn Composer sessions consume large token volumes because each successive turn re-transmits the complete conversation history alongside relevant file contents to maintain continuity. Using broad @codebase queries also injects dozens of code chunks from local vector search into the prompt context. If your repository contains generated bundles, dependencies, or lock files, context quickly balloons into tens of thousands of tokens per turn.

What is the difference between .cursorignore and .cursorindexingignore?

The .cursorignore file completely blocks matching files from both background indexing and runtime model context, preventing the AI from viewing or modifying them even if explicitly mentioned. In contrast, .cursorindexingignore excludes files from background vector embeddings while leaving them available for manual @file mentions in chat. Use .cursorignore for dependencies and build artifacts, and .cursorindexingignore for large reference files that you only need to query occasionally.

How can I keep large documentation accessible without burning Cursor tokens?

Instead of storing large PDF manuals, API specifications, and architecture docs in your local project root where vector indexing ingests raw files into prompt context, store them in a dedicated cloud workspace. Connecting Cursor to an indexed workspace via remote MCP allows the AI assistant to perform targeted semantic searches, retrieving only concise, relevant excerpts into the active prompt window while keeping raw file contents out of your token budget.

Related Resources

Fastio features

Control Cursor Token Spend With an Indexed MCP Workspace

Connect Cursor to an indexed Fast.io workspace via remote MCP to search extensive reference documentation without exhausting prompt context. Monthly plans start with a 30-day free trial (credit card required).