# Gemini Code Assist Limits: Quotas, Daily Caps, and Codebase Indexing

Gemini Code Assist enforces daily developer quotas, including 6,000 code completions, 240 chat requests, and 1,000 to 2,000 CLI requests per day, alongside a 1,000,000 token local context window and 20,000 repository indexing cap. When multi-turn agent sessions exhaust daily request allocations or 24-hour git reindexing cycles delay project context, connecting external workspaces via MCP provides searchable, real-time retrieval for large document corpuses.

Source: https://fast.io/resources/gemini-code-assist-limits/
Author: [Derek Labian](https://fast.io/authors/derek-labian/)
Last reviewed: 2026-09-24

## What Are the Daily Usage Limits and Quotas for Gemini Code Assist?

Google caps individual Gemini Code Assist usage at 6,000 code completions and 240 chat queries per day, according to official Google documentation as of September 2026. These usage ceilings apply directly to individual developers using the tool within supported integrated development environments like Visual Studio Code, JetBrains IDEs, and Cloud Workstations. Understanding these ceilings is important for developers who depend on uninterrupted AI code generation throughout the workday.

Gemini Code Assist limits are Google Cloud project quotas and daily usage caps, including 6,000 code completions, 240 chat queries, and 1,000 CLI requests per day, enforced on developer IDE extensions and repository indexing. Evaluating these restrictions requires distinguishing between three separate operational layers that control developer access:

1. **Interactive IDE assistance.** Real-time tab completions, inline code transformations, and conversational sidebars inside the code editor.
2. **Agent mode and command-line interfaces.** Automated terminal workflows and multi-step reasoning tasks executed through the Gemini CLI.
3. **Google Cloud project quotas.** Account-level and project-level administrative limits managed through the Cloud Quotas console.

In addition to daily allowances, Google applies instantaneous rate limits to prevent server congestion. Gemini for Google Cloud restricts interactive requests to 2 requests per second for each user in a project. If an automated script or rapid keystroke pattern sends requests faster than this threshold, the API responds with temporary throttling errors. Daily limits, by contrast, operate on a strict 24-hour reset cycle aligned with midnight Pacific Time (PT). Once a user reaches their daily ceiling, all subsequent generation requests to that model interface fail until the reset window passes.

The following comparison table details documented Gemini Code Assist limits, request allowances, and indexing boundaries across service tiers as of September 2026.

| Feature or Interface | Target Plan or Role | Enforced Quota or Limit | Reset Schedule | Documented Source |
| :--- | :--- | :--- | :--- | :--- |
| Individual Code Completions | Free Individual Tier | 6,000 requests per day | Midnight PT | codeassist.google |
| Individual Chat Queries | Free Individual Tier | 240 requests per day | Midnight PT | codeassist.google |
| Project Code Requests | Google Cloud Project | 6,000 requests per day | Midnight PT | docs.cloud.google.com/gemini/docs/quotas |
| Console & IDE Cloud Assist Chat | Google Cloud Project | 960 requests per day | Midnight PT | docs.cloud.google.com/gemini/docs/quotas |
| Agent Mode & Gemini CLI | Standard Edition | 1,500 requests per user per day | Midnight PT | docs.cloud.google.com/gemini/docs/quotas |
| Agent Mode & Gemini CLI | Enterprise Edition | 2,000 requests per user per day | Midnight PT | docs.cloud.google.com/gemini/docs/quotas |
| GitHub Pull Request Reviews | GitHub App Installation | At least 100 reviews per day | Midnight PT | docs.cloud.google.com/gemini/docs/quotas |
| Instantaneous Rate Limit | All Cloud Users | 2 requests per second | Sliding 1-second window | docs.cloud.google.com/gemini/docs/quotas |
| Local Codebase Awareness Context | Standard & Enterprise | 1,000,000 token context window | Per request | docs.cloud.google.com/gemini/docs/quotas |
| Code Customization Repositories | Enterprise Edition | 20,000 repositories | 24-hour reindex cycle | docs.cloud.google.com/gemini/docs/quotas |

While 6,000 daily completions provide ample headroom for typical manual coding, automated agent loops and deep repository indexing consume requests at a vastly accelerated pace.

## How Does Gemini Code Assist Enforce Agent Mode and CLI Caps?

Developers frequently encounter quota exhaustion when migrating from interactive IDE code completion to autonomous agent workflows. In standard editing mode, typing code generates single completion calls as the cursor moves. Agent mode and the Gemini CLI operate under a completely different consumption model.

When a developer issues an instruction in agent mode, the assistant does not make a single model call. To complete a complex instruction like refactoring a database client or generating unit tests across multiple files, the assistant conducts an iterative reasoning loop. It reads local directory listings, inspects file buffers, analyzes syntax trees, runs test scripts, evaluates terminal error outputs, and produces code diffs. Each iterative step generates one or more distinct model requests. A single user prompt in agent mode routinely consumes between 5 and 15 model requests behind the scenes.

Google explicitly groups agent mode and CLI usage under a single, unified daily quota. Quotas for requests from Gemini Code Assist agent mode and Gemini CLI are combined. For users on the Standard Edition, Google enforces a cap of 1,500 requests per user per day. Users on the Enterprise Edition receive an expanded allowance of 2,000 requests per user per day.

Crucially, Google aggregates these daily request limits across all model versions and families used with the Gemini CLI or agent mode, including both Gemini Pro and Gemini Flash. Developers cannot circumvent a depleted daily quota by switching from an advanced model to a lighter model in their configuration settings. Once the user reaches the maximum number of requests for that day, no further requests can be processed through those interfaces until the quota resets at midnight Pacific Time.

Google also monitors consumption velocity. Requests are limited per user per minute and remain subject to service availability during periods of peak demand. When regional data centers experience high traffic, background model requests in agent mode may experience queue delays or transient HTTP 429 status codes. Individual developers participating in the Google Developer Program or holding Google AI subscriptions can access premium limits by linking their personal license credentials to an active Google Cloud project.

## How Do Local Codebase Awareness and Code Customization Indexing Work?

Providing relevant code suggestions requires grounding generative models in your private code and team standards. Gemini Code Assist provides two distinct grounding mechanisms, each governed by its own architectural boundaries and indexing limits: local codebase awareness and enterprise code customization.

Local codebase awareness operates entirely within the developer's immediate editing environment. The IDE extension analyzes open editor tabs, active file buffers, and user-specified workspace folders to assemble context for prompt payloads. Gemini Code Assist supports a context window of 1,000,000 tokens for local codebase awareness. This expanded context allows the model to process substantial chunks of reference code alongside prompt instructions. Developers control which files contribute to local context by defining an `.aiexclude` file in the project root, preventing sensitive credentials, build artifacts, or vendor libraries from loading into prompt memory.

However, local codebase awareness has practical limits. The extension scans files directly on your local workstation. When working inside massive monorepos containing thousands of files, local file discovery degrades editor responsiveness and risks hitting context truncation thresholds during complex multi-file refactoring tasks.

To support large-scale enterprise environments, Google provides code customization for organizations on Gemini Code Assist Enterprise. Instead of relying on local workstation memory, code customization indexes private source code repositories hosted on GitHub, GitLab, or Bitbucket through Google Cloud Developer Connect. Google maintains a dedicated, single-tenant index environment for each organization, supporting a ceiling of 20,000 connected repositories. To ensure suggestions reflect recent updates, Google automatically reindexes registered repositories every 24 hours.

While enterprise code customization addresses repository scale, it introduces three significant operational constraints:

* **24-hour indexing latency.** Source code updates committed throughout the day are not reflected in code customization suggestions until the next overnight reindexing pass finishes. Fast-moving teams working on newly introduced internal libraries must wait up to a day for full semantic grounding.
* **The non-code documentation gap.** Software development involves far more than raw source code. Teams maintain architectural decision records, OpenAPI specifications, technical requirements, compliance checklists, and database migration runbooks in external formats like PDF, Markdown, Word documents, and spreadsheets. Git repository indexing cannot parse or organize this broader knowledge base.
* **Ecosystem isolation.** Enterprise repository indexing in Google Cloud operates exclusively within Gemini Code Assist. If your team uses diverse tools, such as Claude Code for terminal tasks, Cursor for feature development, and custom Python agents for data analysis, those external tools cannot query Google's private repository index.

## What Happens When You Exceed Gemini Code Assist Quotas?

When an engineer or automated pipeline surpasses configured Gemini Code Assist thresholds, Google immediately blocks further consumption. The exact failure behavior depends on which quota boundary was breached:

### 1. Instantaneous Concurrency Errors
Exceeding the 2 requests per second concurrency limit returns an immediate rate-limit error. In development environments, this presents as transient UI toast notifications in the IDE or HTTP 429 Too Many Requests responses in terminal logs. Because this limit is evaluated across a sliding one-second window, pauses in typing or brief exponential backoffs in scripts allow requests to resume almost immediately.

### 2. Daily Request Cap Exhaustion
Reaching the daily request cap produces a hard operational halt. In the IDE chat panel, Gemini displays messages indicating that your daily quota has been exhausted. In terminal workflows, the Gemini CLI returns a gRPC status code of `RESOURCE_EXHAUSTED` (code 8), noting that the quota metric for daily requests has been reached. When this occurs, all further requests to that interface fail until the quota automatically resets at midnight Pacific Time.

### 3. Pull Request Review Quotas on GitHub
Organizations using the Gemini Code Assist GitHub app receive an operational quota of at least 100 pull request reviews per day per installation. The exact review capacity varies based on the size of the codebase and the number of model calls required to evaluate each diff. If an active engineering team opens dozens of complex pull requests that consume all allocated model calls, automated reviews cease functioning until the daily reset.

### How to Audit and Adjust Quotas in Google Cloud Console
Project administrators can inspect active usage and submit quota adjustment requests through the Google Cloud console:

1. Sign in to the Google Cloud console and select your active project.
2. Navigate to **IAM & Admin** > **Quotas & System Limits**.
3. In the filter bar, enter `cloudaicompanion.googleapis.com` or search for "Gemini for Google Cloud".
4. Locate the specific quota metric, such as **Requests per day for Gemini Code Assist**.
5. Review your current percentage utilization against the assigned limit.

If your team consistently exhausts daily request allowances, administrators can select the target metric and click **Edit Quotas** to submit a formal increase request with business justification. Keep in mind that while project-level daily quotas can often be raised for enterprise accounts, certain system limits, including per-second concurrency caps, are fixed architectural constraints that cannot be modified.

## How Can You Offload Large Document Corpuses to Fast.io MCP?

When teams reach local context limits or encounter the 24-hour latency of git repository indexing, they require a decoupled storage architecture. Cramming complete technical manuals into prompt context windows or waiting for overnight repository indexing passes wastes engineering velocity.

Engineers typically evaluate two traditional paths for external knowledge retrieval:

1. **Bespoke vector databases.** Deploying standalone vector databases like Chroma, Pinecone, or Qdrant paired with custom chunking and embedding pipelines. While flexible, this approach demands continuous maintenance, infrastructure provisioning, and custom client integration code.
2. **Cloud data stores.** Configuring enterprise search solutions like Vertex AI Search or Google Cloud Storage buckets. These services introduce recurring search indexing fees and complex IAM permission schemes that add friction to developer onboarding.

An alternative production pattern separates persistent document storage from language models entirely, placing team knowledge into an external workspace accessed dynamically through the Model Context Protocol (MCP).

In this decoupled pattern, documentation libraries, architectural decision records, and API schemas are stored in shared, organization-owned [Fast.io workspaces](/product/workspaces/). Files can be uploaded directly or imported from cloud platforms like Google Drive, Dropbox, Box, or Microsoft OneDrive. Cloud synchronization functions for Dropbox, Box, and OneDrive, while Google Drive imports today with two-way sync coming soon.

When Intelligence Mode is enabled on the workspace, incoming documents are automatically indexed on arrival for hybrid search, combining full-text search, semantic vector retrieval, and metadata filtering without requiring an external vector database.

Instead of attaching multi-megabyte files directly to Gemini prompts, developers and coding assistants connect to the workspace through the remote Fast.io MCP server at `https://mcp.fast.io/mcp` (or with API key authentication at `https://mcp.fast.io/mcp/key`). Using the consolidated `storage` tool with the `search` action, the assistant queries the indexed workspace in real time and retrieves only the precise excerpts needed to answer the question, injecting a few hundred tokens into the prompt context.

To connect your coding assistant to a Fast.io workspace, add the remote MCP endpoint to your client configuration:

```json
{
  "mcpServers": {
    "fastio": {
      "url": "https://mcp.fast.io/mcp/key",
      "headers": {
        "Authorization": "Bearer YOUR_FASTIO_API_KEY"
      }
    }
  }
}
```

This architecture delivers critical operational advantages for development teams:

* **Instant indexing without 24-hour delays.** When an engineer updates an architectural document or API schema, the file is parsed and indexed immediately, eliminating overnight reindexing delays.
* **Structured attribute filtering with Metadata Views.** Teams can configure [Fast.io Metadata Views](/product/document-data-extraction/) to automatically extract typed schema attributes, such as API versions, service names, and compliance tags, from unstructured documents. Assistants query these extracted fields via MCP to filter files before retrieving full text.
* **Cross-tool neutral ground.** Because Fast.io provides a standard MCP interface, the same indexed workspace can be queried simultaneously by Gemini Code Assist, Claude Code, Cursor, and custom Python agent scripts.
* **Complete version history and audit trails.** Every document modification retains full per-file version history alongside an append-only audit log, ensuring team members can trace exactly when reference specifications changed.

Monthly plans start with a trial of up to 30 days (credit card required); annual plans have no trial. Creating an account is free; doing real work requires an organization on a paid subscription. Plans are Starter at `$9.99/mo`, Business at `$49.99/mo`, and Enterprise at `$199.99/mo`. By decoupling knowledge storage from model context, teams bypass local token ceilings and maintain reliable access to enterprise documentation across every tool in their engineering stack.

## Frequently asked questions

### What are the daily quota limits for Gemini Code Assist?

Gemini Code Assist enforces daily quotas based on user tier and interface. Individual developers receive a daily allowance of 6,000 code completions and 240 chat requests. For agent mode and the Gemini CLI, Google enforces a combined daily cap of 1,500 requests per user for Standard Edition and 2,000 requests per user for Enterprise Edition. Google Cloud project environments support 6,000 code generation requests and 960 Cloud Assist chat requests daily per user.

### When do Gemini Code Assist daily quotas reset?

Gemini Code Assist daily quotas reset automatically at midnight Pacific Time (PT). Daily caps are aggregated across all model families used within an interface, meaning users cannot reset their quota mid-day by switching from Gemini Pro to Gemini Flash. Once the daily limit is reached, all subsequent requests through that interface fail until the reset timestamp passes.

### What is the difference between local codebase awareness and code customization?

Local codebase awareness operates inside the developer's IDE, analyzing open tabs and local workspace folders within a 1,000,000 token context window. Code customization is an enterprise feature that indexes 20,000 private repositories connected via Google Cloud Developer Connect. Code customization indexes repositories in a dedicated Google Cloud environment and automatically reindexes them every 24 hours.

### Can you increase Gemini Code Assist quotas in Google Cloud?

Project administrators can request quota increases for certain project-level metrics through the Google Cloud console under IAM & Admin > Quotas & System Limits by filtering for Gemini for Google Cloud. While project-level daily request caps can often be adjusted upon submitting a business justification, fixed system limits, such as the 2 requests per second concurrency ceiling, cannot be increased.

### How can you connect Gemini to large documentation sets without hitting context limits?

Rather than attaching large files directly to prompts, teams can store documents in an external intelligent workspace like Fast.io. By enabling Intelligence Mode on the workspace, documents are automatically indexed for hybrid search. Coding assistants connect through [Fast.io storage for agents](/storage-for-agents/) over Streamable HTTP to retrieve relevant text passages on demand, keeping prompt token consumption minimal while preserving source fidelity.

## Sources

- [Google Cloud Documentation: Quotas and limits for Gemini for Google Cloud](https://docs.cloud.google.com/gemini/docs/quotas) — Quotas for requests from Gemini Code Assist agent mode and Gemini CLI are combined across all model versions.
- [Google Cloud Documentation: Quotas and limits for Gemini for Google Cloud](https://docs.cloud.google.com/gemini/docs/quotas) — Gemini Code Assist on GitHub enforces a quota of at least 100 pull request reviews per day per installation.

## About Fast.io

Fast.io provides shared workspaces where people and AI agents work on the same files, with built-in semantic search and citation-backed chat over what they hold. Agents reach it through a remote MCP server at https://mcp.fast.io/mcp, a REST API at https://api.fast.io/current/, and a command line client published on npm as @vividengine/fastio-cli.
