# Gemini Image Upload Limits: File Size, Image Counts, and Multimodal Tokens

The Gemini image upload limit allows up to 3,600 images per prompt in developer APIs, but operational ceilings vary across environments. Inline base64 requests cap out at `20MB`, the temporary File API staging buffer accommodates assets up to `2GB` each, and consumer chat interfaces limit prompts to 10 files. Understanding fixed 258-token tiling mechanics, image formats, and context consumption prevents runtime errors and token exhaustion.

Source: https://fast.io/resources/gemini-image-upload-limit/
Author: [Tom Langridge](https://fast.io/authors/tom-langridge/)
Last reviewed: 2026-09-26

## What Are the Exact Google Gemini Image Upload Limits Across Environments?

Google Gemini developer APIs accept up to 3,600 images per request in Gemini 1.5 and 2.0, with inline payloads limited to 20MB and File API staging supporting up to 2GB per file, tokenized at a fixed cost of 258 tokens per image. While Gemini models provide expansive multimodal context windows spanning one million to two million tokens, developers frequently encounter upload rejections when confusing consumer chat constraints with developer API architecture.

Google Gemini models support a maximum of 3,600 image files per request in developer APIs. Inline image data limits total Gemini API request payload size to 20MB. Beyond these baseline figures, image handling splits across distinct technical layers, each governed by its own validation boundaries:

* Gemini Web App (`gemini.google.com`): The consumer browser interface limits standard user prompts to `10 files` per prompt. Non-video files carry an individual upload cap of `100MB`, while video files can reach up to `2GB`.
* Gemini API Inline Payloads: When transmitting base64-encoded image strings directly inside the `contents` or `input` array of a JSON request, Google's API gateway enforces a strict `20MB` ceiling across the entire HTTP request payload.
* Gemini API File API: For larger image datasets or high-resolution graphics, developers upload assets to Google's temporary Files API staging service. The File API accepts individual image assets up to `2GB` and provides up to `20GB` of ephemeral storage per project.
* Vertex AI Enterprise Ingestion: Enterprises deploying Gemini models on Google Cloud Vertex AI submit images via Google Cloud Storage buckets (`gs://` URIs) with support for up to `3,600 images` per request. Direct console browser uploads in the Vertex AI studio interface restrict individual files to `7MB`.

Understanding these operational boundaries prevents unexpected request failures during inference:

| Environment | Maximum Image File Size | Maximum Images per Prompt | Ingestion Method | Token Handling Mechanism | Common Error Code |
|---|---|---|---|---|---|
| Gemini Web Apps | 100MB per file | 10 files | Browser drag-and-drop | Automated context ingestion | Browser upload alert |
| Gemini API (Inline base64) | 20MB total request | Context window bound | Raw base64 string in JSON | 258 tokens per 768x768 tile | HTTP 413 Payload Too Large |
| Gemini API (File API) | 2GB per file | 3,600 images | Files API upload (`client.files.upload`) | 258 tokens per 768x768 tile | HTTP 400 INVALID_ARGUMENT |
| Vertex AI (Cloud Storage) | Object storage bound | 3,600 images | Cloud Storage URI (`gs://`) | 258 tokens per 768x768 tile | Bucket permission or extraction error |
| Vertex AI (Console UI) | 7MB per file | 1 file per upload | Web console file picker | Automated visual encoding | Console file size error |

### Supported Image Formats and MIME Type Requirements

The Gemini API accepts five primary raster and container image formats. Each submitted image must specify a supported standard MIME type in the payload:

* PNG (`image/png`): Ideal for diagrams, application screenshots, and text-heavy visual graphics where lossless compression prevents optical artifacting.
* JPEG (`image/jpeg`): Standard photographic format, providing compact file sizes for natural scenes and web graphics.
* WEBP (`image/webp`): Modern web format providing both lossy and lossless compression with smaller byte footprints than equivalent PNG or JPEG files.
* HEIC (`image/heic`): High Efficiency Image Container format used by Apple iOS and macOS devices.
* HEIF (`image/heif`): High Efficiency Image File format used across modern mobile camera systems.

Native support for HEIC and HEIF is a key technical differentiator for Google Gemini. Many alternative multimodal LLMs require client applications to decode and convert HEIC photos to JPEG before API transmission. Gemini parses HEIC and HEIF containers natively, removing client-side conversion overhead for mobile developer pipelines.

Unsupported formats include raw vector graphics such as SVG (`image/svg+xml`), animated GIF streams, and uncompressed TIFF files. Submitting an SVG file throws an `INVALID_ARGUMENT` error because Gemini's vision encoder requires discrete raster pixel grids rather than XML coordinate vectors. Applications handling SVG diagrams must rasterize vectors to high-resolution PNG format prior to API submission.

## How Does Gemini Calculate Multimodal Tokens and Image Tiling?

Unlike text tokens generated by byte-pair encoding algorithms, Google Gemini processes visual inputs through a vision transformer that standardizes images into uniform patch representations. Understanding token consumption is critical when submitting batches of images, as image tokens directly consume prompt context capacity.

In Gemini 1.5, Gemini 2.0, and Gemini 3 models, Google standardizes image inputs using a base unit of `258 tokens`:

* Base Resolution Floor: Any image where both dimensions are less than or equal to `384 pixels` consumes exactly `258 tokens` in the model context window.
* High-Resolution Tiling: Images exceeding `384 pixels` along either axis are divided into a dynamic grid of `768 x 768 pixel` tiles. Each individual tile consumes `258 tokens`.

Google calculates the exact number of tiles using the following mathematical sequence:

1. Determine the crop unit size: Compute `floor(min(width, height) / 1.5)`.
2. Calculate dimension tiles: Divide both image width and height by the computed crop unit size.
3. Compute total tiles: Multiply the horizontal and vertical tile counts together.
4. Calculate final image tokens: Multiply the total tile count by `258 tokens`.

Consider an engineering team submitting a standard landscape photograph with dimensions of `960 x 540 pixels`. The minimum dimension is `540 pixels`. Dividing `540` by `1.5` yields a crop unit size of `360 pixels`. Dividing width (`960 / 360 = 2.66`) and rounding up gives 3 horizontal tiles. Dividing height (`540 / 360 = 1.5`) gives 2 vertical tiles. The image divides into `3 * 2 = 6 tiles`, consuming `1,548 tokens` (`6 * 258`).

In Gemini 3 models, Google introduced the `media_resolution` parameter, providing explicit programmatic control over multimodal token allocation. Setting `media_resolution="low"` restricts the vision encoder to a single `258-token` representation regardless of input pixel dimensions. This optimization lowers latency and reduces token consumption for high-throughput image classification tasks where fine text reading is unnecessary.

### Context Window Saturation in Multi-Image Pipelines

Google's documented limit of `3,600 images` per request represents a structural API ceiling, but context window capacity governs actual practical throughput. Calculating batch token consumption ensures that multi-image prompts fit within available model context windows:

* Small Thumbnails (`384 x 384` pixels): At `258 tokens` per image, submitting `3,600 images` consumes `928,800 tokens`. This fits comfortably within the standard one-million-token context window of Gemini 1.5 Flash and Gemini 2.0 Flash, leaving approximately `71,200 tokens` for instructions and response generation.
* Multi-Tile High-Resolution Images: If full-resolution mobile photos average `6 tiles` (`1,548 tokens` each), a batch of `600 images` consumes `928,800 tokens`. Submitting `3,600` such images would require `5,572,800 tokens`, far exceeding a one-million or two-million token context window.

Prompt structure also impacts visual comprehension. Google's documentation advises developers to place the text prompt instruction before the image array when submitting multimodal inputs. Placing text instructions prior to binary image references improves instruction adherence during complex visual question answering and image extraction tasks.

## When to Choose Inline Payloads Versus the Google File API

When building applications on the Gemini API, developers must choose between passing images inline as base64-encoded strings or uploading files to Google's temporary Files API staging service. Selecting the wrong pipeline causes immediate transport errors or unnecessary network latency.

Inline data passing embeds binary image data directly into the JSON request body using the `inline_data` field. While convenient for single-turn scripts, base64 encoding substantially inflates raw binary file size. A `16MB` JPEG image on disk expands to more than `21MB` of base64 text characters when serialized into JSON. This exceeds Google's `20MB` inline request ceiling and triggers an immediate `HTTP 413 Request Entity Too Large` error from the API gateway.

The Google File API (`client.files.upload`) solves payload bloat by decoupling file transport from model inference. When using the File API:

1. The application transmits the raw binary file to Google's staging endpoint.
2. The File API stores the image in an ephemeral cloud buffer and returns an active `file.uri`.
3. The application passes the lightweight `file.uri` inside the `contents` parameter of `client.models.generate_content`.
4. The inference engine pulls the image directly from the internal staging buffer during execution.

Files uploaded to the File API are retained for `48 hours` before automated deletion. The File API is entirely free of separate storage fees, but project storage is capped at a `20GB` staging ceiling across all active assets. File API uploads are write-only staging buffers: client applications cannot download original files back from the API, meaning primary storage must remain in dedicated workspaces.

### Implementing Multi-Image File API Ingestion in Python

The following Python implementation demonstrates staging multiple images through Google's official SDK, verifying active status, and submitting them within a single multimodal prompt:

```python
import os
from google import genai
from google.genai.errors import APIError

def analyze_image_batch(image_paths: list[str], prompt_text: str) -> str:
    client = genai.Client()
    staged_files = []
    try:
        for path in image_paths:
            file_size_mb = os.path.getsize(path) / (1024 * 1024)
            if file_size_mb > 2000.0:
                raise ValueError(f"File {path} exceeds staging threshold.")
            print(f"Staging {path} ({file_size_mb:.2f} MB)...")
            staged_file = client.files.upload(
                file=path,
                config={"mime_type": "image/jpeg"}
            )
            staged_files.append(staged_file)
        request_contents = [prompt_text] + staged_files
        response = client.models.generate_content(
            model="gemini-2.0-flash",
            contents=request_contents
        )
        return response.text
    except APIError as exc:
        if exc.code == 413:
            print("Payload exceeded inline boundaries. Route through Files API.")
        elif exc.code == 400:
            print(f"Ingestion rejected by Gemini API: {exc.message}")
        raise
    finally:
        for staged_file in staged_files:
            try:
                client.files.delete(name=staged_file.name)
            except Exception:
                pass
```

Using the Files API keeps network request payloads lightweight and prevents client-side serialization timeouts when handling large collections of high-resolution images.

## Why Do Image Uploads Fail and How Can You Fix Them?

Production vision pipelines encounter several recurring failure modes when processing diverse visual inputs. Understanding how Gemini handles metadata, orientation, and resolution prevents silent extraction degradation and runtime API crashes.

The most frequent image ingestion failure modes include:

* EXIF Orientation Flags Ignored: Smartphone cameras record physical orientation inside EXIF metadata tags rather than rotating the underlying pixel matrix. While web browsers automatically read EXIF tags and display photos right-side up, vision transformers process raw raster arrays. When an image with an unapplied rotation tag enters Gemini, the model receives sideways or upside-down pixels, causing severe optical character recognition degradation and inaccurate spatial reasoning.
* Extreme Aspect Ratio Truncation: Architectural floor plans, panoramic photos, and full-page mobile screenshots with aspect ratios exceeding `16:1` trigger aggressive downscaling. During tiling, Google's scaling pipeline compresses narrow dimensions, rendering small typography and line annotations unreadable.
* Ephemeral File Expiration in Long Conversations: Multi-turn agent conversations that reference a `file.uri` over multiple days break when Google's `48-hour` File API retention window expires. Once expired, subsequent prompts referencing that URI fail with an `HTTP 404 File Not Found` error.
* Color Space Conversion Defects: Images encoded in non-standard CMYK print color profiles or 16-bit uncompressed RAW formats trigger decoding errors in Google's vision parser. Images should always be converted to standard sRGB 8-bit color spaces before transmission.

### Normalizing Image Orientation and Dimensions with Pillow

To avoid orientation errors and excessive tiling costs, engineering teams should preprocess images locally before API dispatch. The following Python routine inspects EXIF metadata, transposes rotated pixels, standardizes color profiles to sRGB, and validates file boundaries:

```python
import os
from PIL import Image, ImageOps

def preprocess_image_for_gemini(source_path: str, target_path: str, max_dimension: int = 2048) -> str:
    with Image.open(source_path) as img:
        img = ImageOps.exif_transpose(img)
        if img.mode in ("RGBA", "P"):
            img = img.convert("RGB")
        elif img.mode == "CMYK":
            img = img.convert("RGB")
        width, height = img.size
        if max(width, height) > max_dimension:
            scale_ratio = max_dimension / max(width, height)
            new_size = (int(width * scale_ratio), int(height * scale_ratio))
            img = img.resize(new_size, Image.Resampling.LANCZOS)
            print(f"Rescaled from {width}x{height} to {new_size[0]}x{new_size[1]}")
        img.save(target_path, format="JPEG", quality=88, optimize=True)
    file_size_mb = os.path.getsize(target_path) / (1024 * 1024)
    print(f"Preprocessed image saved: {target_path} ({file_size_mb:.2f} MB)")
    return target_path

preprocess_image_for_gemini("mobile_photo.heic", "normalized_photo.jpg")
```

Normalizing orientation and downsampling redundant resolutions protects vision pipelines from OCR errors and controls token expenditure across large image workloads.

## Scaling Multi-Image Pipelines with Intelligent Workspaces

When engineering teams move beyond simple single-image prompts to autonomous multi-agent pipelines, attaching hundreds of raw image files directly to model prompts quickly becomes unsustainable. Passing large volumes of image binaries directly into an LLM context window exhausts token limits, increases network latency, and multiplies recurring API costs.

Engineering teams typically evaluate conventional storage options before discovering their operational limitations:

* Local File Storage: Storing images on local server disks works for local evaluation scripts, but isolates files on a single machine. Other agents and human teammates cannot access the files, and version history is absent.
* Raw Cloud Object Storage (Amazon S3 or Google Cloud Storage): Object storage handles petabyte-scale image libraries, but functions as passive bit storage. S3 provides no semantic search, no visual indexing, and no structured querying. Teams must build and maintain custom embedding pipelines, vector databases, and retrieval servers.
* Google Drive Folders: Common for human desktop sharing, but Drive's API rate limits, complex authentication, and background sync timeouts create friction for automated agent loops.

Fast.io provides an intelligent workspace platform built specifically for agentic teams collaborating on complex file collections. Instead of repeatedly stuffing raw image files into model prompts, teams store visual assets in shared, organization-owned workspaces in [Fast.io workspaces for AI agents](/storage-for-agents/).

Files can be uploaded directly or imported from existing repositories. Fast.io supports cloud sync for Dropbox, Box, and OneDrive; Google Drive imports today, with sync coming soon. Once files arrive in a workspace, enabling Intelligence Mode automatically indexes documents and visual assets for hybrid search, combining exact full-text matching, semantic meaning-based search, and metadata value filtering.

For workflows requiring structured data extraction from visual documents, receipts, and photos, Fast.io provides [Metadata Views](/product/document-data-extraction/). Metadata Views turn unstructured file collections into queryable, typed database tables without requiring manual OCR template configuration. Autonomous agents can create views, trigger extraction, and filter extracted attributes directly over MCP.

AI assistants connecting through the remote Fast.io Model Context Protocol (MCP) server at `https://mcp.fast.io/mcp/tools` query the workspace intelligence layer directly. When an assistant needs information from a multi-thousand-image catalog, it retrieves only the relevant assets or extracted attributes, completely bypassing vendor prompt upload ceilings and recurring visual token overhead.

Every organization starts with a 30-day free trial, which requires a credit card. Subscriptions are available on the Starter plan at `$29/mo`, Business at `$99/mo`, and Enterprise at `$299/mo` on [Fast.io pricing](/pricing/), providing team workspaces with version history and a detailed activity log for human-agent collaboration.

## Frequently asked questions

### How many images can you upload to Google Gemini?

Google Gemini models support a maximum of 3,600 image files per request in developer APIs across Gemini 1.5, Gemini 2.0, and Gemini 3 models, provided the combined token count fits inside the model context window. In consumer Gemini web apps, prompts are restricted to a maximum of 10 files.

### What is the maximum image file size for Gemini?

Maximum file size depends on the upload method. Inline API requests are capped at 20MB for the entire request payload. The Google File API accepts individual image assets up to 2GB for staging, while the consumer Gemini web app limits non-video files to 100MB each.

### How many tokens does an image take in Gemini?

Gemini calculates image tokens using a base rate of 258 tokens for images within 384x384 pixels. Larger images are divided into 768x768 pixel tiles based on the crop unit formula floor(min(width, height) / 1.5), with each tile consuming 258 tokens.

### What image formats does Google Gemini support?

The Gemini API supports five image MIME types: PNG (image/png), JPEG (image/jpeg), WEBP (image/webp), HEIC (image/heic), and HEIF (image/heif). Unsupported formats such as SVG or TIFF must be converted to PNG or JPEG before ingestion.

### Why did my image upload fail with an HTTP 413 error in the Gemini API?

An HTTP 413 Payload Too Large error occurs when an inline base64 image request exceeds Google's 20MB request body ceiling. Base64 encoding expands raw image files by approximately one-third, meaning files larger than 15MB on disk should always be uploaded using the Google File API instead.

### Does Google Gemini support HEIC and HEIF photos from smartphones?

Yes, Google Gemini provides native support for HEIC and HEIF image containers in both the Gemini web app and the Gemini developer API. Developers do not need to convert smartphone photos to JPEG before submitting them to Gemini models.

## Sources

- [Google AI for Developers: Image Understanding](https://ai.google.dev/gemini-api/docs/image-understanding): Google Gemini models support a maximum of 3,600 image files per request in developer APIs.
- [Google AI for Developers: Image Understanding](https://ai.google.dev/gemini-api/docs/image-understanding): Inline image data limits total Gemini API request payload size to 20MB.

## About Fast.io

Fast.io provides shared workspaces where people and AI agents work on the same files, with built-in semantic search and citation-backed chat over what they hold. Agents reach it through a remote MCP server, a REST API at https://api.fast.io/current/, and a command line client published on npm as @vividengine/fastio-cli. MCP setup is at https://mcp.fast.io/docs: Claude and most MCP clients connect to https://mcp.fast.io/mcp/tools, ChatGPT to https://mcp.fast.io/mcp/operations, and coding agents to https://mcp.fast.io/mcp/code.
