# OpenAI File Size Limits: API Endpoints, Whisper, and Workarounds

OpenAI file size limits vary by endpoint, from 512 MB on the Files API to 25 MB on Whisper speech-to-text. While individual files have hard caps, token ceilings and storage quotas create additional operational bottlenecks. Understanding these constraints helps engineering teams structure preprocessing pipelines or offload large corpora to external searchable workspaces.

Source: https://fast.io/resources/openai-file-size-limit/
Author: [Tom Langridge](https://fast.io/authors/tom-langridge/)
Last reviewed: 2026-09-25

## What Are the OpenAI File Size Limits Across API Endpoints?

As of September 2026, individual files uploaded to the OpenAI Files API can be up to 512 MB, with project storage capped at 2.5 TB. Meanwhile, the OpenAI Speech-to-Text Transcriptions API accepts audio files up to 25 MB per transcription request. These operational thresholds define how developers must prepare data before passing it to language models, audio transcribers, or multimodal processors.

OpenAI file size limits are endpoint-specific constraints that govern the maximum file size (such as `512 MB` for the Files API, `25 MB` for Whisper, and `20 MB` for Vision) that developers can upload to OpenAI's infrastructure.

Many developers assume that OpenAI maintains a single global file upload ceiling across its platform. In practice, limits differ across REST routes, ingestion pipelines, and user interfaces. Uploading a training set through the Files API operates under different technical boundaries than attaching an audio interview to Whisper or indexing a technical manual inside an Assistants API vector store.

The following reference table outlines the documented file size limits, accepted formats, and primary constraints across OpenAI services:

| Endpoint or Service | Documented File Limit | Supported Formats | Primary Operating Constraint | Date Checked |
|---|---|---|---|---|
| OpenAI Files API (`/files`) | 512 MB per file | JSONL, PDF, TXT, DOCX, audio, images | 2.5 TB project storage ceiling; 1,000 req/min | September 2026 |
| Whisper Audio API (`/audio/transcriptions`) | 25 MB per request | mp3, mp4, mpeg, mpga, m4a, wav, webm | Entire HTTP multipart request body; no duration cap | September 2026 |
| Assistants API Vector Stores | 512 MB per file | PDF, DOCX, TXT, HTML, JSON, MD | Up to 5,000,000 tokens per file; 10,000 files per store | September 2026 |
| OpenAI Batch API (`/batches`) | 200 MB per file | `.jsonl` only | 50,000 requests per batch file; 2,000 batches/hour | September 2026 |
| Vision and Multimodal Inputs | 20 MB per image | PNG, JPEG, WEBP, non-animated GIF | 512 MB total payload per request; 30,000 patch limit | September 2026 |
| ChatGPT Web Interface | 512 MB per file | Office docs, PDFs, data, text | Spreadsheets capped near 50 MB; 25 GB user storage | September 2026 |

Understanding these boundaries prevents runtime exceptions and structural data loss. When file payloads exceed these limits, requests fail immediately with HTTP 413 Payload Too Large or service-level quota errors.

## What Is the Single-File Limit on the OpenAI Files API and ChatGPT?

The OpenAI Files API (`POST /files`) provides persistent asset storage for fine-tuning, batch processing, and assistant retrieval. Under the official specification, individual files uploaded to the OpenAI Files API can be up to 512 MB while project storage reaches 2.5 TB.

While `512 MB` represents the maximum transport size for a single file upload, downstream consumers enforce stricter semantic constraints depending on file structure:

### 1. Document Token Ceilings
For text documents, PDFs, and Markdown files intended for the Assistants API or code interpreter, OpenAI enforces a hard ceiling of `2,000,000` tokens per file. A dense text document containing raw database logs or compressed JSON exports will easily surpass two million tokens long before reaching the `512 MB` physical boundary. When this occurs, the API rejects file processing during extraction.

### 2. Tabular Data and Spreadsheets
Structured data files such as CSV, TSV, and XLSX spreadsheets present a unique bottleneck. While the Files API accepts CSV files up to the full `512 MB` boundary, ChatGPT and code interpreter environments routinely fail on tabular sheets larger than approximately `50 MB`. Processing spreadsheets requires parsing individual rows into memory buffers. A CSV file containing millions of narrow numeric rows can exhaust memory allocation during analysis, triggering execution timeouts.

### 3. Storage Quotas and Rate Limits
Storage limits operate across multiple administrative tiers:

* **Project Storage Quotas**: Modern OpenAI API accounts group resources by project, with default capacity set to `2.5 TB` across all stored files.
* **Organization Storage Quotas**: Legacy organization accounts or specific tiered plans enforce a `100 GB` default storage cap across all uploaded files.
* **Upload Rate Limits**: The `/files` endpoint enforces a rate limit of `1,000` upload requests per minute per authenticated user, preventing uncontrolled parallel flooding.
* **Batch File Caps**: The Batch API accepts `.jsonl` input files up to `200 MB`, with a cap of `50,000` individual API requests per file.

When an upload exceeds either the single-file threshold or the project capacity limit, the API returns a structured error object indicating that the payload exceeds permitted parameters.

## What Is the File Size Limit for OpenAI Whisper Speech-to-Text?

Audio transcription through OpenAI models runs under much tighter limits than general document storage. The OpenAI Speech-to-Text Transcriptions API accepts audio files up to 25 MB per transcription request.

This `25 MB` limit applies to the complete HTTP request body sent to `POST /audio/transcriptions` or `POST /audio/translations`. If your audio file is near the boundary, adding multipart form metadata or verbose model parameters can push the total request size over `25 MB`, causing an immediate failure.

The audio endpoint officially supports nine formats: `flac`, `mp3`, `mp4`, `mpeg`, `mpga`, `m4a`, `ogg`, `wav`, and `webm`.

### Audio Duration Versus File Size
OpenAI does not impose a static cap on the duration of an audio file. The constraint is purely byte size. Consequently, the length of audio you can transcribe in a single request depends entirely on compression and bitrate:

* **Uncompressed WAV (16-bit, 44.1 kHz, Stereo)**: Reaches the `25 MB` ceiling in roughly two and a half minutes of playback.
* **Standard MP3 (128 kbps, Stereo)**: Reaches `25 MB` at approximately twenty-six minutes of recorded audio.
* **Voice-Optimized MP3 (64 kbps, Mono)**: Extends `25 MB` capacity to roughly fifty-two minutes.
* **Compressed Opus or OGG (32 kbps, Mono)**: Fits over one hundred minutes of speech within a single `25 MB` payload without noticeable degradation in transcription accuracy.

### Splitting Long Recordings with FFmpeg

When working with multi-hour conference calls, depositions, or podcast archives, audio files inevitably exceed `25 MB`. Developers must slice the recording into chunks prior to upload.

Using FFmpeg, you can segment a long recording into twenty-minute chunks formatted as lightweight mono MP3 files:

```bash
ffmpeg -i long_recording.wav -vn -ar 16000 -ac 1 -b:a 48k -f segment -segment_time 1200 chunk_%03d.mp3
```

For cleaner transcription boundaries, avoid cutting strictly on fixed time intervals, which can slice spoken words in half. Instead, use silence detection filters to segment files during natural conversational pauses:

```bash
ffmpeg -i long_recording.mp3 -af silencedetect=noise=-30dB:d=0.5 -f null -
```

After identifying silence timestamps, slice the source file along those boundaries, transcribe each segment independently, and concatenate the resulting transcript text while adjusting timestamp offsets.

## How Do Assistants API Vector Stores Enforce File and Token Ceilings?

The Assistants API uses vector stores to power its `file_search` tool, allowing models to query external documentation through semantic retrieval. While individual files uploaded to a vector store adhere to the standard `512 MB` file ceiling, vector indexing introduces distinct operational constraints.

Each vector store supports up to `10,000` files. During vector ingestion, OpenAI extracts text, splits content into passages, and generates embeddings. The ingestion engine enforces a maximum limit of `5,000,000` tokens per file.

When attaching files to vector stores, developers must account for several structural rules:

### 1. File Batching Limits
Attaching thousands of files individually via repeated API calls will trigger rate limit exceptions. OpenAI provides a dedicated file batch endpoint (`/vector_stores/{vector_store_id}/file_batches`) designed to process file collections asynchronously. Organization-level attachment limits cap ingestion throughput at `2,000` attached files per minute.

### 2. Supported Formats for Parsing
Vector stores accept standard document types including plain text (`.txt`), Markdown (`.md`), PDF (`.pdf`), Microsoft Word (`.docx`), HTML (`.html`), and JSON (`.json`). Binary executables, spreadsheets with complex formulas, and image-heavy archives cannot be processed directly by `file_search`.

### 3. Ingestion Latency and Failure Modes
Processing massive documents through OpenAI's internal parser introduces variable latency:

* **Parsing Timeouts**: Complex PDF files containing nested vector graphics, large embedded tables, or non-standard font encodings frequently fail during parsing. The vector store marks the file status as `failed` with a corresponding error code.
* **Token Overflows**: If a structured JSON file or dense technical text expands past `5,000,000` tokens during parsing, the indexing pipeline halts and rejects the asset.
* **Sync Delays**: Ingestion is asynchronous. Assistants cannot query new vector store files until the batch indexing status transitions from `in_progress` to `completed`.

Managing these parsing behaviors requires continuous error checking and custom retry logic to ensure that failed document chunks do not leave blind spots in the agent's knowledge retrieval.

## How Do You Bypass OpenAI File Size Limits for Massive Corpora?

When application requirements exceed native file upload ceilings, developers employ several architectural workarounds. Managing multi-gigabyte codebases, multi-hour media libraries, or terabyte-scale enterprise document archives requires decoupling raw storage from model invocation.

### Strategy 1: Client-Side Partitioning and Batching
For datasets exceeding the `512 MB` file limit, the traditional workaround is programmatic client-side partitioning. Large JSONL files used for fine-tuning or batch completions can be split into smaller segments using shell utilities or Python scripts:

```python
def split_jsonl_file(source_path, max_bytes=180 * 1024 * 1024):
    part_num = 1
    current_size = 0
    out_file = open(f"part_{part_num}.jsonl", "w", encoding="utf-8")
    with open(source_path, "r", encoding="utf-8") as src:
        for line in src:
            line_bytes = len(line.encode("utf-8"))
            if current_size + line_bytes > max_bytes:
                out_file.close()
                part_num += 1
                current_size = 0
                out_file = open(f"part_{part_num}.jsonl", "w", encoding="utf-8")
            out_file.write(line)
            current_size += line_bytes
    out_file.close()
```

This script ensures that each generated slice remains safely below the `200 MB` Batch API ceiling, allowing automated batch jobs to submit sequential requests without encountering file size exceptions.

### Strategy 2: Pre-Processing and Audio Downsampling
Rather than uploading raw 24-bit studio recordings to Whisper, standardize your ingestion pipeline to downsample audio to 16 kHz mono Opus or MP3 before sending requests. Downsampling voice recordings reduces payload volume substantially without sacrificing word error rates, allowing one-hour meetings to comfortably fit inside the `25 MB` limit.

### Strategy 3: Multi-Agent Storage Coordination
When teams build systems where multiple autonomous agents collaborate, relying on individual OpenAI file uploads causes immediate friction. If Agent A extracts data from a large document and Agent B needs to verify the findings, uploading the same file twice wastes bandwidth, consumes project storage quotas, and duplicates indexing overhead.

Traditional consumer cloud drives such as Dropbox, Box, or Google Drive resolve file storage for humans, but they present significant drawbacks for AI agents. They lack native Model Context Protocol interfaces, enforce strict user-centric permissions, and require complex synchronization logic that introduces latency into agent loops.

## When Should You Index Corpora in Fastio Instead of Direct OpenAI Uploads?

As document collections expand into gigabytes and terabytes, uploading entire files directly to language model APIs becomes inefficient. The practical solution is to separate physical document storage and indexing from the language model itself.

Fastio provides an intelligent cloud workspace designed specifically for agentic teams and human collaboration. Rather than pushing multi-hundred-megabyte files across API endpoints, your team stores documents in a shared Fastio workspace. Fastio leaves OpenAI's native upload limits exactly where they are; what it adds is a searchable, persistent home for the files that do not fit into direct API payloads.

### Universal Cloud Import and Continuous Indexing
Getting large corpora into Fastio requires no local disk transfers. With cloud import, teams can pull entire directory trees from Dropbox, Box, OneDrive, and Google Drive (with cloud sync available for Dropbox, Box, and OneDrive, and Google Drive sync coming soon). Public web files can also be ingested directly via URL import.

Once files land in a Fastio workspace, Intelligence Mode indexes the content automatically. There is no need to stand up separate vector databases, manage custom chunking scripts, or monitor token limits. Fastio provides hybrid search, combining exact full-text keyword retrieval with semantic meaning-based search across documents, code files, spreadsheets, and PDFs.

### Structured Data Extraction with Metadata Views

When dealing with unstructured documents at scale, extracting specific values through chat prompts consumes immense context. Fastio solves this with [Metadata Views](/product/document-data-extraction/).

Metadata Views transform workspace documents into a structured, queryable database:

* **Schema Definition in Plain English**: Describe the desired data fields in natural language (such as agreement dates, monetary values, invoice numbers, or compliance terms).
* **Automated Typing**: The system automatically designs typed schemas supporting Text, Integer, Decimal, Boolean, URL, JSON, and Date & Time fields.
* **Universal Document Compatibility**: Works across complex PDFs, scanned forms, Word files, spreadsheets, and presentation decks.
* **Programmatic Querying**: AI agents can create schemas, trigger field extraction, and retrieve tabular records via MCP without downloading or reading the underlying files.

### Connecting AI Agents via Fastio MCP Server
Autonomous agents interact with Fastio workspaces through the official remote Model Context Protocol server. The Fastio MCP server runs remotely over Streamable HTTP at `https://mcp.fast.io/mcp` (or `https://mcp.fast.io/mcp/key` when passing an authorization header), eliminating local runtime dependencies or complex package installations.

Add Fastio to your agent configuration:

```json
{
  "mcpServers": {
    "fastio": {
      "type": "http",
      "url": "https://mcp.fast.io/mcp/key",
      "headers": {
        "Authorization": "Bearer YOUR_FASTIO_API_KEY"
      }
    }
  }
}
```

When an agent needs information from a massive technical manual or an extensive contract repository, it calls Fastio MCP search tools. The workspace searches indexed content and returns precise text excerpts along with source citations and page numbers. The agent receives the exact grounding context it needs in a few hundred tokens, completely bypassing OpenAI's `512 MB` file upload limit and saving context capacity.

### Team Collaboration, Version History, and Governance
Fastio bridges autonomous agent execution with team oversight:

* **Per-File Version History**: Every update made by human editors or AI agents is tracked in full version history. Previous revisions can be inspected or restored at any point.
* **Append-Only Audit Log**: Every file upload, metadata extraction, search query, and download is preserved in an immutable audit record for compliance.
* **Collaborative Notes**: Humans and AI agents can co-edit project documentation, meeting summaries, and technical specs using Agent Intents to coordinate writing slots without conflicts.
* **Transparent Workspace Pricing**: Monthly plans start with a trial of up to 30 days (credit card required). Ongoing subscriptions are organized into Starter, Business, and Enterprise tiers. Explore technical agent architecture at [Fast.io Storage for Agents](/storage-for-agents/), check onboarding documentation at [fast.io/llms.txt](https://fast.io/llms.txt), or review plan options on the [Fast.io Pricing page](/pricing/).

## Frequently asked questions

### What is the maximum file size you can upload to OpenAI?

The maximum file size for individual file uploads to the OpenAI Files API and ChatGPT is `512 MB`. However, specific tools enforce lower thresholds. The Whisper Speech-to-Text API limits audio uploads to `25 MB` per request, the Batch API restricts JSONL files to `200 MB`, and multimodal image uploads are capped at `20 MB` per image. Text documents in the Assistants API are also constrained by a `2,000,000` token limit.

### What is the file size limit for OpenAI's Whisper API?

The Whisper API enforces a strict `25 MB` file size limit per request for audio transcriptions and translations. Because the limit applies to the full HTTP request body, developers working with longer recordings must compress audio to lower bitrates, convert uncompressed files to mono MP3 or Opus formats, or slice recordings into chunks under `25 MB` using FFmpeg before uploading.

### How do you bypass the 512 MB individual file limit on the OpenAI Files API for large datasets?

To bypass the `512 MB` limit for large datasets, developers can partition datasets into smaller chunked files using client-side scripts, downsample high-bitrate media, or decouple storage from the model. Instead of uploading entire multi-gigabyte corpora directly to OpenAI, teams store documents in an external workspace like Fastio, index content automatically, and query relevant excerpts via the Model Context Protocol.

### What happens when an uploaded file exceeds the token limit in the Assistants API?

If a document uploaded to the Assistants API exceeds the `2,000,000` token limit, OpenAI's document parser rejects the file during extraction, causing the ingestion process to fail. Even if the file is well under the `512 MB` physical ceiling, dense text files, raw code repositories, or large JSON exports must be split into smaller topical documents before attachment.

### Why do CSV and spreadsheet files fail below maximum file limits in ChatGPT?

While the Files API accepts files up to `512 MB`, tabular datasets such as CSV and XLSX files often fail near `50 MB` in ChatGPT and code interpreter environments. Spreadsheets require row-by-row memory allocation during Python parsing. Large files with complex headers or millions of rows can exhaust available container memory and trigger execution timeouts.

### How does external MCP indexing compare to OpenAI vector store uploads?

OpenAI vector stores parse and store files directly inside OpenAI infrastructure, subject to a `512 MB` per-file limit, a `5,000,000` token per-file ceiling, and an organization ingestion rate of `2,000` files per minute. External MCP indexing through Fastio stores files in a shared team workspace, indexes documents automatically for hybrid search, and allows agents to retrieve cited passages dynamically, preserving model context.

## Sources

- [OpenAI OpenAPI Specification: Files Endpoint](https://raw.githubusercontent.com/openai/openai-openapi/master/openapi.yaml) — Individual files uploaded to the OpenAI Files API can be up to 512 MB, with project storage capped at 2.5 TB.
- [OpenAI Developer Documentation: Speech to Text Guide](https://developers.openai.com/api/docs/guides/speech-to-text) — The OpenAI Speech-to-Text Transcriptions API accepts audio files up to 25 MB.

## About Fast.io

Fast.io provides shared workspaces where people and AI agents work on the same files, with built-in semantic search and citation-backed chat over what they hold. Agents reach it through a remote MCP server at https://mcp.fast.io/mcp, a REST API at https://api.fast.io/current/, and a command line client published on npm as @vividengine/fastio-cli.
