# Whisper File Size Limit: The 25 MB Cap, Chunking, and Cloud Workspaces

The OpenAI Whisper API enforces a strict 25 MB file size limit for audio and video uploads across all supported file formats. Transcribing long recordings requires reducing audio bitrates, chunking files into smaller segments, or querying transcripts in persistent cloud workspaces. This guide explains Whisper limits, practical FFmpeg chunking commands, and enterprise storage patterns for large audio archives.

Source: https://fast.io/resources/whisper-file-size-limit/
Author: [Tom Langridge](https://fast.io/authors/tom-langridge/)
Last reviewed: 2026-09-19

## What is the Whisper file size limit?

The OpenAI Whisper API enforces a strict 25 MB file size limit for audio and video uploads across all supported file formats. Whether you send a raw WAV recording, an MP3 podcast, or an MP4 video clip, any single HTTP request with a payload exceeding the 25 MB Whisper API upload limit is rejected immediately by OpenAI servers.

When an upload exceeds the 25 MB OpenAI Whisper file limit, the API returns an HTTP 413 status code with the error message:

```json
{
  "error": {
    "message": "Maximum content size limit exceeded",
    "type": "server_error",
    "param": null,
    "code": null
  }
}
```

This 25 MB file upload limit exists because OpenAI's Whisper API processes transcriptions through stateless HTTP multipart form requests. A single unified upload endpoint must guard against hanging TCP connections, buffer exhaustion, and unbounded memory consumption across thousands of concurrent API requests.

### Supported formats and raw file sizes

As documented in OpenAI's official speech-to-text guide, the Whisper API accepts seven file formats:

- `mp3` (MPEG-1 Audio Layer III)
- `mp4` (MPEG-4 Part 14, audio or video)
- `mpeg` (MPEG-1 or MPEG-2 audio or video)
- `mpga` (MPEG-1 Layer 1/2 audio stream)
- `m4a` (MPEG-4 audio container, typically AAC or ALAC)
- `wav` (Waveform Audio File Format, uncompressed PCM)
- `webm` (WebM media container, Opus or Vorbis audio)

The central friction point for developers and data teams is that audio recordings rapidly exceed the 25 MB Whisper API upload limit. An uncompressed one-hour stereo interview recorded in WAV format consumes hundreds of megabytes of disk space. Even when recorded in mono, raw uncompressed speech demands substantial storage that quickly breaches the request ceiling. Long-form corporate recordings, such as all-hands meetings, customer research calls, quarterly earnings conferences, and legal depositions, regularly produce large media files that cannot be uploaded without prior processing.

### Three ways to transcribe files larger than 25 MB

Because OpenAI does not offer an enterprise quota exception or a paid toggle to raise the 25 MB single-file ceiling, teams must adopt one of three practical strategies:

1. **Downsample and compress:** Convert stereo audio to mono, downsample the sample rate to 16 kHz, and compress into a low-bitrate format such as 64 kbps MP3 or Opus. This reduces file size substantially and allows extended speech to fit beneath the 25 MB OpenAI Whisper upload limit.
2. **Split audio into discrete chunks:** Slice long recordings into time-based segments using [FFmpeg](https://ffmpeg.org), transcribe each segment individually through the API, and stitch the text back together while preserving boundary context.
3. **Run Whisper locally or in hybrid cloud workspaces:** Use open-source implementations such as `faster-whisper` on dedicated compute to remove file size limits completely, while storing raw media, transcripts, and metadata in unified, queryable cloud workspaces on [Fast.io](/product/workspaces/).

### Comparing transcription architectures

The following table compares file constraints, deployment requirements, and operational characteristics across speech recognition approaches as of September 2026:

| Platform or Model | Single-File Limit | Audio Duration Limit | Key Operational Constraint | Primary Deployment Model | Checked Date |
| :--- | :--- | :--- | :--- | :--- | :--- |
| **OpenAI Whisper API (`whisper-1`)** | 25 MB | No duration limit (bounded by file size) | Strict HTTP 413 error if payload exceeds 25 MB | Hosted Cloud API ($0.006 per minute) | September 2026 |
| **OpenAI `gpt-transcribe`** | 25 MB | No duration limit (bounded by file size) | Same 25 MB payload limit, supports keywords hints | Hosted Cloud API ($0.0045 per minute) | September 2026 |
| **Local Whisper (`whisper` / `faster-whisper`)** | Unlimited | Unlimited | Bounded by local RAM, VRAM, and processing time | Self-Hosted / Dedicated GPU | September 2026 |
| **Deepgram Nova-3 API** | 2 GB | Up to 480 minutes per file | Supports direct URLs and large batch uploads | Hosted Cloud API | September 2026 |
| **AssemblyAI Speech-to-Text** | 5 GB | Up to 10 hours per file | Requires asynchronous upload URL pattern | Hosted Cloud API | September 2026 |

## How to compress audio files before Whisper upload

Before writing complex file-splitting pipelines, the fastest way to work around the 25 MB Whisper API upload limit is to strip unnecessary data from the audio file. Most raw recordings contain multichannel sound and frequencies that speech recognition models ignore.

### Understanding Whisper's acoustic input requirements

Whisper does not benefit from high-resolution studio audio. Internally, the Whisper neural network resamples all incoming audio to 16,000 Hz (16 kHz) mono. The neural network computes an 80-channel log-Mel spectrogram using 25-millisecond windows with a 10-millisecond stride.

Uploading a high-resolution stereo FLAC file or a 24-bit uncompressed WAV does not improve transcription accuracy. It only exhausts the 25 MB OpenAI Whisper file limit within minutes. Stripping extra audio channels, downsampling to 16 kHz, and encoding with a modern codec allows you to pack hours of intelligible voice into a fraction of the space.

### Bitrate, duration, and file size capacity

File size is a direct mathematical product of bitrate and duration. The table below illustrates how much speech audio fits into the 25 MB OpenAI Whisper file limit across standard audio formats:

| Audio Format and Encoding | Typical Bitrate | Approximate Duration per 25 MB | Practical Use Case |
| :--- | :--- | :--- | :--- |
| **Uncompressed WAV (16-bit, 44.1 kHz mono)** | ~705 kbps | ~4 to 5 minutes | Raw studio recording, requires compression |
| **Standard MP3 (Stereo)** | 128 kbps | ~26 minutes | Standard music or podcast audio |
| **Optimized Voice MP3 (Mono)** | 64 kbps | ~52 minutes | Ideal balance of voice clarity and compact size |
| **AAC / M4A (Mono)** | 32 kbps | ~104 minutes | High-efficiency voice encoding for long meetings |
| **Opus / WebM (Mono)** | 24 kbps | ~138 minutes | Maximum recording duration within the upload cap |

For pure speech, 64 kbps MP3 or 32 kbps AAC delivers crisp, intelligible voice recognition. That means a standard 45-minute team standup or customer interview can easily fit into a single API request once compressed properly.

### Practical FFmpeg compression commands

FFmpeg is the industry-standard tool for processing media. To convert an oversized WAV or uncompressed audio recording into a Whisper-ready 64 kbps mono MP3, run:

```bash
ffmpeg -i input_recording.wav -vn -ar 16000 -ac 1 -b:a 64k output_whisper.mp3
```

Here is what each flag controls:
- `-vn`: Disables video recording streams, stripping thumbnails or cover art.
- `-ar 16000`: Sets the audio sampling rate to 16,000 Hz, matching Whisper's native acoustic model.
- `-ac 1`: Downmixes stereo or surround sound channels into a single mono channel.
- `-b:a 64k`: Restricts the constant audio bitrate to 64 kilobits per second.

### Extracting audio from video files directly

The Whisper API accepts `mp4`, `mpeg`, and `webm` video files directly, but doing so is inefficient. A ten-minute screen recording in high-definition video format often grows dramatically larger than pure speech because video frames take up the vast majority of the payload.

Instead of uploading the video container to the API, extract only the audio stream before calling Whisper:

```bash
ffmpeg -i product_demo.mp4 -vn -c:a libmp3lame -b:a 64k extracted_speech.mp3
```

This command extracts the sound track, ignores the video frames, and creates an MP3 that is a fraction of the size of the original video container, easily staying under the 25 MB OpenAI Whisper file limit.

## How to split long audio files into chunks with FFmpeg

When recordings extend past one hour, even aggressive audio compression cannot prevent a file from breaching the 25 MB Whisper API upload limit. A two-hour executive meeting or a three-hour university lecture requires programmatic chunking.

Splitting audio requires careful handling. If an arbitrary split divides a word in half (for example, slicing between the syllables of a spoken word), Whisper will hallucinate or drop that word entirely.

### Splitting by time segments with FFmpeg

The most dependable way to slice audio without re-encoding is to use FFmpeg's `segment` muxer. This command splits a long MP3 file into exact 10-minute (600-second) chunks:

```bash
ffmpeg -i long_conference.mp3 -f segment -segment_time 600 -c copy chunk_%03d.mp3
```

By specifying `-c copy`, FFmpeg slices the bitstream at keyframes without re-encoding the audio, completing the entire split in seconds while preserving original audio fidelity.

### Preserving context across chunk boundaries

One common issue with segmented transcription is lost context at segment boundaries. A speaker might begin a sentence at the end of `chunk_001.mp3` and finish it in `chunk_002.mp3`.

The Whisper API provides a `prompt` parameter designed specifically to maintain continuity. The Whisper API prompt parameter accepts a 224-token context window of prior text. By passing the final sentence or paragraph of the previous chunk into the prompt for the next chunk, you teach Whisper the ongoing sentence structure, punctuation style, and domain terms.

### Automated chunking and transcription script

Here is a Python script using the official `openai` package to iterate through audio chunks, pass trailing text context, and build a unified transcript:

```python
import os
from pathlib import Path
from openai import OpenAI

client = OpenAI()

def transcribe_chunks(chunk_directory: str) -> str:
    chunk_paths = sorted(Path(chunk_directory).glob("chunk_*.mp3"))
    full_transcript = []
    previous_context = ""
    for chunk_path in chunk_paths:
        print(f"Transcribing {chunk_path.name}...")
        with open(chunk_path, "rb") as audio_file:
            response = client.audio.transcriptions.create(
                model="whisper-1",
                file=audio_file,
                prompt=previous_context[-200:] if previous_context else "",
                response_format="text"
            )
            chunk_text = response.strip()
            full_transcript.append(chunk_text)
            previous_context = chunk_text
    separator = chr(10) + chr(10)
    return separator.join(full_transcript)

if __name__ == "__main__":
    final_text = transcribe_chunks("./audio_chunks")
    with open("full_transcript.txt", "w", encoding="utf-8") as out:
        out.write(final_text)
    print("Transcription complete.")
```

This pattern ensures that specialized vocabulary, proper nouns, and grammatical phrasing continue smoothly across 10-minute segment breaks.

## When to choose local Whisper over the Whisper API

When teams process thousands of hours of audio each month, the 25 MB Whisper API upload limit and recurring API fees prompt many engineering teams to consider self-hosting open-source Whisper.

OpenAI open-sourced the original Whisper weights under the MIT license, allowing anyone to execute speech recognition locally or on private cloud servers.

### Advantages of self-hosting Whisper

1. **Zero file size limits:** Local Whisper implementations take file paths directly from disk. You can pass massive WAV files or multi-hour audio streams directly into the model without slicing or compression.
2. **Predictable hardware costs:** Cloud speech APIs meter every minute of audio processed, meaning high-volume transcription pipelines generate linear operational expenses. In contrast, running a dedicated local GPU server incurs a fixed hardware or instance cost regardless of transcription volume.
3. **Data isolation:** Sensitive legal, healthcare, or financial recordings stay on local infrastructure without transiting third-party HTTP endpoints.

### Tradeoffs and operational overhead

Self-hosting Whisper introduces infrastructure maintenance that cloud APIs eliminate:

- **GPU memory requirements:** Running the highest-accuracy large models requires dedicated GPU hardware with ample video memory (VRAM). While smaller models like `base` or `small` run on CPU, their word error rates (WER) are noticeably higher on accented or reverberant speech.
- **Inference speed and batching:** The official `openai/whisper` Python repository is unoptimized for high-concurrency production. Teams must deploy modern inference engines like `faster-whisper` (built on CTranslate2) or `whisper.cpp` to achieve rapid inference speeds that transcribe hour-long recordings in minutes.
- **Cold starts and scaling:** Maintaining dedicated GPU instances creates ongoing idle costs during quiet periods and queue backlog during traffic spikes.

For most teams, the managed Whisper API remains the most reliable path for occasional or variable transcription workloads, provided they automate the pre-upload compression and chunking steps.

## Why enterprise audio archives require intelligent cloud workspaces

Most engineering tutorials solve the Whisper file size limit by writing a localized script and stopping there. But in an enterprise environment, audio processing is not an isolated one-off task. Organizations routinely manage hundreds of gigabytes of customer discovery calls, team standups, product demos, and executive interviews.

The real challenge emerges after the audio is chunked and transcribed: where do the large media files live, how do team members search through thousands of hours of speech, and how can autonomous AI agents access that knowledge without blowing past their own context limits?

### The storage and context bottleneck

Handling enterprise audio archives involves three fundamental hurdles:

1. **Media files overwhelm standard tools:** Cloud storage services built for human file synchronization struggle when automated pipelines dump massive video and audio archives into folders.
2. **Context window exhaustion:** Just as Whisper enforces a 25 MB upload ceiling, LLMs like Claude, GPT-4, and Gemini have strict context windows. You cannot paste 50 complete transcripts into an LLM prompt without hitting context limits, degrading retrieval quality, and incurring steep token fees.
3. **Siloed transcription output:** When scripts save transcripts to local developer laptops, the rest of the company cannot search or query those insights.

### Intelligent cloud workspaces for audio workflows

Intelligent cloud workspaces solve this bottleneck by combining persistent file storage with an automated intelligence layer. Rather than treating storage as a passive bucket, modern platforms index files on arrival so both humans and AI agents can query them immediately.

Fast.io provides an enterprise workspace platform designed specifically for collaborative human and agent workflows:

- **Chunked uploads without size roadblocks:** Raw audio and video files upload reliably regardless of whether they are standard voice memos or massive multi-gigabyte recordings, bypassing client-side upload limits. Learn more about media management on [Fast.io media solutions](/product/media/).
- **Intelligence Mode and built-in RAG:** Once files land in a workspace, [Intelligence Mode](/product/ai/) auto-indexes transcripts, summaries, and meeting notes for semantic and full-text search. AI assistants can ask specific questions and retrieve verified answers backed by document citations without ingesting entire audio corpora into prompt context.
- **Remote Model Context Protocol (MCP) tooling:** Autonomous coding and research agents (such as Claude Code, Cursor, Codex, and OpenClaw) connect directly to Fast.io workspaces through Streamable HTTP at `https://mcp.fast.io/mcp` or `https://mcp.fast.io/mcp/key`. For integration patterns and tool definitions, see [Storage for AI Agents](/storage-for-agents/). Agents can read audio metadata, search transcripts, and retrieve relevant excerpts on demand.
- **Collaborative Notes:** Human team members and AI agents can co-edit meeting notes, synthesize themes across hundreds of interviews, and track action items in real time.
- **Granular permissions and audit logging:** Control access at the organization, workspace, folder, and file level, supported by an append-only audit log that records every upload, edit, and agent interaction.
- **Branded shares and portals:** Deliver client-ready audio archives, video reels, and searchable transcripts via branded Send, Receive, and Exchange links that support password protection, recipient access controls, and link expiration.

By pairing local or API-based Whisper chunking with persistent cloud workspaces, teams turn raw, oversized voice recordings into an organized, searchable asset library accessible to every team member and AI assistant.

## Frequently asked questions

### What is the file size limit for the OpenAI Whisper API?

The OpenAI Whisper API enforces a strict maximum file size limit of 25 MB per request. Any upload exceeding the 25 MB OpenAI Whisper file limit is rejected immediately with an HTTP 413 Payload Too Large error.

### How do you transcribe audio files larger than the 25 MB Whisper upload limit?

You can transcribe audio files larger than the 25 MB Whisper upload limit by compressing the audio down to 64 kbps mono MP3, splitting the file into 10-minute segments using FFmpeg, or self-hosting open-source Whisper models on dedicated compute.

### Can Whisper transcribe video files directly?

Yes, the Whisper API accepts video containers including MP4, MPEG, and WebM directly, provided the file is under the 25 MB Whisper API upload limit. However, extracting the audio track with FFmpeg before uploading is recommended to avoid wasting size quota on video frames.

### What audio formats does the Whisper API support?

The Whisper API supports seven input formats: MP3, MP4, MPEG, MPGA, M4A, WAV, and WebM. All formats share the same 25 MB OpenAI Whisper file upload limit.

### Does the Whisper API have a maximum duration limit for audio?

No, OpenAI does not enforce an explicit audio duration limit. The constraint is strictly based on total request payload size. Highly compressed audio using modern codecs can fit over an hour of clear voice within the 25 MB Whisper API upload limit.

### What HTTP error code is returned when an audio file exceeds the 25 MB Whisper upload limit?

When a file exceeds the 25 MB Whisper API upload limit, the API returns an HTTP 413 status code with the message 'Maximum content size limit exceeded'. The request fails before transcription begins.

### How does reducing audio bitrate affect Whisper transcription accuracy?

Because Whisper internally downsamples all audio to 16 kHz mono before generating spectrograms, reducing speech bitrates to 64 kbps MP3 or 32 kbps AAC rarely degrades accuracy. Only extreme compression below 24 kbps risks loss of phoneme clarity.

### How can AI agents search and query transcribed audio archives?

AI agents can connect to cloud workspaces like Fast.io using the remote Model Context Protocol (MCP) server. Once transcripts are stored in an intelligent workspace, agents query specific passages using semantic search with citations instead of loading multi-gigabyte audio files into context.

## Sources

- [OpenAI: Speech to Text Guide](https://developers.openai.com/api/docs/guides/speech-to-text) — The OpenAI Whisper API enforces a 25 MB maximum file upload limit across all supported audio and video formats.

## About Fast.io

Fast.io provides shared workspaces where people and AI agents work on the same files, with built-in semantic search and citation-backed chat over what they hold. Agents reach it through a remote MCP server at https://mcp.fast.io/mcp, a REST API at https://api.fast.io/current/, and a command line client published on npm as @vividengine/fastio-cli.
