Whisper File Size Limit: The 25 MB Cap, Chunking, and Cloud Workspaces
The OpenAI Whisper API enforces a strict 25 MB file size limit for audio and video uploads across all supported file formats. Transcribing long recordings requires reducing audio bitrates, chunking files into smaller segments, or querying transcripts in persistent cloud workspaces. This guide explains Whisper limits, practical FFmpeg chunking commands, and enterprise storage patterns for large audio archives.
What is the Whisper file size limit?
The OpenAI Whisper API enforces a strict 25 MB file size limit for audio and video uploads across all supported file formats. Whether you send a raw WAV recording, an MP3 podcast, or an MP4 video clip, any single HTTP request with a payload exceeding the 25 MB Whisper API upload limit is rejected immediately by OpenAI servers.
When an upload exceeds the 25 MB OpenAI Whisper file limit, the API returns an HTTP 413 status code with the error message:
{
"error": {
"message": "Maximum content size limit exceeded",
"type": "server_error",
"param": null,
"code": null
}
}
This 25 MB file upload limit exists because OpenAI's Whisper API processes transcriptions through stateless HTTP multipart form requests. A single unified upload endpoint must guard against hanging TCP connections, buffer exhaustion, and unbounded memory consumption across thousands of concurrent API requests.
Supported formats and raw file sizes
As documented in OpenAI's official speech-to-text guide, the Whisper API accepts seven file formats:
mp3(MPEG-1 Audio Layer III)mp4(MPEG-4 Part 14, audio or video)mpeg(MPEG-1 or MPEG-2 audio or video)mpga(MPEG-1 Layer 1/2 audio stream)m4a(MPEG-4 audio container, typically AAC or ALAC)wav(Waveform Audio File Format, uncompressed PCM)webm(WebM media container, Opus or Vorbis audio)
The central friction point for developers and data teams is that audio recordings rapidly exceed the 25 MB Whisper API upload limit. An uncompressed one-hour stereo interview recorded in WAV format consumes hundreds of megabytes of disk space. Even when recorded in mono, raw uncompressed speech demands substantial storage that quickly breaches the request ceiling. Long-form corporate recordings, such as all-hands meetings, customer research calls, quarterly earnings conferences, and legal depositions, regularly produce large media files that cannot be uploaded without prior processing.
Three ways to transcribe files larger than 25 MB
Because OpenAI does not offer an enterprise quota exception or a paid toggle to raise the 25 MB single-file ceiling, teams must adopt one of three practical strategies:
- Downsample and compress: Convert stereo audio to mono, downsample the sample rate to 16 kHz, and compress into a low-bitrate format such as 64 kbps MP3 or Opus. This reduces file size substantially and allows extended speech to fit beneath the 25 MB OpenAI Whisper upload limit.
- Split audio into discrete chunks: Slice long recordings into time-based segments using FFmpeg, transcribe each segment individually through the API, and stitch the text back together while preserving boundary context.
- Run Whisper locally or in hybrid cloud workspaces: Use open-source implementations such as
faster-whisperon dedicated compute to remove file size limits completely, while storing raw media, transcripts, and metadata in unified, queryable cloud workspaces on Fast.io.
Comparing transcription architectures
The following table compares file constraints, deployment requirements, and operational characteristics across speech recognition approaches as of September 2026:
Related guides
- OpenAI Assistant File Limits: Upload Caps, Size Limits, and SolutionsOpenAI assistant file limits balance responsiveness against overhead, restricting direct attachments while capping...
- NotebookLM PDF Limit: Word Caps, Page Limits, and Gemini Notebook RulesGoogle's Gemini Notebook, formerly NotebookLM, enforces a strict limit of 500,000 words and 200MB per uploaded PDF,...
- ChatGPT PDF Limit: File Size, Page Count, and Large Document WorkaroundsThe ChatGPT PDF limit enforces a 512MB file size ceiling, a 2 million token extraction threshold, and a restriction of...
- ChatGPT Attachment Limit: File Sizes, Counts, and Large Corpus WorkaroundsChatGPT caps message attachments at 20 files, enforces a 512 MB file ceiling, and silently truncates documents beyond 2...
- ChatGPT Projects File Limit: Plan Caps, Workarounds, and Large-Corpus SearchThe ChatGPT Projects file limit restricts workspace knowledge to 5 files on Free, 25 on Plus, and 40 on Enterprise,...
- Custom GPT File Limit: Knowledge Base Caps and Large-Corpus SearchOpenAI restricts Custom GPT Knowledge bases to a hard ceiling of 10 files, 512MB per file, and 2 million tokens per...
More on this subject: Agent File and Document Workflows (218 guides)
How to compress audio files before Whisper upload
Before writing complex file-splitting pipelines, the fastest way to work around the 25 MB Whisper API upload limit is to strip unnecessary data from the audio file. Most raw recordings contain multichannel sound and frequencies that speech recognition models ignore.
Understanding Whisper's acoustic input requirements
Whisper does not benefit from high-resolution studio audio. Internally, the Whisper neural network resamples all incoming audio to 16,000 Hz (16 kHz) mono. The neural network computes an 80-channel log-Mel spectrogram using 25-millisecond windows with a 10-millisecond stride.
Uploading a high-resolution stereo FLAC file or a 24-bit uncompressed WAV does not improve transcription accuracy. It only exhausts the 25 MB OpenAI Whisper file limit within minutes. Stripping extra audio channels, downsampling to 16 kHz, and encoding with a modern codec allows you to pack hours of intelligible voice into a fraction of the space.
Bitrate, duration, and file size capacity
File size is a direct mathematical product of bitrate and duration. The table below illustrates how much speech audio fits into the 25 MB OpenAI Whisper file limit across standard audio formats:
For pure speech, 64 kbps MP3 or 32 kbps AAC delivers crisp, intelligible voice recognition. That means a standard 45-minute team standup or customer interview can easily fit into a single API request once compressed properly.
Practical FFmpeg compression commands
FFmpeg is the industry-standard tool for processing media. To convert an oversized WAV or uncompressed audio recording into a Whisper-ready 64 kbps mono MP3, run:
ffmpeg -i input_recording.wav -vn -ar 16000 -ac 1 -b:a 64k output_whisper.mp3
Here is what each flag controls:
-vn: Disables video recording streams, stripping thumbnails or cover art.-ar 16000: Sets the audio sampling rate to 16,000 Hz, matching Whisper's native acoustic model.-ac 1: Downmixes stereo or surround sound channels into a single mono channel.-b:a 64k: Restricts the constant audio bitrate to 64 kilobits per second.
Extracting audio from video files directly
The Whisper API accepts mp4, mpeg, and webm video files directly, but doing so is inefficient. A ten-minute screen recording in high-definition video format often grows dramatically larger than pure speech because video frames take up the vast majority of the payload.
Instead of uploading the video container to the API, extract only the audio stream before calling Whisper:
ffmpeg -i product_demo.mp4 -vn -c:a libmp3lame -b:a 64k extracted_speech.mp3
This command extracts the sound track, ignores the video frames, and creates an MP3 that is a fraction of the size of the original video container, easily staying under the 25 MB OpenAI Whisper file limit.
How to split long audio files into chunks with FFmpeg
When recordings extend past one hour, even aggressive audio compression cannot prevent a file from breaching the 25 MB Whisper API upload limit. A two-hour executive meeting or a three-hour university lecture requires programmatic chunking.
Splitting audio requires careful handling. If an arbitrary split divides a word in half (for example, slicing between the syllables of a spoken word), Whisper will hallucinate or drop that word entirely.
Splitting by time segments with FFmpeg
The most dependable way to slice audio without re-encoding is to use FFmpeg's segment muxer. This command splits a long MP3 file into exact 10-minute (600-second) chunks:
ffmpeg -i long_conference.mp3 -f segment -segment_time 600 -c copy chunk_%03d.mp3
By specifying -c copy, FFmpeg slices the bitstream at keyframes without re-encoding the audio, completing the entire split in seconds while preserving original audio fidelity.
Preserving context across chunk boundaries
One common issue with segmented transcription is lost context at segment boundaries. A speaker might begin a sentence at the end of chunk_001.mp3 and finish it in chunk_002.mp3.
The Whisper API provides a prompt parameter designed specifically to maintain continuity. The Whisper API prompt parameter accepts a 224-token context window of prior text. By passing the final sentence or paragraph of the previous chunk into the prompt for the next chunk, you teach Whisper the ongoing sentence structure, punctuation style, and domain terms.
Automated chunking and transcription script
Here is a Python script using the official openai package to iterate through audio chunks, pass trailing text context, and build a unified transcript:
import os
from pathlib import Path
from openai import OpenAI
client = OpenAI()
def transcribe_chunks(chunk_directory: str) -> str:
chunk_paths = sorted(Path(chunk_directory).glob("chunk_*.mp3"))
full_transcript = []
previous_context = ""
for chunk_path in chunk_paths:
print(f"Transcribing {chunk_path.name}...")
with open(chunk_path, "rb") as audio_file:
response = client.audio.transcriptions.create(
model="whisper-1",
file=audio_file,
prompt=previous_context[-200:] if previous_context else "",
response_format="text"
)
chunk_text = response.strip()
full_transcript.append(chunk_text)
previous_context = chunk_text
separator = chr(10) + chr(10)
return separator.join(full_transcript)
if __name__ == "__main__":
final_text = transcribe_chunks("./audio_chunks")
with open("full_transcript.txt", "w", encoding="utf-8") as out:
out.write(final_text)
print("Transcription complete.")
This pattern ensures that specialized vocabulary, proper nouns, and grammatical phrasing continue smoothly across 10-minute segment breaks.
Manage and query large audio transcripts in one shared workspace
Store large media files with chunked uploads, auto-index transcripts for semantic search, and connect your AI agents through the Fast.io remote MCP server. Every organization starts with a 14-day free trial, which requires a credit card.
When to choose local Whisper over the Whisper API
When teams process thousands of hours of audio each month, the 25 MB Whisper API upload limit and recurring API fees prompt many engineering teams to consider self-hosting open-source Whisper.
OpenAI open-sourced the original Whisper weights under the MIT license, allowing anyone to execute speech recognition locally or on private cloud servers.
Advantages of self-hosting Whisper
- Zero file size limits: Local Whisper implementations take file paths directly from disk. You can pass massive WAV files or multi-hour audio streams directly into the model without slicing or compression.
- Predictable hardware costs: Cloud speech APIs meter every minute of audio processed, meaning high-volume transcription pipelines generate linear operational expenses. In contrast, running a dedicated local GPU server incurs a fixed hardware or instance cost regardless of transcription volume.
- Data isolation: Sensitive legal, healthcare, or financial recordings stay on local infrastructure without transiting third-party HTTP endpoints.
Tradeoffs and operational overhead
Self-hosting Whisper introduces infrastructure maintenance that cloud APIs eliminate:
- GPU memory requirements: Running the highest-accuracy large models requires dedicated GPU hardware with ample video memory (VRAM). While smaller models like
baseorsmallrun on CPU, their word error rates (WER) are noticeably higher on accented or reverberant speech. - Inference speed and batching: The official
openai/whisperPython repository is unoptimized for high-concurrency production. Teams must deploy modern inference engines likefaster-whisper(built on CTranslate2) orwhisper.cppto achieve rapid inference speeds that transcribe hour-long recordings in minutes. - Cold starts and scaling: Maintaining dedicated GPU instances creates ongoing idle costs during quiet periods and queue backlog during traffic spikes.
For most teams, the managed Whisper API remains the most reliable path for occasional or variable transcription workloads, provided they automate the pre-upload compression and chunking steps.
Why enterprise audio archives require intelligent cloud workspaces
Most engineering tutorials solve the Whisper file size limit by writing a localized script and stopping there. But in an enterprise environment, audio processing is not an isolated one-off task. Organizations routinely manage hundreds of gigabytes of customer discovery calls, team standups, product demos, and executive interviews.
The real challenge emerges after the audio is chunked and transcribed: where do the large media files live, how do team members search through thousands of hours of speech, and how can autonomous AI agents access that knowledge without blowing past their own context limits?
The storage and context bottleneck
Handling enterprise audio archives involves three fundamental hurdles:
- Media files overwhelm standard tools: Cloud storage services built for human file synchronization struggle when automated pipelines dump massive video and audio archives into folders.
- Context window exhaustion: Just as Whisper enforces a 25 MB upload ceiling, LLMs like Claude, GPT-4, and Gemini have strict context windows. You cannot paste 50 complete transcripts into an LLM prompt without hitting context limits, degrading retrieval quality, and incurring steep token fees.
- Siloed transcription output: When scripts save transcripts to local developer laptops, the rest of the company cannot search or query those insights.
Intelligent cloud workspaces for audio workflows
Intelligent cloud workspaces solve this bottleneck by combining persistent file storage with an automated intelligence layer. Rather than treating storage as a passive bucket, modern platforms index files on arrival so both humans and AI agents can query them immediately.
Fast.io provides an enterprise workspace platform designed specifically for collaborative human and agent workflows:
- Chunked uploads without size roadblocks: Raw audio and video files upload reliably regardless of whether they are standard voice memos or massive multi-gigabyte recordings, bypassing client-side upload limits. Learn more about media management on Fast.io media solutions.
- Intelligence Mode and built-in RAG: Once files land in a workspace, Intelligence Mode auto-indexes transcripts, summaries, and meeting notes for semantic and full-text search. AI assistants can ask specific questions and retrieve verified answers backed by document citations without ingesting entire audio corpora into prompt context.
- Remote Model Context Protocol (MCP) tooling: Autonomous coding and research agents (such as Claude Code, Cursor, Codex, and OpenClaw) connect directly to Fast.io workspaces through Streamable HTTP at
https://mcp.fast.io/mcporhttps://mcp.fast.io/mcp/key. For integration patterns and tool definitions, see Storage for AI Agents. Agents can read audio metadata, search transcripts, and retrieve relevant excerpts on demand. - Collaborative Notes: Human team members and AI agents can co-edit meeting notes, synthesize themes across hundreds of interviews, and track action items in real time.
- Granular permissions and audit logging: Control access at the organization, workspace, folder, and file level, supported by an append-only audit log that records every upload, edit, and agent interaction.
- Branded shares and portals: Deliver client-ready audio archives, video reels, and searchable transcripts via branded Send, Receive, and Exchange links that support password protection, recipient access controls, and link expiration.
By pairing local or API-based Whisper chunking with persistent cloud workspaces, teams turn raw, oversized voice recordings into an organized, searchable asset library accessible to every team member and AI assistant.
Sources
References used to verify factual claims in this guide.
-
The OpenAI Whisper API enforces a 25 MB maximum file upload limit across all supported audio and video formats.
Frequently Asked Questions
What is the file size limit for the OpenAI Whisper API?
The OpenAI Whisper API enforces a strict maximum file size limit of 25 MB per request. Any upload exceeding the 25 MB OpenAI Whisper file limit is rejected immediately with an HTTP 413 Payload Too Large error.
How do you transcribe audio files larger than the 25 MB Whisper upload limit?
You can transcribe audio files larger than the 25 MB Whisper upload limit by compressing the audio down to 64 kbps mono MP3, splitting the file into 10-minute segments using FFmpeg, or self-hosting open-source Whisper models on dedicated compute.
Can Whisper transcribe video files directly?
Yes, the Whisper API accepts video containers including MP4, MPEG, and WebM directly, provided the file is under the 25 MB Whisper API upload limit. However, extracting the audio track with FFmpeg before uploading is recommended to avoid wasting size quota on video frames.
What audio formats does the Whisper API support?
The Whisper API supports seven input formats: MP3, MP4, MPEG, MPGA, M4A, WAV, and WebM. All formats share the same 25 MB OpenAI Whisper file upload limit.
Does the Whisper API have a maximum duration limit for audio?
No, OpenAI does not enforce an explicit audio duration limit. The constraint is strictly based on total request payload size. Highly compressed audio using modern codecs can fit over an hour of clear voice within the 25 MB Whisper API upload limit.
What HTTP error code is returned when an audio file exceeds the 25 MB Whisper upload limit?
When a file exceeds the 25 MB Whisper API upload limit, the API returns an HTTP 413 status code with the message 'Maximum content size limit exceeded'. The request fails before transcription begins.
How does reducing audio bitrate affect Whisper transcription accuracy?
Because Whisper internally downsamples all audio to 16 kHz mono before generating spectrograms, reducing speech bitrates to 64 kbps MP3 or 32 kbps AAC rarely degrades accuracy. Only extreme compression below 24 kbps risks loss of phoneme clarity.
How can AI agents search and query transcribed audio archives?
AI agents can connect to cloud workspaces like Fast.io using the remote Model Context Protocol (MCP) server. Once transcripts are stored in an intelligent workspace, agents query specific passages using semantic search with citations instead of loading multi-gigabyte audio files into context.
Related Resources
Manage and query large audio transcripts in one shared workspace
Store large media files with chunked uploads, auto-index transcripts for semantic search, and connect your AI agents through the Fast.io remote MCP server. Every organization starts with a 14-day free trial, which requires a credit card.