AI & Agents

OpenAI Batch API Limits: File Sizes, Request Quotas, and Token Enqueue Caps

OpenAI Batch API limits enforce a maximum file size of 200 MB, a cap of 50,000 requests per input file, and tier-specific enqueued token pools. Understanding file partitioning methods, queue capacity thresholds, and output reconciliation enables engineering teams to execute high-volume workloads reliably without hitting quota rejections.

Derek Labian 15 min read Updated
OpenAI Batch API enforces strict constraints on input file size, request counts, and enqueued prompt tokens.

What Are the Documented OpenAI Batch API Limits?

OpenAI caps every individual batch input file at 50,000 separate requests and 200 MB in uncompressed JSON Lines (JSONL) format, delivering a 50% cost discount compared to synchronous APIs across a guaranteed 24-hour execution window. For engineering teams evaluating offline workloads, understanding these operational boundaries determines how source datasets are prepared, staged, and submitted.

The Batch API operates independently of standard synchronous endpoints. Instead of sending an HTTP request and keeping a client socket open while waiting for model inference, developers package thousands of individual requests into a single file formatted as JSON Lines. That file is uploaded through the Files API and submitted to the batch processing queue as documented in the OpenAI Batch API documentation. Because the batch execution pool is completely decoupled from synchronous request queues, it allows engineering teams to execute massive offline workloads without drawing down real-time rate limits. However, scaling batch jobs introduces a distinct set of operational boundaries.

The table below summarizes the core limits, service parameters, and operational constraints governing the OpenAI Batch API:

Parameter Platform Limit Scope and Conditions Date Checked Source
Maximum File Size 200 MB (JSONL) Per uploaded batch input file; 100 MB practical target September 2026 OpenAI Platform Docs
Maximum Requests per Batch 50,000 requests Maximum separate API calls in a single input file September 2026 OpenAI Platform Docs
Embeddings Input Limit 50,000 total inputs Cumulative text input strings across all requests in an embeddings batch September 2026 OpenAI Platform Docs
Batch Creation Rate 2,000 batches per hour Maximum new batch creation calls per organization account September 2026 OpenAI Platform Docs
Turnaround Window (SLA) 24 hours Documented execution window before remaining items expire September 2026 OpenAI Platform Docs
Token Cost Discount 50% discount Applied to both input and output tokens across supported models September 2026 OpenAI Platform Docs
Output Token Limit No explicit cap Output file scales dynamically to contain all completed outputs September 2026 OpenAI Platform Docs
Model Uniformity 1 model per batch All requests within an input file must target the same model September 2026 OpenAI Platform Docs

Structural Rules for Batch JSONL Files

Every request in a batch input file must exist as a standalone, valid JSON object on its own line. Standard JSON arrays wrapping multiple requests will fail validation immediately. Each line specifies three mandatory components: a custom identifier string (custom_id), the target HTTP method (POST), and the relative endpoint URL.

The custom_id string is critical for data integrity. Because OpenAI distributes execution across parallel compute nodes, output lines in the completed results file do not arrive in the same order as the input file. Applications must correlate responses back to source records using this identifier rather than line indices.

File Size Restrictions and Format Requirements

While the OpenAI Batch API accepts single input files up to 200 MB, experienced teams often design ingestion pipelines targeting 100 MB slices. The 200 MB file limit applies to the raw byte size of the uncompressed JSONL file uploaded to the Files API with purpose: "batch". Gzip, zip, or tar compression formats are rejected. Smaller slices reduce network transfer retry penalties and simplify local memory management during pre-processing.

Request Caps and Supported Endpoints

OpenAI caps each batch at exactly 50,000 requests. For /v1/embeddings, an additional constraint applies: the total number of embedding inputs across all requests within the file cannot exceed 50,000 items. If you send 5,000 requests where each request contains an array of 20 strings, the batch reaches 100,000 inputs and fails validation.

The Batch API supports key production endpoints, including /v1/chat/completions, /v1/embeddings, /v1/moderations, /v1/completions, and /v1/responses. However, every single request in a batch file must specify the exact same model. You cannot mix small evaluation calls targeting lightweight models with heavy reasoning calls in the same file.

How Do Enqueued Token Limits Restrict Batch Throughput Across Tiers?

Batch API rate limits do not follow the standard Requests Per Minute (RPM) or Tokens Per Minute (TPM) counters used for synchronous calls. Instead, OpenAI regulates batch throughput through enqueued prompt token limits.

An enqueued token limit defines the maximum volume of input tokens that can sit waiting in uncompleted batch queues for a given model at any single point in time. When an application submits a batch, OpenAI calculates the total prompt tokens across all requests in the file and adds that number to the organization's active queue counter. Those tokens remain reserved against your quota until the batch reaches a terminal state: completed, failed, cancelled, or expired. Only when the job finishes are those tokens released back into your available pool.

If a newly submitted batch would push your organization's total queued tokens beyond your tier ceiling, the API returns a token_limit_exceeded error and refuses to create the job.

The table below outlines typical enqueued token thresholds across OpenAI usage tiers:

Usage Tier Account Status Typical Enqueued Token Limit Parallel Batch Feasibility Date Checked Source
Tier 1 Initial funded account 3,000,000 tokens 1 to 2 small batches September 2026 OpenAI Platform
Tier 2 Early production account 10,000,000 tokens 3 to 5 standard batches September 2026 OpenAI Platform
Tier 3 Established production account 40,000,000 tokens Multiple concurrent production batches September 2026 OpenAI Platform
Tier 4 High-volume production account 80,000,000 tokens Large parallel processing queues September 2026 OpenAI Platform
Tier 5 High-scale enterprise account 200,000,000+ tokens Continuous enterprise backlogs September 2026 OpenAI Platform

Why Enqueued Token Caps Block Concurrent Jobs

The most frequent bottleneck in automated batch systems is queue saturation. Consider a Tier 2 organization with a 10,000,000 enqueued token limit submitting a batch containing 40,000 requests, each with a 200-token prompt. That single batch reserves 8,000,000 tokens.

Because that batch occupies most of the available pool, submitting a second batch of equal size will immediately fail with a rate limit error. Because batch turnaround can take several hours depending on cluster load, the entire ingestion pipeline stalls unless the orchestration layer monitors active queue capacity.

Managing Shared Model Family Pools

OpenAI groups certain related models into shared rate limit pools. When models share a pool, enqueued batch tokens for one model directly reduce the available queue headroom for all sibling models in that family.

Before submitting multi-model pipelines, review the organization limits in OpenAI platform settings. If your evaluation workflow runs multiple variations of a base model concurrently, calculate the aggregate prompt token footprint to ensure the shared queue ceiling is not exceeded.

Calculating Queue Headroom Before Submission

To prevent submission failures, production pipelines should calculate total prompt tokens client-side using a tokenizer before uploading files. Maintain an internal ledger of active batch jobs and their token commitments.

When the active queue approaches the tier threshold, pause new submissions and poll the status of running jobs. Once a running job returns a terminal status, the released capacity permits submitting the next chunk without triggering unhandled API exceptions.

How to Split Large Corpora Exceeding File and Request Limits

When raw corpora exceed 100 MB, 200 MB, or 50,000 requests, engineering teams must partition their data into compliant OpenAI batch input files. Naive chunking based purely on line counts can produce files that violate byte-size boundaries if individual lines contain extensive document context. Conversely, splitting strictly by file size can produce chunks that exceed the 50,000 request limit if records are concise.

A reliable preparation pipeline applies dual-constraint partitioning: it tracks both cumulative byte size and total line count simultaneously, closing the current chunk whenever either threshold is approached. To maintain a safe operating margin against edge-case serialization expansions, set target ceilings below the hard platform maximums, keeping request counts and file sizes comfortably within boundaries.

The following Python script demonstrates how to split an oversized JSONL dataset into safe, compliant batch slices using standard library tools:

import json
from pathlib import Path

def partition_batch_file(
    input_path: Path,
    output_dir: Path,
    max_requests: int = 45000,
    max_bytes: int = 90 * 1024 * 1024,
) -> list:
    output_dir.mkdir(parents=True, exist_ok=True)
    created_files = []
    chunk_index = 1
    current_requests = 0
    current_bytes = 0
    current_lines = []
    ### Helper function to write completed chunk
    def flush_chunk():
        nonlocal chunk_index, current_requests, current_bytes, current_lines
        if not current_lines:
            return
        chunk_filename = output_dir / f"batch_chunk_{chunk_index:03d}.jsonl"
        with open(chunk_filename, "w", encoding="utf-8") as f:
            f.writelines(current_lines)
        created_files.append(chunk_filename)
        chunk_index += 1
        current_requests = 0
        current_bytes = 0
        current_lines = []
    ### Stream input file record by record
    with open(input_path, "r", encoding="utf-8") as src:
        for line in src:
            if not line.strip():
                continue
            line_bytes = len(line.encode("utf-8"))
            if (current_requests + 1 > max_requests) or (current_bytes + line_bytes > max_bytes):
                flush_chunk()
            current_lines.append(line)
            current_requests += 1
            current_bytes += line_bytes
    ### Write final trailing chunk
    flush_chunk()
    return created_files

This implementation reads the source data line by line, preventing high memory usage even when handling multi-gigabyte source archives. Each emitted chunk remains safely beneath the OpenAI Batch API input file limits of 50,000 requests and 200 MB, allowing automated submission scripts to cycle through generated files sequentially.

Preserving Record Traceability Across Slices

When partitioning datasets into multiple files, maintain a persistent registry mapping each record's primary key to its assigned batch chunk and internal custom_id. If an individual chunk fails validation or experiences an unexpected network timeout during upload, the registry allows targeted reprocessing of that specific slice without regenerating the entire dataset.

Pre-Flight Schema and Syntax Validation

Before initiating an upload to the Files API, validate each JSON line locally against the OpenAI batch JSON schema. A single malformed JSON object, trailing comma, or unrecognized parameter in an input file causes the entire batch job to fail immediately during OpenAI's initial validation phase. Running a client-side syntax check prevents wasted upload bandwidth and unnecessary queue delays.

Fastio features

Manage Document Corpora for OpenAI Batch Processing

Store, partition, and query large datasets in shared workspaces with built-in search and remote MCP tools before queueing offline batches. Monthly plans start with a trial of up to 30 days (credit card required); annual plans have no trial.

How to Handle Asynchronous Job Failures and Partial Completions

Batch jobs transition through multiple discrete states during their lifecycle: validating, failed, in_progress, finalizing, completed, expired, cancelling, and cancelled. Understanding how errors surface at each stage is essential for building resilient automation.

Failures in the Batch API fall into two distinct categories: batch-level rejections and request-level execution errors. Treating these two failure types identically in code will lead to either data loss or infinite retry loops.

Batch-level rejections occur during the validating stage. If your input file contains invalid JSON syntax on line 12,000, specifies an unsupported model name, or violates the 50,000 request cap, the entire batch status flips to failed. In this scenario, OpenAI executes zero requests. The job produces no output file, and the error details appear directly within the errors array of the batch object retrieved via the API.

Conversely, request-level errors occur during the in_progress stage. If an input file is syntactically valid but request 4,210 contains a prompt that triggers a content moderation filter or exceeds the model context window, the overall batch continues processing. When the job completes, OpenAI generates two separate file artifacts: an output file (output_file_id) containing successful responses and an error file (error_file_id) recording individual failed requests alongside their specific HTTP status codes and error messages.

Reconciling Out-of-Order Responses

Because OpenAI evaluates requests in parallel across distributed worker clusters, the sequence of records in the result file will not match the sequence in your input file. A completed output line contains the custom_id, the HTTP response code, and the response body.

Your ingestion worker must stream the output file line by line, parse the custom_id, and update the corresponding database record or downstream document store. Do not rely on line number correspondence between input and output files.

Handling the 24-Hour Expiration Window

OpenAI provides a 24-hour turnaround window for batch execution. In rare situations where cluster demand spikes or a model family experiences transient compute shortages, requests that remain unexecuted after 24 hours expire.

When a batch reaches the expired state, OpenAI generates an output file containing any requests that finished successfully before the deadline. Orchestration pipelines must detect the expired status, parse the partial output file, identify which custom_id records were never completed, and package the remainder into a fresh batch file for resubmission.

Coordinating Persistent Workspace Storage for Batch Ingestion

Managing tens of thousands of documents across multi-file batch runs requires durable storage that bridges local preparation, model execution, and team review. Storing massive prompt archives solely on ephemeral developer laptops or isolated object buckets creates operational blind spots when jobs fail or when team members need to inspect raw results.

Modern agentic teams organize high-volume AI pipelines using Fast.io intelligent workspaces. In Fast.io, files are stored in organization-owned workspaces with granular folder-level permissions, per-file version history, and an append-only audit log. This architecture provides a centralized, traceable repository where source datasets, partitioned batch chunks, and completed JSONL output files remain accessible to both human operators and automated agents.

When preparing batch files, agents can access shared files directly using the remote Fast.io Model Context Protocol (MCP) server at https://mcp.fast.io/mcp. Rather than writing custom file transfer code, an AI agent interacting with Fast.io can inspect raw document collections, trigger structured metadata extraction using Metadata Views, and generate formatted JSONL batch slices ready for the OpenAI Files API.

Once OpenAI finishes processing a batch, the output file can be downloaded and stored directly back into the Fast.io workspace. Because Fast.io includes Intelligence Mode, stored documents and generated model responses are automatically indexed for hybrid full-text and semantic search. Team members can query thousands of model completions using natural language through the workspace interface or via MCP, identifying edge cases and verifying completion quality without writing complex SQL scripts or database queries.

For teams collaborating with external stakeholders or review teams, Fast.io provides branded shares (Send, Receive, and Exchange) with expiring links and access controls. An engineering team can securely share batch output summaries or review packets without granting external parties direct access to production OpenAI credentials or internal database clusters.

Production Best Practices for OpenAI Batch Orchestration

Running high-volume batch workloads reliably requires operational discipline around token budgeting, polling frequency, and storage hygiene. Implementing the following architectural practices prevents pipeline interruptions and minimizes operational friction:

  • Implement Client-Side Token Auditing: Never submit an uncounted batch to the API. Tokenize prompt texts using a compatible tokenizer library prior to generating JSONL chunks. Track cumulative enqueued tokens against your tier limit in an internal database or caching layer to avoid unexpected token_limit_exceeded rejections.
  • Apply Exponential Backoff for Status Polling: Because batch jobs are designed for offline execution over hours rather than seconds, aggressive polling burns unnecessary rate limit capacity. Configure status check workers to poll on an exponential schedule, starting at 30 seconds and widening to 5-minute intervals.
  • Automate OpenAI Storage Cleanup: Uploaded batch files and generated output files count against your organization's storage quota on OpenAI servers. Once an output file has been retrieved, verified, and safely written to your Fast.io AI storage workspace, delete both the input file and output file from OpenAI using the Files API.
  • Maintain Idempotent State Tracking: Ensure every record submitted in a batch carries a globally unique custom_id that encodes both the dataset identifier and record primary key. If a network blip causes a batch submission retry, idempotent keys prevent duplicate processing and conflicting database writes.
  • Monitor Shared Model Headroom: If multiple teams or automated workflows within your company share an OpenAI organization key, coordinate batch scheduling to prevent one team's offline batch from consuming the entire enqueued token pool for a critical model family.

Sources

References used to verify factual claims in this guide.

  1. 1 OpenAI API Documentation Accessed

    The OpenAI Batch API provides a 50% cost discount compared to synchronous APIs, with a single batch including up to 50,000 requests and an input file up to 200 MB in size. For /v1/embeddings batches, the total number of embedding inputs across all requests within the file cannot exceed 50,000 items.

Frequently Asked Questions

What is the maximum file size for the OpenAI Batch API?

The OpenAI Batch API enforces a maximum input file size of 200 MB for uncompressed JSON Lines (JSONL) files. In practice, many engineering teams target smaller slices to simplify memory handling and reduce network retry overhead.

How many tokens can you enqueue in the OpenAI Batch API?

Enqueued token limits vary by organization usage tier and model family. Tier 1 accounts typically have a 3,000,000 enqueued prompt token cap, while Tier 5 accounts can enqueue 200,000,000 or more tokens across pending batches.

What happens if a batch exceeds the 50,000 request limit?

If a single JSONL file contains more than 50,000 requests, OpenAI rejects the batch during the initial validation stage. The batch status transitions to failed, and no requests in the file are executed.

Does the OpenAI Batch API offer a cost discount compared to synchronous endpoints?

Yes, the OpenAI Batch API provides a 50% cost discount compared to synchronous APIs on both input prompt tokens and output completion tokens, in exchange for a 24-hour turnaround window.

What happens if individual requests fail inside a valid batch?

If a batch file passes initial syntax validation, individual request errors do not fail the entire batch. Successful responses are written to an output file, while failed requests are written to a separate error file detailing the HTTP error code and message.

Related Resources

Fastio features

Manage Document Corpora for OpenAI Batch Processing

Store, partition, and query large datasets in shared workspaces with built-in search and remote MCP tools before queueing offline batches. Monthly plans start with a trial of up to 30 days (credit card required); annual plans have no trial.