OpenAI Vector Store File Limits: Capacities, Pricing, and Workarounds
OpenAI caps each vector store file at 512 MB and 5,000,000 tokens, and bills storage at $0.10 per GB per day once a project passes the first free gigabyte. A 10,000-file ceiling per store is widely cited and appears in Microsoft's Azure OpenAI documentation, though OpenAI's own current Retrieval guide publishes no file-count limit. Understanding chunking boundaries, expiration policies, and external retrieval architectures helps engineering teams scale document access without unexpected billing or ingestion failures.
What Are the OpenAI Vector Store File Limits and Capacities?
OpenAI's Retrieval guide caps each vector store file at 512 MB and 5,000,000 tokens, limits file and batch creation to 300 requests per minute per store, and bills storage at $0.10 per GB per day once a project passes the first free gigabyte. When an enterprise corpus exceeds those boundaries, automated ingestion halts with unhandled batch errors, while daily storage fees accumulate on parsed chunks and embeddings.
A 10,000-file ceiling per vector store is widely repeated and is stated in Microsoft's Azure OpenAI Assistants documentation. It is worth knowing about, but OpenAI's own Retrieval guide no longer publishes a file-count ceiling, so treat 10,000 as a planning assumption to verify against your account rather than a figure OpenAI currently documents.
The hosted File Search tool inside OpenAI's Responses API manages document parsing, chunking, and embedding automatically. Developers do not need to configure an external embedding model or vector database to start asking questions about uploaded documents. That convenience comes with structural constraints that become apparent once an application moves from prototype testing into production workloads.
The table below summarizes the core technical specifications and limits governing OpenAI vector stores:
Attachment Boundaries Across Assistants and Threads
OpenAI structures document access around two attachment scopes: assistant-level resources and thread-level resources.
An assistant can link to at most one vector store inside its tool resources configuration, serving as a static knowledge base across all runs. Individual threads can also attach one vector store for session-specific uploads. When a run executes, File Search queries both stores in parallel and reranks the results. While this dual-store model suits single-agent chat, it limits workflows that need dynamic access across many file repositories.
Supported Document Formats and Token Limits
OpenAI vector stores accept standard document and developer formats, including PDF, DOCX, TXT, MD, JSON, CSV, Python, and JavaScript. Uploads are validated against the maximum file size of 512 MB for an OpenAI vector store. Even compact files fail ingestion if they exceed the 5,000,000 tokens per file ceiling for an OpenAI vector store. Compressed logs, dense codebases, and large JSON dumps often reach the token limit well before hitting disk size thresholds.
Storage Billing and Capacity Planning
Storage is billed rather than capped. OpenAI charges on the total storage used across all of a project's vector stores, measured on parsed chunks and their embeddings, with the first gigabyte free and everything beyond it at $0.10 per GB per day. That billing model is separate from API rate limits like Tokens Per Minute (TPM), so a project can be well inside its throughput budget and still accrue meaningful daily index costs.
Related guides
- AnythingLLM Context Window: Configuration, Chunking Limits, and Document ManagementThe AnythingLLM context window determines how many tokens of chat history, system instructions, and vector search...
- How to Configure Claygent RAG Document Storage for GTM ResearchSales intelligence teams struggle to scale automated B2B research because of LLM hallucination rates in agentic...
- GPT-4o Mini Context Window: Token Limits, TPM Tiers, and Large-File WorkaroundsThe GPT-4o mini context window is the 128,000-token total input capacity supported by OpenAI's lightweight model,...
- How to Secure Vector Stores for AI AgentsVector stores are the memory layer for AI agents, and attackers know it. RAG poisoning, embedding manipulation, and...
- AWS Bedrock Context Window: Token Limits, Model Capacities, and Memory ArchitectureThe AWS Bedrock context window defines the maximum sequence of input and output tokens a hosted foundation model...
- Google AI Studio Context Window: Token Limits, File Uploads, and Persistent StorageThe Google AI Studio context window supports massive token capacities alongside a generous file upload ceiling via the...
More on this subject: Agent Memory and Storage (220 guides)
How Parsing, Chunking, and Embedding Work Under the Hood
Adding a file to an OpenAI vector store triggers an automated pipeline: document parsing, token-based chunking, vector embedding generation, and index persistence. Default chunking slices documents into 800-token segments with a 400-token overlap to maintain contextual continuity across chunk boundaries.
OpenAI documents that semantic similarity in this pipeline is computed with cosine similarity over text-embedding-3-small embeddings, alongside a keyword similarity score. Because the embedding model is fixed, tuning retrieval quality means adjusting chunk size, overlap, and ranking options rather than swapping in a different embedding model.
Why Complex Documents Fail Ingestion
While the platform supports files up to 512 MB, real-world document processing frequently encounters operational limits:
- Complex PDF Parsing Timeouts. Scanned documents, layered diagrams, and high-resolution images demand heavy CPU time during text extraction, leading to parsing timeouts.
- Table Structure Flattening. The default parser flattens tabular data into plain text strings, stripping row and column alignment across 800-token chunks.
- Token Density Memory Spikes. Unformatted text and minified code files can exhaust worker memory during tokenization.
- Opaque Ingestion Telemetry. Failed uploads return generic error messages without page-level traces, requiring manual file debugging.
Configuring Custom Chunking Parameters
Developers can override default chunking settings when attaching files to a vector store. The API accepts a chunking strategy configuration defining maximum chunk size between 100 and 4,096 tokens, with chunk overlap adjustable up to half the chunk size. For concise FAQs, 300-token chunks minimize context dilution. For legal contracts spanning multiple pages, 1,500-token chunks keep interdependent clauses intact.
Batch Ingestion and Asynchronous Status Tracking
Sequential file uploads introduce network latency. OpenAI provides a file batch endpoint that accepts batches of file IDs in a single request. Processing runs asynchronously, requiring client code to poll the batch resource's file_counts object until in-progress operations reach zero. Failed files must be queried separately to inspect specific error codes.
How Much Does OpenAI Vector Store Storage Cost in Practice?
OpenAI vector store pricing differs fundamentally from completion token billing. While language models charge per token during generation, vector stores assess a recurring daily rate to maintain the search index.
Vector store storage charges a recurring daily maintenance rate after the initial complimentary gigabyte allowance. This daily fee applies to the aggregate volume of parsed text chunks and vector embeddings across all active vector stores in an organization project.
The table below details daily and monthly storage costs across different data volumes:
At approximately three dollars per gigabyte monthly, hosted vector storage costs significantly more than raw object storage. Teams pay a premium for managed chunking, automated embeddings, and serverless index querying.
The Compounding Cost of Inactivity Policies
OpenAI lets you set an expiration policy on a vector store with expires_after, and when a store expires its files are deleted and billing for them stops. That makes expiration the main lever on runaway index cost. Leave it unset on stores created per session and abandoned sessions accumulate ongoing storage charges across thousands of users; set it too aggressively and later requests fail unless the application recreates and re-indexes the store.
Calculating Total Retrieval Expenses
Total File Search expenses include persistent daily storage fees plus runtime search execution charges. In addition, every retrieved chunk injected into the completion prompt incurs standard model input token fees. In high-traffic applications with frequent queries, prompt token usage from retrieved context often exceeds base vector storage costs.
Managing Expiration Anchors in Production
Production systems must configure explicit expiration anchors to avoid runaway billing. Vector stores accept an expiration policy anchored to last_active_at with a defined day count. Ephemeral customer chats benefit from a 3-day expiration, whereas corporate knowledge bases should omit expiration while enforcing storage monitoring alerts.
Why Vector Stores Fail and Create Bottlenecks at Scale
Pushing a single OpenAI vector store toward tens of thousands of files reveals operational friction points across ingestion latency, rate limits, and context retrieval precision.
Planning Around a File-Count Ceiling
Microsoft's Azure OpenAI documentation states a 10,000-file ceiling per vector store, and the same figure circulates widely for OpenAI's own service, though OpenAI's current Retrieval guide does not publish one. Either way, vector stores provide no native partitioning, so applications accumulating ongoing records should implement an external sharding mechanism rather than assume a single store will absorb them.
Ingestion Throttles
OpenAI caps file and file batch creation at 300 requests per minute per vector store. Large document ingestions require rate limiting to avoid HTTP 429 errors. Batching helps: OpenAI recommends batch creation for higher-throughput ingestion into a single store and accepts up to 500 files in one batch request, which reduces contention against that per-minute ceiling.
Chat Attachment Limits Compared Across Vendors
Vendor file limits differ across chat and retrieval interfaces. For example, Anthropic Claude enforces per-message limits in standard chat and project-level caps in Claude Projects, where total capacity is constrained by the active context window. Reaching context limits is why large document collections require a dedicated external retrieval layer.
Diagnosing Silent Ingestion Failures
Batch ingestion does not fail as a whole when individual files error out; the overall batch status marks as completed. Production scripts must inspect each file object's status and review specific error messages, such as unsupported encodings or token limit violations.
Context Window Saturation Versus Retrieval Precision
Injecting multiple retrieved chunks into every model prompt consumes thousands of context tokens. Loading excessive context inflates inference latency and risks attention dilution. Optimizing precision requires pre-filtering search scopes and applying strict reranking score thresholds before passing chunks to the model.
Store and search large document corpora without vector store ceilings
Index extensive document collections in an intelligent workspace to query files over remote MCP without per-store file caps or daily per-gigabyte vector fees. Monthly plans start with a trial of up to 30 days (credit card required); annual plans have no trial.
What Architectural Workarounds Exist When a Vector Store Runs Out of Room?
When document archives outgrow a single vector store, engineering teams must adopt alternative architectures. Four primary strategies address single-store capacity boundaries:
1. Document Concatenation and Archive Merging
Merging multiple small files into consolidated text archives reduces overall file count. However, this approach sacrifices source citation granularity, requires expensive re-indexing when single records change, and risks semantic blending across chunk boundaries.
2. Multi-Vector Store Routing Layers
Teams can partition documents across multiple vector stores categorized by department or domain. An initial classification prompt routes user queries to the appropriate vector store. The major limitation is that assistants attach only one vector store at a time, preventing unified cross-store searches during a single run.
3. Dedicated Vector Databases
Dedicated vector databases like Pinecone, Qdrant, Milvus, and Weaviate scale to millions of vectors with custom embedding models and metadata filters. The downside is operational complexity: engineering teams must build and maintain parsing, chunking, and index synchronization pipelines.
4. Decoupling Storage into Intelligent Cloud Workspaces
An effective alternative separates persistent document storage from the LLM runtime. In Fastio intelligent workspaces, documents are stored in shared workspaces with no per-store file-count ceiling. Workspace Intelligence provides hybrid search, combining exact keyword matching with semantic retrieval. AI agents query documents over the remote Fastio Model Context Protocol (MCP) server at https://mcp.fast.io/mcp, receiving cited passages without storing files inside model infrastructure.
Evaluating Multi-Store Router Tradeoffs
Router architectures add latency and token overhead to every request. If classification misidentifies the topic, the assistant queries the wrong vector store and returns empty results. Routing also complicates access control when different users require distinct permissions across document collections.
Hybrid Search Advantages Over Pure Vector Retrieval
For technical domains requiring exact alphanumeric lookups (such as part numbers, error codes, and contract identifiers), pure vector search often misses exact matches. Hybrid search engines score exact keyword occurrences alongside conceptual meaning to preserve retrieval accuracy. OpenAI's own search combines cosine similarity over embeddings with a keyword similarity score for this reason, and exposes weights for both under ranking_options.
How to Connect AI Agents to Scaled Workspaces via Remote MCP
Connecting an intelligent workspace to an agent pipeline eliminates file upload bottlenecks and recurring vector storage fees. Fastio provides a consolidated MCP toolset over Streamable HTTP and SSE, enabling agents built with OpenAI, Claude, or open-source frameworks to query documents directly.
Configuring Fastio Remote MCP for AI Agents
AI agents authenticate to Fastio using scoped API keys. The configuration below registers the remote MCP endpoint:
{
"mcpServers": {
"fastio": {
"url": "https://mcp.fast.io/mcp/key",
"headers": {
"Authorization": "Bearer YOUR_FASTIO_API_KEY"
}
}
}
}
Connected agents can search indexed documents, retrieve specific passages, and inspect file metadata without downloading entire files into context memory.
Automated Ingestion via Cloud Import and Sync
Populating an agent workspace does not require custom ingestion scripts. Fastio supports cloud import from Google Drive, Dropbox, Box, and OneDrive via OAuth, preserving folder structures and file names. Folders linked from Dropbox, Box, and OneDrive support on-demand and scheduled sync, while Google Drive provides one-time import with scheduled sync coming soon.
Structured Document Extraction with Metadata Views
For tabular extraction across large document sets, Fastio offers Metadata Views. Users define target fields in natural language, and the system extracts structured columns (Text, Number, Date, Boolean) into an interactive spreadsheet view that agents can query directly over MCP.
Team Collaboration and Workspace Governance
Fastio supports hybrid workflows where human team members inspect files, previews, and metadata views in the web UI, while AI agents interact programmatically via MCP or the REST API at https://api.fast.io/current/. Every file maintains version history, and real-time updates stream via WebSocket activity feeds.
Creating an account on Fastio is free. Setting up active workspaces for teams and agents requires an organization on a paid subscription. Monthly plans start with a trial of up to 30 days that requires a credit card, with plans structured as Starter at $9.99/mo, Business at $49.99/mo, and Enterprise at $199.99/mo. Explore options on Fast.io pricing. Subscriptions include persistent storage, seats, and built-in workspace intelligence.
Python Agent Tool Calling Example
Developers building agent workflows with the official openai Python SDK can register workspace retrieval tools directly into function definitions:
import os
from openai import OpenAI
client = OpenAI(api_key=os.environ.get("OPENAI_API_KEY"))
tools = [
{
"type": "function",
"function": {
"name": "storage",
"description": "Calls the consolidated Fastio storage tool with the search action to find relevant document excerpts.",
"parameters": {
"type": "object",
"properties": {
"action": {
"type": "string",
"description": "The storage action to invoke, such as search, list, or details."
},
"query": {
"type": "string",
"description": "The search query to match against document contents and metadata."
}
},
"required": ["action", "query"]
}
}
}
]
response = client.chat.completions.create(
model="gpt-4o",
messages=[
{"role": "system", "content": "You are an assistant with access to an intelligent workspace."},
{"role": "user", "content": "What are the termination notice requirements across our vendor contracts?"}
],
tools=tools
)
When the model triggers a tool call, the application executes the search against the Fastio MCP endpoint, returns matching passages with citations, and streams the grounded response.
Command-Line Scripting with Fastio CLI
For shell scripting, backups, and CI/CD pipelines, Fastio provides a command-line tool via @vividengine/fastio-cli on npm:
npm install --global @vividengine/fastio-cli
fastio auth login
fastio upload file --workspace enterprise-docs ./quarterly-brief.pdf
The CLI facilitates automated batch uploads and workspace administration without requiring custom API clients.
Sources
References used to verify factual claims in this guide.
-
The maximum file size for an OpenAI vector store is 512 MB, with up to 5,000,000 tokens per file. OpenAI bills vector storage at $0.10 per GB per day beyond the first free gigabyte across all of a project's stores. OpenAI limits file and file batch creation to 300 requests per minute per vector store, and accepts up to 500 files per batch request. OpenAI chunks each vector store file into 800-token segments with 400 tokens of overlap by default.
-
A 10,000-file ceiling per vector store is documented by Microsoft for Azure OpenAI, not by OpenAI's own current Retrieval guide.
Frequently Asked Questions
What is the file limit for an OpenAI vector store?
OpenAI's Retrieval guide caps each file at 512 MB and 5,000,000 tokens, and limits file and batch creation to 300 requests per minute per store. It does not currently publish a maximum file count per vector store. A 10,000-file ceiling is widely cited and appears in Microsoft's Azure OpenAI documentation, so treat it as a planning assumption to verify against your own account.
How much does OpenAI vector store storage cost?
OpenAI gives the first gigabyte of vector storage across all stores at no cost and charges $0.10 per GB per day beyond it, billed on the size of parsed chunks and their embeddings. A 100 GB collection therefore runs close to $300 per month. Additional costs apply for tool query calls and model prompt tokens.
How do you work around the OpenAI vector store file limit?
Engineering teams either consolidate smaller files before upload, implement a multi-store routing architecture that distributes queries across multiple vector stores, or store their corpus in an intelligent workspace like Fastio that provides hybrid search over remote MCP without file count ceilings.
Can you attach multiple vector stores to a single OpenAI assistant?
An OpenAI assistant can attach at most one vector store in its tool resources configuration. However, an individual conversation thread can also attach one vector store. During execution, the model searches across both the assistant-level and thread-level stores simultaneously, retrieving relevant candidate chunks up to defined search limits.
What happens when an OpenAI vector store hits a file ceiling?
Subsequent upload attempts or file batch associations fail with an API error. The platform does not automatically partition files into a new store or overwrite older documents. Applications must handle this error explicitly by creating a new vector store or removing inactive files.
How do vector store expiration policies work?
OpenAI lets you set an expiration policy on a vector store using the `expires_after` parameter, anchored to the store's last activity with a day count you choose. When a store expires, its associated files are deleted and you stop being charged for them. Expiration is therefore the main control on runaway index cost for stores created per conversation.
How does Fastio compare to OpenAI vector stores for large document collections?
Fastio provides persistent, shared cloud workspaces where files are automatically indexed for hybrid search (exact full-text plus semantic meaning) with no per-store file-count ceiling and no daily per-gigabyte index storage charge. Agents connect over the remote Fastio MCP server to search and retrieve source-grounded excerpts, while human team members can view files, inspect previews, and audit changes in the same workspace.
Related Resources
Store and search large document corpora without vector store ceilings
Index extensive document collections in an intelligent workspace to query files over remote MCP without per-store file caps or daily per-gigabyte vector fees. Monthly plans start with a trial of up to 30 days (credit card required); annual plans have no trial.