# Parquet File Size Limits: Format Specs, Row Group Sizing, and AI Analytics

The Apache Parquet specification imposes no format-level maximum file size, but query engines and worker memory make moderate files the practical sweet spot. Tuning row group boundaries preserves predicate pushdown benefits while avoiding memory exhaustion during decompression. For AI data agents analyzing columnar tables, querying external workspaces via MCP and pushdown engines prevents context window failures.

Source: https://fast.io/resources/parquet-file-size-limit/
Author: [Derek Labian](https://fast.io/authors/derek-labian/)
Last reviewed: 2026-10-03

## What Is the Maximum Size of an Apache Parquet File?

The Apache Parquet specification enforces no format-level maximum file size on individual files, leaving storage ceilings entirely to underlying file systems, but analytical query engines encounter performance degradation and memory failures when single files or row groups stray outside documented sizing windows. A Parquet file size limit refers to practical storage and query performance thresholds in Apache Parquet columnar files, where target file sizes of 128MB to 1GB prevent both the 'small file problem' and Out-Of-Memory exceptions during decompression.

To understand why format specifications differ from operational limits, consider the physical layout of an [Apache Parquet](https://parquet.apache.org/) file. A Parquet file begins and ends with a 4-byte magic number (`PAR1`). The interior structure consists of three sequential tiers:

1. **Row Groups:** Logical horizontal partitions of the data table containing rows grouped together.
2. **Column Chunks:** The data for a specific column within a row group, laid out contiguously on storage media.
3. **Data Pages:** Indivisible units within a column chunk where individual values, definition levels, and repetition levels are encoded and compressed.

At the tail of the file sits the File Metadata Footer, written after all row groups finish serialization. The metadata footer encodes column schemas, key-value metadata, row counts, and per-column chunk offsets using Apache Thrift binary protocols. Thrift metadata fields rely on signed 32-bit integers, which theoretically permit billions of row groups and vast data streams. However, storage engines hit physical barriers long before reaching integer limits.

When engineers deploy Parquet across production architectures, two opposing failure modes define operational boundaries:

* **The Small File Problem:** Generating thousands of files under `32 MB` overwhelms metadata catalogs and query planners. Distributed engines like Amazon Athena and Apache Spark must initiate separate HTTP GET requests or file handle open calls for each file. File system listing latency and footer parsing times quickly eclipse the time spent reading raw data.
* **The Oversized File Problem:** Creating individual files that measure several gigabytes eliminates query parallelism. When a query engine assigns work units by file rather than by row group split, single massive files create pipeline stragglers that starve executor threads and exhaust system memory buffers during decompression.

The table below illustrates format boundaries versus practical production sizing across analytical engines and storage platforms:

| Platform or Protocol | Documented Maximum File Size | Recommended Target File Size | Primary Processing Limit |
| :--- | :--- | :--- | :--- |
| Apache Parquet Specification | No format-enforced limit | `128 MB` to `1 GB` | Footer metadata parsing and decompression memory |
| Amazon Athena | 5 TB (Amazon S3 object limit) | `128 MB` to `512 MB` | Amazon S3 GET request rate limits and planner latency |
| Snowflake External Stages | 5 TB (cloud storage limit) | `100 MB` to `250 MB` (compressed) | Micro-partition alignment and scan parallelism |
| DuckDB (In-Process) | Available storage capacity | `128 MB` to `1 GB` | System memory allocations and thread vector buffers |
| Claude Chat File Upload | 500 MB per file (max 20 files) | Query via MCP instead | Text context window bounds and binary token expansion |
| Claude Projects Context | 30 MB per file | Query via MCP instead | Overall project context window token ceiling |

Understanding these format fundamentals allows engineering teams to structure data pipelines that avoid both extreme fragmentation and memory bottlenecks.

## Why Analytical Engines Target Sizing Between 128 MB and 1 GB

Analytical query engines consistently recommend target Parquet file sizes between `128 MB` and `1 GB` because this window balances I/O throughput against distributed task coordination. Whether executing queries on cloud data warehouses like Snowflake and Google BigQuery, distributed engines like Presto and Trino, or local embedded engines like DuckDB, this range optimizes four structural performance variables.

### 1. Cloud Storage Request Latency and Rate Limiting
Modern analytical architectures decouple compute from storage, hosting Parquet files on object storage tiers such as Amazon S3, Google Cloud Storage, or Azure Blob Storage. Cloud object storage exposes inherent network latency for initial byte requests. Each separate object requires an HTTP GET request and an authentication handshake.

If an analytical workload queries a dataset divided into 10,000 files of `2 MB` each, the engine must issue 10,000 independent network requests. When queries scale up, high request volumes can trigger storage rate limits, such as Amazon S3 request rate thresholds per partitioned prefix. Conversely, when that same dataset is compacted into 20 files of `1 GB` each, the engine issues only 20 requests, fetching data in sustained sequential read bursts that maximize network interface card throughput.

### 2. File Metadata Overhead and Compression Ratios
Every Parquet file contains its own metadata footer containing schema definitions, column chunk offsets, encoding dictionaries, and page index statistics. When file sizes drop below `32 MB`, this metadata block represents a disproportionate percentage of total file weight, sometimes occupying nearly a third of the storage footprint.

In contrast, within a dataset sized at `512 MB` or `1 GB`, the metadata footer occupies a fraction of 1% of the file. Larger files also provide compression algorithms like Snappy and ZSTD with wider reference windows, allowing dictionary encoding and run-length encoding to identify repeated patterns across broader row populations.

### 3. Distributed Query Planning and Executor Parallelism
Query planning engines (such as the Catalyst optimizer in Apache Spark or the Presto coordinator) examine table metadata to divide execution plans into discrete splits. When a table contains millions of small files, query compilation takes minutes before a single byte of data is processed by worker nodes.

However, sizing files beyond `2 GB` presents the opposite hazard: loss of fine-grained parallelism. If an extract-transform-load pipeline produces a single `50 GB` file containing only one row group, a query engine cannot distribute that file across multiple compute instances. Sizing files in the `128 MB` to `512 MB` range provides distributed workers with a uniform distribution of independent tasks, preventing straggler nodes from delaying query completion.

### 4. Vectorized Execution and In-Process Engine Memory
Embedded analytics engines, led by DuckDB, process Parquet files using vectorized execution pipelines that process batches of 2,048 tuples at a time. DuckDB maps Parquet column chunks directly into memory-efficient vectors.

When DuckDB reads a file sized between `128 MB` and `1 GB`, it uses HTTP range requests or memory-mapped file handles to read only the specific byte ranges corresponding to queried columns. A data analyst querying three columns out of a 100-column table only downloads and decompresses a tiny percentage of the total file size, allowing analytical workflows to run on modest local hardware without exhausting operating system RAM. Teams running multi-agent analysis can configure shared cloud storage using [Fastio workspace storage for agents](/storage-for-agents/).

## Row Group Sizing and Memory Allocation Tradeoffs

While file size dictates storage distribution and network transfer efficiency, row group sizing governs memory consumption and query filtering performance. A row group represents a horizontal slice of table rows containing individual column chunks for every column declared in the schema.

Official Apache Parquet configurations recommend large row groups between 512MB and 1GB so column chunks support sequential IO. The rationale behind this specification recommendation stems from traditional distributed file systems like Hadoop HDFS, where aligning a row group with an HDFS block allowed an entire block to be scanned sequentially without cross-node network hops.

However, modern cloud query engines have adjusted this baseline. In Amazon Athena performance tuning, the default size for Apache Parquet row groups is 128 MB. Cloud engines favor moderate row groups because multi-tenant serverless nodes must balance concurrent user queries without overflowing worker memory buffers.

Choosing the right row group target requires evaluating the architectural tradeoffs between write-time memory, read-time memory, and predicate pushdown efficiency:

* **Large Row Groups (`512 MB` to `1 GB`):**
  * **Memory Overhead:** Writing large row groups requires holding all uncompressed column data in memory until the entire row group threshold is reached. In wide tables with hundreds of columns, write buffers can easily exceed several gigabytes per thread, causing Out-Of-Memory exceptions in distributed worker containers.
  * **Compression Efficiency:** Column chunks are large, maximizing dictionary encoding reuse and compression ratios.
  * **Query Granularity:** If a query filters for specific records, the engine must decompress and evaluate large blocks of rows, increasing scan latency if min/max statistics cannot rule out the entire row group.

* **Standard Row Groups (`128 MB` to `256 MB`):**
  * **Memory Overhead:** Read and write memory allocations remain modest, fitting comfortably within standard container memory limits.
  * **Data Skipping:** Smaller row groups yield more granular min/max statistics. If data is sorted on common filter keys (such as date or customer ID), query engines can skip entire row groups without reading their contents from disk.
  * **Throughput Balance:** Provides sufficiently large byte runs for sequential I/O while protecting query workers from heap exhaustion.

* **Undersized Row Groups (`<64 MB`):**
  * **Metadata Bloat:** Each row group writes its own metadata block into the file footer. Thousands of small row groups inflate the footer size, slowing down query planners and degrading performance.
  * **Compression Degradation:** Encoders reset dictionaries between row groups, forfeiting compression efficiency and bloating file sizes.

Engineers can inspect the physical row group structure of any Parquet file using DuckDB in a terminal:

```sql
-- Inspect row group counts and physical byte distributions
SELECT
    row_group_id,
    row_group_num_rows,
    row_group_bytes,
    row_group_compressed_bytes,
    (row_group_bytes / row_group_compressed_bytes)::DECIMAL(5,2) AS compression_ratio
FROM parquet_metadata('analytics_export.parquet');
```

Running this inspection reveals whether your pipeline generates balanced columnar chunks or suffers from row group fragmentation.

## Why Uploading Large Parquet Files to AI Assistants Fails

As organizations deploy autonomous AI assistants like Claude, ChatGPT, and coding agents to perform business intelligence and automated data reporting, engineers routinely attempt to provide Parquet files directly to language models. This interaction immediately triggers platform barriers and structural incompatibilities.

Anthropic documents clear mechanical boundaries for file uploads: a standard chat conversation accepts up to `20` files at up to `500 MB` each, while Claude Projects accepts individual files up to `30 MB` each with an unlimited file count bounded by Claude's context window. While a user can physically attach a `200 MB` file into a standard chat prompt without triggering an upload size rejection, the assistant cannot parse or reason across the raw binary data.

Three structural barriers prevent large language models from reading raw Parquet files directly:

### 1. Binary Columnar Encoding Versus Token Ingestion
Large language models operate on discrete text tokens parsed by tokenizer algorithms (such as Byte-Pair Encoding). Apache Parquet is a binary format that stores data in columnar byte streams encoded with bit-packing, run-length compression, and binary dictionary pages. A language model cannot ingest a binary stream of Snappy-compressed columnar chunks directly into its attention mechanism. When an assistant attempts to parse a raw Parquet file, it encounters unreadable binary headers (`PAR1`) and encoded byte arrays, resulting in parsing errors.

### 2. Context Window Explosion from Format Conversion
To bypass binary parsing errors, analysts often convert Parquet files into text formats like CSV or JSON before uploading them to the model. This conversion triggers severe token inflation. A `100 MB` Parquet file compressed with ZSTD frequently expands to `800 MB` or more of uncompressed text.

A tabular dataset of half a million records converted to JSON expands across tens of millions of tokens, exceeding context allowances by orders of magnitude. Because leading frontier model context windows cap out between 128,000 and 1,000,000 tokens, the converted dataset exceeds context allowances by orders of magnitude. Even where long-context windows exist, stuffing raw datasets into prompts incurs prohibitive inference costs and degrades retrieval precision.

### 3. Context Window Starvation in Project Knowledge
In platforms like Claude Projects, the `30 MB` per file limit allows attaching documentation and code. However, all attached project files share the same fixed context window that also holds system prompts, custom instructions, and conversation history. Uploading large multi-megabyte structured extracts into project knowledge starves the model of working memory, causing degraded reasoning and hallucinated aggregations.

When datasets exceed small reference extracts, providing raw files directly to the model fails. The solution is not larger context windows, but an external storage and computation substrate where files remain indexed and queryable.

## Connecting AI Data Agents to Columnar Storage via MCP

Modern data teams solve the file limits barrier by decoupling columnar data storage from the AI model's context window. Instead of attaching massive files to prompt windows, teams host Parquet datasets in an external workspace and connect AI assistants using the Model Context Protocol (MCP).

In this architecture, the AI assistant acts as an analytical director rather than a raw compute engine. The files reside in a shared, org-owned Fastio workspace. Fastio provides persistent storage with per-file version history, granular permissions, and a detailed activity log. The assistant connects to Fastio through a remote MCP server hosted at `https://mcp.fast.io/mcp/code` over Streamable HTTP. You can review the protocol specifications on [Fastio workspace storage for agents](/storage-for-agents/).

When data files land in the workspace, teams configure access through two mechanisms:

1. **Structured Extraction with Metadata Views:** For documentation, schemas, and data dictionaries accompanying the dataset, [Fastio Metadata Views](/product/document-data-extraction/) extract typed metadata into a live, queryable spreadsheet without manual OCR rules or parsing code.
2. **Pushdown Analytical Execution:** The AI assistant uses its consolidated MCP toolset to inspect file directories, fetch table schemas, and dispatch targeted SQL queries to an execution engine like DuckDB.

To configure an AI assistant like Cursor to access workspace files via MCP, add the server configuration to `~/.cursor/mcp.json` (or `.cursor/mcp.json` in a project):

```json
{
  "mcpServers": {
    "fastio-workspace": {
      "url": "https://mcp.fast.io/mcp/code"
    }
  }
}
```

Sign in with OAuth in the browser when Cursor connects. The Review Permissions screen lets you select Read Only or Read & Write access and choose which organizations and workspaces the connection can reach. For Claude Code, run `claude mcp add --transport http fastio-workspace https://mcp.fast.io/mcp/code`, then run `/mcp` inside Claude Code to sign in. Setup steps are documented on the [Fastio MCP documentation](https://mcp.fast.io/docs).

Once connected, the agent executes data analysis through an efficient pushdown loop:

* **Step 1 (Schema Inspection):** The user asks, "What was our total cloud infrastructure spend by provider in Q2?" The agent calls Fastio workspace tools to inspect file metadata and retrieve table column names.
* **Step 2 (Targeted Query Execution):** Instead of downloading the full `800 MB` file, the agent uses an in-process query engine (like DuckDB or a cloud data warehouse API) to execute a targeted aggregate query against the remote storage location:

```sql
SELECT
    cloud_provider,
    ROUND(SUM(invoice_amount), 2) AS total_spend
FROM 'fastio://workspaces/data-lake/q2_billing.parquet'
GROUP BY cloud_provider
ORDER BY total_spend DESC;
```

* **Step 3 (Precise Context Ingestion):** The query engine processes the Parquet file using row group min/max statistics, reads only the `cloud_provider` and `invoice_amount` column chunks, and returns an aggregate table of five rows.
* **Step 4 (Synthesis):** The AI assistant receives the concise five-row summary into its context window (consuming minimal token overhead) and generates a complete, accurate response with zero hallucination.

By keeping the raw Parquet file in external storage and querying row groups on demand, the assistant analyzes datasets of any physical size without encountering prompt limits or context degradation. Review developer instructions on [Fastio agent onboarding](https://fast.io/llms.txt) to see how agents interact with workspace storage.

## Best Practices for Parquet Compaction and Multi-Agent Collaboration

Operating high-volume Parquet pipelines and multi-agent analytical workspaces requires proactive engineering to prevent file bloat, data corruption, and concurrency conflicts. Implementing established maintenance patterns keeps data pipelines fast and reliable.

### 1. Automated Compaction Routines
Streaming pipelines (such as web event loggers, IoT telemetry collectors, or CDC database replicators) write micro-batches every few seconds, generating thousands of small files. To prevent the small file problem:

* Implement scheduled compaction jobs using Apache Iceberg `OPTIMIZE` commands or DuckDB scheduled tasks to merge files smaller than `64 MB` into consolidated `256 MB` to `512 MB` partitions.
* Partition data by moderate-cardinality keys (such as `year` and `month`). Avoid partitioning by high-cardinality attributes like `user_id` or `timestamp`, which creates millions of isolated single-record directories.

### 2. Data Sorting for Predicate Pushdown
Parquet row group headers store minimum and maximum values for every column chunk. When data is written in random order, min/max intervals overlap across all row groups, forcing query engines to scan every row group in the file.

Before writing Parquet datasets, sort the records by your primary query filter column (such as `tenant_id` or `created_at`). When sorted data is written, row groups contain discrete, non-overlapping value ranges. Analytical engines like Athena and DuckDB can then skip the vast majority of physical row groups during query execution, reducing disk I/O and query latency.

### 3. Selecting the Optimal Compression Codec
Parquet supports several compression codecs, each balancing compression ratio against CPU decompression speed:

* **Snappy:** The historical default in Hadoop ecosystems. Snappy provides rapid decompression speeds with modest CPU overhead, but yields lower compression ratios.
* **ZSTD (Zstandard):** The modern production standard. ZSTD provides substantially higher compression ratios than Snappy while maintaining competitive decompression speeds. Setting ZSTD to compression level 3 offers the best overall throughput for analytical workloads.
* **Gzip:** Delivers high compression ratios but suffers from slow decompression speeds and high CPU utilization. Avoid Gzip for interactive analytics.

### 4. Managing Multi-Agent Concurrency and Workspace Handoffs
When multiple AI agents and human analysts collaborate within a shared data workspace, uncoordinated writes can cause race conditions and corrupted file versions.

In [Fastio workspaces](/product/workspaces/), teams coordinate concurrent activity using built-in collaboration controls:

* **Advisory File Leases:** Agents acquire advisory file locks (`lock-acquire`, `lock-status`, `lock-release`) through the storage MCP tools before initiating file transformations. Leases signal active write operations to other workspace participants without hard-locking storage.
* **Per-File Version History:** Every file modification creates a distinct version in history. If an automated script overwrites a dataset incorrectly, team members can review changes in the audit log and restore previous versions.
* **Clean Ownership Transfer:** When an automated data engineering agent completes pipeline setup, it transfers workspace ownership to a human administrator. Monthly plans start with a 30-day free trial that requires a credit card. Plans are Starter at `$9.99/mo`, Business at `$49.99/mo`, and Enterprise at `$199.99/mo`. Explore plan details on the [Fastio pricing page](/pricing/).

Applying these operational standards keeps Parquet datasets performant, cost-effective, and fully accessible to AI assistants and human teams alike.

## Frequently asked questions

### What is the maximum size of an Apache Parquet file?

The Apache Parquet specification imposes no hard upper limit on single file size. Physical file limits are determined by the underlying storage layer, such as Amazon S3's `5 TB` object ceiling or local file system bounds. In production analytics, query engines recommend keeping files between `128 MB` and `1 GB` to avoid memory bottlenecks and maintain distributed parallelism.

### What is the optimal Parquet file size for performance?

The optimal Parquet file size for query engines like Amazon Athena, Snowflake, and DuckDB is between `128 MB` and `512 MB`, extending up to `1 GB` for large batch workloads. Files smaller than `64 MB` cause the small file problem with excessive metadata overhead, while files larger than `2 GB` reduce parallel split distribution.

### How large should a Parquet row group be?

Recommended Parquet row group sizes range from `128 MB` to `512 MB`. While the official Apache Parquet specification recommends large row groups between `512 MB` and `1 GB` for sequential I/O, modern query engines like Amazon Athena set the default row group size to `128 MB` to lower memory buffer requirements and improve predicate pushdown filtering.

### Can AI assistants read large Parquet files directly?

AI assistants cannot parse raw binary Parquet files uploaded to chat. Claude accepts chat uploads up to `500 MB` per file (max `20` files) and Claude Projects accepts files up to `30 MB` each, but Parquet is a compressed columnar binary format that models cannot deserialize as text tokens. Instead, teams host files in external workspaces and connect assistants via MCP.

### How do you fix the small file problem in Parquet data lakes?

You resolve the small file problem by running automated compaction jobs that merge fragmented files into target `128 MB` to `512 MB` files. In modern table formats like Apache Iceberg, execute the OPTIMIZE command with REWRITE DATA. In SQL engines like Amazon Athena or DuckDB, use CREATE TABLE AS SELECT queries to consolidate micro-partitions.

### What is the difference between Parquet file size and row group size?

A Parquet file is a complete physical file stored on disk, whereas a row group is an internal logical partition of rows within that file containing columnar chunks. A single `512 MB` Parquet file typically contains four `128 MB` row groups, enabling query engines to skip irrelevant row groups using column min/max statistics.

## Sources

- [Apache Parquet Documentation: Configurations](https://parquet.apache.org/docs/file-format/configurations/): Official Apache Parquet configurations recommend large row groups between 512MB and 1GB so column chunks support sequential IO.
- [AWS Documentation: Performance tuning data optimization techniques](https://docs.aws.amazon.com/athena/latest/ug/performance-tuning-data-optimization-techniques.html): In Amazon Athena performance tuning, the default size for Apache Parquet row groups is 128 MB.

## About Fast.io

Fast.io provides shared workspaces where people and AI agents work on the same files, with built-in semantic search and citation-backed chat over what they hold. Agents reach it through a remote MCP server, a REST API at https://api.fast.io/current/, and a command line client published on npm as @vividengine/fastio-cli. MCP setup is at https://mcp.fast.io/docs: Claude and most MCP clients connect to https://mcp.fast.io/mcp/tools, ChatGPT to https://mcp.fast.io/mcp/operations, and coding agents to https://mcp.fast.io/mcp/code.
