AI & Agents

AWS Lambda File Size Limits: Payload, Package, and Ephemeral Storage Caps

An AWS Lambda file size limit encompasses the execution constraints of AWS Lambda functions, specifically the 50MB direct zip upload limit, 250MB unzipped code limit, 10GB ephemeral /tmp storage ceiling, and 6MB synchronous invocation payload threshold. Processing files larger than these boundaries requires decoupling payloads using Amazon S3, response streaming, or external workspaces. Autonomous AI agents can query large corpora via MCP without hitting Lambda memory or payload caps.

Tom Langridge 15 min read Updated
AWS Lambda limits synchronous invocation payloads to 6MB while supporting up to 10GB of configurable ephemeral /tmp storage.

What Are the AWS Lambda File Size Limits?

An AWS Lambda file size limit encompasses the execution constraints of AWS Lambda functions, specifically the 50MB direct zip upload limit, 250MB unzipped code limit, 10GB ephemeral /tmp storage ceiling, and 6MB synchronous invocation payload threshold. As documented in the AWS Lambda Developer Guide, AWS Lambda limits synchronous invocation payloads strictly to 6 MB for request and response bodies. Attempting to transmit a payload exceeding 6 MB directly to a Lambda function causes the runtime to terminate the invocation immediately with an HTTP 413 Payload Too Large error and an unhandled RequestEntityTooLargeException.

Understanding AWS Lambda boundaries requires separating four architectural layers that engineering teams frequently conflate:

  1. Invocation Payload: The volume of request and response data passed directly through the function event object during execution. Synchronous invocations cap payloads at 6 MB, whereas direct asynchronous invocations allow up to 1 MB.

  2. Deployment Package: The bundle of compiled code, libraries, and custom runtimes uploaded to AWS Lambda. Direct archive uploads through the AWS Console, AWS CLI, or SDKs are capped at 50 MB zipped. When staged in Amazon S3, uncompressed archives and attached Lambda Layers can total up to 250 MB. For larger workloads, container images stored in Amazon Elastic Container Registry (Amazon ECR) support up to 10 GB.

  3. Ephemeral Disk Storage (/tmp): The local scratch disk attached to each execution environment. The default allocation is 512 MB, but administrators can configure storage up to 10,240 MB (10 GB) in 1-MB increments.

  4. Execution Memory (RAM): The volatile memory allocated to function execution, configurable from 128 MB to 10,240 MB. Because AWS Lambda allocates CPU capacity in direct proportion to configured memory, processing large files in memory often requires increasing memory allocation even if the file fits on disk.

The following comparison table outlines the primary file size and storage limits across AWS Lambda execution layers:

Resource Dimension Default Quota Maximum Configurable Limit Primary Failure Mechanism Verification
Synchronous Invocation Payload 6 MB 6 MB (Hard limit) HTTP 413 Payload Too Large / RequestEntityTooLargeException Verified October 2026
Asynchronous Invocation Payload 1 MB 1 MB (Hard limit; 256 KB via SQS/EventBridge) Message rejection or silent Dead Letter Queue routing Verified October 2026
Response Streaming Payload 6 MB (Standard) 200 MB (Streamed, throttled to 2 MBps after 6 MB) Stream cutoff or socket termination Verified October 2026
Direct Deployment Package (.zip) 50 MB (Zipped) 50 MB (Via Console, API, or SDK) InvalidParameterValueException / Upload rejected Verified October 2026
S3-Staged Deployment Package 50 MB 250 MB (Unzipped code plus layers) Unzipped size limit exceeded during deployment Verified October 2026
Container Image Deployment None 10 GB (Uncompressed image via ECR) Image manifest validation failure Verified October 2026
Ephemeral Storage (/tmp) 512 MB 10,240 MB (10 GB in 1-MB increments) ENOSPC: no space left on device Verified October 2026
Function Execution Memory 128 MB 10,240 MB (10 GB in 1-MB increments) Out of memory / Runtime exited without response Verified October 2026

Confusing these boundaries leads to fragile architectures. A common misconception is assuming that increasing Lambda execution memory to 10 GB or configuring /tmp to 10 GB enables the function to accept a 50 MB HTTP upload directly. Because the 6 MB invocation payload limit is an immutable hard boundary, file transfer pipelines must decouple network transit from execution storage.

AWS Lambda Payload Limits: Synchronous, Asynchronous, and Streaming Boundaries

The 6 MB synchronous payload limit is the most frequent obstacle in serverless file processing pipelines. When invoking a Lambda function synchronously via the AWS SDK, API Gateway, Application Load Balancers, or Lambda Function URLs, the combined size of the request payload cannot exceed 6,291,456 bytes. The identical 6 MB ceiling applies to the response payload returned by the handler.

Architectures fronted by Amazon API Gateway introduce additional confusion because API Gateway enforces a maximum request payload size of 10 MB. If a client transmits an 8 MB payload to API Gateway, the gateway accepts the HTTP connection, validates the request, and attempts to invoke the backend Lambda integration. Lambda immediately rejects the request with an HTTP 413 error because 8 MB exceeds Lambda's 6 MB ceiling. The client receives an error, and the function handler never executes.

Handling binary data within JSON payloads introduces another hidden payload inflation. Because JSON does not natively support raw binary buffers, applications encode binary data as Base64 strings. Base64 encoding expands raw byte volume. A binary file approaching 5 MB breaches the 6 MB ceiling once converted into Base64 text within a JSON request body. Consequently, attempting to transmit media or documents in synchronous JSON events triggers immediate payload violations.

Direct asynchronous invocations through the AWS Lambda API (InvocationType: Event) support request payloads up to 1 MB. When invoked asynchronously, Lambda places the event into an internal queue and returns an HTTP 202 Accepted status code immediately. However, upstream event sources impose much stricter constraints:

  • Amazon EventBridge enforces a maximum event payload of 256 KB.
  • Amazon Simple Queue Service (Amazon SQS) caps standard and FIFO messages at 256 KB.
  • Amazon Simple Notification Service (Amazon SNS) limits message payloads to 256 KB.

When Lambda functions trigger from SQS queues or EventBridge rules, the practical asynchronous payload limit is 256 KB rather than 1 MB. Passing larger data through these services requires payload offloading or S3-backed message pointers.

To alleviate response size constraints for web clients, AWS Lambda supports response streaming up to 200 MB for Node.js and custom runtimes. Using the awslambda.streamifyResponse wrapper, functions stream response chunks directly through Lambda Function URLs without buffering the entire output in memory. AWS provides uncapped bandwidth for the first 6 MB of the stream, after which throughput is throttled to 2 MBps for the remainder of the 200 MB response. While response streaming resolves egress bottlenecks for large reports, documents, or media files, it remains strictly unidirectional: it provides no mechanism to stream large request payloads into a Lambda function.

Deployment Package Size Caps: Direct Zip, Unzipped Packages, and Containers

Packaging application code and runtime dependencies for AWS Lambda requires managing strict artifact size ceilings. When creating or updating a function through the AWS Management Console, the AWS CLI, or an AWS SDK, direct .zip archive uploads are limited to 50 MB. Attempting to upload a compressed archive of 51 MB through direct API calls returns an InvalidParameterValueException.

To deploy zip archives larger than 50 MB, developers must stage the archive in an Amazon S3 bucket within the same AWS Region. The deployment command then passes the S3Bucket and S3Key parameters instead of transmitting the zip bytes directly. Staging the artifact in S3 bypasses the 50 MB upload threshold, but the uncompressed contents of the package must remain within the 250 MB ceiling.

The 250 MB unzipped limit applies to the combined file footprint of the function code, custom runtimes, and all attached Lambda Layers. AWS Lambda permits up to five layers per function. During cold start container initialization, Lambda extracts the function archive into the /var/task directory and extracts attached layers into /opt. If the cumulative uncompressed footprint of /var/task and /opt exceeds 250 MB, the deployment fails.

The following AWS CLI command demonstrates the error returned when updating a function with an uncompressed archive exceeding the size limit:

aws lambda update-function-code \
    --function-name process-document \
    --s3-bucket build-artifacts-prod \
    --s3-key lambdas/process-document.zip

When the unzipped package exceeds limits, the command returns an InvalidParameterValueException stating that unzipped size must be smaller than 262144000 bytes.

Python and Node.js applications that rely on heavy third-party packages frequently breach the 250 MB boundary. Data science packages such as NumPy, SciPy, Pandas, PyTorch, and headless browser binaries like Chromium routinely expand to hundreds of megabytes. Teams can reduce package volume using several optimization techniques:

  • Stripping shared library symbols: Running strip --strip-unneeded *.so across compiled C and Rust extensions removes debug symbols and reduces binary file footprint.
  • Purging unused assets: Removing __pycache__ folders, .pyc files, test directories (tests/), documentation (doc/), and unneeded native platform binaries prior to zipping.
  • Omitting development dependencies: Excluding build tools, linters, and type definitions from production deployment bundles.

When an application footprint cannot be compressed below 250 MB, container images provide an alternative packaging model. AWS Lambda supports container images up to 10 GB in size, packaged using Docker or OCI specifications and stored in Amazon ECR. AWS uses optimized caching and block-level fetching to deploy container layers across microVM hosts efficiently. However, container images with large files experience longer initial cold start latencies, and pushing multi-gigabyte images slows down automated deployment pipelines.

Fastio features

Index and Query Large Corpora Beyond Serverless Payload Caps

Store large datasets in persistent workspaces with automatic semantic search, per-file version history, and remote MCP tooling for autonomous AI agents. Monthly plans start with a 30-day free trial.

Ephemeral Storage: Configuring 512MB to 10GB for File Processing in /tmp

Every AWS Lambda execution environment includes a dedicated scratch directory mounted at /tmp. As documented in the AWS Lambda Developer Guide, AWS Lambda allows configuring ephemeral /tmp storage between 512 MB and 10,240 MB in 1-MB increments. The default allocation is 512 MB. Allocating additional scratch space allows functions to download, extract, and manipulate files up to 10 GB without attaching external block volumes.

The first 512 MB of /tmp storage is included in standard Lambda execution pricing. Storage configured above 512 MB is billed based on gigabyte-seconds: the amount of extra storage allocated multiplied by the execution duration in milliseconds. Because pricing scales with allocation, functions should configure only the disk capacity required for their specific workload.

Understanding the lifecycle of /tmp is essential for application stability. The /tmp volume is strictly ephemeral and tied to a single microVM execution environment. If an execution environment handles multiple warm invocations sequentially, files written to /tmp during the first invocation remain present during subsequent executions. However, AWS Lambda can terminate or replace warm environments at any moment without warning. When incoming traffic scales horizontally, new concurrent instances receive clean, empty /tmp directories. Applications must never rely on /tmp for state persistence across invocations or treat it as a shared cache between parallel workers. Data stored in /tmp is encrypted at rest using an AWS-managed key.

The primary architectural pattern for processing files larger than the 6 MB payload limit is the Claim Check pattern using Amazon S3 and /tmp. Instead of routing file bytes through HTTP requests, the application isolates file storage from function invocation:

  1. Pre-Signed Upload: A client requests an upload authorization from an API. The backend generates an Amazon S3 pre-signed URL with an expiration window and returns it to the client.
  2. Direct S3 Upload: The client performs an HTTP PUT request directly against the S3 pre-signed URL, transmitting files of several gigabytes without involving serverless compute.
  3. Event Trigger: Amazon S3 emits an s3:ObjectCreated:* event, which invokes the Lambda function asynchronously. The event payload contains only S3 bucket and object key metadata, consuming minimal network bytes.
  4. Local Disk Processing: The Lambda function reads the object from S3, streams the byte data into /tmp, executes transformation logic (such as video transcoding, image compression, or file parsing), and uploads the output to a destination S3 bucket.
  5. Cleanup: The function removes temporary files from /tmp before completing to preserve disk headroom for subsequent warm invocations.

When streaming large files into /tmp, developers must balance disk capacity against function RAM. If a function is configured with 10 GB of /tmp but only 512 MB of memory, attempting to read a multi-gigabyte file into memory using file.read() causes an immediate memory exhaustion crash. File processing pipelines must use stream-based IO, piping chunks directly from network sockets to disk to maintain a low, constant memory footprint.

Why Large File Sets Create Bottlenecks for AI Agents in Serverless Workflows

Serverless architectures are increasingly used to host and support autonomous AI coding agents, data extraction pipelines, and automated developer tooling. Agents including Claude Code, Cursor, Codex, OpenClaw, and Cline inspect project repositories, analyze system logs, and summarize large collections of documentation. However, combining autonomous agents with serverless runtimes exposes severe friction points around file handling and context constraints.

The fundamental challenge is that Large Language Models operate on token context rather than block storage. While an AWS Lambda function with 10 GB of /tmp can download large archive files, passing that raw unstructured data into an AI model creates steep operational bottlenecks.

This constraint is evident in the documented file limits of leading AI platforms. As documented in Anthropic Claude help documentation, a standard chat accepts a maximum of twenty files at up to 500 MB each, while a Claude Project accepts individual files up to 30 MB each with no fixed file-count limit, subject to the overall context window. Claude Projects has no fixed project file count ceiling, so the context window itself forms the practical operational boundary. Passing extensive documentation sets or multi-gigabyte codebases directly into an agent context window triggers rapid degradation:

  • Token Exhaustion and High Latency: Passing hundreds of thousands of words into an LLM burns token allowances rapidly and increases inference latency, turning rapid agent iterations into sluggish batch jobs.
  • Lost-in-the-Middle Degradation: When LLM context windows are flooded with massive, unstructured file dumps, retrieval accuracy declines. Models struggle to locate critical edge-case specifications buried deep within irrelevant boilerplate code.
  • Serverless Execution Churn: If a Lambda function downloads and extracts a 1 GB documentation archive on every invocation, the function spends most of its 15-minute execution window performing repetitive network downloads and chunking operations, inflating cloud bills.
  • Ephemeral Context Loss: Because Lambda /tmp storage is wiped whenever execution environments rotate, parsed embeddings or intermediate extraction indices cannot be reused by subsequent agent runs without dedicated external storage.

Attempting to solve this problem by piping multi-megabyte files through synchronous Lambda invocations causes immediate HTTP 413 failures. Conversely, storing raw archives in Amazon S3 leaves the AI agent with the burden of downloading, parsing, and chunking files on every interaction. To operate efficiently, AI agents require an intelligent persistence layer that indexes data on arrival and exposes structured search endpoints.

How to Connect Serverless Agent Workflows to Persistent Workspaces

Fast.io resolves the architectural mismatch between serverless compute limits and AI agent context requirements. Instead of altering AWS Lambda quotas, Fast.io provides shared, organization-owned workspaces that act as the persistent collaboration and intelligence layer for both AI agents and human engineers.

In this architecture, heavy datasets, code repositories, and project assets reside in Fast.io workspaces rather than passing through fragile Lambda invocation payloads. Fast.io supports cloud import from Google Drive, Dropbox, Box, and OneDrive, as well as direct URL imports. Cloud Sync is available for Dropbox, Box, and OneDrive on a schedule or on demand, with Google Drive imports operating today and sync coming soon.

When Intelligence Mode is enabled on a workspace, incoming documents are automatically indexed for hybrid search, which combines full-text keyword matching with semantic vector search. Rather than downloading multi-gigabyte files into Lambda /tmp or stuffing raw text into prompt contexts, an AI agent connects to Fast.io through the remote Model Context Protocol (MCP) server at https://mcp.fast.io/mcp/tools over Streamable HTTP, signing in with OAuth in the browser or sending an Authorization: Bearer <api key> header on the connection, as detailed in our guide on storage for agents. The agent issues targeted semantic queries and retrieves only the precise text passages and source citations required for its immediate task.

For structured extraction from technical specifications, invoices, contracts, and architecture diagrams, Fast.io provides Metadata Views. Metadata Views convert unstructured documents into a live, queryable relational spreadsheet. Users describe the desired fields in natural language, and the system extracts typed schemas (text, numbers, booleans, dates, and JSON) without requiring manual OCR templates. Autonomous agents can create Views, trigger automated extraction, and inspect structured data directly through MCP tooling.

The following Python script illustrates how a serverless function or autonomous agent queries a Fast.io workspace for relevant technical context using the standard httpx HTTP library:

import httpx

FASTIO_WORKSPACE_ID = "ws_serverless_docs_102"
FASTIO_API_KEY = "your-fastio-api-key"
SEARCH_ENDPOINT = f"https://api.fast.io/current/workspace/{FASTIO_WORKSPACE_ID}/storage/search/"

def query_workspace_context(query_text: str) -> list[dict]:
    headers = {
        "Authorization": f"Bearer {FASTIO_API_KEY}",
        "Content-Type": "application/json"
    }
    params = {
        "search": query_text,
        "limit": 5
    }
    with httpx.Client(timeout=30.0) as client:
        response = client.get(SEARCH_ENDPOINT, headers=headers, params=params)
        response.raise_for_status()
        results = response.json().get("data", {}).get("items", [])
        return [
            {
                "file_name": item.get("name"),
                "file_id": item.get("id"),
                "score": item.get("relevance_score"),
                "snippet": item.get("matching_snippet")
            }
            for item in results
        ]

To coordinate concurrent agent and human workflows, Fast.io workspaces provide per-file version history, a detailed activity log, and advisory file locks (lock-acquire, lock-status, lock-release) accessible via MCP. Teams can also collaborate on system designs and documentation in real time using Collaborative Notes.

Engineering teams can evaluate persistent workspaces for their serverless agent pipelines with a 14-day trial. Monthly plans start with a 30-day free trial that requires a credit card, providing scalable workspace capacity across Starter, Business, and Enterprise plans.

Sources

References used to verify factual claims in this guide.

  1. AWS Lambda limits synchronous invocation payloads strictly to 6 MB for request and response bodies, restricting payload and file size limit thresholds. AWS Lambda allows configuring ephemeral /tmp storage between 512 MB and 10,240 MB.

Frequently Asked Questions

What is the maximum file size you can process in AWS Lambda?

The maximum file size you can process in AWS Lambda depends on storage configuration. While synchronous payloads are strictly capped at `6 MB`, an AWS Lambda function can process files up to `10 GB` by configuring ephemeral `/tmp` storage between 512 MB and 10,240 MB and streaming data directly from Amazon S3. For datasets exceeding `10 GB`, functions must mount Amazon Elastic File System (Amazon EFS) or process data in streaming chunks.

How do I bypass the AWS Lambda 6MB payload limit?

To bypass the 6MB AWS Lambda payload limit, implement the S3 Claim Check pattern. Instead of transmitting file bytes directly in the invocation payload, clients upload the file to Amazon S3 using a pre-signed URL. S3 then triggers the Lambda function asynchronously with an event containing only the bucket name and object key. For large egress data, enable response streaming to return payloads up to `200 MB` to web clients.

Can AWS Lambda download and process a file sized at 1 GB in ephemeral storage?

Yes, AWS Lambda can download and process a file sized at 1 GB if its ephemeral `/tmp` storage is configured above `1,024 MB` (up to the `10,240 MB` maximum limit). The download must stream the bytes directly from an external source, such as Amazon S3 or an HTTP endpoint, directly to the `/tmp` disk volume. The function code should use streaming IO rather than loading the entire file into memory to avoid exceeding RAM limits.

What happens when an invocation payload exceeds 6MB in AWS Lambda?

When an AWS Lambda synchronous invocation payload exceeds 6,291,456 bytes (`6 MB`), the runtime immediately terminates the request and returns an HTTP 413 Payload Too Large status code with a RequestEntityTooLargeException error. The function execution handler never runs, and no execution metrics or function logs are generated in Amazon CloudWatch.

Does increasing Lambda memory increase ephemeral /tmp storage?

No, memory allocation and ephemeral `/tmp` storage are configured independently in AWS Lambda. Memory can be set between `128 MB` and `10,240 MB`, while ephemeral `/tmp` storage is configured separately between 512 MB and 10,240 MB. Increasing memory scales CPU power and RAM but does not increase disk space in the `/tmp` directory.

How do AI agents process multi-gigabyte documents without hitting Lambda payload caps?

AI agents process multi-gigabyte documents by connecting to an external persistent workspace rather than passing files through serverless payloads or LLM context prompts. By staging documents in a Fast.io workspace with Intelligence Mode enabled, files are automatically indexed for hybrid semantic search. AI agents query the workspace via the remote Model Context Protocol server (detailed in our guide on [storage for agents](/storage-for-agents/)) to retrieve relevant passages with source citations without loading raw files into context.

Related Resources

Fastio features

Index and Query Large Corpora Beyond Serverless Payload Caps

Store large datasets in persistent workspaces with automatic semantic search, per-file version history, and remote MCP tooling for autonomous AI agents. Monthly plans start with a 30-day free trial.