# How to Configure MinIO Object Storage for Dify AI Workflows

A Dify MinIO integration connects self-hosted Dify LLM orchestration instances to private S3-compatible MinIO object storage for persisting document datasets, vector indexes, and agent artifacts. While local Docker volumes drop files during container teardowns, object storage decouples application compute from dataset persistence. This guide explains how to configure Dify environment variables for MinIO, eliminate upload bottlenecks, and manage large agent knowledge bases.

Source: https://fast.io/resources/dify-minio-storage-integration/
Author: [Tom Langridge](https://fast.io/authors/tom-langridge/)
Last reviewed: 2026-09-24

## Why Production Dify Workflows Require Dedicated Object Storage

Deploying self-hosted Dify with local filesystem storage works for initial prototyping, but it creates an operational failure the moment teams scale to multi-container workers, high-concurrency agent workflows, or production knowledge bases. Local Docker volumes fail to share state across decoupled API and Celery worker replicas, trigger silent upload failures when container storage fills, and trap agent execution artifacts inside ephemeral instances.

A Dify MinIO integration connects self-hosted Dify LLM orchestration instances to private S3-compatible MinIO object storage for persisting document datasets, vector indexes, and agent artifacts.

In a default local installation, Dify relies on Apache OpenDAL configured with local filesystem storage (`OPENDAL_SCHEME=fs`) or standard Docker volume mounts. Dify is an enterprise-grade LLM orchestration engine composed of several distinct container services: the core Web and API backend, background Celery asynchronous workers, a Redis cache, PostgreSQL metadata storage, vector engines (such as Weaviate, Qdrant, or Milvus), and an isolated code execution sandbox. 

When users or automated agents upload source documents into Dify Knowledge Bases (Datasets), the ingestion process breaks across isolated container boundaries under local filesystem storage:

1. Asynchronous Ingestion Splits. The Dify API container receives the raw document upload over HTTP. However, text extraction, segmentation, chunking, and embedding generation are dispatched to background Celery worker containers. If these containers do not share an identical, high-performance network volume with synchronized POSIX permissions, worker tasks fail with missing file errors.
2. Ephemeral Storage Risks. Container lifecycle events, such as upgrading Dify via Docker Compose, rebuilding images, or scaling horizontally across Kubernetes pods, risk wiping locally mounted directories if volume paths are misconfigured.
3. Concurrency Contention. High-throughput multi-agent workflows generate concurrent reads and writes. Local disk I/O quickly saturates when multiple agents upload large PDF datasets, generate output images, and write intermediate workflow artifacts simultaneously.

MinIO solves these architectural vulnerabilities by providing a dedicated, high-performance, S3-compatible object storage layer running directly within your private network or on-premises infrastructure. Because MinIO speaks the standard AWS S3 API protocol, Dify treats MinIO as an S3 endpoint, gaining enterprise resilience, data immutability, and horizontal scalability without leaking proprietary data to public cloud vendors.

## How to Configure Dify Environment Variables for MinIO S3 Storage

Connecting self-hosted Dify to MinIO requires configuring S3 storage parameters in your deployment environment. When deploying Dify using Docker Compose, platform configuration is managed within the `docker/.env` file, cloned from `docker/.env.example`.

To switch Dify storage from local disk to MinIO, set the primary storage provider and populate the S3 connection parameters:

```env
STORAGE_TYPE=s3
S3_ENDPOINT=http://minio:9000
S3_BUCKET_NAME=difyai
S3_ACCESS_KEY=your_minio_access_key
S3_SECRET_KEY=your_minio_secret_key
S3_REGION=us-east-1
S3_ADDRESS_STYLE=path
S3_USE_AWS_MANAGED_IAM=false
```

Understanding how Dify parses each parameter ensures a clean deployment:

* `STORAGE_TYPE`: Instructs Dify to use the S3 storage adapter rather than local disk or alternative cloud providers. Setting this to `s3` activates the underlying S3 driver.
* `S3_ENDPOINT`: Defines the network address of your MinIO instance. If MinIO runs as a service in the same `docker-compose.yaml` file, use the internal service name and port (`http://minio:9000`). If MinIO runs on an external server, supply the fully qualified domain name or internal IP address (for example, `https://minio.internal.net:9000`).
* `S3_BUCKET_NAME`: Specifies the target bucket where Dify stores uploaded documents, generated audio, and workflow artifacts. Dify defaults to `difyai`. Create this bucket in MinIO before launching Dify containers, as Dify does not always auto-create target buckets.
* `S3_ACCESS_KEY` and `S3_SECRET_KEY`: The authentication credentials used to sign S3 API requests. In MinIO, these correspond to your root user credentials or a dedicated service account identity.
* `S3_REGION`: The S3 region string. While MinIO is self-hosted and does not enforce geographic regions, underlying S3 SDKs require a non-empty string. Setting `us-east-1` prevents client initialization exceptions.
* `S3_ADDRESS_STYLE`: Controls URL construction. For MinIO, specify `path` (or `auto`) so requests format as `http://minio:9000/difyai/object-key`. Using virtual-hosted styling (`difyai.minio:9000`) fails in local Docker networks because DNS cannot resolve bucket subdomains without custom wildcard routing.
* `S3_USE_AWS_MANAGED_IAM`: Set to `false`. This disables AWS IAM metadata service lookups, forcing Dify to authenticate using the provided static access and secret keys.

To run MinIO alongside Dify within the same Docker network, define the MinIO service in `docker/docker-compose.yaml`:

```yaml
services:
  minio:
    image: quay.io/minio/minio:RELEASE.2024-09-22T00-33-43Z
    container_name: dify-minio
    restart: always
    command: server /data --console-address ":9001"
    ports:
      - "9000:9000"
      - "9001:9001"
    environment:
      MINIO_ROOT_USER: your_minio_access_key
      MINIO_ROOT_PASSWORD: your_minio_secret_key
    volumes:
      - minio_data:/data
    networks:
      - ssrf_proxy_network
      - default

volumes:
  minio_data:
    driver: local
```

After updating `.env` and `docker-compose.yaml`, restart the Dify stack:

```bash
docker compose down
docker compose up -d
```

Verify that the `api` and `worker` containers initialize without storage errors:

```bash
docker compose logs api
docker compose logs worker
```

## How to Resolve Upload Bottlenecks, Nginx Limits, and CORS Policies

Most introductory guides stop at basic environment variables, leaving engineering teams to encounter severe production failures when real users upload multi-megabyte datasets or when browser frontends attempt to render stored assets. Three specific operational issues frequently disrupt Dify MinIO deployments:

### 1. Nginx 413 Payload Errors on Large Document Uploads

The Dify platform restricts document uploads to `15 MB` via the `UPLOAD_FILE_SIZE_LIMIT` parameter. When knowledge base curators upload extensive technical manuals, legal disclosures, or image-heavy presentations, the application rejects the request.

However, simply increasing `UPLOAD_FILE_SIZE_LIMIT` is insufficient. Dify routes all frontend traffic through an internal Nginx reverse proxy. Nginx enforces its own upload barrier via `NGINX_CLIENT_MAX_BODY_SIZE`, which defaults to `100M`. If you configure `UPLOAD_FILE_SIZE_LIMIT=120` without modifying Nginx, Nginx intercepts the upload before it ever reaches Dify or MinIO, returning an immediate HTTP 413 Request Entity Too Large error.

To support large file processing, update both variables in `docker/.env`:

```env
UPLOAD_FILE_SIZE_LIMIT=100
NGINX_CLIENT_MAX_BODY_SIZE=150M
```

### 2. Cross-Origin Resource Sharing (CORS) Failures

When Dify workflows generate preview URLs or when browser clients directly fetch documents stored in MinIO, cross-origin security checks can block the response. This manifests as missing image previews, failed audio playback in conversational apps, or blocked file downloads in the Dify Console.

To fix CORS across your architecture, configure policies in both Dify and MinIO. In Dify's `.env`, specify permitted origins:

```env
CONSOLE_CORS_ALLOW_ORIGINS=*
WEB_API_CORS_ALLOW_ORIGINS=*
```

For production deployments, replace the wildcard `*` with your explicit frontend domains (for example, `https://dify.example.com`).

Next, configure the CORS policy directly on the MinIO bucket using the official MinIO Client (`mc`):

```bash
mc alias set localminio http://localhost:9000 your_minio_access_key your_minio_secret_key

cat << 'EOF' > cors.json
[
  {
    "AllowedOrigins": ["https://dify.example.com", "http://localhost:3000"],
    "AllowedMethods": ["GET", "PUT", "POST", "DELETE", "HEAD"],
    "AllowedHeaders": ["*"],
    "ExposeHeaders": ["ETag", "Content-Type", "Content-Length"]
  }
]
EOF

mc cors set localminio/difyai cors.json
mc cors get localminio/difyai
```

### 3. Securing Private Buckets and Configuring FILES_URL

A common security mistake is making the entire MinIO bucket public via `mc anonymous set download localminio/difyai`. While this resolves file viewing issues, it exposes proprietary knowledge base documents to unrestricted public browsing.

Keep the MinIO bucket private. Dify manages authenticated access by generating signed URLs for internal workers or proxying file requests through the Dify backend. Ensure `FILES_URL` in `docker/.env` points to your public Dify API endpoint rather than your private MinIO host:

```env
FILES_URL=https://api.dify.example.com/files
```

## Why Knowledge Base Workflows Face Storage Synchronization Bottlenecks

Configuring MinIO provides reliable raw object persistence for Dify, but it reveals an architectural limitation in how LLM agents consume enterprise knowledge. MinIO is a passive blob store: it stores bytes on disk, but it does not index document semantics, detect internal relationships, or provide structured search.

When a Dify agent workflow references a Knowledge Base, the system undergoes an intensive ingestion cycle:

1. Document Retrieval. Celery workers pull the entire raw document from MinIO into temporary container memory.
2. Text Extraction. Open-source parsers extract raw text from PDFs, Word documents, or HTML files.
3. Segmentation and Chunking. Text is partitioned into token windows governed by `INDEXING_MAX_SEGMENTATION_TOKENS_LENGTH` (defaulting to 4,000 tokens).
4. Vector Embedding. Every chunk is transmitted over the network to an embedding model API, converting text into high-dimensional vectors.
5. Indexing. Vectors are written into Weaviate, Milvus, or Qdrant for semantic similarity searches.

This process introduces severe retrieval bottlenecks in enterprise environments. When business documents change in primary corporate storage, Dify has no awareness of the update until an engineer or curator manually re-uploads the file and re-runs the entire embedding pipeline.

Furthermore, many organizations already maintain large document collections across enterprise repositories such as Dropbox, Box, Google Drive, OneDrive, and SharePoint. Relying on basic web scraping or rudimentary native connectors to bridge these systems creates sluggish workflows: native connectors often attempt to pull entire folder structures, exhaust API rate quotas, or fail on large discovery exports.

This is where Fast.io provides a modern architecture for agentic teams. Fast.io functions as an intelligent workspace platform that unifies persistent cloud storage with built-in indexing and remote Model Context Protocol (MCP) access.

Instead of manually exporting files from corporate silos into MinIO, teams keep their existing storage. Fast.io syncs folders directly from Dropbox, Box, or OneDrive (one-way or two-way, on a schedule or on demand), while Google Drive imports today with sync coming soon; synchronization is never real-time.

Once ingested into a Fast.io workspace, files are automatically indexed for hybrid search, combining exact full-text keyword matching with deep semantic vector search. Autonomous AI agents and Dify workflows connect to Fast.io through the remote MCP server at `https://mcp.fast.io/mcp` (or `https://mcp.fast.io/mcp/key` with Bearer authentication), and legacy Server-Sent Events at `https://mcp.fast.io/sse`. Instead of pulling massive raw files across the network and forcing Dify workers to re-chunk them, agents query the workspace index directly through MCP and receive targeted, citation-backed excerpts.

The performance advantages of pre-indexed workspaces over traditional storage connectors are substantial. In head-to-head testing published at [Fast.io Benchmarks](https://fast.io/benchmarks/), Fast.io was measured the fastest and lowest cost of the providers tested.

## How to Structure Document Intelligence with Metadata Views and Team Governance

Production AI workflows require more than unstructured text search. Enterprise teams operating agent fleets in Dify, Cursor, or Claude Cowork frequently deal with structured business documents: legal agreements with explicit execution dates, financial statements with line items, insurance claims with policy numbers, and engineering specifications with strict typing.

Passing raw text chunks from MinIO or a vector database forces the downstream LLM to parse messy strings, increasing hallucination rates and processing costs. To solve this challenge, Fast.io provides [Metadata Views](/product/document-data-extraction/).

Metadata Views convert unstructured workspace documents into a live, queryable database. Users describe the target fields in natural language, and Fast.io automatically generates a typed schema supporting Text, Integer, Decimal, Boolean, URL, JSON, and Date & Time. The platform inspects matching documents across the workspace (including PDFs, spreadsheets, Word documents, scanned forms, and presentations) and populates a filterable, sortable spreadsheet without requiring custom OCR pipelines or fragile regex parsers.

Dify custom tools and autonomous agents read those extracted fields through the Fast.io REST API. One call lists the files in a workspace with their typed fields, optionally scoped to a folder or filtered by category, MIME type or file extension:

```
GET https://api.fast.io/current/workspace/{workspace_id}/metadata/eligible/?category=Legal&page_size=100
```

The agent receives precise, typed JSON containing exactly the records required to execute its task. Teams can add new columns to an existing Metadata View at any time without reprocessing underlying files.

Beyond structured extraction, Fast.io delivers the operational governance necessary for human-agent collaboration:

* Per-File Version History. Every update written by an automated agent or human team member is preserved in full version history. If an agent script writes invalid metadata or corrupts a file, engineers can inspect diffs and restore previous states instantly.
* Append-Only Audit Log. All read, write, search, and extraction actions are permanently recorded with timestamps and identity attribution, providing complete visibility into agent activities across workspaces.
* Collaborative Notes. Built-in Notes provide a shared document coordinated through Agent Intents, where developers and AI agents collaborate on system prompts, workflow design documents, and operational runbooks.
* Scoped Ownership Transfer. Development teams and AI consultants can build an organization, configure workspaces, and integrate agent data pipelines for a client, then transfer organizational ownership to the client upon project completion while retaining administrative permissions.

Getting started with Fast.io is straightforward. Monthly plans start with a trial of up to 30 days (credit card required); annual plans have no trial.

| Plan Tier | Monthly Price | Included Storage | AI Tokens / Credits | Workspace Seats |
| :--- | :--- | :--- | :--- | :--- |
| Starter | $9.99/mo | 250 GB | 100,000 credits | 3 seats |
| Business | $49.99/mo | 5 TB | 600,000 credits | 10 seats |
| Enterprise | $199.99/mo | 25 TB | 3,000,000 credits | 30 seats |

Plans start with the entry Starter tier, scaling to Business and Enterprise for larger enterprise deployments. You can review plan details and feature comparisons on the [pricing page](/pricing/).

## Frequently asked questions

### How do I configure MinIO storage in self-hosted Dify?

To configure MinIO in self-hosted Dify, edit your docker/.env file to set STORAGE_TYPE=s3. Then configure S3_ENDPOINT pointing to your MinIO instance (for example, http://minio:9000), set S3_BUCKET_NAME to your pre-created bucket, provide S3_ACCESS_KEY and S3_SECRET_KEY, set S3_REGION=us-east-1, and specify S3_ADDRESS_STYLE=path. Restart Dify with docker compose down && docker compose up -d to apply the changes.

### Can Dify use private S3 buckets for knowledge bases?

Yes, Dify is designed to work with private S3 and MinIO buckets. Dify authenticates using access and secret keys to read and write document datasets. Internal workers access files directly over the S3 API, while file access for end users is securely proxied through the Dify API backend or governed via temporary pre-signed URLs, preventing public exposure of your raw documents.

### How does MinIO storage compare to cloud S3 in Dify performance?

MinIO delivers significantly lower latency and eliminates outbound cloud egress fees when deployed in the same local network or Kubernetes cluster as Dify. Cloud S3 introduces variable internet latency during document chunking and vector indexing. However, self-hosted MinIO requires managing local disk provisioning, backups, and high availability, whereas cloud S3 provides managed multi-region durability out of the box.

### Why does Dify return an HTTP 413 error during file uploads?

An HTTP 413 Request Entity Too Large error occurs when an uploaded file exceeds the Nginx reverse proxy limit. While Dify restricts document size through the `UPLOAD_FILE_SIZE_LIMIT` parameter (defaulting to `15 MB`), Nginx enforces its own ceiling via `NGINX_CLIENT_MAX_BODY_SIZE` (defaulting to `100M`). If you increase the application upload limit beyond `100 MB` without raising `NGINX_CLIENT_MAX_BODY_SIZE` to match, Nginx intercepts and rejects the payload before it reaches Dify or MinIO.

### How do I fix CORS errors between Dify and MinIO?

CORS errors occur when browser clients attempt to preview or download files directly from MinIO. To resolve this, configure CONSOLE_CORS_ALLOW_ORIGINS and WEB_API_CORS_ALLOW_ORIGINS in Dify's .env file, and apply a CORS policy to your MinIO bucket using the MinIO Client tool: mc cors set <alias>/<bucket> cors.json, permitting GET, PUT, POST, and HEAD methods from your Dify domain.

### How does connecting Fast.io via MCP improve agent dataset retrieval over raw object storage?

MinIO acts as raw object storage, requiring Dify Celery workers to repeatedly pull full documents across the network for text parsing and vector indexing. Fast.io provides intelligent workspaces that automatically index synced corporate documents for hybrid full-text and semantic search. AI agents connect via remote MCP to retrieve surgical 200-word passages on demand, dramatically cutting token usage and eliminating local indexing pipelines.

## Sources

- [Dify Documentation: Environment Variables](https://docs.dify.ai/getting-started/install-self-hosted/environments) — S3 endpoint address is required for non-AWS S3-compatible services like MinIO in self-hosted Dify environments.

## About Fast.io

Fast.io provides shared workspaces where people and AI agents work on the same files, with built-in semantic search and citation-backed chat over what they hold. Agents reach it through a remote MCP server at https://mcp.fast.io/mcp, a REST API at https://api.fast.io/current/, and a command line client published on npm as @vividengine/fastio-cli.
