AI & Agents

How to List Files with the Dropbox API: Pagination, Cursors, and Agent Workspaces

Listing files through the Dropbox API requires managing cursor-based pagination across the /files/list_folder and /files/list_folder/continue endpoints. Without proper cursor tracking and rate-limit backoff, deep directory traversals quickly exhaust API quotas. Synchronizing storage folders into an intelligent workspace provides an alternative path, allowing AI agents to query indexed files directly without recursive polling.

Derek Labian 12 min read Updated
Navigating directory listings, cursor pagination, and agent workspace storage

How to Initialize Folder Listings with the Dropbox Endpoint

When an autonomous agent needs to locate a document in a deep Dropbox directory, issuing recursive listing calls forces the application to traverse every subdirectory sequentially, consuming API quotas and hitting rate limits long before reaching the relevant file. The alternative is separating raw cloud storage from the agent's query layer, using indexed workspaces that resolve file queries through direct semantic search.

Listing files with the Dropbox API requires calling the /files/list_folder endpoint to obtain an initial batch and pagination cursor, followed by repeated /files/list_folder/continue requests until has_more is false. Understanding this two-stage interaction is necessary for any developer building integration scripts, backup workers, or agentic storage tools.

The Dropbox API v2 operates on an RPC architecture over HTTPS. Rather than issuing GET requests with URL query parameters, client applications send POST requests containing a JSON payload to https://api.dropboxapi.com/2/files/list_folder. Requests require an OAuth 2.0 bearer token in the Authorization header and the files.metadata.read permission scope, as documented in the official Dropbox API documentation.

import requests

DROPBOX_LIST_URL = "https://api.dropboxapi.com/2/files/list_folder"
ACCESS_TOKEN = "YOUR_DROPBOX_ACCESS_TOKEN"

headers = {
    "Authorization": f"Bearer {ACCESS_TOKEN}",
    "Content-Type": "application/json",
}

payload = {
    "path": "/Engineering/Specs",
    "recursive": False,
    "include_media_info": False,
    "include_deleted": False,
    "include_has_explicit_shared_members": False,
    "include_mounted_folders": True,
    "limit": 500,
}

response = requests.post(DROPBOX_LIST_URL, headers=headers, json=payload)
response.raise_for_status()
initial_data = response.json()

The path parameter specifies the folder path to evaluate. An empty string "" targets the user account's root namespace. Specifying a named path such as "/Engineering/Specs" bounds the listing to that specific subfolder. Dropbox paths are case-insensitive for routing, but returned entries retain their original casing in path_display.

Request Payload Parameters and Path Conventions

Configuring the initial request payload requires tuning parameters to match your application's data requirements:

  • path: The directory location to enumerate. Paths must start with a leading forward slash, such as "/Finance". Avoid trailing slashes; passing "/Finance/" results in a malformed path error (400 Bad Request). For root listings, supply an empty string "".
  • recursive: A boolean flag. When set to false, the API returns only immediate child files and folders within the target path. When set to true, the API traverses all nested folders recursively.
  • limit: An integer specifying the maximum entries returned per page. Dropbox API limits list_folder responses to 2,000 entries per page, which represents the approximate upper limit. Setting a lower threshold, such as 500 or 1,000, helps prevent server-side request timeouts on dense directories.
  • include_deleted: A boolean flag. Set this to true only if you maintain a local database mirror that needs to track deletions. For standard file discovery, leave this set to false.
  • include_mounted_folders: When true, includes shared team folders and external shares mounted within the user account.

If the initial request succeeds, Dropbox returns an HTTP 200 payload containing three core attributes: entries (an array of file and folder metadata objects), cursor (an opaque string capturing pagination state), and has_more (a boolean indicating whether more items remain).

How Cursor Pagination Maintains Traversal State Across Pages

The Dropbox API does not support offset-based pagination. You cannot supply query parameters like offset=100 or page=2. Instead, the API uses a cursor-based pagination model. The cursor string returned by the initial /files/list_folder call acts as a stateful snapshot token, identifying the exact point in the directory stream where retrieval paused.

When a folder contains more entries than fit in the initial response, has_more returns true. The client must then pass the returned cursor to the continuation endpoint: https://api.dropboxapi.com/2/files/list_folder/continue.

import requests

CONTINUE_URL = "https://api.dropboxapi.com/2/files/list_folder/continue"
ACCESS_TOKEN = "YOUR_DROPBOX_ACCESS_TOKEN"

def list_complete_directory(initial_cursor):
    headers = {
        "Authorization": f"Bearer {ACCESS_TOKEN}",
        "Content-Type": "application/json",
    }
    cursor = initial_cursor
    has_more = True
    collected_files = []
    while has_more:
        response = requests.post(
            CONTINUE_URL,
            headers=headers,
            json={"cursor": cursor},
        )
        response.raise_for_status()
        page_data = response.json()
        for entry in page_data.get("entries", []):
            if entry.get(".tag") == "file":
                collected_files.append({
                    "name": entry.get("name"),
                    "path": entry.get("path_display"),
                    "size": entry.get("size"),
                    "modified": entry.get("server_modified"),
                    "hash": entry.get("content_hash"),
                })
        cursor = page_data.get("cursor")
        has_more = page_data.get("has_more", False)
    return collected_files

The continuation endpoint accepts only the cursor in its JSON payload. You cannot modify the folder path, recursive setting, or include flags during continuation calls. The original query parameters remain locked inside the cursor state.

Handling Cursor Expiration and State Resets

Cursors remain valid for active pagination loops, but they are not permanent database keys. A cursor can expire if an application delays between requests or if heavy background writes reorganize the folder tree.

When a cursor expires or becomes invalid, /files/list_folder/continue returns an HTTP 409 Conflict status with a specific error tag: path/not_found/ or reset. An expired cursor cannot be refreshed or recovered. The application must handle this exception, discard the invalidated cursor, and restart directory enumeration from /files/list_folder.

If your application stores cursors to monitor long-term changes, persist the creation timestamp alongside the cursor. When an update check fails with a reset error, schedule a complete directory sweep to re-establish a fresh baseline.

Differentiating File, Folder, and Deleted Metadata Entries

Every object in the entries array contains a ".tag" discriminator string that establishes the entry type:

  • file: Designates an active document or asset. Includes name, id (a persistent identifier formatted as id:a4nz...), size in bytes, server_modified timestamp, path_display, and content_hash.
  • folder: Designates a subdirectory. Includes name, id, path_display, and optional sharing metadata like shared_folder_id. Folders do not carry size or content hash attributes.
  • deleted: Indicates that an item was deleted. Appears only when include_deleted was set to true. Contains the file name and path markers.

Dropbox calculates content_hash by taking SHA-256 hashes of individual binary block chunks, concatenating the resulting digests, and running a final SHA-256 hash over the combined bytes. This value differs from standard whole-file SHA-256 or MD5 hashes.

Why Directory Traversals Trigger Rate Limits and Timeouts

Running large-scale file listings across production accounts frequently encounters infrastructural barriers. The Dropbox API applies rate limits to prevent resource exhaustion, returning an HTTP 429 Too Many Requests response when an application exceeds acceptable call frequencies.

Repeated directory traversals across large agent corpora frequently trigger HTTP 429 rate limits. When autonomous agents poll folders repeatedly to check for updated dependencies or completed research outputs, they quickly consume available call budgets.

import time
import random
import requests

def execute_with_rate_limit_backoff(url, headers, payload, max_retries=6):
    for attempt in range(max_retries):
        response = requests.post(url, headers=headers, json=payload)
        if response.status_code == 200:
            return response.json()
        if response.status_code == 429:
            retry_header = response.headers.get("Retry-After")
            wait_time = int(retry_header) if retry_header else 2 ** attempt
            jitter = random.uniform(0.5, 1.5)
            time.sleep(wait_time + jitter)
            continue
        if response.status_code in (500, 502, 503, 504):
            backoff = (2 ** attempt) + random.uniform(0.1, 1.0)
            time.sleep(backoff)
            continue
        response.raise_for_status()
    raise RuntimeError("Max retries exceeded while calling Dropbox API.")

When an HTTP 429 arrives, client code should parse the Retry-After header. This header provides the integer number of seconds the client must pause before retrying. Combining this duration with randomized jitter prevents multiple workers from synchronizing their retry spikes.

Avoiding Lock Contention and Deep Traversal Timeouts

Rate limits in Dropbox are not purely volumetric. The platform also issues HTTP 429 and 503 responses due to lock contention. If an automated script writes files to a directory while another thread issues an extensive recursive listing on the same folder tree, the backend pauses read operations to preserve filesystem consistency.

Setting recursive: true on parent folders containing tens of thousands of items creates severe operational bottlenecks. The server must traverse deep directory hierarchies to compile the initial response page, often triggering HTTP 504 Gateway Timeout errors.

To avoid recursive timeouts on large repositories, adopt a queue-based breadth-first traversal. List only the top-level directory with recursive: false, gather subfolder paths, and query each subfolder independently. This approach isolates failures to individual folders without failing the entire tree.

Why Recursive Directory Traversal Breaks Agent Workflows

Directly connecting AI agents to raw cloud storage APIs creates a clear architectural mismatch. Autonomous agents such as Claude Code, Codex, and custom developer bots do not need raw directory listings. They need specific document contents, semantic context, and structured data.

When an agent relies on direct Dropbox API listings, it encounters two primary operational penalties:

  1. Context Window Saturation: A recursive directory listing of a project folder holding thousands of files generates hundreds of kilobytes of JSON text. Dumping raw file lists into an LLM prompt burns thousands of context tokens on directory paths, timestamps, and hashes before the model processes a single line of actual content.
  2. Polling Latency and Quota Depletion: Setting agents to poll Dropbox continuously for newly generated files consumes rate limit allowances and introduces multi-second network delays into agent response loops.

The alternative pattern separates storage management from the agent's real-time query loop. Instead of writing custom listing logic inside agent scripts, teams connect their existing storage to an intelligent workspace platform. Fast.io provides Cloud Sync for Dropbox, Box, and OneDrive (one-way or two-way, on a schedule or on demand, never continuous, live, or real-time; SharePoint libraries through the OneDrive connector; Google Drive is import today with sync coming soon). For organizations evaluating storage architectures, reviewing Dropbox alternatives reveals how dedicated workspace platforms handle high-concurrency agent interactions.

In benchmark evaluations of storage connectors for agentic workflows, the connector comparison is published at Fast.io benchmark report and that page is the only place its numbers live. Fastio was measured the fastest and the lowest cost of the providers tested.

By synchronizing Dropbox folders into an intelligent workspace on a schedule or on demand, files are indexed once upon arrival. The agent queries an indexed repository directly, eliminating recursive file crawls and preserving token capacity for reasoning.

Fastio features

Query Synced Cloud Storage Without API Traversal Bottlenecks

Connect your Dropbox storage to an intelligent workspace with scheduled sync and a consolidated MCP toolset. Starts with a 30-day free trial.

How to Connect AI Agents to Synced Storage via Model Context Protocol

When files synchronize from Dropbox into a Fast.io workspace, AI agents interact with them through the Model Context Protocol (MCP) rather than custom REST wrappers. ChatGPT and Codex connect through the Fastio plugin, while Fast.io hosts a remote MCP server over Streamable HTTP, providing dedicated endpoints configured for different client architectures:

  • https://mcp.fast.io/mcp/tools for Claude applications, Claude Cowork, and general conversational runtimes.
  • https://mcp.fast.io/mcp/code for coding agents including Claude Code, Cursor, and Gemini CLI.
  • https://mcp.fast.io/mcp/operations as an alternative for ChatGPT and Codex (such as for the Codex IDE extension or when plugins are blocked).

Interactive applications connect via OAuth with visual permission reviews. Automated server agents, background worker daemons, and framework runtimes authenticate by passing a scoped API key as an HTTP header: Authorization: Bearer YOUR_FASTIO_API_KEY. Complete setup steps are available at Fast.io MCP setup documentation, while tool schemas and interaction recipes are maintained in the Fast.io MCP skill guide.

{
  "mcpServers": {
    "fast-io": {
      "url": "https://mcp.fast.io/mcp/tools",
      "headers": {
        "Authorization": "Bearer YOUR_FASTIO_API_KEY"
      }
    }
  }
}

On the /mcp/tools endpoint, operations are split into read and management tools. File operations use storage for read-only actions and storage_manage for mutations. On /mcp/code, agents call search and execute for retrieval and execute_manage for write operations.

Replacing File Scans with Semantic Search and Metadata Views

Inside an intelligent workspace, agents no longer need to list files sequentially to locate relevant information.

Enabling Intelligence Mode on a workspace automatically indexes incoming documents for hybrid search, combining keyword matching with semantic vector retrieval. Agents query files by topic, concept, or document meaning rather than guessing directory paths, and search responses include source document citations.

For structured extraction across documents like legal agreements, financial receipts, or technical datasheets, Fast.io provides Metadata Views. Metadata Views uses natural language field descriptions to automatically construct a structured schema (supporting Text, Integer, Decimal, Boolean, URL, JSON, and Date & Time formats). Rather than requiring agents to download full documents and extract data with bespoke parsers, the system extracts properties into a sortable spreadsheet view that agents query directly via MCP.

When to Choose Direct API Polling Versus Synced Agent Workspaces

Selecting the right approach between raw Dropbox API integration and a synchronized workspace depends on your operational requirements and workflow complexity.

Direct Dropbox API calls work well for discrete, procedural scripts. If your workflow involves a single script that uploads an export archive once a day or downloads a specific file by known path, writing a direct Python client with standard pagination and backoff logic is practical and requires no additional services.

However, when multiple AI agents or blended human-agent teams collaborate on shared documents, direct API listings create maintenance friction. Agentic workflows require content indexing, version history, audit logging, and concurrency control.

Fast.io workspaces provide advisory file leases through the storage_manage tool using lock-acquire and lock-release, alongside lock-status on the storage tool. These leases signal active agent edits without preventing concurrent writes. Every file maintains full version history, allowing previous revisions to be reviewed or restored if an agent generates unexpected output.

Monthly plans start with a 30-day free trial, which requires a credit card. Plans are Starter at $9.99/mo, Business at $49.99/mo, and Enterprise at $199.99/mo on Fast.io subscription pricing. Credits meter AI operations like semantic search and document extraction, while storage and member seats are bundled with each subscription tier.

Sources

References used to verify factual claims in this guide.

  1. 1 Dropbox for Python Documentation Accessed

    The Dropbox API files_list_folder method begins returning folder contents and returns a cursor with has_more to retrieve subsequent entries.

Frequently Asked Questions

How do I list all files in a Dropbox folder using the API?

To list all files in a Dropbox folder, make an initial POST request to /files/list_folder with the target folder path, setting recursive to true if subfolders are required. The response returns an initial batch of items, a cursor string, and a has_more boolean. When has_more is true, pass the cursor to repeated /files/list_folder/continue requests in a loop until has_more returns false.

What is the difference between list_folder and list_folder/continue in Dropbox API?

The /files/list_folder endpoint initializes directory listing by taking configuration parameters like path, recursive, and limits to return the first page of files and a cursor. The /files/list_folder/continue endpoint accepts only that cursor payload to fetch subsequent pages. Traversal parameters cannot be passed to the continuation endpoint because they are encoded within the cursor.

How do I handle pagination when listing files in Dropbox?

Handle pagination by checking the has_more field returned in each Dropbox API response. Run a while loop that calls /files/list_folder/continue with the latest cursor as long as has_more is true. In each iteration, collect the returned entries, extract the new cursor, and update your termination condition when has_more evaluates to false.

What happens when a Dropbox list_folder cursor expires?

A Dropbox cursor can expire or reset if extensive folder reorganizations occur or if too much time passes between pagination requests. The continuation endpoint returns an error indicating the cursor is invalid. Applications cannot repair a broken cursor; they must discard it and restart the listing process from /files/list_folder.

How can AI agents search Dropbox files without triggering HTTP 429 rate limits?

Instead of having AI agents repeatedly poll Dropbox via recursive directory traversals, you can configure Cloud Sync to synchronize Dropbox folders into a Fast.io workspace on a schedule or on demand. AI agents then query the workspace through the remote MCP server using semantic search and Metadata Views, avoiding rate limits and saving context tokens.

Related Resources

Fastio features

Query Synced Cloud Storage Without API Traversal Bottlenecks

Connect your Dropbox storage to an intelligent workspace with scheduled sync and a consolidated MCP toolset. Starts with a 30-day free trial.