# How to Access SharePoint Files with Microsoft Graph API and AI Agents

The Microsoft Graph API for SharePoint provides programmatic endpoints to access sites, lists, and document libraries across Microsoft 365 tenants. Connecting AI agents directly to Microsoft Graph requires multi-step identifier resolutions, managing broad tenant permissions, and consuming context tokens on directory traversals. Synchronizing SharePoint libraries into indexed agent workspaces via remote Model Context Protocol servers offers a token-efficient retrieval pattern.

Source: https://fast.io/resources/microsoft-graph-api-sharepoint/
Author: [Derek Labian](https://fast.io/authors/derek-labian/)
Last reviewed: 2026-10-07

## Why Microsoft Graph Requires Multi-Step Resolution for SharePoint

Autonomous agents querying enterprise storage fail before retrieving their first document when pointed directly at Microsoft Graph endpoints. Unlike flat object stores where an object key resolves in a single request, SharePoint organizes enterprise files inside a rigid three-tier hierarchy: tenant, site collection, document library drive, and drive item. An AI agent inspecting a repository must resolve site identifiers, discover associated drives, locate folders, and paginate file collections before reading a single sentence of text. For production agents, this multi-step resolution burns API quotas, consumes context tokens on directory listings, and introduces latency into every retrieval step.

The Microsoft Graph API for SharePoint provides programmatic endpoints to access sites, lists, and document libraries across Microsoft 365 tenants.

Engineering teams building autonomous workflows often assume that connecting an LLM to SharePoint is as simple as supplying an API key and a folder path. In practice, Microsoft Graph treats SharePoint as an interconnected web of site collections, where document libraries are exposed as drive abstractions. When an agent attempts to locate a document by its human-readable path, it must translate that path through multiple layers of the Microsoft Graph resource model.

### The Microsoft Graph SharePoint Hierarchy

To access a file programmatically, an application traverses four distinct resource levels:

1. **Tenant**: The top-level Microsoft 365 organization boundary authenticated through Microsoft Entra ID.
2. **Site Collection (`site`)**: A SharePoint site or sub-site containing lists, document libraries, and permissions.
3. **Document Library (`drive`)**: A specialized document library exposed as a OneDrive-compatible drive container.
4. **Drive Item (`driveItem`)**: An individual folder or file containing metadata, permissions, and binary content streams.

Each level requires distinct endpoint paths and identifier resolutions:

```http
// Step 1: Resolve the site collection by hostname and relative path
GET https://graph.microsoft.com/v1.0/sites/{hostname}:/{server-relative-site-path}
Authorization: Bearer {token}

// Step 2: Retrieve the document libraries (drives) inside the site
GET https://graph.microsoft.com/v1.0/sites/{site-id}/drives
Authorization: Bearer {token}

// Step 3: List items inside the root of the document library
GET https://graph.microsoft.com/v1.0/drives/{drive-id}/root/children
Authorization: Bearer {token}

// Step 4: Download the binary file stream
GET https://graph.microsoft.com/v1.0/drives/{drive-id}/items/{item-id}/content
Authorization: Bearer {token}
```

### The Token Penalty of Directory Traversal

When an AI agent interacts with Microsoft Graph directly, every directory inspection returns verbose JSON payloads containing OData context metadata, entity tags (`eTag`, `cTag`), creation timestamps, user identity hashes, and URLs. A single folder containing forty documents can generate several kilobytes of structural JSON.

If an agent executes recursive directory scans to locate a specific agreement, policy brief, or technical drawing, it floods its context window with directory listings rather than substantive document content. This structural overhead consumes available reasoning context, slows down agent execution loops, and increases inference costs.

Instead of burning tokens on raw directory traversal, teams building production systems often separate file synchronization from agent interaction. By synchronizing SharePoint libraries into an intelligent workspace, such as [Fast.io storage for agents](/storage-for-agents/), files are automatically indexed for semantic retrieval, allowing agents to query content directly without parsing directory hierarchies.

## Authentication Scopes and Microsoft Entra ID Permissions

Deploying autonomous AI agents against Microsoft Graph requires a clear security architecture. Because background agents operate unattended without a human clicking through browser prompts, they cannot use standard delegated authentication flows. Instead, agents authenticate through Microsoft Entra ID using the OAuth 2.0 Client Credentials Grant.

### Configuring Application Authentication

To authenticate an agent daemon, administrators create an App Registration in the Microsoft Entra admin center. The registration provides three required credentials:

- **Application (Client) ID**: The unique identifier for the registered agent application.
- **Directory (Tenant) ID**: The identifier for the Microsoft 365 tenant hosting the SharePoint sites.
- **Client Secret**: A secure credential string generated for server-to-server token requests.

The agent requests an access token by sending a POST request to the Microsoft Entra token endpoint:

```http
POST https://login.microsoftonline.com/{your-tenant-id}/oauth2/v2.0/token
Content-Type: application/x-www-form-urlencoded

client_id=your-client-id
&client_secret=your-client-secret
&grant_type=client_credentials
&scope=https://graph.microsoft.com/.default
```

Upon successful validation, Microsoft Entra ID returns a JSON Web Token containing the authorized application roles. The agent includes this token as a Bearer credential in the `Authorization` header of every subsequent Graph request.

### The Permissions Tradeoff: Tenant-Wide vs. Granular Scopes

When configuring application permissions for Microsoft Graph, security teams face a difficult tradeoff between operational breadth and least-privilege security.

#### Tenant-Wide Scopes (`Sites.Read.All` and `Files.Read.All`)

The simplest implementation path grants the application either `Files.Read.All` or `Sites.Read.All` application permissions. These scopes allow the agent to read every file, document library, and list across all site collections in the entire tenant.

While this broad access eliminates permission errors when querying diverse libraries, it introduces severe security liabilities. If an agent with tenant-wide read access suffers from prompt injection or credential exposure, the blast radius encompasses all corporate documents, financial records, and employee data.

#### The Principle of Least Privilege: `Sites.Selected`

To protect corporate data, Microsoft Entra ID provides the `Sites.Selected` application permission. When an application is assigned `Sites.Selected`, it possesses zero access to any SharePoint site by default.

An administrator must explicitly grant permissions to designated site collections using Microsoft Graph API calls:

```http
POST https://graph.microsoft.com/v1.0/sites/{site-id}/permissions
Authorization: Bearer {admin-token}
Content-Type: application/json

{
  "roles": ["read"],
  "grantedToIdentitiesV2": [
    {
      "application": {
        "id": "your-client-id",
        "displayName": "AI Document Processing Daemon"
      }
    }
  ]
}
```

Using `Sites.Selected` restricts the agent's reach strictly to authorized project repositories. However, it introduces constraints: tenant-wide search endpoints (such as `/search/query`) often fail when an application only holds `Sites.Selected`, requiring agents to query each approved site collection individually.

## How to Query SharePoint Files and Document Libraries with Microsoft Graph

Reading documents from SharePoint via Microsoft Graph API involves resolving site identifiers, enumerating document library drives, and retrieving file streams. Following a structured procedure helps developers and AI engineers build reliable retrieval routines.

### Step-by-Step Procedure to Access SharePoint Files

1. **Obtain an Access Token**: Authenticate with Microsoft Entra ID using the OAuth 2.0 client credentials grant to receive a valid bearer token.
2. **Resolve the Target Site Collection ID**: Look up the site collection by hostname and server-relative path using `GET /sites/{hostname}:/{path}`.
3. **Identify the Document Library Drive ID**: Query `GET /sites/{site-id}/drives` to locate the target document library.
4. **Locate the Target Drive Item**: Query the folder contents or execute a search query using `GET /drives/{drive-id}/root/search(q='{query}')`.
5. **Download the File Content**: Retrieve the binary stream using `GET /drives/{drive-id}/items/{item-id}/content`.

### Implementation Example in Python

The following script demonstrates how to authenticate, resolve a SharePoint document library, search for a document, and retrieve its raw text content:

```python
import os
import requests

TENANT_ID = os.environ.get("AZURE_TENANT_ID", "your-tenant-id")
CLIENT_ID = os.environ.get("AZURE_CLIENT_ID", "your-client-id")
CLIENT_SECRET = os.environ.get("AZURE_CLIENT_SECRET", "your-client-secret")
HOSTNAME = "contoso.sharepoint.com"
SITE_PATH = "/sites/engineering"

// 1. Acquire OAuth 2.0 token
token_url = f"https://login.microsoftonline.com/{TENANT_ID}/oauth2/v2.0/token"
token_payload = {
    "client_id": CLIENT_ID,
    "client_secret": CLIENT_SECRET,
    "grant_type": "client_credentials",
    "scope": "https://graph.microsoft.com/.default"
}
token_response = requests.post(token_url, data=token_payload)
token_response.raise_for_status()
access_token = token_response.json()["access_token"]
headers = {"Authorization": f"Bearer {access_token}"}

// 2. Resolve site identifier
site_endpoint = f"https://graph.microsoft.com/v1.0/sites/{HOSTNAME}:{SITE_PATH}"
site_data = requests.get(site_endpoint, headers=headers).json()
site_id = site_data["id"]

// 3. List drives to locate the document library
drives_endpoint = f"https://graph.microsoft.com/v1.0/sites/{site_id}/drives"
drives_data = requests.get(drives_endpoint, headers=headers).json()
drive_id = drives_data["value"][0]["id"]

// 4. Search for a specific file within the drive
search_endpoint = f"https://graph.microsoft.com/v1.0/drives/{drive_id}/root/search(q='specification')"
search_results = requests.get(search_endpoint, headers=headers).json()
items = search_results.get("value", [])

if items:
    target_item_id = items[0]["id"]
    // 5. Fetch file content
    content_endpoint = f"https://graph.microsoft.com/v1.0/drives/{drive_id}/items/{target_item_id}/content"
    content_response = requests.get(content_endpoint, headers=headers)
    document_bytes = content_response.content
```

### Search Limitations in Native Microsoft Graph

While Microsoft Graph supports search queries within document libraries (`/root/search(q='...')`), native search presents severe limitations for autonomous agent architectures:

- **Lexical Keyword Matching**: Microsoft Graph search relies on lexical keyword matching. It does not perform semantic vector matching, causing queries with synonyms or conceptual questions to miss relevant documents.
- **Unchunked Binary Output**: When an agent searches for information, Graph returns entire file objects. If the matching document is an eighty-page PDF, the agent receives the entire binary file. The agent must parse the document, extract text, split it into chunks, and identify the relevant passage locally.
- **Rate Throttling Under Parallel Tool Calls**: When autonomous agents run multi-step reasoning plans, they frequently issue parallel search requests. Microsoft Graph aggressively throttles bursts of requests with HTTP 429 status codes, forcing agents into backoff loops.

## Architectural Alternative: Connecting AI Agents to SharePoint via Remote MCP

Most enterprises already store their critical documents in SharePoint, OneDrive, Dropbox, Google Drive, or Box. Rearchitecting storage systems or migrating corporate repositories simply to accommodate AI agents is neither practical nor necessary.

The enterprise architectural pattern keeps SharePoint as the authoritative system of record while introducing an intelligent workspace as the agent retrieval layer.

### The Fast.io Remote MCP Architecture

Instead of pointing an LLM directly at raw Microsoft Graph endpoints, organizations synchronize target SharePoint document libraries into a Fast.io workspace.

Cloud Sync connects directly to OneDrive and SharePoint document libraries, synchronizing folders one-way or two-way on a schedule or on demand. Synchronizations run reliably in the background without real-time overhead. For teams working across cloud environments, Dropbox and Box sync on the same scheduled or on-demand model; Google Drive supports direct cloud import today, with sync coming soon.

Once documents land in a Fast.io workspace, the platform's Intelligence Mode automatically indexes their contents for hybrid search. This index combines full-text lexical search, semantic embeddings, and metadata value filtering without requiring a separate vector database or external chunking pipeline.

Autonomous agents connect to the workspace through a remote Model Context Protocol (MCP) server over Streamable HTTP:

- **Coding Agents**: Claude Code, Cursor, Devin, and VS Code connect to `https://mcp.fast.io/mcp/code`.
- **General Personal Agents**: Claude apps, OpenClaw, and Hermes Agent connect to `https://mcp.fast.io/mcp/tools`.
- **ChatGPT and Codex**: Connect via the Fastio plugin in the plugin directory, or through `https://mcp.fast.io/mcp/operations` as a custom MCP server.

Detailed setup guides for human administrators are available at `https://mcp.fast.io/docs`, while autonomous agents can inspect the live tool definitions at `https://mcp.fast.io/skill.md`.

```json
{
  "mcpServers": {
    "fastio": {
      "url": "https://mcp.fast.io/mcp/tools",
      "headers": {
        "Authorization": "Bearer your-fastio-api-key"
      }
    }
  }
}
```

### Eliminating the Token Tax

When an agent needs information from an indexed SharePoint library, it no longer traverses directories or downloads entire binary streams. The agent issues a structured MCP tool call:

```json
{
  "tool": "storage",
  "action": "search",
  "parameters": {
    "query": "What are the termination notice requirements in the master services agreement?",
    "files_scope": ["legal-docs"]
  }
}
```

The MCP server returns concise, semantic passages with direct document citations. Instead of consuming thousands of tokens on directory trees and PDF binaries, the agent receives only the exact paragraphs required to answer the query.

In connector evaluations across cloud platforms, SharePoint was not measured in Claude Cowork connector benchmarks; Fastio was measured fastest and lowest cost among tested providers (https://fast.io/benchmarks/).

### Structured Document Extraction with Metadata Views

In addition to semantic text retrieval, enterprise workflows frequently require extracting structured fields from unstructured SharePoint documents. [Fast.io Metadata Views](/product/document-data-extraction/) convert collections of contracts, invoices, and technical specs into queryable tabular databases.

Users define required extraction columns in natural language, such as contract values, effective dates, counterparty names, or liability caps. Intelligence Mode populates a typed schema across PDFs, Word documents, and spreadsheets without manual OCR template rules. Autonomous agents can query these structured records directly through MCP without parsing raw file streams.

Organizations can deploy intelligent workspaces starting with a 30-day free trial on monthly plans, which requires a credit card.

| Plan Tier | Monthly Price | Storage Allowance | Included Seats | Monthly Credits |
| :--- | :--- | :--- | :--- | :--- |
| **Starter** | $9.99/mo | 250 GB | 3 seats | 100,000 credits |
| **Business** | $49.99/mo | 5 TB | 10 seats | 600,000 credits |
| **Enterprise** | $199.99/mo | 25 TB | 30 seats | 3,000,000 credits |
| **Credit Overage** | $10 per 100,000 credits | Available on all plans | Billed as consumed | Meters AI intelligence |

## Throttling, Governance, and Multi-Agent Production Best Practices

Building production-grade AI agent systems requires accounting for rate limits, concurrency conflicts, and organizational governance. Operating an autonomous agent fleet against Microsoft Graph introduces operational challenges that teams must actively mitigate.

### Handling Microsoft Graph Rate Limits and Throttling

Microsoft Graph enforces tenant-level and application-level request thresholds. When an agent fires parallel requests to inspect multiple folders or download concurrent files, SharePoint returns HTTP 429 Too Many Requests responses.

Production integration code must inspect the `Retry-After` HTTP header returned in 429 responses. This header specifies the exact number of seconds the agent must wait before retrying the operation. Implementing exponential backoff with randomized jitter prevents retry storms from extending throttling windows.

SharePoint also enforces a list view threshold of 5,000 items on document libraries. Queries that filter or sort on unindexed columns across large libraries fail with threshold errors. When applications execute batch operations, Microsoft Graph batch requests can combine up to 20 individual requests into one JSON object. Each sub-request inside the batch is evaluated against service-level throttling limits.

### Multi-Agent Write Safety and Advisory Locks

When multiple agents or human teammates collaborate inside shared document environments, uncoordinated writes can cause silent data loss. Fast.io provides multi-layered conflict avoidance for collaborative agent environments:

- **Per-File Version History**: Every file written to a workspace preserves complete version history. If an agent updates a document or generates a revised draft, previous versions remain intact and restorable through the UI or API.
- **Append-Only Audit Logs**: All workspace interactions, file reads, exports, and permission modifications are recorded in an append-only audit log, providing security teams with full visibility into agent actions.
- **Advisory File Leases**: To coordinate concurrent edits without hard lockouts, agents can acquire advisory leases via the `lock-acquire` and `lock-release` actions on the `storage_manage` MCP tool (and check status with `lock-status` on the `storage` tool). Other agents inspecting the file can observe active lease holders, preventing conflicting edits while allowing concurrent versioning.
- **Workspace Activity Feed**: Rather than repeatedly polling SharePoint or storage APIs to detect new files, agents can subscribe to workspace updates using WebSocket connections or the HTTP long-poll endpoint (`GET /current/activity/poll/{entity_id}?wait=95&lastactivity={timestamp}`). The feed notifies agents immediately when files are added or modified, eliminating polling overhead.

## Frequently asked questions

### How do I get files from SharePoint using Microsoft Graph API?

To get files from SharePoint using Microsoft Graph API, authenticate against Microsoft Entra ID using OAuth 2.0 to receive a bearer token. Next, resolve the target site collection ID using `GET /sites/{hostname}:/{path}` and find the document library drive ID using `GET /sites/{site-id}/drives`. Finally, retrieve the file metadata or stream its raw content using `GET /drives/{drive-id}/items/{item-id}/content`.

### What permissions are required to access SharePoint via Microsoft Graph?

Unattended AI agents and background daemons use application permissions. While `Files.Read.All` and `Sites.Read.All` provide broad tenant-wide read access across all site collections, enterprise security best practices recommend using `Sites.Selected`. With `Sites.Selected`, an administrator explicitly grants read or write access to specific site collections via the `/sites/{site-id}/permissions` endpoint.

### How do AI agents search SharePoint files efficiently?

Rather than forcing agents to crawl directory structures over REST APIs and download raw binary files, teams synchronize SharePoint libraries into an intelligent workspace on a schedule or on demand. The workspace automatically indexes documents for semantic and full-text search, allowing agents to query relevant passages with citations through a remote Model Context Protocol (MCP) server.

### What is the difference between Sites.Read.All and Sites.Selected in Microsoft Graph?

`Sites.Read.All` grants an application read access to all document libraries, lists, and site collections across an entire Microsoft 365 tenant. `Sites.Selected` grants no default access to any site collection; an administrator must explicitly authorize the application on designated site collections using Microsoft Graph permissions calls.

### Can AI agents interact with SharePoint using the Model Context Protocol (MCP)?

Yes. Agents can connect to synchronized SharePoint documents through a remote MCP server using Streamable HTTP. Coding agents connect to `https://mcp.fast.io/mcp/code`, while general personal agents connect to `https://mcp.fast.io/mcp/tools`. The agent queries indexed documents using structured semantic search tools rather than parsing raw SharePoint REST payloads.

### How does cloud sync handle SharePoint document libraries?

SharePoint document libraries are synchronized through the OneDrive cloud connector. Cloud Sync operates one-way or two-way on a schedule or on demand, without continuous or real-time background overhead. Fast.io supports Cloud Sync for OneDrive, Dropbox, and Box; Google Drive supports direct cloud import today, with sync coming soon.

## Sources

- [Microsoft Learn: Combine multiple HTTP requests using JSON batching](https://learn.microsoft.com/en-us/graph/json-batching): Microsoft Graph batch requests can combine up to 20 individual requests into one JSON object.

## About Fast.io

Fast.io provides shared workspaces where people and AI agents work on the same files, with built-in semantic search and citation-backed chat over what they hold. Agents reach it through a remote MCP server, a REST API at https://api.fast.io/current/, and a command line client published on npm as @vividengine/fastio-cli. MCP setup is at https://mcp.fast.io/docs: Claude and most MCP clients connect to https://mcp.fast.io/mcp/tools, ChatGPT to https://mcp.fast.io/mcp/operations, and coding agents to https://mcp.fast.io/mcp/code.
