AI & Agents

Amazon Bedrock SharePoint Connector: Setup, Limitations, and Alternatives

The Amazon Bedrock SharePoint connector allows Amazon Bedrock Knowledge Bases to ingest and index documents stored in Microsoft SharePoint Online for RAG workflows. While native crawling indexes document libraries directly, teams face scheduled sync delays, Graph API rate limits, and OpenSearch Serverless collection fees. Evaluating these constraints helps teams decide when native connectors work and when indexed cloud workspaces offer a simpler retrieval path.

Derek Labian 14 min read Updated
Amazon Bedrock SharePoint connector data ingestion architecture

How the Amazon Bedrock SharePoint Connector Works

Connecting generative AI models to enterprise SharePoint repositories usually fails at the ingestion boundary rather than the retrieval layer. When teams point an Amazon Bedrock Knowledge Base at a corporate SharePoint tenant, they expect conversational answers grounded in current company documents, but the native connector relies on periodic batch crawling that leaves hours of latency between file edits and model awareness. Combined with Microsoft Graph API rate limits on large document libraries and persistent vector database running costs, engineering teams frequently find that native repository ingestion creates more operational friction than the retrieval pipeline solves.

The Amazon Bedrock SharePoint connector allows Amazon Bedrock Knowledge Bases to ingest and index documents stored in Microsoft SharePoint Online for RAG workflows. Rather than requiring developers to manually write custom export scripts, stage files into Amazon S3 buckets, and trigger embedding jobs, Amazon Web Services provides a managed ingestion pipeline that connects directly to Microsoft 365. The service crawls document libraries, extracts text from standard file formats, partitions content into chunks, generates vector embeddings using models such as Amazon Titan Text Embeddings, and indexes those vectors into an Amazon OpenSearch Serverless vector collection.

In a typical enterprise deployment, documents do not originate inside AWS. Teams create, edit, and organize departmental files inside SharePoint Online, OneDrive for Business, Google Drive, Box, or Dropbox. When building retrieval augmented generation systems on AWS Bedrock, the core requirement is allowing foundation models like Anthropic Claude or Amazon Titan to retrieve grounded excerpts from those enterprise documents without requiring employees to change where they collaborate.

The native architecture splits into three main layers: identity federation through Microsoft Entra ID, ingestion orchestration handled by Amazon Bedrock Knowledge Bases, and vector search powered by OpenSearch Serverless. When a user or automated agent submits a query to the Bedrock runtime, the system converts the question into a vector embedding, queries OpenSearch Serverless for semantically relevant chunks, and injects the retrieved text into the model prompt as verified context.

While this managed flow removes the initial burden of writing glue code, running it in production surfaces distinct architectural constraints. Data synchronization does not occur in real time, Microsoft Graph API rate limits throttle high-volume crawlers, and the underlying vector database accumulates infrastructure costs regardless of whether users run queries. Understanding each phase of the pipeline is essential before committing production workloads to native ingestion.

How to Configure Entra ID Permissions for Bedrock Ingestion

Setting up native SharePoint indexing in Bedrock requires Microsoft Entra ID app permissions, tenant admin consent, and OpenSearch vector store provisioning. Because Bedrock runs outside your Microsoft 365 tenant, the crawler must authenticate as a registered enterprise application with appropriate API permissions to inspect sites, list libraries, and stream document contents.

Configuring the integration requires coordination between an AWS administrator and a Microsoft 365 tenant administrator across multiple portals:

  1. Register an application in Microsoft Entra ID to establish an identity for the AWS Bedrock crawler.
  2. Configure Microsoft Graph and SharePoint REST API application permissions.
  3. Obtain tenant admin consent for the configured permission scopes.
  4. Store the Application Client ID, Directory Tenant ID, and Client Secret in AWS Secrets Manager.
  5. Create an Amazon Bedrock Knowledge Base with an Amazon OpenSearch Serverless vector collection.
  6. Add a SharePoint data source pointing to your site collection URL and reference the secret.

Security scope represents the first major operational decision. Microsoft historically allowed broad app-only access through Azure Access Control Services (ACS), but Microsoft retired ACS authentication in April 2026. All new SharePoint data source connections must use Microsoft Entra ID OAuth 2.0 client credentials.

When assigning permissions in Entra ID, administrators must choose between tenant-wide read access (Sites.Read.All) and scoped site access (Sites.Selected). While Sites.Read.All simplifies configuration by granting the crawler access to every site collection across the entire organization, it violates the principle of least privilege. In enterprise environments, corporate security teams rarely permit third-party cloud services to hold tenant-wide document read rights.

Using Sites.Selected restricts the application so it can only access specific SharePoint sites that an administrator explicitly grants through PowerShell or the Microsoft Graph API. After registering the application, an administrator executes a grant command defining read access for the target site collection:

POST https://graph.microsoft.com/v1.0/sites/{site-id}/permissions
Content-Type: application/json

{
  "roles": ["read"],
  "grantedToIdentities": [{
    "application": {
      "id": "YOUR_ENTRA_APP_CLIENT_ID",
      "displayName": "Bedrock-SharePoint-Crawler"
    }
  }]
}

Once Entra ID credentials are generated, you store the client secret inside AWS Secrets Manager as plain key-value pairs (clientId, clientSecret, and tenantId). The IAM execution role assigned to the Amazon Bedrock Knowledge Base must include explicit permission to decrypt and read that secret, alongside read-write access to the OpenSearch Serverless vector index.

Ingestion Constraints and Microsoft Graph API Throttling

The native Amazon Bedrock connector relies on periodic web crawling rather than change data capture or event notifications. AWS Bedrock SharePoint data sources require scheduled ingestion synchronization rather than real-time file updates. When a team member edits a policy manual, updates a technical specification, or uploads a contract to SharePoint, those modifications remain invisible to the retrieval model until the next synchronization cycle runs and completes indexing.

During ingestion, Bedrock sends thousands of programmatic requests to Microsoft Graph and SharePoint REST endpoints to enumerate folder hierarchies, check metadata timestamps, and download raw file binaries. This process directly encounters Microsoft 365 protection mechanisms.

Microsoft protects SharePoint Online infrastructure through multi-tiered rate limiting across application, user, and tenant boundaries. When a background ingestion process generates high request volumes, Microsoft servers respond with HTTP 429 status codes to preserve interactive tenant performance. Queries using app-only permissions in SharePoint Online are throttled at 25 requests per second. When an application hits this threshold, Microsoft includes a Retry-After header indicating how many seconds the crawler must pause before issuing another call.

During initial full crawls across large document libraries, this rate limiting creates significant operational bottlenecks. If an organization maintains tens of thousands of PDF manuals, spreadsheets, and word processing documents, the initial synchronization can take multiple days to finish. If the crawler does not manage exponential backoff with randomized jitter, repeated requests during active throttle windows extend the penalty duration.

File format and payload constraints further restrict native crawling:

  • The simple upload API for Microsoft Graph only supports files up to 250 MB in size. While downloads handle larger binaries through chunked drive item requests, Bedrock Knowledge Bases impose practical document parsing thresholds that skip oversized files.
  • SharePoint Lists are completely excluded from native ingestion. The connector indexes document libraries and standard modern site pages, but structured data stored in custom SharePoint lists is ignored.
  • Embedded multimodal content within documents, such as complex architecture diagrams, nested spreadsheet charts, and raster graphics, is discarded during standard text chunking pipelines.
  • Password-protected documents, legacy binary formats, and OneNote notebooks fail extraction and produce synchronization error logs that require manual review in the AWS CloudWatch console.

OpenSearch Serverless Overhead and Infrastructure Costs

Competitor posts omit the operational overhead of managing OpenSearch Serverless collections and handling Graph API throttling during initial full crawls. When tutorials demonstrate connecting Bedrock to SharePoint in a few console clicks, they often gloss over the continuous infrastructure footprint required to support that connection.

For the native SharePoint connector, Amazon Bedrock Knowledge Bases require Amazon OpenSearch Serverless as the vector storage engine. While OpenSearch Serverless eliminates the manual maintenance of operating hardware clusters, it introduces specific financial and architectural constraints.

OpenSearch Serverless meters capacity using OpenSearch Compute Units (OCUs), split between indexing OCUs and search OCUs. AWS enforces baseline minimum OCU allocations to ensure high availability and index durability. Even when an organization runs zero user queries over a weekend or holiday period, the provisioned collection continues to consume compute capacity units around the clock. For enterprise teams running proof-of-concept projects or internal departmental tools with sporadic usage, these base compute unit charges often exceed the cost of the foundation model invocations themselves.

A frequent administrative pitfall is the lifecycle separation between Bedrock Knowledge Bases and OpenSearch collections. When an administrator deletes an Amazon Bedrock Knowledge Base from the AWS Management Console, the associated OpenSearch Serverless collection is not deleted. The underlying collection, its vector index, and its associated network and encryption policies remain active in the account. Unless an engineer explicitly navigates to the Amazon OpenSearch Service dashboard and removes the orphaned collection, the account continues to accrue compute capacity charges indefinitely.

Private enterprise networking adds another layer of complexity. Corporate security standards typically forbid vector databases from exposing public internet endpoints. Connecting a Bedrock Knowledge Base to an OpenSearch Serverless collection inside a private network requires provisioning AWS PrivateLink VPC endpoints, authoring OpenSearch data access policies, and configuring security group ingress rules that allow the Bedrock service principal to communicate with the vector index.

Fastio features

Connect AI Agents to Your Enterprise Files

Sync SharePoint document libraries into intelligent workspaces with built-in search and consolidated MCP tools. Every organization starts with a 30-day free trial, which requires a credit card. Plans are Starter at $29/mo, Business at $99/mo, and Enterprise at $299/mo.

Syncing Cloud Storage into Workspaces with Consolidated MCP Tools

Because native cloud connectors demand extensive infrastructure configuration and endure scheduled crawl delays, many development teams adopt an alternate pattern: keeping existing cloud drives as the authoring surface while using an intelligent workspace layer to supply real-time context to AI agents.

In this model, teams do not build a custom crawler or maintain a standalone vector database. The team continues storing and updating operational files inside SharePoint, OneDrive, Dropbox, Box, or Google Drive. Through Fastio Cloud Sync, target folders synchronize into a centralized, org-owned Fastio workspace. Cloud Sync operates for Dropbox, Box, and OneDrive, running one-way or two-way on a schedule or on demand. Because SharePoint document libraries reside on Microsoft 365 storage backends, teams connect SharePoint libraries directly through the OneDrive connector. Google Drive is import today with sync coming soon, and direct URL imports allow pulling files without intermediate local disk input or output.

Once documents enter the workspace, Fastio Intelligence Mode automatically indexes the content on arrival. There are no vector collections to size, no compute capacity floors to monitor, and no OpenSearch network policies to debug. Text extraction, semantic indexing, and full-text keyword indexing happen automatically as documents land.

Rather than tying retrieval to a single cloud provider's proprietary API, Fastio exposes a consolidated Model Context Protocol (MCP) toolset. AI agents connect to the Fastio MCP server over Streamable HTTP using the address for their environment: https://mcp.fast.io/mcp/code for coding agents such as Claude Code and Cursor, and https://mcp.fast.io/mcp/tools for Claude apps and general runtimes like OpenClaw, while ChatGPT and Codex connect through the Fastio plugin, with https://mcp.fast.io/mcp/operations as an alternative custom MCP server. Interactive clients sign in with OAuth in the browser and carry no API key, while headless code and custom enterprise runtimes send an Authorization: Bearer <api key> header on the connection, as detailed in the Fastio setup documentation.

Rather than forcing an agent to download entire multi-gigabyte document libraries to answer a specific question, the agent invokes semantic search tools directly through MCP. The agent submits a natural language query, receives concise text excerpts with exact document citations, and constructs verified answers while consuming minimal context tokens.

When evaluating connector efficiency across platforms, independent measurements matter. The connector comparison is published at Fast.io Benchmarks and that page is the only place its numbers live. Fastio was measured the fastest and the lowest cost of the providers tested, providing agentic teams with an optimized retrieval substrate that avoids the latency and expense of manual vector pipeline maintenance.

When to Choose Native Bedrock Ingestion Versus Fastio Workspaces

Choosing between the native Amazon Bedrock SharePoint connector and an intelligent workspace architecture depends on where your models run, who manages the infrastructure, and how quickly your agents require access to updated documents.

The following comparison summarizes how each architecture addresses enterprise knowledge retrieval:

Evaluation Dimension AWS Bedrock SharePoint Connector Fastio Workspace MCP Architecture
Primary Data Source SharePoint Online document libraries SharePoint (via OneDrive connector), Box, Dropbox, Drive
Ingestion Mechanism Scheduled batch crawl via Microsoft Graph Cloud Sync (on a schedule or on demand) and URL import
Sync Frequency Periodic batch sync (manual or scheduled) Scheduled or on demand folder sync
Vector Storage Amazon OpenSearch Serverless collection Built-in Intelligence Mode indexing
Infrastructure Overhead High (Entra ID app, Secrets Manager, OpenSearch) Low (Zero vector DB configuration or maintenance)
Agent Interface AWS Bedrock Knowledge Base Retrieve API Consolidated remote MCP server (Streamable HTTP / SSE)
Model Portability Locked to Amazon Bedrock hosted models Works with Claude, GPT-4, Gemini, Bedrock, or local LLMs
Structured Data Layer Text chunk extraction only Metadata Views for schema-typed data extraction

Beyond basic document search, real enterprise workflows require structured data extraction and human collaboration. Fastio provides Metadata Views, which turn unstructured documents like vendor agreements, financial filings, and technical specifications into typed, queryable tables. AI automatically identifies fields such as contract parties, renewal dates, and line item totals without requiring brittle regular expressions or template models. Agents can inspect and query these structured views through MCP tools to perform analytical reasoning across entire document sets.

Team coordination also requires clear operational guardrails. When multiple agents and human colleagues collaborate within shared workspaces, Fastio maintains per-file version history and a detailed activity log tracking every view, edit, and download. Advisory file locking allows agents to signal active edits through MCP storage tools without blocking other team members, while version history preserves every revision. When an external contractor or autonomous agent builds out a workspace, Fastio ownership transfer allows handing administrative control over to an enterprise stakeholder while preserving workspace continuity.

Sources

References used to verify factual claims in this guide.

  1. 1 Microsoft Learn Accessed

    Queries using app-only permissions in SharePoint Online are throttled at 25 requests per second.

  2. 2 Microsoft Learn Accessed

    The simple upload API for Microsoft Graph only supports files up to 250 MB in size.

Frequently Asked Questions

How do I connect Amazon Bedrock to SharePoint?

To connect Amazon Bedrock to SharePoint, register an enterprise application in Microsoft Entra ID with appropriate Microsoft Graph and SharePoint REST API permissions, grant tenant admin consent, and store the credentials in AWS Secrets Manager. Then, create an Amazon Bedrock Knowledge Base backed by an Amazon OpenSearch Serverless collection, add Microsoft SharePoint as a data source using your site URL, and execute an initial synchronization crawl.

What are the limitations of the Bedrock SharePoint connector?

The Amazon Bedrock SharePoint connector is limited to document libraries and modern site pages, leaving SharePoint Lists unsupported. It does not extract embedded charts or diagrams from documents, cannot parse OneNote notebooks, and enforces file size limits. Ingestion relies on scheduled batch crawling rather than real-time updates, and large repositories frequently encounter Microsoft Graph API rate limits.

How often does Amazon Bedrock sync with SharePoint?

Amazon Bedrock does not sync with SharePoint in real time. Synchronization runs on demand through the AWS Management Console or on a recurring schedule configured with Amazon EventBridge Scheduler or AWS Step Functions. Changes made in SharePoint Online only appear in model retrieval results after a scheduled sync completes.

Why does Amazon Bedrock require OpenSearch Serverless for SharePoint data sources?

Amazon Bedrock requires a vector database to index document text embeddings and perform semantic similarity searches. For the native SharePoint connector, AWS designates Amazon OpenSearch Serverless as the supported vector store service to store chunk embeddings, document metadata, and access control mappings.

How does Microsoft Graph API throttling affect Bedrock Knowledge Base crawls?

During full crawls across large document collections, the Bedrock connector makes rapid, repeated API calls to enumerate directories and download files. Microsoft Graph API enforces rate limits to protect tenant performance, returning HTTP 429 Too Many Requests responses that force the crawler to pause and retry, extending ingestion times.

Can AI agents query SharePoint documents without managing an OpenSearch collection?

Yes. Teams can sync SharePoint libraries into Fastio workspaces using Cloud Sync through the OneDrive connector. Fastio Intelligence Mode indexes documents automatically on arrival, allowing agents to query context and retrieve citations over a remote Model Context Protocol (MCP) server without provisioning or managing a vector database.

Related Resources

Fastio features

Connect AI Agents to Your Enterprise Files

Sync SharePoint document libraries into intelligent workspaces with built-in search and consolidated MCP tools. Every organization starts with a 30-day free trial, which requires a credit card. Plans are Starter at $29/mo, Business at $99/mo, and Enterprise at $299/mo.