# NotebookLM PDF Limit: Word Caps, Page Limits, and Gemini Notebook Rules

Google's Gemini Notebook, formerly NotebookLM, enforces a strict limit of 500,000 words and 200MB per uploaded PDF, with no fixed ceiling on page count. Free standard accounts can import up to 50 sources per notebook, while premium tiers support up to 600 sources. This guide details verified document limits, why PDF imports fail, and practical architectures for querying large document archives.

Source: https://fast.io/resources/notebooklm-pdf-limit/
Author: [Tom Langridge](https://fast.io/authors/tom-langridge/)
Last reviewed: 2026-09-19

## What is the NotebookLM PDF limit?

The NotebookLM PDF limit is governed by a 500,000-word cap and a 200MB file size ceiling per source, requiring unencrypted, OCR-processed text for successful document grounding. As documented in Google's official Gemini Notebook help center (Google renamed NotebookLM to Gemini Notebook on July 16, 2026), every PDF imported into a notebook must strictly adhere to these twin thresholds. Every uploaded PDF source file that exceeds either the 500,000 words limit or 200MB ceiling is rejected during ingestion.

NotebookLM calculates document boundaries strictly by lexical volume and binary payload weight, meaning there is no fixed page count limit. A short PDF containing high-resolution raster images can easily breach the 200MB ceiling, while a massive plain-text regulatory filing uploads smoothly provided the total words stay under the 500,000 limit.

### The twin boundaries: word count and raw file size

Google enforces two distinct physical boundaries on every uploaded file:

1. **The 500,000-word ceiling:** When an uploaded PDF source arrives, Google's document service extracts the text layer and counts the total words. Content extending beyond the 500,000 words limit is either truncated or causes the entire upload to fail.
2. **The 200MB binary size cap:** Regardless of how few words a document contains, uploaded source files cannot exceed the 200MB size ceiling. This limit guards against transmission timeouts, memory exhaustion, and server-side parsing bottlenecks.

Understanding the interaction between these two constraints determines how you prepare research corpora. Text-dense documents reach the word count ceiling first, whereas presentations, architectural blueprints, and scanned reports encounter the file size ceiling long before approaching half a million words.

### Supported file formats alongside PDF

While PDF remains the most common document format for research notebooks, Gemini Notebook accepts a broad spectrum of source formats. All uploaded source documents share the identical 500,000 words and 200MB limits:

- **Portable Document Format (PDF):** Standard documents, reports, whitepapers, and books. Must contain an extractable text layer.
- **Google Workspace files:** Google Docs, Google Slides (limited to 100 slides), and Google Sheets (currently limited to 100,000 tokens).
- **Microsoft Office files:** Word documents (.docx) and PowerPoint decks (.pptx).
- **Text and code files:** Plain text (.txt), Markdown (.md), and Comma-Separated Values (.csv).
- **Digital publications:** Open EPUB files without digital rights management (DRM), and eligible purchased Google Play Books.
- **Audio recordings:** MP3, WAV, and related voice formats, automatically transcribed into uploaded text sources under the 500,000 words limit.
- **Web and media links:** Clean HTML web URLs and public YouTube URLs with user-submitted or auto-generated captions.

## Does NotebookLM have a PDF page limit?

NotebookLM has no page limit for PDFs. Google enforces restrictions exclusively through a 500,000-word limit and a 200MB file size ceiling per source. A PDF document can contain 50 pages, 500 pages, or 1,200 pages, as long as the total word count stays under 500,000 words and the uncompressed file size remains below 200MB.

Competitors falsely cite arbitrary page counts, often claiming that NotebookLM caps documents at 500 pages. That assertion confuses typical document formatting assumptions with the actual ingestion pipeline. Because page count is purely a visual rendering artifact that changes with margins, font size, line spacing, and graphics, language models do not evaluate documents by page boundaries.

### Word density across common document types

The number of pages you can upload depends entirely on text density. To evaluate whether a specific PDF will fit beneath the 500,000-word limit, review typical page densities across publishing categories:

- **Standard paperback or textbook:** A typical trade book contains roughly 250 to 300 words per page. At 300 words per page, a single uploaded PDF source can hold 1,600 printed pages before reaching 500,000 words.
- **Double-spaced academic manuscripts:** Academic papers formatted with standard double-spacing contain approximately 250 words per page. A long dissertation or journal archive totals roughly 250,000 words, easily fitting within a single NotebookLM source.
- **Dense single-spaced legal filings:** Court transcripts, statutory codes, and corporate contracts formatted in tight single-spaced typography pack dense text into every page. In this layout, a substantial regulatory volume can span hundreds of pages before approaching the half-million-word boundary.
- **Technical specifications and manual drafts:** Technical guides with schematics, tables, and code snippets often average only 150 to 200 words per page. You can import technical documentation spanning over a thousand pages, provided the embedded diagrams do not inflate the total file size beyond the 200MB ceiling.

### Words versus tokens in the Gemini model architecture

Software engineers and data scientists often think in terms of model tokens rather than natural language words. The underlying Gemini models powering Gemini Notebook process tokenized vectors rather than raw strings.

In standard English prose, one word corresponds to roughly 1.3 to 1.4 tokens. A 500,000 words PDF represents approximately 650,000 to 700,000 tokens of context. Because Google built the Gemini 1.5 model family with native million-token context windows, a single 500,000 words source fits comfortably within the model's active attention span during query grounding.

## How NotebookLM source limits differ across Google AI tiers

Each individual PDF imported into NotebookLM cannot exceed 500,000 words or 200MB across all account tiers, but the total number of sources allowed in a single notebook expands with higher subscriptions. Free NotebookLM accounts cap total notebook capacity at 50 documents, while premium tiers expand up to 600 sources.

Google updated Gemini Notebook usage limits on September 2, 2026, transitioning from simple request counters to compute-based allocations that balance prompt complexity, chat depth, and media synthesis. While the per-source word and file size caps remain constant across all tiers, upgrading your plan multiplies total notebook capacity and daily generation limits.

### Source capacity and generation quotas by subscription plan

The following table details verified account limits across standard free tiers and premium Google AI plans as of September 2026:

| Subscription Plan | Per-Source Word Cap | Per-Source File Size | Max Sources per Notebook | Daily Chat Quota | Checked Date |
| :--- | :--- | :--- | :--- | :--- | :--- |
| **Standard (Free Account)** | 500,000 words | 200 MB | 50 sources | 50 chats/day | September 2026 |
| **Google AI Plus** | 500,000 words | 200 MB | 100 sources | 200 chats/day | September 2026 |
| **Google AI Pro** | 500,000 words | 200 MB | 300 sources | 500 chats/day | September 2026 |
| **Google AI Ultra (20 TB)** | 500,000 words | 200 MB | 500 sources | 2,500 chats/day | September 2026 |
| **Google AI Ultra (30 TB)** | 500,000 words | 200 MB | 600 sources | 5,000 chats/day | September 2026 |

### Calculating total corpus capacity

Multiplying the per-source cap by the maximum source allowance reveals the total theoretical knowledge base a single notebook can reference:

- **Free accounts:** Free standard accounts with 50 sources multiplied by 500,000 words provide up to 25 million words of grounded reference material.
- **Plus accounts:** 100 sources yields an aggregate corpus of 50 million words.
- **Pro accounts:** 300 sources expands total capacity to 150 million words.
- **Ultra accounts:** 600 sources reaches an extraordinary 300 million words in a single research workspace.

While a theoretical 25-million-word or 300-million-word repository offers massive retrieval potential, managing hundreds of distinct sources introduces cognitive overhead. When querying large notebooks, select specific sources manually in the source drawer to ensure the model focuses its attention on the exact documents relevant to your query.

## Why PDF uploads fail in NotebookLM and how to fix them

When a PDF fails to import into NotebookLM, the application typically displays a generic error message stating that the source could not be added. Because the user interface does not always pinpoint the exact failure mechanism, understanding the five most common causes helps you resolve the issue quickly.

### 1. Document exceeds the 500,000-word ceiling

If your uploaded PDF exceeds the allowed 500,000 words limit per source, the ingestion parser will reject the file or stop indexing after hitting the ceiling. This frequently occurs with omnibus legislation, multi-volume case law compilations, and monolithic technical manuals.

To check the exact word count of a PDF before uploading on macOS or Linux, extract the raw text and run a word count in your terminal:

```bash
pdftotext sample_document.pdf - | wc -w
```

If the command returns a figure greater than 500,000 words, split the PDF into logical volumes before uploading.

### 2. File size exceeds the 200MB threshold

High-resolution photographs, full-color diagrams, and uncompressed background textures inflate PDF file sizes rapidly. A slide deck or annual report containing uncompressed print-quality images can easily exceed the 200MB ceiling even when containing relatively few words.

You can downsample images without losing readable text quality using Ghostscript:

```bash
gs -sDEVICE=pdfwrite -dCompatibilityLevel=1.4 -dPDFSETTINGS=/ebook    -dNOPAUSE -dQUIET -dBATCH -sOutputFile=compressed_document.pdf input_document.pdf
```

Using the `/ebook` preset compresses images to screen resolution, which preserves crisp typography for text extraction while keeping the file comfortably beneath the 200MB ceiling.

### 3. Password protection and digital rights management

Google explicitly documents that copy-protected PDFs cannot be imported into NotebookLM. If a document requires a password to open, or if the PDF permission flags disable content copying and text extraction, the ingestion engine aborts.

To verify whether your PDF carries encryption restrictions, inspect it with `pdfinfo`:

```bash
pdfinfo sample_document.pdf | grep "Encrypted"
```

If the document is encrypted, decrypt it using `qpdf` (provided you have lawful access to the file):

```bash
qpdf --decrypt encrypted_input.pdf decrypted_output.pdf
```

### 4. Image-only scanned pages without OCR

Scanned documents from physical archives, historical records, and paper contracts are often stored as raw image containers. While the document appears readable to human eyes, the underlying PDF contains zero extractable text strings.

NotebookLM relies on text extraction rather than running optical character recognition (OCR) on every page of every uploaded file. If `pdftotext` yields zero words, run an OCR pre-processing tool such as `ocrmypdf` to inject a searchable text layer:

```bash
ocrmypdf --deskew --clean input_scan.pdf searchable_output.pdf
```

### 5. Corrupted cross-reference tables and broken fonts

PDFs generated by legacy export software or incomplete downloads often suffer from corrupted cross-reference (xref) tables or malformed font dictionaries. Web browsers can often render these files by ignoring minor syntax faults, but programmatic ingestion parsers will fail.

You can repair corrupted PDF xref tables and linearize documents using `qpdf`:

```bash
qpdf --replace-input --linearize corrupted_file.pdf
```

## How to prepare and split oversized PDFs for NotebookLM

When working with research materials that exceed 500,000 words or 200MB, the most effective solution is to divide the document into modular, coherent sections. Rather than splitting pages randomly, organizing chunks around logical topic boundaries produces superior AI retrieval results.

### Designing semantic document chunks

Language models perform grounding most effectively when each source represents a self-contained subject. When dividing oversized documents, follow these structural guidelines:

- **Split by chapters or modules:** For technical books and training manuals, separate each chapter into an independent PDF.
- **Split by regulatory subparts:** For statutory compilations and compliance manuals, partition files by title, article, or regulatory domain.
- **Preserve metadata headers:** Ensure the first page of each split PDF includes the master title, publication date, and author attribution. This ensures the model retains full provenance context when generating citations.

### Splitting PDFs with Python and pypdf

Using Python and the lightweight `pypdf` library, you can automate document splitting based on target page ranges:

```python
from pypdf import PdfReader, PdfWriter

def split_pdf_by_pages(input_path: str, chunk_size: int = 150) -> None:
    reader = PdfReader(input_path)
    total_pages = len(reader.pages)
    
    for start in range(0, total_pages, chunk_size):
        end = min(start + chunk_size, total_pages)
        writer = PdfWriter()
        
        for page_num in range(start, end):
            writer.add_page(reader.pages[page_num])
            
        output_filename = f"split_part_{start + 1}_to_{end}.pdf"
        with open(output_filename, "wb") as f:
            writer.write(f)
        print(f"Generated {output_filename} with {end - start} pages")

if __name__ == "__main__":
    split_pdf_by_pages("mammoth_archive.pdf", chunk_size=150)
```

This modular approach allows you to import multi-thousand-page corporate archives into NotebookLM as separate uploaded source files while remaining safely within both the 500,000 words limit and 200MB ceiling.

### Command-line extraction with Poppler utilities

If you prefer terminal-based shell scripts, Poppler's `pdfseparate` utility extracts page ranges cleanly:

```bash
pdfseparate -f 1 -l 200 master_record.pdf part_1_%d.pdf
pdfunite part_1_*.pdf part_1_complete.pdf
rm part_1_*.pdf
```

## Querying document collections that exceed NotebookLM limits

NotebookLM is an exceptional tool for personal research, study guides, and ad-hoc synthesis. However, engineering teams and knowledge-intensive enterprises quickly encounter systemic friction when attempting to manage company-wide document archives inside personal notebooks.

Manual file uploads do not scale across thousands of contracts, client records, and engineering specifications. Furthermore, personal notebooks lack shared organizational ownership, programmatic ingestion pipelines, and multi-agent accessibility.

### Comparing document limits across AI platforms

Different AI platforms adopt fundamentally different architectures for handling large document collections:

- **Anthropic Claude Projects:** Claude manages documents through conversation context. A chat accepts multiple attached files, while Projects allow document uploads without an arbitrary fixed file-count cap. However, total project knowledge is bounded by Claude's active context window. Reaching that context ceiling forces users to swap documents in and out of the project repository manually.
- **Google Gemini Notebook (NotebookLM):** Offers a high single-source limit of 500,000 words and 200MB, supporting up to 50 sources on free plans and 600 sources on Ultra plans. However, it operates as a closed research interface without programmatic read-write APIs or remote tool connections for external coding agents.

When organizations manage massive document collections spanning gigabytes of data, neither manual prompt stuffing nor isolated notebooks provide a scalable operating model.

### The persistent workspace approach on Fast.io

The scalable path for large corpora is to store documents in persistent cloud workspaces that combine large-file capacity with an intelligent retrieval layer. Instead of struggling with per-source word caps or prompt context ceilings, organizations store their full document libraries in shared workspaces that both humans and autonomous AI agents can query on demand.

Fast.io provides a modern workspace platform built specifically for collaborative human and agent workflows:

- **Chunked uploads for multi-gigabyte collections:** Upload large PDF libraries, technical documentation sets, and media archives without client-side file size roadblocks or timeout failures.
- **Direct cloud import:** Import existing archives directly from Google Drive, Dropbox, Box, and OneDrive without routing data through local machines. Cloud sync is available for Dropbox, Box, and OneDrive, while Google Drive offers import today with sync coming soon.
- **Intelligence Mode and built-in RAG:** Once documents land in a workspace, [Intelligence Mode](/product/ai/) auto-indexes text, presentations, and spreadsheets for semantic and full-text search. Both humans and AI agents can query the entire corpus, receiving verifiable answers backed by source citations without stuffing hundreds of thousands of words into prompt context.
- **Structured document extraction with Metadata Views:** Turn document collections into queryable, structured databases with [Metadata Views](/product/document-data-extraction/). Define schema fields in natural language (contract values, effective dates, policy numbers, counterparties), and the system extracts structured data from PDFs, scanned forms, and spreadsheets into sortable, filterable views.
- **Remote Model Context Protocol (MCP) tooling:** Autonomous coding agents (such as Claude Code, Cursor, Codex, and OpenClaw) connect directly to Fast.io workspaces via Streamable HTTP at `https://mcp.fast.io/mcp` or `https://mcp.fast.io/mcp/key` (with legacy SSE available at `https://mcp.fast.io/sse`). For setup patterns, see [Storage for AI Agents](/storage-for-agents/). Instead of attaching massive PDFs to prompt context, agents search workspace files dynamically and retrieve only the precise passages required to answer a question.
- **Collaborative Notes and team coordination:** Humans and AI agents can co-edit notes, synthesize research takeaways, and organize action items in real time.
- **Granular permissions and audit logging:** Manage access control across organizations, workspaces, folders, and files, backed by an append-only audit log that tracks every document interaction.
- **Branded shares and portals:** Deliver client-ready research dossiers, project files, and documents through customizable Send, Receive, and Exchange links with expiration controls and recipient access rules.

By pairing local research tools like NotebookLM with persistent, intelligent workspaces, teams maintain access to infinite document archives while equipping their AI agents with fast, citation-backed retrieval.

## Frequently asked questions

### Is there a page limit for PDFs in NotebookLM?

No, NotebookLM does not enforce a page limit on PDFs. Google evaluates documents strictly by word count and raw file size. A PDF can contain 50 pages or 1,200 pages, provided it does not exceed the 500,000 words limit and stays under 200MB.

### What is the maximum word count per source in NotebookLM?

The maximum word count per source in NotebookLM is 500,000 words. This ceiling applies to all individual source types, including PDFs, Google Docs, Word documents, text files, and audio transcripts.

### What is the maximum file size for a PDF in NotebookLM?

The maximum file size for an uploaded PDF in NotebookLM is 200MB. Uploaded files that exceed the 200MB source size limit must be compressed or split into smaller segments before ingestion.

### Why is my PDF not importing into NotebookLM?

A PDF upload usually fails because the file exceeds 500,000 words, exceeds 200MB in size, is password-protected or encrypted, contains scanned pages without an OCR text layer, or has corrupted cross-reference tables.

### How many total PDFs can you upload to NotebookLM?

Standard free accounts can upload up to 50 sources per notebook. Google AI Plus expands capacity to 100 sources, Google AI Pro supports up to 300 sources, and Google AI Ultra tiers support up to 600 sources per notebook.

### Can NotebookLM read scanned PDFs without an OCR text layer?

No, NotebookLM requires an extractable text layer to process and ground document content. If you upload a pure image scan, run OCR software like ocrmypdf before uploading to inject readable text.

### What happens if a PDF exceeds the 500,000-word limit?

If a PDF exceeds 500,000 words, the upload will fail or the parser will stop reading beyond the 500,000-word mark. Excess text will not be indexed or referenced by the AI assistant.

### How do NotebookLM limits compare to Claude Projects?

NotebookLM allows up to 500,000 words and 200MB per file with up to 50 to 600 sources per notebook. In contrast, Claude Projects accepts files up to 30MB each with no fixed file-count cap, but total project knowledge is bounded by Claude's context window.

## Sources

- [Google: Gemini Notebook Help](https://support.google.com/gemininotebook/answer/16269187?hl=en) — Google NotebookLM limits each uploaded PDF source to 500,000 words or 200MB file size ceiling with no page limit, rejecting copy protected files and oversized uploads.
- [Google: Gemini Notebook Help](https://support.google.com/gemininotebook/answer/16215270?hl=en) — Free standard accounts include up to 50 sources per notebook with up to 500,000 words and 200MB per source, expanding up to 600 sources on Google AI tiers.

## About Fast.io

Fast.io provides shared workspaces where people and AI agents work on the same files, with built-in semantic search and citation-backed chat over what they hold. Agents reach it through a remote MCP server at https://mcp.fast.io/mcp, a REST API at https://api.fast.io/current/, and a command line client published on npm as @vividengine/fastio-cli.
