Turn your documents
into a live database
Metadata extraction reads PDFs, contracts, invoices, images, and handwritten documents, then records what it finds as named fields. Add the fields you care about as columns and your workspace becomes sortable and filterable. No templates, no OCR rules.
Found 4 contracts expiring this quarter ($283,000 total). Flagged 4 overdue invoices totaling $51,450.
Contracts Q2Video guide
See metadata extraction in action
Watch how to extract structured data from any document in minutes.
Extract what other tools can't even see
Traditional IDP tools need rigid templates and only handle clearly printed fields. Metadata extraction pulls structured data, handwritten notes, and even subjective judgments that require reading between the lines.
Reads like a human
Pulls handwritten data, inferred judgments, and loosely structured fields, not just perfectly formatted values.
Every file type
PDFs, images, Word docs, spreadsheets, presentations, scanned pages. If Fastio can read it, extraction can read it.
Incremental by design
Add new columns without reprocessing. Re-extract specific fields on demand when files change.
Penalty amounts, policy exclusions, handwritten totals, or roof types. If a human could read it, AI can extract it.
Every file in the workspace, in one table.
Metadata is a view of the whole workspace, not something you set up per project. Every file is a row from the moment it lands, and the columns are whatever fields you care about.
- Narrow the list by folder, or by the categories AI assigns
- Search metadata across every file in the workspace at once
- Columns are yours to add, reorder, and remove at any time
Nothing to create and nothing to name. Open Metadata on a workspace and the table is already there.
AI finds the fields. You pick the columns.
Extraction reads each file and records what it finds as named, typed fields. Adding a column means picking one of those fields, or naming a new one and letting AI extract it across the workspace.
- Search the fields already found in your files, typed as text, number, date, URL, or JSON
- Name a field of your own and check Extract this field with AI
- Extract Missing fills the gaps, Re-extract All refreshes the lot
A title, a summary, a date, an amount. The contents become facts you can read without opening the file.
Your data, fully queryable
Extracted data isn't useful if you can't work with it. Sort, filter, and query across every field, then click any row to open the original document behind it. Your files become a live database, but you never lose the source.
- Filter builder for any field
- Toggle and reorder columns
- Edit and re-extract inline
Questions the folder could never answer become a filter, like which penalty notices come due in the next 30 days.
Metadata is for your agents first.
Fields exist so agents work better with your files: finding the right ones, understanding what is in them, and answering questions without re-reading everything. Your team gets the same fields through the same table.
- Agents match on what the fields say, not on filenames alone
- Questions start from a short list, not the whole workspace
- Anything an agent can narrow down, you can narrow down by hand
Pair with Ripley or your own agent for Q and A over your extracted fields.
From a folder of penalty notices
to a sortable database
A property developer drops a folder of municipal penalty notices into Fastio. Every notice becomes a row, and the fields worth tracking become columns.
Thousands of scanned notices, stop-work orders, and zoning correspondence. Some are clean PDFs, some are photographed pages, and the penalty amount is buried deep in the document. No easy way to ask which sites are past a deadline.
+ hundreds more files
AI reads the notices and records what it finds, so roughly 10 typed fields turn up ready to add as columns. One loosely formatted field, case officer, gets named by hand and extracted across thousands of pages with everything else.
Now sortable, filterable, and ready for workflows.
Same pattern works for bank statements, insurance policies, job-site videos, or any other document type.
Find files by what's inside them
Once a field is extracted, it becomes something you can search on. Query files directly by their extracted values, with number ranges, date windows, and text matches, then click straight through to the source document behind every result.
"Every invoice where amount > 10,000"
"Contracts with renewal_date in the next 30 days"
"Penalty notices where response_deadline is past due"
This is the other half of metadata. Extraction turns documents into typed fields. Search by value turns those fields into a way to pull up exactly the files you need, even across thousands of documents, without opening a single one.
Numeric and date fields compare and range. Text fields match. Every result links back to its original file.
Built for every industry's documents
Teams across property, insurance, field services, and finance use metadata extraction to turn manual review processes into searchable knowledge.
Legal & Compliance
Extract case numbers, property addresses, zoning, penalty amounts, and response deadlines from every municipal notice.
Insurance & Claims
Pull exclusions, endorsements, and form codes from every policy in your book.
Media & Creative
Tag every job-site video with equipment, project type, roof type, and whether the footage is aerial.
Finance & Accounting
Extract P&L line items by year, plus entity, period, balances, and deposits from bank statements.
Storage that scales with your documents
Every plan includes metadata extraction. Start on Starter with 1 TB of storage and scale to 50 TB on Growth. Every plan includes a 14-day trial so you can test real workspaces first.
Start on the Starter plan
Extract metadata on Starter with 1 TB of storage, then scale up as your library grows. Every plan includes a 14-day trial.
Your data stays yours
Extracted data lives in your workspace. No shared models, no leaked training data.
Extraction in seconds
Per-file extraction typically completes in seconds. Batch jobs run async with real-time progress.
Your files already have the answers. Make them queryable.
Add your first column in under a minute. Start on the Starter plan with 1 TB of storage.