n8n
operations

Document & PDF Automation Pack

8 workflows that get data out of PDFs, scans, and invoices and into a system you can actually query

What This Pack Does

Extracts text and structured fields from PDFs, scans, and invoices, stores them with embeddings, and answers natural-language questions about their contents with citations.

Who It's For

Operations and finance teams re-keying data from documents by hand, and anyone sitting on a document archive they cannot search.

Expected Outcome

Turn incoming PDFs, scans, and invoices into searchable, queryable data with no manual re-keying

Highlights

  • Pipeline: Ingest → Extract → Structure → Store → Query
  • Three grades of extraction: digital text, OCR for scans, structured fields for invoices
  • End-to-end invoice processing from Gmail to Google Sheets
  • Natural-language document querying with source citations you can verify
  • Automatic summarisation of anything dropped into Google Drive

Included Workflows(8)

Extract

Structure

Store

Query

Setup Requirements

An n8n instance and OpenAI API key are the baseline. OCR needs a Mistral API key, document storage and querying need a Supabase project with pgvector enabled, and the invoice workflows need Gmail and Google Sheets access.

Example Use Case

A bookkeeping firm receives 400 supplier invoices a month as email attachments in mixed formats. The Gmail invoice workflow files them automatically; digital PDFs go through text extraction while scanned ones go through Mistral OCR; the structured extraction workflow pulls supplier, date, line items, and total into Google Sheets. Everything is stored in Supabase, so when a client asks what they spent with a supplier last quarter, the query workflow answers with citations pointing at the source invoices.

Primary Integrations

OpenAI
Mistral OCR
Supabase
Google Drive
Google Sheets
Gmail
LlamaParse
Vertex AI

Related Packs