PDF Extraction (auto text/OCR)
Extract text, tables, and metadata from PDFs. Auto-detects native text vs scanned image pages and routes to pdfplumber or Tesseract OCR.
AlexHT Hung
@alex-ht
Install
$ openclaw skills install @alex-ht/pdf-extractionPDF Extraction (auto text / OCR)
Extract content from PDFs without deciding whether each page is selectable text or a scan.
| Page type | Tool |
|---|---|
| Native text | pdfplumber (text + tables) |
| Scanned / image | PyMuPDF + Tesseract OCR |
Install
# System OCR engine (required for scanned pages)
# Ubuntu/Debian:
sudo apt install tesseract-ocr tesseract-ocr-eng
# Optional Traditional Chinese:
# sudo apt install tesseract-ocr-chi-tra
# CLI (from this skill folder)
pip install -e .
# or: python3 -m pip install -e .
Quick start
# Auto-detect text vs OCR per page → print text
pdf-extract document.pdf
# Write to file
pdf-extract document.pdf -o out.txt
# See which mode each page will use
pdf-extract document.pdf --analyze-only
# Tables + Markdown
pdf-extract document.pdf --tables --format markdown -o out.md
# JSON (includes per-page mode)
pdf-extract document.pdf --format json -o out.json
# Force mode
pdf-extract scan.pdf --mode ocr --ocr-lang eng
pdf-extract text.pdf --mode text
# Page range
pdf-extract doc.pdf --pages 1-3,5
# Also: python -m pdf_extract document.pdf
Auto mode rules
For each page (--mode auto, default):
- Try native text via pdfplumber; count chars and embedded images
- Enough text → text
- Sparse text (default < 40 non-whitespace chars) or image-heavy → ocr
Tune with --min-text-chars. Override with --mode text or --mode ocr.
Agent usage
When the user provides a PDF and wants content extracted:
- Prefer
pdf-extract <file>(auto mode). Do not ask whether it is text or scanned. - Use
--analyze-onlyif you only need routing diagnostics. - Use
--tableswhen tables matter (works best on native-text pages). - For Chinese scans, set
--ocr-lang eng+chi_traifchi_trais installed. - Surface stderr summary lines (which pages used text vs ocr) when useful.
CLI reference
| Flag | Meaning |
|---|---|
-o, --output | Write to file (default stdout) |
-f, --format | text | json | markdown |
--mode | auto | text | ocr |
--pages | e.g. 1-3,5 |
--tables | Extract tables on text pages |
--layout | Preserve native text layout |
--meta | Include PDF metadata |
--ocr-lang | Tesseract langs (default eng) |
--ocr-dpi | OCR render DPI (default 200) |
--analyze-only | Classification JSON only |
-q, --quiet | Suppress progress on stderr |
Dependencies
- Python 3.10+:
pdfplumber,pymupdf,Pillow(seepyproject.toml) - System:
tesseract(and language packs as needed)
Top skills in this category
Nano Pdf
@steipeteEdit PDFs with natural-language instructions using the nano-pdf CLI.
Word / DOCX
@ivangdavilaCreate, inspect, and edit Microsoft Word documents and DOCX files with reliable styles, numbering, tracked changes, tables, sections, and compatibility check...
Excel / XLSX
@ivangdavilaCreate, inspect, and edit Microsoft Excel workbooks and XLSX files with reliable formulas, dates, types, formatting, recalculation, and template preservation...
Markdown Converter
@steipeteConvert documents and files to Markdown using markitdown. Use when converting PDF, Word (.docx), PowerPoint (.pptx), Excel (.xlsx, .xls), HTML, CSV, JSON, XML, images (with EXIF/OCR), audio (with transcription), ZIP archives, YouTube URLs, or EPubs to Markdown format for LLM processing or text analysis.
Powerpoint / PPTX
@ivangdavilaCreate, inspect, and edit Microsoft PowerPoint presentations and PPTX decks with reliable layouts, templates, placeholders, notes, charts, and visual QA. Use...