SPEC.md
Defines business rules, features, user flows, and out-of-scope items for a batch PDF-to-Obsidian transcription tool using Claude or Gemini APIs.
What this file does
Defines business rules, features, user flows, and out-of-scope items for a batch PDF-to-Obsidian transcription tool using Claude or Gemini APIs.
When to use it
- You are building a CLI tool that processes archival PDFs into structured markdown
- You need to document entity extraction rules and output metadata fields for a transcription pipeline
- You want to specify confidence thresholds and fallback behavior for OCR pipelines
- You are planning a theme consolidation post-processing step across a corpus of notes
Assumes this stack
SPEC.md
For technical implementation details, architecture, and developer documentation, see AGENTS.md.
Table of Contents
Overview
This tool batch-processes historical PDFs using the Claude or Gemini APIs to produce Obsidian-compatible markdown transcriptions with structured metadata. It solves the problem of manually transcribing archival documents at scale — a common bottleneck in academic historical research — by submitting documents in bulk and writing output files ready for direct use in an Obsidian vault.
Three processing pipelines are available:
- Image-based (Claude Batch API): PDFs sent directly as base64; Claude transcribes from vision. 50% cost savings via batch API.
- Vision OCR (macOS only): PDFs rasterized locally, OCR'd via Apple Vision framework, text sent to Claude for correction and entity extraction. Significantly cheaper input costs for typed documents.
- Gemini: Sequential processing with free-tier rate limiting available.
The primary audience is academic historians and archivists who maintain large collections of primary source PDFs and want searchable, linked transcriptions with consistent entity metadata.
Users & Roles
Researcher (Primary User)
- An individual academic or archivist managing a personal or project-level collection of archival PDFs
- Goals: transcribe documents accurately, extract consistent metadata (people, places, organizations, themes), integrate with Obsidian for research notes
- Runs all scripts locally from the command line; no multi-user or permission model
Business Rules
Source Path Parsing
- All processors expect PDFs organized in a five-level directory hierarchy:
Archive / Collection / Box / Folder / file.pdf - The source citation in each output file is derived from this hierarchy in reverse order:
"Folder, Box, Collection, Archive" - If fewer than 5 path parts are present, the script falls back to the parent directory path
Entity Extraction
- Only extract entities explicitly mentioned in the document body/content
- Do not extract people from letterhead, headers, footers, officer lists, or organizational rosters
- Include the author/sender and recipient when identifiable from the letter text
- Mark unclear text as
[illegible]rather than guessing - If uncertain about an entity, omit it
OCR Correction (Vision Pipeline)
- Fix only demonstrable OCR character substitutions — do not silently alter historical content
- When correction is certain (clear substitution + context confirms it): correct silently
- When reconstruction relies on inference: mark as
[reconstructed: ...] - When text is truly unrecoverable: mark as
[illegible] - Paratext (stamps, fax headers, marginal annotations): note their presence but exclude from transcription body; annotate inline (e.g.
[marginal notation present: "2:00pm"])
OCR Confidence Threshold
- The Vision OCR pipeline reports a per-document average confidence score (0.0–1.0) across all text observations on all pages
- Documents below
--confidence-threshold(default 0.5) are flagged in a summary report at run end - Flagged documents are still processed normally — the threshold only triggers a log entry, not skipping
- Flagged documents may benefit from reprocessing with the image-based pipeline (using a vision model rather than OCR text)
Output Metadata Fields
Required YAML frontmatter fields in every output file:
title,creator,publication,source,date,doc_typepeople,organization,locations,themes(all as[[wiki-link]]format lists)tags:to-doandsource/primary/{doc_type}(e.g.source/primary/letter)added(today's date in ISO format:YYYY-MM-DD)
When --skip-claude is used in the Vision pipeline, only source is populated from the path; all other fields are left blank.
Obsidian Templates
The templates/ directory provides two Obsidian note templates for manual note creation:
archival_template.md— general archival documents; usescreatorandpublicationfields; tags default tosource/primary/<item>correspondence_template.md— letters and correspondence; usesauthorandrecipientfields instead ofcreator/publication; tags default tosource/primary/letter
These are reference templates for Obsidian's Templates plugin, not used by the scripts directly.
Transcription Fidelity
- Preserve original spelling, punctuation, and grammar — no modernization or correction
- For multi-page documents, insert HTML page-break comments:
<!-- page 1 -->,<!-- page 2 -->, etc. - Apply document-type-specific boilerplate rules:
- Letters: skip letterhead (addresses, phone numbers, officer/staff lists); preserve date, sender org, recipient
- Newspaper articles: skip mastheads, column headers, page numbers, ad copy; preserve dateline, headline, byline, body
- Reports/memoranda: skip cover page boilerplate; preserve title, date, authoring office, body
- When skipping non-content sections, note the omission inline:
[letterhead omitted],[masthead omitted], etc.
Model Parameters
- Requests use
temperature: 0.1for factual accuracy max_tokens: 8192per document- Single-document test mode uses standard (non-batch) pricing
Theme Consolidation (Post-Processing)
- Run after a full batch completes, not mid-batch
- Consolidation is conservative: Claude only merges themes that clearly refer to the same concept
- When
--applyis used, original themes are preserved inoriginal_themesfrontmatter field and athemes_consolidated: trueflag is set - Backup
.bakfiles are created by default before any file updates
Features
Feature: Batch PDF Processing (Image-Based)
Description: Submit a directory of PDFs to Claude's Batch API in a single job. PDFs are encoded as base64 and sent to Claude's vision model. Polls for completion and writes one Obsidian markdown file per PDF.
Functionality:
- Recursively finds all
.pdffiles in the input directory - Encodes each PDF as base64 and submits as a single batch job
- Polls batch status every N seconds (configurable, default 60)
- Parses Claude's structured YAML + transcription response per document
- Writes output
.mdfiles with YAML frontmatter, Overview, Images embed, Notes, Connections, and Transcription sections - Displays estimated batch cost before submission
Edge Cases:
- YAML parsing failure: logs a warning, saves what it can, continues
- Empty API response: logs the failure, counts toward
failedtotal - No PDF files found in input directory: exits with message
- Missing API key: exits with error before submission
Feature: Vision OCR Pipeline (macOS only)
Description: Rasterize PDFs locally via PyMuPDF, extract text using Apple's Vision framework OCR, then send only the text to Claude for correction and entity extraction. Substantially cheaper than the image-based pipeline for typed documents because input tokens are text rather than images.
Functionality:
- Rasterizes each PDF page to a PNG image at configurable DPI (default 200; use 300 for degraded docs)
- Runs
VNRecognizeTextRequest(accurate level, language correction enabled) on each page image - Reports per-page and per-document average OCR confidence scores
- Documents below
--confidence-thresholdare flagged in a summary at run end - Sends OCR text to Claude Batch API as plain text content (not base64 image)
- Displays estimated cost based on actual OCR character count
- Writes same Obsidian markdown output format as the image-based pipeline
--skip-claude mode:
- Skips the Claude API call entirely (no API key required)
- Writes raw OCR text directly into the Transcription section with minimal frontmatter
- Useful for inspecting OCR quality before committing to a Claude batch, or when Vision confidence is high enough that correction isn't needed
--ocr-out FILE (test script only):
- Saves raw OCR text to a file for offline inspection before the Claude step
Edge Cases:
- Zero text extracted from a page: logged as a warning; document still submitted/written
- Low confidence document: flagged and logged, processing continues normally
- Non-macOS system: exits immediately with a clear error message
Feature: Single Document Testing
Description: Process one PDF using the standard (non-batch) API for quick testing of prompts and output format before committing to a large batch. Available for both the image-based and Vision OCR pipelines.
Functionality:
- Vision version displays raw OCR text and per-page confidence scores before the Claude step
- Displays raw API response, parsed YAML metadata, entity counts, transcription preview (first 500 chars), and token usage/cost breakdown
- Accepts same model, API key, and pipeline-specific options as the corresponding batch processor
Feature: Theme Consolidation
Description: After batch processing, scan all output markdown files, collect all unique themes, and use Claude to identify variants of the same concept. Optionally rewrite files with canonical theme names.
Functionality:
- Extracts themes from YAML frontmatter of all
.mdfiles in a directory - Strips
[[wiki-link]]brackets for analysis, re-applies them on output - Generates a
theme_analysis.mdreport with counts, groups, and reduction statistics --applyflag rewrites frontmatter with canonical themes; preserves originals inoriginal_themes--backupflag (default: on) creates.bakfiles before any writes
User Flows
Flow 1: Vision OCR Batch (Recommended for Typed Documents)
Goal: Transcribe typed archival PDFs cheaply using local OCR + Claude text correction
Steps:
- User installs Vision dependencies:
make install-vision - User organizes PDFs in the expected hierarchy:
Archive/Collection/Box/Folder/*.pdf - User sets API key in
.envor environment - Optionally, user tests one document first:
make test-vision PDF=./sample.pdf- Raw OCR text and per-page confidence scores are displayed
- Claude's corrected output and cost estimate are shown
- User runs:
make process-vision IN=./pdfs OUT=./transcriptions - Script OCRs each PDF locally, logs confidence scores, submits text batch to Claude
- On completion, script retrieves results and writes one
.mdfile per PDF - Any flagged low-confidence documents are listed at the end for manual review
- User opens output directory in Obsidian; entities auto-link via
[[wiki-link]]format
Confidence-based routing:
- High confidence docs (≥0.5): proceed through Vision OCR pipeline
- Low confidence docs flagged in summary: consider reprocessing with
make process(image-based)
Flow 2: Image-Based Batch Processing
Goal: Transcribe a folder of archival PDFs (any quality) using Claude's vision model
Steps:
- User organizes PDFs in the expected hierarchy:
Archive/Collection/Box/Folder/*.pdf - User sets
ANTHROPIC_API_KEYin.envor environment - User runs:
make process IN=./pdfs OUT=./transcriptions - Script encodes each PDF, displays estimated cost, and submits batch job
- Script polls status every 60 seconds, printing progress counts
- On completion, script retrieves results and writes one
.mdfile per PDF - User opens output directory in Obsidian; entities auto-link via
[[wiki-link]]format
Flow 3: OCR-Only (No Claude)
Goal: Quickly dump OCR text to Obsidian notes without any API cost
Steps:
- User runs:
make process-vision IN=./pdfs OUT=./transcriptionswith--skip-claude(or:python batch_pdf_processor_vision.py --input ./pdfs --output ./out --skip-claude) - Script OCRs each PDF, writes minimal
.mdfiles with raw OCR as transcription - User reviews output in Obsidian; selectively reprocesses poor-quality documents with Claude
Flow 4: Consolidate Themes Across a Corpus
Goal: Standardize inconsistent theme tags produced by independent per-document processing
Steps:
- User runs theme analysis first (no file changes):
make consolidate DIR=./transcriptions - User reviews
theme_analysis.mdreport to verify proposed consolidations - User applies changes:
make consolidate-apply DIR=./transcriptions - Script rewrites frontmatter of each
.mdfile with canonical themes;.bakbackups created - If results are unsatisfactory, user restores from
.bakfiles
Out of Scope
Not Currently Implemented
- GUI or web interface
- Spatial column sorting for newspaper OCR (bounding box layout analysis) — groundwork laid with per-page rasterization, but column ordering not implemented
- Duplicate document detection
- CSV/JSON export of extracted metadata
- Progress bar during batch polling
- Multi-language support
- Controlled vocabulary file for theme consolidation (manual taxonomy input)
- Automatic fallback from Vision OCR to image-based pipeline for low-confidence documents
Architectural Constraints
- Vision OCR pipeline is macOS-only (Apple Vision framework dependency)
- No multi-user support — single researcher, local CLI only
- No database — all state is in the filesystem (markdown files)
- Batch jobs may take up to 24 hours; the script blocks while polling
Open Questions
Product
-
Q: Should the output section order (Overview, Images, Notes, Connections, Transcription) be configurable?
- Status: Not currently configurable; hardcoded in
create_obsidian_document
- Status: Not currently configurable; hardcoded in
-
Q: Should low-confidence documents be automatically routed to the image-based pipeline within the same run?
- Status: Currently just logged; manual reprocessing required
Technical
-
Q: How should very large PDFs (exceeding Claude's context window) be handled?
- Status: Currently truncated; no chunking or splitting implemented
-
Q: Should newspaper column bounding boxes be spatially sorted for correct reading order?
- Status: Not implemented; Vision returns observations in approximate top-to-bottom order which may not respect column layout
Last Updated: 2026-03-21 This document is maintained for AI agent context and onboarding.
What's inside
8 major sections: overview, users, business rules, features, user flows, out of scope, open questions, and a table of contents
Change this for your project
- Replace
hepplerj/claude-transcribewith your own repository name in the AGENTS.md link - Replace
Archive / Collection / Box / Folder / file.pdfwith your own directory hierarchy if different - Replace
ANTHROPIC_API_KEYwith your own environment variable name if using a different provider - Replace
--confidence-thresholddefault 0.5 with your own threshold value
Where it goes
Keep in docs/ or alongside the feature. Agents read it to implement against a defined contract.
Worth borrowing
- Separating business rules (entity extraction, OCR correction) from feature descriptions and user flows
- Using a confidence threshold that flags low-quality OCR without blocking processing, enabling manual review
- Providing a
--skip-claudemode for quick OCR inspection before committing to API costs
Related Documents
GPU Selection Guide for Large Language Models (LLMs)
Guides GPU selection for LLM inference, fine-tuning, and training by mapping model sizes, precision levels, and budgets to VRAM requirements.
Community AI Agent Skills Discovery Sources
Catalogs 50+ platforms, repositories, directories, and communities for discovering and sharing AI agent skills across multiple coding tools.
ReleaseKit - Technical Requirements Document
Specifies a Go library and CLI for release automation with conventional commit parsing, validation checks, and workflow orchestration.
api_llm Specification
Defines a workspace of thin HTTP API clients for major LLM providers with no abstraction layer and explicit developer control.