Doc2Markdown
Lightweight document utility designed to convert files to Markdown (MD), built specifically for intelligent agents (e.g., OpenClaw, ClaudeCode) to read and p...
Haoyt27
@haoyt27
What This Skill Does
Converts a wide range of document formats (docx, pdf, pptx, xlsx, images, epub, and more) into Markdown files or packages, preserving structure, tables, and images. It uploads the file to a cloud service for parsing, polls for completion, and downloads the result to the source file's directory.
Replaces manual document-to-Markdown conversion and complex parsing libraries by providing a single CLI command that handles dozens of formats with no external dependencies.
When to Use It
- Convert a PDF or Word document to Markdown for AI agents to read and analyze
- Extract content from a PowerPoint presentation as a Markdown file
- Convert an Excel spreadsheet to Markdown to inspect tabular data
- Transform an image of a document (JPG/PNG) into readable Markdown text
- Export an EPUB or CHM ebook to Markdown for further processing
- Batch convert multiple document formats to Markdown for a unified content pipeline
Install
$ openclaw skills install @haoyt27/doc2markdowndoc2markdown
Document conversion assistant that automatically converts documents to Markdown (MD), saving output to the same directory as the source file. Designed to help intelligent agents read and process document content in various formats.
Quick Start
# Convert document (auto-polls for 60s, downloads if complete, returns doc ID if timeout)
node scripts/doc2markdown.js convert <file_path> # Downloads MD package
node scripts/doc2markdown.js convert <file_path> --md # Downloads single MD file
# Check status and download (for documents that exceeded timeout)
node scripts/doc2markdown.js check <doc_id> <original_file_path> # Downloads MD package
node scripts/doc2markdown.js check <doc_id> <original_file_path> --md # Downloads single MD file
Capabilities
- Supported formats: docx, doc, pdf, ppt, pptx, xls, xlsx, jpg, jpeg, png, ceb, teb, caj, odt, ofd, cebx, odp, ott, wps, ods, et, dps, epub, chm, sdc, sdd, sdw, mobi, etc.
- Preserves document structure, tables, and images
- No API Key or account required, zero external dependencies
- Downloaded ZIP files are extracted to
{doc_id}_{filename}/under the source file's parent directory; single MD files are saved directly there
When to Use
- User requests to "read", "extract", "convert", or "view" a document
- User provides a document path and asks about its content
- User needs to summarize or analyze a document
- User needs to convert document content to Markdown package
Download Modes
This tool supports two download modes:
--mdmode: Downloads a single merged MD file to the source file's parent directory. Images are not included- MD package: Downloads and extracts a ZIP package to
{doc_id}_{filename}/in the source file's parent directory. Includes image files and tables, tables are rendered in HTML format
Choosing the Right Mode
| User Intent | Example Phrases | Mode to Use |
|---|---|---|
| Read / view / analyze a document | "read this file", "what's in this doc", "summarize this PDF" | --md (single MD file) |
| Explicitly convert to MD | "convert to MD", "export as markdown", "转成MD" | MD package (default, no --md); use --md only if user specifically asks for a single file |
Workflow
convert — Convert Document
- Invoke file parsing service
- Auto-poll conversion status (up to 60 seconds)
- Completes within 60s → Auto-download to source file directory
- Exceeds 60s → Return doc ID for subsequent
checkquery
check — Query and Download
- Provide the previously returned doc ID
- Download if complete, otherwise continue polling for 60 seconds
- Prompt to retry later if still not complete
Data & Privacy
convertuploads files to the docchain cloud service (lab.hjcloud.com) for parsing. Results are returned as a ZIP archive and extracted locally.- All transfers use HTTPS encryption.
- Users should ensure that documents do not contain sensitive or confidential information unless they have verified the service's data handling practices.
- Service endpoint: https://lab.hjcloud.com/llmdoc
Feedback & Support
For parsing errors, format issues, or other problems, please submit an issue on GitHub: https://github.com/wct-lab/docchain-skills
Top skills in this category
Nano Pdf
@steipeteEdit PDFs with natural-language instructions using the nano-pdf CLI.
Word / DOCX
@ivangdavilaCreate, inspect, and edit Microsoft Word documents and DOCX files with reliable styles, numbering, tracked changes, tables, sections, and compatibility check...
Excel / XLSX
@ivangdavilaCreate, inspect, and edit Microsoft Excel workbooks and XLSX files with reliable formulas, dates, types, formatting, recalculation, and template preservation...
Markdown Converter
@steipeteConvert documents and files to Markdown using markitdown. Use when converting PDF, Word (.docx), PowerPoint (.pptx), Excel (.xlsx, .xls), HTML, CSV, JSON, XML, images (with EXIF/OCR), audio (with transcription), ZIP archives, YouTube URLs, or EPubs to Markdown format for LLM processing or text analysis.
Powerpoint / PPTX
@ivangdavilaCreate, inspect, and edit Microsoft PowerPoint presentations and PPTX decks with reliable layouts, templates, placeholders, notes, charts, and visual QA. Use...