Ocr Skills for OpenClaw Agents
67 skills in the OpenClaw catalogue are tagged Ocr, ranked here by downloads over the last 30 days so the list reflects what people are installing now rather than what accumulated the most downloads years ago.
Markdown Converter
@steipeteConvert documents and files to Markdown using markitdown. Use when converting PDF, Word (.docx), PowerPoint (.pptx), Excel (.xlsx, .xls), HTML, CSV, JSON, XML, images (with EXIF/OCR), audio (with transcription), ZIP archives, YouTube URLs, or EPubs to Markdown format for LLM processing or text analysis.
18948k1.9k/30dOCR - Local (No API Key)
@shaw555Extract text from images using Tesseract.js OCR (100% local, no API key required). Supports Chinese (simplified/traditional) and English.
2019k886/30dImage Ocr
@xejraxExtract text from images using Tesseract OCR
1216k763/30dTencent COS
@shawnminh腾讯云对象存储(COS)和数据万象(CI)集成技能。覆盖文件存储管理、AI处理和知识库三大核心场景。 存储场景:上传文件到云端、下载云端文件、批量管理存储桶文件、获取文件签名链接分享、查看文件元信息、查询数据万象及子服务开通状态。 图片处理场景:图片质量评估打分、AI超分辨率放大、AI智能裁剪、二维码/条形码识别、添加文字水印、获取图片EXIF信息、缩放、裁剪、旋转、格式转换。 文档处理场景:Word/Excel/PPT等办公文档转PDF、文档预览。 媒体处理场景:视频智能封面提取、视频转码、视频截帧、获取媒体信息。 内容审核场景:图片/视频/音频/文本/文档内容审核,检测违规内容。 智能语音
49.6k750/30d多模态图像识别
@linkfox-ai基于多模态AI的图片识别与分析。当用户想分析、描述、从图片URL中提取信息、image recognition, image analysis, image description, image content understanding, OCR text recognition, visual Q&A时触发此技能。当用户提到图片识别、图片分析、图片描述、识别图片内容、分析产品图、从图片中读取文字、描述图片、提取视觉内容或理解照片内容时触发。当用户提供图片URL并就其视觉内容提问时,即使未明确说"图片识别",也应触发此技能。
11.3k393/30dmineru document extractor
@mineru-extractMinerU document extraction — convert PDFs, scanned documents, images, Word (DOC/DOCX), PowerPoint (PPT/PPTX), Excel (XLS/XLSX), and web pages into clean Mark...
95.0k377/30d百度文档解析pipeline-parser
@maglanyulan调用百度文档解析API解析文档。支持PDF、Word、Excel、PPT、图片等18+格式。提取文本、表格、版面分析、OCR识别及RAG文档分块。当用户需要解析文档、提取文本/表格、分析文档结构、处理扫描件时使用。触发词:文档解析、PDF解析、Word解析、表格提取、OCR、文档分析、提取文本、文档结构、扫描识别。
01.2k375/30dsocial_security_card_ocr
@scnet-sugon仅在用户明确提及“社保卡”、“社会保障卡”、“医保卡”、“社保卡识别”等特定词汇时触发,用于识别社会保障卡上的姓名、社保号码、身份证号等核心信息。严禁用于通用 OCR 或其他非社保卡类证件识别。
01.1k329/30dPDF万能大师
@ford828全能 PDF 智能处理技能,覆盖 25 项能力五大核心(Office 高保真互转、扫描件 OCR、PDF 直接编辑、加密水印权限、批量合并拆分重命名)、九大增值(表格提取、图文解析、摘要、脱敏、元数据清洗、自动命名、版本对比、文档问答、规格对比矩阵)、四大扩展(智能压缩、排版保留翻译、合同风险审查、发票验真归档)、七大壁垒(电子签名审批流、批量内容替换、表单创建、图纸测量、翻页电子书、印刷级导出、密文级永久删除)。当用户上传 PDF 并提出转换、识别、编辑、加密、压缩、翻译、审查、归档、签名、替换、表单、测量、电子书、印刷、删除等需求,或提到"转Word""OCR""加水印""压缩到xxMB"
0831328/30dMongolian AI for Codex
@youteacherasiaUse the Mongol AI API for Mongolian translation, script conversion, conversation, composition, OCR, ASR, TTS, and Word/PDF translation. Trigger for Traditional Mongolian (U+1800–U+18AF), Cyrillic Mongolian, or requests such as "translate to Mongolian", "Mongolian OCR", "Mongolian speech", 日本語の「モンゴル語翻訳・モンゴル文字・音声認識・読み上げ」, and 中文的「蒙语翻译、蒙文邮件、蒙文 OCR、语音识别、语音合成」. Requests send text, images, audio, or documents to https://mongol.open-idea.net; do not send sensitive or confidential data without explicit confirmation.
0659316/30dscnet-ocr
@scnet-sugon将图片中的文字、通用文字识别, 票据混贴识别, 印章文字识别, 表格识别, 居民身份证, 银行卡, 社保卡, 户口本, 出生医学证明, 往来港澳通行证, 往来台湾通行证, 台湾居民来往大陆通行证, 港澳居民来往内地通行证, 中国香港身份证, 外国人永久居留身份证, 结婚证, 不动产权证书, 机动车行驶证正页, 机动车行驶证副页, 机动车驾驶证正页, 机动车驾驶证副页, 中国护照, 学历证书, 学历证书电子注册备案表, 学位证书。营业执照, 社会团体法人登记证书, 工会法人资格证书, 宗教活动场所登记证, 民办非企业单位登记证书, 事业单位法人证书, 统一社会信用代码证书, 财务票据统一识别,
01.2k294/30dpersonal_card_ocr
@scnet-sugon将图片中的文字、身份证、银行卡、社保卡、户口本、出生医学证明、往来港澳通行证、往来台湾通行证、台湾居民来往大陆通行证、港澳居民来往内地通行证、中国香港身份证、外国人永久居留身份证、结婚证、不动产权证书、机动车行驶证正页、机动车行驶证副页、机动车驾驶证正页、机动车驾驶证副页、中国护照、学历证书、学历证书电子注册备案表、学位证书等信息识别并提取出来。本技能仅在用户明确要求对图片进行 OCR 识别,并同意将图片上传到第三方 OCR 服务(Scnet)处理时触发。只处理用户主动提供的本地图片路径,不会主动扫描或枚举文件系统。
01.2k292/30dexpense_invoice_ocr
@scnet-sugon支持识别企业财务报销场景的19 种常见票据,包括:增值税发票,增值税卷票,出租车发票,火车票,航空运输电子客票行程单,机动车销售统一发票,定额发票,过路过桥费发票,医疗发票,税收完税证明,船票,非税票据,通用机打发票,汽车票,值税通行费发票,网约车行程单,银联POS签购单,医疗住院发票,医疗费用结算单识别。仅在用户明确要求识别某张本地票据图片时触发,调用前必须确认用户同意上传该文件到 Scnet 远程 OCR 服务。
01.5k284/30dInsurance Claims Intelligence
@gechengling提供多模态医疗票据OCR识别、智能判责、反欺诈检测和全险种覆盖的保险理赔智能分析与自动化支持。
11.5k283/30dbirth_medical_cert_ocr
@scnet-sugon仅在用户明确提及“出生医学证明”、“出生证明”、“医学出生证明”、“新生儿证明”等特定词汇时触发,用于识别出生医学证明上的核心信息(新生儿姓名、出生日期、父母信息、证件号码等)。严禁用于通用 OCR 或其他非出生医学证明类证件识别。
01.1k281/30dInvoices
@ivangdavilaFiles, checks, and audits the invoices you receive: OCR and e-invoice XML, duplicate and fraud checks, VAT deduction, and a searchable archive. Use when an invoice, bill, or receipt arrives as a PDF, photo, email attachment, or portal download and has to be filed; when asked where an old invoice is, or what was spent with a supplier last quarter; when preparing a VAT return or an export for an accountant; when an invoice looks wrong, duplicated, or larger than usual; when a supplier's bank details changed on the document; when a recurring invoice never arrived; and when a credit note, reverse charge, import VAT, or a foreign currency has to be booked correctly. Covers Factur-X/ZUGFeRD, XRechnung, Peppol, supplier normalization, and retention mandates. Not for issuing invoices to your own clients (`invoice`), tracking personal spending (`expenses`), chasing a client who owes you money (`clients`), or building a billing system (`billing`).
22.3k279/30dTencent Cloud COS
@shawnminh腾讯云对象存储(COS)和数据万象(CI)集成技能。覆盖文件存储管理、AI处理和知识库三大核心场景。 存储场景:上传文件到云端、下载云端文件、批量管理存储桶文件、获取文件签名链接分享、查看文件元信息、查询数据万象及子服务开通状态。 图片处理场景:图片质量评估打分、AI超分辨率放大、AI智能裁剪、二维码/条形码识别、添加文字水印、获取图片EXIF信息、缩放、裁剪、旋转、格式转换。 文档处理场景:Word/Excel/PPT等办公文档转PDF、文档预览。 媒体处理场景:视频智能封面提取、视频转码、视频截帧、获取媒体信息。 内容审核场景:图片/视频/音频/文本/文档内容审核,检测违规内容。 智能语音
13.2k278/30dPaddleOCR Text Recognition
@bobholamovicUse this skill whenever the user wants text extracted from images, photos, scans, screenshots, or scanned PDFs. Returns exact machine-readable strings with l...
124.2k276/30dPDF Reader
@panpeter2024Extract text from PDF files with automatic OCR fallback for scanned/image-based PDFs. Use when: (1) a user sends a PDF file and the framework did not auto-in...
11.9k262/30dNanonets OCR
@shhdwiDocument extraction API by Nanonets. Convert PDFs and images to markdown, JSON, or CSV with confidence scoring. Use when you need to OCR documents, extract invoice fields, parse receipts, or convert tables to structured data.
215.0k256/30dMarkItDown Skill
@karmanvermaOpenClaw agent skill for converting documents to Markdown. Documentation and utilities for Microsoft's MarkItDown library. Supports PDF, Word, PowerPoint, Excel, images (OCR), audio (transcription), HTML, YouTube.
03.5k255/30dWiseOCR
@wisediagPDF & Image OCR — Convert a single PDF or image to Markdown via WiseDiag cloud API, with high-accuracy text extraction, table recognition, and multi-column l...
52.1k249/30d文件OCR解析
@coolalam当用户需要将扫描版 PDF、图片、OFD、票据、合同扫描件、法院文书扫描件或其他 Agent 无法直接读取的文件解析为文本/Markdown 时使用。适用于 OCR 识别、扫描件识别、图片转文字、图片转 Markdown、PDF 转 Markdown、文件解析失败、文档无法读取、提取结果为空、乱码、扫描版文件解析不了、agent 解析不了文件、平台无法解析文件、需要调用得理 OCR 接口等场景。优先使用当前 Agent/平台原生解析能力;只有原生解析失败、结果明显不可用,或用户明确要求使用得理 OCR 时才调用接口。
0547219/30dMistral OCR
@yzdameExtract text, tables, and images from PDFs or images using Mistral OCR API and output in Markdown, JSON, or HTML formats.
52.7k204/30d微信QQ自动发消息
@smallccwcWindows 平台微信和 QQ 自动发消息工具。支持搜索联系人、发送消息、截图OCR分析、智能回复建议(需用户确认后发送)。
52.7k201/30d夸克扫描王-OCR 文字识别/文件扫描/转 Office Alibaba-Quark-Scanking-All
@yescan-ai夸克扫描王官方图片处理中心:OCR 文字识别(身份证/社保卡/驾驶证/行驶证/港澳台通行证/学位证/营业执照、增值税发票/火车票/英文发票、医疗检验报告/药检报告、表格/公式/手写体/试卷习题/商品图/通用文字)、拍照翻译、画质增强(去水印/去阴影/去手写/去底色/去摩尔纹、裁剪矫正、高清修复、试卷增强、合同增强、素描/线稿)、图片转 Word/Excel/PDF、AI 证件照生成。
0919195/30dMinerU Doc Parser
@mineru-extractMinerU AI document parser — intelligent document extraction powered by AI. Parse PDFs, scanned documents, images, Word files, PowerPoint slides, and web page...
112.4k190/30dPdf Toolkit
@youpele52Run a local script to work with PDF files, DOCX documents, OCR, and text-to-speech. Use the read tool to load this SKILL.md, then exec the uv run command ins...
11.7k185/30dVeryfi Documents AI
@dbiruliaReal-time OCR and data extraction API by Veryfi (https://veryfi.com). Extract structured data from receipts, invoices, bank statements, W-9s, purchase orders...
161.9k180/30dPDF Extract
@compdf-younaExtract PDF extracts structured data from PDFs and images, including tables, OCR text, images, and stamps, built on ComPDF data extraction and AI document ex...
991.5k179/30dmake-to-markdown
@ebandao777-oss工业级RAG Markdown物料生成技能,使用 markitdown 将各类文档和文件转换为 Markdown 格式。支持 .doc/.ppt 老格式自动预处理(Word/PowerPoint COM / LibreOffice)。启动时自动检测 OS/版本/能力,按平台选择最佳执行路径。触发词:转 Markdown / 转换文档 / markitdown / 文档转 md / 批量转换。当需要将 PDF、Word (.docx/.doc)、PowerPoint (.pptx/.ppt)、Excel (.xlsx, .xls)、HTML、CSV、JSON、XML、图片(含 EXIF/OCR)、音频(含语音转写)、ZIP 压缩包、YouTube 链接或 EPub 电子书转换为 Markdown 格式,为知识库提供统一的"通用语言"时触发此技能。
0419177/30dComPDF Conversion CLI
@compdf-younaMUST use for ANY PDF or image format conversion task — converting PDF and images (JPG/JPEG/PNG/BMP/TIFF/TIF/WEBP/JPEG2000) to 10 formats (Word, Excel, PPT, H...
1011.5k175/30dNutrient OpenClaw
@jdrhyneUse the pinned Nutrient OpenClaw plugin to convert, OCR, extract, redact, watermark, sign, or inspect the last-known local credit record for documents. Route only OpenClaw document-processing requests to its declared tools. Treat processing as an external, credit-consuming DWS transfer that requires a bounded estimate and action-time confirmation for every invocation.
43.7k170/30dbank_check_ocr
@scnet-sugon仅在用户明确提及“银行支票”、“支票号码”、“支票金额”、“支票识别”等特定词汇时触发,用于识别银行支票上的关键信息(号码、日期、大小写金额、签章等)。不适用于通用 OCR 或非支票类图像识别。
0519169/30dfinancial_bill_ocr
@scnet-sugon支持金融单据识别,支持识别多种金融单据,包括银行承兑汇票、电子银行承兑汇票、商业承兑汇票、电子商业承兑汇票、银行支票、银行回单、进账单、电汇凭证、支款凭证、移动支付账单、财政授权支付凭证、海关专用缴款书、海关进/出口货物报关单、国际汇票、商业发票、原产地证明、货物运输保险单、装箱单、提单,结构化提取关键信息。
0466169/30dMongol AI Skill 蒙古语AI技能
@youteacher需在 OpenClaw 配置 MONGOL_AI_SKILL_API_KEY(https://mongol.open-idea.net)。提供蒙古语 API 服务:zh/mw/mn 互译、蒙古语对话与创作、TTS、ASR、OCR、Word/PDF 文档翻译。输出语言以用户需求为准;传统蒙古文(mw)须经外部 AP...
21.9k166/30dmed-chronic-disease-review
@unisound-llm门诊慢病审核(糖尿病/高血压)。输入 OCR 结果数组 JSON,由内部医疗大模型输出审核结论与原因(原始 JSON + 自然语言结论)。
0847164/30dmobile_pay_bill_ocr
@scnet-sugon仅在用户明确提及“支付宝账单”、“微信支付记录”、“交易截图”、“支付明细”、“移动支付账单”等特定词汇时触发,用于识别并结构化提取移动支付交易截图中的时间、商户、金额、收支类型等信息。不适用于通用OCR或非支付类票据识别。
0510164/30dMarkdown convert
@compdf-younaProcess, convert, edit, and extract data from PDF files using the ComPDF Cloud API. Supports format conversion (Word, Excel, Image), page manipulation (merge...
1021.2k154/30dUnlimited-OCR PDF & Document Parsing
@aidenwu0209Convert long documents to complete Markdown with Unlimited-OCR. Supports images, scanned PDFs, OFD, Office and text files through Baidu Cloud, plus local image/PDF inference through SGLang or an OpenAI-compatible server. Use for OCR, PDF-to-Markdown, Chinese/CJK text, tables, formulas, reading order, multi-page scans, invoices, reports, papers, and structured document extraction.
0152153/30dDeepRead OCR
@uday390AI-native OCR platform that turns documents into high-accuracy data in minutes. Using multi-model consensus, DeepRead achieves 97%+ accuracy and flags only u...
76.1k145/30dPdf To Structured
@datadrivenconstructionExtract structured data from construction PDFs. Convert specifications, BOMs, schedules, and reports from PDF to Excel/CSV/JSON. Use OCR for scanned documents and pdfplumber for native PDFs.
94.8k140/30dglkvm
@duzefuRemotely control a target host through the GLKVM IP-KVM HTTP API - keyboard/mouse input, screenshots and OCR, Fingerbot physical button control, and ATX power management. Also manages the GLKVM device itself (reboot, firmware upgrade) and its virtual MSD storage (remote ISO download and mounting). N
01.0k134/30dGeneral Text Recognition OCR - 通用文字识别
@jisuapi图片通用文字 OCR,支持中英文及多语种。当用户说:这张图里的字提取成文本、截图 OCR 一下,或类似通用识图问题时,使用本技能。
101.4k117/30dsm-ocr-scanner
@kaarl92Perform OCR on image files (jpg, png, bmp, gif, tiff) using the system's `tesseract` binary and return extracted plain text.
1740115/30dComPDF Toolkit
@compdf-younaAll-in-one PDF workflow for document tasks.
0113115/30dComPDF Convert Images To Documents
@compdf-younaTurn images into editable business documents.
0112114/30dComPDF Ocr
@compdf-younaRecognize text from scanned PDFs and images.
0105108/30dID Card Recognition OCR - 身份证识别
@jisuapi对身份证等证件图 OCR,返回姓名、号码等字段。当用户说:身份证照片识别一下信息、证件图转文字,或类似证件 OCR 时,使用本技能。
101.0k99/30dBank Card Recognition OCR - 银行卡识别
@jisuapi对银行卡图片 OCR,返回卡号、银行与卡类型等。当用户说:银行卡卡号从照片里读出来、识别这张卡哪家银行,或类似银行卡 OCR 时,使用本技能。
1094697/30dPDF Analysis
@mzlzycaAnalyze the structure, layout, and content of PDF documents using MinerU. Returns structured output preserving headings, tables, images, formulas, and docume...
090094/30dVIN Recognition OCR - VIN识别
@jisuapi对车架号/VIN 图片做识别并返回 VIN 及品牌厂家等信息。当用户说:拍了一张车架号照片帮我识别、从图里读出 VIN,或类似 VIN 图片识别时,使用本技能。
111.0k92/30dDeepSeek Harness Unlimited-OCR GUI Setup
@aidenwu0209Install and configure the native Unlimited-OCR plugin for DeepSeek Harness (DSH) from its Settings GUI, using Baidu Cloud or a local SGLang/OpenAI-compatible service. Use for long-document OCR; PDF, OFD, Office, text, and scanned-image to Markdown; tables, formulas, and reading order; or DSH provider, credential, local inference, GUI setup, verification, and troubleshooting.
08691/30dPaddleOCR OCR & Document Parsing Setup
@aidenwu0209Install and configure two PaddleOCR Agent Skills for text recognition and structured document parsing in Codex, Claude Code, GitHub Copilot, Cursor, OpenCode, OpenClaw, and other compatible agents. Use for OCR and image-to-text from screenshots, photos, scans, and PDFs; Chinese/CJK text and bounding boxes; PDF-to-Markdown/JSON; tables, formulas, layout, and reading order; or endpoint, token, installation, and troubleshooting help.
08690/30dDeepSeek Harness PaddleOCR GUI Setup
@aidenwu0209Install and configure the native PaddleOCR plugin for DeepSeek Harness (DSH) from the Settings → PaddleOCR GUI. Use for OCR and image-to-text from screenshots, scans, and PDFs; Chinese/CJK text; PDF-to-Markdown; structured document parsing with tables, formulas, layout, and reading order; or DSH endpoint, credential, GUI setup, verification, and troubleshooting.
08287/30dPDF Extraction (auto text/OCR)
@alex-htExtract text, tables, and metadata from PDFs. Auto-detects native text vs scanned image pages and routes to pdfplumber or Tesseract OCR.
08184/30dMinerU PDF Parser
@easonai-5589用 MinerU API 解析 PDF/Word/PPT/图片为 Markdown,支持公式、表格、OCR。适用于论文解析、文档提取。
95.8k80/30dBook PDF to Structured JSON
@yasmineliu把整本 PDF 重建为可审计、可上传验证的 JSON/TXT 电子版
25874/30dChat Order OCR|聊天訂單整理工具
@qpooqp777Chat order extraction and grouping from pasted LINE/Facebook messages or OCR text. Use for parsing +1, +2, 一份, 盒, 個等數量、按姓名分組排序、產生可人工覆核的 JSON 或 CSV 名單,以及建立或更新相關的聊天訂單整理工具。
05859/30dPDF Converter
@compdf-younaPDF conversion toolkit featuring AI layout analysis and OCR. Converts PDFs to Word, Markdown, JSON, PPT, CSV, HTML, and XML for seamless LLM data processing.
9593158/30d
Related topics
- Marketing255
- Xiaohongshu186
- Api Integration185
- Aigc184
- Ai-hive181
- Crawler172
- Competitive-analysis170
- Content-acquisition167
- Mcp162
- Json161
- Agent-skills150
- Pdf132
- Image-generation129
- Email125
- Audio122
- Ecommerce111
- Health110
- Github105
- News105
- Remote-sensing89
- Geo81
- Toolkit80
- Web Search79
- Browser76
- Calendar76
- Git74
- Video-generation74
- Crypto72
- Stock72
- Twitter72
- Documentation70
- Douyin69
- Competitor-analysis66
- Content-analysis65
- Trading65
- Mental-models63
- Trend-tracking63
- Prompt62
- Home61