Transcription Skills for OpenClaw Agents
26 skills in the OpenClaw catalogue are tagged Transcription, ranked here by downloads over the last 30 days so the list reflects what people are installing now rather than what accumulated the most downloads years ago.
Markdown Converter
@steipeteConvert documents and files to Markdown using markitdown. Use when converting PDF, Word (.docx), PowerPoint (.pptx), Excel (.xlsx, .xls), HTML, CSV, JSON, XML, images (with EXIF/OCR), audio (with transcription), ZIP archives, YouTube URLs, or EPubs to Markdown format for LLM processing or text analysis.
18948k1.9k/30dYoutube
@grpaivaSearch YouTube videos, get channel info, fetch video details and transcripts using YouTube Data API v3 via MCP server or yt-dlp fallback.
69.2k580/30d语音转文字 ElevenLabs STT
@dlazyaiElevenLabs scribe_v1 speech-to-text with auto language detection and optional speaker diarization. Suitable for subtitles, transcription, and meeting notes. ElevenLabs scribe_v1 语音转文字,支持自动语种识别与说话人分离,适合字幕、转录与会议记录。
01.8k531/30d录音转写 Fun ASR
@dlazyaiAlibaba Bailian Fun-ASR recording transcription. Supports Chinese, English and other languages, with auto language detection and speaker diarization. Suitable for subtitles, transcription, and meeting notes. 阿里云百炼 Fun-ASR 录音文件识别,支持中英文及多语种,自动语种识别与说话人分离。适合字幕、转录与会议记录。
01.7k503/30dLocal Whisper
@araa47Local speech-to-text using OpenAI Whisper. Runs fully offline after model download. High quality transcription with multiple model sizes.
1213k387/30ddeyo
@casatwyUse only when the current user explicitly asks to use Deyo to transcribe one provided URL or one exact local audio/video file path, or explicitly asks for Deyo install, status, or troubleshooting. Do not trigger from a mere Deyo mention, ambient context, an implicit attachment, directory browsing, a glob, stdin, a batch request, or inferred permission to log in, install software, or read files.
01.3k332/30dAudiolla
@psyb0tConnect to a user-deployed audiolla server to perform stem separation, mastering, MIR analysis, DSP transforms, and loudness normalization on audio files.
01.1k326/30dListen
@ivangdavilaRepairs garbled speech-to-text input: fixes mistranscribed names, numbers, and commands in voice-dictated messages. Use when a message arrived by voice and a word breaks the sentence, dictation mangles proper nouns, jargon, amounts, times, or email addresses, the user says "no, I said X" or repeats themselves, transcripts contain filler, spoken punctuation, or hallucinated sentences, the user dictates an email or document by voice, or an STT engine (Whisper or cloud speech) needs vocabulary tuning for recurring terms. Not for transcribing audio files or for typed-text typos.
21.9k298/30dFaster Whisper
@theplasmakLocal speech-to-text using faster-whisper. 4-6x faster than OpenAI Whisper with identical accuracy; GPU acceleration enables ~20x realtime transcription. SRT...
58.5k288/30dAzure语音转写免费版
@thcjp使用Azure AI进行批量语音转文字,支持基础转写与时间戳,适合个人用户处理音频。Use when 需要提升效率、自动化流程、批量处理、工作流优化时使用。不适用于需要人工创意判断的任务。适用于独立开发者、企业团队和自动化工作流场景。支持中文交互,无需复杂配置即开即用。输出结果可直接使用,减少二次加工成本。
0423268/30dDaDaScribe advanced speech-to-text transcription & translation
@fablauTranscribe audio and video with the DaDaScribe AI service (YouTube URLs, direct links, or local files). Supports 100+ languages, speaker diarization with named speakers, translation to up to 5 languages, and returns .txt transcripts plus .srt subtitles. Use whenever the user asks to transcribe, capt
0434262/30dMarkItDown Skill
@karmanvermaOpenClaw agent skill for converting documents to Markdown. Documentation and utilities for Microsoft's MarkItDown library. Supports PDF, Word, PowerPoint, Excel, images (OCR), audio (transcription), HTML, YouTube.
03.5k255/30dSpeech is Cheap Transcribe
@ilyakamFast, affordable automatic speech-to-text transcription supporting 100 languages, speaker diarization, word timestamps, and customizable output formats.
53.8k247/30dYouTube Transcript with Speaker Diarization
@patelnavGenerate speaker-aware YouTube transcripts through diarize. Use when generic transcript skills are not enough and you need attributed speakers in TXT, JSON,...
0327244/30dVideo Summary
@lifei68801Video summarization for Bilibili, Xiaohongshu, Douyin, and YouTube. Extract insights from video content through transcription and summarization.
63.1k196/30dOATDA Transcribe Audio
@devcsdeTranscribe audio to text using OATDA's unified audio API. Triggers when the user wants speech-to-text, transcription of meetings, podcasts, voice notes, subt...
0922188/30dVideoLens.io
@shadoprizmTurn videos into professional timestamped reports. Use for YouTube summaries, tutorials, meetings, bugs, UX, privacy, and creator QA.
1427147/30dTranscribe Media
@niuzbTranscribe local and public media with captions first
06675/30dYouTube Summarizer
@abe238Automatically fetch YouTube video transcripts, generate structured summaries, and send full transcripts to messaging platforms. Detects YouTube URLs and provides metadata, key insights, and downloadable transcripts.
68.6k69/30dwatch-cli
@sonpiazWatch any social video → get an architecture diagram, working component, runnable notebook, or step-by-step cheat sheet — automatically.
05868/30dwhisperx
@niuzbWhisperX provides local speech-to-text transcription using OpenAI Whisper, with high-quality offline recognition, no API key required, word-level timestamps,...
035666/30dvideo-analyzer
@tamakooooo鏅鸿兘鍒嗘瀽 Bilibili/YouTube/鏈湴瑙嗛锛岀敓鎴愯浆鍐欍€佽瘎浼板拰鎬荤粨銆傛敮鎸佸叧閿抚鎴浘鑷姩宓屽叆銆?
094261/30dAzure Ai Voicelive Py
@thegovindBuild real-time voice AI applications using Azure AI Voice Live SDK (azure-ai-voicelive). Use this skill when creating Python applications that need real-time bidirectional audio communication with Azure AI, including voice assistants, voice-enabled chatbots, real-time speech-to-speech translation, voice-driven avatars, or any WebSocket-based audio streaming with AI models. Supports Server VAD (Voice Activity Detection), turn-based conversation, function calling, MCP tools, avatar integration, and transcription.
23.0k53/30dAzure Ai Transcription Py
@thegovindAzure AI Transcription SDK for Python. Use for real-time and batch speech-to-text transcription with timestamps and diarization. Triggers: "transcription", "speech to text", "Azure AI Transcription", "TranscriptionClient".
12.9k51/30dAzure语音转写专业版
@thcjp企业级Azure语音转写工具,支持实时流式转写、说话人分离、批量处理、自定义模型及多语言混合转写。
03341/30dpodcast-transcribe
@mteng27Download podcast audio from RSS feeds and transcribe to text using AuralWise API. This skill should be used when the user wants to download podcast episodes, convert podcast audio to text transcripts, or batch-process a podcast library for searchable content. Triggers include downloading podcasts, podcast transcription, audio-to-text conversion, RSS feed downloading, or any request involving podcast audio acquisition and speech-to-text conversion. Covers RSS feed discovery, audio downloading, AuralWise API transcription, and AI-generated content overviews with book references, key concepts, and searchable keywords.
01717/30d
Related topics
- Marketing256
- Xiaohongshu186
- Api Integration185
- Aigc184
- Ai-hive181
- Crawler172
- Competitive-analysis170
- Content-acquisition167
- Json162
- Mcp162
- Agent-skills150
- Ecommerce133
- Pdf133
- Image-generation130
- Email126
- Audio122
- Health110
- Github108
- News105
- Remote-sensing89
- Geo81
- Toolkit80
- Web Search80
- Browser76
- Calendar76
- Git75
- Video-generation74
- Stock73
- Crypto72
- Twitter72
- Documentation71
- Douyin69
- Ocr67
- Trading67
- Competitor-analysis66
- Content-analysis65
- Mental-models63
- Trend-tracking63
- Prompt62
- Home61