Qwen3-tts
Local text-to-speech using Qwen3-TTS-12Hz-1.7B-CustomVoice. Use when generating audio from text, creating voice messages, or when TTS is requested. Supports 10 languages including …
paki81
@paki81
What This Skill Does
Local text-to-speech tool using Qwen3-TTS-12Hz-1.7B-CustomVoice model. Generates audio from text with support for 10 languages, 9 premium speaker voices, and instruction-based voice control (emotion, tone, style). Runs entirely offline after initial model download.
Replaces cloud-based TTS services like ElevenLabs by providing a fully offline, privacy-preserving alternative with multilingual voice control.
When to Use It
- Generate Italian speech from text with natural voice
- Create audio files with specific emotional tone (e.g., enthusiastic, calm)
- Produce voice messages for apps or notifications using different speaker voices
- Batch convert written content to speech for accessibility or audio content
- Test TTS output offline without relying on cloud APIs or internet connection
Install
$ openclaw skills install @paki81/qwen-ttsQwen TTS
Local text-to-speech using Hugging Face's Qwen3-TTS-12Hz-1.7B-CustomVoice model.
Quick Start
Generate speech from text:
scripts/tts.py "Ciao, come va?" -l Italian -o output.wav
With voice instruction (emotion/style):
scripts/tts.py "Sono felice!" -i "Parla con entusiasmo" -l Italian -o happy.wav
Different speaker:
scripts/tts.py "Hello world" -s Ryan -l English -o hello.wav
Installation
First-time setup (one-time):
cd skills/public/qwen-tts
bash scripts/setup.sh
This creates a local virtual environment and installs qwen-tts package (~500MB).
Note: First synthesis downloads ~1.7GB model from Hugging Face automatically.
Usage
scripts/tts.py [options] "Text to speak"
Options
-o, --output PATH- Output file path (default: qwen_output.wav)-s, --speaker NAME- Speaker voice (default: Vivian)-l, --language LANG- Language (default: Auto)-i, --instruct TEXT- Voice instruction (emotion, style, tone)--list-speakers- Show available speakers--model NAME- Model name (default: CustomVoice 1.7B)
Examples
Basic Italian speech:
scripts/tts.py "Benvenuto nel futuro del text-to-speech" -l Italian -o welcome.wav
With emotion/instruction:
scripts/tts.py "Sono molto felice di vederti!" -i "Parla con entusiasmo e gioia" -l Italian -o happy.wav
Different speaker:
scripts/tts.py "Hello, nice to meet you" -s Ryan -l English -o ryan.wav
List available speakers:
scripts/tts.py --list-speakers
Available Speakers
The CustomVoice model includes 9 premium voices:
| Speaker | Language | Description |
|---|---|---|
| Vivian | Chinese | Bright, slightly edgy young female |
| Serena | Chinese | Warm, gentle young female |
| Uncle_Fu | Chinese | Seasoned male, low mellow timbre |
| Dylan | Chinese (Beijing) | Youthful Beijing male, clear |
| Eric | Chinese (Sichuan) | Lively Chengdu male, husky |
| Ryan | English | Dynamic male, rhythmic |
| Aiden | English | Sunny American male |
| Ono_Anna | Japanese | Playful female, light nimble |
| Sohee | Korean | Warm female, rich emotion |
Recommendation: Use each speaker's native language for best quality, though all speakers support all 10 languages (Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, Italian).
Voice Instructions
Use -i, --instruct to control emotion, tone, and style:
Italian examples:
"Parla con entusiasmo""Tono serio e professionale""Voce calma e rilassante""Leggi come un narratore"
English examples:
"Speak with excitement""Very happy and energetic""Calm and soothing voice""Read like a narrator"
Integration with OpenClaw
The script outputs the audio file path to stdout (last line), making it compatible with OpenClaw's TTS workflow:
# OpenClaw captures the output path
cd skills/public/qwen-tts
OUTPUT=$(scripts/tts.py "Ciao" -s Vivian -l Italian -o /tmp/audio.wav 2>/dev/null)
# OUTPUT = /tmp/audio.wav
Performance
- GPU (CUDA): ~1-3 seconds for short phrases
- CPU: ~10-30 seconds for short phrases
- Model size: ~1.7GB (auto-downloads on first run)
- Venv size: ~500MB (installed dependencies)
Troubleshooting
Setup fails:
# Ensure Python 3.10-3.12 is available
python3.12 --version
# Re-run setup
cd skills/public/qwen-tts
rm -rf venv
bash scripts/setup.sh
Model download slow/fails:
# Use mirror (China mainland)
export HF_ENDPOINT=https://hf-mirror.com
scripts/tts.py "Test" -o test.wav
Out of memory (GPU): The model automatically falls back to CPU if GPU memory insufficient.
Audio quality issues:
- Try different speaker:
--list-speakers - Add instruction:
-i "Speak clearly and slowly" - Check language matches text:
-l Italianfor Italian text
Model Details
- Model: Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice
- Source: Hugging Face (https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice)
- License: Check model card for current license terms
- Sample Rate: 16kHz
- Output Format: WAV (uncompressed)
Top skills in this category
Nano Banana Pro
@steipeteGenerate/edit images with Nano Banana Pro (Gemini 3 Pro Image). Use for image create/modify requests incl. edits. Supports text-to-image + image-to-image; 1K/2K/4K; use --input-image.
AdMapix
@fly0pantsAdMapix raw data layer for ad creatives, apps, rankings, downloads/revenue, and market metadata. Returns structured JSON from the AdMapix API; the calling ag...
API Gateway
@byungkyuConnect to external services through Maton-managed API routes. Use this skill only after the user names the target app, account, and task. Start with read/list calls when possible and follow the app-specific reference before any change.
YouTube Watcher
@michaelgatharaFetch and read transcripts from YouTube videos. Use when you need to summarize a video, answer questions about its content, or extract information from it.
SuperDesign
@mpociotExpert frontend design guidelines for creating beautiful, modern UIs. Use when building landing pages, dashboards, or any user interface.