CosyVoice logo

CosyVoice

Free

Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.

AI ChatbotsFreeFree tier
Inputs: text, audioOutputs: audio
Type
Open Source
Company
Alibaba Group

About CosyVoice

CosyVoice is a family of multi-lingual large voice generation models by the FunAudioLLM Team (Alibaba Group). The latest iteration, CosyVoice 3, is designed for zero-shot multilingual speech synthesis in the wild, surpassing its predecessor in content consistency, speaker similarity, and prosody naturalness. Key advancements include a novel speech tokenizer trained via supervised multi-task learning (ASR, emotion recognition, language identification, audio event detection, speaker analysis), a differentiable reward model for post-training, dataset scaling from 10K to 1M hours covering 9 languages and 18 Chinese dialects, and model scaling from 0.5B to 1.5B parameters. CosyVoice 3 supports zero-shot in-context generation, mixed-lingual, emotionally expressive, Chinese dialect, cross-lingual, and instructed voice generation, as well as target speaker fine-tuning and hotfix capability.

Key Features

Novel speech tokenizer via supervised multi-task training (ASR, emotion recognition, language identification, audio event detection, speaker analysis)
Differentiable reward model for post-training applicable to LLM-based speech synthesis
Training data scaled from 10K to 1M hours, covering 9 languages and 18 Chinese dialects
Model scaling from 0.5B to 1.5B parameters for improved performance
Zero-shot in-context generation for content consistency, speaker similarity, and prosody naturalness
Mixed-lingual, emotionally expressive, Chinese dialect, and cross-lingual voice generation
Post-training hotfix capability for instructed voice generation and target speaker fine-tuning

Pros & Cons

Pros
  • High content consistency, speaker similarity, and prosody naturalness compared to prior versions
  • Broad language coverage: 9 languages and 18 Chinese dialects
  • Scalable architecture from 0.5B to 1.5B parameters for flexible deployment
  • Open source with full-stack ability: inference, training, and deployment
  • Novel post-training techniques improve adaptability and controllability
Cons
  • Large model size (1.5B parameters) may require significant computational resources for inference and training
  • Still primarily a research-stage model; production deployment may need additional optimization
  • Limited documentation on real-time latency and streaming capabilities for in-the-wild use

Best For

Zero-shot multilingual speech synthesis in diverse domainsMixed-lingual TTS combining multiple languages in one utteranceEmotionally expressive voice generation for interactive applicationsChinese dialect voice generation (e.g., 18 dialects supported)Cross-lingual voice cloning (e.g., English speaker speaking Chinese)Target speaker fine-tuning for personalized voice assistantsInstructed voice generation with post-training hotfix support

Alternatives to CosyVoice

FAQ

What languages does CosyVoice 3 support?
CosyVoice 3 supports 9 languages and 18 Chinese dialects, covering diverse linguistic domains.
Is CosyVoice open source?
Yes, CosyVoice is an open-source model providing inference, training, and deployment full-stack ability.
How does CosyVoice 3 improve over CosyVoice 2?
CosyVoice 3 introduces a novel speech tokenizer via multi-task training, a differentiable reward model for post-training, and scales training data to 1M hours and model parameters to 1.5B for better content consistency, speaker similarity, and prosody naturalness.
Can CosyVoice 3 generate emotionally expressive speech?
Yes, emotionally expressive voice generation is supported, enabled by the speech emotion recognition task in the tokenizer training.