Alibaba's Qwen Audio 3.0 TTS Plus Tops the Competition in Text-to-Speech Rankings
Alibaba's latest text-to-speech model, Qwen-Audio-3.0-TTS-Plus, has taken the lead on Artificial Analysis' Speech Arena leaderboard for provider voices. The model achieved an Elo score of 1,236, placing it just ahead of SpeechifyAI's Simba 3.2, which scored 1,234. Other top contenders include Gemini 3.1 Flash TTS with a score of 1,214 and Sonic 3.5 at 1,207.
Model Versions and Capabilities
The Qwen-Audio-3.0-TTS-Plus comes in two versions. The Flash variant is designed for real-time interaction, boasting a latency of approximately 300 milliseconds. The Plus version, on the other hand, focuses on delivering high-quality speech output. The model supports 16 languages, including less commonly covered ones such as Tagalog, Malay, Thai, and Vietnamese, as well as several Chinese dialects.
Users can control the speaking style using natural language commands. The model also supports nonverbal cues through tags like "[angry]" or "[giggles]." Alibaba claims that the model handles noisy or echo-heavy reference recordings better than previous versions when cloning voices.
Performance and Pricing
Stay ahead of the AI curve
The most important updates, news, and content — delivered weekly.
No spam. Unsubscribe anytime.
Speed is a notable weakness for the model. It processes text at a rate of 16 characters per second, which is significantly slower than competitors like Sonic 3.5 (120 characters per second) and Simba 3.2 (30.2 characters per second). Pricing for the model is set at $27.60 per million characters through Alibaba Cloud Model Studio. A collection of audio samples is available for listening.
Background on Alibaba and Qwen
Alibaba is a Chinese multinational technology conglomerate specializing in e-commerce, cloud computing, and artificial intelligence. The Qwen series of AI models, developed by Alibaba's DAMO Academy, includes large language models, vision models, and audio models. Qwen-Audio-3.0-TTS-Plus represents the latest advancement in the company's text-to-speech capabilities, building on previous iterations that have been well-received in the AI community.
Artificial Analysis Leaderboard
Artificial Analysis is an independent benchmarking platform that evaluates AI models across various tasks, including text-to-speech. Its Speech Arena leaderboard ranks provider voices based on Elo scores derived from human evaluations. The leaderboard is widely used by developers and researchers to compare the quality and performance of different TTS models.

