AI Models

Alibaba Qwen Audio 3.0 TTS Plus Tops Speech Arena Leaderboard

Alibaba's Qwen-Audio-3.0-TTS-Plus has claimed the top spot on Artificial Analysis' Speech Arena leaderboard for provider voices, achieving an Elo score of 1,236. The model narrowly beats SpeechifyAI's Simba 3.2 (1,234) and other competitors like Gemini 3.1 Flash TTS and Sonic 3.5. Available in Flash and Plus versions, it supports 16 languages and offers natural language style control, though its speed lags behind rivals.

Neura News

Neura News

Neura Market Editorial

July 21, 20262 min read

Originally reported by the-decoder.com

Alibaba Qwen Audio 3.0 TTS Plus Tops Speech Arena Leaderboard

Alibaba's Qwen Audio 3.0 TTS Plus Tops the Competition in Text-to-Speech Rankings

Alibaba's latest text-to-speech model, Qwen-Audio-3.0-TTS-Plus, has taken the lead on Artificial Analysis' Speech Arena leaderboard for provider voices. The model achieved an Elo score of 1,236, placing it just ahead of SpeechifyAI's Simba 3.2, which scored 1,234. Other top contenders include Gemini 3.1 Flash TTS with a score of 1,214 and Sonic 3.5 at 1,207.

Model Versions and Capabilities

The Qwen-Audio-3.0-TTS-Plus comes in two versions. The Flash variant is designed for real-time interaction, boasting a latency of approximately 300 milliseconds. The Plus version, on the other hand, focuses on delivering high-quality speech output. The model supports 16 languages, including less commonly covered ones such as Tagalog, Malay, Thai, and Vietnamese, as well as several Chinese dialects.

Users can control the speaking style using natural language commands. The model also supports nonverbal cues through tags like "[angry]" or "[giggles]." Alibaba claims that the model handles noisy or echo-heavy reference recordings better than previous versions when cloning voices.

Performance and Pricing

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

Speed is a notable weakness for the model. It processes text at a rate of 16 characters per second, which is significantly slower than competitors like Sonic 3.5 (120 characters per second) and Simba 3.2 (30.2 characters per second). Pricing for the model is set at $27.60 per million characters through Alibaba Cloud Model Studio. A collection of audio samples is available for listening.

Background on Alibaba and Qwen

Alibaba is a Chinese multinational technology conglomerate specializing in e-commerce, cloud computing, and artificial intelligence. The Qwen series of AI models, developed by Alibaba's DAMO Academy, includes large language models, vision models, and audio models. Qwen-Audio-3.0-TTS-Plus represents the latest advancement in the company's text-to-speech capabilities, building on previous iterations that have been well-received in the AI community.

Artificial Analysis Leaderboard

Artificial Analysis is an independent benchmarking platform that evaluates AI models across various tasks, including text-to-speech. Its Speech Arena leaderboard ranks provider voices based on Elo scores derived from human evaluations. The leaderboard is widely used by developers and researchers to compare the quality and performance of different TTS models.

Related on Neura Market

More from Neura News

AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google has released three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and knowledge work with 17% fewer output tokens and lower costs. The 3.5 Flash-Lite is the fastest in the series at 350 tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber model, available only to governments and trusted partners via CodeMender, focuses on finding and fixing cybersecurity vulnerabilities. Google also noted that Gemini 3.5 Pro is being tested with partners and that pre-training for Gemini 4 has begun.

Jul 21·5 min read
AI Models

Alibaba Qwen-Image-3.0 renders infographics and tiny text in one pass

Alibaba's Qwen team released Qwen-Image-3.0, an image generator designed for practical applications like newspaper layouts and complex infographics. The model processes prompts of up to 4,500 tokens and can render legible text as small as ten pixels, mathematical formulas, and twelve languages in a single pass. It is currently available through invite-only API access, with plans to integrate it into first-party apps like Qwen Chat soon.

Jul 21·4 min read
AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google DeepMind has introduced three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and multimodal performance with 17% fewer output tokens and lower cost. The 3.5 Flash-Lite is the fastest in its series at 350 output tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber, fine-tuned for cybersecurity, will be available exclusively to governments and trusted partners via the CodeMender agent.

Jul 21·6 min read