AI Models

Alibaba Qwen Audio 3.0 TTS Plus Tops Speech Arena Leaderboard

Alibaba's Qwen-Audio-3.0-TTS-Plus has claimed the top spot on Artificial Analysis' Speech Arena leaderboard for provider voices, achieving an Elo score of 1,236. The model narrowly beats SpeechifyAI's Simba 3.2 (1,234) and other competitors like Gemini 3.1 Flash TTS and Sonic 3.5. Available in Flash and Plus versions, it supports 16 languages and offers natural language style control, though its speed lags behind rivals.

Neura News

Neura News

Neura Market Editorial

July 21, 20262 min read
Alibaba Qwen Audio 3.0 TTS Plus Tops Speech Arena Leaderboard

Alibaba's Qwen Audio 3.0 TTS Plus Tops the Competition in Text-to-Speech Rankings

Alibaba's latest text-to-speech model, Qwen-Audio-3.0-TTS-Plus, has taken the lead on Artificial Analysis' Speech Arena leaderboard for provider voices. The model achieved an Elo score of 1,236, placing it just ahead of SpeechifyAI's Simba 3.2, which scored 1,234. Other top contenders include Gemini 3.1 Flash TTS with a score of 1,214 and Sonic 3.5 at 1,207.

Model Versions and Capabilities

The Qwen-Audio-3.0-TTS-Plus comes in two versions. The Flash variant is designed for real-time interaction, boasting a latency of approximately 300 milliseconds. The Plus version, on the other hand, focuses on delivering high-quality speech output. The model supports 16 languages, including less commonly covered ones such as Tagalog, Malay, Thai, and Vietnamese, as well as several Chinese dialects.

Users can control the speaking style using natural language commands. The model also supports nonverbal cues through tags like "[angry]" or "[giggles]." Alibaba claims that the model handles noisy or echo-heavy reference recordings better than previous versions when cloning voices.

Performance and Pricing

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

Speed is a notable weakness for the model. It processes text at a rate of 16 characters per second, which is significantly slower than competitors like Sonic 3.5 (120 characters per second) and Simba 3.2 (30.2 characters per second). Pricing for the model is set at $27.60 per million characters through Alibaba Cloud Model Studio. A collection of audio samples is available for listening.

Background on Alibaba and Qwen

Alibaba is a Chinese multinational technology conglomerate specializing in e-commerce, cloud computing, and artificial intelligence. The Qwen series of AI models, developed by Alibaba's DAMO Academy, includes large language models, vision models, and audio models. Qwen-Audio-3.0-TTS-Plus represents the latest advancement in the company's text-to-speech capabilities, building on previous iterations that have been well-received in the AI community.

Artificial Analysis Leaderboard

Artificial Analysis is an independent benchmarking platform that evaluates AI models across various tasks, including text-to-speech. Its Speech Arena leaderboard ranks provider voices based on Elo scores derived from human evaluations. The leaderboard is widely used by developers and researchers to compare the quality and performance of different TTS models.

Related on Neura Market

More from Neura News

AI Models

fal Opens Developer Access to Meta's Muse Image Model

fal has launched developer and enterprise access to Meta's Muse Image agentic image generation and editing model, available via the Meta Model API at $0.01 per image. Muse Image uses a planner-plus-diffuser architecture with web search, code execution, and self-refinement to improve accuracy on complex prompts. The model supports generation, editing, and reference-driven composition, and is priced to make high-volume workloads economically feasible.

Sep 1·5 min read
General

EFF Urges Courts to Reject AI Copyright Expansion, Citing History of Tech Panics

The Electronic Frontier Foundation has filed amicus briefs in two major generative AI copyright cases, urging courts to reject expanded protections based on what it calls hype and speculation. Drawing parallels to past panics over player pianos and VTRs, the EFF argues that AI tools are general purpose and capable of non-infringing uses, and that distorting copyright law would harm creativity and the public.

Sep 1·5 min read