Alibaba's New Voice Model Takes Top Spot

The new Qwen Audio 3.0 TTS Plus model from Alibaba is now the top performer on the Speech Arena leaderboard. It works in 16 different languages and gives users precise control over the tone of the voice. You can change how it sounds by using simple tags like angry or by describing the style in plain English.
While the quality is high, the model is not very fast. It generates speech at a rate of 16 characters per second. This is much slower than other popular options like Sonic 3.5 and Simba 3.2.
This technology represents a big step forward for Alibaba in the voice space. Users will have to decide if the expressive output is worth the wait compared to faster competitors.
Comments (0)
No comments yet. Be the first!
More AI news
NewsNew AI Tool Fixes Mistakes in Financial Research
Researchers developed a new framework to stop AI from learning from its own bad data in financial trading models.
NewsPerplexity adds local AI to Mac apps
The new update lets your Mac handle some AI tasks directly on your computer.
NewsMeta Releases Muse Image Model on Fal
Meta has launched a new AI model on the Fal platform that plans and edits images on its own.