News · Science · 9 July 2026
Artificial Analysis launched a new Controlled Voice Arena Leaderboard for comparing text-to-speech models.

The leaderboard uses the same set of 8 cloned voices across all models: 2 US male, 2 US female, 2 UK male, and 2 UK female voices. Each model is tested on the same 1–2 minute voice samples.
This makes the comparison more controlled because it separates voice preference from overall model quality.
The top overall model is Cartesia Sonic 3.5 with 1,122 Elo, followed by ElevenLabs Eleven v3 at 1,088 and Inworld Realtime TTS-2 Research Preview at 1,070.
Cartesia Sonic 3.5 also leads both the US accent and UK accent categories.
Among open-weight models, Fish Audio S2 Pro leads with 1,034 Elo, followed by Mistral Voxtral TTS at 1,024 and Resemble AI Chatterbox at 930.
First posted to our Telegram.