ModelsMedia Generation 🇺🇸 27.07.2026 15:06

Mistral AI Unveils Voxtral TTS — New Multilingual Speech Synthesis Model

MistralMistral ElevenLabsElevenLabs
Mistral AI has released Voxtral TTS, a speech synthesis model with 4 billion parameters that supports 9 languages, emotionally expressive speech, low latency, and easy voice adaptation. The model is available via API and Mistral Studio, starting at $0.016 per 1,000 characters.
Mistral AI has launched Voxtral TTS, its first speech synthesis model with 4 billion parameters. It delivers realistic, emotionally colored speech in 9 languages, accounts for dialects, has low latency (70 ms for a typical sample), and is easy to adapt to new voices. The model surpasses ElevenLabs Flash v2.5 in naturalness with comparable latency and matches ElevenLabs v3 in quality. Voxtral TTS supports zero-shot voice adaptation both within a single language and cross-lingually (for example, a French accent in English speech). The architecture includes a transformer decoder with 3.4 billion parameters, an acoustic transformer (390 million), and a proprietary neural audio codec (300 million). Fine-tuning to a specific voice requires a sample of at least 3 seconds. The model is available via API (from $0.016 per 1,000 characters) and in Mistral Studio, as well as open weights on Hugging Face under the Creative Commons Attribution Non Commercial 4.0 license.
Source: Mistral AI — original
Our earlier posts on this topic ↓
Fresh news