ModelsOpen Source 🇺🇸 27.07.2026 16:05

Mistral AI Releases Voxtral Transcribe 2 — Ultra-Low Latency Speech Recognition Models

MistralMistral Mistral AIMistral AI
Mistral AI has introduced the Voxtral Transcribe 2 family of models, including Voxtral Mini Transcribe V2 for batch transcription and Voxtral Realtime for live applications. Voxtral Realtime is distributed with open weights under the Apache 2.0 license. The models support 13 languages, diarization, and contextual control, while Voxtral Mini Transcribe V2 offers low cost ($0.003/min) and high accuracy.
Mistral AI has announced the release of Voxtral Transcribe 2, two new speech recognition models focused on high-quality transcription, diarization, and minimal latency. The family includes Voxtral Mini Transcribe V2 for batch processing and Voxtral Realtime for live applications. Voxtral Realtime uses a new streaming architecture, processing audio as it arrives, and allows latency to be tuned down to sub-200 ms. The model has 4 billion parameters, supports 13 languages (including English, Chinese, Hindi, Spanish, Arabic, French, Portuguese, Russian, German, Japanese, Korean, Italian, and Dutch), and is distributed with open weights under the Apache 2.0 license on the Hugging Face Hub. Voxtral Mini Transcribe V2 achieves a word error rate of around 4% on FLEURS and costs $0.003 per minute, which the company claims is the best price-performance ratio among transcription APIs. The model outperforms GPT-4o mini Transcribe, Gemini 2.5 Flash, Assembly Universal, and Deepgram Nova in accuracy, processes audio approximately three times faster than ElevenLabs Scribe v2 at the same accuracy, and is five times cheaper. Voxtral Mini Transcribe V2 supports diarization (generating speaker labels with precise start and end times), contextual guidance (up to 100 words for spelling correction), per-second timestamps, noise robustness, processing of audio up to three hours, and the same 13 languages. Also launched is an audio playground in Mistral Studio that allows testing transcription with diarization and timestamps. Voxtral Mini Transcribe V2 is available via API at $0.003 per minute, Voxtral Realtime at $0.006 per minute, and in the form of open weights. Both models support deployment in compliance with GDPR and HIPAA through on-premises or private cloud installations.
Source: Mistral AI — original
Our earlier posts on this topic ↓
Fresh news