ModelsOpen Source 🇺🇸 24.07.2026 00:05

Best Open ASR Models of 2026: WER, Languages, Latency and Licenses Compared

CohereCohere IBMIBM NVIDIANVIDIA Alibaba/QwenAlibaba/Qwen MistralMistral KyutaiKyutai MetaMeta OpenAIOpenAI
The open speech recognition (ASR) market is no longer a monoculture of Whisper: over the past 12 months, many competitive models have emerged. The leaders of the Open ASR Leaderboard are separated by less than one percentage point in WER, so the decisive factors have become license, language coverage, streaming support, and cost per hour of audio. The article compares the main models in detail across these parameters and provides recommendations for selection.
Over the past 12 months, the open speech recognition market has moved beyond being a Whisper monoculture. In March 2026, Cohere released Transcribe (2B parameters, Apache 2.0), which topped the Hugging Face Open ASR Leaderboard with an average Word Error Rate (WER) of 5.42%. Five weeks later, IBM released Granite Speech 4.1 2B with a score of 5.33%, followed by ARK-ASR-3B and MOSS-Transcribe-preview-2B, which achieved even lower values. The top models on the leaderboard are now separated by less than one percentage point in WER, making the ranking not the sole criterion for selection. Important factors include license, language coverage, streaming support, and cost per hour of audio. The article notes that the average WER on the leaderboard is not a fixed value: models were tested on different datasets. For example, Cohere Transcribe scored 5.42% on eight English test sets, including TED-LIUM, while ARK-ASR-3B scored 5.04% on seven sets without TED-LIUM, inflating its average. If Cohere is recalculated on those same seven sets, its WER would be 5.84%, and Granite Speech 4.1 2B would be 5.65%. Some models, like MOSS-Transcribe-preview-2B, were fine-tuned with reinforcement learning on training sections of the leaderboard, inflating their results. Private test sets from Appen change the ranking order: on them, zoom/scribe_v1 takes first place. Notable models include: Cohere Transcribe — 2B, Apache 2.0, 14 languages, downloaded over 620,000 times in a month; Granite Speech 4.1 2B — Apache 2.0, 6 languages, ASR + bidirectional speech translation, Real Time Factor (RTFx) 231.29; Canary-Qwen-2.5B — 2.5B, Creative Commons Attribution 4.0 International (CC-BY-4.0), English, WER 5.63%, RTFx 418; Qwen3-ASR-1.7B — Apache 2.0, 52 languages and dialects, WER 5.76%. In the throughput category, leaders include: Parakeet TDT 0.6B v3 (0.6B, CC-BY-4.0, RTFx 3332.74, 25 European languages, WER 6.32%), Granite Speech 4.1 2B-NAR (non-autoregressive, RTFx ~1820), and Qwen3-ASR-0.6B (52 languages, RTFx ~2000). For streaming: Voxtral Mini 4B Realtime 2602 (Apache 2.0, 13 languages, latency from 80 ms, runs on 16 GB GPU) and Kyutai STT (CC-BY-4.0, English/French, latency 0.5 s, with built-in semantic Voice Activity Detection (VAD)). Meta Omnilingual ASR (Apache 2.0) natively supports 1600+ languages and up to 5400+ via zero-shot learning. Whisper large-v3 (Massachusetts Institute of Technology (MIT) license, 99 languages) remains the best choice for projects where minimal license burden is critical. Among research novelties: diffusion-gemma-asr-small (parallel diffusion denoising generation, 42M trained parameters) and MOSS-Transcribe-Diarize 0.9B (Apache 2.0, 50+ languages, performs recognition and diarization in one pass). License division: Apache 2.0 (Cohere Transcribe, Granite Speech, Qwen3-ASR, Voxtral Realtime, Omnilingual ASR, ARK-ASR, MOSS-Transcribe), MIT (Whisper large-v3), and CC-BY-4.0 (Canary-Qwen, Parakeet, Kyutai STT). Meta splits licenses: Apache 2.0 for models, CC-BY for corpus. It is recommended to choose a model in the following order: license, language coverage, streaming or batch mode, then measure WER on your own audio data and calculate cost per hour of audio on your hardware. The main takeaway: an open model with 2B parameters under a permissive license now surpasses what closed APIs offered 18 months ago, and the choice boils down to a procurement decision rather than a research one.
Source: MarkTechPost — original
Our earlier posts on this topic ↓
Fresh news