⚡ BREAKING
Qwen2.5 Omni: Sees, Hears, Speaks, Writes – All in One
Alibaba/Qwen
Alibaba Qwen has released Qwen2.5-Omni, a flagship multimodal model capable of processing text, images, audio, and video, as well as generating real-time speech. The model uses the Thinker-Talker architecture and surpasses similar models across all modalities.
Alibaba Qwen has unveiled Qwen2.5-Omni, a new flagship model in the Qwen series designed for comprehensive multimodal perception. The model processes text, images, audio, and video, generating streaming responses in the form of text or natural speech. The Thinker-Talker architecture includes a Thinker that understands input and a Talker that generates speech. The model is available on Hugging Face, ModelScope, DashScope, and GitHub. Qwen2.5-Omni outperforms Qwen2.5-VL-7B and Qwen2-Audio on audio tasks, is comparable on video and images, and achieves state-of-the-art results on OmniBench. The model also demonstrates high accuracy in speech recognition, translation, audio understanding, and speech generation. Future plans include improving adherence to voice commands and integrating more modalities.
Source: Alibaba Qwen —
original
