Models RSS

Models 🇨🇳

Qwen2.5 Omni: Sees, Hears, Speaks, Writes – All in One

Alibaba Qwen has released Qwen2.5-Omni, a flagship multimodal model capable of processing text, images, audio, and video, as well as generating real-time speech. The model uses the Thinker-Talker architecture and surpasses similar models across all modalities.

Alibaba/QwenAlibaba/Qwen
Alibaba Qwen28.07 · 00:03
Models 🇨🇳

Alibaba Qwen Releases QVQ-Max: Visual Reasoning Model with Proofs

Alibaba Qwen officially releases QVQ-Max, its first version of a visual reasoning model. The model is capable of not only recognizing images and videos but also analyzing them, conducting reasoning, and offering solutions. In the MathVision benchmark, QVQ-Max demonstrates consistent accuracy improvement as the length of the thought process increases.

Alibaba/QwenAlibaba/Qwen
Alibaba Qwen28.07 · 00:03
Fresh news