Digest · past 24h
- Qwen2.5 Omni: Sees, Hears, Speaks, Writes – All in One
Alibaba Qwen has released Qwen2.5-Omni, a flagship multimodal model capable of processing text, images, audio, and video, as well as generating real-time speech. The model uses the Thinker-Talker architecture and surpasses similar models across all modalities.
Source: dnb66 · Digest —
original
