⚡ BREAKING
ModelsMedia Generation 🇺🇸 27.07.2026 14:04

Google DeepMind Introduces Gemini Omni: A Multimodal Model for Video Creation and Editing

Google/DeepMindGoogle/DeepMind DeepMindDeepMind
Google DeepMind has announced Gemini Omni, a new model capable of generating video based on any combination of inputs: images, audio, video, and text. The first version, Gemini Omni Flash, is already available in the Gemini app, Google Flow, and YouTube Shorts, and will soon be available for developers via API.
Google DeepMind has introduced Gemini Omni, a new multimodal model that combines reasoning and creative capabilities. The model accepts images, audio, video, and text as input, and produces high-quality videos based on the real-world knowledge of Gemini. The first version, Gemini Omni Flash, is now available in the Gemini app, Google Flow, and YouTube Shorts (free for YouTube Shorts and YouTube Create App users). It is also available globally to Google AI Plus, Pro, and Ultra subscribers. In the coming weeks, the model will be accessible to developers and enterprise customers via the API. Key features include: the ability to edit videos through natural language dialogue (maintaining consistency in characters, physics, and scenes); video generation with improved physics (gravity, kinetic energy, hydrodynamics); leveraging Gemini's knowledge of history, science, and cultural context for realistic storytelling; support for any combination of input modalities (image, text, video, audio; audio currently limited to voice references). For safety, an invisible digital watermark called SynthID has been implemented, along with content verification tools through the Gemini app, Chrome, and Google Search. In the future, the model will be able to generate images and audio.
Source: Google DeepMind — original
Our earlier posts on this topic ↓
Fresh news