Media GenerationModels 🇷🇺 06.08.2026 12:01

MiniMax H3: The Neural Network That Directs, Does Sound, and Edits Itself

MiniMax has released H3 (also known as Hailuo 3.0), a multimodal AI model that generates video with synchronized audio, performs edits following reference videos, and even handles green screen removal. The model is open-sourced under a license, with weights available on Hugging Face, and supports up to 2K resolution via API. It aims to unify text, image, video, and audio processing into a single context.
MiniMax has introduced H3, also branded as Hailuo 3.0 or Hailuo 03, which is a multimodal model capable of generating 4-15 second videos at 24 fps with native stereo audio at 32 kHz, and up to 2K resolution via API or 768p when self-hosted. The model can take up to 9 reference images, 3 videos, and 3 audio files simultaneously, and can edit videos to match the timing and style of a reference clip. It also removes green screens and adjusts lighting. The architecture compresses sources into language descriptions (reducing ~100k tokens to 4k), uses a redesigned VAE tokenizer with a fourfold increase in effective sequence length, and includes separate computation for understanding and generation. MiniMax released the model's weights on Hugging Face on August 2-3, 2026, several days after the API launch on July 31, 2026. Demonstrations show the model handling complex prompts for editing, advertising, and seamless transitions, but it still struggles with small text on video, a common weakness in generative video models.
Source: Habr — хаб ИИ — original
Our earlier posts on this topic ↓
Fresh news