ModelsMedia Generation 🇨🇳 31.07.2026 13:01

Video Post-Production at Risk: MiniMax H3 Turns Hand-drawn into Effects, the Multimodal 'Coding Moment' Arrives

MiniMaxMiniMax
MiniMax releases its first open-source video model, H3, which integrates editing, typography, transitions, pacing, background music, and visual effects end-to-end. Users can input text and get a ready-to-publish video, with 2K resolution by default. The model is praised as a new SOTA in video editing and supports full-modal input, including text, images, audio, and video.
MiniMax has officially released its next-generation video model, MiniMax H3, which is also its first open-source video model. H3 breaks the previous paradigm where video models only generated raw footage, by integrating editing logic, typography, transition design, rhythm control, background music, and visual effects end-to-end. Users can input text and receive a finished video ready for publication, with default 2K resolution. The model has been well-received by overseas developers and topped the Artificial Analysis leaderboard for video editing. The author tested H3 and found it highly responsive to user intent, even following complex multi-shot prompts accurately. H3 also demonstrates significant improvement in generating text in videos, which has been a common failure point for AI video models. It can generate text that remains consistent across frames and even understands the relationship between text and visual content. Additionally, H3 can take an audio clip as input for dialogue, with voice synced to the emotion and rhythm of the visuals, building on MiniMax's expertise in multimodal voice models. The pricing is low, with the cost per second at 2K resolution being less than one-third of mainstream models. H3 supports full-modal input, meaning all types of media are part of the model's context. It uses a custom caption model and an H3-Omni Transformer architecture for efficient multimodal processing. MiniMax's approach is based on the emergence of multimodal plus language, connecting isolated modalities through language. The open-source nature of H3 allows enterprises to privately deploy and fine-tune the model for custom content workflows, potentially leading to a burst of IP vitality. MiniMax has a history of open-sourcing models, from the M2 series to now H3, aligning with its vision of 'Intelligence with Everyone'. The release marks a shift of video generation into a 'Coding moment', where it becomes a commercially viable production tool.
Сокращения
SOTA = State of the Art — самый современный
2K = 2K resolution (2048×1080) — разрешение 2K
Source: QbitAI 量子位 — original
Our earlier posts on this topic ↓
Fresh news