MiniMax H3 becomes first open video model to top video editing ranking
MiniMax
ByteDance
MiniMax has released the weights of its video model H3, making it the first open model to reach the top of a video ranking. According to Artificial Analysis, H3 ranks first in video editing, second in text-to-video, and third in image-to-video. The 33-billion-parameter model processes text, images, videos, and audio jointly, generating clips with stereo sound, but some components remain closed.
MiniMax has made the weights of its video model H3 publicly available, marking the first time an open model has topped a video ranking. On Artificial Analysis, H3 ranks first in video editing, second in text-to-video, and third in image-to-video. The 33-billion-parameter model processes text, images, videos, and audio jointly, producing a clip of four to 15 seconds with stereo sound, and according to the model card, it can handle up to nine reference images, three video clips, and three audio clips per prompt. However, two components remain closed: the module for 2K resolution and H3-Context-IR, which converts prompts and reference material into a structured intermediate form. Locally in ComfyUI, this results in 768p output, but the context processing can be replicated using the published prompting guides. The open weights also allow fine-tuning, for example on one's own footage, characters, or a consistent look, though the license permits commercial use only for companies with revenue under 20 million dollars. Simultaneously, ByteDance released the closed Seedance 2.5, which delivers 30-second clips with audio.
Source: The Decoder (DE) —
original
