MiniMax has introduced H3, an open-weights model for generating video from text, images, video and audio inputs. The company says the model can produce up to 15-second 2K videos with native stereo audio, while fal and ComfyUI have announced deployment and workflow support.
MiniMax has released H3, an open-weights general-purpose multimodal video generation model designed to work with text, images, video and audio as inputs.
According to MiniMax’s announcement, H3 can generate video clips up to 15 seconds long at 2K resolution and produces native stereo audio alongside the video output. The company describes the model as a unified system intended to handle multiple generation tasks and input modalities rather than separating them into distinct tools.
The inclusion of native stereo audio is a notable part of H3’s positioning. Video-generation systems often require users to create or add sound through a separate process; MiniMax says H3 generates audiovisual output directly.
MiniMax also says the model accepts combined context from text, images, video and audio. That capability could support workflows such as animating an input image, using reference material to guide a generated scene, or combining visual direction with audio context. The company’s materials do not establish how the model performs across every type of prompt or reference input, so real-world results will depend on the task, hardware and workflow used.
ComfyUI said it added day-zero support for MiniMax H3. Its announced workflows cover text-to-video, image-to-video and reference-to-video generation, as well as native stereo-audio output. ComfyUI is a node-based interface widely used for local and customizable generative-media workflows.
Fal also lists MiniMax H3 as an open-weights multimodal video model. Its model page describes support for text, image, video and audio inputs, along with video generation up to 15 seconds in 2K with native stereo sound.
The open-weights designation means developers and researchers can access the model weights under the terms set by MiniMax, rather than relying solely on a closed hosted product. However, open weights do not automatically determine the hardware requirements, licensing obligations or commercial-use conditions for every deployment. Users should consult MiniMax’s published documentation and license materials before adopting the model.
The release puts MiniMax H3 into a fast-moving market for AI video tools, where prompt adherence, motion consistency, editing control, generation speed and compute costs are key practical factors. MiniMax, fal and ComfyUI have outlined the model’s capabilities and integration options, but independent benchmarking will be needed to assess performance against other video models under comparable settings.
For now, H3’s combination of multimodal inputs, 2K output, short-form video generation and integrated stereo audio makes it a new option for developers building local or hosted AI video workflows.
MiniMax introduces H3 MiniMax has released H3, an open weights general purpose multimodal video generation model designed to work with text, images, video and audio as inputs.
According to MiniMax’s announcement, H3 can generate video clips up to 15 seconds long at 2K resolution and produces native stereo audio alongside the video output.
The company describes the model as a unified system intended to handle multiple generation tasks and input modalities rather than separating them into distinct tools.
Continue reading