Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
MiniMax Releases H3, an Open-Weights Multimodal Video Model With Native Audio · News · Kaino
MiniMax Releases H3, an Open-Weights Multimodal Video Model With Native Audio
Kaino
8h agoAug 3, 2026, 12:00 AM0 views

MiniMax Releases H3, an Open-Weights Multimodal Video Model With Native Audio

MiniMax has introduced H3, an open-weights model for generating video from text, images, video and audio inputs. The company says the model can produce up to 15-second 2K videos with native stereo audio, while fal and ComfyUI have announced deployment and workflow support.

agentsminimax

MiniMax introduces H3

MiniMax has released H3, an open-weights general-purpose multimodal video generation model designed to work with text, images, video and audio as inputs.

According to MiniMax’s announcement, H3 can generate video clips up to 15 seconds long at 2K resolution and produces native stereo audio alongside the video output. The company describes the model as a unified system intended to handle multiple generation tasks and input modalities rather than separating them into distinct tools.

Video generation with audio

The inclusion of native stereo audio is a notable part of H3’s positioning. Video-generation systems often require users to create or add sound through a separate process; MiniMax says H3 generates audiovisual output directly.

MiniMax also says the model accepts combined context from text, images, video and audio. That capability could support workflows such as animating an input image, using reference material to guide a generated scene, or combining visual direction with audio context. The company’s materials do not establish how the model performs across every type of prompt or reference input, so real-world results will depend on the task, hardware and workflow used.

Availability in ComfyUI and fal

ComfyUI said it added day-zero support for MiniMax H3. Its announced workflows cover text-to-video, image-to-video and reference-to-video generation, as well as native stereo-audio output. ComfyUI is a node-based interface widely used for local and customizable generative-media workflows.

Fal also lists MiniMax H3 as an open-weights multimodal video model. Its model page describes support for text, image, video and audio inputs, along with video generation up to 15 seconds in 2K with native stereo sound.

The open-weights designation means developers and researchers can access the model weights under the terms set by MiniMax, rather than relying solely on a closed hosted product. However, open weights do not automatically determine the hardware requirements, licensing obligations or commercial-use conditions for every deployment. Users should consult MiniMax’s published documentation and license materials before adopting the model.

What remains to be tested

The release puts MiniMax H3 into a fast-moving market for AI video tools, where prompt adherence, motion consistency, editing control, generation speed and compute costs are key practical factors. MiniMax, fal and ComfyUI have outlined the model’s capabilities and integration options, but independent benchmarking will be needed to assess performance against other video models under comparable settings.

For now, H3’s combination of multimodal inputs, 2K output, short-form video generation and integrated stereo audio makes it a new option for developers building local or hosted AI video workflows.

Key takeaways
  • 1

    MiniMax introduces H3 MiniMax has released H3, an open weights general purpose multimodal video generation model designed to work with text, images, video and audio as inputs.

  • 2

    According to MiniMax’s announcement, H3 can generate video clips up to 15 seconds long at 2K resolution and produces native stereo audio alongside the video output.

  • 3

    The company describes the model as a unified system intended to handle multiple generation tasks and input modalities rather than separating them into distinct tools.

Continue reading

Latest from Kaino News

Story pulse

Freshness

8h ago

Views

0

Reading

2 min

Byline

Kainotomic Team

Utilities

Topics

agentsminimax

Sources

Reference material and original reporting used in this story.

fal

Published Aug 3, 2026, 12:00 AM

View source