Skip to main content
Models & Technology

MiniMax H3 General-Purpose Multimodal Video Model to Be Open-Sourced on August 3, Supporting up to 15s at 2K Resolution

MiniMax H3 supports unified understanding of multimodal contexts composed of text, images, video, and sound, and can output audio and video with native stereo sound, supporting up to 15s at 2K resolution.

MiniMax H3 General-Purpose Multimodal Video Model to Be Open-Sourced on August 3, Supporting up to 15s at 2K Resolution

MiniMax announced today the launch of the MiniMax H3 general-purpose multimodal video model. ModelScope shows that the model will be officially open-sourced at 0:00 Beijing time on August 3.

MiniMax H3 General-Purpose Multimodal Video Model to Be Open-Sourced on August 3, Supporting up to 15s at 2K Resolution

MiniMax H3 supports unified understanding of multimodal contexts composed of text, images, video, and sound, and can output audio and video with native stereo sound, supporting up to 15s at 2K resolution.

Benefiting from technologies including Contextual Omni Representation, H3-VAE, H3-Omni Transformer, and In-context Regeneration, MiniMax says it can offer the best cost-performance ratio in the industry. It provides a default resolution of 2K, with a per-second price at 2K resolution of less than one-third that of mainstream models, while its price at 768P resolution is half that of mainstream models at 720P.

According to the official announcement, the model has the following highlights:

Native multimodal understanding and generation: Supports various inputs, including text, images, audio, and video; understands the people, actions, sounds, emotions, shots, styles, and expressive intent in different materials; and naturally integrates multiple reference sources.

Precise multimodal editing and control: Supports multidimensional editing of people, objects, scenes, sounds, and rhythm, while offering fine-grained instruction following.

Commercial-grade content generation for multiple scenarios: Designed for real-world commercial content scenarios such as film and television, advertising, branding, e-commerce, and games, covering text subtitles, brand information, creative effects, product showcases, UI / UX dynamic demonstrations, game visuals, and stylized expression.