
稀宇科技 today announced the launch of the MiniMax H3 multimodal generative model. It supports unified understanding of multimodal contexts consisting of text, images, video, and audio, and can output audiovisual content with native stereo sound. It supports resolutions of up to 15s and 2K.

According to feedback from preliminary invitation-only testing, MiniMax H3 offers commercial-grade content generation capabilities across multiple scenarios, performing well in instruction following, the presentation of text and brand information, and V2V Motion Transfer (video-to-video motion transfer). It is claimed to enable precise and controllable multimodal content editing and generation for applications such as advertising, e-commerce, branding, product design, UI / UX, and gaming.
In terms of core technology, MiniMax H3 adds Caption capabilities and uses a customized dedicated model and a unified multimodal understanding pipeline. Most source materials require 100K Tokens of inference, ultimately producing an average of approximately 4K Tokens.
Meanwhile, MiniMax H3's 2K video output does not use a conventional super-resolution module. Instead, it uses the H3 base model to regenerate its own low-resolution results in context, maximizing the reuse of the generative capabilities already possessed by the H3 base model. It also recovers information that traditional super-resolution solutions cannot restore through “guessing.”
In addition, MiniMax H3 will make its model weights available in the coming days. IT之家 will continue to follow the developments and bring you the latest news as soon as possible.
