MiniMax releases the full-modal generation model H3 – one model that handles images, videos, and audio all in one.
2026-07-31 10:27:00
According to CoinMeta, MiniMax has released a full-modal generation model called H3, which is capable of understanding text, images, videos, and audio simultaneously, and can generate or edit videos based on natural language instructions. Users can instruct H3 to learn camera movements from a video, use characters from another image, and refer to the sound from a third audio file. H3 can generate videos up to 15 seconds in length, supporting 2K resolution and native stereo sound. The official price for 2K videos with API is 0.13 dollars per second, while generating a 15-second video costs about 1.95 dollars; 768p videos cost 0.09 dollars per second. A single task allows for up to 9 images, 3 videos, and 3 audio files, with a total of no more than 12 files. MiniMax plans to make the model's weights available in the coming days.
Bullish 0
Bearish 0
Source:Internet
This content is for market information only and does not constitute investment advice.