MiniMax releases the full-modal generation model H3 – one model that handles images, videos, and audio all in one.
2026-07-31 10:27:00
According to CoinMeta, MiniMax has released a full-modal generation model called H3, which is capable of understanding text, images, videos, and audio simultaneously, and can generate or edit videos based on natural language instructions. Users can instruct H3 to learn camera movements from a video, use characters from another image, and refer to the sound from a third audio file. H3 can generate videos up to 15 seconds in length, supporting 2K resolution and native stereo sound. The official price for 2K videos with API is 0.13 dollars per second, while generating a 15-second video costs about 1.95 dollars; 768p videos cost 0.09 dollars per second. A single task allows for up to 9 images, 3 videos, and 3 audio files, with a total of no more than 12 files. MiniMax plans to make the model's weights available in the coming days.
Source:Internet
This content is for market information only and does not constitute investment advice.
Follow HQYC official accounts to stay updated
Hot Articles
Refresh

Ethereum: Quantum sells another 1,000 ETH to invest in AI data centers
30m ago

The decline in AI concept stocks has triggered a deleveraging trend.
1h ago

nft: Why has Bitfinex's ecosystem token LEO remained low-key for so long?
1h ago

Anthropic disclosed that Claude had unauthorized access to three companies' systems during testing.
3h ago

Web3: Foreign media: XRP under pressure, ZEC attempts to stabilize, HYPE hype cools down.
4h ago

