---
type: "post"
slug: "minimax-announces-minimax-h3"
title: "MiniMax H3 unifies multimodal understanding and generation"
description: "MiniMax H3 handles text, images, video, and audio in one model and generates video with stereo sound at up to 2K."
retrieval_nugget: "MiniMax H3 handles text, images, video, and audio in one model and generates video with stereo sound at up to 2K."
published_at: "2026-08-16"
updated_at: "2026-08-16"
record_date: "2026-08-16"
date_kind: "scheduled_at"
topics: ["multimodal","video-generation","open-models","minimax"]
entities: ["minimax.io"]
editorial_format: "field_note"
basket_id: "64af3bcb-1c2d-42a9-a664-91510a61d75a"
basket_revision: 1
source_urls: ["https://www.minimax.io/blog/minimax-h3"]
visual_decision: "text_only"
recovery_incident: "NR-2026-08-15-HERMES-SITE-COPY"
schema_version: "newruntime-agent-readable-v0.2"
stable_id: "post:minimax-announces-minimax-h3"
status: "published"
visuals: []
editorial_provenance: {"schema_version":"newruntime-editorial-copy-v1","content_status":"source_grounded_final","final_copy_sha256":"sha256:f59886f95b07c9d0f144f6b19f4102de9c885c4f26a98957611a2ae244b6f7b0","reviewed_at":"2026-08-15T20:30:00.000Z","source_evidence_count":1,"verified_claim_count":2}
routes: {"html":"https://newruntime.com/posts/minimax-announces-minimax-h3/","markdown":"https://newruntime.com/posts/minimax-announces-minimax-h3.md","json":"https://newruntime.com/posts/minimax-announces-minimax-h3.json"}
---

# MiniMax H3 unifies multimodal understanding and generation

## Retrieval answer

MiniMax H3 handles text, images, video, and audio in one model and generates video with stereo sound at up to 2K.

MiniMax introduced H3 as one model for understanding and generating text, images, video, and audio. Its video output includes native stereo sound, lasts up to 15 seconds, and is offered at up to 2K resolution. The company says it plans to release weights subject to applicable laws and regulations.

The architecture uses a new H3-VAE tokenizer that MiniMax says increases effective sequence length fourfold. A training system that separates understanding and generation workloads reportedly improved end-to-end throughput by nearly 30 percent. For 2K output, the base model regenerates its lower-resolution result in context instead of relying on a separate super-resolution model.

Those are vendor claims and need independent evaluation, especially for temporal consistency, audio alignment, and real pricing. The notable shift is the shared context boundary: generation modes that were separate services are being pulled into one instruction-following runtime, which may simplify orchestration while making evaluation more complex.

## Source

- [MiniMax](https://www.minimax.io/blog/minimax-h3)
