MiniMax H3 and the Case for Unified Multimodal Context
The phrase sounds abstract until you sit down with an actual brief. Then it turns into something concrete: how many different things can you hand the model at once, and can it keep track of which one is responsible for what.
A typical request contains four or five separate requirements that have nothing to do with each other. The character must look like this person. She must move the way this other clip moves. The camera must behave like that reference. Her voice must sound like this recording. The whole thing must cut like the campaign you delivered last quarter.
Writing all of that in a prompt does not work, and everyone who has tried knows why. Below are the situations where handing over the material instead changes the outcome.
Building a cast that survives eighty episodes
Serialised work is the hardest test, because errors accumulate. A lead whose face shifts slightly every episode is not a visual problem by episode three, it is a continuity problem by episode thirty.
The approach that holds is assembling a reference package before production starts rather than describing anything.
For each principal character: three or four stills covering front, three-quarter and profile, under neutral light. A short clip showing how that person carries themselves, walks, gestures. One audio sample fixing the voice and speaking rhythm. Keep the package fixed and attach it to every request for the run of the series.
Minimax H3 will take written notes, photographs, footage and recordings together in one submission, twelve files being the ceiling, and works across the whole set instead of processing each type on its own. That capacity is what makes the package viable — you are not choosing between supplying a face or supplying a voice.
Extend the same discipline to recurring locations and to plot-relevant props. For animated or comic-style series, add frames that fix line weight, palette and shading, since style drift damages those formats the way face drift damages live-action ones.
Separating who someone is from how they move
This is the scenario that unified input handles and single-input tools cannot.
A brief calls for a specific character performing a specific action with a specific camera treatment. Conventionally these are three separate battles: get the character right, then get the movement right without losing the character, then get the camera right without losing either.
With multiple references, they stop competing. Stills carry identity. A performance clip carries how the body moves, how the timing sits, how the gesture lands. A second clip can carry the camera language independently — the speed of a push-in, when a move settles.
Practically this means you can reference a piece of choreography, a fight beat, a product handling gesture, or a particular walk-and-turn, and apply it to your own character rather than approximating the description in words.
Getting the product exactly right
Commercial work has a requirement most creative work does not: the object on screen has to match the object that ships. Correct colourway, current packaging, the label the customer will actually receive.
Generative models invent by default. Left to a text description, you get a plausible bottle rather than your bottle, and the difference is visible to the one person guaranteed to look closely.
The working method is a division of responsibility. Photograph the product properly once, then supply those stills as the source of truth. Everything else — the setting, the light, the atmosphere, the movement — gets generated around it. That division is also what makes the result defensible in a client review, because you can point to what came from reference and what did not.
The same logic applies to a brand’s typography and lockup, which have documented specifications and cannot be approximated.
Carrying a voice across languages and lines
Audio references do more than set a tone.
Because voice characteristics come from a sample and spoken content can be rewritten inside footage that already exists, a delivered piece can produce alternate lines without the artist, the room or the camera changing. Script notes stop meaning reshoots.
The same mechanism handles markets. Spoken output covers eleven languages, and a territory edition becomes a revision of approved footage rather than a dubbing booking with its own mixing schedule. For a title with dozens of episodes, or a campaign running across regions, that difference determines which markets are worth entering at all.
Teams running this continuously rather than occasionally move it into the Minimax H3 API so versions generate against an approved master inside an existing delivery process.
Matching work you have already delivered
An underused application: referencing your own output.
Campaigns run for months and accumulate house style — a cutting rhythm, a colour feel, a way the camera moves. Briefing a new batch to match usually means writing a style document that nobody reads accurately.
Supplying a finished piece as reference transfers that character directly. New assets inherit the pacing and grade of approved work, which matters for sequels, seasonal refreshes, and any situation where new material has to sit beside old material without looking newer.
Fixing what the references missed
No package covers everything, and the second half of unified input is the ability to correct.
H3 modifies people, objects, environments, sound and timing inside material already produced, with the rest of the frame staying put. In serialised production this handles what surfaces constantly: a prop that appears an episode too early, an outfit detail that stopped matching partway through, a background object that contradicts an earlier scene, a line rewritten after a script note.
The model holds the top position for video editing on Artificial Analysis, above Seedance 2.0, and that ranking reflects how much of the untouched frame survives an edit rather than how good a fresh generation looks.
Keep interventions contained. Swapping a prop or adjusting a delivery behaves well. Replacing a lead or exchanging an entire environment moves so much of the image that regenerating from the reference package is faster.
What the approach does not cover
Sequence and pace are decided by a person. In serialised and advertising work especially, where a beat lands is the product.
Any one generation covers four to fifteen seconds and tops out at 2K, so a completed piece gets built from a stack of them. Continuity across shots is managed rather than automatic — reference packages get you most of the way and someone still watches for what slipped.
And no reference produces a photograph of a physical object as it exists. Categories selling on material truth still need a camera.
Why it pays off
Reference packages take time to assemble, which is the main objection. The return arrives on the second and third use, not the first.
What it costs per second undercuts comparable systems, Seedance 2.0 among them; the Minimax h3 pricing page carries the current figures. Combined with reusable references, that changes what a production run looks like: the expensive part happens once, at the start, and every subsequent episode, variant or market version draws on it rather than rebuilding from a written description.