Minimax H3 Prompt Guide — How to Write Prompts That Actually Work (2026)
The complete Minimax H3 prompt engineering guide. Learn timestamp format, camera moves, audio syntax, shot budgets, and see real before/after examples for Hailuo AI.
Table of Contents
Why Most H3 Prompts Produce Garbage
Minimax H3 (the model behind Hailuo AI) is one of the strongest text-to-video models available in 2026. It generates video and audio simultaneously, handles multi-shot sequences, and supports camera direction.
But most people write prompts like this:
a beautiful sunset over the ocean with waves
And get generic, flat, 6-second clips that look like stock footage from 2019.
The problem isn't the model. It's the prompt. H3 is a precision tool — it responds to structure, specificity, and filmmaking vocabulary. Give it vague input, get vague output.
This guide covers everything that actually matters for H3 prompts: the timestamp system, camera language, shot budgets, audio direction, and the specific patterns that consistently produce better results. Every technique is based on real testing across hundreds of generations.
Want to skip the reading and just get optimized prompts? Use our free Minimax H3 Prompt Generator — paste your idea, get a production-ready prompt.
The Core Formula: 6 Layers That Matter
Every good H3 prompt follows the same structure, whether it's one shot or five. Think of it as layers — each one you add gives the model more to work with.
The formula:
Subject + Action + Scene + Visual Style + Camera + Audio
Here's what each layer does:
1. Subject — Who or what is in the frame. Don't say "a person." Say "a woman in her 30s wearing a black leather jacket, silver rings on both hands." The more specific you are about appearance, the less the model guesses.
2. Action — What happens. Use active verbs with direction and speed. "Walks" is weak. "Strides confidently toward camera, coat swaying" is strong. Motion description is the single most impactful thing you can add.
3. Scene — Where it happens. Time of day, weather, environment details. "A street" is nothing. "A rain-soaked alley in Shinjuku at 2 AM, neon signs reflecting in puddles" gives the model a complete world.
4. Visual Style — The overall look. Pick ONE and commit. "Photorealistic, shot on 35mm film" or "Studio Ghibli anime style" or "cyberpunk neon aesthetic." Mixing styles (e.g. "photorealistic anime") confuses the model.
5. Camera — How the camera moves. Use actual filmmaking terms: "slow dolly in," "orbit shot," "crane up reveal." One camera move per shot. More than one = muddy results.
6. Audio — What you hear. H3 generates audio natively — this is its superpower. Describe sounds naturally: "soft piano melody plays," "rain patters on glass," "distant city traffic." Don't skip this layer.
The minimum viable prompt covers layers 1-3. But the best results come from hitting all six. The prompt generator handles this automatically.
Timestamps: The Shot Structure System
H3's timestamp system is what separates it from most video models. Instead of one continuous description, you can structure your video as multiple shots with precise timing.
Format:
MM:SS.mmm–MM:SS.mmm [Shot description]
Real example (6-second video, 2 shots):
00:00.000–00:03.000 Close-up of espresso pouring from a brass pot into a white ceramic cup. Steam curls upward in slow motion, backlit by warm morning light. Soft jazz piano.
00:03.000–00:06.000 Camera slowly pulls back to reveal a café table with a croissant and open book. Ambient café chatter, cups clinking. Warm color grading, photorealistic.
Rules that matter:
- Ranges must be sequential. No gaps, no overlaps.
- First shot starts at
00:00.000. Last shot ends at your total duration. - Each shot should be 2-5 seconds. Shorter than 1.5s doesn't give the model enough frames.
- More shots = more chances for visual inconsistency between cuts. The model handles continuity, but it's not perfect.
Shot budget by duration:
| Duration | Shots | Notes |
|---|---|---|
| 4-6s | 1-2 | Single continuous shot often works best |
| 7-10s | 2-3 | Sweet spot for storytelling |
| 11-15s | 3-5 | Plan carefully — more risk of inconsistency |
Common timestamp mistakes:
- Overlapping ranges (e.g.
00:02.000–00:04.000followed by00:03.000–00:06.000) - Shots under 1.5 seconds (model can't render meaningful motion)
- Too many shots for short videos (4 shots in a 5-second video = chaos)
- Forgetting to end at the exact total duration
Camera Moves That Actually Work
H3 understands filmmaking vocabulary. Use the right terms and you get cinematic camera work. Use vague terms and the camera sits still.
High-reliability moves (these work consistently):
slow dolly in— Camera moves forward toward subject. Great for building intimacy or tension.slow dolly out— Camera pulls back. Good for reveals.tracking shot— Camera follows subject laterally. Natural for walking/running scenes.orbit shot/orbit 360°— Camera circles the subject. Product shots, hero reveals.crane up/crane down— Vertical camera movement. Reveals landscapes, shows scale.static shot— No movement. Lets the subject's action be the focus.pan left/pan right— Camera pivots horizontally. Scanning environments.zoom in/zoom out— Lens zoom, not physical movement. More dramatic, less natural.
Advanced moves (work but less predictable):
crash zoom— Rapid zoom into detail. Impact moments.dolly zoom— Simultaneous dolly + zoom creating a disorienting effect (Vertigo effect).whip pan— Fast camera swing. Transition between subjects.steadicam follow— Smooth handheld tracking. Behind-the-subject following.pull focus rack— Shift focus from foreground to background or vice versa.bird's eye view— Directly overhead looking down.dutch angle— Tilted frame. Unease, tension.
The one-move rule: Use ONE camera move per shot. Slow dolly in with orbit and crane up will confuse the model. If you want multiple moves, split them across shots:
00:00.000–00:03.000 Slow dolly in toward the subject...
00:03.000–00:06.000 Camera orbits the subject 180 degrees...
Framing keywords (combine with moves):
extreme close-up, medium close-up, wide establishing shot, over-the-shoulder, POV shot, low angle, high angle
Visual Style: Pick One, Commit
Style keywords set the entire look of your video. The most important rule: pick one style and stick with it.
Photorealistic styles:
photorealistic— Default realistic lookshot on 35mm film— Film grain, organic color, slight imperfectionsanamorphic— Wide aspect feel, oval bokeh, lens flaresIMAX— Massive scale, ultra-sharp detailmacro photography— Extreme close-up of tiny detailseditorial fashion— High-end magazine look, deliberate compositiondocumentary— Observational, natural, unpolished
Stylized:
Studio Ghibli— Lush hand-painted anime, pastoral scenesPixar 3D— Smooth 3D animation, expressive charactersanime— Japanese animation stylecyberpunk— Neon lights, rain, dark urban futurismsteampunk— Victorian + industrial machineryvaporwave— Pastel gradients, retro 80s/90s aesthetic
Artistic:
watercolor— Soft edges, transparent color washesoil painting— Thick brushstrokes, rich colorclaymation— Stop-motion clay figuresblack and white— Monochromefilm noir— High contrast, dramatic shadows, 1940s moodretro VHS— Scan lines, color bleeding, analog distortion
What NOT to do:
- Don't combine incompatible styles: "photorealistic anime" or "watercolor hyperrealistic"
- Don't add style keywords everywhere — one mention is enough
- Don't say "high quality" or "4K" — these are meta-instructions, not styles. Describe what you SEE instead.
Dialogue & Sound Design
H3 generates audio natively. This is a major advantage over models that output silent video. You can direct both dialogue and ambient sound.
Dialogue syntax:
<d>[English] Your dialogue text here</d>
Place the <d> tag inline within your shot description:
00:00.000–00:04.000 A man in a lab coat stands in front of a whiteboard, gesturing with a marker. <d>[English] The key insight is that nobody reads the documentation.</d> Warm office lighting, medium shot, static camera.
Dialogue rules:
- Keep dialogue short — 1-2 sentences per shot maximum
- The language tag
[English]tells the model which language to synthesize - Dialogue quality is decent but not ElevenLabs-tier. For high-quality VO, generate silent video and add TTS separately
- You can have multiple characters speak across different shots
Sound design — describe it naturally:
Instead of technical audio tags, just write what you hear:
rain patters against the windowa deep bass note pulsesbirds chirping in the distance, wind rustling leavescrowd cheering, stadium echosoft lo-fi hip hop beat playsthunder rumbles, then silence
H3 picks up these audio cues and generates matching sound. The more specific your sound descriptions, the better the audio matches the visuals.
Music direction:
cinematic orchestral swell— Big dramatic momentssoft piano melody— Emotional, intimate scenesupbeat electronic beat— Energy, product revealsjazz quartet plays— Café, lounge, sophisticated moodno music, only ambient sounds— Documentary, ASMR, realism
Progressive Enhancement: Building Better Prompts Layer by Layer
Here's how the same prompt improves as you add dimensions. Each row adds one layer.
Starting idea: "a cat on a windowsill"
| Layer | Prompt | What it adds |
|---|---|---|
| Subject only | A cat on a windowsill | Nothing for the model to work with |
| + Action | A tabby cat stretches lazily on a windowsill, then yawns | Motion and behavior |
| + Scene | ...on a rain-streaked windowsill in a cozy apartment, grey afternoon light | Environment, weather, time |
| + Style | ...Photorealistic, shot on 35mm film with shallow depth of field | Visual look and texture |
| + Camera | ...Static close-up, rack focus from raindrops on glass to cat's face | Camera behavior |
| + Lighting | ...Soft diffused overcast light, warm lamp glow from behind camera | Mood and depth |
| + Audio | ...Rain patters against the window, distant thunder, cat purrs softly | Sound design |
The full prompt:
00:00.000–00:06.000 A tabby cat stretches lazily on a rain-streaked windowsill in a cozy apartment, then yawns wide showing tiny teeth. Grey afternoon light filters through the glass. Static close-up, rack focus from raindrops on glass to cat's face. Soft diffused overcast light with warm lamp glow from behind camera. Rain patters against the window, distant thunder, cat purrs softly. Photorealistic, shot on 35mm film with shallow depth of field.
This isn't about writing essays. It's about giving the model enough information to make decisions. Every vague word is a coin flip. Every specific word is a direction.
T2VA, I2VA, and Ref2VA — When to Use Each Mode
H3 supports multiple input modes. Each works differently and needs a different prompting approach.
T2VA (Text-to-Video+Audio) — The default. Pure text prompt generates both video and audio.
- Best for: Original scenes, creative ideas, anything you can describe
- Prompt approach: Describe everything — the model has no visual reference
- Tip: The more specific you are, the more control you have
I2VA (Image-to-Video+Audio) — Upload a still image, H3 animates it into video.
- Best for: Product shots, portraits, artwork you want to bring to life
- Prompt approach: DON'T re-describe the image — describe how the scene EVOLVES. Add motion, camera movement, and audio.
- Bad prompt: "A woman standing in a field wearing a blue dress" (just re-describes the image)
- Good prompt: "The woman walks forward, wind catching her hair. Camera slowly tracks her from the side. Grass rustles, birds sing in the distance."
- Tip: The uploaded image is frame 1. Your prompt describes what happens next.
Ref2VA (Reference-to-Video+Audio) — Upload a reference image for style or subject guidance, but generate a new scene.
- Best for: Style transfer, maintaining character consistency across videos, brand aesthetics
- Prompt approach: Describe a NEW scene that uses the reference's style/subject. The output won't be a direct animation of the reference.
- Example: Upload a neon-lit cyberpunk city image → prompt describes a character walking through a DIFFERENT cyberpunk street
- Tip: Good for series where you want visual consistency across multiple clips
Which mode for which use case:
| Use case | Mode | Why |
|---|---|---|
| Original creative video | T2VA | No image needed, full creative control |
| Animate a product photo | I2VA | Product IS the starting frame |
| Bring a portrait to life | I2VA | Face is the starting frame |
| Match a visual style | Ref2VA | Style guides the look, scene is new |
| Character consistency | Ref2VA | Character reference carries over |
| Storyboard to video | I2VA | Each storyboard frame = one shot |
What to Avoid — The Don'ts List
These are the patterns that consistently produce worse results. Avoid all of them.
Don't use transition language. H3 handles cuts between shots automatically. Writing "dissolve to..." or "fade to black..." or "smooth transition into..." wastes tokens and confuses the model. Just end one shot and start the next.
Don't request text overlays. AI video models can't reliably render text, logos, or watermarks. If you need title cards, add them in post.
Don't write slideshow prompts. Each shot needs continuous motion within it. "Show the product. Then show the ingredients. Then show the packaging." = a slideshow. Instead: "The product rotates slowly on a turntable, camera orbits to reveal the ingredients label, then continues to show the back packaging." = motion.
Don't mix incompatible styles. "Photorealistic anime" sends contradictory signals. Pick one visual style per video.
Don't write shots shorter than 1.5 seconds. The model needs time to render meaningful motion. A 1-second shot produces a blurry mess.
Don't specify FPS or frame counts. "Render at 24fps" or "exactly 72 frames" — the model doesn't work that way. Control timing with timestamp ranges.
Don't use meta-instructions. "Make it viral," "high quality," "cinematic masterpiece," "trending on social media" — these are instructions to a person, not a video model. The model only understands what to SHOW and what to HEAR. Describe the actual content.
Don't overlap timestamp ranges. Each range must start where the previous one ended. Overlapping timestamps create undefined behavior.
Don't forget audio. H3 generates audio natively. If you don't describe sound, you get random ambient noise. Take 10 seconds to add "soft piano plays in background" or "city traffic ambience" and the audio quality jumps dramatically.
5 Real Before & After Prompt Transformations
These are real transformations showing exactly what changes make a difference.
1. Product Showcase
Before:
show a sneaker rotating
After:
00:00.000–00:03.000 A matte white sneaker with neon green accents sits on a glossy black turntable, rotating slowly. Camera orbits at eye level. Hard studio spotlight from above creates dramatic shadows. Clean white background.
00:03.000–00:06.000 Crash zoom into the sole tread pattern, then rack focus to the knit texture of the upper. Subtle bass-heavy electronic pulse. Hyperrealistic product photography style, 8K detail.
What changed: Added product details (color, material), turntable + orbit camera, studio lighting, crash zoom for detail, audio, style reference.
2. Nature Scene
Before:
a beautiful forest with sunlight
After:
00:00.000–00:06.000 Dense ancient redwood forest, early morning. Thick fog weaves between massive tree trunks. A single beam of golden sunlight breaks through the canopy, illuminating floating dust particles and spider webs. Slow crane up from forest floor to canopy level. Birds call in the distance, leaves rustle softly. Cinematic, anamorphic lens, muted green and gold color palette.
What changed: Specific tree type, time of day, fog + light interaction, camera movement with direction, audio layer, lens style.
3. Character Introduction
Before:
a detective walking down a dark alley
After:
00:00.000–00:04.000 A weathered detective in his 50s — trench coat, loosened tie, cigarette smoke trailing — walks slowly through a narrow alley. Rain drips from fire escapes above. Camera tracks from behind at shoulder height, slight handheld sway. Footsteps echo on wet concrete.
00:04.000–00:07.000 He stops under a single flickering streetlight, turns his head to look at something off-screen. Close-up of his face — exhaustion, suspicion. Hard rim light from the streetlamp. Film noir, high contrast black and white, 1940s mood. A distant siren wails.
What changed: Character appearance + personality, weather, specific camera angle + style, multi-shot with tonal shift, sound design, film noir commitment.
4. Food / ASMR
Before:
pouring honey on pancakes
After:
00:00.000–00:03.000 Extreme close-up of golden honey pouring from a wooden dipper onto a tall stack of fluffy buttermilk pancakes. The honey flows in a thick, slow ribbon — catching warm backlight. Shallow depth of field, only the honey stream is in focus.
00:03.000–00:06.000 Camera pulls back slightly as honey pools between pancake layers and drips down the side. A pat of butter melts slowly on top. Steam rises. Gentle sizzle and honey dripping sounds — no music. Macro food photography style, warm golden tones.
What changed: Specific honey vessel, pancake type, macro lens behavior, physics of honey flow, ASMR audio approach (no music), food photography style.
5. Fantasy / Cinematic
Before:
a dragon flying over mountains
After:
00:00.000–00:05.000 A massive obsidian dragon with iridescent scales soars through a mountain pass at sunset. Wings beat slowly, each downstroke sending snow swirling off the peaks below. Aerial tracking shot from the side, matching the dragon's speed. Wind howls, wings thud with bass-heavy impact.
00:05.000–00:09.000 The dragon banks sharply and dives toward a glacial lake, its reflection visible in the still water. Camera follows the dive from above. As it pulls up at the last second, water erupts in a massive spray. Epic orchestral brass crescendo. Cinematic fantasy, inspired by Lord of the Rings cinematography, golden hour light.
What changed: Dragon details (material, scale), physics (snow displacement, water spray), multi-shot with action arc, aerial camera matching, specific sound design per shot, LOTR as style anchor.
Aspect Ratios & Duration Sweet Spots
16:9 (Landscape) — Best for cinematic content, YouTube, presentations, desktop viewing.
- Gives the most horizontal space for wide shots, landscapes, and multi-subject frames
- Camera movements like pan and tracking shots feel most natural here
- Sweet spot duration: 6-10 seconds
9:16 (Vertical) — Best for TikTok, Instagram Reels, YouTube Shorts, Stories.
- Vertical frame emphasizes single subjects, faces, products
- Close-ups and medium shots work better than wide shots (less horizontal room)
- Camera moves: dolly in/out and tilt up/down work great. Pan left/right has limited space.
- Sweet spot duration: 5-8 seconds (matches platform algorithms)
1:1 (Square) — Best for Instagram feed, Facebook ads, product displays.
- Center-weighted composition — put the subject in the middle
- Works well for product turntables, food shots, portraits
- Sweet spot duration: 5-6 seconds
Duration guidance:
- 4-5s: Quick moments, reactions, product reveals. One shot.
- 6-8s: The sweet spot. Enough for a mini-story. 1-2 shots.
- 9-12s: Full scenes with setup and payoff. 2-3 shots.
- 13-15s: Extended sequences. Needs careful shot planning. 3-5 shots.
Cost note: Longer videos cost more credits/tokens on most platforms. If you're iterating on a concept, start with 5-6 second versions, nail the prompt, then extend.
Try It — Free Prompt Optimizer
Everything in this guide is baked into our free Minimax H3 Prompt Generator. Paste your rough idea, pick your duration and aspect ratio, and get a production-ready prompt in seconds.
The optimizer uses Claude AI to apply all seven dimensions, structure your shots with proper timestamps, add camera direction, and include audio cues — following every rule covered in this guide.
It's free, no credits charged. Sign in with an inReels account and optimize as many prompts as you want.
Already using inReels for video generation? Your optimized H3 prompts work directly in the text-to-video and UGC ads tools.
More resources:
- AI Video Effects Generator — apply effects like Cake-ify, Melt, Explode to any image
- Free Faceless Video Generator — create faceless content for TikTok and YouTube
- Best AI Reel Makers 2026 — comparison of the top AI video tools
Start Creating Video Ads Today
Create UGC-style video ads in minutes. No creators, no waiting, no expensive production. Perfect for TikTok, Instagram Reels, YouTube Shorts, and more.
Try inReels Free →No credit card required