GuidesUpdated August 9, 2026

Seedance 2.5 Prompt Guide — How to Write Prompts That Actually Work (2026)

The complete Seedance 2.5 prompt engineering guide. Learn timestamp format, staged video structure, audio bracket syntax, reference binding, and see real before/after examples for Dreamina.

Why Most Seedance 2.5 Prompts Fail

Seedance 2.5 (the model behind ByteDance's Dreamina) is one of the most capable video generation models available in 2026. It generates videos up to 30 seconds, supports multi-reference materials, handles staged storytelling, and has a bracket-based audio system.

But most people write prompts like this:

a beautiful cinematic video of a chef cooking

And get flat, generic output that ignores half of what they imagined.

Here's the thing about Seedance 2.5 — it's unusually literal. It does what you ask, but only if you say it. Every detail you leave out becomes a decision the model makes without you. The model isn't broken. Your prompt is underspecified.

This guide covers everything that actually matters for Seedance 2.5 prompts: the core formula, timestamp pacing, staged video structure, audio bracket syntax, reference binding, and the specific patterns that consistently produce better results.

Want to skip the reading and just get optimized prompts? Use our free Seedance 2.5 Prompt Generator — paste your idea, get a production-ready prompt.

The Core Formula: 5 Layers That Matter

Every good Seedance 2.5 prompt follows the same structure. Think of it as layers — each one you add gives the model more to work with, and less to guess.

The formula:

Subject + Action/Event + Scene & Environment + Visual Style + Camera + Audio

Here's what each layer carries:

1. Subject + Action — The only non-optional part. Who or what is doing what. Don't say "a person cooking." Say "a ceramic artist finishes a pale blue cup, lifts it from the wheel, and places it in the center of a wooden shelf." Specific subjects with specific actions.

2. Scene & Environment — Where it happens. Location, time of day, weather, spatial relationships, background state. "A studio" is nothing. "A pottery studio at dawn, soft morning light through the window, workbench tidy" gives the model a complete setting.

3. Visual Style — This is where most people go wrong. "Cinematic" is a mood word that leaves the model free. "Soft morning light through a window, visible sensor grain, no colour grading" is an observable instruction. Describe what you actually see — lighting direction, color temperature, material texture, grain.

4. Camera — Shot size, angle, movement, and focus subject. "Begin with a medium shot of the wheel-throwing process, slowly push in toward the cup's surface texture, then cut to a frontal view of the shelf." Use filmmaking terms.

5. Audio — What you hear. Seedance uses bracket syntax: ( ) for music, < > for sound effects, { } for dialogue. More on this in the audio section.

Three levels of specification (same subject):

LevelPrompt
MinimumA ceramic artist lifts a finished cup from the wheel.
MiddleA ceramic artist lifts a finished cup from the wheel in a studio at dawn. Medium shot, slow push-in toward the cup.
FullA ceramic artist finishes a pale blue cup in a studio at dawn, lifts it from the wheel, and places it in the center of a wooden shelf. Soft morning light enters through the window. The wet clay has a delicate sheen. Begin with a medium shot, slowly push in toward the cup's surface texture, then cut to a frontal view of the shelf. Retain the low hum of the pottery wheel and the friction of clay.

You can omit any layer you don't need. But every layer you drop is a decision the model makes for you. The prompt generator handles all five layers automatically.

End States: The Rule That Changes Everything

This is the single most important concept for Seedance 2.5 that most people miss. Every shot, every stage, every timestamp range needs a directly visible end state.

What an end state is:

An end state describes what is physically visible when that segment ends. Character positions, prop locations, scene state.

Good end states (directly visible):

  • "Cup centered on wooden shelf, hand has left frame"
  • "Bouquet lies flat on workbench, ribbon bow facing camera"
  • "Both characters stand behind the workbench inspecting the finished order"
  • "Only the clear glass remains in the center of the table"

Bad end states (moods, not visuals):

  • "Feels peaceful"
  • "The atmosphere shifts to tension"
  • "Mood becomes contemplative"
  • "The scene conveys warmth"

End states are what keep Seedance from drifting. Without them, the model decides where characters end up, where props go, and what the scene looks like. With them, each segment lands exactly where you need it to.

For multi-stage videos, end states are critical because each stage's end state becomes the next stage's starting point. If Stage 1 doesn't specify where the character is standing, Stage 2 can't reliably continue from there.

Timestamps: Integer-Second Pacing

Seedance 2.5 uses integer-second timestamps — simpler than some other models, and effective for controlling pacing.

Format:

0-5 seconds: [Shot description]. End state: [visible state].
5-10 seconds: [Shot description]. End state: [visible state].

Real example (10-second product video):

0-5 seconds: A matte white sneaker with neon green accents sits on a glossy
black turntable, rotating slowly under hard studio spotlight. Camera orbits
at eye level. Clean white background.
End state: sneaker has completed one full rotation.

5-10 seconds: Camera pushes in to extreme close-up on the sole tread pattern,
then rack focus to the knit texture of the upper.
End state: upper texture fills the frame, sharp and detailed.

Three timestamp patterns:

PatternWhen to useExample
Time rangeMost shots0-5 seconds: the shutter is closed.
Exact time pointOne critical momentAt 5 seconds, the camera whip-pans left.
Relative timingDelayed reactionThree seconds after the button press, lights turn off.

Rules that matter:

  • Ranges must be consecutive and non-overlapping. 0-5 then 4-9 breaks. 0-5 then 5-10 works.
  • Ranges are a time budget, not frame-accurate edit points. Actions may occur slightly before or after a boundary.
  • Too little content in a range gives the model freedom to invent motion. "0-8 seconds: she stands there" will not produce 8 seconds of standing.
  • Too much content in a range causes excessive cutting or dropped events. Don't pack five actions into a 3-second range.
  • Never demand a frequency. "Complete three actions in one second" will not work.

When to use timestamps vs stages:

DurationStructure
5sSingle shot, no timestamps needed
10s2-3 timestamp ranges
15-20sTimestamps or stages (your choice)
25-30sAlways use [Stage] structure

Staged Video: Up to 30 Seconds

For videos over 15 seconds, Seedance 2.5's staged structure is the most reliable way to maintain consistency across a longer sequence. Each stage gets one primary event, one end state, and a carry-forward statement.

The structure:

[Generation Goal]
Generate a <type> video. The central subject is <subject>, and the primary
event is <story summary>.

[Stage 1]
Initial state: <where characters/props/scene start>.
Primary event: <one action or event>.
End state: <directly visible positions and states>.

[Stage 2]
Continue from the previous stage: <what must not change>.
Primary event: <one action or event>.
End state: <directly visible positions and states>.

[Maintain Consistency]
Keep <character identity, clothing, prop ownership, spatial direction,
and audio relationships> consistent.

Critical rules for stages:

1. One primary state change per stage. Two events in one stage is where sequences start dropping beats. If your stage has an "and then" in the middle, split it.

2. End states must be directly visible. A position, an object in a hand, a visible scene state. Never a mood.

3. Each stage restates what must not change. "Continue from the previous stage: both characters retain the same identities and clothing, and the florist still holds the bouquet."

4. Close with the consistency block. Character count, clothing, prop ownership, spatial direction, and audio relationships all drift without it.

Real example — 30-second flower shop video:

[Generation Goal]
Generate an instructional video showing a flower shop's order-packing
process. <Florist> and <Store Assistant> arrange, wrap, and hand off
a bouquet together.

[Stage 1]
Initial state: <Florist> stands behind the workbench. Loose flower stems,
scissors, and wrapping paper lie on the tabletop.
Primary event: <Florist> arranges the stems and trims them to length.
End state: <Florist> holds the bouquet in the left hand, and the scissors
are back on the right side of the workbench.

[Stage 2]
Continue from the previous stage: both characters retain the same
identities and clothing, and <Florist> still holds the bouquet.
Primary event: <Store Assistant> unfolds the wrapping paper. <Florist>
places the bouquet inside and ties it with a green ribbon.
End state: the wrapped bouquet lies flat in the center of the workbench,
with the ribbon bow facing the camera.

[Stage 3]
Primary event: <Store Assistant> picks up the bouquet and places it on
the pickup shelf.
End state: the bouquet is centered on the pickup shelf, and both
characters stand behind the workbench inspecting the finished order.

[Maintain Consistency]
Keep <Florist> and <Store Assistant>'s identities, clothing, workbench
orientation, scissors position, and bouquet ownership consistent.

Audio Bracket Syntax

Seedance 2.5 has a structured audio system using different bracket types for different audio categories. This is one of its strongest features — but only if you use the syntax correctly.

The four bracket types:

BracketCategoryExample
( )Music(Soft, rhythmic piano music plays in the background)
< >Sound effects<A bell rings in the distance>
{ }Dialogue{Hello, welcome back.}
【 】Subtitles【Chapter One: Departure】

Music ( )

Describe the score — mood, tempo, instrumentation.

(Soft, rhythmic piano music plays in the background)
(Upbeat electronic beat with bass pulse)
(No music)

Use (No music) or (No BGM) to explicitly suppress background music.

Sound effects < >

Discrete, source-linked events. One effect per bracket.

<The metal shutter rattles upward>
<Steam hisses from the wand>
<Footsteps reverberate on stone>

Don't stack five effects into one bracket — give each its own.

Dialogue { }

The spoken words go inside the braces. Everything else — delivery style, speaker, language — goes outside, immediately before the braces.

The barista says: {First one's always the best one.}

The full dialogue formula:

Dialogue language + regional variety + delivery style + speaker + {text}

Example:

Dialogue language: American English. The barista says in natural,
conversational American English: {First one's always the best one.}

Name the language before the line when dialogue is not in Chinese. If the model speaks English text in Chinese, reinforce with the full formula including regional variety.

All four brackets in one prompt:

A barista opens the shutter of a small corner café at dawn, wipes down the
counter, and starts the first pour of the morning.

(Soft, rhythmic piano music plays in the background)
<The metal shutter rattles upward>
<Steam hisses from the wand>
Dialogue language: American English. The barista says in natural,
conversational American English: {First one's always the best one.}

Reference Binding: @Image, @Video, @Audio

When you upload images, videos, or audio as reference materials, you must tell Seedance exactly what each one contributes. This is where most multi-reference prompts fail.

The single most common failure:

@Images 1 through 4 define four characters respectively.

This never works. It doesn't say which image is which character.

The right way:

@Image 1 defines the ceramic artist's facial features, hairstyle, and dark
green apron. Do not use the image background.

@Image 2 defines the wooden workbench, window placement, and morning light
of the pottery studio. Do not use the people in the image.

@Video 1 defines the pacing of throwing clay with both hands, lifting the
cup, and placing it down. Do not use the person's identity, clothing, or
scene from the video.

Key rules for reference binding:

1. Bind each reference individually. State exactly what attributes to use.

2. Always write exclusions. If a reference image contains a background you don't want carried over, say "Do not use the image background." If it has people you don't want, say "Do not use the people in the image."

3. Write mappings in the prompt, not in the image. Text labels inside a reference image are not read as instructions.

4. State the count for multi-view objects. If four images show different angles of the same lamp:

@Image 1 defines the front view of the same folding desk lamp.
@Image 2 defines the left-side structure of the same folding desk lamp.
@Image 3 defines the right-side structure of the same folding desk lamp.
@Image 4 defines the rear structure of the same folding desk lamp.

All four images define one folding desk lamp. The output must contain
only one lamp throughout.

Without the count statement, you get four separate lamps instead of four views of one lamp.

5. Don't re-describe motion a reference video already carries. If your reference video already shows the exact motion you want, don't restate it — the description may conflict with the reference.

Reference limits:

TypeLimitBest range
ImagesUp to 30 (each ≤4K)1-8 distinct subjects
VideosUp to 10 (30s combined)1-5 subjects, 5-10s each
AudioUp to 10 (30s combined)Only relevant clips
Total50 materials max

Camera Moves and Framing

Seedance 2.5 understands standard cinematography terms. Basic shot sizes and camera movements can be written directly. For niche or ambiguous terms, also describe the visible result.

Basic terms (write directly):

CategoryTerms
Shot sizeextreme wide, wide, medium, close-up, extreme close-up
Camera movementpush in, pull out, pan, lateral move, follow shot, orbit, tilt up, handheld shake
Camera positionlow angle, overhead view, first-person view

Popular techniques (write directly, but add detail):

  • dolly zoom — State which subject to preserve and whether the background moves closer or farther.
  • one-take shot — List the subjects, spaces, and events the camera passes through, in order.
  • bullet time — State the action to freeze or slow, and the camera's orbit direction.
  • FPV — State the flight/traversal path, speed, and turns.
  • bounce speed ramp — State where the action accelerates, decelerates, or rebounds, and its final resting state.

When a term is niche or ambiguous:

Use the formula: cinematography term + target subject + visual change + direction/speed

Rack focus: shift focus smoothly from the leaves in the foreground to the
person in the background. The leaves gradually blur while the person's face
changes from soft to sharp.
Tracking shot: move horizontally at the same speed as <Skateboarder>,
keeping the subject sharp while the roadside wall forms horizontal motion
blur from right to left.
Natural vignette: darken the four corners gradually while keeping the
brightness and skin tone of <Pianist> in the center natural, without a
black border.

If the frame has multiple subjects, state which one the camera follows or revolves around, where movement begins, and where it ends.

Visual Style: Observable, Not Adjectival

This is where Seedance prompts diverge from most other AI models. "Cinematic" is not a style instruction — it's a mood word that the model interprets freely.

The principle: describe what you actually see on screen.

Mood word (weak)Observable instruction (strong)
CinematicSoft morning light through a window, visible sensor grain, no colour grading
VibrantSaturated primary colors, hard studio lighting, clean white background
MoodyLow-key lighting, single practical light source from the left, deep shadows on the right
WarmGolden hour backlight, long shadows across a mountain ridge, amber-toned skin
ProfessionalThree-point studio lighting, shallow depth of field, 85mm portrait lens bokeh

Lighting keywords that work:

golden hour, warm low-angle sunlight, soft morning light, hard studio spotlight, top-down spotlight, soft diffused light, backlight silhouette, rim light, neon glow, candlelight, moonlight, volumetric light rays, Rembrandt lighting, high key, low key, chiaroscuro

Style keywords that work:

photorealistic, hyperrealistic, 35mm film, anamorphic, documentary, raw handheld footage, visible sensor grain, anime, watercolor, claymation, cyberpunk, film noir, product photography

Emotional direction — observable cues, not adjectives:

When directing emotional performance, describe visible and audible cues:

After confirming the curtain call is over, the actor exhales softly.
The shoulders gradually relax, a restrained smile appears, and the eyes
slowly well with tears, but the actor never turns to leave.

Not: "The actor feels relieved and emotional."

Two to four clear cues per emotional transition is usually enough. Don't catalogue every facial muscle.

What to Avoid

Every model has its failure modes. These are the patterns that consistently produce bad output in Seedance 2.5:

1. Don't put generation parameters in the prompt. Aspect ratio and duration are set on the generation page or through the API. Writing "16:9 aspect ratio" or "10 seconds" in your prompt text does nothing — or worse, it confuses the model.

2. Don't use mood words as style. "Cinematic," "vibrant," "moody," "professional" — these all leave room for interpretation. Describe the lighting, color, material, and texture you actually want to see.

3. Don't batch-reference. @Images 1 through 4 define four characters never maps correctly. Bind each reference individually with explicit attribute statements and exclusions.

4. Don't skip exclusions. If a reference image contains a background or person you don't want in the output, you must say so. "Do not use the image background." The model doesn't infer your intent.

5. Don't put multiple primary events in one stage. One state change per stage. Two events = dropped beats. If there's an "and then" in your stage, split it.

6. Don't write moods as end states. "Feels peaceful" is not an end state. "Cup centered on shelf, hand has left frame" is.

7. Don't overlap timestamp ranges. 0-5 then 4-9 breaks. Use 0-5 then 5-10. Ranges must be consecutive and non-overlapping.

8. Don't demand frequencies. "Complete three actions in one second" will never work. Give each action enough time budget.

9. Don't re-describe reference motion. If your reference video already carries the exact motion, don't rewrite it in the prompt. The description conflicts with what the model sees in the reference.

Before & After Examples

Here's what each optimization adds. Every "after" prompt follows the core formula with observable style, camera direction, audio brackets, and end states.


Example 1: Simple scene

Before:

a potter making a cup

After:

A ceramic artist finishes a pale blue cup in a studio at dawn, lifts it
from the wheel, and places it in the center of a wooden shelf.

Soft morning light enters through the window. The wet clay has a delicate
sheen, and the workbench remains tidy.

Begin with a medium shot of the wheel-throwing process, slowly push in
toward the cup's surface texture, then cut to a frontal view of the shelf.

Retain the low hum of the pottery wheel, the friction of clay, and subtle
indoor ambience.

What changed: specific action sequence, observable lighting, material texture, camera moves with shot sizes, natural audio description.


Example 2: Scene with dialogue

Before:

café opening in the morning

After:

A barista opens the shutter of a small corner café at dawn, wipes down
the counter, and starts the first pour of the morning.

Warm low sunlight through the front glass, steam rising, visible sensor
grain. Medium shot, slow push-in toward the cup, then cut to the street.

(Soft, rhythmic piano music plays in the background)
<The metal shutter rattles upward>
<Steam hisses from the wand>
Dialogue language: American English. The barista says in natural,
conversational American English: {First one's always the best one.}

What changed: specific actions, observable visual style (not "cinematic"), all four audio bracket types demonstrated, dialogue with language reinforcement.


Example 3: Product video with timestamps

Before:

show a sneaker product video

After:

0-5 seconds: A matte white sneaker with neon green accents sits on a glossy
black turntable, rotating slowly under hard studio spotlight. Camera orbits
at eye level. Clean white background.
End state: sneaker has completed one full rotation.

5-10 seconds: Camera pushes in to extreme close-up on the sole tread
pattern, then rack focus to the knit texture of the upper.
End state: upper texture fills the frame, sharp and detailed.

(Subtle bass-heavy electronic pulse)
<Soft turntable motor hum>
Hyperrealistic product photography style, 8K detail.

What changed: timestamp ranges with end states, product details, camera orbit + push-in + rack focus, studio lighting, audio brackets, style as texture description.


Example 4: 30-second staged video

Before:

a 30 second flower shop video

After: (uses [Stage] structure — see the staged video section for the full prompt)

What changed: [Generation Goal] + 3 stages with single events and visible end states + [Maintain Consistency] block for character/prop/spatial continuity.

Generation Modes and Parameters

Seedance 2.5 supports several generation modes. Some parameters lock automatically depending on the mode.

Text-to-Video (T2V)

The default mode. Write a prompt, set aspect ratio and duration, generate.

  • Aspect ratio: settable
  • Duration: settable (5-30s)

Image-to-Video (I2V)

Upload one or more images as references. The prompt describes how the scene evolves from the reference materials.

  • Use @Image 1 defines... to specify what each image contributes
  • State exclusions for what NOT to use from each image
  • Don't re-describe the image — describe the evolution
  • Aspect ratio: settable (or inherits from first-frame image)
  • Duration: settable

Video Editing

Upload a source video and modify it. The source is the "sole editing master."

  • Aspect ratio: locked to source video
  • Duration: locked to source (±0.3s drift from transition-frame handling)
  • The prompt must specify: edit goal, source video role, edit scope, target material, and content to preserve

Video Extension (Forward/Backward)

Extend an existing video by adding new content before or after it.

  • Aspect ratio: locked to source video
  • Duration: settable (for the new segment only)
  • Must describe boundary frame matching for continuity

Locked vs settable parameters:

ModeAspect RatioDuration
Text-to-videoSettableSettable
Image-to-videoSettable (or inherits first frame)Settable
Video editingLocked to sourceLocked to source ±0.3s
Video extensionLocked to sourceSettable (new segment)
First/last frameLocked to first imageSettable

The practical consequence: decide your aspect ratio on the first generation. Every downstream edit and extension inherits it and cannot change it. If you need a different ratio, regenerate or reframe in post.

Try It: Free Prompt Generator

Everything in this guide is built into our free Seedance 2.5 Prompt Generator. Paste a rough idea, pick your duration and aspect ratio, and get a production-ready Seedance 2.5 prompt with:

  • Observable visual style (not mood words)
  • Integer-second timestamps with end states
  • Staged structure for 20-30s videos
  • Audio bracket syntax — ( ) music, < > SFX, { } dialogue
  • Camera direction with filmmaking vocabulary

No credits charged. No paywall. Sign in with a free inReels account and optimize as many prompts as you want.

If you're generating videos with inReels, the same prompt engineering principles apply across all our tools — from UGC ad generation to faceless videos.

Also check out our Minimax H3 Prompt Guide for the other major video model, or use the H3 Prompt Generator for Hailuo AI prompts.

Start Creating Video Ads Today

Create UGC-style video ads in minutes. No creators, no waiting, no expensive production. Perfect for TikTok, Instagram Reels, YouTube Shorts, and more.

Try inReels Free →

No credit card required

Questions? Chat with us

We typically reply instantly